Executed notebooks evaluating the phase pickers and the event classifier ahead of the 2026 re-run. Every figure and table here was produced by the notebook linked beside it.
Our workers overloaded EarthScope's credentials endpoint: 351,735 rejected token requests over four hours, peaking at 5,104/min, and 3,925 requests for temporary networks that omitted the year and could only be refused. Five defects in one file, each linked to the line that caused it and the line that replaced it, with a self-test that runs on any EarthScope account.
What breaks when a few hundred workers read an archive whose credentials are issued per network and per year. Eight defects and what each one generalises to, how to find them before the archive operator does, and what the archive side can do to make a client's job possible.
The picks are public-read Parquet on S3: no account, no credentials, no
database. Mirror one month of the finished western campaign,
plot what is in it, fetch the original waveforms from FDSN with ObsPy and
draw the picks on the record, then the three rules that keep 1.3 billion
picks readable.
Choose stations by place and time from the station table, mirror every partition that covers them, query the mirror with DuckDB, check which station-days were actually processed, and export for PyOcto or GaMMA. Worked on the Monte Cristo Range M6.5, May–June 2020.
The full metric set for all five benchmarks — detection at three threshold treatments, onset-time statistics, confidence calibration, phase swaps, duplicate picks — with what each one means and which are identifiable when the reference is an operator bulletin rather than a labelled test set. Written to be read by people who evaluate models for a living and by people who pick phases for a living. Section 8 is the script that computes all of it on your own two CSV files.
Every benchmark below in one place, drawn from the result tables the notebooks export: recall at a shared threshold and at matched pick budgets across seven sequences on two continents, timing, the ocean-bottom studies, and whether the stored catalogue reproduces. The figures for the paper.
Do the three weight sets hold up away from the sequence they were first checked on? Ridgecrest, San Simeon, Monte Cristo, Mendocino 2024 and Monroe WA — four regions, three archives, catalogs from 2003 to 2024.
Kaikōura 2016, Norcia 2016 and Thessaly 2021 against GeoNet, INGV and NOA analyst picks: does the fine-tune travel, and which weight should the global campaign run?
Re-picks the campaign's own station-days through the production code path, but reading waveforms over ObsPy/FDSN instead of the S3 buckets, and compares pick for pick. 59,298 of 59,315 match exactly — and the station-days with no picks at all turn out to be the interesting ones.
Twenty-one aftershocks in thirty minutes, several separated by seconds, scored against 533 analyst S picks. Includes a per-station breakdown and a matching-tolerance sweep.
Does one weight set produce physically sensible picks at all? Record section, per-station waveforms, and an S−P timing check against hypocentral distance. The first thing to run after changing weights.
What the three SeisBench OBS models buy over land models on ocean-bottom
data — Cascadia, the Alaska Peninsula, and the Blanco transform — whether the
hydrophone channel earns its place, and how the obs campaign's
stored picks compare with analyst-checked arrivals on the AACSE array and with
the real-time catalogue at Axial Seamount.
Deploying on ocean-bottom data? The offset diagnosis and station-selection gotchas are written up in obs_deployment_notes.md.
Does the PNW-trained QuakeXNet separate earthquakes, explosions and surface events in Alaska — and how much does the analysis window placement matter?
The picker. At the shared 0.3 threshold everyone uses,
original appears to recover far more S arrivals than the others.
That reading is an artefact of the threshold, not a property of the weights.
| S recall at a shared 0.3 | quakescope2026 | jma_wc | original | instance | analyst S |
|---|---|---|---|---|---|
| Ridgecrest | 0.52 | 0.49 | 0.73 | 0.18 | 298 |
| Mendocino 2024 | 0.62 | 0.68 | 0.74 | 0.56 | 113 |
| San Simeon | 0.75 | 0.75 | 0.94 | 0.75 | 16 |
| Monte Cristo | 0.42 | 0.50 | 0.67 | 0.50 | 12 |
Holding the threshold fixed does not hold the operating point fixed.
original emits close to twice as many S picks at 0.3, so it sits
further along the recall curve and collects both more recall and more extra
detections. Matched on pick budget, the three are within a few points and no
ordering survives across sequences:
| Ridgecrest · S picks emitted | quakescope2026 | jma_wc | original |
|---|---|---|---|
| 287 | 0.548 | 0.536 | 0.540 |
| 370 | 0.663 | 0.639 | 0.628 |
| 535 | 0.775 | 0.777 | 0.768 |
instance — what the 2025 campaign actually ran — is a
separate case. On Mendocino it matches the others on budget and is in
places the best of the four. On Ridgecrest it never reaches their budgets at
all: with its threshold on the floor it emits 246 S picks where the others reach
684 and 832. No threshold recovers that, so it is a ceiling rather than
a calibration offset — and the two have opposite implications, since tuning
fixes one and cannot fix the other.
| Ridgecrest · S picks emitted | at 0.02 (floor) | at 0.1 | at 0.3 |
|---|---|---|---|
| original | 832 | 646 | 520 |
| quakescope2026 | 684 | 459 | 291 |
| jma_wc | 638 | 447 | 279 |
| instance | 246 | 149 | 98 |
Ridgecrest is the densest and closest-in sequence here, so this reads as the incumbent weight struggling specifically with heavily overlapping near-field aftershocks, while remaining fine at regional distance.
The three are also not one lineage. original is Zhu et al.,
trained on Northern California. jma_wc is a different architecture
— PhaseNetWC, double the filters per layer — trained on Japanese JMA data, and
quakescope2026 is fine-tuned from it. Nothing here descends from
original.
Offshore. SeisBench ships three ocean-bottom pickers.
PickBlue is a constructor returning the obs weights on
either a PhaseNet or an EQTransformer backbone — both four-component, including a
hydrophone — while OBSTransformer is OBS-trained but takes only three.
Across 138 windows on three deployments the OBS models lead the land models by
roughly 5–15 points of detection, and no single one wins everywhere.
| Detection rate | pickblue phasenet | pickblue eqt | obs transformer | quakescope 2026 | original |
|---|---|---|---|---|---|
| Cascadia (7D) | 0.74 | 0.79 | 0.61 | 0.72 | 0.65 |
| AACSE (XO) | 0.78 | 0.72 | 0.89 | 0.78 | 0.67 |
| Blanco (X9) | 0.53 | 0.53 | 0.51 | 0.53 | 0.49 |
obstransformer
competes without a hydrophone at all points the same way: the advantage comes
from training on ocean-bottom data, not from the fourth channel.
Offshore, against published picks. The obs
campaign's stored picks, scored against Barcheck's analyst-checked AACSE arrivals
(65 OBS stations, 2018) and against the University of Washington's real-time picks
at Axial Seamount (7 stations, 2015–2025). Recall is on station-days the campaign
holds picks for; the reference's silence is never counted as a false positive.
Re-scored on the same AACSE windows against analyst picks at 1 s, the five models
keep the ordering of the table above with wider gaps: obstransformer
0.81 P / 0.87 S, the PickBlue pair 0.74–0.77 P, original 0.49 P.
| Recall of the reference's picks | P | S | tolerance | reference |
|---|---|---|---|---|
| AACSE 2018, 65 OBS stations | 0.86 | 0.86 | 1 s | analyst, manual picks |
| AACSE 2018, 30 land stations, same weight | 0.91 | 0.85 | 1 s | analyst, manual picks |
| Axial 2015–2025, short-period EH | 0.21 | 0.04 | 0.5 s | UW automatic picks |
| Axial 2015–2025, broadband HH | 0.30 | 0.01 | 0.5 s | UW automatic picks |
obs weights were trained on;
recall reaches 0.5 for M 0.5–1.5 and the campaign catalogue should not be read
below M 0.
Outside the United States. Kaikōura 2016 (M7.8), Norcia 2016
(M6.5) and Thessaly 2021 (M6.3), scored against the operators' own analyst picks.
jma_wc beats quakescope2026 on every sequence and phase
at the shared threshold and at matched pick budgets, with timing identical to
within 6 ms, so the fine-tune is not the more general weight. instance
is the most efficient at any common budget on all three, not only in Italy.
| Recall at a matched pick budget | quakescope2026 | jma_wc | instance |
|---|---|---|---|
| Kaikōura P (1,367 picks) | 0.61 | 0.62 | 0.72 |
| Norcia P (982) | 0.69 | 0.71 | 0.79 |
| Thessaly P (1,439) | 0.68 | 0.71 | 0.76 |
| Kaikōura S (1,034) | 0.52 | 0.52 | 0.70 |
| Norcia S (975) | 0.67 | 0.69 | 0.86 |
| Thessaly S (1,181) | 0.41 | 0.41 | 0.46 |
jma_wc. The
fine-tune loses 1–2 points of recall at every budget and gains nothing on
timing. instance emits half the picks at 0.3 and tops out below the
others' ceilings, so it is the better choice only where the target is the
operator's catalogue rather than every arrival, and only at a threshold set for
it. Thessaly's S recall is capped near 0.5 for every weight, a property of the
reference rather than of any model.
The classifier. QuakeXNet agrees with the Alaska catalog 78% of the time when the analysis window matches the training convention, and 16% when it does not — same events, same waveforms, cut differently. It is deferred for the 2026 campaign on that basis.
docs/rerun_2026/; arrivals for those events
are then harvested from every archive that located them, restricted to manually
reviewed picks. The report's provenance section covers what the catalogs do and
do not record. Sample sizes still differ by more than an order of magnitude
between sequences — San Simeon rests on 16 S picks against Ridgecrest's 298 —
so read the analyst column alongside every recall.
Notebooks live in tutorials/. To rebuild a report:
pixi install --environment tutorials |
pixi run -e tutorials install-kernel |
pixi run -e tutorials jupyter nbconvert --to notebook --execute … |
Full commands are in reports/README.md. Reports are committed pre-rendered, so publishing them needs no data access, no model weights, and no compute in CI.