QuakeScope model reports

Executed notebooks evaluating the phase pickers and the event classifier ahead of the 2026 re-run. Every figure and table here was produced by the notebook linked beside it.

Operations

EarthScope credential audit — 2026-09-04 incident

Our workers overloaded EarthScope's credentials endpoint: 351,735 rejected token requests over four hours, peaking at 5,104/min, and 3,925 requests for temporary networks that omitted the year and could only be refused. Five defects in one file, each linked to the line that caused it and the line that replaced it, with a self-test that runs on any EarthScope account.

Fixed and merged · validated by two bounded dry runs · campaigns still stopped

Reading an archive behind per-object credentials

What breaks when a few hundred workers read an archive whose credentials are issued per network and per year. Eight defects and what each one generalises to, how to find them before the archive operator does, and what the archive side can do to make a client's job possible.

Engineering note · drawn from the incident above

Reports

Campaign dashboard — live

Picks per station and per day, progress by campaign, compute and spend. Rebuilt hourly from S3 and Batch.

Reading the pick catalogue — start here

The picks are public-read Parquet on S3: no account, no credentials, no database. Mirror one month of the finished western campaign, plot what is in it, fetch the original waveforms from FDSN with ObsPy and draw the picks on the record, then the three rules that keep 1.3 billion picks readable.

Runs anonymously · open in Colab · ~5 min

Downloading a region and a period

Choose stations by place and time from the station table, mirror every partition that covers them, query the mirror with DuckDB, check which station-days were actually processed, and export for PyOcto or GaMMA. Worked on the Monte Cristo Range M6.5, May–June 2020.

Runs anonymously · ~10 min · reference: docs/data_access.md

2026 launch briefing

Why we are re-running, what the benchmarks and waveforms show, and what it costs

Scoring a picker: every metric, on every benchmark

The full metric set for all five benchmarks — detection at three threshold treatments, onset-time statistics, confidence calibration, phase swaps, duplicate picks — with what each one means and which are identifiable when the reference is an operator bulletin rather than a labelled test set. Written to be read by people who evaluate models for a living and by people who pick phases for a living. Section 8 is the script that computes all of it on your own two CSV files.

Computed from the studies' exported picks · definitions in benchmark_metrics.py · run it yourself with score_picks.py · ~2 min

The picker benchmarks, consolidated

Every benchmark below in one place, drawn from the result tables the notebooks export: recall at a shared threshold and at matched pick budgets across seven sequences on two continents, timing, the ocean-bottom studies, and whether the stored catalogue reproduces. The figures for the paper.

Reads docs/benchmark/results/ · no inference · ~1 min · protocol and caveats in docs/benchmark/README.md

Pickers across five earthquake sequences

Do the three weight sets hold up away from the sequence they were first checked on? Ridgecrest, San Simeon, Monte Cristo, Mendocino 2024 and Monroe WA — four regions, three archives, catalogs from 2003 to 2024.

5 stations per sequence · scored against analyst picks · ~25 min to run

Pickers outside the United States

Kaikōura 2016, Norcia 2016 and Thessaly 2021 against GeoNet, INGV and NOA analyst picks: does the fine-tune travel, and which weight should the global campaign run?

3 sequences · 18 stations · 4 weight sets · ~15 min to run

Did the western campaign pick what it says it picked?

Re-picks the campaign's own station-days through the production code path, but reading waveforms over ObsPy/FDSN instead of the S3 buckets, and compares pick for pick. 59,298 of 59,315 match exactly — and the station-days with no picks at all turn out to be the interesting ones.

52 station-days · mainshock and quiet · independent data path · ~90 min to run

S-recall benchmark, dense Ridgecrest aftershocks

Twenty-one aftershocks in thirty minutes, several separated by seconds, scored against 533 analyst S picks. Includes a per-station breakdown and a matching-tolerance sweep.

8 stations · 4 weight sets · ~5 min to run

Picker smoke test

Does one weight set produce physically sensible picks at all? Record section, per-station waveforms, and an S−P timing check against hypocentral distance. The first thing to run after changing weights.

5 stations · 2 events · ~2 min to run

Ocean-bottom pickers offshore

What the three SeisBench OBS models buy over land models on ocean-bottom data — Cascadia, the Alaska Peninsula, and the Blanco transform — whether the hydrophone channel earns its place, and how the obs campaign's stored picks compare with analyst-checked arrivals on the AACSE array and with the real-time catalogue at Axial Seamount.

3 deployments · 5 models · 65 AACSE stations against 30,618 analyst arrivals · ~10 min to run

Deploying on ocean-bottom data? The offset diagnosis and station-selection gotchas are written up in obs_deployment_notes.md.

Classifier transfer to Alaska

Does the PNW-trained QuakeXNet separate earthquakes, explosions and surface events in Alaska — and how much does the analysis window placement matter?

32 events · 2 stations · ~30 min to run

What they currently show

The picker. At the shared 0.3 threshold everyone uses, original appears to recover far more S arrivals than the others. That reading is an artefact of the threshold, not a property of the weights.

S recall at a shared 0.3quakescope2026jma_wcoriginalinstanceanalyst S
Ridgecrest0.520.490.730.18298
Mendocino 20240.620.680.740.56113
San Simeon0.750.750.940.7516
Monte Cristo0.420.500.670.5012

Holding the threshold fixed does not hold the operating point fixed. original emits close to twice as many S picks at 0.3, so it sits further along the recall curve and collects both more recall and more extra detections. Matched on pick budget, the three are within a few points and no ordering survives across sequences:

Ridgecrest · S picks emittedquakescope2026jma_wcoriginal
2870.5480.5360.540
3700.6630.6390.628
5350.7750.7770.768
Thresholds belong to the weight set, not to the pipeline. Carrying 0.3 across a change of weights silently moves the operating point and changes catalog completeness with it. Choose the target — a pick budget, an extra-detection rate, a recall floor — then solve for the threshold that hits it, separately for each weight set.

instance — what the 2025 campaign actually ran — is a separate case. On Mendocino it matches the others on budget and is in places the best of the four. On Ridgecrest it never reaches their budgets at all: with its threshold on the floor it emits 246 S picks where the others reach 684 and 832. No threshold recovers that, so it is a ceiling rather than a calibration offset — and the two have opposite implications, since tuning fixes one and cannot fix the other.

Ridgecrest · S picks emittedat 0.02 (floor)at 0.1at 0.3
original832646520
quakescope2026684459291
jma_wc638447279
instance24614998

Ridgecrest is the densest and closest-in sequence here, so this reads as the incumbent weight struggling specifically with heavily overlapping near-field aftershocks, while remaining fine at regional distance.

The three are also not one lineage. original is Zhu et al., trained on Northern California. jma_wc is a different architecture — PhaseNetWC, double the filters per layer — trained on Japanese JMA data, and quakescope2026 is fine-tuned from it. Nothing here descends from original.

Offshore. SeisBench ships three ocean-bottom pickers. PickBlue is a constructor returning the obs weights on either a PhaseNet or an EQTransformer backbone — both four-component, including a hydrophone — while OBSTransformer is OBS-trained but takes only three. Across 138 windows on three deployments the OBS models lead the land models by roughly 5–15 points of detection, and no single one wins everywhere.

Detection ratepickblue
phasenet
pickblue
eqt
obs
transformer
quakescope
2026
original
Cascadia (7D)0.740.790.610.720.65
AACSE (XO)0.780.720.890.780.67
Blanco (X9)0.530.530.510.530.49
The hydrophone is not where the gain comes from. Re-running the four-component model on 94 detected windows with the channel withheld moved mean P confidence by +0.0002 — helping 32 windows, hurting 35. Its sampling rate spans 100 Hz to 10 Hz across these deployments and the effect is indistinguishable from noise at any of them. That obstransformer competes without a hydrophone at all points the same way: the advantage comes from training on ocean-bottom data, not from the fourth channel.

Offshore, against published picks. The obs campaign's stored picks, scored against Barcheck's analyst-checked AACSE arrivals (65 OBS stations, 2018) and against the University of Washington's real-time picks at Axial Seamount (7 stations, 2015–2025). Recall is on station-days the campaign holds picks for; the reference's silence is never counted as a false positive. Re-scored on the same AACSE windows against analyst picks at 1 s, the five models keep the ordering of the table above with wider gaps: obstransformer 0.81 P / 0.87 S, the PickBlue pair 0.74–0.77 P, original 0.49 P.

Recall of the reference's picksPStolerancereference
AACSE 2018, 65 OBS stations0.860.861 sanalyst, manual picks
AACSE 2018, 30 land stations, same weight0.910.851 sanalyst, manual picks
Axial 2015–2025, short-period EH0.210.040.5 sUW automatic picks
Axial 2015–2025, broadband HH0.300.010.5 sUW automatic picks
Where AACSE disagrees, the archive is the reason. Median residuals are +0.02 s (P) and +0.04 s (S). Two OBS stations score zero because the nearest campaign pick sits hours from the analyst's, drifting month to month — a clock correction the archive does not carry — and 5% of the analyst arrivals fall on station-days the campaign never read, in runs of consecutive days inside shards that reported themselves complete. Axial is the opposite case: the archive is fine, but the catalogue there is mostly sub-M0 events with S half a second behind P, outside what the obs weights were trained on; recall reaches 0.5 for M 0.5–1.5 and the campaign catalogue should not be read below M 0.

Outside the United States. Kaikōura 2016 (M7.8), Norcia 2016 (M6.5) and Thessaly 2021 (M6.3), scored against the operators' own analyst picks. jma_wc beats quakescope2026 on every sequence and phase at the shared threshold and at matched pick budgets, with timing identical to within 6 ms, so the fine-tune is not the more general weight. instance is the most efficient at any common budget on all three, not only in Italy.

Recall at a matched pick budgetquakescope2026jma_wcinstance
Kaikōura P (1,367 picks)0.610.620.72
Norcia P (982)0.690.710.79
Thessaly P (1,439)0.680.710.76
Kaikōura S (1,034)0.520.520.70
Norcia S (975)0.670.690.86
Thessaly S (1,181)0.410.410.46
For the global deployment: stay on jma_wc. The fine-tune loses 1–2 points of recall at every budget and gains nothing on timing. instance emits half the picks at 0.3 and tops out below the others' ceilings, so it is the better choice only where the target is the operator's catalogue rather than every arrival, and only at a threshold set for it. Thessaly's S recall is capped near 0.5 for every weight, a property of the reference rather than of any model.

The classifier. QuakeXNet agrees with the Alaska catalog 78% of the time when the analysis window matches the training convention, and 16% when it does not — same events, same waveforms, cut differently. It is deferred for the 2026 campaign on that basis.

Recall is measured against analyst picks, which are not exhaustive. Unmatched model picks are reported as extra detections, never as false positives — in a dense sequence most of them are real earthquakes nobody had time to work through. Which earthquakes count is fixed by the team's curated reviewed-event lists in docs/rerun_2026/; arrivals for those events are then harvested from every archive that located them, restricted to manually reviewed picks. The report's provenance section covers what the catalogs do and do not record. Sample sizes still differ by more than an order of magnitude between sequences — San Simeon rests on 16 S picks against Ridgecrest's 298 — so read the analyst column alongside every recall.

Reproducing

Notebooks live in tutorials/. To rebuild a report:

pixi install --environment tutorials
pixi run -e tutorials install-kernel
pixi run -e tutorials jupyter nbconvert --to notebook --execute …

Full commands are in reports/README.md. Reports are committed pre-rendered, so publishing them needs no data access, no model weights, and no compute in CI.