Downloading the QuakeScope pick catalogue for a region and a period¶

The first notebook reads one network-month. This one does what a study needs: choose stations by place and time, mirror every partition that covers them, query the mirror, check which station-days were actually processed, and export for an associator. Anonymous throughout; nothing here needs an AWS account.

Worked example: the Monte Cristo Range, Nevada, M6.5 of 2020-05-15, with a 1.5 degree box and the two months from May 1 to June 30. Change the four values in section 2 and everything below follows.

bucket s3://quakescope-picks-2026, us-east-2, public-read
catalogue western: 1986 to 2026, PhaseNet original, P and S thresholds 0.2 (three campaign eras and a repair pass, all under one prefix)
layout <campaign>/picks/network=<NET>/year=<YYYY>/month=<MM>/<shard_id>.parquet, plus manifests/, runs/, stations.parquet
reference docs/data_access.md

Runtime: about ten minutes on a laptop, most of it the download; re-runs skip files already mirrored.

1. Setup¶

In [1]:
import datetime, json, time
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path

import duckdb
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import s3fs

BUCKET = "quakescope-picks-2026"
fs = s3fs.S3FileSystem(anon=True)
MIRROR = Path("quakescope_mirror")          # local tree, same layout as the bucket
MIRROR.mkdir(exist_ok=True)

CATALOGUE = "western"                       # one prefix, years 1986 to 2026

C_P, C_S = "#2a78d6", "#eb6834"
plt.rcParams.update({"figure.dpi": 110, "axes.spines.top": False, "axes.spines.right": False})

2. Choose a region and a period¶

Four values. The box is in degrees; the period is inclusive of both dates.

In [2]:
CENTER_LAT, CENTER_LON, HALF_DEG = 38.169, -117.850, 1.5          # Monte Cristo Range M6.5, 2020-05-15
START, END = datetime.date(2020, 5, 1), datetime.date(2020, 6, 30)

LAT0, LAT1 = CENTER_LAT - HALF_DEG, CENTER_LAT + HALF_DEG
LON0, LON1 = CENTER_LON - HALF_DEG, CENTER_LON + HALF_DEG
print(f"box {LAT0:.2f} to {LAT1:.2f} N, {LON0:.2f} to {LON1:.2f} E, {START} to {END}")
box 36.67 to 39.67 N, -119.35 to -116.35 E, 2020-05-01 to 2020-06-30

Stations in the box, operating in the period¶

stations.parquet is the table the catalogue was planned from. start_date and end_date are Parquet dates, so they compare directly with a datetime.date; a station still operating carries 3000-01-01 rather than a null, so a filter like end_date >= START keeps it instead of dropping it.

Until 2026-09-29 these two columns were a YYYY.DDD float, which is kept beside them as start_yearday and end_yearday. Do not decode that float by formatting it as a string: str(2010.21) is "2010.21" and reads as day 21, where the value means day 210.

In [3]:
stations = pd.read_parquet(f"s3://{BUCKET}/{CATALOGUE}/stations.parquet", storage_options={"anon": True})
in_box = stations[stations.latitude.between(LAT0, LAT1) & stations.longitude.between(LON0, LON1)]
# Dates compare with dates. No decoding, which is the point of the change.
sel = in_box[(in_box.start_date <= END) & (in_box.end_date >= START)].copy()
NETWORKS = sorted(sel.network_code.unique())
print(f"{len(in_box):,} station-locations in the box, {len(sel):,} operating in the period, "
      f"{len(NETWORKS)} networks: {NETWORKS}")
print(sel.network_code.value_counts().to_string())
553 station-locations in the box, 166 operating in the period, 10 networks: ['BK', 'CI', 'GS', 'IM', 'LB', 'NC', 'NN', 'NP', 'SN', 'ZG']
network_code
NN    49
NC    44
NP    34
IM    13
SN     7
BK     6
CI     6
LB     3
GS     2
ZG     2
In [4]:
fig, ax = plt.subplots(figsize=(6.5, 6))
for net, g in sel.groupby("network_code"):
    ax.scatter(g.longitude, g.latitude, s=14, label=f"{net} ({len(g)})")
ax.plot(CENTER_LON, CENTER_LAT, marker="*", ms=16, color="#16150f", mec="w", ls="none", label="mainshock")
ax.set_xlim(LON0, LON1); ax.set_ylim(LAT0, LAT1); ax.set_aspect(1 / np.cos(np.radians(CENTER_LAT)))
ax.set_xlabel("longitude"); ax.set_ylabel("latitude")
ax.legend(fontsize=7.5, ncol=2, frameon=False, loc="lower left")
ax.set_title(f"{len(sel)} station-locations operating {START} to {END}", loc="left", fontsize=10.5)
fig.tight_layout()
No description has been provided for this image

3. Which partitions cover them¶

A partition is (campaign, network, year, month). The era decides the campaign; a network partition holds every station of that network, so the station filter is applied to the rows after the download (section 5).

In [5]:
def partitions(networks, start, end):
    out = []
    d = start.replace(day=1)
    while d <= end:
        for net in networks:
            out.append((CATALOGUE, net, d.year, d.month))
        d = (d.replace(day=28) + datetime.timedelta(days=4)).replace(day=1)
    return out

PARTS = partitions(NETWORKS, START, END)
print(f"{len(PARTS)} partitions, e.g. {PARTS[:3]}")
20 partitions, e.g. [('western', 'BK', 2020, 5), ('western', 'CI', 2020, 5), ('western', 'GS', 2020, 5)]

4. Mirror them¶

One listing per partition, then every missing file in parallel. s3fs fetches concurrently and lands at 15 to 20 MB/s from a laptop; the AWS CLI is the same speed (aws s3 sync --no-sign-request <s3 prefix> <local dir>, one line per partition, printed below for people who prefer a shell).

In [6]:
def sync(camp, net, year, month):
    # Mirror one partition; return (files fetched, bytes fetched, files already present).
    prefix = f"{camp}/picks/network={net}/year={year}/month={month:02d}/"
    local = MIRROR / prefix
    local.mkdir(parents=True, exist_ok=True)
    try:
        remote = fs.ls(f"{BUCKET}/{prefix}", detail=True)      # one LIST
    except FileNotFoundError:                                  # no picks were written for this partition
        remote = []
    have = {p.name for p in local.glob("*.parquet")}
    todo = [o for o in remote if Path(o["name"]).name not in have]
    if todo:
        fs.get([o["name"] for o in todo], [str(local / Path(o["name"]).name) for o in todo])
    return len(todo), sum(o["size"] for o in todo), len(have)

t0 = time.time(); n_new = b_new = n_have = 0
for camp, net, year, month in PARTS:
    n, b, h = sync(camp, net, year, month)
    n_new += n; b_new += b; n_have += h
print(f"fetched {n_new:,} files, {b_new / 1e6:,.0f} MB in {time.time() - t0:.0f} s; {n_have:,} already mirrored")
empty = [p for p in PARTS if not any((MIRROR / f"{p[0]}/picks/network={p[1]}/year={p[2]}/month={p[3]:02d}").glob("*.parquet"))]
print(f"{len(empty)} of {len(PARTS)} partitions hold no picks at all: {sorted({p[1] for p in empty})}")

print("\nequivalent shell commands:")
for camp, net, year, month in PARTS[:3]:
    p = f"{camp}/picks/network={net}/year={year}/month={month:02d}/"
    print(f"aws s3 sync --no-sign-request s3://{BUCKET}/{p} {MIRROR}/{p}")
print(f"... ({len(PARTS)} partitions)")
fetched 1,597 files, 299 MB in 30 s; 0 already mirrored
2 of 20 partitions hold no picks at all: ['ZG']

equivalent shell commands:
aws s3 sync --no-sign-request s3://quakescope-picks-2026/western/picks/network=BK/year=2020/month=05/ quakescope_mirror/western/picks/network=BK/year=2020/month=05/
aws s3 sync --no-sign-request s3://quakescope-picks-2026/western/picks/network=CI/year=2020/month=05/ quakescope_mirror/western/picks/network=CI/year=2020/month=05/
aws s3 sync --no-sign-request s3://quakescope-picks-2026/western/picks/network=GS/year=2020/month=05/ quakescope_mirror/western/picks/network=GS/year=2020/month=05/
... (20 partitions)

5. Query the mirror¶

DuckDB over the local tree, with the partition keys read from the paths. The station filter is a join against the selection table, so nothing outside the box is counted. Everything comes back aggregated; the raw rows are only pulled where they are needed.

In [7]:
con = duckdb.connect()
con.register("sel", sel[["id", "network_code", "station_code", "latitude", "longitude", "elevation"]])
con.sql(f'''
    CREATE OR REPLACE VIEW picks AS
    SELECT p.*, s.latitude, s.longitude
    FROM read_parquet('{MIRROR}/{CATALOGUE}/picks/**/*.parquet', hive_partitioning=1) p
    JOIN sel s ON p.tid = s.id
    WHERE p.peak >= TIMESTAMP '{START}' AND p.peak < TIMESTAMP '{END + datetime.timedelta(days=1)}'
''')

t0 = time.time()
totals = con.sql("SELECT pha, count(*) AS picks, count(DISTINCT tid) AS stations, round(avg(conf), 3) AS mean_conf "
                 "FROM picks GROUP BY pha ORDER BY pha").df()
print(f"{time.time() - t0:.1f} s")
totals
0.3 s
Out[7]:
pha picks stations mean_conf
0 P 2184746 95 0.636
1 S 2111152 90 0.566
In [8]:
per_station = con.sql('''
    SELECT tid, any_value(cha) AS band,
           count(*) FILTER (pha = 'P') AS p_picks,
           count(*) FILTER (pha = 'S') AS s_picks,
           count(*) FILTER (conf >= 0.5) AS confident,
           count(DISTINCT date_trunc('day', peak)) AS days_with_picks
    FROM picks GROUP BY tid ORDER BY p_picks DESC
''').df()
print(f"{len(per_station)} stations with picks of {len(sel)} selected")
per_station.head(12)
95 stations with picks of 166 selected
Out[8]:
tid band p_picks s_picks confident days_with_picks
0 NN.MCM08. HH 170503 189137 238925 36
1 NN.MC02. HN 150834 151159 214875 37
2 NN.MCM05. HH 137194 153690 194558 36
3 GS.MCA04.00 HH 132567 124616 189088 45
4 NN.MCA06. HH 119784 139576 175237 41
5 NN.MCM07. HH 117041 142072 166067 37
6 NN.MC01. HN 114583 113225 171055 43
7 NN.MC03. HN 109340 111622 163052 43
8 NN.BRS2. HH 64544 64939 99097 47
9 IM.NV11. SH 57078 64875 68279 61
10 NN.LHV. HH 39352 43351 50219 58
11 IM.NV31. BH 36330 38569 46554 61
In [9]:
daily = con.sql('''
    SELECT date_trunc('day', peak) AS day, pha, count(*) AS n
    FROM picks WHERE conf >= 0.5 GROUP BY day, pha ORDER BY day
''').df().pivot(index="day", columns="pha", values="n").fillna(0)

fig, ax = plt.subplots(figsize=(11, 3.4))
for ph, c in (("P", C_P), ("S", C_S)):
    ax.plot(daily.index, daily[ph], lw=1.4, color=c, label=ph)
ax.axvline(pd.Timestamp("2020-05-15"), color="#8a8a8a", lw=1, ls="--")
ax.set_yscale("log"); ax.set_ylabel("picks per day, conf >= 0.5"); ax.legend(frameon=False)
ax.set_title("the box, all selected stations: the M6.5 on May 15 and its aftershocks", loc="left", fontsize=10.5)
fig.tight_layout()
No description has been provided for this image

The threshold is yours to choose¶

Everything at or above 0.2 is stored. How much a higher floor removes depends on the station and the sequence, so look before choosing.

In [10]:
sweep = con.sql('''
    SELECT t.thr, count(*) FILTER (p.conf >= t.thr AND pha = 'P') AS P, count(*) FILTER (p.conf >= t.thr AND pha = 'S') AS S
    FROM picks p, (SELECT unnest([0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9]) AS thr) t
    GROUP BY t.thr ORDER BY t.thr
''').df()
sweep["P_kept"] = (sweep.P / sweep.P.iloc[0]).round(3); sweep["S_kept"] = (sweep.S / sweep.S.iloc[0]).round(3)
sweep
Out[10]:
thr P S P_kept S_kept
0 0.2 2184746 2111152 1.000 1.000
1 0.3 1812098 1731536 0.829 0.820
2 0.4 1561320 1439077 0.715 0.682
3 0.5 1341337 1180497 0.614 0.559
4 0.6 1186801 966724 0.543 0.458
5 0.7 1042478 743221 0.477 0.352
6 0.8 873141 477713 0.400 0.226
7 0.9 608600 169616 0.279 0.080

6. Coverage: which station-days were actually processed¶

A station-day with no rows may mean the archive held nothing, or that the campaign never read it, and the picks alone cannot tell the two apart. The public manifests help: each shard's manifest lists the station-days it processed with their pick counts (records: tid, cha, yr, doy, npks), so a station-day present there with npks: 0 was read and yielded nothing above 0.2.

Two things to know before trusting them. Whether a station-day has picks is read from the Parquet, never from the manifest: a shard that was preempted and resumed by another worker writes a manifest covering only the resuming attempt, while the first attempt's files stay in the bucket, so the manifest under-reports. And the shards found here are the ones named by the files we mirrored (<shard_id>[-NNN].parquet); a shard that processed our stations and wrote nothing into these partitions leaves no filename behind. Being exhaustive means fetching every manifest of the campaign (72,505 GETs for western, about twenty minutes); this cell does the cheap version.

In [11]:
shard_ids = set()
for camp, net, year, month in PARTS:
    for p in (MIRROR / f"{camp}/picks/network={net}/year={year}/month={month:02d}").glob("*.parquet"):
        sid = p.stem
        if sid[-4] == "-" and sid[-3:].isdigit():           # strip the -001 sequence suffix
            sid = sid[:-4]
        shard_ids.add((camp, sid))
print(f"{len(shard_ids)} shards wrote into these partitions")

def fetch_manifest(key):
    camp, sid = key
    try:
        return camp, json.loads(fs.cat(f"{BUCKET}/{camp}/manifests/{sid}.json"))
    except FileNotFoundError:
        return camp, None

t0 = time.time()
with ThreadPoolExecutor(32) as ex:
    manifests = list(ex.map(fetch_manifest, sorted(shard_ids)))
print(f"fetched in {time.time() - t0:.0f} s; {sum(m is None for _, m in manifests)} missing")

rec = pd.DataFrame([dict(campaign=camp, **r) for camp, m in manifests if m for r in m["records"]])
rec["day"] = pd.to_datetime(rec.yr.astype(str) + rec.doy.astype(str).str.zfill(3), format="%Y%j").dt.date
rec = rec[rec.tid.isin(sel.id) & (rec.day >= START) & (rec.day <= END)]
print(f"{len(rec):,} station-day records for our stations in the period, "
      f"{(rec.npks == 0).sum():,} of them with zero picks")
423 shards wrote into these partitions
fetched in 9 s; 0 missing
4,046 station-day records for our stations in the period, 171 of them with zero picks
In [12]:
# Expected station-days: each selected station on each day of the period it was operating.
days = pd.date_range(START, END, freq="D").date
expected = pd.DataFrame([(s.id, d) for s in sel.itertuples() for d in days
                         if s.start_date <= d <= s.end_date], columns=["tid", "day"])

# "picks" comes from the Parquet itself; the manifests add "read, no pick".
have_picks = con.sql("SELECT DISTINCT tid, CAST(date_trunc('day', peak) AS DATE) AS day FROM picks").df()
have_picks["day"] = pd.to_datetime(have_picks.day).dt.date
cov = (expected.merge(have_picks.assign(has_picks=True), on=["tid", "day"], how="left")
               .merge(rec[["tid", "day", "npks"]].drop_duplicates(["tid", "day"]), on=["tid", "day"], how="left"))
cov["status"] = np.select([cov.has_picks.eq(True), cov.npks.notna()], ["picks", "read, no pick"], "not processed")
summary = cov.status.value_counts()
print(f"{len(expected):,} expected station-days on {sel.id.nunique()} stations")
print((summary / len(cov)).round(3).to_string())

by_station = cov.groupby("tid").status.value_counts().unstack(fill_value=0)
never = by_station[by_station.get("picks", 0) + by_station.get("read, no pick", 0) == 0]
print(f"\n{len(never)} of {len(by_station)} stations have no processed day at all in the period")
print(never.head(10).to_string())
9,533 expected station-days on 166 stations
status
not processed    0.515
picks            0.467
read, no pick    0.018

71 of 166 stations have no processed day at all in the period
status      not processed  picks  read, no pick
tid                                            
BK.HELL.               61      0              0
BK.HELL.00             61      0              0
BK.HELL.S0             61      0              0
BK.KCC.                61      0              0
CI.GRA.2C              61      0              0
CI.MLAC.2C             61      0              0
CI.TIN.2C              61      0              0
IM.NVAR.FB             61      0              0
LB.TPH.                61      0              0
NC.MCY.                61      0              0
In [13]:
order = by_station.sort_values("picks", ascending=False).index if "picks" in by_station else by_station.index
code_of = {"not processed": 0, "read, no pick": 1, "picks": 2}
grid = cov.pivot(index="tid", columns="day", values="status").reindex(order).apply(lambda c: c.map(code_of)).astype(float)

fig, ax = plt.subplots(figsize=(12, 0.09 * len(grid) + 1.4))
im = ax.imshow(grid.values, aspect="auto", cmap=plt.matplotlib.colors.ListedColormap(["#e7e7e2", "#f6c77a", "#2a78d6"]),
               vmin=-0.5, vmax=2.5, interpolation="nearest")
ax.set_yticks([]); ax.set_ylabel(f"{len(grid)} stations")
xt = np.arange(0, len(days), 10); ax.set_xticks(xt); ax.set_xticklabels([str(days[i]) for i in xt], fontsize=8)
cb = fig.colorbar(im, ax=ax, ticks=[0, 1, 2], fraction=0.02, pad=0.01)
cb.ax.set_yticklabels(["not processed", "read, no pick", "picks"], fontsize=8)
ax.set_title("station-day coverage from the manifests (stations sorted by days with picks)", loc="left", fontsize=10.5)
fig.tight_layout()
No description has been provided for this image

The rectangular gaps, one shard (20 days) wide, across blocks of stations that have picks on both sides are the pattern to look for: a shard that did not process those station-days. The next cell lists them.

In [14]:
# Runs of consecutive unprocessed days on stations that otherwise have picks.
gaps = []
for tid, g in cov.sort_values("day").groupby("tid"):
    if (g.status == "picks").sum() < 10:
        continue
    run = []
    for row in g.itertuples():
        if row.status == "not processed":
            run.append(row.day)
        elif run:
            gaps.append((tid, run[0], run[-1], len(run))); run = []
    if run:
        gaps.append((tid, run[0], run[-1], len(run)))
gaps = pd.DataFrame(gaps, columns=["tid", "first_day", "last_day", "days"])
long_gaps = gaps[gaps.days >= 5].sort_values(["first_day", "tid"])
print(f"{len(long_gaps)} gaps of 5+ days on stations that have picks; by start day:")
print(long_gaps.groupby(["first_day", "last_day"]).agg(stations=("tid", "size"), days=("days", "first")).to_string())
15 gaps of 5+ days on stations that have picks; by start day:
                       stations  days
first_day  last_day                  
2020-05-01 2020-05-07         1     7
           2020-05-14         3    14
           2020-05-15         2    15
2020-05-18 2020-06-30         8    44
2020-06-12 2020-06-18         1     7

Gaps that start and end on the same days for many stations are shard boundaries (the queue is cut into 20-day shards of about 30 stations). A block of stations skipped for a whole shard is an archive outage, a shard the campaign has set aside (embargoed or awaiting review, listed on the dashboard), or a silent skip of the kind documented in western_pick_validation.html section 10. A gap of one to three days at the start of a shard window on stations that are otherwise complete is the signature of a shard resumed after a preemption, where the resuming worker's first write replaced the first attempt's first checkpoint file. Either is worth reporting (https://github.com/SeisSCOPED/QuakeScope/issues) with this table.

Read "not processed" with care. Most of it is station-locations the planning table lists but the archive holds nothing for: the table carries every location code the metadata ever declared (BK.HELL., BK.HELL.00 and BK.HELL.S0 are three rows for one instrument), and metadata outlives instruments. Some of it is the campaign's own silent skips, documented in western_pick_validation.html section 10. The manifests cannot separate the two; the archive can, by asking it for the day. Stations with picks on most days and a few missing ones are the rows worth asking about.

7. Export¶

Two shapes are useful downstream. A per-station-day summary for QC and coverage statements, and the picks themselves in the columns an associator expects. Both are written under export/.

In [15]:
EXPORT = Path("export"); EXPORT.mkdir(exist_ok=True)

cov.to_csv(EXPORT / "coverage_station_days.csv", index=False)
per_station.to_csv(EXPORT / "picks_per_station.csv", index=False)

# The picks, filtered to the box and the period, one file, all columns. Parquet keeps the types.
con.sql(f"COPY (SELECT * EXCLUDE (latitude, longitude) FROM picks ORDER BY peak) TO '{EXPORT}/picks.parquet' (FORMAT PARQUET)")
n = con.sql(f"SELECT count(*) FROM '{EXPORT}/picks.parquet'").fetchone()[0]
print(f"picks.parquet: {n:,} rows, {(EXPORT / 'picks.parquet').stat().st_size / 1e6:.0f} MB")
picks.parquet: 4,295,898 rows, 132 MB

For an associator¶

PyOcto takes a DataFrame with station, time, phase and probability; GaMMA takes id, timestamp, type, prob and amp. Both want the station table alongside. The mapping is the whole conversion; a confidence floor of 0.5 is a reasonable starting point for association and is not the catalogue's floor.

In [16]:
picks_df = pd.read_parquet(EXPORT / "picks.parquet")
assoc = picks_df[picks_df.conf >= 0.5]

pyocto_picks = pd.DataFrame({"station": assoc.tid.str.rstrip("."),   # PyOcto keys stations as NET.STA[.LOC]
                             "time": assoc.peak, "phase": assoc.pha, "probability": assoc.conf})
pyocto_stations = sel.rename(columns={"id": "id", "latitude": "latitude", "longitude": "longitude", "elevation": "elevation"})
pyocto_stations = pyocto_stations.assign(id=pyocto_stations.id.str.rstrip("."))[["id", "latitude", "longitude", "elevation"]]

gamma_picks = pd.DataFrame({"id": assoc.tid, "timestamp": assoc.peak, "type": assoc.pha.str.lower(),
                            "prob": assoc.conf, "amp": assoc.amp})

pyocto_picks.to_parquet(EXPORT / "pyocto_picks.parquet"); pyocto_stations.to_csv(EXPORT / "pyocto_stations.csv", index=False)
gamma_picks.to_parquet(EXPORT / "gamma_picks.parquet")
print(f"{len(assoc):,} picks at conf >= 0.5 for association")
pyocto_picks.head()
2,521,834 picks at conf >= 0.5 for association
Out[16]:
station time phase probability
4 CI.TIN 2020-05-01 00:00:46.408 P 0.708237
7 NC.MEM 2020-05-01 00:01:20.590 S 0.639534
14 CI.GRA 2020-05-01 00:05:00.018 P 0.809767
15 CI.GRA 2020-05-01 00:05:01.958 S 0.609104
18 NN.LHV 2020-05-01 00:06:19.318 P 0.834368

Running the associator is a separate step and a separate notebook: seisbench_pyocto_ncedc.ipynb shows PyOcto on picks of this shape.

8. Scaling up¶

The recipe above is the recipe at any size; only the byte counts change.

A whole network-year. Replace the month loop with the year:

aws s3 sync --no-sign-request s3://quakescope-picks-2026/western/picks/network=UW/year=2019/ quakescope_mirror/western/picks/network=UW/year=2019/

The whole western catalogue, 42 GB in 417,000 objects (S3 listing, 2026-09-18), about 90 minutes on a fast link:

aws s3 sync --no-sign-request s3://quakescope-picks-2026/western/picks/ quakescope_mirror/western/picks/
aws s3 cp   --no-sign-request s3://quakescope-picks-2026/western/stations.parquet quakescope_mirror/western/

Where the bytes are, by network (same listing):

network objects GB network objects GB
CI 51,561 8.94 UO 10,136 1.31
UW 66,847 5.30 AZ 9,540 1.10
PB 14,441 5.27 CC 7,119 0.94
NC 101,863 3.41 YN 3,861 0.87
BG 5,495 2.01 SB 4,140 0.86
BK 21,426 1.77 TA 6,719 0.80
NN 12,120 1.47 75 others 6.6

And by year: under 0.2 GB a year before 2003, 1.2 GB a year by 2008, 1.6 to 1.8 GB a year through the 2010s and 2.5 to 3.0 GB a year from 2019. The pre-2010 era is small because the archives hold fewer stations then, not because it was picked differently.

Keeping a mirror current. Embargoed years fill in as EarthScope opens them and compaction will rename objects; aws s3 sync re-run against the same local tree fetches only what changed. Record the date of any pull you publish from.


Reference for the layout, the rules and the known limits: docs/data_access.md. Progress of every campaign: the dashboard.