QuakeScope · Incident report · 2026-09-04
Five defects in one file turned every EarthScope refusal into thousands of repeat requests. Each links to the line that caused it and the line that replaced it. A self-test reproduces the measurements on your own account.
The fix is deployed. Nothing is running. Those are different
things. All three campaign job definitions now carry the fixed build,
their credentials are split across two roles rather than one, and every
campaign target is 0 — so no worker exists and none will
until somebody raises a target deliberately. c5846f3, the
build that caused this, remains quarantined: the fleet workflow reads the
image out of AWS Batch before submitting anything and refuses to launch
that tag.
Read from the account rather than from our notes, because a document claiming a deployment it never made is how this went wrong once already:
| Campaign | Job definition | Image | Target |
|---|---|---|---|
| western | quakescope_2026_western:20 | quakescope:3c74c81 | 0 |
| global | quakescope_2026_global:6 | quakescope:3c74c81 | 0 |
| obs | quakescope_2026_obs:17 | quakescope:3c74c81 | 0 |
| quarantined | quakescope:c5846f3 — the incident build, refused by the fleet workflow | — | |
So the question that matters is not what the fixed client does — it is
what stops the old one running again. Until this incident the answer was
three zeros in a configuration file and our intention to leave them there.
Our fleet is held at its target by a scheduled workflow that anyone with
write access to the repository can trigger, including from a phone. One
dispatch would have put 208 workers back on c5846f3.
That is now a control rather than a promise. The fleet workflow resolves
each campaign's job definition against Batch, live, and refuses to launch
any image on a
quarantine list
that c5846f3 is on. It fails before submitting anything, on
the scheduled top-up and the manual dispatch alike, and it reads the image
Batch would actually pull rather than trusting the configuration file to
describe itself. Lifting it takes a reviewable commit.
One gap we would rather state than have you find: the guard covers the
fleet controller, which is how workers are launched. A direct
submit-job against the queue would bypass it. Say the word and
we will disable the queues outright.
Beyond the client fix: the Batch role held AmazonS3FullAccess
and now reaches one bucket, the platform's permissions are split from the
container's so a worker cannot read the secret it is handed, the
catalogue bucket has versioning so a bad delete is recoverable, and the
fleet workflow now counts your token-endpoint traffic every 15 minutes
and stops every campaign by itself if it exceeds a threshold. That last
one is the gap that mattered: we did not detect this, you told us.
With your go-ahead we are running one bounded test before any campaign
restarts. It is 2 workers on quakescope_2026_global:3,
which is the merged build 833743e, against a queue of
8 shards and 87 station-days written for this purpose — only
4F and 3J, two of the temporary codes in your
year-less table. Temporary codes are the year-scoped path that produced
the requests you flagged, so that is the path worth testing first, and
three of the eight shards straddle a year boundary and so need two
scoped credentials rather than one.
Every campaign target is 0 and we will tell you before any of them goes above zero.
Your report: our workers were denial-of-servicing the credentials endpoint,
400/403/404 answers were being
retried, and some requests asked for a temporary network without a year.
All three are correct. Mechanism and logs for each are below.
Everything is in one file,
sb_catalog/src/s3_helper.py,
at c5846f3, the build that was running.
CompositeS3ObjectHelper.get_filesystem() is called
once per day per network by the listing loop and
once per object by the read path: 11,315 calls for one
network in one year-long, 30-station shard. Nothing remembered a refusal,
so nearly every call became a request.
The account block increased the load. A rejected refresh token surfaces from
earthscope_sdk as InvalidRefreshTokenError, which
has no .response attribute, so it passed our HTTP status
handling into a generic retry loop that re-ran the refresh grant five times
per scope. After the block, every network took that path, on 208 workers.
CloudWatch log group /aws/batch/job, us-east-2.
Each count is one POST to
login.earthscope.org/oauth/token. In this window all returned
invalid_grant · user is blocked.
HTTP Request: POST
https://login.earthscope.org/oauth/token lines in the exported
events, one per request.
D2 in one worker's log, spanning 88 ms. The 403
triggers the year-less retry; your 400 body states the rule:
14:59:50.125 INFO HTTP Request: GET https://api.earthscope.org/beta/user/
credentials/aws/s3-miniseed-v2?network=FDSN%3A7D&year=2025
"HTTP/1.1 403 Forbidden"
14:59:50.126 WARNING Scope {'network': 'FDSN:7D', 'year': 2025} refused:
EarthScope returned 403 ... on role s3-miniseed-v2
14:59:50.126 WARNING EarthScope: 7D denied under 'network+year' scoping;
retrying as 'network'.
14:59:50.157 INFO HTTP Request: GET https://api.earthscope.org/beta/user/
credentials/aws/s3-miniseed-v2?network=FDSN%3A7D
"HTTP/1.1 400 Bad Request"
14:59:50.158 WARNING Scope {'network': 'FDSN:7D'} refused: HTTP 400
{"detail":"Temporary networks require a year. See
https://docs.fdsn.org/projects/source-identifiers/en/
latest/network-codes.html#transitional-mapping."}
7D is a temporary code, so the second request could only
return 400. The client sent it because it read a
403 as a scoping error. 3,925 of our requests carried
a temporary code and no year — 24.4% of the 16,070 refusals we were
given, and every one of them could only be answered 400.
Two codes account for all of them:
| Code | Requests |
|---|---|
| 7D | 3,442 |
| 2F | 483 |
| every other temporary code | 0 |
An earlier version of this page put the year-less total at 20,024 and
spread it over eight codes, and the token storm at 260,322 requests over
66 minutes. Both came from filter_log_events substring
counts, which count every log line that mentions a scope — the
request, the warning that follows it, and the retry — rather than
one line per request. That inflates some numbers and, because the storm
was measured over too narrow a window, understates others.
Every figure on this page is now derived by parsing the
HTTP Request: lines themselves, one per outbound request,
from an export of all 10,797,274 events in the window. On that basis the
year-less total is 3,925 rather than 20,024, only 7D
and 2F ever sent such a request, and the token storm is
351,735 requests over four hours rather than 260,322 over one.
The peak, 5,104/min at 15:55, is unchanged. We would rather you had the
smaller number and the larger one both right than a tidy story.
Both builds replayed against the same scripted responses: one shard of calls (365 days × 30 stations) at a single network. Only the client differs. The harness counts calls that would have left the process. No traffic reached your servers.
| Scenario | c5846f3 | 713698e | Change |
|---|---|---|---|
| 403 — no access to this network | 11,316 | 1 | ÷11,316 |
| 404 — no such network-year | 11,315 | 1 | ÷11,315 |
| 400 — scope missing its year | 11,316 | 2 | ÷5,658 |
| Rejected refresh token, six networks | 10,980 | 1 | ÷10,980 |
| Year-less requests for temporary codes | 2,400 | 0 | eliminated |
| Retry sleeps (5 s each, no jitter) | 10,980 | 0 | 15.2 h → 0 |
The last row is wall-clock. The old client spent 15.2 hours per shard per worker sleeping between retries, which is why it could saturate the endpoint without finishing a shard.
Against the live endpoint with our blocked account, the fixed client made 2 requests for 200 calls. The old client sends about 1,000.
Each block links to the exact lines. Red: c5846f3, the build
that was running. Green: 713698e on
PR #32.
Successful credentials were cached by scope. Refusals were not, so each of the 11,315 calls per shard re-ran the exchange.
try:
self.credentials[key] = self.get_es_credential(net, year)
except RuntimeError as exc:
logger.warning(f"Scope {scope} refused: {exc}")
if not self.escalate_scope_mode(net):
raise
return self.get_es_filesystem(net, year)
self.set_es_filesystem(key)
verdict = self.es_verdict(net, year)
if verdict is not None:
raise verdict # no request is sent
try:
self.credentials[key] = self.get_es_credential(net, year)
except ES_TERMINAL as exc:
self.es_record_verdict(net, year, exc)
...
The store is at module scope, not on the helper: a worker claims shard after shard in one process and built a fresh helper each time, discarding per-instance memory every few minutes. It keeps the exception class and message rather than the instance, so no worker pins the frames the exception passed through.
_ES_STATE = {
"refused": {}, # scope key -> terminal verdict, never re-asked
"auth_failed": None, # 401: not scoped to a network, so it stops everything
"exchanges": {}, # scope key -> [monotonic stamps], for the rate limit
"denials": {}, # scope key -> consecutive AccessDenied count
"client": None, # the one earthscope_sdk client this process uses
}
Also
es_verdict and es_record_verdict
and the terminal-type list
ES_TERMINAL.
Every 4xx has a type in that tuple, including
EarthScopeRequestRefused
for statuses not yet seen, so none can slip past the store.
Source of the year-less temporary-network requests. On any denial the client flipped a network's scoping to the alternative. For a temporary code, the alternative drops the year.
tried.add(current)
other = "network" if current == "network+year" else "network+year"
if other in tried:
return False
self.es_scope_mode[net] = other
tried.add(current)
if current != "network":
return False # nothing safe left to ask
if "network+year" in tried:
return False
self.es_scope_mode[net] = "network+year"
It was reachable from a 403: the class raised for one
subclassed RuntimeError, so the handler in D1 caught it. The
read path called it once per object.
Escalation now only adds the year, only on a 400, at
most once per network. A request that cannot succeed is answered
locally:
if self.es_scope_mode.get(net) == "network+year" and "year" not in scope:
exc = EarthScopeScopeIncomplete(
f"Refusing to ask for {scope} without a year: {net} is a "
f"temporary FDSN code, those are reused across experiments, "
f"and EarthScope authorises them per year. A year-less request "
f"can only ever return 400, so it is not sent."
)
self.es_record_verdict(net, year, exc)
raise exc
The workflow was not reusing credentials at the library level: a fresh
EarthScopeClient was built inside the exchange function.
for attempt in range(1, ES_CREDENTIAL_ATTEMPTS + 1):
try:
with EarthScopeClient() as client:
return client.user.get_aws_credentials(
role=EARTHSCOPE_ROLE, ...)
for attempt in range(1, ES_CREDENTIAL_ATTEMPTS + 1):
try:
return self.es_client().user.get_aws_credentials(
role=EARTHSCOPE_ROLE,
ttl_threshold=self.ttl_threshold,
**scope,)
In earthscope-sdk 1.8.0, _UserService caches
issued credentials in _aws_creds_by_key, an instance
attribute; the on-disk cache present in 1.3.x was removed. A client per
call discarded that cache and forced a round trip, plus an OAuth refresh
when the access token had not been persisted.
es_client()
now returns one client per process.
Produced the traffic in the chart above. 401 and
403 were special-cased; everything else read
exc.response.status_code. An AuthFlowError has no
.response, so it reached the retry loop and slept.
resp = getattr(exc, "response", None) # None for InvalidRefreshTokenError
detail = f"{type(exc).__name__}"
if resp is not None: # ... so every 4xx check is skipped
...
logger.warning(f"EarthScope credential request failed for {scope} "
f"({attempt}/{ES_CREDENTIAL_ATTEMPTS}): {detail}. "
f"Sleeping 5 seconds.")
time.sleep(5) # x5, each attempt re-runs the refresh grant
HTTP Request: POST https://login.earthscope.org/oauth/token "HTTP/1.1 403 Forbidden"
error during token refresh (1 attempts): {"error":"invalid_grant",
"error_description":"user is blocked"}
EarthScope credential request failed for {'network': 'FDSN:4F', 'year': 2019}
(4/5): InvalidRefreshTokenError. Sleeping 5 seconds.
if isinstance(exc, AuthFlowError):
# Every remaining AuthFlowError is about OUR credentials, not about
# this scope: InvalidRefreshTokenError, NoRefreshTokenError, the
# device-code errors. None is congestion, and none is fixed by asking
# again - but they carry no `.response`, so the status-code branch
# below never saw them and they fell into the retry loop.
raise EarthScopeAuthFailed(...) # terminal, and process-wide
A rejected token is not a fact about a network, so
EarthScopeAuthFailed is recorded in
_ES_STATE["auth_failed"] rather than per scope and
es_verdict() checks it first: one rejection stops every
network in the process. The remaining retry loop handles
429 and 5xx only, with
exponential backoff, full jitter and Retry-After.
The previous fixed 5-second sleep returned every affected worker at the
same instant.
PermissionError and FileNotFoundError both
subclass OSError, and an OSError handler sat
before both, so neither could run. Present since June.
except OSError as e:
if e.errno == 5:
return obspy.Stream()
# AccessDenied has errno None -> falls through,
# returns nothing, `while True` goes round again
except PermissionError as e: # unreachable
...
except (FileNotFoundError, ...): # unreachable
except PermissionError as e:
denied += 1
scope_denied = self.s3helper.note_access_denied(net, year)
...
except FileNotFoundError:
return obspy.Stream()
except OSError as e:
if e.errno == 5: ...
s3fs maps AccessDenied to
PermissionError with a message and no errno. A
denied object fell out of the OSError branch without
returning, the while True loop repeated, and the object was
re-HEADed at socket speed until the 900-second station-day
timeout.
Across the full five-day CloudWatch retention, the log lines that
only the PermissionError branch can emit do not
appear:
0 "Credential refreshed after access denied" ← PermissionError branch
0 "Access denied ... times for" ← same branch, budget spent
0 "Not authorized to access this resource" ← OSError errno==5 branch
18 "Timeout after" ← the 900 s station-day cap
The ES_DENIED_ATTEMPTS budget that bounds credential
refreshes therefore never executed. During the blocked window workers
failed earlier, at the credential stage, so this defect contributed
little to the traffic in the chart. The 18 station-day timeouts are its
observable symptom.
A test
parses the handler chain and asserts each narrow type precedes
OSError.
note_access_denied
counts per scope, resetting only on a successful read.
_es_throttle
refuses more than three exchanges of the same scope in five minutes and
sends nothing. A credential is valid for an hour.
netyear_sweep
probes about 1,600 planned network-years and now takes a
--pace, default 0.25 s.
Nothing in our AWS account grants access to your archive.
s3-miniseed-v2 is an EarthScope role alias, not an IAM role in
our account, and the credentials that read data are issued by you. That
determines what you need to run the test: nothing from us.
network and year. Our AWS account reads the
refresh token from Secrets Manager into the container environment. A tester
with their own EarthScope login reproduces the upper lane and needs no
access to the lower one.
The whole of the AWS-side configuration, none of it needed to run the test. This page is public, so account and secret identifiers are described rather than printed. Ask and we will send the full ARNs directly.
AmazonS3FullAccess on the job role is broader than needed. It
grants nothing against your archive, since those reads use the credentials
you issue. We are scoping it to our own buckets in a separate change.
scripts/earthscope_selftest.py replays a full shard of calls and counts what would leave the process. Offline mode sends nothing and needs no EarthScope account. Live mode uses your account and enforces a hard request cap.
Python 3.10+. Verified from a clean virtualenv: no AWS credentials, no access to our account.
git clone https://github.com/SeisSCOPED/QuakeScope.git
cd QuakeScope
python -m venv .venv && source .venv/bin/activate
pip install "earthscope-sdk>=1.8.0" s3fs obspy numpy
Six assertions, each replaying 11,315 get_filesystem() calls
against a scripted response. Exit code 0 only if all pass.
python scripts/earthscope_selftest.py --offline
CURRENT - this working tree
--------------------------------------------------------------
PASS 403 is asked once requests sent: 1 expected: 1
PASS 404 is asked once requests sent: 1 expected: 1
PASS 400 is corrected once, by adding the year
requests sent: 2 expected: 2
PASS a blocked account stops every network at once
requests sent: 1 expected: 1
PASS no yearless request for a temporary network
requests sent: 0 expected: 0
PASS no retry sleeps at all requests sent: 0 expected: 0
ALL CHECKS PASSED
Pass --baseline a copy of the old file and the script runs
both side by side. It counts the retry sleeps instead of serving them, so
the run finishes in seconds rather than days.
git show c5846f3:sb_catalog/src/s3_helper.py > /tmp/old_s3_helper.py
python scripts/earthscope_selftest.py --offline --baseline /tmp/old_s3_helper.py
BASELINE - the version deployed during the incident
--------------------------------------------------------------
FAIL 403 is asked once requests sent: 11316 expected: 1
FAIL 404 is asked once requests sent: 11315 expected: 1
FAIL 400 is corrected once requests sent: 11316 expected: 2
[no year -> year=2019 -> year=2019 -> year=2019]
FAIL a blocked account stops everything requests sent: 10980 expected: 1
FAIL no yearless request for a temporary network
requests sent: 2400 expected: 0
[2406 requests total, 2400 of them yearless]
FAIL no retry sleeps at all requests sent: 10980 expected: 0
[10,980 sleeps totalling 54,900 s = 15.2 h per shard, per worker]
Device-code flow through the SDK. No extra package, no password typed into
our script. It prints a URL and a code; you approve it in a browser and
the SDK writes tokens to
~/.earthscope/default/tokens.json.
python scripts/earthscope_selftest.py --login
The client also reads ES_OAUTH2__REFRESH_TOKEN from the
environment, which is how our Batch jobs are configured:
export ES_OAUTH2__REFRESH_TOKEN=<your own refresh token>
Pick three networks: one you can read, one you cannot (403),
and one FDSN code that does not exist (404). The script
drives 200 calls at each and prints every outbound request with its query
string. It aborts if it exceeds --max-requests.
python scripts/earthscope_selftest.py --live \
--allowed-network AV --allowed-year 2019 \
--denied-network LH \
--missing-network 5A --missing-year 2018
Watch the request count per case and the
yearless temporary-network requests total. A run against our
blocked account, which is why it stops at the token exchange:
a network you CANNOT read (403): LH
-> GET https://api.earthscope.org/beta/user/credentials/aws/
s3-miniseed-v2?network=FDSN%3ALH
-> POST https://login.earthscope.org/oauth/token
2 request(s) for 200 calls (EarthScopeAuthFailed)
total requests sent this run: 2 (cap 4)
yearless temporary-network requests: 0 (must be 0)
ALL CHECKS PASSED
The old build sends about 1,000 POSTs for the same 200 calls.
The container is public and pullable anonymously, and now carries the self-test itself:
docker run --rm ghcr.io/seisscoped/quakescope:<tag> selftest --offline
Live mode takes the same flags, with
-e ES_OAUTH2__REFRESH_TOKEN for your own token. Use a tag
from 833743e onwards; earlier images do not contain the
script.
An earlier version of this page gave this as
docker run … python scripts/earthscope_selftest.py,
which could not have worked: the image is built from
sb_catalog/ so scripts/ was never in it, and
the entrypoint is python -m src.picker, which would have
swallowed those arguments. The build context now takes in the whole
repository and the picker gained a selftest subcommand.
Nothing here is a request. The defects were in our client, not in your API, and the fix has been merged and exercised against your endpoint under the bounded conditions described above.
Two dry runs. The first, 2 workers over 8 shards of 4F and
3J, sent 8 credential requests across 5 scopes and
0 year-less requests, and cached every refusal. It also found a
defect of our own making: our fix reused one SDK client per process, and
the SDK caches its synchronous runner on first use, so a credential
exchange from inside the read loop raised before any request was sent.
Three shards that crossed a year boundary were lost to it. That is fixed,
and the same eight shards then completed 8/8, issuing the
second-year credentials that had never once left the process.
The second read real data: 6 shards, 1,500 station-days of 4F
in 2015 and 2017, 1,240,055 picks written. It spent 1 h 45 min and
sent 15 credential requests over 2 scopes, four of them mid-read
renewals when the one-hour credential expired — the case that would have
failed before the runner fix, and the one that covers most of a real
campaign.
The three campaign job definitions still name c5846f3 and the
fleet workflow refuses to launch it. We will not raise a campaign target
without telling you first.