QuakeScope · Engineering note · 2026-09-05
Eight defects from one incident, and what each generalises to. If you are building a client that reads a large scientific archive whose credentials are issued per slice of the data rather than once per session, this is the list of things that will go wrong.
On 2026-09-04 a seismology archive operator told us our processing fleet was denial-of-servicing their credentials endpoint. It was: 351,735 token requests in four hours, peaking at 5,104 per minute, every one rejected. A few hundred AWS Batch workers, reading waveform data whose access credentials are scoped to one network and one year at a time.
Five defects in one file caused it. Fixing those introduced two more, and one of our own fixes nearly introduced a third. None of the eight was exotic; all of them are available to anyone building this shape of client. The full incident report is here; this note is the part that transfers.
Coarse credentials are easy: authenticate once, read everything, renew on a timer. Granular credentials invert the economics. If access is scoped per network and per year, then a job spanning many networks and many years needs many credentials, and the code that fetches them sits inside the innermost read loop rather than in startup.
That relocation is the whole story. A credential exchange in startup runs once and its bugs are obvious. The same call inside a loop that runs once per object runs 11,315 times per shard — and any failure to remember an answer becomes a load multiplier with the loop count as its coefficient.
A 4xx is a verdict on the request. A 429 or 5xx is a statement about the service. Only the second kind is worth retrying. Every one of our eight defects is a variation on failing to hold that line.
Cache refusals as hard as you cache successes, keyed by the exact scope that was refused, and never send that scope again for the life of the process. A refusal is more valuable to cache than a success, because failures repeat harder — a network you cannot read fails on every object, forever, while a success is used once and expires.
Successes were cached by scope. Refusals were not, so each of 11,315 calls per shard re-ran the exchange.
On denial the client flipped a network to "the other" scoping. For codes that are reused between experiments that meant dropping the year — producing a request that can only ever be refused, which is what the operator noticed first.
The vendor SDK caches issued credentials on the client instance. Building one per call discarded that and forced a round trip plus a token refresh.
The rejected-token error had no .response attribute, so it
slipped past every status-code branch into a generic retry loop that re-ran
the token grant five times per scope, on a fixed five-second sleep. When the
account was blocked, every network took that path on every worker at once.
PermissionError and FileNotFoundError both
subclass OSError, and an except OSError sat before
both, making them dead code. The S3 filesystem layer maps
AccessDenied to PermissionError with no
errno, so a denied object fell out of the broad branch without
returning and the enclosing while True re-requested it at network
speed until a timeout. One shard spent 447 minutes on a single object.
We moved the refusal cache to process scope so it would outlive a work unit. The learned "this network needs a year" flag stayed per work unit. The cache key is built from that flag, so unit 1 filed a refusal under one key, learned better, and succeeded; unit 2 rebuilt the old key, found unit 1's refusal, and gave up on the network entirely.
The SDK chooses a synchronous execution strategy the first time it is used
and caches it: no running event loop gets an in-thread runner, a running loop
gets a background-thread one. Our first credential exchange came from
synchronous setup; every later one came from inside an async read loop. So the
cached choice was wrong for every call after the first, and raised a
bare RuntimeError before any request was sent — carrying
no status, therefore not terminal, therefore retried five times.
Pinning the runner required importing two SDK-private modules. Unguarded,
that would have raised ImportError inside the retry loop from #4
— breaking every exchange, far worse than the bug being fixed.
The most useful thing we built is a harness that replays one work unit against scripted responses and counts outbound requests, running the current and the broken build side by side. It turns "we fixed it" into a number, and it keeps working as a regression detector: CI asserts the old build still fails every check, so if the harness stops seeing the bug, the harness is broken.
| scenario | before | after |
|---|---|---|
| 403 — no access to this network | 11,316 | 1 |
| 404 — no such network-year | 11,315 | 1 |
| 400 — scope missing its year | 11,316 | 2 |
| rejected token, six networks | 10,980 | 1 |
| retry sleeps | 15.2 h | 0 |
Defect 5 had a test asserting handler order by parsing the source. It passed throughout and told us nothing, because the branch had never executed — not in five days of production logs, not in either dry run. Only a test that drives the read loop with a filesystem that refuses shows it: the broken build was still looping after 201 requests, the fixed one returns after three.
Defect 7 was invisible to our offline checks by construction: they call the read path synchronously, so the SDK always picks correctly there. No number of additional offline checks would have found it. Some classes of bug require a real, bounded run against the real service.
Our first published figures came from substring counts over log lines, which match the request, the warning about it, and every retry. We reported 20,024 year-less requests across eight network codes. Parsing the actual request lines gives 3,925, from two codes — five of the eight had made no request at all. The ratio survived the correction, because numerator and denominator inflated together, which is exactly what makes this error hard to notice.
We did not detect this. The people we were overloading told us. Nothing in the stack alarmed on outbound request rate, and the log group that held the only evidence had five-day retention — shorter than the investigation. A metric filter on the outbound-request line, an alarm wired to stop the fleet, and retention longer than your incident-response window are all cheap next to the alternative.
After stopping the fleet we wrote publicly that it would stay stopped. Nothing enforced that: the controller was a workflow any collaborator could trigger, and every job definition still pointed at the build that caused the incident. The control is a quarantine list checked against live infrastructure before anything is submitted, which fails the run rather than trusting a config file to describe itself.
Process-wide caches multiply by worker count and by processes per worker: 208 machines at four processes each is 832 independent caches, each of which must learn every refusal once, and all of which reset on preemption. "One request per refused scope" is true per process and misleading per fleet.
The defects above suggest what a client will get wrong, and therefore what the service can usefully do about it:
Make refusals unambiguous. Distinct codes for "no access" and "no such thing" let a client cache correctly. We conflated them for two weeks, because an unscoped credential lists successfully and only fails at the read — an asymmetry that reads like a bug in the reader.
Say when a retry could succeed. Retry-After on 429 and
5xx, and nothing resembling it on 4xx, tells a careful client exactly what
to do.
Publish the scoping rule as data. We inferred which identifiers need a year qualifier from prose documentation. A wrong guess in either direction costs a request per network, and a client cannot discover the truth except by being refused.
Rate-limit rather than trust. A 429 with a budget is a control the client cannot forget to implement. Our worst hour would have been one rejected minute.