QuakeScope · Engineering note · 2026-09-05

Reading an archive behind per-object credentials

Eight defects from one incident, and what each generalises to. If you are building a client that reads a large scientific archive whose credentials are issued per slice of the data rather than once per session, this is the list of things that will go wrong.

On 2026-09-04 a seismology archive operator told us our processing fleet was denial-of-servicing their credentials endpoint. It was: 351,735 token requests in four hours, peaking at 5,104 per minute, every one rejected. A few hundred AWS Batch workers, reading waveform data whose access credentials are scoped to one network and one year at a time.

Five defects in one file caused it. Fixing those introduced two more, and one of our own fixes nearly introduced a third. None of the eight was exotic; all of them are available to anyone building this shape of client. The full incident report is here; this note is the part that transfers.

The shape of the problem

Coarse credentials are easy: authenticate once, read everything, renew on a timer. Granular credentials invert the economics. If access is scoped per network and per year, then a job spanning many networks and many years needs many credentials, and the code that fetches them sits inside the innermost read loop rather than in startup.

That relocation is the whole story. A credential exchange in startup runs once and its bugs are obvious. The same call inside a loop that runs once per object runs 11,315 times per shard — and any failure to remember an answer becomes a load multiplier with the loop count as its coefficient.

The one rule

A 4xx is a verdict on the request. A 429 or 5xx is a statement about the service. Only the second kind is worth retrying. Every one of our eight defects is a variation on failing to hold that line.

Cache refusals as hard as you cache successes, keyed by the exact scope that was refused, and never send that scope again for the life of the process. A refusal is more valuable to cache than a success, because failures repeat harder — a network you cannot read fails on every object, forever, while a success is used once and expires.

Eight defects

1 · Asymmetric caching

Successes were cached by scope. Refusals were not, so each of 11,315 calls per shard re-ran the exchange.

Generalises to: if you cache the happy path, ask what happens to the unhappy one. This is the single highest-leverage fix in the list.

2 · A retry that broadened the request

On denial the client flipped a network to "the other" scoping. For codes that are reused between experiments that meant dropping the year — producing a request that can only ever be refused, which is what the operator noticed first.

Generalises to: a retry that changes the request may only ever make it more specific. Broadening on failure turns one bad request into a class of them.

3 · A fresh SDK client per call

The vendor SDK caches issued credentials on the client instance. Building one per call discarded that and forced a round trip plus a token refresh.

Generalises to: know what your SDK caches and where. "Construct it fresh each time" is the safe-looking default that disables the library's own protections.

4 · An exception with no status code

The rejected-token error had no .response attribute, so it slipped past every status-code branch into a generic retry loop that re-ran the token grant five times per scope, on a fixed five-second sleep. When the account was blocked, every network took that path on every worker at once.

Generalises to: status-code dispatch that assumes an HTTP-shaped exception. Errors about your credentials must stop everything, not be retried per scope. And a fixed sleep synchronises a fleet — use exponential backoff with full jitter.

5 · A shadowed exception handler

PermissionError and FileNotFoundError both subclass OSError, and an except OSError sat before both, making them dead code. The S3 filesystem layer maps AccessDenied to PermissionError with no errno, so a denied object fell out of the broad branch without returning and the enclosing while True re-requested it at network speed until a timeout. One shard spent 447 minutes on a single object.

Generalises to: Python's exception hierarchy makes later handlers silently unreachable. Order narrow before broad — and note that a loop whose handler falls through without returning is an infinite loop wearing a disguise.

6 · Two caches, one key space, different lifetimes

We moved the refusal cache to process scope so it would outlive a work unit. The learned "this network needs a year" flag stayed per work unit. The cache key is built from that flag, so unit 1 filed a refusal under one key, learned better, and succeeded; unit 2 rebuilt the old key, found unit 1's refusal, and gave up on the network entirely.

Generalises to: state read together must share a lifetime. If A computes the key for B, scope A and B identically. Two caches with different lifetimes over a shared key space will disagree, and single-iteration tests cannot see it. Found by the archive operator's engineers reading our fix, not by us.

7 · A cached runner that outlived its context

The SDK chooses a synchronous execution strategy the first time it is used and caches it: no running event loop gets an in-thread runner, a running loop gets a background-thread one. Our first credential exchange came from synchronous setup; every later one came from inside an async read loop. So the cached choice was wrong for every call after the first, and raised a bare RuntimeError before any request was sent — carrying no status, therefore not terminal, therefore retried five times.

Generalises to: the same bug as #6, one layer up. Also: a bare exception type is a classification failure. Anything reaching a retry loop without a status will be retried, so the loop's default must be "give up loudly".

8 · A fix that depended on a private API

Pinning the runner required importing two SDK-private modules. Unguarded, that would have raised ImportError inside the retry loop from #4 — breaking every exchange, far worse than the bug being fixed.

Generalises to: depending on a private API is sometimes right, but it must degrade. Guard the import, fall back to the supported path, and log why.

How to find these before the operator does

Count requests that would leave the process

The most useful thing we built is a harness that replays one work unit against scripted responses and counts outbound requests, running the current and the broken build side by side. It turns "we fixed it" into a number, and it keeps working as a regression detector: CI asserts the old build still fails every check, so if the harness stops seeing the bug, the harness is broken.

Requests reaching the archive, one work unit, one network
scenariobeforeafter
403 — no access to this network11,3161
404 — no such network-year11,3151
400 — scope missing its year11,3162
rejected token, six networks10,9801
retry sleeps15.2 h0

Test the branch, not the shape

Defect 5 had a test asserting handler order by parsing the source. It passed throughout and told us nothing, because the branch had never executed — not in five days of production logs, not in either dry run. Only a test that drives the read loop with a filesystem that refuses shows it: the broken build was still looping after 201 requests, the fixed one returns after three.

An offline harness cannot see context-dependent bugs

Defect 7 was invisible to our offline checks by construction: they call the read path synchronously, so the SDK always picks correctly there. No number of additional offline checks would have found it. Some classes of bug require a real, bounded run against the real service.

Measure with a parser, not a grep

Our first published figures came from substring counts over log lines, which match the request, the warning about it, and every retry. We reported 20,024 year-less requests across eight network codes. Parsing the actual request lines gives 3,925, from two codes — five of the eight had made no request at all. The ratio survived the correction, because numerator and denominator inflated together, which is exactly what makes this error hard to notice.

Instrument the thing you could be blamed for

We did not detect this. The people we were overloading told us. Nothing in the stack alarmed on outbound request rate, and the log group that held the only evidence had five-day retention — shorter than the investigation. A metric filter on the outbound-request line, an alarm wired to stop the fleet, and retention longer than your incident-response window are all cheap next to the alternative.

A promise is not a control

After stopping the fleet we wrote publicly that it would stay stopped. Nothing enforced that: the controller was a workflow any collaborator could trigger, and every job definition still pointed at the build that caused the incident. The control is a quarantine list checked against live infrastructure before anything is submitted, which fails the run rather than trusting a config file to describe itself.

Count processes, not machines

Process-wide caches multiply by worker count and by processes per worker: 208 machines at four processes each is 832 independent caches, each of which must learn every refusal once, and all of which reset on preemption. "One request per refused scope" is true per process and misleading per fleet.

If you operate the archive

The defects above suggest what a client will get wrong, and therefore what the service can usefully do about it:

Make refusals unambiguous. Distinct codes for "no access" and "no such thing" let a client cache correctly. We conflated them for two weeks, because an unscoped credential lists successfully and only fails at the read — an asymmetry that reads like a bug in the reader.

Say when a retry could succeed. Retry-After on 429 and 5xx, and nothing resembling it on 4xx, tells a careful client exactly what to do.

Publish the scoping rule as data. We inferred which identifiers need a year qualifier from prose documentation. A wrong guess in either direction costs a request per network, and a client cannot discover the truth except by being refused.

Rate-limit rather than trust. A 429 with a budget is a control the client cannot forget to implement. Our worst hour would have been one rejected minute.