QuakeScope · Incident report · 2026-09-04

EarthScope Credential Audit

Five defects in one file turned every EarthScope refusal into thousands of repeat requests. Each links to the line that caused it and the line that replaced it. A self-test reproduces the measurements on your own account.

Fleet
Stopped208 workers terminated, queue empty 16:37 UTC, all campaign targets 0. The fix is not yet deployed — see below.
Measured peak
5,104POSTs/min to the token endpoint at 15:55. 351,735 in total, all rejected.
Year-less requests
3,925Requests with a temporary FDSN code and no year. 24.4% of all refusals.
After the fix
1 requestPer refused scope, per process. The replayed shard sent 11,316.

Current deployment state

The fix is deployed. Nothing is running. Those are different things. All three campaign job definitions now carry the fixed build, their credentials are split across two roles rather than one, and every campaign target is 0 — so no worker exists and none will until somebody raises a target deliberately. c5846f3, the build that caused this, remains quarantined: the fleet workflow reads the image out of AWS Batch before submitting anything and refuses to launch that tag.

Read from the account rather than from our notes, because a document claiming a deployment it never made is how this went wrong once already:

EarthScope-credentialed job definitions · us-east-2 · read from AWS Batch, 2026-09-05
Campaign Job definition Image Target
westernquakescope_2026_western:20quakescope:3c74c810
globalquakescope_2026_global:6quakescope:3c74c810
obsquakescope_2026_obs:17quakescope:3c74c810
quarantinedquakescope:c5846f3 — the incident build, refused by the fleet workflow—

So the question that matters is not what the fixed client does — it is what stops the old one running again. Until this incident the answer was three zeros in a configuration file and our intention to leave them there. Our fleet is held at its target by a scheduled workflow that anyone with write access to the repository can trigger, including from a phone. One dispatch would have put 208 workers back on c5846f3.

That is now a control rather than a promise. The fleet workflow resolves each campaign's job definition against Batch, live, and refuses to launch any image on a quarantine list that c5846f3 is on. It fails before submitting anything, on the scheduled top-up and the manual dispatch alike, and it reads the image Batch would actually pull rather than trusting the configuration file to describe itself. Lifting it takes a reviewable commit.

One gap we would rather state than have you find: the guard covers the fleet controller, which is how workers are launched. A direct submit-job against the queue would bypass it. Say the word and we will disable the queues outright.

What changed on our side since

Beyond the client fix: the Batch role held AmazonS3FullAccess and now reaches one bucket, the platform's permissions are split from the container's so a worker cannot read the secret it is handed, the catalogue bucket has versioning so a bad delete is recoverable, and the fleet workflow now counts your token-endpoint traffic every 15 minutes and stops every campaign by itself if it exceeds a threshold. That last one is the gap that mattered: we did not detect this, you told us.

The dry runs, and what follows them

With your go-ahead we are running one bounded test before any campaign restarts. It is 2 workers on quakescope_2026_global:3, which is the merged build 833743e, against a queue of 8 shards and 87 station-days written for this purpose — only 4F and 3J, two of the temporary codes in your year-less table. Temporary codes are the year-scoped path that produced the requests you flagged, so that is the path worth testing first, and three of the eight shards straddle a year boundary and so need two scoped credentials rather than one.

Every campaign target is 0 and we will tell you before any of them goes above zero.

What happened

Your report: our workers were denial-of-servicing the credentials endpoint, 400/403/404 answers were being retried, and some requests asked for a temporary network without a year. All three are correct. Mechanism and logs for each are below.

Everything is in one file, sb_catalog/src/s3_helper.py, at c5846f3, the build that was running. CompositeS3ObjectHelper.get_filesystem() is called once per day per network by the listing loop and once per object by the read path: 11,315 calls for one network in one year-long, 30-station shard. Nothing remembered a refusal, so nearly every call became a request.

The account block increased the load. A rejected refresh token surfaces from earthscope_sdk as InvalidRefreshTokenError, which has no .response attribute, so it passed our HTTP status handling into a generic retry loop that re-ran the refresh grant five times per scope. After the block, every network took that path, on 208 workers.

What your endpoint saw

CloudWatch log group /aws/batch/job, us-east-2. Each count is one POST to login.earthscope.org/oauth/token. In this window all returned invalid_grant · user is blocked.

2,500 / min 5,000 / min peak 5,104 terminated 16:35 → 15:30 16:00 16:30 16:40 UTC
Token-endpoint requests per minute, 2026-09-04. Step at 15:54: a fleet top-up. Drop at 16:35: termination of all 208 workers. 351,735 requests, 12:21–16:36 UTC; sustained above 100/min from 13:56, mean 2,193/min over those 160 minutes, peak 5,104/min at 15:55. All rejected. Counted by parsing the HTTP Request: POST https://login.earthscope.org/oauth/token lines in the exported events, one per request.

D2 in one worker's log, spanning 88 ms. The 403 triggers the year-less retry; your 400 body states the rule:

/aws/batch/job · quakescope_2026_obs/default/e48ec7b6…2026-09-04 14:59:50 UTC
14:59:50.125 INFO     HTTP Request: GET https://api.earthscope.org/beta/user/
                      credentials/aws/s3-miniseed-v2?network=FDSN%3A7D&year=2025
                      "HTTP/1.1 403 Forbidden"
14:59:50.126 WARNING  Scope {'network': 'FDSN:7D', 'year': 2025} refused:
                      EarthScope returned 403 ... on role s3-miniseed-v2
14:59:50.126 WARNING  EarthScope: 7D denied under 'network+year' scoping;
                      retrying as 'network'.
14:59:50.157 INFO     HTTP Request: GET https://api.earthscope.org/beta/user/
                      credentials/aws/s3-miniseed-v2?network=FDSN%3A7D
                      "HTTP/1.1 400 Bad Request"
14:59:50.158 WARNING  Scope {'network': 'FDSN:7D'} refused: HTTP 400
                      {"detail":"Temporary networks require a year. See
                       https://docs.fdsn.org/projects/source-identifiers/en/
                       latest/network-codes.html#transitional-mapping."}

7D is a temporary code, so the second request could only return 400. The client sent it because it read a 403 as a scoping error. 3,925 of our requests carried a temporary code and no year — 24.4% of the 16,070 refusals we were given, and every one of them could only be answered 400. Two codes account for all of them:

Year-less requests by temporary FDSN code · 2026-09-04, whole incident
CodeRequests
7D3,442
2F483
every other temporary code0

These figures were corrected on 2026-09-05

An earlier version of this page put the year-less total at 20,024 and spread it over eight codes, and the token storm at 260,322 requests over 66 minutes. Both came from filter_log_events substring counts, which count every log line that mentions a scope — the request, the warning that follows it, and the retry — rather than one line per request. That inflates some numbers and, because the storm was measured over too narrow a window, understates others.

Every figure on this page is now derived by parsing the HTTP Request: lines themselves, one per outbound request, from an export of all 10,797,274 events in the window. On that basis the year-less total is 3,925 rather than 20,024, only 7D and 2F ever sent such a request, and the token storm is 351,735 requests over four hours rather than 260,322 over one. The peak, 5,104/min at 15:55, is unchanged. We would rather you had the smaller number and the larger one both right than a tidy story.

The replay measurement

Both builds replayed against the same scripted responses: one shard of calls (365 days × 30 stations) at a single network. Only the client differs. The harness counts calls that would have left the process. No traffic reached your servers.

Requests that would reach EarthScope · one shard, one network
Scenario c5846f3 713698e Change
403 — no access to this network11,3161÷11,316
404 — no such network-year11,3151÷11,315
400 — scope missing its year11,3162÷5,658
Rejected refresh token, six networks10,9801÷10,980
Year-less requests for temporary codes2,4000eliminated
Retry sleeps (5 s each, no jitter)10,980015.2 h → 0

The last row is wall-clock. The old client spent 15.2 hours per shard per worker sleeping between retries, which is why it could saturate the endpoint without finishing a shard.

Against the live endpoint with our blocked account, the fixed client made 2 requests for 200 calls. The old client sends about 1,000.

The five defects

Each block links to the exact lines. Red: c5846f3, the build that was running. Green: 713698e on PR #32.

D1

A refusal was never remembered

Successful credentials were cached by scope. Refusals were not, so each of the 11,315 calls per shard re-ran the exchange.

403, one shard 11,316→1 20 shards, one worker 7,320→1
try:
    self.credentials[key] = self.get_es_credential(net, year)
except RuntimeError as exc:
    logger.warning(f"Scope {scope} refused: {exc}")
    if not self.escalate_scope_mode(net):
        raise
    return self.get_es_filesystem(net, year)
self.set_es_filesystem(key)
verdict = self.es_verdict(net, year)
if verdict is not None:
    raise verdict          # no request is sent

try:
    self.credentials[key] = self.get_es_credential(net, year)
except ES_TERMINAL as exc:
    self.es_record_verdict(net, year, exc)
    ...

The store is at module scope, not on the helper: a worker claims shard after shard in one process and built a fresh helper each time, discarding per-instance memory every few minutes. It keeps the exception class and message rather than the instance, so no worker pins the frames the exception passed through.

s3_helper.py#L172-L178process-wide store
_ES_STATE = {
    "refused": {},        # scope key -> terminal verdict, never re-asked
    "auth_failed": None,  # 401: not scoped to a network, so it stops everything
    "exchanges": {},      # scope key -> [monotonic stamps], for the rate limit
    "denials": {},        # scope key -> consecutive AccessDenied count
    "client": None,       # the one earthscope_sdk client this process uses
}

Also es_verdict and es_record_verdict and the terminal-type list ES_TERMINAL. Every 4xx has a type in that tuple, including EarthScopeRequestRefused for statuses not yet seen, so none can slip past the store.

D2

Escalation removed scope instead of adding it

Source of the year-less temporary-network requests. On any denial the client flipped a network's scoping to the alternative. For a temporary code, the alternative drops the year.

Year-less temporary requests, replayed shard 2,400→0 observed in production 3,925 in production
tried.add(current)
other = "network" if current == "network+year" else "network+year"
if other in tried:
    return False
self.es_scope_mode[net] = other
tried.add(current)
if current != "network":
    return False          # nothing safe left to ask
if "network+year" in tried:
    return False
self.es_scope_mode[net] = "network+year"

It was reachable from a 403: the class raised for one subclassed RuntimeError, so the handler in D1 caught it. The read path called it once per object.

Escalation now only adds the year, only on a 400, at most once per network. A request that cannot succeed is answered locally:

s3_helper.py#L766-L772never leaves the process
if self.es_scope_mode.get(net) == "network+year" and "year" not in scope:
    exc = EarthScopeScopeIncomplete(
        f"Refusing to ask for {scope} without a year: {net} is a "
        f"temporary FDSN code, those are reused across experiments, "
        f"and EarthScope authorises them per year. A year-less request "
        f"can only ever return 400, so it is not sent."
    )
    self.es_record_verdict(net, year, exc)
    raise exc
D3

A new SDK client on every call

The workflow was not reusing credentials at the library level: a fresh EarthScopeClient was built inside the exchange function.

for attempt in range(1, ES_CREDENTIAL_ATTEMPTS + 1):
    try:
        with EarthScopeClient() as client:
            return client.user.get_aws_credentials(
                role=EARTHSCOPE_ROLE, ...)
for attempt in range(1, ES_CREDENTIAL_ATTEMPTS + 1):
    try:
        return self.es_client().user.get_aws_credentials(
            role=EARTHSCOPE_ROLE,
            ttl_threshold=self.ttl_threshold,
            **scope,)

In earthscope-sdk 1.8.0, _UserService caches issued credentials in _aws_creds_by_key, an instance attribute; the on-disk cache present in 1.3.x was removed. A client per call discarded that cache and forced a round trip, plus an OAuth refresh when the access token had not been persisted. es_client() now returns one client per process.

D4

Auth failures fell into the generic retry loop

Produced the traffic in the chart above. 401 and 403 were special-cased; everything else read exc.response.status_code. An AuthFlowError has no .response, so it reached the retry loop and slept.

Rejected token, six networks 10,980→1 observed 351,735 POSTs
s3_helper.py#L538-L572c5846f3 · what a blocked account reached
resp = getattr(exc, "response", None)   # None for InvalidRefreshTokenError
detail = f"{type(exc).__name__}"
if resp is not None:                    # ... so every 4xx check is skipped
    ...
logger.warning(f"EarthScope credential request failed for {scope} "
               f"({attempt}/{ES_CREDENTIAL_ATTEMPTS}): {detail}. "
               f"Sleeping 5 seconds.")
time.sleep(5)                           # x5, each attempt re-runs the refresh grant
/aws/batch/job · quakescope_2026_global2026-09-04 16:32:58 UTC
HTTP Request: POST https://login.earthscope.org/oauth/token "HTTP/1.1 403 Forbidden"
error during token refresh (1 attempts): {"error":"invalid_grant",
                                          "error_description":"user is blocked"}
EarthScope credential request failed for {'network': 'FDSN:4F', 'year': 2019}
    (4/5): InvalidRefreshTokenError. Sleeping 5 seconds.
if isinstance(exc, AuthFlowError):
    # Every remaining AuthFlowError is about OUR credentials, not about
    # this scope: InvalidRefreshTokenError, NoRefreshTokenError, the
    # device-code errors. None is congestion, and none is fixed by asking
    # again - but they carry no `.response`, so the status-code branch
    # below never saw them and they fell into the retry loop.
    raise EarthScopeAuthFailed(...)   # terminal, and process-wide

A rejected token is not a fact about a network, so EarthScopeAuthFailed is recorded in _ES_STATE["auth_failed"] rather than per scope and es_verdict() checks it first: one rejection stops every network in the process. The remaining retry loop handles 429 and 5xx only, with exponential backoff, full jitter and Retry-After. The previous fixed 5-second sleep returned every affected worker at the same instant.

D5

A shadowed handler spun a tight loop on the access point

PermissionError and FileNotFoundError both subclass OSError, and an OSError handler sat before both, so neither could run. Present since June.

except OSError as e:
    if e.errno == 5:
        return obspy.Stream()
    # AccessDenied has errno None -> falls through,
    # returns nothing, `while True` goes round again
except PermissionError as e:        # unreachable
    ...
except (FileNotFoundError, ...):   # unreachable
except PermissionError as e:
    denied += 1
    scope_denied = self.s3helper.note_access_denied(net, year)
    ...
except FileNotFoundError:
    return obspy.Stream()
except OSError as e:
    if e.errno == 5: ...

s3fs maps AccessDenied to PermissionError with a message and no errno. A denied object fell out of the OSError branch without returning, the while True loop repeated, and the object was re-HEADed at socket speed until the 900-second station-day timeout.

Across the full five-day CloudWatch retention, the log lines that only the PermissionError branch can emit do not appear:

filter_log_events · /aws/batch/job · full 5-day retentiondead code, confirmed
       0  "Credential refreshed after access denied"   ← PermissionError branch
       0  "Access denied ... times for"                 ← same branch, budget spent
       0  "Not authorized to access this resource"      ← OSError errno==5 branch
      18  "Timeout after"                               ← the 900 s station-day cap

The ES_DENIED_ATTEMPTS budget that bounds credential refreshes therefore never executed. During the blocked window workers failed earlier, at the credential stage, so this defect contributed little to the traffic in the chart. The 18 station-day timeouts are its observable symptom.

A test parses the handler chain and asserts each narrow type precedes OSError.

Smaller changes in the same pass

  • Per-scope denial accounting. The refresh budget lived in a local that reset on every object. note_access_denied counts per scope, resetting only on a successful read.
  • A local rate limit. _es_throttle refuses more than three exchanges of the same scope in five minutes and sends nothing. A credential is valid for an hour.
  • Paced diagnostics. netyear_sweep probes about 1,600 planned network-years and now takes a --pace, default 0.25 s.
  • Quieter logs. A cached verdict is logged once per network rather than once per day.

Which credentials are which

Nothing in our AWS account grants access to your archive. s3-miniseed-v2 is an EarthScope role alias, not an IAM role in our account, and the credentials that read data are issued by you. That determines what you need to run the test: nothing from us.

ISSUED BY EARTHSCOPE refresh token ES_OAUTH2__ REFRESH_TOKEN login.earthscope .org /oauth/token api.earthscope.org /beta/user/ credentials/aws/ s3-miniseed-v2 temporary AWS credentials ~1 h, scoped S3 access point earthscope- mseed-v2-… POST 200 sigv4 access token ?network=FDSN:XX&year=YYYY 400 / 403 / 404 are answered here OUR AWS ACCOUNT SeisBenchBatchRole Secrets Manager one secret ARN supplies the token, and nothing else
Every credential that reads archive data is issued by EarthScope, scoped by network and year. Our AWS account reads the refresh token from Secrets Manager into the container environment. A tester with their own EarthScope login reproduces the upper lane and needs no access to the lower one.

The whole of the AWS-side configuration, none of it needed to run the test. This page is public, so account and secret identifiers are described rather than printed. Ask and we will send the full ARNs directly.

Job & execution role
one IAM role, shared by the Batch job and its ECS execution
Attached
AmazonECSTaskExecutionRolePolicy, AmazonS3FullAccess, inline QuakeScopeEarthScopeSecretRead
Inline policy
secretsmanager:GetSecretValue on exactly one secret — the EarthScope refresh token, by full ARN
Region
us-east-2 (same as the access point; reads are not cross-region)

AmazonS3FullAccess on the job role is broader than needed. It grants nothing against your archive, since those reads use the credentials you issue. We are scoping it to our own buckets in a separate change.

Reproduce it yourself

scripts/earthscope_selftest.py replays a full shard of calls and counts what would leave the process. Offline mode sends nothing and needs no EarthScope account. Live mode uses your account and enforces a hard request cap.

  1. Get the code and its dependencies

    Python 3.10+. Verified from a clean virtualenv: no AWS credentials, no access to our account.

    setup~3 min
    git clone https://github.com/SeisSCOPED/QuakeScope.git
    cd QuakeScope
    
    python -m venv .venv && source .venv/bin/activate
    pip install "earthscope-sdk>=1.8.0" s3fs obspy numpy
  2. Run the offline checks — no account, no traffic

    Six assertions, each replaying 11,315 get_filesystem() calls against a scripted response. Exit code 0 only if all pass.

    offlinesends nothing
    python scripts/earthscope_selftest.py --offline
    expected outputabridged
    CURRENT - this working tree
    --------------------------------------------------------------
      PASS  403 is asked once                   requests sent: 1   expected: 1
      PASS  404 is asked once                   requests sent: 1   expected: 1
      PASS  400 is corrected once, by adding the year
                                                requests sent: 2   expected: 2
      PASS  a blocked account stops every network at once
                                                requests sent: 1   expected: 1
      PASS  no yearless request for a temporary network
                                                requests sent: 0   expected: 0
      PASS  no retry sleeps at all              requests sent: 0   expected: 0
    
    ALL CHECKS PASSED
  3. Run the same checks against the build that caused this

    Pass --baseline a copy of the old file and the script runs both side by side. It counts the retry sleeps instead of serving them, so the run finishes in seconds rather than days.

    A/Bstill sends nothing
    git show c5846f3:sb_catalog/src/s3_helper.py > /tmp/old_s3_helper.py
    python scripts/earthscope_selftest.py --offline --baseline /tmp/old_s3_helper.py
    expected outputbaseline · failures are the bug
    BASELINE - the version deployed during the incident
    --------------------------------------------------------------
      FAIL  403 is asked once                   requests sent: 11316  expected: 1
      FAIL  404 is asked once                   requests sent: 11315  expected: 1
      FAIL  400 is corrected once               requests sent: 11316  expected: 2
            [no year -> year=2019 -> year=2019 -> year=2019]
      FAIL  a blocked account stops everything  requests sent: 10980  expected: 1
      FAIL  no yearless request for a temporary network
                                                requests sent:  2400  expected: 0
            [2406 requests total, 2400 of them yearless]
      FAIL  no retry sleeps at all              requests sent: 10980  expected: 0
            [10,980 sleeps totalling 54,900 s = 15.2 h per shard, per worker]
  4. Sign in with your own EarthScope account

    Device-code flow through the SDK. No extra package, no password typed into our script. It prints a URL and a code; you approve it in a browser and the SDK writes tokens to ~/.earthscope/default/tokens.json.

    loginyour account, not ours
    python scripts/earthscope_selftest.py --login

    The client also reads ES_OAUTH2__REFRESH_TOKEN from the environment, which is how our Batch jobs are configured:

    alternativehow our workers are configured
    export ES_OAUTH2__REFRESH_TOKEN=<your own refresh token>
  5. Run the live checks against your own access

    Pick three networks: one you can read, one you cannot (403), and one FDSN code that does not exist (404). The script drives 200 calls at each and prints every outbound request with its query string. It aborts if it exceeds --max-requests.

    livecapped at 12 requests
    python scripts/earthscope_selftest.py --live \
        --allowed-network AV --allowed-year 2019 \
        --denied-network  LH \
        --missing-network 5A --missing-year 2018

    Watch the request count per case and the yearless temporary-network requests total. A run against our blocked account, which is why it stops at the token exchange:

    actual runblocked account · 200 calls
      a network you CANNOT read (403): LH
        -> GET https://api.earthscope.org/beta/user/credentials/aws/
               s3-miniseed-v2?network=FDSN%3ALH
        -> POST https://login.earthscope.org/oauth/token
        2 request(s) for 200 calls  (EarthScopeAuthFailed)
    
      total requests sent this run: 2 (cap 4)
      yearless temporary-network requests: 0  (must be 0)
    
    ALL CHECKS PASSED

    The old build sends about 1,000 POSTs for the same 200 calls.

If you would rather not install anything

The container is public and pullable anonymously, and now carries the self-test itself:

docker run --rm ghcr.io/seisscoped/quakescope:<tag> selftest --offline

Live mode takes the same flags, with -e ES_OAUTH2__REFRESH_TOKEN for your own token. Use a tag from 833743e onwards; earlier images do not contain the script.

An earlier version of this page gave this as docker run … python scripts/earthscope_selftest.py, which could not have worked: the image is built from sb_catalog/ so scripts/ was never in it, and the entrypoint is python -m src.picker, which would have swallowed those arguments. The build context now takes in the whole repository and the picker gained a selftest subcommand.

What we did next

Nothing here is a request. The defects were in our client, not in your API, and the fix has been merged and exercised against your endpoint under the bounded conditions described above.

Two dry runs. The first, 2 workers over 8 shards of 4F and 3J, sent 8 credential requests across 5 scopes and 0 year-less requests, and cached every refusal. It also found a defect of our own making: our fix reused one SDK client per process, and the SDK caches its synchronous runner on first use, so a credential exchange from inside the read loop raised before any request was sent. Three shards that crossed a year boundary were lost to it. That is fixed, and the same eight shards then completed 8/8, issuing the second-year credentials that had never once left the process.

The second read real data: 6 shards, 1,500 station-days of 4F in 2015 and 2017, 1,240,055 picks written. It spent 1 h 45 min and sent 15 credential requests over 2 scopes, four of them mid-read renewals when the one-hour credential expired — the case that would have failed before the runner fix, and the one that covers most of a real campaign.

The three campaign job definitions still name c5846f3 and the fleet workflow refuses to launch it. We will not raise a campaign target without telling you first.