The watch needs an alert contract.
No broad model collapse is proven. The sharper update is operational: define the user-visible task failure, its baseline, and the action before turning a cause-like signal into a degradation headline.
Calm status narrows the outage lane
OpenAI’s live status says the service is fully operational, and its history shows a run of resolved incidents through 17 September, including GPT-5.6 paid-plan errors, API model errors, and ChatGPT Work errors. Anthropic’s page lists no incident on 18 September after resolving Fable/Mythos errors on 15 September. xAI says no incident is declared.
That is useful for a reader deciding whether to retry or wait. It is not a clean bill of health for every account or task. OpenAI explicitly notes that aggregate availability can vary by tier, model, and API feature; a calm page cannot expose hidden routing, context pressure, allowance state, or a bad artifact.
A historical postmortem shows why “the model got worse” is too cheap
Anthropic’s 17 September 2025 postmortem is the best new causal warning in this review. It says three infrastructure bugs intermittently degraded Claude responses. A context-window routing error was amplified by a load-balancing change; separate deployment and TPU compiler issues caused output corruption and wrong token-selection behavior. Anthropic also says its evaluations were too noisy and did not connect the spike in online reports to the infrastructure change quickly enough.
This is historical provider evidence, not a 2026 prevalence estimate. Its value is the distinction: a real user-facing quality decline can be produced by the serving system without proving that the underlying weights changed. The audience’s “it feels worse” observation can therefore be both real and misattributed.
The method upgrade: symptom, outcome, action
NIST’s TEVV-Athlon draft frames evaluation around organizational objectives, Events, Tools, and Blocks that produce evidence about measurement concepts and real-world outcomes. Google SRE’s incident guide adds the operating rule: alert on end-to-end, user-facing symptoms rather than internal causes, and make each alert actionable.
Together they force a better watch. “The route changed” is a lead. “This fixed task fell below its 80% completion threshold on the same surface, and the operator should pin, retry, or switch” is an alert. Without the second sentence, the pool is collecting interesting causes, not monitoring a decision.
The audience is already asking for the receipt
A Claude community hub updated 18 September asks contributors to include prompts and responses, platform, time, and screenshots. It separates performance from usage-limit reports and links a problem log that compares category volume with a recent four-week average. That is a useful product signal: people understand that a complaint becomes more valuable when the task and conditions are visible.
It is not a denominator. The hub is volunteer-run and self-selected, the log has unknown coverage, and neither exposes provider-served model, route, account ledger, or completed-work rate.
The live degradation ledger remains mixed
The recent 73% five-hour-window report still supports a real allowance and usable-access lead, not a weight verdict. The earlier Astra capacity-conditioned output report still supports a narrow user-facing quality lead, not a served-model or causal verdict. Windows surface failures, status incidents, safety disclosures, cross-sectional benchmarks, and community complaints remain separate lanes. A quiet status page cannot erase them; none can be stacked into a broad model decline without a matched task series.
The decision receipt
| Record | Minimum fields | Action |
|---|---|---|
| Symptom | Fixed task, surface, baseline, threshold, date, and completed outcome. | Decide whether an alert should fire. |
| Serving path | Requested and provider-served model, fallback, plan, route, region, capacity state. | Separate model from routing, entitlement, and rollout. |
| Execution | Context, instructions, effort, tools, retries, completion boundary, and artifact. | Separate model behavior from orchestration and surface failure. |
| Economics | Corrections, time, cache, allowance movement, cost, meter, and ledger status. | Separate work-window loss from answer quality. |
| Next move | Retry, wait, pin, switch, replay, or no action; state what would stop the alert. | Make the watch useful instead of merely dramatic. |
What would change the call?
- Run one fixed task across affected and control accounts or providers with the same surface, context, tools, and rubric.
- Capture provider-served model, route, capacity state, and allowance movement rather than relying on the requested label.
- Score the completed artifact, correction burden, and recovered work separately from tokens, cache, time, and errors.
- Join the task receipt to a provider ledger where possible; otherwise label the gap.
- Call broad capability decline only after repeated negative movement survives outage, routing, product, safety, access, context, and tool explanations.