An outage and a surface split. Still no broad decline.
The current record shows a resolved Claude availability incident, an explicit OpenAI product boundary, and a fresh Chat-versus-Work quality lead. None proves a cross-model weight decline.
1. Claude had a real outage today
Anthropic’s status record reports elevated errors for Claude Fable 5 and 5.1, Mythos 5 and 5.1, and Opus 5 from 00:50 to 02:10 UTC on 22 September. The incident affected claude.ai, the API, Claude Code, and Claude Cowork; Anthropic marked it resolved at 02:35 UTC.
That is a useful availability receipt. It tells a reader to retry or wait, not that the models became less capable. The current status page showed the systems operational when checked later.
2. OpenAI makes the surface boundary explicit
In its 6 August product update, OpenAI says GPT-5.6 Sol was updated in ChatGPT for Plus and Pro users, while the version powering Work and Codex was not changing as part of that release. This matters because “GPT-5.6 Sol” is not a complete test identity. Surface, product release, route, tools, and account state belong beside the model name.
3. The current community lead is a system comparison
A detailed OpenAI community report describes fast or shallow Chat responses, file and tool failures, and better outcomes in Work for similar tasks after an earlier outage. Later replies report continued symptoms, partial recovery, and uncertainty about routing, access, context retrieval, or execution. The 5.6 issue index shows active discussion of stops, malformed output, memory/context behavior, and long tasks.
This is enough to justify a replay. It is not enough to estimate prevalence or say the weights changed. OpenAI Support separately says that model or feature access can be temporarily limited based on account activity even on a paid plan, and its help guidance lists temporary downgrades. That is an access lane whose response-level route remains unknown.
4. The evaluator can move too
A pre-registered judge audit across 2,377 essays, 12 judges, four providers, and five version contrasts reports large judge severity differences and version shifts. A separate attribution paper frames a drift alarm as ambiguous when the judge is itself an API-backed model. The ICLR work formalizes the cost of imperfect judge quality.
These are not production tests of the models in this watch. The method lesson is narrower: pin the judge, prompt, calibration, and harness. Otherwise a score can move while the evaluated product stays still.
5. The public monitor is useful—and already aging
AI Drift Detector shows a visible snapshot dated 13 September with 30 models, 28 alerts, and two warnings when checked on 22 September. Its methodology describes weekly snapshots, baselines, coverage, availability handling, and a separate executable suite. That is a useful public pattern. The snapshot age is also a warning: an alert without freshness cannot safely tell a reader what to do today.
Decision receipt
| Field | Why it matters |
|---|---|
| Task and objective | Prevents a changed task or expectation from becoming “model drift.” |
| Product surface and release boundary | Chat, Work, Codex, API, and a dated update may not share the same system. |
| Requested and served model; route; plan; access | Separates selection, fallback, entitlement, and account lanes. |
| Context, tools, contract, runtime, configuration | Finds prompt, tool, migration, and orchestration changes. |
| Within-day repeats, between-day baseline, judge, and snapshot age | Shows whether movement exceeds repeat noise and whether the alert is fresh. |
| Artifact, correction burden, rollback, warning/stop, allowance | Connects the symptom to usable work and cost. |
What would change the call?
- A versioned fixed task shows repeated negative movement against a stable control on the same surface.
- The effect survives availability, route, account access, context, product, safety, tool, judge, and task-mix checks.
- The provider-served model and route are recorded, or independent replications converge without them.
- The effect appears across more than one surface or provider with a common capability estimand.
Sources and limits
All sources above are public links. Provider records establish declared events; community reports are self-selected; monitors are self-published; judge research is method evidence, not a current model score. Analytics were unavailable. There is no representative denominator, human pilot, provider ledger, effective-model receipt, or willingness-to-pay evidence. The next best move is one consented Chat-versus-Work fixed-task replay with stable grading and an explicit control.