The system changed; “worse” still needs a controlled receipt.
No broad model collapse is proven. The fresh evidence says the surrounding work system is moving in visible ways: provider incidents, a tool-contract migration, competing drift estimands, and quota-driven switching.
OpenAI’s recent history is a service and economic cluster
The OpenAI status history checked today lists a resolved paid-plan GPT-5.6 and GPT-5.6 Instant conversation incident on 15 September, elevated errors across API models on 17 September, managed Agents API delays or session-start failures on 14 September, and higher-than-expected charges for OpenAI-hosted Agent API containers from 18–19 September with refunds under review.
That cluster matters because it gives the reader several concrete alternatives to “the model got worse”: model-specific availability, API service error, managed-session access, and billing or guardrail failure. The events are real provider-declared incidents. They do not measure a fixed task’s completed artifact, and they do not prove a weight change. OpenAI’s aggregate current status can be fully operational while a historical or account-specific path was still costly or unavailable.
Google changed the Antigravity tool contract
Google’s 17 September Gemini API notes say Antigravity 09-2026 replaces 05-2026. For local tools, the contract changes parameter casing and file edit, read, list, and search operations; the older agent version is scheduled to shut down on 5 October. An integration that assumes the old names or shapes can fail, loop, or produce a worse artifact even if the underlying model weights are unchanged.
This is not a reason to dismiss a bad agent session. It is a reason to capture the tool-contract version, actual tool calls, errors, and final artifact before assigning the cause to capability. Remote users who only read output may have a different exposure than local integrations that parse function calls.
The counterweights disagree because they measure different work
Anthropic’s 18 September internal retrospective reports Claude Code session success rising across four task types to 88–92% by September, with open-ended success moving from about 26% to about 91% since March. It is useful counterweight, not neutral proof: Anthropic selects the sample and workload, a Claude judge scores success, and the page notes that work mix can move the series.
AI Drift Detector reports 28 alerts and 2 warnings across 30 models in its 13 September board snapshot. Its methodology describes a baseline-relative weekly monitor plus a fixed 50-task executable software-engineering suite with hidden tests, retries, and coverage. That is a meaningful repeat signal, but a one-week-old all-model alert pattern can reflect harness, calibration, release, or task effects.
Pondral reports a maximum engine-level change of 0.9 percentage points between its 1 September and 2 August AI-search visibility runs. This is a stability counterweight on a different estimand. The correct conclusion is not “one monitor is wrong”; it is “coding drift, search visibility, provider-internal success, and user workflow quality must not be collapsed into one score.”
Users are switching before they can name the cause
Current Codex quota discussions describe reduced or nonlinear allowance value and fast depletion. A reset thread includes users comparing Claude and Gemini, including a report that Gemini 3.8 Flash needed less back-and-forth for some tasks; another limits thread asks for transparent credits. A separate “done” thread describes moving providers because of quota and task-fit frustration.
These threads show a real audience decision: users buy a usable work window, not a model label. They still do not supply a fixed task, effective-model receipt, plan-aligned denominator, provider ledger, or controlled task outcome. Record the switching reason, the replacement, the completed artifact, and the allowance movement before treating a switch as evidence of capability decline.
The market already teaches symptom-first checks
NextReset and Codex Pulse both separate latency, limits, early stopping, tools, context, runtime, and task outcome, and recommend fixed-task replays with exact client and model details. This is competitive evidence of a useful vocabulary, not proof of adoption. The public pool’s edge should be the explanation: which lane the evidence touches, what it cannot establish, and which reversible choice is now justified.
The decision receipt, revised
| Record | Minimum fields | Decision |
|---|---|---|
| Task and outcome | Fixed task, baseline, threshold, completed artifact, self-check, correction burden. | Was the work actually worse? |
| Serving and access | Requested and served model, route, plan, capacity, entitlement, managed-surface error. | Did the task run under the intended system? |
| Contract and execution | Tool-contract/runtime version, tool lifecycle, context, skills, descendant agents, response count, task budget, warning and stop state. | Did a migration or hidden work explain the symptom? |
| Benchmark context | Suite, judge, coverage, baseline, release date, and monitor estimand. | Can this public alert be compared with another? |
| Economics and action | Allowance, cost, provider-ledger state, switching reason, replacement outcome, retry/stop/pin/switch rule. | What should the operator do next? |
What would change the call?
- Replay one fixed task across affected and control accounts or providers with the same context, tools, contract, budget, and rubric.
- Capture provider-served model, route, entitlement, capacity, warning, stop state, and allowance movement.
- Score the artifact and correction burden separately from tokens, cache, time, and provider incident state.
- Align any public monitor by task, harness, judge, coverage, baseline, and release window before comparing alerts.
- Call broad capability decline only after repeated negative movement survives service, routing, product, safety, access, context, and tool explanations.