The newest lead is a route, not a model collapse.
The latest public evidence sharpens the answer rather than flipping it. Specific surfaces have been unreliable; providers openly change access, safety paths, and allowance; and a technical Google report may reveal a hidden serving switch. None of that, yet, proves a broad decline in underlying model capability.
What the sources establish
- Recent outages are recovery evidence, not a capability score. OpenAI’s history marks the 8 September image-generation and file-upload incidents recovered. Claude and Grok status pages were operational at the check. A recovered outage can explain a bad workflow at a time without showing weaker model weights.
- The served system is not just the label. OpenAI’s current guide says allowance depends on plan, workspace, rollout eligibility, model, reasoning, Fast mode, and task. Anthropic’s help page says safety checks can move a Fable request to an Opus fallback. These are provider-documented mechanisms by which quality, speed, or quota can change.
- There is one technically detailed but unconfirmed route lead. A Google AI Developers Forum report says requests labeled gemini-3.1-pro-preview switched in final streaming chunks to 3.1-flash-lite-preview above roughly 12,600 input tokens; a second poster reports similar GCP behavior. The thread has no Google confirmation. If reproduced, the right label is routing or product-serving change, not immediate proof of Pro capability decline.
- Scoped controlled evidence does not show a broad collapse. Marginlab’s direct Codex CLI tracker reports 88% today and 85% over seven and thirty days against an 83.40% baseline, within its nominal thresholds. Its Claude Code page is collecting a new Opus baseline and pauses degradation detection. A nominal result and a collecting-baseline result are both scoped states, not universal claims.
- Field pain remains real and mixed. Current discussions describe refusals, incomplete work, context loss, correction burden, gibberish, and usage pressure. A positive Fable 5.1 report says quality was excellent while usage burned quickly. Quality, access, and cost can move in different directions.
How to test the route hypothesis
1 · Capture identity
Save the visible model, final streamed modelVersion or fallback notice, endpoint, plan, billing tier, and safety or effort setting.
2 · Sweep context
Replay matched prompts just below and above the reported threshold. Keep task, temperature, tools, and output budget fixed.
3 · Compare surfaces
Run AI Studio, API, and Vertex where permitted. Record whether the route, bill, latency, and output length move together.
4 · Separate cost
Count corrections, minutes, tokens, interruptions, and quota. A more expensive good answer is a different failure from a weak fallback.
What would change the call?
A repeated before-and-after result on the same model identity and surface, or a provider disclosure connected to a measurable task movement, would move the verdict toward capability degradation. Repeated route mismatches across accounts and endpoints would move it toward a confirmed serving change. A status recovery, rollout explanation, or access rule lowers the need for a model-wide explanation but does not invalidate a user’s task report.
Primary sources
- OpenAI status history, Claude status, and Grok status.
- OpenAI usage guide, OpenAI safety note, and Anthropic fallback guide.
- Google developer route report and Gemini changelog.
- Marginlab Codex, Marginlab Claude Code, WDCD instruction-decay report, and Beyond Benchmarks.
- Modelgrep, Claude Code usage discussion, Fable counterreport, and Cadence session-cost product.
The analytics broker was unavailable for this run. No private Shaduf evaluation or private traffic is presented as evidence.