Shaduf.
AI Model Degradation Watch/Incidents resolved; rollout and route changes still make quality hard to generalize
Research report — AI Model Degradation Watch
Research report · run 9 September 2026

The outage story expired. The route story did not.

The latest public evidence does not support “popular models are broadly getting worse.” It supports a more precise and more useful call: specific surfaces have been unreliable; providers are changing what users receive; and community pain is real but not yet a controlled capability verdict.

Question: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Verdict: different, recently reconfigured, and intermittently unreliable on specific surfaces; not proven broadly worse. Confidence is high for resolved declared incidents and documented changes, medium for narrow user-facing signals, and low for a general capability decline.

What the sources establish

  1. Declared outages resolved. OpenAI marked its 8 September image and file incidents recovered; Claude and Grok status pages were operational at the latest check. The 3 September cluster is a historical reliability event, not a current model test.
  2. The comparison target moved. OpenAI's gradual Astra/GPT-5.6 access, Anthropic Fable 5.1, and Google's Gemini 3.8 release mean a day-over-day output difference can be a model, route, plan, effort, safety, or product change.
  3. One controlled surface is nominal. Marginlab's direct Codex CLI tracker reports 88% today and 85% over seven and thirty days versus an 83.40% frozen baseline, below its significance thresholds. That is meaningful evidence against a current regression on that scoped coding surface, not a universal all-model result.
  4. Cross-sectional comparisons show competition. Artificial Analysis, Arena, and an independent Terminal-Bench analysis show current models near one another or tied in their stated settings. These snapshots cannot see a hidden route changing over time.
  5. Users still report costly failures. Current Claude, Grok, and ChatGPT discussions describe refusals, incomplete work, context loss, correction burden, gibberish, and plan pressure. Positive or normal counterreports and missing controls keep the broad claim unresolved.

What would change the call?

A repeatable before-and-after result on the same model identity and product surface, or a provider disclosure connecting a change to a measurable task movement, would move the verdict toward capability degradation. Repeated reports that match task, plan, route, timing, and output would raise confidence in a surface regression. A status recovery, new rollout, or access explanation lowers the need for a model-wide explanation but does not invalidate the task report.

Reader protocol

1 · Pin it

Record provider, surface, visible model, plan, effort, task, time, and failure cost.

2 · Check it

Open status and release notes. Look for rollout, safety, limit, or access changes.

3 · Replay it

Use a fresh thread or stable local task. Count correction turns, missing context, refusal, latency, and task success.

4 · Decide

Retry or wait for a reversible incident; pin or switch only after matched failures repeat or production evidence moves.

Primary sources

The analytics broker was unavailable for this run. No private Shaduf evaluation or private traffic is presented as evidence.

Search published pools, pages, reports, and evidence.