Shaduf.Research preview
AI Model Degradation Watch/OpenAI reports a Work Mode incident; Opus 5.5 regression remains unproven
OpenAI reports a Work Mode incident; Opus 5.5 regression remains unproven | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 5 October 2026 · Public-source review

No broad model decline is established. Check the affected product first.

OpenAI reports elevated Work Mode errors and says it is monitoring recovery after mitigation. Anthropic documents cases where safety checks can change which Claude model answers. Neither is evidence that the underlying model broadly lost capability.

Current call: broad degradation remains unproven. OpenAI’s 5 October incident is scoped to ChatGPT Work and Codex in ChatGPT Desktop. For Claude Opus 5 or 5.5, check the per-response model label and any fallback notice before comparing answers.

A confirmed service problem, not a model test

OpenAI’s status record says elevated Work Mode errors began at 07:03 UTC. The 07:14 update said errors continued and scheduled tasks might be affected. At 09:50 UTC, OpenAI said the issue was mitigated and recovery was being monitored. The listed affected components are ChatGPT Work and Codex in ChatGPT Desktop. This is evidence of a service incident on those surfaces; it does not measure answer quality or identify a cause.

A selected model may not be the model that answered

Anthropic’s help page says Opus 5 and 5.5 safety checks inspect each request and the context the model reads. Some flagged requests can fall back to an earlier model. Claude says it shows a switch notice and labels the responding model; after an automatic switch, the model selector remains on the fallback for the rest of the conversation. The documented categories are narrow, including some cybersecurity work, biology requests for Opus 5.5, and a small set of frontier-model-development tasks. The API configures fallback differently and does not switch automatically by default. Record the responding model, not only the model you initially selected.

Opus 5.5 monitoring still does not establish a regression

BridgeBench · Oct 4

96.5% of launch power

The score is up from 94.2% on 2 October and within BridgeBench’s stated 90–110% band. The board shows five tests. “Power” combines task performance, tokens, and cost; tasks and weights are private.

LiveNerf · Oct 5

Two follow-up days

Its ten-day baseline is complete and 12 of 30 daily collections are reported. There is no post-baseline comparison yet. The first possible decision remains around 24 October.

Marginlab · last updated Sep 29

Detection remains paused

The page says its new Opus 5.5/high baseline is still collecting. Its displayed daily and rolling rates do not form a published before-and-after result.

LiveNerf estimates a realized minimum detectable shift of 6.6 percentage points per ten-day window, versus 7.5 predicted. Its test is a fixed question panel through one headless Claude Code subscription path. It does not cover every Claude product or everyday workflow.

One user kept a task receipt

A 2 October ClaudeAI post says the author reran a Godot game prompt with the same reference images in a fresh workspace after an earlier 23 September run. The author lists rendering and feature differences. The comparison is worth trying to reproduce, but the project files, run logs, and exact responding model for both runs were not independently checked. It is one task report, not evidence of a broad or confirmed regression.

If your task fails

  1. Check the provider’s status for the exact product surface and time.
  2. Save the prompt, files, settings, context, output artifact, and corrections.
  3. For Claude, save the fallback notice and model label shown on the response.
  4. After recovery, replay the same task with the same tools and settings against a saved baseline. Compare the finished work and repair effort.

Sources

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.