Shaduf.Research preview
AI Model Degradation Watch/Opus 5.5 drew “Claude is back” reports. That is not a recovery test.
Opus 5.5 drew “Claude is back” reports. That is not a recovery test. | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 28 September 2026 · Public-source review

Opus 5.5 drew “Claude is back” reports. That is not a recovery test.

A new Claude version has positive provider, evaluator, and user signals. They say something about Opus 5.5 and task fit; they do not show that an unchanged model declined and recovered.

Call: no broad cross-model core-capability decline is proven. Anthropic’s launch claims and METR’s preliminary evaluation point toward an improvement for the new version on selected measures. A high-attention Reddit post and one creator’s blind comparison describe better fit for some work. This evidence does not establish a before-and-after recovery, a typical user experience, or what happened on any one account.
Provider release

New version, several changes

Anthropic announced Opus 5.5 on 22 September and reports selected benchmark gains, improved communication, faster output, lower prices and typical task costs, and higher five-hour plan limits.

External evaluation

Modest gain on five tasks

METR’s preliminary assessment says Opus 5.5 likely improves modestly over Fable 5.1 on its selected AI-R&D tasks, not by a large amount.

Different benchmarks

GPT-6 evidence still differs by task

Bug Hunt’s small coding slice is negative for GPT-6 Sol; Braintrust’s separate 175-example API comparison is nearly tied. Neither is a fixed production time series.

What the release evidence says

Anthropic’s release page compares Opus 5.5 with Opus 5 and reports gains on selected benchmarks along with changes to communication, speed, price, typical task cost, and plan limits. These are provider claims. The comparison cannot tell us whether a user’s improvement came from capability, the product surface, access, cost, or a different task mix. A score from a new version is not a record of an older version getting worse.

METR’s 22 September public summary describes an unpaid, preliminary predeployment evaluation using API access over ten business days and five AI-R&D tasks. METR judges that Opus 5.5 likely improved modestly over Fable 5.1 on this evaluation and was not a large jump. The summary is narrow, not a production replay; Anthropic had an opportunity to review and edit the summary. It does not estimate ordinary chat quality or every use case.

What users noticed

The 22 September r/ClaudeAI thread describes Opus 5.5 as more collaborative after frustration with earlier releases. Replies discuss response length, tone, tasks, and plan use. The thread is useful for finding a task to replay, but its votes are not unique customers, switches, or repeat use.

In a How I AI blind comparison, host Claire Vo describes trying multiple models across email, PRDs, frontend and backend work, agents, SVGs, and video. The episode description says her preference differed by task and an LLM judge disagreed with her. This is one practitioner’s sample, not a representative or longitudinal test.

Keep the outage separate. Anthropic’s status page currently reports systems operational and no incidents for 27–28 September. It also records a resolved 80-minute incident on 22 September affecting requests to Opus 5, Fable 5/5.1, and Mythos 5/5.1. The status record does not connect that incident to the positive reports. OpenAI’s status page was also operational when checked.

What this changes in the wider answer

Recent results for GPT-6 Sol remain mixed: Bug Hunt Bench reports a lower mean than GPT-5.6 Sol on a small planted-bug coding test, while Braintrust reports a near-tie across 175 API examples under a different task mix and setup. These are cross-sectional comparisons, not repeated tests of the same deployed system.

The new Claude evidence moves the current picture toward a useful version-transition and task-fit question. It does not change the broad verdict: public evidence reviewed here does not establish that popular models generally lost core capability. A release can improve one task, alter tone, reduce typical cost, speed output, or change access without resolving a user’s separate experience on another surface.

A practical replay

  1. Check the provider status page and wait until a declared incident is resolved.
  2. Record exact version, product surface, plan, route, and requested and served model when visible.
  3. Replay the same task with the same context, tools, effort, and budget; keep the final artifact.
  4. Score outcome, corrections, speed, cost, and usable access separately. Repeat enough to see run variation.
  5. Use a version-change report as a lead. Do not call it recovery unless the same system and task are measured before and after.

Sources

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.