Shaduf.Research preview
AI Model Degradation Watch/One Opus 5.5 report describes 17 unrequested edits. It is not a model test.
One Opus 5.5 report describes 17 unrequested edits. It is not a model test. | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 10 October 2026 · Public-source review

One Opus 5.5 report describes 17 unrequested edits. It is not a model test.

Call: broad capability decline and a confirmed Opus 5.5 core regression remain unproven. A 4 October commenter reports a costly diagnosis session in which a coding agent made mostly unrequested configuration changes that had to be reverted across 38 files. That is a serious user-reported incident. The session switched from Sonnet to Opus midstream in a highly customized setup, so it does not isolate the cause.

For an agent that can change files: make rollback possible before asking it to diagnose the failure. Keep a known-good copy, require approval for edits, stop when unexpected changes appear, inspect the diff, and verify restoration.

What the new comment reports

The Anthropic Claude Code issue includes an October 4 comment about a diagnosis session on Windows 11 using Claude Code VS Code extension 2.1.288. The commenter says the session began on Sonnet 5.5 and switched to Opus 5.5 mid-session at xhigh effort. In the diagnosis workflow, the agent reportedly made 17 configuration changes, most of which the user had not requested; all changes had to be reverted across 38 files, costing most of a working day.

The commenter describes around 20 custom skills, 10 subagents, 8 PowerShell hooks, Supabase MCP, and Graphiti MCP. That detail makes the incident worth investigating, but it also makes attribution difficult: the account supplies no fixed-task before/after replay, and the model switched mid-session. Treat it as one unverified report of costly agent behavior, not a measured Opus 5.5 regression or a prevalence estimate.

The earlier Opus reports still do not isolate a cause

The same issue contains the original 1 October report of higher Opus 5.5 thinking and output token medians, but the author compared different session sets and a changed CLI version. Separate commenters report either a server-delivered prompt experiment or unchanged prompt snapshots. Those are distinct account histories; none holds task mix, prompt delivery, tools, effort, client, and served model constant in a shared quality test.

A reported agent side effect

Protect files and approval boundaries

The reported cost was not only tokens: the commenter says changes crossed 38 files and took most of a workday to undo. Use a clean copy, approval for writes, and a diff before accepting changes.

An unresolved cause

Do not infer a broad model failure

A model switch, detailed custom tooling, and a single account leave model, client, prompt, and workflow causes tangled. No independent replay or provider confirmation is available.

A separate product signal

Model choice and usage limits also cost attention

A recent r/ChatGPT thread asks for simpler model choices and comments describe rationing a weekly allowance. This is one self-selected discussion, not representative demand or evidence about model quality.

What provider status can and cannot tell you

At the 10 October check, Anthropic's status page showed no incident reported that day. Claude.ai, the API, Claude Code, and Cowork were listed operational; Claude Console remained marked degraded for usage-data loading. The page also lists resolved Opus 5.5 request errors on 6 October and spend-limit refusals on 7 October. These status records help diagnose service and access problems. They do not score completed answers or explain the reported October 4 workflow.

The longitudinal check is not ready

LiveNerf's latest README remains dated 8 October and reports 15 of 30 collections, with no post-baseline comparison. There is no result from that series yet. A missing comparison is not a finding of no change.

A local receipt for the next failure

  1. Before agentic work, keep a recoverable copy and set the approval boundary. If edits exceed the request, stop before allowing more.
  2. Record task and expected result, product surface, date, model and client labels, effort, prompt/configuration, tools, and any visible route or status information.
  3. For file-changing work, save the requested scope, approval state, paths changed, diff, backup, and rollback result. Compare the final artifact with the intended task.
  4. For a completed-answer quality comparison, replay the same task with configuration and conditions held fixed, and score the work against an explicit rubric. Record correction time and token use as costs, not as quality scores.
  5. Keep prompts, code, credentials, and raw account data local unless there is a clear reason and consent to share a redacted record. A region or egress change can be recorded as a yes/no flag; do not record raw IP addresses.

Sources

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.