Shaduf.Research preview
AI Model Degradation Watch/One Opus 5.5 report shows higher token use; cause unresolved
One Opus 5.5 report shows higher token use after Oct 1; cause unresolved | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 8 October 2026 · Public-source review

One Opus 5.5 report shows higher token use after 1 October. The cause is unresolved.

A user issue includes before-and-after session metrics, but the Claude Code version changed on the comparison date. The report does not establish that the model caused the change.

Verdict: no broad model decline or confirmed Opus 5.5 capability regression is established. This one-account report is a useful signal to test, not a verdict.

The new Opus 5.5 report

An issue opened in the Anthropic Claude Code repository on 1 October reports higher thinking-token and output-token counts and worse task judgment. The author compares 29 sessions from 23–30 September (Claude Code 2.1.282–2.1.285) with 8 sessions on 1 October (2.1.286), all at high effort. The author reports median thinking tokens per request rose from 67 to 126, and median output tokens from 447 to 711, and states Mann–Whitney p < 0.003 for three metrics.

The issue reports that the author kept their workflow and starting context similar and that a similar reliability problem appeared in claude.ai that day. But the client version and date changed together. The tasks were not matched, the issue provides no output set or task-quality rubric, and the linked session-metrics gist could not be fetched in this check. The p-values are the author's analysis; they do not isolate a model-side cause. The report is neither a provider-confirmed change nor an independent replication.

Other current signals

SignalWhat the public record showsWhat it does not show
LiveNerf · Opus 5.5As of 8 October, 15/30 daily collections; the baseline is complete, 90 samples ran per reported day on the same harness hash and pinned CLI, and the result table has no post-baseline comparison. Days 5 and 12 used a disclosed budget-guard override. Its preregistration says it aligned the cutoff to 4 October 00:00 UTC and added a regression test before window-one data.There is no published comparison result. The monitor covers a Claude Code subscription path and a fixed academic panel, not every Claude surface or user task. The current code page could not be verified from its six-day-old crawl.
OpenAI · Dot in Codex and WorkOpenAI marked a 7 October incident resolved after reporting Dot turn failures, delayed responses, and reduced proactivity.This is a service/execution record, not a test of a completed model answer.
OpenAI · APAC servicesErrors in ChatGPT Work, conversations, and GPTs for some APAC users were marked recovered at 04:51 UTC on 8 October.The aggregate notice does not show whether a particular account was affected or whether answer quality changed.
OpenAI · GPT-5.6 Instant in FedRAMPElevated errors in FedRAMP workspaces were marked recovered at 00:03 UTC on 8 October.This narrow deployment error does not establish a broader GPT-5.6 capability change.

What would resolve the Opus report?

  1. Record task, exact model label, CLI and client versions, effort, context, tools, output, thinking/output tokens, and correction effort.
  2. Repeat the same task with the same prompt and configuration while holding the client version fixed; if testing a CLI update, compare both versions on matched tasks.
  3. Keep service state, route, and fallback evidence separate from completed-answer quality. Seek independent replication before making a model-wide claim.

For a current outage or missing agent response, check the exact service and region first. For a completed answer that seems worse, preserve the task and replay it after recovery. Do not switch based on a token change or one uncontrolled report alone.

Search published pools, pages, reports, and evidence.