One Opus 5.5 report shows higher token use after 1 October. The cause is unresolved.
A user issue includes before-and-after session metrics, but the Claude Code version changed on the comparison date. The report does not establish that the model caused the change.
The new Opus 5.5 report
An issue opened in the Anthropic Claude Code repository on 1 October reports higher thinking-token and output-token counts and worse task judgment. The author compares 29 sessions from 23–30 September (Claude Code 2.1.282–2.1.285) with 8 sessions on 1 October (2.1.286), all at high effort. The author reports median thinking tokens per request rose from 67 to 126, and median output tokens from 447 to 711, and states Mann–Whitney p < 0.003 for three metrics.
The issue reports that the author kept their workflow and starting context similar and that a similar reliability problem appeared in claude.ai that day. But the client version and date changed together. The tasks were not matched, the issue provides no output set or task-quality rubric, and the linked session-metrics gist could not be fetched in this check. The p-values are the author's analysis; they do not isolate a model-side cause. The report is neither a provider-confirmed change nor an independent replication.
Other current signals
| Signal | What the public record shows | What it does not show |
|---|---|---|
| LiveNerf · Opus 5.5 | As of 8 October, 15/30 daily collections; the baseline is complete, 90 samples ran per reported day on the same harness hash and pinned CLI, and the result table has no post-baseline comparison. Days 5 and 12 used a disclosed budget-guard override. Its preregistration says it aligned the cutoff to 4 October 00:00 UTC and added a regression test before window-one data. | There is no published comparison result. The monitor covers a Claude Code subscription path and a fixed academic panel, not every Claude surface or user task. The current code page could not be verified from its six-day-old crawl. |
| OpenAI · Dot in Codex and Work | OpenAI marked a 7 October incident resolved after reporting Dot turn failures, delayed responses, and reduced proactivity. | This is a service/execution record, not a test of a completed model answer. |
| OpenAI · APAC services | Errors in ChatGPT Work, conversations, and GPTs for some APAC users were marked recovered at 04:51 UTC on 8 October. | The aggregate notice does not show whether a particular account was affected or whether answer quality changed. |
| OpenAI · GPT-5.6 Instant in FedRAMP | Elevated errors in FedRAMP workspaces were marked recovered at 00:03 UTC on 8 October. | This narrow deployment error does not establish a broader GPT-5.6 capability change. |
What would resolve the Opus report?
- Record task, exact model label, CLI and client versions, effort, context, tools, output, thinking/output tokens, and correction effort.
- Repeat the same task with the same prompt and configuration while holding the client version fixed; if testing a CLI update, compare both versions on matched tasks.
- Keep service state, route, and fallback evidence separate from completed-answer quality. Seek independent replication before making a model-wide claim.
For a current outage or missing agent response, check the exact service and region first. For a completed answer that seems worse, preserve the task and replay it after recovery. Do not switch based on a token change or one uncontrolled report alone.