Opus 5.5's 1 October reports now include a server-delivered prompt change.
Call: broad capability decline and a core-model regression remain unproven. The new evidence adds a possible product-configuration change. One commenter reports a server-delivered prompt section and higher token use on an unchanged CLI. Another reports a perceived quality shift while their client and prompt snapshots were unchanged. Neither account is a matched task-quality test.
What issue #98679 now contains
The Anthropic Claude Code issue began with one author's Opus 5.5 High session counts. They report 29 sessions from 23–30 September on CLI 2.1.282–2.1.285 and 8 sessions on 1 October on CLI 2.1.286. Median thinking tokens per request rose from 67 to 126; median output tokens rose from 447 to 711. The author reports worse judgment, but tasks were not matched and the CLI version changed on the comparison date.
A server-delivered prompt may explain part of one account's token increase
The commenter says a “Finishing work” section appeared in a prompt snapshot after local client data refreshed at about 09:05 UTC on 1 October, under an experiment key. In their session groups, mean output tokens per request were 999 without the section and 1,616 with it; thinking tokens were 379 and 594. Their same-day averages were also higher after the refresh. The commenter says the groups are confounded with time and lack tool-schema and request-parameter records.
A second report says the prompt did not change
This commenter says their prompt snapshots were identical between 29 September and 2 October, while perceived quality and tool use changed. Their output-token count stayed in its usual range. They note that their later samples were short question-and-answer sessions rather than build tasks. This is an account report, not a controlled comparison.
The first token comparison still mixes date and client
The 29-versus-8 session comparison is useful as a lead, but not a clean model test. It does not hold task mix constant or show a shared output-quality rubric. The reported p-values do not isolate model weights, serving, prompt delivery, or client effects.
The three histories do not establish one common cause. The server prompt may have changed behavior for one account, but that does not explain a separate report with an unchanged prompt. None of the accounts provides an independently checked, matched task-quality result or a provider confirmation.
Service and access checks are separate
Anthropic's 6 October status incident says elevated Opus 5.5 request errors were investigated from 12:24 UTC and marked resolved at 12:43. It lists Claude.ai, the API, Claude Code, and Claude Cowork. A separate 7 October incident says some organizations were incorrectly paused for reaching spend limits, causing refused requests across several Claude products; it was marked resolved at 21:23 UTC. Both explain possible service or access symptoms, not the quality of a completed answer on 1 October.
Current public checks do not settle capability
ModelSentiment reported an Opus 5.5 Reddit opinion index of 66, unchanged from the previous seven-day window, as of 11:55 UTC on 9 October. The page reports 1,974 opinions across 799 threads and labels the result a sample, not a census. It measures opinion, not task performance. The day's sample was still partial.
LiveNerf's latest README still reports 15 of 30 collections as of 8 October and no post-baseline comparison. Its first two-window decision is not expected before about 24 October. There is no controlled result from that series yet.
Correction cost is part of the user's result
A recent r/Claudeopus report describes a coding task that followed an unrelated API path despite correction. The author says repeated repair consumed time, tokens, and context. This is one unverified account without a prior-version baseline. It shows what a useful local replay should measure; it does not confirm a trend.
What to record before switching
- For an error or refusal, check provider status, surface, region, and account access.
- For a completed answer, record the task and expected result, date, surface, model or route label, client/CLI version, system and custom prompt snapshots, visible experiment indicator, tools and schemas, effort, context, output, correction work, and token use.
- Keep proprietary prompts and code in a local receipt unless there is a clear reason and consent to share them.
- Replay the same task with client and prompt/configuration fixed. Compare task quality and correction cost, not token totals alone.
- Use an independent check for high-stakes work. Do not treat timing or a single report as proof of a model-side change.