Shaduf.Research preview
AI Model Degradation Watch/Opus 5.5 scored 94.2% on BridgeBench; a decline is not established
Opus 5.5 scored 94.2% on BridgeBench; a decline is not established | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 3 October 2026 · Public-source review

Opus 5.5 scored 94.2% of BridgeBench’s launch “power”; a decline is not established.

The 2 October score is 5.8% below launch and 9.6% below the previous test. BridgeBench places it inside its stated normal-variance range. Its “power” measure combines task score, tokens, and cost.

Call: no confirmed Opus 5.5 capability regression, and no broad cross-model decline is established. The result is a commercial board’s composite score, not a capability-only comparison. The task set, actual weights, sample count, and uncertainty are not shown on its result page.

What the 94.2% score measures

BridgeBench’s Opus 5.5 board lists four tests since launch: 100.0% on 22 September, 99.2% on 27 September, 103.8% on 1 October, and 94.2% on 2 October. Its stated normal-variance band is 90–110%. The board reports a 9.6% one-day decrease from 1 to 2 October, but the latest score remains within its own band.

BridgeBench’s method page defines “power” using task score, token use, and cost. Its sample formula is labeled simplified; the actual task set and weights are private. The board does not publish a sample count or confidence interval for the latest result. The observed change could include efficiency or cost movement; it does not isolate answer quality. The displayed 90–110% band is BridgeBench’s stated rule, not a statistical interval shown on the result page.

Do not read 94.2% as “the model is 5.8% less capable.” It is a composite, a sparse set of four displayed checks, and still inside the board’s normal-variance range.

The fixed comparison has no result yet

LiveNerf’s latest visible progress note is dated 1 October. It reports 8 of 30 daily runs collected and 8 of 10 baseline days complete; its results table still lists the baseline as collecting, with no post-baseline comparison. No newer progress count appeared in the repository view checked on 3 October.

LiveNerf measures Opus 5.5 through headless Claude Code on a subscription using a frozen 78-item panel and an Opus 5 control on part of the panel. Its preregistered decision rule calls a change only if a 99% interval excludes zero in two consecutive 10-day windows, the movement is at least three points in both, the harness is unchanged, error rate stays below 5%, and the control does not move the same way. Its stated earliest possible call is around 24 October. This is a narrow served-system test, not a census of Claude products or user workflows.

Service check

OpenAI’s status page said its systems were fully operational at the 3 October check. Anthropic’s status page showed listed services operational and no incidents reported for 2 or 3 October. Provider status is aggregate. It does not verify any one account or measure the quality of completed answers.

What to do if a familiar task fails

  1. Check the specific service and product surface. Save the time, task, model or route if visible, settings, error, output, and last usable artifact.
  2. After the surface is available, replay the same task against your saved baseline. Compare the completed artifact and correction effort, not just answer tone.
  3. Do not switch models based on a social post or launch-relative composite alone. If the task is business-critical, use a small, versioned replay set and keep a control.

Sources

Back to the current verdict · Read the evidence method

Search published pools, pages, reports, and evidence.