Opus 5.5 scored 94.2% of BridgeBench’s launch “power”; a decline is not established.
The 2 October score is 5.8% below launch and 9.6% below the previous test. BridgeBench places it inside its stated normal-variance range. Its “power” measure combines task score, tokens, and cost.
What the 94.2% score measures
BridgeBench’s Opus 5.5 board lists four tests since launch: 100.0% on 22 September, 99.2% on 27 September, 103.8% on 1 October, and 94.2% on 2 October. Its stated normal-variance band is 90–110%. The board reports a 9.6% one-day decrease from 1 to 2 October, but the latest score remains within its own band.
BridgeBench’s method page defines “power” using task score, token use, and cost. Its sample formula is labeled simplified; the actual task set and weights are private. The board does not publish a sample count or confidence interval for the latest result. The observed change could include efficiency or cost movement; it does not isolate answer quality. The displayed 90–110% band is BridgeBench’s stated rule, not a statistical interval shown on the result page.
The fixed comparison has no result yet
LiveNerf’s latest visible progress note is dated 1 October. It reports 8 of 30 daily runs collected and 8 of 10 baseline days complete; its results table still lists the baseline as collecting, with no post-baseline comparison. No newer progress count appeared in the repository view checked on 3 October.
LiveNerf measures Opus 5.5 through headless Claude Code on a subscription using a frozen 78-item panel and an Opus 5 control on part of the panel. Its preregistered decision rule calls a change only if a 99% interval excludes zero in two consecutive 10-day windows, the movement is at least three points in both, the harness is unchanged, error rate stays below 5%, and the control does not move the same way. Its stated earliest possible call is around 24 October. This is a narrow served-system test, not a census of Claude products or user workflows.
Service check
OpenAI’s status page said its systems were fully operational at the 3 October check. Anthropic’s status page showed listed services operational and no incidents reported for 2 or 3 October. Provider status is aggregate. It does not verify any one account or measure the quality of completed answers.
What to do if a familiar task fails
- Check the specific service and product surface. Save the time, task, model or route if visible, settings, error, output, and last usable artifact.
- After the surface is available, replay the same task against your saved baseline. Compare the completed artifact and correction effort, not just answer tone.
- Do not switch models based on a social post or launch-relative composite alone. If the task is business-critical, use a small, versioned replay set and keep a control.