LiveNerf has completed its baseline. It has not reported a comparison.
The Opus 5.5 test has moved into its follow-up period. That is a measurement milestone, not evidence of a decline.
LiveNerf: what has changed
The repository reports that 10 of 30 daily collections were complete as of 3 October. Those 10 days established the baseline; the first follow-up window began on 4 October. The run table still contains only its baseline row. All ten reported collections used the same harness hash and pinned Claude Code CLI; the author disclosed one budget-guard override on day 5.
The series tests Claude Opus 5.5 at high effort through headless Claude Code on a Claude Max subscription. It uses a calibrated 78-question panel with an Opus 5 control on GPQA items. It is a test of one served path and selected academic questions, not the raw API model, other Claude products, or everyday user workflows.
The test has a stated detection limit
The operator estimates that the daily schedule can detect an accuracy shift of about 7.5 percentage points per 10-day window. A pre-baseline validation could not distinguish an Opus 5 substitution from Opus 5.5 at 99% confidence (−3.8 ± 6.3 points in accuracy). The validation therefore cannot support detecting a swap of that size. A report-only audit also found 8 suspect answer keys and 30 ambiguous items; the preregistered sensitivity analysis will show the effect of excluding them.
The pre-registration requires the same-direction change to clear a 99% interval and a three-point threshold in both consecutive 10-day windows, with the same test harness and less than 5% sample error. If the Opus 5 control moves in the same direction, LiveNerf says it will report a harness or platform change instead of attributing the change to Opus 5.5. “No change detected” will not mean that smaller changes are impossible.
A second tracker is still collecting its baseline
Marginlab’s Claude Code tracker says its Opus 5.5/high baseline began on 24 September and that degradation detection is paused until the baseline is established. Its page displays a 77% latest daily pass rate, a 76% seven-day rate, an 81% 30-day rate, and a statistically significant +8.5% 30-day label. The page does not explain how that 30-day summary relates to the new baseline. We do not treat it as a clean Opus 5.5 before/after result. The tracker tests 50 cases daily and uses the latest Claude Code CLI, so CLI changes are part of its monitored system.
BridgeBench and service status
BridgeBench’s Opus 5.5 board still shows a latest test dated 2 October: 94.2% of launch “power,” within the board’s stated 90–110% normal-variance range. Its method page says that “power” combines task performance, token use, and cost and that its task set and weights are private. The result is not a capability-only score.
OpenAI and Anthropic reported listed systems operational at the 4 October check. Provider status is aggregate and does not show the quality of a completed response or confirm any individual account.
What to do if a familiar task fails
- Check the affected product surface and its current status.
- Save the task, time, model or route if visible, plan, settings, output, and the last usable result.
- After recovery, replay the same task against your saved baseline. Compare the completed work and correction effort.
- For a business-critical workflow, keep a small versioned task set and a control. Do not act on one daily score alone.