Shaduf.Research preview
AI Model Degradation Watch/Space Pages recovered; Opus 5.5 sentiment fell, but no task verdict is ready
Space Pages recovered; Opus 5.5 task evidence remains incomplete | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 2 October 2026 · Public-source review

Space Pages recovered. Reddit opinion about Opus 5.5 declined; matched task evidence is incomplete.

The service status changed overnight. New Reddit opinion signals raise a question about Opus 5.5; its public repeated task test is still collecting the baseline.

Call: broad capability decline remains unproven. OpenAI marked ChatGPT Space Pages resolved at 23:44 UTC on 1 October and said affected services had recovered. LiveNerf reported 8/10 baseline days and no post-baseline result. A public Reddit opinion monitor shows recent negative movement, but measures discussion, not model output.

The Space Pages incident moved to resolved

OpenAI’s incident record said at 16:51 UTC on 1 October that some users could still see slow or timed-out results, especially when creating pages or spaces. At 23:44 it changed the status to Resolved and said impacted services had fully recovered. That is a provider-level service update. It does not prove every account recovered, and it is not a measure of completed-answer quality.

Check the affected feature in your account before repeating lost work. Save the prompt, error, timestamp, settings, route if visible, and last usable artifact. After recovery, replay the same task before attributing a new failure to capability.

Opus 5.5 has a sentiment signal, not a capability result

The author of modelsentiment reported in a Reddit post that daily Opus 5.5 opinion scores were 71–73 on 25–28 September, 69 on 29 September, 58 on 30 September, and 55 so far on 2 October. The live page separately showed a positive seven-day index of 67 from 4,637 opinions across 1,366 threads (95% interval 65–70). These are different windows, not one continuous statistic.

The monitor’s method page says it samples recent items from 25 AI subreddits and uses a model to identify and score opinions. In its reported manual audit, reviewers judged 124 of 157 scorable examples to be the author’s own opinion (79%; 95% interval 72–85%). It also says the 27 September scoring-model change shifted earlier series by up to 6.5 points. The series describes Reddit opinion. It does not measure whether Opus 5.5’s answers changed.

New Reddit posts also disagree. A LiveNerf discussion contains reports of worse performance. Another author says their own use has not worsened and points to harder tasks as one possible explanation. Neither post provides a matched task, baseline, route, settings, and independently checked outputs.

The repeated test is not ready to call

LiveNerf’s repository reported 8 of 30 daily runs collected as of 1 October, with 8 of 10 baseline days complete. It uses a 78-item calibrated panel on Opus 5.5 through headless Claude Code on a subscription. There is no post-baseline comparison result.

The preregistration calls for a ten-day baseline followed by two ten-day decision windows. It requires a 99% interval excluding zero in both windows, a change of at least three points, the same harness, and no matching control movement. The earliest possible call is around 24 October. Daily baseline movement is not a result.

Do not switch on a sentiment score. If a familiar task changed, keep the task, route, settings, output, and correction burden. Compare a new completed result with the same dated baseline. Treat service errors, account access, sentiment, and model output as separate records.

What to watch next

  • Check whether Space Pages errors recur after the provider’s resolved update.
  • Wait for the complete LiveNerf baseline and its preregistered decision windows.
  • Compare social opinion with the same time window, service status, and matched task outcomes before calling a model change.

Sources

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.