Shaduf.Research preview
AI Model Degradation Watch/Space Pages still showed degraded status after a recovery notice. Model decline is separate.
Space Pages stayed degraded after a recovery notice | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 1 October 2026 · Public-source review

Space Pages still showed degraded status after a recovery notice. Model decline is a separate question.

OpenAI’s current status records a Space Pages service problem; a new public Opus 5.5 time series is still establishing its baseline. Neither proves broad model decline.

Call: broad decline in deployed popular models’ core capability remains unproven. At the 12:24 UTC check, OpenAI’s ChatGPT Space Pages page still showed Monitoring / Degraded performance. A 30 September recovery note was followed by a later mitigation update. The incident page does not say whether errors persisted or returned. Check the affected product surface before retrying; do not treat a failed Page/tool session as a completed-answer quality test.

OpenAI’s status changed after the recovery notice

UTCRecorded statusWhat it meansWhat remains unknown
30 Sep, 02:20OpenAI opened a Space Pages incident.Users may experience errors creating or interacting with Pages, using Page tools, or connecting to live Page sessions.The entry does not quantify affected accounts or task impact.
30 Sep, 10:34Monitoring update: “All impacted services are recovered”; monitoring continued.A provider recovery statement at that point in the timeline.It is not a guarantee of recovery for every account or later time.
30 Sep, 15:00Later update: mitigation applied; recovery monitored.The status sequence no longer presents an unqualified recovered state.The page does not say whether impact persisted, returned, or changed in scope.
1 Oct, 12:24 checkThe incident still showed Monitoring / Degraded performance.Check the current Space Pages incident before retrying time-sensitive work.No model-quality measurement or account-level denominator is reported.

Anthropic’s status page displayed all systems operational and no incidents reported for 1 October. That is a provider snapshot, not proof of every user’s experience or of response quality.

One r/chatgptplus user also described a Space error that remained after refresh on 1 October. That single, unverified report is consistent with the service issue, but it does not estimate how many accounts were affected or confirm the underlying cause.

Do not turn a service status into a model verdict. Space creation, Page tools, and live sessions are product operations. A failed request is not a valid sample of model quality. A completed answer that seems worse needs a matched task, baseline, and visible model and settings.

The first public Opus 5.5 time series has no verdict yet

LiveNerf’s repository reported seven of 30 daily collections through 30 September. Its first ten days are the launch baseline; no post-baseline comparison result was posted at this check. The project says the first possible call is around 24 October, after the next two ten-day windows can be compared.

The preregistration describes Opus 5.5 served through headless Claude Code on a subscription, a calibrated 78-item panel, and an Opus 5 control on GPQA items. Its call rule requires a 99% interval excluding zero in two consecutive ten-day windows, an effect of at least three points, and no matching control movement. This is a useful, bounded series—not evidence of decline today, not an API test, and not a sample of every user workflow.

A second tracker uses a different score

BridgeBench’s Nerf Bench last displayed tests dated 27 September. It showed Claude Opus 5.5 at 99.2% and GPT-6 Astra at 102.8% relative to their respective launch baselines; its stated normal range is 90–110%. The board tracked two models. BridgeBench’s method page says its score combines task performance, token use, and cost while keeping tasks and weights private. Without visible sample sizes and task details, the displayed number cannot be independently recomputed or combined with LiveNerf’s different instrument.

What to do now

  1. Check the live Space Pages incident and wait for a resolved status before repeating a lost or incomplete Page task.
  2. Save the request, error, time, Page/tool state, and last usable artifact.
  3. After recovery, replay the same task with the same model, route, context, and settings. Compare valid completed outputs against a dated baseline.
  4. Keep the resulting quality comparison separate from service availability, access, usage limits, and product changes.

What would change this call

  • A provider update explaining the later Space Pages status and when it was resolved.
  • Repeated, matched post-recovery tests showing a task-quality change on the same deployed surface.
  • A completed longitudinal result with its panel, baseline, control, coverage, effect size, and uncertainty disclosed.

Sources

Back to the current verdict · Read how we separate availability and capability

Search published pools, pages, reports, and evidence.