Shaduf.Research preview
AI Model Degradation Watch/OpenAI resolves a Work desktop issue and a separate Agent response issue
OpenAI fixes a Work desktop issue; no broad model decline is established | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 7 October 2026 · Public-source review

OpenAI resolved a Work desktop bug and a separate Agent response issue.

Call: broad model decline is not established. The Work issue blocked new ChatGPT Work threads on a specified desktop-app version. A separate Workspace Agents issue affected response delivery. These are product availability findings, not evidence of model capability change.

Suggested action: If affected by the Work desktop issue, update the app, as OpenAI advises. For a weak completed answer, preserve the task, settings, surface, displayed model, output, and correction work; replay the same task against a dated baseline after recovery. This action is suggested, not attempted or verified by a reader.

Two resolved product issues

ChatGPT Work desktop

New threads failed on app version 26.1002.51308

OpenAI's status page says the version was affected. It records updated fixes for Linux/macOS and later Windows, then marks the issue resolved and advises affected users to update. The rendered page showed 11:53 AM on 7 October; it did not specify a time zone.

Workspace Agents

Some users might not see agent responses

OpenAI's separate incident page says responses could be missing during the issue and marks services recovered on 7 October. A missing response is not a low-quality completed answer.

ChatGPT quality complaint

No same-task replay has been posted

The 6 October Reddit post reports context and reasoning trouble on a selected GPT-5.6 Sol High setting. It has no prompt, output, exact client, route, or verified served-model record. It is not linked to either incident.

Current public evaluations

BridgeBench lists its latest Opus 5.5 test on 4 October at 96.5% of launch power, within its own 90–110% band. This is a vendor composite that includes task performance, tokens, and cost.

LiveNerf reports 13 of 30 daily collections by 6 October, with no post-baseline comparison. Issue #9, now marked closed, reports a potential window-boundary defect: day 1 started at 22:10 UTC, while the day 11–20 window starts at midnight. Under the issue's described calculation, that first comparison window can still be labelled baseline, leaving only one eligible window for a rule requiring two consecutive ones. The page showed no linked branch, pull request, or resolution note. The README results table still shows the baseline row only. This is a method concern, not evidence that Opus 5.5 changed. Do not interpret no result as a null or stable outcome until a code fix and valid comparison are verified.

Do not connect incidents to the complaint without a task trace. The new Work and Agent records are for different product functions. Provider status is aggregate. No record identifies the Reddit user's account, route, or result.

How to test a completed-task quality report

  1. Save the date, task, exact product surface/client, settings, context, and any displayed responding-model label.
  2. Keep the original response and note corrections, retries, or tool failures.
  3. Check the incident page for that exact feature. Update the Work desktop app if the affected version applies.
  4. After the service is stable, replay the same task with the same input and settings. Compare against a dated baseline before switching models.

Sources

Back to current verdict · Read the method

Search published pools, pages, reports, and evidence.