OpenAI resolved a Work desktop bug and a separate Agent response issue.
Call: broad model decline is not established. The Work issue blocked new ChatGPT Work threads on a specified desktop-app version. A separate Workspace Agents issue affected response delivery. These are product availability findings, not evidence of model capability change.
Two resolved product issues
New threads failed on app version 26.1002.51308
OpenAI's status page says the version was affected. It records updated fixes for Linux/macOS and later Windows, then marks the issue resolved and advises affected users to update. The rendered page showed 11:53 AM on 7 October; it did not specify a time zone.
Some users might not see agent responses
OpenAI's separate incident page says responses could be missing during the issue and marks services recovered on 7 October. A missing response is not a low-quality completed answer.
No same-task replay has been posted
The 6 October Reddit post reports context and reasoning trouble on a selected GPT-5.6 Sol High setting. It has no prompt, output, exact client, route, or verified served-model record. It is not linked to either incident.
Current public evaluations
BridgeBench lists its latest Opus 5.5 test on 4 October at 96.5% of launch power, within its own 90–110% band. This is a vendor composite that includes task performance, tokens, and cost.
LiveNerf reports 13 of 30 daily collections by 6 October, with no post-baseline comparison. Issue #9, now marked closed, reports a potential window-boundary defect: day 1 started at 22:10 UTC, while the day 11–20 window starts at midnight. Under the issue's described calculation, that first comparison window can still be labelled baseline, leaving only one eligible window for a rule requiring two consecutive ones. The page showed no linked branch, pull request, or resolution note. The README results table still shows the baseline row only. This is a method concern, not evidence that Opus 5.5 changed. Do not interpret no result as a null or stable outcome until a code fix and valid comparison are verified.
How to test a completed-task quality report
- Save the date, task, exact product surface/client, settings, context, and any displayed responding-model label.
- Keep the original response and note corrections, retries, or tool failures.
- Check the incident page for that exact feature. Update the Work desktop app if the affected version applies.
- After the service is stable, replay the same task with the same input and settings. Compare against a dated baseline before switching models.