OpenAI marked incidents recovered; one GPT-5.6 quality report has no replay data.
Call: broad model decline is unproven. A Reddit user says a ChatGPT interaction on a selected GPT-5.6 Sol High setting felt weaker than two days earlier. The post appeared after OpenAI reported recent service recoveries, but it has no matched task, output, or verified served-model record. Replay the same task after recovery; timing alone does not establish a cause.
What OpenAI’s timeline says
Work Mode recovered at 02:10 UTC
OpenAI’s incident page says elevated errors began at 07:03 UTC on 5 October. It lists ChatGPT Work and Codex in ChatGPT Desktop and marks them recovered early on 6 October.
Separate conversation errors recovered at 00:01 UTC
A second incident record says some ChatGPT conversations saw elevated errors on 5 October and were fully recovered at 00:01 UTC on 6 October.
Another incident recovered at 03:17 UTC
This record lists ChatGPT conversations, Work, Image Generation, dots, and Space. It says investigation began at 02:41 and recovery completed at 03:17.
OpenAI’s history also lists a wider ChatGPT, Codex, API, and Agents API item as recovered at 08:26 UTC. Its detail page displays earlier investigation updates dated 29 September, so that item’s start and relation to today’s reports are unclear; we do not use it as a separate, confirmed 6 October start.
A fresh complaint, not a controlled comparison
A 6 October r/ChatGPT post says the selected GPT-5.6 Sol High interaction felt like an older, weaker system, with context and reasoning symptoms. The post appeared several hours after the reported recovery updates. It does not include a prompt, conversation link, task output, exact client, route, or verified model identity. Replies include similar complaints and users saying their own ChatGPT use seemed normal. This is one user report. It does not measure prevalence or model capability.
The Opus 5.5 trackers have not changed their last reported result
BridgeBench still shows 96.5% of launch power on 4 October, above its 2 October low and within its 90–110% band. Its measure combines task performance, tokens, and cost. LiveNerf last reports 12 of 30 daily collections by 5 October, with two follow-up days and no post-baseline comparison. These Claude monitors do not answer the new GPT-5.6 complaint.
If your task still fails
- Check status for the exact product feature and time; do not infer from the model name alone.
- Save a privacy-safe task definition, product surface/client, date, selected model and effort, context, response, and any model label shown.
- Once the surface is stable, replay the same task with the same inputs and settings; compare the finished result and correction work with your dated baseline.
- Only then decide whether to change settings, route, model, or provider. For urgent or high-stakes work, use an independent check instead of relying on an unverified response.