A resolved outage, a route lead, and no broad decline.
OpenAI had a short ChatGPT Conversations incident. The strongest new quality lead is a same-account Chat-versus-Work execution difference. Neither proves that popular model weights got worse.
1. OpenAI had a declared availability incident
OpenAI’s status record reports elevated error rates for ChatGPT Conversations across Plus and Pro plans beginning at 09:13 UTC on 23 September. It moved to monitoring at 09:45 and marked the incident resolved at 09:54 UTC.
This is an availability receipt. It supports retry or wait. It does not identify a model-weight change or tell us how a particular account’s task quality changed. Claude’s status page showed all systems operational, and xAI’s status page showed no incidents declared at the check.
2. The strongest new lead is execution class
A detailed OpenAI community report claims a same-account split: ordinary Chat GPT-5.6 Thinking Extended runs were around 25–26 minutes, while a historical Chat run exceeded 100 minutes and current Work still received long worker handoff. The report preserves native metadata and distinguishes product_experience=chat from product_experience=work.
This is a stronger replay lead than a generic “it feels worse” post. It is still one account, not a prevalence estimate. The route predicate is not provider-confirmed, and the tasks, contexts, and serving conditions are not a controlled matched series. OpenAI’s August product note independently says its Sol update changed Chat while the Work/Codex version was not changing in that release. Surface and execution class belong in the test identity.
3. Access and provenance reports need separate lanes
A fresh Pro20x account report describes capacity errors, automatic model switching, slower tasks, and weaker perceived responses despite reported low seven-day usage. Another community thread reports request-level Astra traces accompanied by Luna response model IDs. These are access and provenance leads. The rate, cause, and quality effect are unverified.
A separate tool-call UI report describes calls appearing outside the Worked-for section without a generation error. A UI or audit regression can increase correction cost without a capability change.
4. An independent counterweight is positive but not longitudinal
The RoboDojo paper reports 2,100 trials in which GPT-6 Astra reached 22.48% average success and exceeded the 40 public policies in its comparison. It also reports a gap between semantic understanding and precision or dynamic control. This is useful cross-sectional context. It does not test a Chat or Work route over time, so it cannot resolve the current production lead.
5. What readers are asking for
Fresh forum threads ask for per-task token and credit usage, acceptance boundaries, completed-artifact proof, and correction-cost accounting. Public observability guides from Currai and LangChain use the same broader vocabulary: traces, tools, tokens, latency, cost, evaluation, human review, and regression tests. The useful product question is not only “what happened?” but “which next action changed the loss?”
Decision receipt
| Field | Why it matters |
|---|---|
| Task, objective, and acceptance | Prevents scope drift or a changed goal from becoming a model verdict. |
| Surface, execution class, release boundary | Chat, Work, Codex, API, foreground, and worker paths may not be the same system. |
| Requested and served model; route; plan; access | Separates selection, fallback, entitlement, and account lanes. |
| Context, tools, contract, runtime, configuration | Finds prompt, tool, migration, and orchestration changes. |
| Budget, warning/stop state, allowance, and cost | Separates capability from a task that could not use the available window. |
| Artifact, correction burden, judge, repeats, and control | Connects the symptom to usable work and tests whether movement exceeds noise. |
What would change the call?
- A versioned fixed task shows repeated negative movement against a stable control on the same surface.
- The effect survives availability, route, access, context, product, safety, tool, judge, and task-mix checks.
- The provider-served model and route are recorded, or independent replications converge without them.
- The effect appears across more than one surface or provider with a common capability estimand.
Sources and limits
Provider records establish declared events; community reports are self-selected; the RoboDojo paper is a separate task suite; observability guides are market positioning. Analytics were unavailable. There is no representative denominator, human pilot, provider ledger, effective-model receipt, or willingness-to-pay evidence. The next best move is one exposure-labeled, consented Chat-versus-Work fixed-task replay with stable grading and an explicit control.