Shaduf.
AI Model Degradation Watch/The rollout moved before the verdict did
The rollout moved before the verdict did | AI Model Degradation Watch
AI model watch · 24 September 2026 · 12:00 UTC

The rollout moved before the verdict did.

OpenAI put GPT-6 Sol and Luna in Work and Codex, not ordinary Chat. Work then had a declared incident, and a community trace shows client and server model labels disagreeing. The evidence points to surface and provenance—not a broad collapse.

Call: no broad cross-model core-capability decline is proven. “Different” is the stronger call for the most concrete new OpenAI lane. The product surface, release boundary, route, capacity, and identity receipt moved before a capability verdict was available.

1. The product changed by surface

OpenAI’s 22 September release says GPT-6 Sol and GPT-6 Luna became available in ChatGPT Work and Codex and through the API, while they were not yet available in ordinary Chat. The same release says production Chat can differ from API or research evaluation because system prompts and tools differ. That makes a Chat-versus-Work comparison a different-system comparison unless the release boundary and tool set are recorded.

OpenAI’s status record separately lists elevated ChatGPT Work errors across Plus, Pro, Business, Enterprise, and Education plans and marks them resolved. Claude’s status history records a resolved 22 September error incident across multiple models and surfaces. The current OpenAI, Claude, and xAI status pages were calm at today’s check. These records establish availability and rollout, not weaker weights.

2. The sharpest new lead is an identity conflict

A detailed ChatGPT Web community report says one response showed GPT-5.6 in request and server metadata and in native attribution, while a browser DOM field changed to a GPT-5.4 slug after completion. The author reports intermittent behavior, reproduction after regeneration, and a fresh-conversation control that stayed on GPT-5.6.

This does not prove that a lower model served the response. It may be a client rendering or metadata bug. It does prove that one model label is not a sufficient receipt. The next replay must capture the identity fields across the response lifecycle and compare the final artifact, not just the label.

3. “Worse” can be a shrinking work window

Conversation capacity

A community thread collects reports of a permanent maximum-length warning after unusually little visible dialogue, including a one-request research case. Files, search results, and tool output may be part of the hidden accounting; the formula is not public. Source

Allowance

A Work user reports similar development tasks consuming three to four times more allowance after 17–18 September and says per-run tokens or credits are unavailable. This is usable-window evidence, not capability evidence. Source

Instrumentation demand

Feature requests ask for a conversation-capacity meter, handover warning, and protection from oversized tool or binary payloads. Users are asking for a budget receipt because the hidden work changes the decision. Meter request · Payload request

4. Counterweights keep the headline honest

Anthropic’s Opus 5.5 release reports broad improvement over Opus 5, lower cost, and better coding and knowledge-work performance. Early community reports are not uniform: some describe a major improvement, while others report weaker judgment or prompt following than older models. That disagreement is a reason to preserve task, effort, plan, and baseline—not a reason to average the comments into a score.

Measurement can move even when the model does not. The official Prompt Drift Lab reports large score changes from harmless prompt variations and keeps judge versions in the audit chain. DriftBench studies 236,985 prompt-response pairs across serving configurations and finds that infrastructure changes can create output drift. These studies do not explain today’s reports, but they set the control standard.

What would change the call?

  1. Run one fixed task across ordinary Chat and Work, or an affected and control account, after the rollout.
  2. Record release boundary, surface, execution class, requested model, server model, native attribution, client attribution, final response metadata, route, system prompt, tools, context, and hidden I/O.
  3. Separate availability failures from valid outcomes. Record budget, allowance movement, warning or stop state, final artifact, correction burden, and acceptance result.
  4. Use a stable judge or human rubric and a control task. If negative movement persists across the same served identity and survives product, route, context, budget, and instrument checks, the capability hypothesis becomes serious.
Reader move: retry or wait for a declared incident; audit the receipt when identity or budget is unclear; replay before switching; stop the task when the work window is not bounded.

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.