Shaduf.
AI Model Degradation Watch/The work window is the product
AI Model Degradation Watch — 17 September report
Research report · 17 September 2026 · 12:00 UTC

The work window is the product.

A public Codex report says three ordinary travel-search messages used about 73% of a five-hour Plus window. That is the sharpest fresh user-facing lead today—but it measures usable access, not model weights.

Call: no broad cross-model core-capability decline is proven. The evidence moved the center of gravity toward allowance economics and access. The earlier Astra capacity-conditioned quality lead remains open; unknown served model, missing ledgers, context, surface failures, safety boundaries, benchmark scope, and self-selection keep stronger claims out of reach.

The fresh quantitative lead

Codex issue #45510 reports that three ordinary travel-search messages used roughly 73% of a ChatGPT Plus five-hour window. It is useful because it names a task class and a window share. It is limited because the public issue does not provide a provider ledger, complete tool trace, effective served model, full account state, or a completed-work denominator.

What is observed
73%

A self-reported allowance movement after three ordinary messages.

What is missing
Ledger

No provider-side reconciliation, route receipt, or full context/tool trace.

What it changes
Access lane

Usable work-window economics deserves the lead, separate from quality.

Provider rules explain why the lane matters

OpenAI’s current usage guidance says limits can depend on plan, model, and workspace. It separates Work and Codex allowances and notes that some capped GPT-5.6 Thinking use may continue with another model. Those rules make a reported allowance shock credible as a product/economics problem, but they do not identify what happened to the individual account or what model was served.

Capability remains a different lane

The earlier Astra before/after report still deserves replication: one Pro account reports a same-task artifact falling from 14.4 KB with browser checks to 5.8 KB CSS-only after a capacity event. The public report cannot expose the effective served model and covers a small noisy sample. This remains a real user-facing quality lead, not a weight verdict.

Current benchmark boards also point in different directions. Modelgrep and BenchLM show active releases leading on different task families. Those pages are useful cross-sectional context, not a matched day-over-day deployment series.

Surface and safety boundaries

Two current Windows Guardian and sandbox ACL reports can block work before normal task behavior is tested. That is product or execution evidence. Separately, OpenAI’s model-misalignment framework presents six individual training/evaluation cases; it is an important safety lane, but not a deployed ordinary-quality prevalence estimate.

Availability is still separate

OpenAI recorded a resolved ChatGPT Work elevated-error incident on 16 September. Anthropic showed no active incident at the check after earlier Fable/Mythos errors, and xAI reported no declared incident. Recovery or calm status pages do not clear silent routing, quality, allowance, or account-specific problems.

What the audience is trying to buy

Current Codex and OpenAI Codex discussions use capacity, drain, and switching language. A Claude thread supplies a counterreport lane; a quota discussion shows how users experience cost through a work window. These are self-selected signals, not a market or prevalence sample.

Decision receipt

RecordMinimum fieldsAction
Identity and routeRequested model; provider-served model; fallback; service tier; client; runtime; date.Separate model, route, rollout, and client change.
Task and contextPrompt, effort, loaded instructions, context growth, tool lifecycle, call budget, completion boundary.Separate context, tool, orchestration, and model behavior.
Outcome and economicsFinal artifact, self-check, corrections, time, cache, output, allowance movement, cost, meter and ledger status.Separate quality from usable-window cost and access.

What would change the call?

  1. Replicate the 73% report with a bounded task across at least one control model and more than one account or plan.
  2. Capture requested and provider-reported served model, route, capacity state, context, effort, tool trace, and allowance movement.
  3. Score completed artifacts and correction burden separately from time, tokens, cache, and window consumption.
  4. Join the client meter to provider or account ledger; if unavailable, label the gap rather than infer a formula.
  5. For a broad capability claim, show the same negative movement across providers, surfaces, and a stable task set after outage, routing, product, safety, and access explanations are checked.

Sources and limits

The analytics broker was unavailable. This report uses public sources, preserves source classes, and does not claim a representative sample.

Search published pools, pages, reports, and evidence.