Shaduf.
AI Model Degradation Watch/The budget is part of the truth
AI Model Degradation Watch — 19 September report
Research report · 19 September 2026 · 12:00 UTC

The work window is part of the truth.

No broad model collapse is proven. The sharper update is operational: a detailed agent task can consume a large work window without a task-wide budget warning, while new managed-surface and entitlement failures block quality judgment before the model speaks.

Call: no broad cross-model core-capability decline is proven. OpenAI, Claude, and xAI status pages are calm. That narrows declared outage evidence. It does not reveal a silent route, an account entitlement, an aggregate task budget, or the quality of a fixed task.

A task used 242.3 million reported tokens without a task-wide stop

An OpenAI Codex issue describes a Windows Codex Desktop Pro task using gpt-5.6-sol at high reasoning with subagents. The reporter says the task produced 1,545 responses, 242,318,802 total tokens, 237,203,328 cached-input tokens, and 513,284 output tokens over 2 hours 43 minutes 44 seconds. The account’s weekly usage moved by 11 percentage points.

The most important product detail is not the large number alone. The report says there was no task-wide aggregate descendant-usage warning, projected-cost notice, or automatic stop. The task included long context, screenshots, attachments, and a serial review loop, so its exact allowance outcome needs provider reconciliation. But the user-facing failure is already real: a bounded task could consume a large work window without a receipt that let the operator choose to stop.

This is not a core-capability comparison. It measures orchestration, context, allowance, and safeguard behavior. The model may have done good work and still produced a bad product experience.

Fresh Codex issues show access before quality

A new Work Cloud issue reports an HTTP/2 401 from a platform-managed remote Git operation while local Git, TLS, and network checks worked. A separate Pro issue says Codex has been unresponsive for more than 1.5 days on any prompt. The current issue index also shows fresh Windows context, session, and memory-pressure reports.

These are not small details around a benchmark. If the managed surface cannot authenticate, or the client cannot respond, the user cannot generate the artifact whose quality would answer the standing question. The correct label is access, transport, entitlement, or client failure until a fixed task actually runs.

Status is calm, which is useful but incomplete

OpenAI says fully operational. Claude says all systems operational at the check. xAI declares no incident, and Google Cloud reports no broad severe incident on its aggregate status view. These records help a reader decide whether a declared outage is active. They do not clear a managed route, an account entitlement, a hidden fallback, a context boundary, or a task budget.

One Gemini path appears to recover

A Google AI Developers Forum thread reports external-URL file inputs failing 100% with 403 for a paid API tier while inline or base64 data worked. Google staff replied that external URL fetching is supported. The reporter later said thorough testing worked in every case. That pattern is valuable because it separates a narrow backend or permission path from model generation. It is a recovery lead, not evidence that Gemini got worse or better overall.

The audience is describing the receipt

A Google entitlement thread shows users checking plan, country, account, version, and quota before purchasing higher limits, with a second user reporting the same ineligibility state. A current Claude community thread reports coding, verbosity, task-fit, skill, context, and cost complaints, while comments include counterreports and requests for statistics. An earlier Claude thread likewise mixes literary-task complaints with counterreports and quota or memory explanations.

These are good discovery sources. They tell the product what fields people need to explain a bad work session. They do not tell us how often the experience occurs, which effective model was served, or whether the task would fail under a control.

Private observability raises the bar

Langfuse positions traces, evaluations, datasets, human annotation, cost and latency dashboards, and alerts as an observability workflow. Braintrust positions production traces, live scoring, thresholds, alerts, and regression datasets. These vendor pages confirm adjacent product vocabulary: teams want to know what happened, score it, and act before users notice. They are not adoption or effectiveness data.

The public pool should not pretend to replace their instrumentation. Its edge is the explanation after the trace: which evidence lane is implicated, what remains unknown, and whether the safe next move is retry, wait, audit, stop, pin, switch, or replay.

The decision receipt

RecordMinimum fieldsAction
SymptomFixed task, surface, baseline, threshold, date, and completed outcome.Decide whether an alert should fire.
Serving and accessRequested and served model, fallback, plan, route, region, capacity, entitlement, and managed-surface error.Separate model from routing, access, rollout, and purchase state.
Budget and executionDescendant agents, response count, cached and uncached tokens, task budget, warning and stop state, context, tools, retries, and completion boundary.Separate hidden work and orchestration from answer quality.
Outcome and economicsFinal artifact, self-check, corrections, elapsed time, allowance movement, cost, and ledger status.Join quality to the usable work window.
Next moveRetry, wait, pin, switch, stop, replay, or no action; state what would end the alert.Make the watch useful instead of merely dramatic.

What would change the call?

  1. Run one fixed task across affected and control accounts or providers with the same surface, context, tools, budget, and rubric.
  2. Capture provider-served model, route, entitlement, capacity, descendant work, warning, stop state, and allowance movement rather than relying on the requested label.
  3. Score the completed artifact, correction burden, and recovered work separately from tokens, cache, time, and errors.
  4. Join the task receipt to a provider ledger where possible; otherwise label the gap.
  5. Call broad capability decline only after repeated negative movement survives outage, routing, product, safety, access, context, and tool explanations.

Sources and limits

Search published pools, pages, reports, and evidence.