Shaduf.
AI Model Degradation Watch/The quota problem is not Astra-only
AI Model Degradation Watch — 15 September report
Research report · 15 September 2026 · 12:00 UTC
NO BROAD DECLINE PROVEN · USABLE ACCESS SHARPER

The quota problem is not Astra-only.

OpenAI’s status pages record fresh Work Mode and GPT-5.6 paid-plan errors. A new Sol control says one unfinished task exhausted two five-hour windows. The more defensible diagnosis is a broken or expensive assembled work system—not broad model-weight decline.

Question: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Verdict: no broad cross-model core-capability decline is proven. But “the model is fine” is too lazy a reassurance: availability, allowance accounting, context, tool lifecycle, and task-slice quality can all make a workday worse. Astra’s benchmark picture remains mixed; the new Sol report makes the usable-window lane less model-specific.

1. OpenAI confirms a current availability lane

OpenAI’s status history, checked at 12:00 UTC, lists resolved elevated errors from GPT-5.6 and GPT-5.6 Instant on paid plans on 15 September. It also lists a resolved Work Mode incident: some Plus users had trouble starting or resuming tasks or had limited workspace tools and files. An Agents API incident was resolved on 14 September.

That is useful confirmation, not a quality score. The status footer warns that aggregate availability may not match an individual customer’s tier, model, or feature. A recovered error is evidence about the product lane; it does not say the weights changed.

Declared lane
Work Mode

Task starts, resumes, files, and workspace tools were affected for some users before mitigation.

Declared lane
GPT-5.6

Paid-plan errors were elevated and then resolved inside the recorded incident window.

What it cannot say
Quality

No status record here measures reasoning, coding, factuality, or temporal capability.

2. The control that weakens an Astra-only story

A fresh open Codex issue reports GPT-5.6 Sol at high reasoning on CLI 0.154.0, ChatGPT Plus, and a bounded implementation task. The author says the same unfinished task consumed a full five-hour window twice. The second session records 240,196 total tokens, 197,230 input tokens, 10,705,408 cached input tokens, 42,966 output tokens, and 16,923 reasoning tokens; the visible weekly allowance moved from 53% to 37%.

This is not proof that Sol is bad, Astra is good, or the provider charged incorrectly. It is a valuable control against a convenient narrative: a usable-window failure can be caused by model effort, cached context, retries, runtime orchestration, allowance accounting, or the task itself. The report asks for server reconciliation and explicitly does not settle those causes.

Do not collapse these sentences: “my task stopped,” “my allowance burned,” “the service had an incident,” and “the model got less capable.” They may share a workday without sharing a cause.

3. The meter is part of the product

A separate Windows Plus issue reports roughly 21 percentage points of weekly allowance disappearing while the account was idle, a reset date changing, and a subsequent usage-limit rejection. That is an account-level reconciliation lead. Without the provider ledger, it could be delayed accounting, another surface, meter synchronization, or an entitlement transition.

OpenAI’s usage guide makes the confounder explicit: Work and Codex share an allowance whose use depends on plan, workspace, model, task, inputs, outputs, reasoning effort, Fast mode, and multistep work. It also says higher effort can use more without always producing a better result. A local token total is therefore not a provider ledger, and a smaller transcript is not automatically a cheaper completed outcome.

4. Astra still has a real but scoped mixed benchmark picture

The previous check remains relevant. Artificial Analysis reports Astra ahead of Sol on several intelligence and coding measures, while about 45 Elo lower on GDPval-AA v2 and lower on presentation quality. The Astra-26 arena shows the same model’s research accuracy moving with search backend and call budget.

Those are real task-slice and harness results. They are not a fixed before-and-after series showing Astra got worse. The right public call is mixed capability evidence plus sharper usable-access evidence, not a universal scoreboard.

5. The audience is asking for receipts

A current Claude performance hub says it is often the subreddit’s highest-traffic post, asks for prompts, responses, platform, time, and screenshots, and publishes logs and workarounds. That is a demand signal for structured incident detail, not a population estimate.

The adjacent products split the job. Polairity exposes provider, model, and harness context but currently shows only four signals from one person in its 30-day read. HypeBench tracks attention and says attention is not quality. NerfedOrNot makes crowd votes memorable while warning they are not proof. Benchmark Radar explicitly says sightings cannot establish benchmark validity, representativeness, stable versions, or production value.

The opening is obvious: give readers the speed of a check-in, then show the fields that let a quality team decide whether to retry, wait, pin, audit, or switch.

6. The practical call

QuestionCallLimit
Are popular model weights broadly worse?Not proven.No matched, cross-provider longitudinal capability series clears that bar.
Did a current service lane fail?Yes, provider records say so.Status confirms availability, not user prevalence or quality.
Is usable-window pain Astra-specific?No. A Sol control reports it too.One account and one task; no server-ledger reconciliation.
Does Astra have negative task slices?There is a cross-sectional lead.GDPval and presentation-quality results need fixed replication; no time series.
What should a reader capture?System tuple plus outcome.Model, route, context, effort, cache, tools, artifact, allowance, time, and next action.

When a workflow fails, name the lane before naming the cause: capability slice, context, tool, surface, access, meter, artifact, or still unclear. Preserve the objective. Record the visible and effective model, route, runtime, effort, tool stages, final artifact, corrections, retries, cached input, allowance, time, and cost. Switch when the work remains unsafe, incomplete, unavailable, or uneconomic; do not turn that rational decision into proof of a broad capability decline.

What would change the call?

A matched Sol/Astra long-task replay with fixed context, effort, runtime, tool trace, final artifact, cached input, and server ledger would resolve the usable-window attribution. A fixed independent replication of the GDPval and presentation-quality gap would upgrade the narrow capability lead. A prevalence sample across Work Mode plans and regions would show whether the recovered incidents left a persistent user-facing tail. A fixed broad task set showing the same decline across surfaces would be needed for a general capability call.

Primary sources

The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.

Search published pools, pages, reports, and evidence.