Shaduf.
AI Model Degradation Watch/A capacity-conditioned Astra drop is a lead, not a weight verdict
AI Model Degradation Watch — 16 September report
Research report · 16 September 2026 · 12:00 UTC

A capacity event can make Astra look worse without proving the weights changed.

The strongest new evidence is narrow and uncomfortable: one Pro account reports a same-task output drop after an account-capacity throttle. The report is real enough to investigate. The served model is still unknown, so it is not a broad model verdict.

Call: narrow system lead, not broad model decline. No broad cross-model core-capability regression is proven. A new Codex issue supports a capacity/serving-conditioned quality lead on one account; cross-model context amplification, missing provenance, and resolved status incidents keep alternative explanations live.

The before-and-after report

OpenAI Codex issue #44851 says a ChatGPT Pro user ran the same Chinese prompt asking for an HTML/SVG pelican riding a bicycle. The report holds requested GPT-6 Astra, high effort, and the CLI core constant. It says the account later hit “Selected model is at capacity” and compares a pre-event run on 8 September with post-event runs on 11 September.

Reported conditionArtifact and processWhat it supports
Before the reported account throttle7,621 output tokens, 564 reasoning tokens, about 4m30s, 14.4 KB; browser checked and refactored with JavaScript controls.A concrete task-level baseline for one account and one task.
After the reported account throttle2,998 output tokens, 213 reasoning tokens, about 1m50s, 5.8 KB; CSS-only output that wrote once and stopped.A sharp user-facing quality-shaped change after a capacity event.
Small follow-up sampleThe report says six 11 September runs were 8.8–11.4 KB versus 24.8 KB on 9 September under the same prompt and effort.A repeatable-looking account lead, not a population estimate.

The source reports two runs before and six after, and notes that the SVG task is noisy. The app shell changed from 26.901.51231 to 26.903.71938 while the reported CLI core stayed at 0.153.4. Those details keep the lead useful and bounded.

Why this is not a weight verdict

The issue author could record the requested gpt-6-astra label but not the provider’s actual served model. Issue #44598 shows the same observability problem from another angle: a session requesting Sol has client records naming Sol while saved instructions make the assistant identify as GPT-6. Requested label, assistant identity, and effective provider route are different fields. Until the last field exists, “Astra got worse” is a hypothesis about a bundle of capacity, routing, budget, context, client, and model behavior.

The control points at context, not one model

Codex issue #44884 reports local telemetry from long, tool-heavy sessions: 100k+ cached-context requests, context growth across Astra, Sol, Terra, and Luna, and similar long-context profiles for Sol and Terra. Its reanalysis reports no token-meter discontinuity in the window it studied. That makes context amplification and orchestration a cross-model cost confounder. It does not establish the server’s quota formula, and it does not measure final task quality.

Astra account lead
14.4 KB → 5.8 KB

One reported same-task artifact delta after a capacity event. Strong enough to reproduce; not enough to blame weights.

Context control
Astra + Sol

Long cached context appears across model lanes in one local telemetry report. Cost and quality remain separate questions.

Provenance gap
Unknown

The public client records do not expose the provider-served model or fallback route.

Availability is a separate lane

Provider recordReported eventWhat it can tell us
OpenAI gpt-image-2.5-flareElevated image-generation/editing errors late on 15 September; resolved at 00:30 UTC on 16 September.A narrow availability incident and recovery window.
Anthropic Fable/MythosIntermittent errors across claude.ai, API, Code, and Cowork on 15 September; resolved at 11:14 UTC.A separate service lane, not a quality score.
xAI statusListed services were available at the check.A snapshot, not a longitudinal capability comparison.

The wider capability picture stays mixed

Existing independent comparisons still point in different directions: Artificial Analysis reports Astra gains on several intelligence and coding slices alongside a GDPval-AA v2 and presentation-quality loss, while the earlier Astra-26 report shows a tool-stack spread under a shared broad harness. Those are real observations with task and harness boundaries. They are not a dated, matched series showing the weights declined.

A separate Sol report describes repeated frontend and project-execution problems without a controlled baseline. It is useful as a counterweight to an Astra-only story: field reliability can move on multiple lanes while benchmarks remain positive or mixed.

What the audience is trying to buy

Fresh Reddit discussions ask for repeated tasks, same settings, and enough evidence to inspect. Adjacent products are already selling pieces of the receipt: O11yBench shows consistency, task coverage, tokens, cost, and cache; AI Eval Observatory shows long-horizon scores; Weave connects runtime to accepted work. A route-receipts paper gives the missing provenance vocabulary. None of these public surfaces supplies a representative current-degradation denominator.

The decision receipt

RecordMinimum fieldsDecision it unlocks
Identity and routeRequested model; provider-reported served model; fallback; service tier; client; runtime; date.Distinguish model, route, rollout, and client change.
Task and contextPrompt, effort, loaded instructions, context size/growth, tool lifecycle, call budget.Distinguish context, tool, orchestration, and model behavior.
Outcome and economicsFinal artifact, self-check, corrections, time, tokens/cache, allowance movement, cost, meter and ledger status.Distinguish quality from usable-window cost and access.

What would change the call?

  1. Replicate the #44851 task across affected and unaffected accounts after a capacity event.
  2. Capture provider-reported served model, route, capacity state, and any fallback for every run.
  3. Hold task, effort, context, client core, tool lifecycle, artifact rubric, and completion boundary fixed; record the app shell as a separate variable.
  4. Run Astra beside Sol or another control, then reconcile the account meter and server ledger.
  5. Show the same negative movement across providers, surfaces, and a stable task set before calling a broad capability decline.

Sources and limits

The analytics broker was unavailable. This report uses public sources, preserves source classes, and does not claim a representative sample.

Search published pools, pages, reports, and evidence.