A capacity event can make Astra look worse without proving the weights changed.
The strongest new evidence is narrow and uncomfortable: one Pro account reports a same-task output drop after an account-capacity throttle. The report is real enough to investigate. The served model is still unknown, so it is not a broad model verdict.
The before-and-after report
OpenAI Codex issue #44851 says a ChatGPT Pro user ran the same Chinese prompt asking for an HTML/SVG pelican riding a bicycle. The report holds requested GPT-6 Astra, high effort, and the CLI core constant. It says the account later hit “Selected model is at capacity” and compares a pre-event run on 8 September with post-event runs on 11 September.
| Reported condition | Artifact and process | What it supports |
|---|---|---|
| Before the reported account throttle | 7,621 output tokens, 564 reasoning tokens, about 4m30s, 14.4 KB; browser checked and refactored with JavaScript controls. | A concrete task-level baseline for one account and one task. |
| After the reported account throttle | 2,998 output tokens, 213 reasoning tokens, about 1m50s, 5.8 KB; CSS-only output that wrote once and stopped. | A sharp user-facing quality-shaped change after a capacity event. |
| Small follow-up sample | The report says six 11 September runs were 8.8–11.4 KB versus 24.8 KB on 9 September under the same prompt and effort. | A repeatable-looking account lead, not a population estimate. |
The source reports two runs before and six after, and notes that the SVG task is noisy. The app shell changed from 26.901.51231 to 26.903.71938 while the reported CLI core stayed at 0.153.4. Those details keep the lead useful and bounded.
Why this is not a weight verdict
The issue author could record the requested gpt-6-astra label but not the provider’s actual served model. Issue #44598 shows the same observability problem from another angle: a session requesting Sol has client records naming Sol while saved instructions make the assistant identify as GPT-6. Requested label, assistant identity, and effective provider route are different fields. Until the last field exists, “Astra got worse” is a hypothesis about a bundle of capacity, routing, budget, context, client, and model behavior.
The control points at context, not one model
Codex issue #44884 reports local telemetry from long, tool-heavy sessions: 100k+ cached-context requests, context growth across Astra, Sol, Terra, and Luna, and similar long-context profiles for Sol and Terra. Its reanalysis reports no token-meter discontinuity in the window it studied. That makes context amplification and orchestration a cross-model cost confounder. It does not establish the server’s quota formula, and it does not measure final task quality.
One reported same-task artifact delta after a capacity event. Strong enough to reproduce; not enough to blame weights.
Long cached context appears across model lanes in one local telemetry report. Cost and quality remain separate questions.
The public client records do not expose the provider-served model or fallback route.
Availability is a separate lane
| Provider record | Reported event | What it can tell us |
|---|---|---|
| OpenAI gpt-image-2.5-flare | Elevated image-generation/editing errors late on 15 September; resolved at 00:30 UTC on 16 September. | A narrow availability incident and recovery window. |
| Anthropic Fable/Mythos | Intermittent errors across claude.ai, API, Code, and Cowork on 15 September; resolved at 11:14 UTC. | A separate service lane, not a quality score. |
| xAI status | Listed services were available at the check. | A snapshot, not a longitudinal capability comparison. |
The wider capability picture stays mixed
Existing independent comparisons still point in different directions: Artificial Analysis reports Astra gains on several intelligence and coding slices alongside a GDPval-AA v2 and presentation-quality loss, while the earlier Astra-26 report shows a tool-stack spread under a shared broad harness. Those are real observations with task and harness boundaries. They are not a dated, matched series showing the weights declined.
A separate Sol report describes repeated frontend and project-execution problems without a controlled baseline. It is useful as a counterweight to an Astra-only story: field reliability can move on multiple lanes while benchmarks remain positive or mixed.
What the audience is trying to buy
Fresh Reddit discussions ask for repeated tasks, same settings, and enough evidence to inspect. Adjacent products are already selling pieces of the receipt: O11yBench shows consistency, task coverage, tokens, cost, and cache; AI Eval Observatory shows long-horizon scores; Weave connects runtime to accepted work. A route-receipts paper gives the missing provenance vocabulary. None of these public surfaces supplies a representative current-degradation denominator.
The decision receipt
| Record | Minimum fields | Decision it unlocks |
|---|---|---|
| Identity and route | Requested model; provider-reported served model; fallback; service tier; client; runtime; date. | Distinguish model, route, rollout, and client change. |
| Task and context | Prompt, effort, loaded instructions, context size/growth, tool lifecycle, call budget. | Distinguish context, tool, orchestration, and model behavior. |
| Outcome and economics | Final artifact, self-check, corrections, time, tokens/cache, allowance movement, cost, meter and ledger status. | Distinguish quality from usable-window cost and access. |
What would change the call?
- Replicate the #44851 task across affected and unaffected accounts after a capacity event.
- Capture provider-reported served model, route, capacity state, and any fallback for every run.
- Hold task, effort, context, client core, tool lifecycle, artifact rubric, and completion boundary fixed; record the app shell as a separate variable.
- Run Astra beside Sol or another control, then reconcile the account meter and server ledger.
- Show the same negative movement across providers, surfaces, and a stable task set before calling a broad capability decline.
Sources and limits
- Codex #44851 · Codex #44884 · Codex #44598 · Codex #45801
- OpenAI status history · Anthropic status · xAI status
- O11yBench · AI Eval Observatory · Weave · Route receipts
The analytics broker was unavailable. This report uses public sources, preserves source classes, and does not claim a representative sample.