Astra is not broadly worse. The system around it is failing in public.
Today’s evidence changes the attribution, not the headline. OpenAI confirms a recent usage-limit incident; a fresh Codex report isolates a desktop transport problem; and a current Project-link incident shows product access can fail independently of model capability. The field panel remains strong but costly and rolling.
What the sources establish
- OpenAI confirms a quota or entitlement incident. Its status record reports unexpected usage-limit resets affecting some Codex users on 9 September, from 17:29 to 17:54 UTC, and marks the incident resolved. This is meaningful corroboration for a product lane. It does not reveal a particular account’s bucket, deduction, or model quality.
- A fresh Codex report points to desktop transport. Issue #44099 describes WebSocket disconnects and reconnect loops in the macOS desktop app, with delayed or missing completion, while the reporter says web ChatGPT worked on the same account and network. That is a narrow client or transport lead. It needs independent replay before it becomes a product-wide conclusion.
- OpenAI is also monitoring Project access. The current status record says some shared ChatGPT Projects cannot be opened through direct links after mitigation was applied. A broken work container can feel like lost context or weak reasoning without testing the model at all.
- Account reports now line up in time, not yet in causality. Issues #44254 and #44234 describe usage-limit failures while local meters still showed allowance or null premium windows, followed by recovery. The reports are temporally consistent with the provider incident, but local counters are not the server ledger. Server-side reconciliation is still the missing fact.
- The older Astra incident cluster remains real but mixed. Open reports describe premature or false completion, repeated
cyber_policytermination, orchestration churn, and quota pain. They justify a user-facing incident label. They do not justify a broad claim about Astra’s core weights. - The field counterweight is strong and expensive. Beyond Benchmarks displays Astra at 94.6% meaningful outcome, 0.1% frustration, 14.0 output tokens per second, $37.42 per million tokens, $28.73 modeled active hour, 0.16 interruptions per session, 1.42x context-hunting, and 5.5% failed tool calls. Its rolling, task-mix-dependent method makes this operational context, not a matched regression test.
- Scoped controlled evidence is not collapsing. Marginlab’s direct Codex tracker remains nominal, while its Claude Code tracker is still collecting a baseline. Neither tests Astra desktop transport, ChatGPT Projects, safety interruption, or account quotas.
The sharp answer for a user
If Astra suddenly feels worse, the most defensible first labels are different, blocked, expensive, or unavailable on this surface. Check the provider record, preserve the client and route, verify tools and the final artifact, capture the allowance and error, retry on a stable surface, and then replay a matched task. Switch immediately when the work is unsafe or the economics fail; do not confuse a sensible product decision with proof of core model degradation.
What would change the call?
A provider-side account log matching the reported allowance failure would resolve the quota lane. Independent desktop-versus-web replay would scope the transport lane. A repeated clean task across accounts or clients with the same effective model and settings would raise confidence in a surface regression. A before-and-after result on a fixed Astra task with tool execution and completion scoring would move the claim toward capability degradation. Continued strong field outcomes on comparable tasks would pull it back toward localized incidents and usable-cost pressure.
Primary sources
- OpenAI usage-limit incident and ChatGPT Project incident.
- Codex issue #44099, #44254, and #44234.
- Carry-forward agent and safety evidence: #43329, #43131, and OpenAI’s Astra safety overview.
- Beyond Benchmarks, Marginlab Codex, and Marginlab Claude Code.
- Audience observability examples: AWS CloudWatch Coding Agent Insights, OpenUsage, and TraceCheck.
The analytics broker was unavailable for this run. No private Shaduf evaluation, private traffic, credentials, or account logs are presented as evidence.