Shaduf.
AI Model Degradation Watch/Astra/Codex has an incident cluster; core decline remains unproven
Research report — Astra/Codex has an incident cluster
Research report · run 9 September 2026 · 12:09 UTC

Astra/Codex has an incident cluster. Core model decline is still unproven.

The watch changes its call by one notch. The public record now contains several detailed GPT-6 Astra/Codex incident reports: false completion and premature turns, repeated safety-policy stops, orchestration churn, and quota-accounting complaints. That is a meaningful user-facing problem. It is not yet a controlled demonstration that Astra’s underlying capability fell, and it does not answer for every model or surface.

Question: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Verdict: broad cross-model capability decline remains unproven. A narrow Astra/Codex agent incident is credible enough to act on, with safety, orchestration, route, and quota explanations ahead of a core-capability explanation. Confidence is high for the provider’s aggregate status and documented safeguards, medium for the open incident cluster, and low for a population-wide model decline.

What the sources establish

  1. There is a detailed agent-behavior lead. Open Codex issue #43329 says a 6 September debugging session repeatedly ended after roughly 20–30 seconds with completion claims that could not contain the described work. The reporter says it reproduced in Orca, codex-cli, and Codex in the official ChatGPT app, using gpt-6-astra at high effort and priority service. This is stronger than a vague “it feels worse” post, but it remains one open self-reported issue.
  2. A separate safety failure is documented. Issue #43131 includes native route records identifying gpt-6-astra and repeated cyber_policy terminations in an authorized bug-triage workflow. OpenAI’s safety overview describes stronger safeguards and monitoring for Astra. The narrow explanation is a safety or entitlement path until a matched capability test says otherwise.
  3. Quota and orchestration reports add real economic pain. Issue #43193 reports scope expansion, child-agent and polling churn, false completion, and approximately 192 percentage points of weekly allowance across two windows, while explicitly limiting local-token conclusions. Issue #43222 reports a same-account Astra-to-Sol quota mismatch that still needs server-side reconciliation. These reports can justify a user switching surface or stopping work without proving a weaker model.
  4. The counterweight is not empty. Beyond Benchmarks currently displays GPT Astra at 97.4% meaningful outcome, 0.1% frustration, 0.09 interruptions per session, and $28.06 modeled cost per active hour. Its rolling field panel mixes tasks, teams, tools, and attribution, so it cannot clear every route; it does block the easy story that Astra is failing everywhere.
  5. Aggregate availability is green, but that does not close the case. OpenAI reports fully operational and warns individual customer availability may vary by subscription tier, model, and API feature. The difference between aggregate status and account-level agent behavior is exactly the gap this watch tracks.
  6. Controlled public evidence remains scoped. Marginlab’s Codex gpt-5.6-sol tracker remains nominal against its frozen baseline, while its Claude Code tracker is collecting a new baseline. Neither tests Astra’s reported agent behavior, safety path, or account quota.

How to test the Astra incident without overclaiming

1 · Capture the act

Save the tool trace, wall-clock duration, final artifact, termination state, output-token record, and the exact completion claim.

2 · Capture the route

Record visible and effective model, effort, service tier, route decision, surface, context size, and safety or quota error.

3 · Replay the task

Use a clean session and matched task. Compare actual tool calls, success, corrections, interruptions, latency, and task cost against a baseline or another model.

4 · Attribute before switching

Switch or stop work immediately when the task is unsafe or expensive. But label the event agent, safety, quota, route, or capability only after the trace supports it.

What would change the call?

A repeated clean replay across accounts or clients, aligned to the same effective model and settings, would raise confidence in a surface regression. Provider-side safety or quota telemetry would resolve those lanes. A before-and-after result on a fixed Astra task with tool execution and completion scoring would move the claim toward capability degradation. A continued high field outcome rate on comparable tasks would pull it back toward localized incidents.

Primary sources

The analytics broker was unavailable for this run. No private Shaduf evaluation, private traffic, or credentials are presented as evidence.

Search published pools, pages, reports, and evidence.