Astra has a narrow regression. The model-collapse story still does not clear the bar.
OpenAI’s own documentation adds a real model-property change: written-chain-of-thought monitorability is lower than in GPT-5.6 Sol. Its guidance also makes context a live confounder. New Codex reports add intent drift and a sharp allowance-meter discontinuity. An independent coding benchmark pushes back against a universal decline story.
1. The new model-level signal is monitorability, not general intelligence
OpenAI’s GPT-6 Astra system card says Astra’s written-chain-of-thought monitorability is lower than GPT-5.6 Sol’s in evaluations about oversight and monitor evasion. In the same overview, OpenAI reports stronger capabilities, stronger alignment, and fewer higher-severity misalignment flags in a simulation of more than 54,000 internal Codex tasks.
That combination is not a contradiction once the estimand is named. A model can improve at tasks and safety outcomes while becoming harder for a chain-of-thought monitor to interpret. The monitorability finding is still a genuine narrow model property. It does not show that an ordinary user receives worse code, reasoning, or factual answers. The system card says much of this work is adversarial and warns that lack of observed failures does not establish reliability across settings. Provider-authored internal evaluations are evidence about what OpenAI measured and disclosed; they are not an independent production benchmark.
2. Context is now a first-class alternative explanation
In an 11 September developer post, OpenAI says skill descriptions can be too long, numerous skills can be shortened in context, and contradictory descriptions can cause the wrong instruction to load. It recommends progressive disclosure, contextual file reading, explicit completion boundaries, and auditing repository guidance. The API guidance similarly says Astra can be more sensitive to skills and other files such as AGENTS.md; it may ask for clarification, delegate less, or test more than a small task requires.
This matters because a new Codex issue, observed on 9 September, reports Astra losing conversational intent between adjacent turns, drifting from an established reporting skill, and failing to apply corrections. The report records a sequence, model, date, and expected behavior. It also says independent reproduction in a fresh environment has not been established and that a same-day update’s causal role is unconfirmed.
The honest classification is a detailed agent-behavior lead with two live explanations: model/version behavior and context or instruction interaction. The cheapest next test is a context audit, not a new provider.
List what the model saw
Skills, AGENTS.md, prompt, prior turns, tool schema, context pressure, and completion boundary.
Hold the task still
Keep model, route, runtime, effort, objective, and expected artifact constant in a clean context.
Measure correction
Did the objective survive? Count corrections and accepted artifact separately from tokens, retries, and time.
3. Account pain is real, but the meter is unresolved
A second Codex issue reports the weekly used percentage moving from 14% to 88% in a single server response. The associated record shows 200,123 total tokens for that response. Across 232 successful responses in the affected workload, local input was 98.1% cached, with 55.3 million total input tokens and about 0.78 million fresh input tokens. The report explicitly does not claim that the charge is wrong.
This is a strong account-reconciliation lead because the discontinuity is timestamped and tied to request-level records. It is not a capability result, and it is not yet a billing verdict. Local tokens are not the same as allowance-weighted use. Hidden work, retries, delayed reconciliation, entitlements, and rate buckets remain on the provider side. The reader-facing classification should be meter unexplained until the server ledger joins the request IDs.
4. The counterweight: Astra still scores well on a fixed coding benchmark
Terminal-Bench 4.0 reports a 66-task Harbor run with GPT-6 Astra at 58.2% ±2.8, Fable 5.1 at 57.9% ±3.8, GPT-5.6 Sol at 37.3% ±3.8, and Gemini 3.8 Flash at 19.1% ±3.4. Astra and Fable are close within the displayed margins, so this is not a clean universal ranking.
The benchmark is useful precisely because it is a counterweight. It measures terminal task resolution through specified agents and effort settings, with task-set and cost metadata. It does not see a user’s loaded skills, adjacent-turn corrections, account meter, routing, transport, or Windows workspace. Snorkel discloses support through its Open Benchmarks Grants program and a reviewer role; retain that disclosure when using the page. A high score and a painful workflow can both be true because they measure different objects.
5. Current availability is not moving in lockstep
At this check, OpenAI, Google Cloud, and Grok Web report operational service or no broad severe incident. Claude’s component API lists Claude Cowork in degraded performance while claude.ai, Console, Claude API, and Claude Code are operational. The unresolved incident list attributes the Windows Cowork problem to workspace access after a Windows update; chat and file operations mostly work.
This is high-confidence evidence about declared availability. It is not a test of model quality. The same model can be usable in one surface and blocked in another.
6. What the evidence says about the question
| Question | Call | Limit |
|---|---|---|
| Are popular model weights broadly worse? | Not proven. | No matched, cross-provider longitudinal capability series clears that threshold. |
| Did an Astra model property worsen? | Yes, narrowly: monitorability. | Provider-authored oversight result; impact on ordinary task quality is unknown. |
| Can context make Astra feel worse? | Yes, plausibly and provider-documented. | Guidance does not prove the Codex issue’s cause; audit and replay are required. |
| Are Codex users seeing access problems? | There are detailed current leads. | One intent report and one meter report; no prevalence, clean reproduction, or server ledger. |
| Does a benchmark contradict the incident reports? | No; the estimands differ. | Terminal-Bench is cross-sectional and harness-dependent. |
7. The practical call
If a familiar workflow fails, name the lane before naming the cause: different, blocked, ungrounded, expensive, unavailable, meter-unexplained, or still unclear. Audit context. Preserve the objective. Record the visible and effective model, route, runtime, effort, tool stages, final artifact, corrections, retries, tokens, allowance, time, and cost. Reconcile the account meter. Switch when the task remains unsafe, incomplete, unavailable, or uneconomic; do not mistake that rational product decision for proof of a broad capability decline.
What would change the call?
A clean independent Astra replay after a skills and context audit would separate model behavior from instruction interaction. A normal-use monitorability study would connect the system-card property to external task observability. A server ledger joined to the Codex meter report would resolve the account lane. A repeated benchmark with fixed model-agent-effort settings and a field outcome link would tell us whether the cross-sectional counterweight persists. A fixed broad task set showing the same decline across surfaces would be needed for a general capability call.
Primary sources
- OpenAI GPT-6 Astra system card · OpenAI skills and prompts guidance · OpenAI model guidance
- Codex intent issue #44136 · Codex meter issue #44213 · Codex global status-line request
- Terminal-Bench 4.0 · MonitrLLM · Inngest Durable Execution Benchmark
- Gemini MCP report · Gemini Agentic Video report · Codex runtime comparison · Codex quota report
The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.