Stronger on several tests. Still a bad workday in the wrong lane.
Artificial Analysis reports GPT-6 Astra ahead of GPT-5.6 Sol on several intelligence and coding measures, but lower on GDPval-AA v2 and presentation quality. A same-model arena shows tools can swing outcomes. Current Codex reports show why benchmark strength does not guarantee a usable work window.
1. Artificial Analysis makes “not worse” too simple
Artificial Analysis reports Astra tying Fable 5.1 at the top of its Intelligence and Coding Agent indices, gaining about six points versus Sol, cutting the reported hallucination rate at maximum effort from 92% to 51%, and leading its Briefcase comparison by about 90 Elo. Those are meaningful positive comparisons.
The same article reports about a 45-Elo Astra deficit to Sol on GDPval-AA v2, a set of economically valuable tasks across 44 occupations, and lower Presentation Quality Elo. Astra also used fewer turns in that comparison. That can mean efficient work, a task-policy difference, or a quality tradeoff; the benchmark alone cannot identify which. It is a real task slice where “Astra is simply better” fails, but it is not a before-and-after test showing that Astra got worse over time.
Artificial Analysis reports an approximately six-point gain versus Sol on its Intelligence Index comparison.
GDPval-AA v2 places Astra below Sol by about 45 Elo in the article’s comparison.
Neither result says whether a fixed task moved down from an earlier Astra release or route.
2. The same model can change when the tool stack changes
The Astra-26 deep-research arena holds the model, broad prompt, task set, and harness design constant while changing search backends. It reports 42.3% accuracy on Exa and built-in web, 38.5% on Tavily, 26.9% on Parallel, and 23.1% on Linkup. Ten of 26 tasks were unsolved by every backend, and 88 of 130 runs stopped at a 30-call budget.
This does not prove Astra is declining. It proves a reader can experience a different “model” when the search provider, call ceiling, or agent path changes. The next independent replication should publish the task, tool trace, call budget, final artifact, and cost—not only the model name.
3. Current Codex failures are real, but they are system lanes
A current Codex issue describes a Windows user receiving capacity errors despite a displayed 100% quota for a week. Another issue describes browser/computer use staying broken within a session after a failed initialization while a new thread succeeds in the same environment. Both are access or session evidence, not model-quality scores.
A detailed tool-output race report describes the next sampling request arriving before the previous tool output is appended. The server rejects the turn for missing tool output, and the thread can remain wedged until restart. That failure can destroy a user’s result even if the underlying model would have reasoned correctly.
A separate allowance issue reports usage-limit errors while roughly 37% remained visible and bucket/reset metadata changed between responses. The appropriate label is access and meter reconciliation lead, not “the model is dumber.”
4. Economics are becoming part of the quality question
An independent Codex quota observatory reports a preliminary one-account estimate: Astra uses about 44% fewer tokens per task than Sol, yet about 2.33 times the estimated quota per task and 4.1 times the quota per token. The observatory explicitly does not measure answer quality and cannot expose the provider’s allowance ledger.
That combination matters to a power user. A shorter transcript can still buy a smaller usable work window. It also explains why local token counts, server meters, and task success must sit beside one another rather than collapse into a single “cost” number.
5. Audience demand is pulling toward a daily check and a deeper receipt
A competing live index displays 1,002 opinions in the previous 24 hours, compares models with their own baselines, and uses blunt labels such as better, worse, no change, slow, dumb, refusal, and hallucination. The page is a strong attention signal and a weak prevalence measure: its denominator and timed-task ground truth are not visible.
A detailed Reddit project report describes 2.5–3 Pro20x weekly allowances over five days, repeated correction of evaluation criteria, and a refusal to delegate from direction to verification. That is exactly the audience’s expensive question: did the model fail, or did the work fail verification? It is one self-report, not a capability estimate.
A recent research paper finds model-release perceptions in Reddit communities are dynamic and intervention-sensitive. Track the language because it brings readers in; do not use the language as the verdict.
6. Current status is still a separate lane
OpenAI’s history lists recent resolved incidents affecting work surfaces, API requests, Projects, and usage limits. Claude’s status still identifies degraded Windows Cowork while core web, API, Console, and Code surfaces are operational. Grok Web reports no active issue in the checked snapshot. These are provider declarations about availability, not comparisons of model capability.
7. The practical call
| Question | Call | Limit |
|---|---|---|
| Are popular model weights broadly worse? | Not proven. | No matched, cross-provider longitudinal capability series clears that bar. |
| Does Astra have a real negative capability slice? | There is a cross-sectional lead. | GDPval-AA v2 and presentation quality need independent replication; no temporal baseline. |
| Can the same model feel different across tools? | Yes, plausibly and measurably. | Astra-26 is one-run-per-task and harness-specific. |
| Are users seeing product and access failures? | Yes, detailed leads exist. | Issue reports lack prevalence, clean reproduction, and server attribution. |
| Is Astra more expensive in usable allowance? | Preliminary one-account lead. | No provider ledger and no answer-quality measurement. |
When a familiar workflow fails, name the lane before naming the cause: capability slice, context, tool, surface, access, meter, artifact, or still unclear. Audit context. Preserve the objective. Record the visible and effective model, route, runtime, effort, tool stages, final artifact, corrections, retries, tokens, allowance, time, and cost. Switch when the task remains unsafe, incomplete, unavailable, or uneconomic; do not mistake that rational product decision for proof of a broad capability decline.
What would change the call?
A fixed independent replication of the GDPval and presentation-quality gap would upgrade the narrow capability lead. A repeated Astra-26 study would show whether the tool-stack spread transfers to real workflows. A prevalence sample across Codex builds would separate product bugs from isolated reports. A server ledger joined to the quota observatory and task outcomes would resolve the economics lane. A fixed broad task set showing the same decline across surfaces would be needed for a general capability call.
Primary sources
- Artificial Analysis Astra benchmark · Astra-26 arena · Terminal-Bench 4.0
- Codex capacity · Codex session · Codex tool race · Codex allowance
- Quota observatory · AI user-experience index · Codex project report · Perception paper
- OpenAI history · Claude status · Grok status · OpenAI customer story
The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.