A negative coding slice, a near-tie, and a real outage.
The fresh evidence is more specific, not more universal. Bug Hunt reports a negative GPT-6 Sol coding result against GPT-5.6 Sol. Braintrust reports a near-tie on a broader API suite. OpenAI also records a resolved Codex outage. These are different evidence lanes.
1. The negative result is real but narrow
Bug Hunt Bench reports a 23 September update with 105 planted bugs across two production repositories. Its displayed maximum means are 29.3/105 for GPT-6 Sol across three runs, 43.5/105 for GPT-5.6 Sol across two runs, and 45/105 for GPT-6 Astra across three runs. That is a useful negative coding slice.
The benchmark method and raw data limit the inference. Most rows have one run. Run spread is material. Cost categories are not interchangeable. Context accounting is not fully verified for the GPT-5.6 Sol control, and the board is cross-sectional rather than a time series. The result does not show that GPT-6 Sol regressed from its own prior state, or that the gap transfers to a user surface.
2. A broader API evaluation is nearly tied
Braintrust compares 175 examples across seven task families. It reports GPT-6 Sol at 76.8% problems solved, GPT-5.6 Sol at 77.2%, and GPT-6 Luna at 78.0%. GPT-6 Sol is fastest in the table at 2.2 seconds per case. The page says the top score differences are not statistically significant. The suite uses its own APIs, effort settings, task mix, judge, and outcome definition.
The two studies do not cancel each other. They estimate different things. Together they block a universal downgrade claim and identify the next test: hold task family, harness, effort, route, served identity, tools, context, judge, and artifact review constant, then repeat.
3. The Codex incident is availability evidence
OpenAI's 25 September status incident records a full Codex outage affecting Codex Web, API, CLI, and VS Code extension. The provider identified errors at 22:58 UTC, offered API-key login as a temporary access path at 23:19, and marked services recovered at 23:54. That tells a reader to retry or wait. It does not establish a model capability decline.
OpenAI's release note also says API or research evaluation can differ from production ChatGPT when system prompts and tools differ, and the GPT-6 Sol/Luna rollout is surface-specific. An API benchmark is not automatically a Chat, Work, or Codex result.
4. Price, allowance, and work are separate
OpenAI's current rate-card material keeps API token prices separate from included Work/Codex usage and five-hour or weekly limits. A lower token price is not proof of more accepted subscription work. The relevant receipt joins model, effort, route, allowance, task outcome, correction burden, and accepted artifact.
A high-engagement r/codex thread makes this the user-facing question: did quality, speed, price, and usable allowance move together? The thread is self-selected and its calculations are not an account ledger. It is audience demand, not market evidence.
5. What to record before switching
| Field | Why it matters | Minimum receipt |
|---|---|---|
| Surface and incident | An outage or product path can look like a weak model. | Chat, Work, Codex, API, client, route, incident ID, recovery time. |
| Identity and effort | The label can hide routing or a different work budget. | Requested, served, request/server/native/client/final attribution, effort, tokens when available. |
| Task and harness | Bug Hunt and Braintrust measure different estimands. | Task family, prompt, tools, context, harness version, judge, run count, spread, coverage. |
| Outcome | Speed alone is not completed quality. | Artifact, acceptance, correction burden, failure mode, cost or allowance, control result. |
Next move
- During a declared outage, retry or wait and record the recovery boundary.
- Replay one fixed task on the affected path and a control path with the same tools, context, effort, and judge.
- Repeat enough times to estimate within-condition variance and inspect the final artifact.
- Only upgrade the capability hypothesis if the negative result persists after route, product, budget, context, safety, availability, and measurement checks.