Shaduf.Research preview
AI Model Degradation Watch/Independent tests disagree on the GPT-6 verdict
Independent tests disagree on the GPT-6 verdict | AI Model Degradation Watch
AI model watch · 25 September 2026 · 12:00 UTC

Independent tests disagree on the GPT-6 verdict.

The strongest fresh evidence is mixed. Some independent suites put GPT-6 Sol above GPT-5.6 Sol. A matched ten-task report gives GPT-5.6 Luna better quality and review results while GPT-6 Luna is faster.

Call: no broad cross-model core-capability decline is proven. The next useful action is a deployed replay with effort and route receipts, not a universal downgrade label.

1. Independent tests point in different directions

Sonar used the same framework for a Java comparison. It reports 4,444 tasks in the benchmark and 544 tasks with executable tests. Pass rates were 83.09% for GPT-6 Sol at medium effort, 85.85% for GPT-6 Astra, 81.99% for GPT-5.6 Sol, and 78.66% for GPT-5.5. This is a useful counterweight to a downgrade claim. It is one task family, one evaluation setup, and not a before/after production series.

PolicyBench also puts GPT-6 Sol above GPT-5.6 Sol on its current board. The note says serving shape and tool choice affect the result and records a corrected reference. That limits the claim to the board and its controls.

2. The matched field report is a tradeoff, not a verdict

A ten-task community comparison used the same C++/Python codebase, bounded requirements, and separate worktrees. It reports GPT-6 Luna as faster on six of ten tasks, while GPT-5.6 Luna wins more first-pass and review-quality judgments. The author marks the sample as small. Comments also raise effort choice and same-family judge concerns. Keep speed and accepted artifact quality as separate measures.

Other reports describe omitted items in an agentic reporting workflow and a Terraform refactor that loops or loses context. A separate user reports strong GPT-6 Luna high/xhigh work on a large project. These reports matter as replay intake, but they do not establish prevalence or a common cause.

3. A route or effort mismatch can imitate a capability change

A Codex/API comparison reports different reasoning-token counts and timing for matched planning prompts: the Codex subscription cells show zero reasoning tokens at low and medium effort in the report, while the API/OpenCode path shows nonzero counts; the high-effort cells are also different. The API path is not a direct OpenAI control, and a commenter suggests a Codex effort configuration problem. This is a system lead, not evidence that model weights changed.

OpenAI’s model guidance says effort, clarification behavior, context, and tool-calling choices can change follow-through. The receipt must therefore include requested effort, effective effort when observable, route, tools, context, and final artifact.

4. Status and price are separate lanes

OpenAI’s history lists resolved GPT-6 Astra Pro and ChatGPT Work events from 22–24 September. Claude’s status page is operational after a resolved multi-model event on 22 September. xAI’s page declares no incident. These records establish availability conditions, not capability.

OpenAI’s rate-card material separates API token prices from included Work/Codex use and five-hour or weekly limits. A cheaper token price does not prove that a subscription gives more usable work. The test is cost or allowance per accepted artifact under a matched task.

5. The monitoring market already has a fixed-suite competitor

AI Drift Detector publishes a weekly drift method with a fixed objective coding suite, versioned evaluations, and coverage rules. Its Margin Eval uses 50 executable tasks and treats infrastructure failures as coverage loss rather than model failures. This is useful audience and product intelligence, not evidence about current model quality. This pool should differentiate by joining public evidence to surface, route, budget, causal limits, and a reversible action.

What would change the call?

  1. Run the same task across an affected and control model on the same surface, then repeat across a second surface.
  2. Record release boundary, requested and served model, request/server/native/client/final attribution, route, effort, system prompt, tools, context, and hidden I/O.
  3. Record availability, budget, allowance, warning or stop state, speed, final artifact, correction burden, judge, coverage, and acceptance.
  4. Repeat enough tasks to estimate within-condition variance. If negative movement persists after route, product, budget, context, safety, and measurement checks, upgrade the capability hypothesis.
Reader move: retry or wait during a declared incident; audit the receipt when identity or budget is unclear; replay before switching; stop when the work window is not bounded.

Sources

    ${links}

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.