Benchmark ledger · R3 reconfirmation
Provider claims are still not a verified Union Alpha benchmark
R3 did not run or discover a new benchmark. It preserves R2’s boundary: provider-published Pareto 26.9 figures and a public harness are not independent results for the retired Union Alpha preview, the former Go alias, or a fixed shared snapshot.
Disposition
No qualifying independent, reproducible public comparison has been verified for the Union Alpha preview. The figures below remain provider claims about Pareto 26.9, not a Union Alpha scorecard.
| Published material | What it says | What it cannot prove |
|---|---|---|
| Pareto 26.9 model card | Provider claims 74 DeepSWE and 51 Terminal-Bench 4.0. | Independent verification, raw result provenance, or a Union Alpha preview mapping. |
pareto-evals harness | Documents intended datasets, settings and output workflow. | That the inspected claims came from an archived, auditable run of a particular retired preview or Go alias. |
| OpenRouter catalog telemetry | Current metadata for paid Pareto. | R3 latency, throughput, reliability, quota, SLA, or a behavioral test. |
Sources: Unbiased model card (R3-S10) and pareto-evals (R2-S11). R3 did not re-run the harness.
Do not redraw these as a ranked Union Alpha chart. A qualifying comparison needs a named route/version, task IDs, harness commit, prompts/settings, samples/retries, controlled comparisons and lawful raw or auditable outputs.
Historical leads stay unqualified
| Lead | Visible claim | Why it remains unqualified |
|---|---|---|
| Operator launch post | “DeepSWE result”; “works in every harness.” | No task/version, sample, settings, route or raw runs. S09 |
| Social reposts | Approx. 73–74% DeepSWE; 51–52% Terminal-Bench v4. | No original runner, controlled configuration, retries or raw results. S18 |
Current route scope: current availability audit. Earlier methodology record: R2 and R1.