Shaduf.
Union Alpha/Benchmarks
Benchmark ledger · R3 reconfirmation

Provider claims are still not a verified Union Alpha benchmark

R3 did not run or discover a new benchmark. It preserves R2’s boundary: provider-published Pareto 26.9 figures and a public harness are not independent results for the retired Union Alpha preview, the former Go alias, or a fixed shared snapshot.

Disposition
No qualifying independent, reproducible public comparison has been verified for the Union Alpha preview. The figures below remain provider claims about Pareto 26.9, not a Union Alpha scorecard.
Published materialWhat it saysWhat it cannot prove
Pareto 26.9 model cardProvider claims 74 DeepSWE and 51 Terminal-Bench 4.0.Independent verification, raw result provenance, or a Union Alpha preview mapping.
pareto-evals harnessDocuments intended datasets, settings and output workflow.That the inspected claims came from an archived, auditable run of a particular retired preview or Go alias.
OpenRouter catalog telemetryCurrent metadata for paid Pareto.R3 latency, throughput, reliability, quota, SLA, or a behavioral test.

Sources: Unbiased model card (R3-S10) and pareto-evals (R2-S11). R3 did not re-run the harness.

Do not redraw these as a ranked Union Alpha chart. A qualifying comparison needs a named route/version, task IDs, harness commit, prompts/settings, samples/retries, controlled comparisons and lawful raw or auditable outputs.

Historical leads stay unqualified

LeadVisible claimWhy it remains unqualified
Operator launch post“DeepSWE result”; “works in every harness.”No task/version, sample, settings, route or raw runs. S09
Social repostsApprox. 73–74% DeepSWE; 51–52% Terminal-Bench v4.No original runner, controlled configuration, retries or raw results. S18

Current route scope: current availability audit. Earlier methodology record: R2 and R1.

Search published pools, pages, reports, and evidence.