How we read · updated 22 September 2026
A model name is not a test identity.
A status page, release note, community report, benchmark, judge, and route can all describe the same product while answering different questions.
Expanded receipt: record surface and product-release boundary beside model, route, context, access, budget, artifact, judge, and control.
The evidence ladder
| Lane | What it can show | What it cannot show alone |
|---|---|---|
| Provider status and history | Declared availability, incidents, recoveries, releases, and billing or agent lanes. | Account task quality, prevalence, or a weight change. |
| Product update | Which surface, model experience, plan, or tool path a provider says changed. | Whether every account received the same route or outcome. |
| Repeated time series | Whether movement exceeds within-condition noise on a versioned task and configuration. | Whether the cause is model, route, infrastructure, task mix, judge, or serving. |
| Matched replay | Same task, surface, identity, context, tools, budget, judge, and artifact comparison. | Population-wide prevalence without sampling and controls. |
| Judge or harness check | Whether the measurement process repeats consistently and whether its version moved. | Capability movement by itself. |
| Community report | Lived friction, session language, and hypotheses worth testing. | Representative denominator or causal attribution. |
What would change the call?
Run the same fixed task in regular Chat and Work or across affected and control accounts. Version the task, product surface, configuration, and judge. Separate availability failures from valid outcomes. Capture requested and served model, route, context, tools, access state, budget, warning and stop state, and artifact correction burden. Then join allowance and cost to a provider ledger. A broad decline requires repeated negative movement after routing, product, safety, access, context, orchestration, and measurement explanations are checked.