Shaduf.
How we read · updated 22 September 2026

A model name is not a test identity.

A status page, release note, community report, benchmark, judge, and route can all describe the same product while answering different questions.

Expanded receipt: record surface and product-release boundary beside model, route, context, access, budget, artifact, judge, and control.

The evidence ladder

LaneWhat it can showWhat it cannot show alone
Provider status and historyDeclared availability, incidents, recoveries, releases, and billing or agent lanes.Account task quality, prevalence, or a weight change.
Product updateWhich surface, model experience, plan, or tool path a provider says changed.Whether every account received the same route or outcome.
Repeated time seriesWhether movement exceeds within-condition noise on a versioned task and configuration.Whether the cause is model, route, infrastructure, task mix, judge, or serving.
Matched replaySame task, surface, identity, context, tools, budget, judge, and artifact comparison.Population-wide prevalence without sampling and controls.
Judge or harness checkWhether the measurement process repeats consistently and whether its version moved.Capability movement by itself.
Community reportLived friction, session language, and hypotheses worth testing.Representative denominator or causal attribution.

What would change the call?

Run the same fixed task in regular Chat and Work or across affected and control accounts. Version the task, product surface, configuration, and judge. Separate availability failures from valid outcomes. Capture requested and served model, route, context, tools, access state, budget, warning and stop state, and artifact correction burden. Then join allowance and cost to a provider ledger. A broad decline requires repeated negative movement after routing, product, safety, access, context, orchestration, and measurement explanations are checked.

Read the latest report

Search published pools, pages, reports, and evidence.