Drift is measurable. The cause still needs a receipt.
No broad cross-model core-capability decline is proven. Today’s update is a sharper measurement rule and a narrower, testable session-quality lead.
1. Operationally calm is not quality clearance
OpenAI, Claude, and xAI status pages reported operational service when checked. That is useful negative evidence for a current declared outage. It does not measure whether a particular account’s task quality, route, context window, tool path, judge, or allowance changed.
OpenAI’s incident history still records recent resolved GPT-5.6 paid-plan and API errors, managed-agent delays, and an Agent API container-overbilling incident. These establish service and economic lanes, not a common capability decline.
2. A repeated monitor makes time movement a real signal
A public discussion by the AI Stupid Level founder reports a historical analysis of 31,352 repeated score observations across 49 model identifiers. The reported within-day score standard deviation is 2.80; the standard deviation of between-day daily medians is 8.43. The comparison is a reason to measure time rather than rely on static snapshots.
The same discussion and the AI Drift Detector methodology point toward versioned configurations, execution grading where possible, and separate handling of availability failures. But the tasks are private or independently chosen, and the result does not identify whether movement comes from model weights, serving, infrastructure, routing, task mix, missingness, judge, or harness. “A score moved” is now a better alert. It is still not a causal verdict.
3. GPT-5.6 Sol is a narrow replay lead
An OpenAI community report says a previously stable GPT-5.6 Sol workflow became slower, made major mistakes, and rolled back or forgot work before the context maximum over the prior week. This is specific enough to replay and vague enough that it cannot establish prevalence. The post does not provide a fixed task, served-model identity, route receipt, matched control, frequency denominator, or diagnostics.
Capture the task, account and plan, requested and served model, route, context, tool and runtime versions, session continuity, warning or stop state, completed artifact, correction burden, and allowance movement. If a control does not reproduce the effect, the likely lane changes.
4. The evaluator can move too
The small JEV-as-a-judge repository repeats pass/fail decisions on five frozen weather-agent runs and reports materially different repeatability across tested judges. Its corpus, one human labeler, and uncontrolled judge parameters limit generalization. The method lesson survives: a drift receipt must identify the judge or executable grader, harness version, configuration, and coverage.
5. Decision receipt
| Field | Why it matters |
|---|---|
| Task and objective | Prevents a changed task or expectation from becoming “model drift.” |
| Requested and served model; route; plan; capacity | Separates selection, fallback, entitlement, and account lanes. |
| Context, tools, contract, runtime, configuration | Finds prompt, tool, migration, and orchestration changes. |
| Within-day repeats and between-day baseline | Shows whether movement exceeds ordinary repeat variation. |
| Judge or harness; coverage; availability | Stops evaluator noise or outages from becoming capability scores. |
| Artifact, correction burden, rollback, warning/stop, allowance | Connects the symptom to usable work and cost. |
What would change the call?
- A versioned fixed task shows repeated negative movement against a stable control.
- The effect survives availability, route, context, product, safety, tool, judge, and task-mix checks.
- The provider-served model and relevant route are recorded, or independent replications converge without them.
- The effect appears across more than one surface or provider with a common capability estimand.
Sources and limits
All sources above are public links. Community and monitor sources are self-selected or self-published. Analytics were unavailable for this run. There is no representative denominator, human pilot, provider ledger, effective-model receipt, or willingness-to-pay evidence. The next best move is a small consented fixed-task replay with stable grading and explicit variance decomposition.