Ask which system changed.
“Worse” now arrives as a rollout boundary, a model-label mismatch, an early conversation limit, or a larger allowance bill. The reader needs a replayable surface receipt.
Current audience leads
Work changed before Chat
OpenAI’s GPT-6 Sol/Luna release is scoped to Work, Codex, and API at launch. A Chat-versus-Work comparison that ignores this boundary is not a matched comparison.
One label is not enough
A community trace reports different model values in the client and server/native layers. Verify identity across the response lifecycle before attributing a quality change.
Hidden I/O can end the task
Conversation-capacity and allowance reports describe files, search, tool output, and per-run cost that users cannot see. This is usable-work evidence, not weight evidence.
Opus 5.5 is not one story
Provider claims and early user reactions include strong improvement and specific disappointment. Keep the task and baseline attached to the report.
What to capture
- Release boundary, surface, execution class, plan, account or workspace state, requested and served model, request/server/native/client/final attribution, route, and time.
- System prompt, tools, context, hidden tool/file payload, configuration version, judge or harness, availability, and control.
- Task budget, allowance movement, warning or stop state, interruption, final artifact, correction burden, and acceptance result.
- Switching reason, replacement outcome, and whether the receipt changed the decision.