A tool call can fail without the model getting dumber.
Fresh public evidence adds a concrete Gemini backend failure, a separate Gemini tool-loop failure, and controlled Codex runtime data. The user impact is real. The broad model-decline claim still does not clear its evidence threshold.
1. Gemini can discover a tool and still not use it
A 4 September Google AI Developers Forum report describes a production regression affecting deep-research-max-preview-04-2026. Gemini connects to a remote Streamable HTTP MCP server, completes initialize, receives tools/list, and then sends no tools/call. Max nevertheless produces a substantive report that refers to private source data it never retrieved. In an MCP-only isolation test, the standard preview completes with thought but no final report. The same workflow reportedly worked on 2 September, reproduces in production and development, and an OpenAI control using the same deployment, input, credentials, and data successfully invokes the tool.
The thread later quotes GCP Support saying the issue came from the backend and describing an unintended ablation of the tool-call infrastructure flow. That is meaningful provider-support evidence inside a public technical report, not the same thing as an official incident page. The report shows a product and grounding failure with a dangerous false-completion shape. It does not show a broad change in model weights, prevalence across Gemini, or a fix.
Discovery stopped before execution
initialize and tools/list completed; tools/call did not appear.
Prose looked complete
Max reportedly wrote a substantive answer without fetching the private data it referenced.
Backend is the lead
The support reply points to an orchestration path, but no public fix history or incident ID is shown.
2. A second Gemini path fails inside the tool loop
An 8 September forum report describes gemini-3.8-flash Agentic Video on the Developer API Interactions endpoint. Seven of seven different full-length files reached 23 processing_call events and 22 processing_result events before an internal api_error. There was no final JSON and no usage metadata. Follow-up variants added output limits, a seed, low thinking, and other envelope changes; later non-streaming attempts returned HTTP 400 with a too-many-tool-calls error, while Flex returned a high-demand error.
This is a recurring API and orchestration signature, not a universal 23-call ceiling: earlier successful runs exceeded 23 calls. The missing usage record means the cost is unknown, not zero. The report has no provider response and changed settings across attempts, so the safe classification is a narrow tool-budget, structured-output, backend, or media-surface failure that needs replay—not a core capability regression.
3. Codex runtime changed the usable work window
A public Codex issue compares five stable Pro sessions on CLI 0.147/0.148 with five sessions on 0.150 for high-reasoning workloads. The local telemetry reports 0.499 million tokens per minute versus 0.662 million, and 3.515 model steps per minute versus 4.610. Observed allowance burn rises from 3.027 to 4.740 percentage points per hour, while cache ratios stay close. In a matched approximately 32.8 million-token pair, the newer runtime finishes about 43% faster with similar token totals; a second approximately 22 million-token pair shows a similar speed difference.
That supports a runtime and usable-window change. It does not measure whether the answers improved or worsened, and the author does not claim to know the server’s exact quota formula. A same-account issue adds local token reconciliation across Astra and Sol but still cannot map the records to the provider’s allowance meter. Current Reddit threads describe fast depletion, stuck sessions, and unfinished tasks, with counterreports mixed in.
4. Claude is a surface story today
Claude’s official status APIs show a minor service outage because Claude Cowork is in partial outage. The component snapshot lists claude.ai, Console, Claude API, and Claude Code as operational. An identified Windows Cowork incident remains unresolved: after a Windows update, the workspace cannot reach the local drive while chat and file operations mostly work.
This is high-confidence availability evidence for one component. It is not a model-quality measurement. A user can see a capable model and still lose the environment required to finish work.
5. What the evidence says about the original question
| Question | Current answer | Confidence and limit |
|---|---|---|
| Are popular model weights broadly worse? | Not proven. No matched, cross-provider longitudinal capability series clears that threshold. | Low confidence in any broad decline call because public samples are narrow and unmatched. |
| Are real user workflows failing? | Yes, on specific paths. Gemini MCP grounding, Gemini Agentic Video processing, Claude Cowork Windows access, and Codex usable windows each have concrete signals. | Medium to high by lane; each has different evidence and scope. |
| Is Gemini’s new evidence stronger than an anecdote? | Yes, for a narrow backend/tool hypothesis. The MCP report exposes lifecycle stages, an isolation test, a control, and a quoted support explanation. | Still one public workflow; no public fix, prevalence, or independent replay. |
| Did Codex quality decline? | Not measured. Runtime throughput and allowance burn changed in the public comparison. | Local telemetry, five sessions per cohort, no answer-quality metric or server ledger. |
6. The practical call
If a model feels worse, label the failure before naming the cause: different, blocked, ungrounded, expensive, unavailable, or still unclear. For a tool report, preserve the task and schema, record initialize → list → call → result → final, and check whether the final answer is grounded in the tool result. For a runtime or quota complaint, record the version, step rate, token rate, retries, transport, bucket, allowance, and completed outcome. Replay the smallest safe case on a stable surface before switching.
What would change the call?
An independent Gemini MCP replay after the reported backend change, across accounts and at least one additional surface, would move the backend hypothesis toward a confirmed incident or show it was fixed. A matched Gemini Agentic Video replay with usage metadata and stable settings would locate the tool-loop boundary. A Codex server ledger joined to runtime traces and matched task outcomes would separate quota mechanics from quality. A fixed broad task set showing the same model and context gap across surfaces would be needed to call a general capability regression. None of those conditions is met today.
Primary sources
- Gemini Deep Research MCP report and quoted support update · Gemini Agentic Video report
- Gemini matched-context tool-use report · Google model documentation
- Claude status API · Claude components API · Claude unresolved incidents API
- Codex runtime throughput comparison · Codex same-day quota report · Current Codex community report
- ProcessRaven agent-cost report · The Context Company observability comparison
The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.