Shaduf.
AI Model Degradation Watch/A tool call can fail without the model getting dumber
Research report — Gemini tool failures and usable windows
Research report · 12 September 2026 · 12:00 UTC
NO BROAD CORE-CAPABILITY DECLINE PROVEN

A tool call can fail without the model getting dumber.

Fresh public evidence adds a concrete Gemini backend failure, a separate Gemini tool-loop failure, and controlled Codex runtime data. The user impact is real. The broad model-decline claim still does not clear its evidence threshold.

Question: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Verdict: no broad cross-model capability decline is proven. The strongest new event is a Google AI Developers Forum report in which Gemini Deep Research discovers a remote MCP tool but never calls it; the thread quotes GCP Support attributing that failure to a backend change. A separate Gemini 3.8 Agentic Video workflow repeatedly dies inside the tool loop. Codex telemetry points to faster runtime consumption and shorter usable allowance windows. These are narrow system failures and access costs, not evidence that model weights broadly deteriorated.

1. Gemini can discover a tool and still not use it

A 4 September Google AI Developers Forum report describes a production regression affecting deep-research-max-preview-04-2026. Gemini connects to a remote Streamable HTTP MCP server, completes initialize, receives tools/list, and then sends no tools/call. Max nevertheless produces a substantive report that refers to private source data it never retrieved. In an MCP-only isolation test, the standard preview completes with thought but no final report. The same workflow reportedly worked on 2 September, reproduces in production and development, and an OpenAI control using the same deployment, input, credentials, and data successfully invokes the tool.

The thread later quotes GCP Support saying the issue came from the backend and describing an unintended ablation of the tool-call infrastructure flow. That is meaningful provider-support evidence inside a public technical report, not the same thing as an official incident page. The report shows a product and grounding failure with a dangerous false-completion shape. It does not show a broad change in model weights, prevalence across Gemini, or a fix.

Observed

Discovery stopped before execution

initialize and tools/list completed; tools/call did not appear.

User consequence

Prose looked complete

Max reportedly wrote a substantive answer without fetching the private data it referenced.

Boundary

Backend is the lead

The support reply points to an orchestration path, but no public fix history or incident ID is shown.

2. A second Gemini path fails inside the tool loop

An 8 September forum report describes gemini-3.8-flash Agentic Video on the Developer API Interactions endpoint. Seven of seven different full-length files reached 23 processing_call events and 22 processing_result events before an internal api_error. There was no final JSON and no usage metadata. Follow-up variants added output limits, a seed, low thinking, and other envelope changes; later non-streaming attempts returned HTTP 400 with a too-many-tool-calls error, while Flex returned a high-demand error.

This is a recurring API and orchestration signature, not a universal 23-call ceiling: earlier successful runs exceeded 23 calls. The missing usage record means the cost is unknown, not zero. The report has no provider response and changed settings across attempts, so the safe classification is a narrow tool-budget, structured-output, backend, or media-surface failure that needs replay—not a core capability regression.

3. Codex runtime changed the usable work window

A public Codex issue compares five stable Pro sessions on CLI 0.147/0.148 with five sessions on 0.150 for high-reasoning workloads. The local telemetry reports 0.499 million tokens per minute versus 0.662 million, and 3.515 model steps per minute versus 4.610. Observed allowance burn rises from 3.027 to 4.740 percentage points per hour, while cache ratios stay close. In a matched approximately 32.8 million-token pair, the newer runtime finishes about 43% faster with similar token totals; a second approximately 22 million-token pair shows a similar speed difference.

That supports a runtime and usable-window change. It does not measure whether the answers improved or worsened, and the author does not claim to know the server’s exact quota formula. A same-account issue adds local token reconciliation across Astra and Sol but still cannot map the records to the provider’s allowance meter. Current Reddit threads describe fast depletion, stuck sessions, and unfinished tasks, with counterreports mixed in.

Do not collapse cost into quality. A shorter Pro work window can be commercially painful even if answer quality is unchanged. Record model, runtime, steps, tokens, retries, cache, transport, allowance, final artifact, and task outcome as separate columns.

4. Claude is a surface story today

Claude’s official status APIs show a minor service outage because Claude Cowork is in partial outage. The component snapshot lists claude.ai, Console, Claude API, and Claude Code as operational. An identified Windows Cowork incident remains unresolved: after a Windows update, the workspace cannot reach the local drive while chat and file operations mostly work.

This is high-confidence availability evidence for one component. It is not a model-quality measurement. A user can see a capable model and still lose the environment required to finish work.

5. What the evidence says about the original question

QuestionCurrent answerConfidence and limit
Are popular model weights broadly worse?Not proven. No matched, cross-provider longitudinal capability series clears that threshold.Low confidence in any broad decline call because public samples are narrow and unmatched.
Are real user workflows failing?Yes, on specific paths. Gemini MCP grounding, Gemini Agentic Video processing, Claude Cowork Windows access, and Codex usable windows each have concrete signals.Medium to high by lane; each has different evidence and scope.
Is Gemini’s new evidence stronger than an anecdote?Yes, for a narrow backend/tool hypothesis. The MCP report exposes lifecycle stages, an isolation test, a control, and a quoted support explanation.Still one public workflow; no public fix, prevalence, or independent replay.
Did Codex quality decline?Not measured. Runtime throughput and allowance burn changed in the public comparison.Local telemetry, five sessions per cohort, no answer-quality metric or server ledger.

6. The practical call

If a model feels worse, label the failure before naming the cause: different, blocked, ungrounded, expensive, unavailable, or still unclear. For a tool report, preserve the task and schema, record initializelistcallresultfinal, and check whether the final answer is grounded in the tool result. For a runtime or quota complaint, record the version, step rate, token rate, retries, transport, bucket, allowance, and completed outcome. Replay the smallest safe case on a stable surface before switching.

What would change the call?

An independent Gemini MCP replay after the reported backend change, across accounts and at least one additional surface, would move the backend hypothesis toward a confirmed incident or show it was fixed. A matched Gemini Agentic Video replay with usage metadata and stable settings would locate the tool-loop boundary. A Codex server ledger joined to runtime traces and matched task outcomes would separate quota mechanics from quality. A fixed broad task set showing the same model and context gap across surfaces would be needed to call a general capability regression. None of those conditions is met today.

Primary sources

The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.

Search published pools, pages, reports, and evidence.