Shaduf.
AI Model Degradation Watch/One narrow Gemini tool-use regression joins the ledger; broad decline remains unproven
Research report — Gemini tool-use lead
Research report · 11 September 2026 · 12:00 UTC

One narrow Gemini tool-use regression joins the ledger. Broad decline still fails.

A matched-context API report changes the evidence mix. It is stronger than an anecdote and narrower than a model-wide verdict. The same run also found Claude surface incidents and new Astra/Codex cost and transport leads, making usable access the more honest product frame.

Question: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Verdict: no broad cross-model capability decline is proven. The strongest new signal is a narrow Gemini 3.8 self-initiated tool-selection lead under reflective context. It needs independent replication. Separate provider-confirmed surface incidents and account-level economics now explain many ways a capable model can feel worse.

1. The new capability-shaped lead

A 5 September Google AI Developers Forum report holds the task, save_note tool, structured schema, low thinking setting, and API setup constant. On a 2,199-character reflective user message, Gemini 3.5 Flash calls the tool in 8/8 runs while Gemini 3.8 Flash calls it in 0/8. Longer system guidance raises 3.8 to 3/8; removing prior reasoning raises it from 0/8 to 6/8; explicit “when asked” tool use remains 15/15. This is an unusually useful diagnostic shape because the report exposes repeated cells and ablations.

It still does not prove a general Gemini regression. It is one author, one API configuration, no prevalence study, no independent runner, and no provider confirmation. The behavior may be a model/API/context interaction; consumer surfaces may route or prompt differently. Google’s current documentation describes 3.8 as generally available, long-context, and designed for tool orchestration, while aggregate Cloud status showed no broad severe incident. Those are context and availability facts, not a refutation of the narrow lead.

2. Confirmed product and surface incidents

Claude’s official status page currently lists degraded Claude Cowork on Windows: after a Windows update, the workspace cannot reach the computer drive, while chat and file operations mostly work. A separate incident involving elevated Claude API latency for some US Midwest traffic was resolved on 10 September. This is confirmed product reliability evidence. It does not measure whether Claude’s model weights got worse.

OpenAI’s aggregate status is separate from its recent usage-limit and Project incidents already in the ledger. The distinction matters: a work container, client transport, or entitlement can fail while the underlying model remains capable.

3. Astra/Codex: cost and transport are now first-class lanes

  1. Transport: an open Codex report says Astra CLI sessions can loop on WebSocket reconnects, consume purchased credits, then progress after HTTPS fallback. This is a narrow client and cost lead.
  2. Rate bucket: another report says concurrent Astra threads on one account can receive different rate-limit buckets, with one substitute appearing empty while another bucket remains available. This is a session-scoped routing or entitlement lead, not account-wide model evidence.
  3. Published economics: OpenAI says shared Work/Codex allowance use varies with model, task, input/output, reasoning effort, Fast mode, and multistep work. Its rate card lists Astra output at 1,250 credits per million tokens and Fast at 2.5x. That makes cost mechanisms plausible; it does not reconcile any individual meter.
  4. Cross-provider economics: an Anthropic issue reports roughly 2.04x Fable 5.1 output tokens per turn and a rapid Max-plan meter movement on one stable workload. It is an account-level economics lead, not proof of lower capability.

4. Counterweights

Bug Hunt Bench’s 10 September page reports 105 planted bugs across 63 runs and 23 models. Its displayed maxima are Astra 48/105, Fable 5.1 43/105, and Sol 42/105. This is a useful independent counterweight, but it is cross-sectional: effort, harness, date, and run selection differ.

Beyond Benchmarks’ latest displayed 28-day Astra panel shows 93.1% meaningful outcomes, 0.1% frustration, 14.8 output tokens per second, $43.10 per million tokens modeled cost, $28.74 per active hour, 0.19 interruptions per session, 1.68x context hunting, and 5.1% failed tool calls. The values moved from the prior snapshot, but rolling task mix prevents a trend claim. Marginlab’s Codex tracker remains nominal with non-significant displayed movement; Claude Code is still collecting a baseline.

5. The practical call

If a popular model feels worse, the defensible first labels remain different, blocked, expensive, unavailable, or still unclear. For a tool-use complaint, run a short versus reflective context, preserve the tool schema and effort, count repeated tool calls, inspect final model identity and artifacts, and compare API with the affected product surface. For a quota complaint, capture transport, retries, output tokens, rate bucket, plan, and meter before asking whether the server reconciles them.

What would change the call?

Independent Gemini replication across endpoints, accounts, and a consumer or Vertex surface would move the narrow lead toward a confirmed surface/model interaction. A documented provider fix or time-series drop-and-recovery would strengthen attribution. A fixed task that shows the same model and context gap across surfaces would support a capability claim. For Astra and Fable, provider-side account reconciliation could close the economics lane. Stable field slices or a matched benchmark time series could challenge the broad no-decline call. Today none of those conditions is met.

Primary sources

The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.

Search published pools, pages, reports, and evidence.