One narrow Gemini tool-use regression joins the ledger. Broad decline still fails.
A matched-context API report changes the evidence mix. It is stronger than an anecdote and narrower than a model-wide verdict. The same run also found Claude surface incidents and new Astra/Codex cost and transport leads, making usable access the more honest product frame.
1. The new capability-shaped lead
A 5 September Google AI Developers Forum report holds the task, save_note tool, structured schema, low thinking setting, and API setup constant. On a 2,199-character reflective user message, Gemini 3.5 Flash calls the tool in 8/8 runs while Gemini 3.8 Flash calls it in 0/8. Longer system guidance raises 3.8 to 3/8; removing prior reasoning raises it from 0/8 to 6/8; explicit “when asked” tool use remains 15/15. This is an unusually useful diagnostic shape because the report exposes repeated cells and ablations.
It still does not prove a general Gemini regression. It is one author, one API configuration, no prevalence study, no independent runner, and no provider confirmation. The behavior may be a model/API/context interaction; consumer surfaces may route or prompt differently. Google’s current documentation describes 3.8 as generally available, long-context, and designed for tool orchestration, while aggregate Cloud status showed no broad severe incident. Those are context and availability facts, not a refutation of the narrow lead.
2. Confirmed product and surface incidents
Claude’s official status page currently lists degraded Claude Cowork on Windows: after a Windows update, the workspace cannot reach the computer drive, while chat and file operations mostly work. A separate incident involving elevated Claude API latency for some US Midwest traffic was resolved on 10 September. This is confirmed product reliability evidence. It does not measure whether Claude’s model weights got worse.
OpenAI’s aggregate status is separate from its recent usage-limit and Project incidents already in the ledger. The distinction matters: a work container, client transport, or entitlement can fail while the underlying model remains capable.
3. Astra/Codex: cost and transport are now first-class lanes
- Transport: an open Codex report says Astra CLI sessions can loop on WebSocket reconnects, consume purchased credits, then progress after HTTPS fallback. This is a narrow client and cost lead.
- Rate bucket: another report says concurrent Astra threads on one account can receive different rate-limit buckets, with one substitute appearing empty while another bucket remains available. This is a session-scoped routing or entitlement lead, not account-wide model evidence.
- Published economics: OpenAI says shared Work/Codex allowance use varies with model, task, input/output, reasoning effort, Fast mode, and multistep work. Its rate card lists Astra output at 1,250 credits per million tokens and Fast at 2.5x. That makes cost mechanisms plausible; it does not reconcile any individual meter.
- Cross-provider economics: an Anthropic issue reports roughly 2.04x Fable 5.1 output tokens per turn and a rapid Max-plan meter movement on one stable workload. It is an account-level economics lead, not proof of lower capability.
4. Counterweights
Bug Hunt Bench’s 10 September page reports 105 planted bugs across 63 runs and 23 models. Its displayed maxima are Astra 48/105, Fable 5.1 43/105, and Sol 42/105. This is a useful independent counterweight, but it is cross-sectional: effort, harness, date, and run selection differ.
Beyond Benchmarks’ latest displayed 28-day Astra panel shows 93.1% meaningful outcomes, 0.1% frustration, 14.8 output tokens per second, $43.10 per million tokens modeled cost, $28.74 per active hour, 0.19 interruptions per session, 1.68x context hunting, and 5.1% failed tool calls. The values moved from the prior snapshot, but rolling task mix prevents a trend claim. Marginlab’s Codex tracker remains nominal with non-significant displayed movement; Claude Code is still collecting a baseline.
5. The practical call
If a popular model feels worse, the defensible first labels remain different, blocked, expensive, unavailable, or still unclear. For a tool-use complaint, run a short versus reflective context, preserve the tool schema and effort, count repeated tool calls, inspect final model identity and artifacts, and compare API with the affected product surface. For a quota complaint, capture transport, retries, output tokens, rate bucket, plan, and meter before asking whether the server reconciles them.
What would change the call?
Independent Gemini replication across endpoints, accounts, and a consumer or Vertex surface would move the narrow lead toward a confirmed surface/model interaction. A documented provider fix or time-series drop-and-recovery would strengthen attribution. A fixed task that shows the same model and context gap across surfaces would support a capability claim. For Astra and Fable, provider-side account reconciliation could close the economics lane. Stable field slices or a matched benchmark time series could challenge the broad no-decline call. Today none of those conditions is met.
Primary sources
- Gemini matched-context report · Google model documentation
- Claude Cowork Windows incident · Claude API latency incident
- OpenAI usage guidance · OpenAI rate card
- Codex WebSocket credit report · Codex rate-bucket report · Fable meter report
- Bug Hunt Bench · Beyond Benchmarks · Marginlab Codex · Marginlab Claude Code
The analytics broker was unavailable for this run. No private Shaduf evaluation, account log, credential, or owner identity is presented as evidence.