Are popular AI models getting worse right now?
The maintained answer and the strongest public evidence available at the last completed review.
The public record does not establish a broad decline across the monitored models.
The completed reviews found real service incidents, model-routing changes, price changes, access restrictions, safeguard differences, and product-surface instability. Those events can make a model feel worse without demonstrating that its underlying capability declined.
Public comparisons provide useful cross-model snapshots, but the available records do not form a controlled longitudinal measurement across identical prompts, model versions, inference settings, regions, and product surfaces.
The last completed review therefore keeps the broad answer unchanged: degradation is not established. Individual providers and surfaces still require continued monitoring.
Uncertainty
- Providers can change routing, system behavior, safeguards, and serving configuration without exposing a fixed public snapshot.
- Independent benchmarks often compare different effort levels or hosted configurations and are not a clean degradation time series.
- The most recent completed public-source cutoff is 6 August 2026; a new scheduled refresh is required.
What the current answer means
The pool is not claiming that model behavior is stable everywhere. It is separating several different phenomena that are often collapsed into “the model got worse”: service availability, routing, product defaults, safeguards, access tiers, price changes, and underlying capability.
The public evidence reviewed through 6 August establishes multiple incidents and product changes. It does not establish a comparable, controlled decline across the monitored model families.
Where to look next
Research history preserves the full dated reports. The model table separates provider surfaces, while Evidence exposes the public records behind the current answer.