Shaduf
PoolAI Model Degradation Watch
Public research pool

6 August 2026 public-source review

6 August 2026 public-source review. Latest completed cross-model review; broad degradation remained unestablished.

Shaduf AI Model Degradation Watch

Publication cutoff: 2026-08-06 UTC. Facts first published after this date are excluded. Mutable pages without versioned historical snapshots are identified as current observations, not historical proof.

Coverage: 10 monitored rows: three GPT-5.6 tiers, three Claude rows, one Grok row, and exactly three open-weight/open-source models.

This report uses public vendor documentation, public status histories, independent public comparisons, and attributable community reporting. It contains no Shaduf-run evaluations, prompts, canaries, benchmarks, or experiments.

Current summary

The public record does not establish broad base-model degradation for any monitored model. It does establish rapid changes in model availability, routing, safeguards, fallback behavior, and infrastructure reliability.

Anthropic provides the clearest recent evidence of operational instability. Its status history records elevated errors affecting Claude Mythos 5, Fable 5, Opus 5, and Sonnet 5 on August 5, with additional multi-model and login incidents on August 3–4. These are availability and error-rate incidents, not evidence that model weights became less capable. The status page reported no new incident on August 6 at the time observed. (Claude Status

Claude Fable 5 and Claude Mythos 5 must be monitored as separate service configurations, but not treated as separate base models. Anthropic explicitly states that they share the same underlying model. Fable adds safety classifiers and may fall back to Opus 4.8; Mythos is restricted through Project Glasswing with selected safeguards lifted. (Anthropic launch announcement — 2026-06-09, current model overview

OpenAI’s GPT-5.6 family reached general availability on July 9. Sol, Terra, and Luna are distinct API tiers, but their product exposure differs materially: standard ChatGPT exposes Sol, while Terra and Luna are available through ChatGPT Work, Codex, and the API. ChatGPT may continue with GPT-5.4 Thinking mini after a GPT-5.6 reasoning allowance is reached. OpenAI also recorded a July 9 model-selection capacity incident and a July 17 Codex Sol server-overload incident. (OpenAI GA announcement — 2026-07-09, OpenAI ChatGPT availability documentation, July 17 status incident

Grok 4.5 has a verified API identifier, grok-4.5, and is documented for Grok Build, Cursor, and Office add-ins. The public consumer documentation confirms Grok on the web, iOS, Android, and X, but does not establish that those surfaces use the API model ID unchanged. The current xAI status page reported no declared incident, but it does not expose a comparable model-specific historical incident record. (Grok 4.5 API documentation — updated 2026-07-17, Grok consumer documentation — updated 2026-06-29, xAI status

The three open-weight selections are Kimi K3, GLM-5.2, and DeepSeek V4 Pro. Their weights and licenses are publicly verifiable. Kimi K3 uses a custom Kimi K3 License and is therefore best described as open-weight rather than straightforward OSI open-source software. GLM-5.2 and DeepSeek V4 Pro are published under MIT terms. (Kimi K3 repository, GLM-5.2 official release, DeepSeek V4 release

The most useful conclusion for degradation monitoring is therefore dimensional separation:

  • base-model capability;
  • availability and error rates;
  • latency and capacity;
  • routing, aliases, and fallback;
  • system instructions and safeguards;
  • context handling and compaction;
  • connected tools and product wrappers.

A public benchmark snapshot can describe capability at one point in time. It cannot, by itself, establish degradation.

Monitored model rows

There are ten monitored rows, but only nine independently identifiable base-model families because Fable 5 and Mythos 5 share underlying weights. Sonnet 5 is a current Claude model, but it is not added as an eleventh row; the three Claude slots are Fable, Mythos, and the next distinct current model, Opus 5.

#Monitored rowOfficial identifier and model factsExplicit surfacesChanges, incidents, and routingIndependent public comparison
1GPT-5.6 Sol — distinct API tiergpt-5.6-sol; the gpt-5.6 alias routes to Sol. Current API documentation lists a 1.05M-token context window and 128K maximum output. (API documentation, accessed 2026-08-06API: direct model ID. Standard ChatGPT: Sol powers Medium, High, and Extra High; Sol Pro is a separate ChatGPT plan variant. ChatGPT Work/Codex: available subject to plan and workspace controls.Previewed June 26 and launched GA July 9. OpenAI status recorded model-selection capacity errors July 9 and Codex Sol server-overload errors July 17. ChatGPT can fall back to GPT-5.4 Thinking mini after a reasoning limit. (Preview — 2026-06-26, GA — 2026-07-09, ChatGPT surface rulesAA Intelligence Index snapshot: 58.9. AA AutomationBench snapshot: 51.2. These are hosted public evaluations, not raw-weight measurements. (AA/BenchLM snapshot, AA AutomationBench snapshot
2GPT-5.6 Terra — distinct API tiergpt-5.6-terra; documented as the balanced intelligence/cost tier. 1.05M context, 128K maximum output. (API documentation, accessed 2026-08-06API: direct model ID. Standard ChatGPT: not selectable. ChatGPT Work/Codex: available on eligible plans; Free/Go Codex access is Terra.Released with the GPT-5.6 GA family on July 9. No reviewed OpenAI status entry specifically names Terra; that absence does not establish uninterrupted availability or unchanged quality.AA Intelligence Index: 55.0. AA AutomationBench: 45.6.
3GPT-5.6 Luna — distinct API tiergpt-5.6-luna; documented as the cost-sensitive, high-volume tier. 1.05M context, 128K maximum output. (API documentation, accessed 2026-08-06API: direct model ID. Standard ChatGPT: not selectable. ChatGPT Work/Codex: available on eligible plans.Released with the GPT-5.6 GA family on July 9. No reviewed OpenAI status entry specifically names Luna. Product limits, reasoning settings, safeguards, and fallback behavior remain surface-dependent.AA Intelligence Index: 51.2. AA AutomationBench: 42.2.
4Claude Fable 5 — safeguarded configurationclaude-fable-5; 1M context and 128K maximum output. Anthropic describes it as its most capable widely released model. (Current model overview, accessed 2026-08-06API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry. Anthropic also restored Fable to Claude.ai, Claude Code, and Claude Cowork globally from July 1.Launch: June 9. Fable includes safety classifiers; flagged requests may be handled by Opus 4.8 instead. Anthropic reported over 95% of Fable sessions had no fallback in early data. Status incidents explicitly named Fable on July 25, July 30, and August 5. (Launch — 2026-06-09, redeployment — 2026-06-30/update 2026-07-01, statusAA Intelligence Index: 59.9. AA AutomationBench: 48.6. The effective score may include documented fallback behavior; it should not be treated as an unwrapped Fable-weight score.
5Claude Mythos 5 — restricted same-weight configurationclaude-mythos-5; Anthropic says Mythos shares Fable’s underlying model, specifications, and pricing. (Fable/Mythos documentation, accessed 2026-08-06Project Glasswing only: invitation-based access for approved organizations. No general self-serve access.Launch: June 9. Mythos has selected Fable safeguards lifted, initially especially cyber safeguards. Access was suspended for both Fable and Mythos June 12 and partially restored July 1. Status explicitly named Mythos in the August 5 multi-model incident.No standalone current AA score was published for Mythos in the reviewed comparable snapshot. Fable’s score must not be copied to Mythos without preserving the safeguard and fallback distinction.
6Claude Opus 5 — distinct current Opus-tier modelclaude-opus-5; 1M context, 128K maximum output, thinking enabled by default. (Opus 5 documentation, accessed 2026-08-06API, Amazon Bedrock, Google Cloud, Microsoft Foundry. Product-surface availability is plan- and routing-dependent.Launched July 24. Changes include default thinking, five effort levels, mid-conversation tool changes, and beta server-side fallback. Status recorded Opus 5 incidents July 26–27, July 30, and August 5. (Release notes — 2026-07-24, statusAA Intelligence Index: 60.7 in the August 5 snapshot. Artificial Analysis separately reported a rounded score of 61 on July 24; this is a different dated presentation, not evidence of a change between 60.7 and 61. (AA article — 2026-07-24, current snapshot
7Grok 4.5 — proprietary modelgrok-4.5; 500K context, configurable low/medium/high reasoning, high by default. (API documentation, updated 2026-07-17API: direct grok-4.5. Grok consumer: grok.com, iOS, Android. X: Grok in X. Coding/product: Grok Build default model, Cursor, Office add-ins. The exact consumer/X serving identifier is not publicly established.API release-note entry: July 8; public launch announcement: July 16; EU API availability: July 17. API supports web search, X search, code execution, and function calling. The status page reported no declared incident on August 6. (Announcement — 2026-07-16, release notes, statusAA Intelligence Index snapshot: 53.8. AA AutomationBench: 51.4. Artificial Analysis separately displayed a rounded 54 for Grok 4.5 high; endpoint and snapshot differences prevent treating these as a time series.
8GLM-5.2 — open-source/open-weightOfficial weight repository: zai-org/GLM-5.2. Z.ai states MIT licensing, public Hugging Face/ModelScope weights, and local support through Transformers, vLLM, SGLang, and other frameworks. (Z.ai release — 2026-06-16, Hugging Face repositoryLocal weights: exact repository ID. Hosted: Z.ai and related hosted surfaces. Hosted API routing is not equivalent to self-hosting the weights.Official release: June 16. No dedicated public incident history for the downloadable weights was identified.AA Intelligence Index: 51.1. Vellum HLE: 54.7; GPQA Diamond: 91.2. Vellum combines provider-reported and independently run/community results. (Vellum leaderboard, updated 2026-07-24
9DeepSeek V4 Pro — open-source/open-weightWeight repository: deepseek-ai/DeepSeek-V4-Pro; Hugging Face lists MIT licensing and local Transformers/vLLM deployment. Official API ID: deepseek-v4-pro. (DeepSeek release — 2026-04-24, Hugging Face repository, API model listLocal weights: Hugging Face/ModelScope-compatible deployment. API: deepseek-v4-pro; official API also supports an Anthropic-compatible endpoint.DeepSeek deprecated deepseek-chat and deepseek-reasoner on July 24; those legacy names routed to V4 Flash rather than V4 Pro. This is an alias/routing change, not evidence of V4 Pro degradation. (Official change logAA Intelligence Index: 44.3 for the V4 Pro Max row. Vellum HLE: 48.2; GPQA Diamond: 90.1.
10Kimi K3 — open-weight, custom licenseWeight repository: moonshotai/Kimi-K3; 2.8T total parameters, 104B activated, 1,048,576-token context, native vision. Full weights are published under the Kimi K3 License. (Official Kimi page, Hugging Face repository, licenseLocal weights: Hugging Face, vLLM, SGLang, Docker. Hosted: Kimi.com, Kimi Work, Kimi Code, and Kimi API are documented. The exact public API slug for K3 is not established by the current public model-list page inspected.Kimi announced the model in July; the official page promised full weights by July 27, and the repository is available by the cutoff. Kilo’s July 21 launch report described capacity strain and 429s as serving conditions, explicitly not model-quality evidence. (Kimi announcement, Kilo report — 2026-07-21AA Intelligence Index: 57.1. Vellum HLE: 56.0; GPQA Diamond: 93.5. Arena.ai Frontend Code snapshot reported by Kilo: 1,679.

Dated change and incident timeline

DateEventEvidence and interpretation
April 2026Claude Mythos Preview began through Project Glasswing.Anthropic’s June 9 announcement identifies April as the first Mythos-class release. (Anthropic — 2026-06-09
2026-04-24DeepSeek V4 Preview released with Pro and Flash variants, open weights, and 1M context.Official DeepSeek release documentation. (DeepSeek — 2026-04-24
2026-06-09Claude Fable 5 and Mythos 5 launched.Same underlying model; Fable safeguards and fallback; Mythos restricted access. (Anthropic — 2026-06-09
2026-06-12Anthropic suspended Fable and Mythos access globally after a U.S. government directive.Access disruption, not a reported base-model quality change. (Anthropic statement — 2026-06-12
2026-06-16GLM-5.2 released with MIT licensing and public weights.Official Z.ai release. (Z.ai — 2026-06-16
2026-06-26GPT-5.6 preview began with Sol, Terra, and Luna.API/Codex preview with phased access and differentiated safeguards. (OpenAI — 2026-06-26
2026-06-30 / 2026-07-01Fable and Mythos access restored after export controls were lifted.Fable restored globally across Claude Platform, Claude.ai, Claude Code, and Cowork; Mythos restored to selected U.S. organizations. (Anthropic — 2026-06-30/update 2026-07-01
2026-07-08 / 2026-07-16Grok 4.5 API release-note entry and public announcement.The API entry is dated July 8; the public launch article is dated July 16. (Grok release notes, announcement — 2026-07-16
2026-07-09GPT-5.6 general availability.Same day, OpenAI status recorded multi-model selection-capacity errors. (GA announcement, status incident
2026-07-16Kimi K3 launch reported; Arena.ai Frontend Code snapshot placed it above Fable and Sol.Kilo’s July 21 report dates the leaderboard snapshot to July 16. This is a narrow frontend preference measure. (Kilo — 2026-07-21
2026-07-17Codex GPT-5.6 Sol server-overload incident; Grok 4.5 EU API availability.Availability/routing events, not capability measurements. (OpenAI status, Grok release notes
2026-07-24Claude Opus 5 launched.New distinct model; thinking by default, 1M context, 128K output, fallback and tool-change features. (Claude release notes — 2026-07-24
2026-07-25 to 2026-08-05Repeated Claude availability incidents.Status history names Fable, Mythos, Opus 5, Sonnet 5, and multi-model errors across separate dates. (Claude Status
2026-07-30OpenAI announced GPT-5.6 Terra and Luna price reductions.Commercial change; no weight change was announced. (OpenAI GA page, July 30 update
2026-08-05Latest public AA/BenchLM snapshot verified.Provides a current cross-model comparison, not a longitudinal degradation measurement. (AA Intelligence Index snapshot, AA AutomationBench snapshot
2026-08-06Cutoff date.Claude status displayed no new incident for August 6; xAI status displayed no declared incident at observation time. (Claude Status, xAI Status

Chart 1 — Artificial Analysis Intelligence Index snapshot

Data verified: 2026-08-05. Interpretation: a public hosted-model snapshot on the source’s displayed scale, not a success percentage and not a degradation trend. Mythos is omitted because no standalone comparable score was published.

Claude Opus 5       60.7 |██████████████████████████████
Claude Fable 5      59.9 |█████████████████████████████
GPT-5.6 Sol         58.9 |████████████████████████████
Kimi K3             57.1 |███████████████████████████
GPT-5.6 Terra       55.0 |██████████████████████████
Grok 4.5            53.8 |█████████████████████████
GPT-5.6 Luna        51.2 |████████████████████████
GLM-5.2             51.1 |████████████████████████
DeepSeek V4 Pro     44.3 |█████████████████████

The underlying source labels Kimi K3 as “Closed,” which conflicts with Moonshot’s official public weights and license. The score is retained; the openness classification is taken from the official Kimi repository instead. (BenchLM display-only mirror, Kimi official repository

model,score,unit,source_url,data_verified_date,configuration_note
Claude Opus 5,60.7,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Public AA model row; exact effort/surface not exposed in mirror
Claude Fable 5,59.9,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Effective hosted configuration; Anthropic documents possible Opus 4.8 fallback
GPT-5.6 Sol,58.9,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Public AA model row; exact surface/configuration not exposed in mirror
Kimi K3,57.1,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Hosted public row; official weights are open despite source classification label
GPT-5.6 Terra,55.0,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Public AA model row
Grok 4.5,53.8,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Public AA model row; exact consumer/X surface not established
GPT-5.6 Luna,51.2,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Public AA model row
GLM-5.2,51.1,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Hosted public row; not a local self-hosting measurement
DeepSeek V4 Pro (Max),44.3,AA Intelligence Index score,https://benchlm.ai/benchmarks/artificialanalysis,2026-08-05,Hosted public row; Max configuration label retained

Chart 2 — Selected open-weight models on Vellum public comparisons

Vellum’s page was updated July 24, 2026 and states that its data combines provider-reported results with evaluations by Vellum and the open-source community. Scores are comparable within each metric, not across metrics.

Humanity's Last Exam

Kimi K3            56.0 |████████████████████████████
GLM-5.2            54.7 |███████████████████████████
DeepSeek V4 Pro    48.2 |████████████████████████

GPQA Diamond

Kimi K3            93.5 |███████████████████████████████████████████████
GLM-5.2            91.2 |█████████████████████████████████████████████
DeepSeek V4 Pro    90.1 |████████████████████████████████████████████
model,metric,score,unit,source_url,source_updated_date
Kimi K3,Humanity's Last Exam,56.0,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24
GLM-5.2,Humanity's Last Exam,54.7,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24
DeepSeek V4 Pro,Humanity's Last Exam,48.2,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24
Kimi K3,GPQA Diamond,93.5,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24
GLM-5.2,GPQA Diamond,91.2,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24
DeepSeek V4 Pro,GPQA Diamond,90.1,percent,https://www.vellum.ai/open-llm-leaderboard,2026-07-24

Chart 3 — Frontend Code Arena snapshot

This is a narrow human-preference leaderboard, not a general capability measure.

Kimi K3             1679 |████████████████████████████████
Claude Fable 5      1631 |██████████████████████████████
GPT-5.6 Sol         1618 |█████████████████████████████

The snapshot was dated July 16 in Kilo’s July 21 report.

model,metric,score,source_url,leaderboard_date,publication_date
Kimi K3,Frontend Code Arena,1679,https://blog.kilo.ai/p/kimi-k3,2026-07-16,2026-07-21
Claude Fable 5,Frontend Code Arena,1631,https://blog.kilo.ai/p/kimi-k3,2026-07-16,2026-07-21
GPT-5.6 Sol,Frontend Code Arena,1618,https://blog.kilo.ai/p/kimi-k3,2026-07-16,2026-07-21

What is established

  • GPT-5.6 Sol, Terra, and Luna are official, separately named model tiers with separate API identifiers.
  • Standard ChatGPT, ChatGPT Work, Codex, and the OpenAI API do not expose those tiers identically.
  • Claude Fable 5 and Mythos 5 share underlying weights but differ in safeguards and access.
  • Claude model IDs from the current generation are pinned snapshots; Anthropic states that serving infrastructure, safety classifiers, routers, and sampling logic can still change around fixed weights. (Anthropic model-ID documentation
  • Anthropic recorded multiple recent model-specific and multi-model availability incidents.
  • OpenAI recorded GPT-5.6 model-selection and Codex Sol availability incidents.
  • Grok 4.5 has the verified API ID grok-4.5; exact consumer and X routing is not publicly established.
  • GLM-5.2, DeepSeek V4 Pro, and Kimi K3 have publicly downloadable weights and published licenses.
  • Kimi K3 is open-weight under a custom license, not simply MIT.
  • Public independent snapshots place Opus 5, Fable 5, GPT-5.6 Sol, and Kimi K3 near the same frontier band, but the scores represent particular hosted configurations and evaluation conditions.

What is not established

  • No public fixed-surface, longitudinal evidence reviewed here establishes that any monitored base model became less capable by August 6.
  • A status incident does not establish a model-quality regression.
  • A user complaint, 429, latency increase, refusal, or fallback does not establish degraded model weights.
  • Fable 5’s public score cannot be transferred directly to Mythos 5 because safeguards, access, and fallback differ.
  • Consumer ChatGPT, Claude.ai, Grok, X, Codex, Claude Code, and hosted API endpoints cannot be assumed to use identical system instructions, tool environments, routing, or effort settings.
  • A hosted open-weight score does not establish the behavior of a local deployment of the same weights.
  • The absence of a provider status entry does not prove the absence of an incident.
  • A single cross-model ranking does not establish a universal capability ordering.

Uncertainty

The principal uncertainty is configuration identity. Public comparisons often combine a model name with effort settings, fallback models, provider harnesses, tools, and different context policies. For example, Anthropic explicitly documents Fable fallback to Opus 4.8, while the Kimi technical materials compare Fable with fallback enabled. (Kimi technical repository

Artificial Analysis pages also show small numerical differences across dated or rounded presentations: the August 5 display-only snapshot lists Opus 5 at 60.7 and Grok 4.5 at 53.8, while separate July pages display rounded values of 61 and 54. These differences are not evidence of degradation or improvement without a controlled longitudinal comparison.

Kimi K3’s official announcement confirms availability through Kimi API, but the current public model-list documentation inspected does not expose a K3 API slug. The weight-repository identifier moonshotai/Kimi-K3 is therefore asserted; an API identifier is not invented.

The current pages for OpenAI, Anthropic, xAI, and hosted open-weight services are mutable. A current page can document present behavior while failing to preserve the historical state that existed on an earlier date.

Freshness

  • Cutoff: 2026-08-06 UTC.
  • The latest comparison snapshot used here was verified 2026-08-05.
  • Claude status was observed with entries through 2026-08-06.
  • xAI status was observed on 2026-08-06 with no declared incident.
  • Static release dates used in the timeline are no later than 2026-07-30, except for current status and current documentation observations on August 6.
  • Vellum’s open-weight page was updated 2026-07-24.
  • Kilo’s Frontend Code report was published 2026-07-21 and refers to a July 16 leaderboard snapshot.
  • Pages without a stable historical version are treated as current observations only.

Sources

Official model and release documentation

Official weights and licenses

Status and incident histories

Independent and attributable public comparisons

Search Shaduf

Search published pools, pages, reports, and evidence.