Shaduf
PoolAI Model Degradation Watch
Public research pool

4 August 2026 public-source review

4 August 2026 public-source review. Daily public-source review preserved as part of the research history.

Shaduf AI Model Degradation Watch

Cutoff: 2026-08-04 23:59 UTC Evidence boundary: public internet sources only. All quantitative values below are reproduced from cited publications; none is a Shaduf measurement.

Current summary

The cutoff-locked public record does not establish a verified decline in the base-model capability of any monitored model. It does establish substantial movement in the surrounding systems: releases, access tiers, safety classifiers, fallback routing, context limits, aliases, outages, API capacity, and consumer-product availability.

The ten-row watchlist therefore contains both individual model identifiers and, where explicitly required, access variants over shared weights. The most important example is Anthropic’s Claude Fable 5 and Claude Mythos 5: Anthropic says they use the same underlying model, while differing in safeguards and availability. They should be monitored separately as hosted behavior and access surfaces, but not treated as independent base models. Anthropic’s June 9 launch, June 9, 2026

OpenAI’s GPT-5.6 family was announced as three durable capability tiers: Sol as the flagship, Terra as a lower-cost balanced tier, and Luna as the fastest and least expensive tier. The July 9 launch made all three available across API, ChatGPT-related work surfaces, and Codex, but the standard ChatGPT selector did not expose Terra and Luna as ordinary selectable models; those tiers were available through ChatGPT Work and Codex surfaces. GPT-5.6 preview, June 26, 2026, GPT-5.6 launch, July 9, 2026, ChatGPT model-surface guide, accessed August 6, 2026

Anthropic’s pre-cutoff history is dominated by access and safety changes. Fable and Mythos launched June 9, were suspended globally on June 12 because of an export-control directive, and were redeployed beginning July 1 after controls were lifted. Anthropic also documented Fable’s topic-specific fallback to Opus 4.8. Claude Opus 5 launched July 24, before the cutoff. These facts materially affect observed hosted behavior without proving a change in the underlying weights. Suspension, June 12, 2026, redeployment, June 30, 2026, Anthropic release notes, July 24, 2026

Grok requires an explicit qualification. Grok 4.5 is the selected Grok row because it was the latest dated flagship and coding-oriented launch reviewed before the cutoff, and xAI’s model catalog calls it its flagship. However, xAI also documents Grok 4.20 as a separate, later model family/configuration. The reviewed public material does not pin a single exact Grok 4.5 backend to the X product and Grok consumer applications. API, X, consumer Grok, Grok Build, and Cursor are therefore kept as separate surfaces. xAI release notes, updated July 31, 2026, Grok 4.5 announcement, July 16, 2026, Grok 4.5 model documentation, accessed August 6, 2026, Grok 4.20 model documentation, accessed August 6, 2026

The three open-weight rows are Kimi K3, GLM-5.2, and DeepSeek V4 Pro. Kimi K3 has downloadable weights under a custom Kimi K3 license, not an unrestricted OSI license. GLM-5.2 and DeepSeek V4 Pro have downloadable weights and MIT licenses. Independent public comparisons make Kimi K3 and GLM-5.2 particularly relevant current open-weight selections, while NIST/CAISI’s public summary gives a more cautious assessment of DeepSeek V4 Pro than DeepSeek’s own comparison material. Kimi K3 weights and license, GLM-5.2 model card, DeepSeek V4 Pro model card, NIST/CAISI evaluation, May 1, 2026

The central watch conclusion is consequently narrow:

Public evidence through 2026-08-04 shows product and service volatility, but does not establish base-model quality degradation for the monitored set.

A service outage, increased latency, fallback to another model, changed system instruction, stricter safeguard, reduced context allowance, API rate limit, or consumer routing change can alter a user’s experience while leaving the base weights unchanged.

Monitored set and surface map

The current live catalog pages were used to verify official names and identifiers. Where a page was live and undated, it is not treated as a frozen historical archive; dated release documents anchor the cutoff where available.

#Monitored rowOfficial identityAPI surfaceConsumer / chat surfaceCoding surfaceWeight or access qualification
1GPT-5.6 Solgpt-5.6-solofficial model page, liveDirect API model; July 9 launch made Sol available globally.Standard ChatGPT maps Medium/High/Extra High effort to Sol; Pro exposes Sol Pro.Codex access for paid plans; ChatGPT Work exposes Sol.Sol is the flagship tier. Current live documentation lists 1.05M context and 128k maximum output; historical catalog state is not separately archived.
2GPT-5.6 Terragpt-5.6-terraofficial model page, liveDirect API model.Not selectable in standard ChatGPT; available in ChatGPT Work.Codex surface, including Free/Go and paid-plan distinctions described by OpenAI.Balanced, lower-cost GPT-5.6 tier; separate model identifier.
3GPT-5.6 Lunagpt-5.6-lunaofficial model page, liveDirect API model.Not selectable in standard ChatGPT; available in ChatGPT Work.Codex surface.Fastest and least expensive GPT-5.6 tier; separate model identifier.
4Claude Fable 5claude-fable-5Anthropic model overview, livePublic Anthropic API and cloud-provider access; API identifier is explicitly documented in the June 9 launch.Claude.ai, Claude Code, and Cowork after redeployment beginning July 1.Claude Code and related Anthropic coding surfaces.Fable and Mythos share underlying weights. Fable applies stronger safeguards and can fall back to Opus 4.8 for selected topics.
5Claude Mythos 5claude-mythos-5 is listed in Anthropic’s current overview; overview, liveRestricted Project Glasswing / approved-customer access; not a generally available public API tier at the cutoff.No general consumer availability established by the dated launch material.No general coding-product availability established.Same underlying model as Fable 5, with some safeguards lifted and narrower access. It is a separate monitoring row for access and safeguards, not an independent base-model row.
6Claude Opus 5claude-opus-5Anthropic release notes, July 24, 2026API, Amazon Bedrock, Google Cloud, and Microsoft Foundry.The dated release establishes provider availability; a historical consumer-app mapping is not fully pinned in the reviewed release record.Historical Claude Code mapping is not separately pinned in the dated release record.1M context, 128k maximum output, thinking enabled by default, and an effort ladder are documented in the July 24 release.
7Grok 4.5grok-4.5official model page, livexAI API; aliases include grok-4.5-latest and grok-build-latest.Grok Web/iOS/Android and X are separate surfaces; the reviewed documents do not establish that all use the same Grok 4.5 backend.Grok Build default and Cursor availability were announced; earlier Grok Build documentation used grok-build-0.1.Selected as the latest dated flagship/coding release. Grok 4.20 remains a separate official model/configuration in the uncertainty log.
8Kimi K3kimi-k3Kimi API model selection, liveKimi API.Kimi.com and Kimi Work.Kimi Code exposes k3 and k3-256k.Downloadable multimodal weights are available at Moonshot AI’s Hugging Face repository; license is custom Kimi K3 License, with commercial restrictions.
9GLM-5.2glm-5.2official Z.ai documentation, liveZ.ai API.Z.ai hosted chat/access surfaces.Z.ai coding access and local deployment.Downloadable weights are available in the Hugging Face repository; repository license is MIT.
10DeepSeek V4 Prodeepseek-v4-proDeepSeek release, April 24, 2026DeepSeek API with OpenAI-compatible and Anthropic-compatible interfaces.Exact mapping of the public consumer surface to the V4 Pro checkpoint is not established.No separate coding-product mapping established.Downloadable Base and Instruct weights are available in the Hugging Face repository; license is MIT.

Why these three open-weight models

  • DeepSeek V4 Pro is a major current open-weight model with a 1.6T total / 49B active architecture, 1M context, local weights, and an MIT license. It is included because it is a significant open-weight frontier candidate, while its independent evaluation evidence is useful precisely because it is more cautious than provider-reported comparisons. DeepSeek V4 release, April 24, 2026, NIST/CAISI evaluation, May 1, 2026

Public changes, incidents, and comparisons

RowDated changes and official status evidence through cutoffIndependent or attributable public evidence
GPT-5.6 SolThe July 9 launch established global rollout. OpenAI’s status history recorded a selected-model capacity incident on July 9 and elevated Codex Sol overload on July 17. These are availability/capacity events, not quality evidence. July 9 incident, July 17 incidentOpenAI’s launch table reproduced Artificial Analysis Intelligence Index v4.1 = 58.9 for Sol. The same page reports coding and agentic scores, but they are provider-published comparisons, not a Shaduf result. July 9 launch table
GPT-5.6 TerraCovered by the July 9 selected-model capacity event; no separate Terra-specific quality incident was identified in the reviewed official history. OpenAI status historyOpenAI’s launch table reports Intelligence Index v4.1 = 55.0.
GPT-5.6 LunaCovered by the same July 9 multi-model capacity event; no separate Luna-specific quality incident was identified in the reviewed official history. OpenAI status historyOpenAI’s launch table reports Intelligence Index v4.1 = 51.2.
Claude Fable 5Launched June 9; globally suspended June 12; redeployed beginning July 1. Anthropic’s public status history records multi-model error events on July 29–30 and an elevated-error event across many models on August 4, 20:48–21:59 UTC. Launch, June 9, suspension, June 12, redeployment, June 30, status historyArtificial Analysis reported Fable 5 at 64.9 on June 9; OpenAI’s July 9 table reported 59.9; Artificial Analysis reported 60 on July 24. These are different snapshots/configurations and do not establish degradation. AA June 9, OpenAI July 9, AA July 24
Claude Mythos 5Launched June 9 as a restricted Glasswing access tier, suspended June 12 with Fable, and restored to selected U.S. organizations after June 30. Anthropic launch, redeploymentNo separate public independent score was located that isolates Mythos from Fable. The absence is important: Fable’s public score cannot be relabeled as an independent Mythos score.
Claude Opus 5Launched July 24. Anthropic status records Opus 5 incidents on July 26–27 and multi-model events on July 30. These establish service instability windows, not a change in model capability. Anthropic release notes, status historyArtificial Analysis reported Intelligence Index = 61, GDPval-AA v2 Elo = 1861, and AA-Briefcase = 1720 in its July 24 snapshot. The article notes evaluation-specific settings and fallback behavior. AA Opus 5, July 24
Grok 4.5API availability was announced July 8; the product announcement followed July 16, with Grok Build and Cursor access. The reviewed xAI status pages did not show a Grok 4.5-specific July/August incident; they do show earlier Grok 4.3 context-loss and general-service events. xAI release notes, Grok 4.5 announcement, Grok Web status, Grok in X status, API statusArtificial Analysis reported 54. xAI’s announcement reproduces cross-provider coding figures, including SWE-bench Pro = 64.7, but explicitly compares figures drawn from different developer system cards and leaderboards; those figures are not used as a unified chart. AI Weekly, July 11, xAI announcement
Kimi K3Kimi K3 was announced in Kimi Code release notes July 16; downloadable weights were reported available by July 27. Moonshot’s official status page records search-service errors on July 23, 27, 30, and August 3; those incidents were not K3-specific. Kimi Code release notes, Moonshot statusArtificial Analysis reported 57 on July 17. Tom’s Hardware reported the July 27 weight release and distinguished open-weight from unrestricted open-source licensing. AA July 17, Tom’s Hardware, July 27
GLM-5.2The reviewed Z.ai status history showed no incidents in its visible July 17–31 window; this does not prove that no other incidents existed. A June 17 GitHub issue reported severe API rate limiting and 429 responses; it is community availability evidence, not an official incident or quality regression. Z.ai status, GitHub issue #83Artificial Analysis reported 51 on June 16 and called GLM-5.2 the leading open-weight model in that snapshot. TechRadar reported a first-place result in a single-turn HTML web-design contest; Axios cited security-evaluation evidence and community reports about jailbreak ease. These are task-specific signals. AA June 16, TechRadar, June 25, Axios, June 25
DeepSeek V4 ProReleased April 24. DeepSeek retired the legacy deepseek-chat and deepseek-reasoner aliases on July 24 at 15:59 UTC, with transitional routing to V4-Flash. Official status history records web/API incidents on May 6 and May 8; they are service-level records, not V4 Pro-specific quality events. DeepSeek release, status historyNIST/CAISI reported a substantially more cautious evaluation than provider self-reported comparisons, including SWE-bench Verified = 74 versus cited GPT-5.5 = 81 and Opus 4.6 = 79. Artificial Analysis’s June 16 open-weight comparison listed DeepSeek V4 Pro at 44, behind GLM-5.2 in that snapshot. NIST/CAISI, May 1, AA June 16

Dated change and incident timeline

  • March 10, 2026 — Grok 4.20 and Grok 4.20 Multi-agent documented as live. This is why Grok 4.20 remains a separate ambiguity rather than being silently mapped to Grok 4.5. xAI release notes
  • April 7, 2026 — Grok 4.20 system card. xAI described Grok 4.20 as its latest model at that time and documented deployment on consumer surfaces including grok.com and X. This does not establish the later Grok 4.5 backend mapping. Grok 4.20 system card, April 7
  • April 24, 2026 — DeepSeek V4 Pro release. DeepSeek documented the API identifier, 1M context, reasoning/non-reasoning modes, and legacy alias retirement plan. DeepSeek release, April 24
  • May 1, 2026 — NIST/CAISI evaluation of DeepSeek V4 Pro. The public summary reported results that were more cautious than the provider’s self-reported frontier comparisons. NIST/CAISI
  • May 6 and May 8, 2026 — DeepSeek web/API incidents. The official status history records degraded performance and an unavailability event. These are service events, not model-quality measurements. DeepSeek status history
  • May 28, 2026 — Claude Opus 4.8 predecessor release. Opus 4.8 is used only as a reference comparator and as Anthropic’s documented Fable fallback target; it is not one of the ten monitored rows. Anthropic Opus 4.8 release
  • June 9, 2026 — Claude Fable 5 and Mythos 5 launch. Anthropic stated that Mythos uses the same underlying model as Fable while changing safeguards and access. Anthropic launch
  • June 12, 2026 — Fable and Mythos globally suspended. Anthropic attributed the suspension to an export-control directive and stated that other Anthropic models were unaffected. Anthropic suspension
  • June 16, 2026 — GLM-5.2 identified as leading open-weight model in an Artificial Analysis snapshot. Artificial Analysis
  • June 26, 2026 — GPT-5.6 Sol, Terra, and Luna preview. OpenAI described the three-tier structure and limited preview. OpenAI preview
  • June 30–July 1, 2026 — Fable redeployment. Anthropic stated that controls were lifted June 30 and Fable availability resumed globally beginning July 1; Mythos returned to selected U.S. organizations. Anthropic redeployment
  • July 8, 2026 — Grok 4.5 API availability. xAI documented the API model and launch pricing. xAI release notes
  • July 9, 2026 — GPT-5.6 general availability. OpenAI documented API, ChatGPT, and Codex surfaces. On the same date, OpenAI recorded a selected-model capacity incident lasting approximately 19:44–20:42 UTC. OpenAI launch, OpenAI status incident
  • July 11, 2026 — Independent Grok 4.5 comparison. AI Weekly reported Artificial Analysis’s score and hallucination comparison; this is an independent snapshot, not a longitudinal finding. AI Weekly
  • July 16, 2026 — Grok 4.5 product announcement and Kimi K3 open-weight release. xAI described Grok 4.5 across API, Grok Build, Cursor, and related surfaces; Kimi Code documented K3’s release and open-sourcing. xAI announcement, Kimi Code release notes
  • July 17, 2026 — Kimi K3 independent snapshot and OpenAI Codex Sol overload. Artificial Analysis reported K3 at 57; OpenAI recorded elevated Codex Sol server overload. AA Kimi K3, OpenAI status incident
  • July 24, 2026 — Claude Opus 5 release and DeepSeek legacy-alias retirement. Anthropic released Opus 5; DeepSeek retired the legacy aliases at 15:59 UTC; Artificial Analysis published its Opus 5 comparison snapshot. Anthropic release notes, DeepSeek release, AA Opus 5
  • July 26–27, 2026 — Opus 5 status incidents and Kimi K3 weight availability. Anthropic recorded Opus 5 service events; Tom’s Hardware reported Kimi K3 weights available through Hugging Face. Anthropic status, Tom’s Hardware
  • July 29–30, 2026 — Anthropic multi-model errors and Moonshot search incidents. These records indicate availability volatility across hosted systems; they do not isolate base-model quality. Anthropic status, Moonshot status
  • August 3, 2026 — Moonshot search-service incident and Anthropic multi-model/Sonnet incidents. Neither record is Kimi K3- or Claude-base-quality evidence. Moonshot status, Anthropic status
  • August 4, 2026 — Anthropic elevated errors across many models, approximately 20:48–21:59 UTC. This falls inside the cutoff and is recorded as a service incident. Anthropic status

Comparable public charts

These charts use only same-metric public snapshots. They are not degradation curves. A degradation claim would require repeated measurements using a stable, comparable protocol; the public data below does not provide that.

Chart 1 — Artificial Analysis Intelligence Index v4.1 values reproduced in OpenAI’s July 9 launch table

Approximate visual bars; the CSV is authoritative.

Claude Fable 5   59.9 | ██████████████████████████████
GPT-5.6 Sol      58.9 | █████████████████████████████
Claude Opus 4.8  55.7 | ████████████████████████████
GPT-5.6 Terra    55.0 | ███████████████████████████
GPT-5.6 Luna     51.2 | ██████████████████████████

This is a provider-published comparison table, not an independently rerun dataset. Opus 4.8 is included as a predecessor reference, not as a monitored row. OpenAI launch table, July 9, 2026

model,metric,value,unit,source_url,published_date,source_type,notes
"GPT-5.6 Sol","Artificial Analysis Intelligence Index v4.1",58.9,"index points","https://openai.com/index/gpt-5-6/","2026-07-09","provider-published comparison","OpenAI launch table; not a Shaduf measurement"
"GPT-5.6 Terra","Artificial Analysis Intelligence Index v4.1",55.0,"index points","https://openai.com/index/gpt-5-6/","2026-07-09","provider-published comparison","OpenAI launch table; not a Shaduf measurement"
"GPT-5.6 Luna","Artificial Analysis Intelligence Index v4.1",51.2,"index points","https://openai.com/index/gpt-5-6/","2026-07-09","provider-published comparison","OpenAI launch table; not a Shaduf measurement"
"Claude Fable 5","Artificial Analysis Intelligence Index v4.1",59.9,"index points","https://openai.com/index/gpt-5-6/","2026-07-09","provider-published comparison","Fable is a monitored safeguards/access row"
"Claude Opus 4.8","Artificial Analysis Intelligence Index v4.1",55.7,"index points","https://openai.com/index/gpt-5-6/","2026-07-09","provider-published comparison","Predecessor reference; not one of the ten monitored rows"

Chart 2 — Artificial Analysis July 24 snapshot

Claude Opus 5   61 | ███████████████████████████████
Claude Fable 5  60 | ██████████████████████████████
GPT-5.6 Sol     59 | ██████████████████████████████
Kimi K3         57 | █████████████████████████████
Claude Opus 4.8 56 | ████████████████████████████

Artificial Analysis reports these as a single comparison snapshot. The article discusses effort settings and fallback behavior, so the values should not be read as pure immutable-weight scores. Artificial Analysis, July 24, 2026

model,metric,value,unit,source_url,published_date,source_type,effort_or_surface,notes
"Claude Opus 5","Artificial Analysis Intelligence Index",61,"index points","https://artificialanalysis.ai/articles/opus-5","2026-07-24","independent publisher snapshot","max","AA July 24 comparison snapshot"
"Claude Fable 5","Artificial Analysis Intelligence Index",60,"index points","https://artificialanalysis.ai/articles/opus-5","2026-07-24","independent publisher snapshot","max","Fable hosted configuration includes fallback behavior in the AA context"
"GPT-5.6 Sol","Artificial Analysis Intelligence Index",59,"index points","https://artificialanalysis.ai/articles/opus-5","2026-07-24","independent publisher snapshot","max","AA-reported comparator"
"Kimi K3","Artificial Analysis Intelligence Index",57,"index points","https://artificialanalysis.ai/articles/opus-5","2026-07-24","independent publisher snapshot","not stated","Effort qualifier not stated in the cited comparison"
"Claude Opus 4.8","Artificial Analysis Intelligence Index",56,"index points","https://artificialanalysis.ai/articles/opus-5","2026-07-24","independent publisher snapshot","max","Predecessor reference; not one of the ten monitored rows"

The apparent Fable movement from Artificial Analysis’s June 9 value of 64.9 to 60 on July 24 is not evidence of degradation. The July 9 OpenAI table reports 59.9, and the publications differ in snapshot date, evaluation configuration, effort, fallback, and presentation. AA June 9, OpenAI July 9, AA July 24

What is established

  • The ten monitored rows and their official names, identifiers, or access artifacts are publicly documented, subject to the historical-catalog caveat below.
  • GPT-5.6 Sol, Terra, and Luna are separate OpenAI tiers with distinct model identifiers and distinct intended price/performance positions. Their API, ChatGPT, and Codex surfaces are not interchangeable. OpenAI GPT-5.6 launch, ChatGPT surface guide
  • Claude Fable 5 and Claude Mythos 5 share underlying model weights according to Anthropic. Their safeguards, access rules, and fallback behavior differ. Anthropic Fable/Mythos launch
  • Official status pages document multiple availability and service events before the cutoff. Those events establish that users may have encountered errors, capacity limits, routing failures, or unavailable surfaces. They do not establish that model weights became less capable.
  • Public comparison values exist for only some models and dates. They are snapshots, not a consistent longitudinal dataset.

What is not established

  • No reviewed public source establishes that any monitored base model’s general capability declined before 2026-08-04 23:59 UTC.
  • No reviewed public status incident proves model degradation. The OpenAI, Anthropic, Moonshot, Z.ai, xAI, and DeepSeek records are primarily service-level records.
  • No independent public score was found that isolates Claude Mythos 5 from Fable 5 as an independent base model.
  • No public document reviewed here definitively maps Grok 4.5 to the backend used by every X, Grok Web, iOS, or Android experience.
  • No universal ranking can be inferred from the cited Artificial Analysis, NIST/CAISI, xAI, TechRadar, Axios, or provider-published tables. They use different tasks, harnesses, effort settings, system instructions, fallback policies, and sometimes different provider submissions.
  • The Kimi K3 license should not be described as an unrestricted OSI open-source license merely because the weights are downloadable.
  • A hosted API score cannot be assumed to represent a local deployment of the same open-weight checkpoint. Quantization, serving stack, context limits, prompt templates, tool routing, and provider system instructions can change the observed result.

Uncertainty

  1. Historical model catalogs. Several official model pages are live catalogs accessed after the cutoff. They verify current identifiers, but their historical state cannot always be reconstructed. Dated release notes are preferred for cutoff claims.
  1. Surface routing. A model name may identify a family while a product routes requests through a selected effort level, fallback model, safety classifier, context compaction path, or hidden system instruction.
  1. Fable and Mythos. The weights are shared, but the user-visible behavior is intentionally different. Treating their scores or incidents as independent base-model evidence would overstate what the public record supports.
  1. Grok 4.5 and Grok 4.20. Both are officially documented. Grok 4.5 is the selected row because it was the latest dated flagship/coding release reviewed for the cutoff, not because public evidence proves it dominates Grok 4.20 on every task.
  1. Comparison drift. Fable’s reported values differ across June 9, July 9, and July 24 publications. The record does not provide the identical evaluator configuration needed to interpret that spread as improvement or decline.
  1. Status-history completeness. Provider status pages differ in retention, incident taxonomy, timezone presentation, and whether model-specific incidents are separated from platform incidents. A visible absence of an incident is not proof that no incident occurred.
  1. Community reports. The GLM-5.2 rate-limit issue and security/jailbreak reports are attributable observations, but they are not controlled evidence of model degradation or general behavior.

Freshness

This report is locked to 2026-08-04 23:59 UTC.

Live pages were accessed after the cutoff to verify model identities, documentation, and visible status history. Dated material first published after the cutoff was excluded from conclusions. In particular, current Anthropic release and status pages contain later August 5 material; it is outside the reporting window. Current API catalog prices or fields that differ from July launch documentation are not presented as historical cutoff values.

Where a current page has changed and its earlier state cannot be established, the report says so rather than treating the current page as an archived snapshot.

Sources

Primary model and release documentation

Official status histories

Independent and attributable public comparisons

Search Shaduf

Search published pools, pages, reports, and evidence.