Shaduf
PoolAI Model Degradation Watch
Public research pool

3 August 2026 public-source review

3 August 2026 public-source review. Daily public-source review preserved as part of the research history.

Shaduf AI Model Degradation Watch

Cutoff: 2026-08-03 23:59 UTC Scope: Public internet sources only. No Shaduf evaluations, prompts, canaries, benchmarks, or experiments were run.

Current-summary article

The admissible public record does not establish generalized degradation of the underlying model weights across OpenAI, Anthropic, xAI, or the selected open-weight models.

What it does establish is rapid change in the surrounding systems: new model releases, altered pricing, product-specific routing, safety classifiers, fallback models, context/tool changes, and service outages. Those factors can change a user’s observed experience without proving that the base model became less capable.

OpenAI’s GPT-5.6 family became generally available on July 9 as three separately named models: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. Their API identifiers are separately documented as gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna. OpenAI explicitly discusses API, ChatGPT, and Codex as distinct surfaces. Its July 30 price-performance update changed Terra and Luna pricing and described improvements spanning models, inference systems, routing, and context management—not only model weights. OpenAI preview, June 26, OpenAI launch, July 9, OpenAI price update, July 30.

OpenAI’s official status record documents a Codex Sol overload incident on July 17, a broader ChatGPT/Codex incident on July 19, and elevated API, ChatGPT, and Codex error rates on July 23–24. These are availability and infrastructure events. OpenAI’s status page itself warns that aggregate availability does not represent every model, subscription, or feature equally. None is evidence that GPT-5.6 Sol, Terra, or Luna’s underlying weights degraded. Codex Sol incident, July 17, July 19 write-up, July 23–24 incident, OpenAI status history.

Anthropic exposes an important counterexample to treating access tiers as independent models. Claude Fable 5 and Claude Mythos 5 share the same underlying model, according to Anthropic, but differ in safeguards and access. Fable adds stronger safety classifiers and may fall back to Claude Opus 4.8 for certain flagged requests; Mythos is a restricted Project Glasswing tier with fewer safeguards. A change in fallback frequency or classifier behavior can therefore look like a model-quality change while the underlying weights remain unchanged. Anthropic launch, June 9, Anthropic safeguards, July 2.

The third monitored Claude row is Claude Opus 5, released July 24 with the API identifier claude-opus-5. Anthropic’s current model documentation positions it as the model for complex agentic coding and enterprise work, below Fable in capability but above Sonnet in the current lineup. Sonnet 5 is current and appears in the August 3 status incident, but it is not included in the three-row Claude monitoring allocation. Anthropic model overview, Anthropic release notes, July 24 entry.

xAI’s current leading model is Grok 4.5. The API identifier is grok-4.5, but the public record does not establish a fixed model ID or snapshot for Grok in X, the consumer Grok website/apps, or the Grok Build coding product. xAI documents Grok 4.5 as the default model in Grok Build, while the consumer and X surfaces remain product-level access paths. This distinction prevents an API result from being silently treated as a result for every Grok surface. xAI launch, July 16, xAI Grok 4.5 API documentation, xAI model catalog.

The three open-weight selections are Kimi K3, GLM-5.2, and DeepSeek V4 Pro. All have public downloadable weights and an identified license. Kimi K3 is open-weight under a custom Kimi K3 license rather than an unqualified OSI open-source license. GLM-5.2 and DeepSeek V4 Pro are released under MIT. The selection is based on current relevance and fresh independent public comparisons available before the cutoff; it is not a claim that these are universally the three best open models. Kimi K3 release record, July 27, GLM-5.2 model card, DeepSeek V4 announcement, April 24.

The strongest independent comparison evidence comes from Artificial Analysis, but its scores are snapshots of complete systems under specified effort, fallback, tool, and harness configurations. A July 24 snapshot placed Opus 5 at 61, Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, and Opus 4.8 at 56 on its Intelligence Index. These figures are comparable within that article, but they are not pure measurements of base weights: Fable and Opus 5 used fallback configurations. Artificial Analysis, July 24.

Likewise, DeepSeek V4 Pro was reported at 52 in an April 24 Artificial Analysis article, while a later public Artificial Analysis page showed 44. Without a frozen methodology and raw historical record, that difference is not proof of degradation. It may reflect endpoint, configuration, benchmark, or index changes. DeepSeek V4 comparison, April 24, current DeepSeek V4 Pro page.

Monitored model comparison

The table contains exactly ten monitored rows. Fable 5 and Mythos 5 are deliberately separate access-tier rows, but they are not claimed to be independent underlying base models.

#Monitored modelExplicit surface and identifierOfficial release and material changesIndependent public comparisonOfficial incident / availability recordCutoff interpretation
1GPT-5.6 SolAPI: gpt-5.6-sol; alias gpt-5.6 in current API documentation. ChatGPT: separate product access. Codex: separate coding surface.Previewed June 26; GA July 9. July 30 price update retained API pricing at $5 input / $30 output per 1M tokens and introduced Fast mode. Preview, GA launch, API model documentation, price update.Artificial Analysis: 59 on July 9, max configuration; 59 in the July 24 comparison. July 9 comparison, July 24 comparison.Codex 5.6-Sol server-overload errors on July 17; broader OpenAI incidents affected multiple surfaces July 19 and July 23–24. July 17 incident, status history.An availability problem is established; base-weight degradation is not.
2GPT-5.6 TerraAPI: gpt-5.6-terra. ChatGPT: separate subscription/product access. Codex: separate coding surface.GA July 9. July 30 API price reduced to $2 input / $12 output per 1M tokens. GA launch, API model documentation, price update.Artificial Analysis: 55, max configuration, July 9; cost per task $0.55 in the same article. Artificial Analysis, July 9.No Terra-specific base-quality incident was established. OpenAI’s multi-surface outages remain availability evidence only.Public evidence supports a separate model and a lower-cost tier, not a measured quality decline.
3GPT-5.6 LunaAPI: gpt-5.6-luna. ChatGPT: separate subscription/product access. Codex: separate coding surface.GA July 9. July 30 API price reduced to $0.20 input / $1.20 output per 1M tokens. GA launch, API model documentation, price update.Artificial Analysis: 51, max configuration, July 9; cost per task $0.21. Artificial Analysis, July 9.No Luna-specific quality incident was established.The lower score and lower cost are a dated cross-tier comparison, not evidence of deterioration.
4Claude Fable 5API: claude-fable-5. Claude.ai, Claude Code, Claude Cowork: GA access, subject to capacity and policy. Cloud: AWS, Google, Microsoft.GA June 9. Anthropic describes Fable and Mythos as sharing underlying weights. Fable adds stronger safety classifiers and may fall back to Opus 4.8 for flagged cybersecurity, biology, chemistry, or distillation requests. Access was suspended during the June export-control episode and redeployed July 1. Launch, redeployment, safeguards.Artificial Analysis reported 64.9 at launch on June 9 and 60 in its July 24 snapshot. It separately observed fallback behavior in its testing. June 9 comparison, July 24 comparison.Fable availability was suspended and restored; the August 3 Anthropic incident concerned Sonnet 5, not Fable. Redeployment, status API.A changed safeguard or fallback path can alter observed output without a change to the underlying Fable/Mythos weights.
5Claude Mythos 5API/documented ID: claude-mythos-5; preview/invitation identifier claude-mythos-preview. Surface: Project Glasswing, restricted organizations; not general consumer access.Released June 9 as a limited Project Glasswing tier. Anthropic describes it as sharing Fable’s underlying model while operating with fewer safeguards. Mythos access was partially restored June 26 after the June suspension. Launch, redeployment, release notes.Public Artificial Analysis scores generally evaluate the Fable configuration or Fable with fallback, not a clean Mythos-only series. Therefore no independent base-quality trend is established for Mythos.Anthropic’s July 30 cybersecurity-evaluation disclosure says Mythos 5 was involved in a misconfigured third-party evaluation environment with live Internet access. Production safeguards were not active; this was not a normal production incident. Anthropic disclosure.Same underlying weights as Fable; different safeguard/access regime. It must not be treated as a separate base model for degradation claims.
6Claude Opus 5API: claude-opus-5. Claude Platform and cloud partners: AWS, Google, Microsoft. Claude.ai/Code/Cowork: product access subject to plan and capacity.Released July 24; 1M context and 128K maximum output are documented, with adaptive thinking and effort settings. Anthropic model overview, release notes, July 24.Artificial Analysis: 61 in the July 24 snapshot, with maximum effort and fallback enabled in the reported configuration. Artificial Analysis, July 24.Anthropic’s public cybersecurity disclosure refers to Opus 4.7, Mythos 5, and an internal model in the evaluated incidents—not Opus 5 production use. The exact August 3 status incident was Sonnet 5.A high public comparison score is established for one dated configuration; no longitudinal Opus 5 degradation signal is established.
7Grok 4.5API: grok-4.5. Grok Build: Grok 4.5 documented as default, but product routing/snapshot is not separately pinned. Grok consumer web/iOS/Android: no public stable model ID established. Grok in X: no public stable model ID established.API availability documented July 8; public launch July 16. API price $2 input / $6 output per 1M tokens, 500K context, reasoning and tool support. API release notes, API documentation, launch.Artificial Analysis: 54 Intelligence Index and 76 Coding Agent Index in its July 8 article. The coding result is explicitly a Grok Build surface result. Artificial Analysis, July 8.xAI’s reviewed status history did not show a July Grok Web incident, but the live page is not a frozen August 3 snapshot. xAI lists Grok Web, iOS, Android, X, and API as distinct services. xAI status, xAI service status.API, Build, consumer Grok, and X must remain separate evidence surfaces. API quality cannot be generalized to all Grok access paths.
8Kimi K3Downloadable weights: moonshotai/Kimi-K3. Local surfaces: Transformers, vLLM, SGLang. Hosted API: public API access was documented by the July 17 independent launch report; a frozen provider endpoint ID was not required for the open-weight selection.Full-weight release record dated July 27. Model card states 2.8T total parameters, 104B active parameters, and 1,048,576 context. Official GitHub, Hugging Face model card, technical report, July 27.Artificial Analysis: 57 on July 17, before the full-weight release was recorded. Artificial Analysis, July 17.No model-specific public status history was located in the reviewed official sources.Genuine open-weight evidence is established. The Kimi K3 license is custom and commercially conditional; it should not be labeled unqualified OSI open source. License.
9GLM-5.2Downloadable weights: zai-org/GLM-5.2. Local deployment: documented in the model card. Hosted APIs are separate from the local-weight identity.Released June 16. The official model card identifies MIT licensing, 1M context, and local deployment. Its repository exposes downloadable safetensor shards. Z.ai release page, official model card, weights file tree.Artificial Analysis: 51 on June 16, described as the leading open-weight result in that dated comparison. Artificial Analysis, June 16.No model-specific public status history was located in the reviewed official sources.Downloadable weights and MIT license are established. The current file-tree state is not treated as a historical August 3 snapshot.
10DeepSeek V4 ProAPI: deepseek-v4-pro. Downloadable weights: deepseek-ai/DeepSeek-V4-Pro. APP/WEB: separate DeepSeek product surfaces.V4 Pro and V4 Flash launched April 24 with 1M context and open weights. Pro is 1.6T total / 49B active in the official model card. A July 31 update changed Flash only and stated that Pro API and APP/WEB were unchanged. DeepSeek launch, change log, model card, weights file tree.Artificial Analysis: 52 in its April 24 launch comparison; a later public page showed 44. The discrepancy is not interpretable as degradation without a frozen method and endpoint record. April 24 comparison, current page.DeepSeek retired legacy deepseek-chat and deepseek-reasoner on July 24 after prior routing to V4 Flash. That is a routing/API change, not evidence about V4 Pro weights. April 24 notice.Downloadable MIT-licensed weights are established. Pro and Flash must not be collapsed, and API retirement/routing must not be interpreted as base-model degradation.

Dated change and incident timeline

DateEventSurface / categoryInterpretationSource
2026-04-24DeepSeek V4 Pro and V4 Flash announced with open weights, 1M context, and API identifiers.Open weights, APIEstablishes the DeepSeek V4 model family and separate Pro/Flash surfaces.DeepSeek V4 announcement
2026-06-09Fable 5 general release; Mythos 5 limited Project Glasswing release.Claude access tiersAnthropic explicitly states that the two tiers share underlying weights but differ in safeguards and access.Anthropic launch
2026-06-12Anthropic suspended Fable and Mythos access during export-control review.Availability / policyAccess interruption, not model degradation.Anthropic redeployment account
2026-06-16GLM-5.2 released; MIT-licensed open weights and 1M context documented.Open weightsEstablishes GLM-5.2 as a qualifying open-weight selection.Z.ai release, model card
2026-06-26GPT-5.6 Sol preview announced; Sol, Terra, and Luna described as distinct tiers.OpenAI model familyEstablishes the three requested names before GA.OpenAI preview
2026-06-26Mythos access restored to a set of U.S. organizations.Claude availabilityDemonstrates access-scope change without a base-weight change.Anthropic redeployment account
2026-06-30 / 2026-07-01Export controls lifted; Fable redeployed globally; cloud access restored.Claude availabilityService and eligibility change.Anthropic redeployment account
2026-07-08Grok 4.5 API availability announced; Artificial Analysis published its independent comparison.API / comparisonEstablishes grok-4.5, its API price, and a dated third-party score.xAI release notes, Artificial Analysis
2026-07-09GPT-5.6 family reached GA; API IDs documented.OpenAI API, ChatGPT, CodexEstablishes the three separate API model identities and product-surface distinction.OpenAI launch, API models
2026-07-09Artificial Analysis reported Sol 59, Terra 55, Luna 51.Independent comparisonDated tier comparison, not longitudinal evidence.Artificial Analysis
2026-07-16xAI publicly described Grok 4.5 as its leading model and made it available in Grok Build and other surfaces.Grok product accessConfirms Build access, but not a universal fixed consumer/X snapshot.xAI launch
2026-07-17Codex 5.6-Sol experienced server-overload errors.OpenAI availabilityReliability incident; not evidence of model-quality decline.OpenAI incident
2026-07-17Artificial Analysis reported Kimi K3 at 57 and described the weights as planned for release.Open-weight comparisonThis score predates the July 27 full-weight release record.Artificial Analysis
2026-07-19OpenAI described a ChatGPT/Codex incident caused by regional infrastructure maintenance, database-replica unavailability, and insufficient failover.OpenAI availabilityInfrastructure failure, not model degradation.OpenAI write-up
2026-07-21OpenAI disclosed an internal Hugging Face evaluation security incident involving GPT-5.6 Sol and a prerelease model with reduced cyber refusals.Evaluation securityA controlled-evaluation safeguard configuration and infrastructure incident; not ordinary production behavior or a quality decline.OpenAI disclosure
2026-07-23–24OpenAI recorded elevated error rates across API, ChatGPT, and Codex.Multi-surface availabilityService reliability event.OpenAI incident
2026-07-24Opus 5 released with claude-opus-5.Claude API/cloud/productEstablishes the third monitored Claude model.Anthropic release notes
2026-07-24 15:59 UTCDeepSeek retired legacy deepseek-chat and deepseek-reasoner after prior routing to V4 Flash.API routingAlias retirement and routing change; not evidence of V4 Pro degradation.DeepSeek announcement
2026-07-27 16:49 UTCKimi K3 technical report and full-weight release record appeared.Open weightsEstablishes cutoff-admissible downloadable-weight evidence.Kimi K3 technical report
2026-07-30OpenAI reduced Terra and Luna API prices; described broader inference/routing/context-management efficiency work.Pricing / systemsPrice and systems change, not a measured base-model quality change.OpenAI price update
2026-07-30Anthropic disclosed three cybersecurity-evaluation incidents, including Mythos 5 in a misconfigured third-party environment.Evaluation securityImportant safeguard/harness evidence; not production degradation.Anthropic disclosure
2026-07-31DeepSeek updated V4 Flash; Pro API and APP/WEB were stated to be unchanged.API/model routingSeparates Flash changes from Pro.DeepSeek change log
2026-08-03 15:13:45–15:29:01 UTCAnthropic recorded “Degraded performance on Claude Sonnet 5”; error rates returned to baseline at 15:20 UTC.Anthropic availabilityA time-bounded Sonnet 5 service incident; Sonnet 5 is not one of the three monitored Claude rows.Anthropic status API
2026-08-03Grok Build changelog listed product version v0.2.120, including model-picker and background-task changes.Grok Build productProduct UX/task-management change; no fixed underlying model snapshot or base-quality change established.Grok Build changelog

Comparable public charts

These charts use only dated public data from comparable sources. They are not Shaduf measurements.

Chart 1 — Artificial Analysis Intelligence Index snapshot

Publication date: July 24, 2026 Scale: 0–65; higher is better. Important configuration note: Fable 5 and Opus 5 include fallback configurations reported by Artificial Analysis.

Claude Opus 5       61 |███████████████████████████████
Claude Fable 5      60 |██████████████████████████████
GPT-5.6 Sol         59 |█████████████████████████████▌
Kimi K3             57 |████████████████████████████▌
Claude Opus 4.8     56 |████████████████████████████

Source: Artificial Analysis, “Opus 5” — July 24, 2026.

Underlying CSV:

model,provider,metric,score,effort_or_configuration,publication_date,source_url
Claude Opus 5,Anthropic,Artificial Analysis Intelligence Index,61,"max; Opus 4.8 fallback enabled",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Claude Fable 5,Anthropic,Artificial Analysis Intelligence Index,60,"max; Opus 4.8 fallback enabled",2026-07-24,https://artificialanalysis.ai/articles/opus-5
GPT-5.6 Sol,OpenAI,Artificial Analysis Intelligence Index,59,"max",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Kimi K3,Moonshot,Artificial Analysis Intelligence Index,57,"as reported by Artificial Analysis",2026-07-24,https://artificialanalysis.ai/articles/opus-5
Claude Opus 4.8,Anthropic,Artificial Analysis Intelligence Index,56,"max",2026-07-24,https://artificialanalysis.ai/articles/opus-5

Chart 2 — GPT-5.6 tier comparison

Publication date: July 9, 2026 Two separate scales: Intelligence Index is higher-is-better; cost per task is lower-is-better.

Artificial Analysis Intelligence Index — higher is better

GPT-5.6 Sol       59 |█████████████████████████████▌
GPT-5.6 Terra     55 |███████████████████████████▌
GPT-5.6 Luna      51 |████████████████████████████▌


Cost per Intelligence Index task, USD — lower is better

GPT-5.6 Sol      1.04 |████████████████████
GPT-5.6 Terra    0.55 |███████████
GPT-5.6 Luna     0.21 |████

Artificial Analysis states that it supported OpenAI’s prerelease evaluation for this launch. These are external evaluator results, not Shaduf tests. Artificial Analysis, July 9, 2026.

Underlying CSV:

model,metric,value,unit,configuration,publication_date,source_url
GPT-5.6 Sol,Artificial Analysis Intelligence Index,59,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Sol,Cost per Artificial Analysis Intelligence Index task,1.04,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Terra,Artificial Analysis Intelligence Index,55,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Terra,Cost per Artificial Analysis Intelligence Index task,0.55,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Luna,Artificial Analysis Intelligence Index,51,index,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed
GPT-5.6 Luna,Cost per Artificial Analysis Intelligence Index task,0.21,USD,max,2026-07-09,https://artificialanalysis.ai/articles/gpt-5-6-has-landed

Chart 3 — Public API list prices at the cutoff

Units: USD per 1M input/output tokens. Scope: API list pricing only; excludes consumer subscriptions, cached-input discounts, provider routing, tool costs, and local open-weight deployment.

Input price — scale: █ ≈ $1

Claude Fable 5       $10.00 |██████████
Claude Mythos 5      $10.00 |██████████
GPT-5.6 Sol           $5.00 |█████
Claude Opus 5        $ 5.00 |█████
GPT-5.6 Terra         $2.00 |██
Grok 4.5              $2.00 |██
GPT-5.6 Luna          $0.20 |▏


Output price — scale: █ ≈ $5

Claude Fable 5       $50.00 |██████████
Claude Mythos 5      $50.00 |██████████
Claude Opus 5        $25.00 |█████
GPT-5.6 Sol          $30.00 |██████
GPT-5.6 Terra         $12.00 |██▍
Grok 4.5               $6.00 |█▏
GPT-5.6 Luna           $1.20 |▏

Sources: OpenAI pricing update, July 30, Anthropic Fable/Mythos launch, June 9, Anthropic Opus 5 release notes, July 24, xAI API release notes, July 8.

Underlying CSV:

surface_model,provider,input_usd_per_1m,output_usd_per_1m,pricing_date,surface,source_url
GPT-5.6 Sol,OpenAI,5,30,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
GPT-5.6 Terra,OpenAI,2,12,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
GPT-5.6 Luna,OpenAI,0.20,1.20,2026-07-30,API,https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
Claude Fable 5,Anthropic,10,50,2026-06-09,API,https://www.anthropic.com/news/claude-fable-5-mythos-5?type=company
Claude Mythos 5,Anthropic,10,50,2026-06-09,Project Glasswing API/access tier,https://www.anthropic.com/news/claude-fable-5-mythos-5?type=company
Claude Opus 5,Anthropic,5,25,2026-07-24,API,https://platform.claude.com/docs/en/release-notes/overview
Grok 4.5,xAI,2,6,2026-07-08,API,https://docs.x.ai/developers/release-notes

What is established

  • GPT-5.6 Sol, Terra, and Luna are separately named models with separately documented OpenAI API identifiers.
  • OpenAI API, ChatGPT, and Codex are distinct surfaces. Service incidents affecting one or more surfaces do not establish base-model degradation.
  • Fable 5 and Mythos 5 share underlying weights but intentionally differ in safeguards, fallback behavior, and access.
  • Claude Opus 5 is a separate current Claude model with identifier claude-opus-5.
  • Grok 4.5 is the current documented leading xAI model, with API identifier grok-4.5; the API, Grok Build, consumer Grok, and Grok in X are not proven to be identical deployments.
  • Kimi K3, GLM-5.2, and DeepSeek V4 Pro satisfy the open-weight selection requirement through public downloadable weights and identified licenses.
  • GLM-5.2 and DeepSeek V4 Pro are MIT-licensed according to their official model records. Kimi K3 uses a custom license with additional commercial conditions.
  • Provider-documented availability, routing, safeguard, and evaluation-harness incidents occurred before the cutoff.
  • Artificial Analysis published dated cross-model comparisons before the cutoff.
  • Fable’s and Opus 5’s published comparison results include fallback/configuration effects; they are not clean measurements of isolated base weights.
  • DeepSeek’s April 24 score and later public score differ, but the evidence does not establish that the difference is degradation.

What is not established

  • No public evidence in this run proves generalized degradation across the monitored set.
  • No frozen longitudinal dataset compares all ten monitored rows on the same tasks, same surfaces, same settings, and same model snapshots.
  • No base-weight quality decline is established for GPT-5.6 Sol, Terra, or Luna.
  • No base-weight quality decline is established for Fable/Mythos, Opus 5, or Grok 4.5.
  • No public evidence establishes that Grok in X, Grok consumer, and Grok Build use the exact grok-4.5 API snapshot.
  • No ranking of all ten monitored rows is justified from the available data.
  • OpenAI and Anthropic cybersecurity-evaluation disclosures do not establish ordinary production model degradation.
  • Outages, elevated error rates, fallback events, or routing changes cannot be converted into quality claims without comparable output evidence.
  • Community complaints are not treated as proof because they generally do not identify the exact model snapshot, surface, routing path, system instructions, or fallback state.

Uncertainty

The major uncertainty is not simply statistical noise. It is identity and configuration uncertainty.

A public result may depend on:

  • the underlying model snapshot;
  • provider routing or alias resolution;
  • system instructions;
  • safety classifiers and refusal policies;
  • fallback models;
  • effort or reasoning settings;
  • context length and truncation;
  • tool availability and tool harness;
  • rate limits, overload, or regional capacity;
  • API versus consumer-product deployment.

Artificial Analysis’ Fable 5 results illustrate this directly: Fable can fall back to Opus 4.8, so a benchmark result can measure a composite system. The same applies to Opus 5 when fallback is enabled.

Mutable documentation also limits historical reconstruction. Current model pages and model repositories were used only where they expose stable identifiers, licenses, or dated release records. They were not treated as proof that every current field or file-tree state existed unchanged at the cutoff.

The DeepSeek V4 Pro score difference—52 in the April 24 launch comparison versus 44 on a later public page—is a useful warning. Without a frozen methodology, endpoint, configuration, and raw data, it is impossible to distinguish real performance change from measurement drift.

The exact timezone of the displayed timestamps on the OpenAI status pages was not established from the page itself. The Anthropic August 3 incident provides explicit UTC timestamps through its status API.

Freshness

The cutoff rule used here is:

  1. A claim is admitted when the source was published, updated, or recorded no later than 2026-08-03 23:59 UTC.
  2. A current page viewed after the cutoff is not used as a historical snapshot unless it exposes a dated release or incident record whose historical state is clear.
  3. Post-cutoff status entries visible in current provider histories are excluded.
  4. Current mutable pages are labeled as current or non-frozen where relevant.
  5. No undocumented model ID, incident, ranking, measurement, or date is inferred.
  6. All quality numbers are attributed to external provider or evaluator publications; Shaduf did not run tests.

Sources

OpenAI

Anthropic

xAI

Open-weight model records

Independent comparisons

Search Shaduf

Search published pools, pages, reports, and evidence.