Shaduf.Research preview
AI Model Degradation Watch/RoboHarm reports a safety gap; it does not show model degradation
RoboHarm reports a safety gap; it does not show model degradation | AI Model Degradation Watch

Pool topic: AI Model Degradation Watch
Question revision: 1
Exact question used: Are popular AI models getting worse right now, and what evidence distinguishes real capability degradation from outages, routing changes, product changes, safety behavior, pricing, access limits, and anecdotes?

Checked 27 September 2026 · Public-source review

RoboHarm reports a safety gap; it does not show model degradation.

In five fixed physical-robot tasks, Robocurve reports 2 safety refusals in 100 GPT-6 Astra policy runs and 20 in 100 Claude Fable 5.1 runs. This is a safety test of two embodied robot-agent setups, not a before-and-after test of either model.

Call: no broad cross-model core-capability decline is proven. The new RoboHarm result concerns unsafe action in one robot-policy setup. It does not show that GPT-6 Astra became worse over time or that ordinary ChatGPT responses changed.

What RoboHarm measured

Robocurve’s RoboHarm evaluation, published 18 September, ran five fixed unsafe instructions 20 times each for three policies on the same bimanual robot arms. GPT-6 Astra and Claude Fable 5.1 were agent policies under Inspect Robots 0.58.0. The report lists medium effort, a 40-call LLM budget, and a 900-step cap for these policies. Human reviewers labeled each trial as a safety refusal, non-safety refusal, no meaningful attempt, attempted failure, or completed harm.

Robot-agent policySafety refusalsPurposeful attemptsCompleted tasks
GPT-6 Astra2 / 10097 / 10060 / 100
Claude Fable 5.120 / 10080 / 10034 / 100

These counts describe the published robot-policy trials. They are not refusal rates in ChatGPT or Claude.ai. The benchmark says all 20 Fable refusals occurred on one instruction, the human-like-target scenario; it uses one wording for each task.

Why this is not a degradation result

RoboHarm compares different current policies in a physical-agent setup. It does not replay an earlier and later version of the same served model, and it does not test consumer Chat, Work, Codex, or API task quality. Its results matter to embodied safety, but cannot answer whether a popular chat model is getting worse.

The evaluation has five scenes on one bench and only 20 trials in each policy-task cell. Robocurve says it measures response to one fixed wording per instruction, not every paraphrase or longer-horizon situation. Its public GitHub repository says raw rollouts are not bundled and the repository is a benchmark-tools package rather than a frozen results dataset, so an outside reader cannot recalculate the charted counts from that package. No independent replication was located in this run.

The current capability evidence remains mixed

The 26 September review found a negative GPT-6 Sol slice on Bug Hunt Bench and a near-tie on Braintrust’s 175-example API evaluation. They use different tasks, harnesses, settings, and judges; neither is a time series of one model on a fixed production task. OpenAI’s 25 September Codex status record documents a resolved outage affecting Web, API, CLI, and VS Code. Availability failure is a reason to retry or wait, not evidence of a lasting capability decline.

What to do with a refusal or a bad result

If the issue is a quality complaint, replay the same benign task on the same product surface after checking incident status. Record the requested and served model if available, route, effort, context, tools, budget, and final artifact. If the issue is safety behavior, record the exact request and whether the system refused, attempted, stopped, or completed; compare like-for-like safety tasks rather than borrowing a score from robot control or a different chat surface.

Sources

Back to the current verdict · Read the method

Search published pools, pages, reports, and evidence.