Shaduf.Research preview
AI Web Design & Frontend Skills Catalog/Test a web skill against no skill

Research note · 2 October 2026

Test a web skill against no skill.

Before you add a frontend skill to an agent, ask whether it improves the task you have. WebDev-Skills-Bench v1 compared matched skill instructions with no-skill baselines on 117 skill-project pairs across 50 web projects and four model configurations. In its functional tests, mean Pass@2 was 1.3 to 4.2 percentage points lower with the skill, and token use rose 72% to 394%. Some pairs improved: 17% to 36%, depending on the model. The GPT-5.1 Pass@2 confidence interval included zero.

Decision: Keep a skill as a candidate, not a default. For a substantial task, run the same brief and project with and without it. Keep the instruction only if the result helps your actual task enough to justify its setup and context cost.

Another benchmark found gains

SkillsBench v4 reported a 16.6 percentage-point average gain from curated skill access across 87 tasks in eight domains and 18 model–harness configurations. That study uses different tasks and lets agents discover skills through its harness. The averages do not answer which instruction helps your screen; they explain why the no-skill comparison must match your task and setup.

What the study can and cannot tell you

The researchers injected each skill's SKILL.md into the prompt and mounted supporting files in the workspace. Their tests checked functional behavior with Playwright. That is not the same as a host loading a skill only when relevant. It does not measure design quality, accessibility, or the visual result of a landing page or dashboard. The paper is a preprint; this pool has not reproduced its experiments.

Anthropic's skill-creation instructions also use a no-skill run when evaluating a new skill. That is a method, not a finding that any particular skill helps. None of our sixteen listings has a pool-run comparison. Their source inspections still help you check task fit, requirements, and revision; they do not establish an outcome advantage.

A small comparison worth doing

  1. Use one real task with the same project files, model, and acceptance checks in both runs.
  2. In the no-skill run, give a careful brief with the existing components, tokens, route, and required states.
  3. In the skill run, add only the candidate skill; record its exact revision and loading method.
  4. Compare the working page or component, not only the agent's report. Check function, required states, responsive behavior, and any visual constraint that matters.

A single trial is local evidence, not a general ranking. We have not performed one yet; the pool's comparison fixture remains pending.

Search published pools, pages, reports, and evidence.