arXiv: 2602.12670

SkillsBench for Evaluating Agent Skills

A visual companion focused on one question: when reusable procedural “skills” are added to an agent, do they improve real multi-step task completion — or just add more context tokens?
86 tasks
11 domains
3 conditions
7 agent-model setups
7,308 trajectories
Avg curated lift
+16.2 pp
Self-generated lift
~0 pp
Tasks worse w/ skills
16 / 84
Largest domain gain
+51.9 pp
No skills Curated skills Self-generated

Benchmark structure

The benchmark isolates the effect of procedural artifacts by keeping the task and verifier fixed, then switching only the agent’s access to skills.

Measured object

PASS/FAIL, not judge scores

Each task runs inside a containerized environment with files, optional skill directory, tools, and a deterministic verifier.

Key intervention

Same task, different procedural support

Conditions: No Skills, Curated Skills, and Self-Generated Skills.

Why it matters

Workflow engineering as a variable

This is not mainly a benchmark of base-model intelligence. It asks whether reusable runbooks, templates, and checks change outcomes.

Performance explorer

Mock-but-realistic data based on the episode’s reported patterns: large average curated gains, wide domain variance, and near-zero self-generated benefit.

Baseline / No skills Curated skills Self-generated skills

Task × domain gain heatmap

Each cell is a task. Hover to inspect where curated skills help, do little, or actively hurt. The pattern should look structured, not uniformly positive.

Low → medium → high Hot cells = more risk / larger gain depending on mode

Is a skill just RAG with better branding?

The episode’s core disagreement: facts-in-context vs procedural know-how about when, how, and in what order to act.

RAG emphasizes

Relevant passages, tool docs, and facts to read before acting.

Skills emphasize

Action sequencing, templates, decision points, verification steps, and reusable scripts.

Open issue

If the skill adds more context tokens, how much of the gain is true procedural abstraction versus generic context augmentation?

Production lifecycle the benchmark does not fully cover

Benchmarking the “pleasant middle” is useful — but real systems must discover, select, compose, test, and retire skills over time.

Benchmark-covered center Open research / operational bottleneck Future benchmark target

References

Compact source map for the visual claims and framing around skills, agents, verification, and benchmark design.

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Li et al., 2026 · arXiv:2602.12670
ReAct: Synergizing Reasoning and Acting in Language Models Yao et al., 2022
Between MDPs and Semi-MDPs: Temporal Abstraction in Reinforcement Learning Sutton, Precup, Singh, 1999
Language Agents with Cognitive Architectures Sumers et al., 2023
SWE-bench Jimenez et al., 2024
WebArena Zhou et al., 2024
Terminal-Bench Merrill et al., 2026
Anthropic Skills documentation / product specification Anthropic, 2025
ToolReflection / self-generated data for tool use Approx. 2025
Agent skills surveys and frameworks Approx. 2025
Additional arXiv IDs detected in transcript: none beyond 2602.12670.