The benchmark isolates the effect of procedural artifacts by keeping the task and verifier fixed, then switching only the agent’s access to skills.
Each task runs inside a containerized environment with files, optional skill directory, tools, and a deterministic verifier.
Conditions: No Skills, Curated Skills, and Self-Generated Skills.
This is not mainly a benchmark of base-model intelligence. It asks whether reusable runbooks, templates, and checks change outcomes.
Mock-but-realistic data based on the episode’s reported patterns: large average curated gains, wide domain variance, and near-zero self-generated benefit.
Each cell is a task. Hover to inspect where curated skills help, do little, or actively hurt. The pattern should look structured, not uniformly positive.
The episode’s core disagreement: facts-in-context vs procedural know-how about when, how, and in what order to act.
Relevant passages, tool docs, and facts to read before acting.
Action sequencing, templates, decision points, verification steps, and reusable scripts.
If the skill adds more context tokens, how much of the gain is true procedural abstraction versus generic context augmentation?
Benchmarking the “pleasant middle” is useful — but real systems must discover, select, compose, test, and retire skills over time.
Compact source map for the visual claims and framing around skills, agents, verification, and benchmark design.