Paper-to-Score Pipeline
Click a stage to highlight where failure pressure moves from missing specification to execution fragility and finally to claim-level judgment.
PaperBench measures whether a frontier agent can survive the ugly middle stretch between PDF and reproduced result: recovering missing details, writing the repo, running experiments, handling failures, and passing a clean rerun under judgment.
Scratch replication is the benchmark’s core move: the agent gets paper plus addendum, builds a repo from zero, and only wins if a fresh Ubuntu 24.04 A10 rerun still holds together.
Click a stage to highlight where failure pressure moves from missing specification to execution fragility and finally to claim-level judgment.
The benchmark is only partly about coding. The brittle part is chained orchestration under delayed feedback.
PaperBench is intentionally specific: accessible dependencies, single-machine runs, author collaboration, and clean rerunability.
The judge does not read the whole repo holistically. It scores leaf requirements one at a time, then rolls them upward through a weighted tree.
Illustrative leaf layout shaped to the paper description: Code Development, Execution, and Result Match converge into one aggregate replication score.
Swap between where the agent feels burden and where the judge feels reliable. The transcript’s concern shows up as a gap.
If the repo is large, the judge sees only the ten most relevant files. Hover the stack to see what survives retrieval and what falls out of context.
PaperBench sits above artifact-backed reruns and below open-ended discovery. It is best read as a scratch-replication rung, not a full model of frontier research practice.
Bubble size approximates time horizon. Toggle colors between environment-access burden and judge-load dependence.
PaperBench compresses more of the research loop than engineering-only benchmarks, but it still stops short of open-ended idea generation.
The transcript’s debate turns on whether scratch replication is “advanced engineering” or “partial proxy for research.” PaperBench lives in the overlap.
The headline is humbling: best full-task agent performance stays low, proxy code-writing looks better than the full benchmark, and humans gain ground over long horizons.
Headline mode uses transcript values directly. Category mode is an illustrative pattern that matches the reported “code better than execution or result match” shape.
Illustrative trajectories follow the transcript: the agent jumps early, humans start slower, and the crossover lands around 24 hours.
High judge agreement matters, but cheap proxy grading is both tempting and noisy. That tradeoff is part of the benchmark story, not a side detail.
Primary paper, benchmark foils, judge-evaluation context, reproducibility work, and prior AI Post Transformers episodes referenced in the discussion.