AI Post Transformers • Visual Companion

PaperBench: Can AI Replicate AI Research?

PaperBench measures whether a frontier agent can survive the ugly middle stretch between PDF and reproduced result: recovering missing details, writing the repo, running experiments, handling failures, and passing a clean rerun under judgment.

arXiv 2504.01848 ICML 2024 oral + spotlight slice 20 papers 12 topics 8,316 gradable tasks Transcript IDs: 2504.01848

Scratch replication is the benchmark’s core move: the agent gets paper plus addendum, builds a repo from zero, and only wins if a fresh Ubuntu 24.04 A10 rerun still holds together.

Paper-to-Score Pipeline

Click a stage to highlight where failure pressure moves from missing specification to execution fragility and finally to claim-level judgment.

Selected Stage

The benchmark is only partly about coding. The brittle part is chained orchestration under delayed feedback.

Curation Funnel

PaperBench is intentionally specific: accessible dependencies, single-machine runs, author collaboration, and clean rerunability.

The judge does not read the whole repo holistically. It scores leaf requirements one at a time, then rolls them upward through a weighted tree.

Weighted Rubric Tree

Illustrative leaf layout shaped to the paper description: Code Development, Execution, and Result Match converge into one aggregate replication score.

Requirement Heatmap

Swap between where the agent feels burden and where the judge feels reliable. The transcript’s concern shows up as a gap.

Top-10 File Context Bottleneck

If the repo is large, the judge sees only the ten most relevant files. Hover the stack to see what survives retrieval and what falls out of context.

PaperBench sits above artifact-backed reruns and below open-ended discovery. It is best read as a scratch-replication rung, not a full model of frontier research practice.

Where PaperBench Sits in the Evaluation Ladder

Bubble size approximates time horizon. Toggle colors between environment-access burden and judge-load dependence.

Capability Rung

PaperBench compresses more of the research loop than engineering-only benchmarks, but it still stops short of open-ended idea generation.

Interpretation Lens

The transcript’s debate turns on whether scratch replication is “advanced engineering” or “partial proxy for research.” PaperBench lives in the overlap.

The headline is humbling: best full-task agent performance stays low, proxy code-writing looks better than the full benchmark, and humans gain ground over long horizons.

Scoreboard

Headline mode uses transcript values directly. Category mode is an illustrative pattern that matches the reported “code better than execution or result match” shape.

48-Hour Time Horizon

Illustrative trajectories follow the transcript: the agent jumps early, humans start slower, and the crossover lands around 24 hours.

Judge Signal vs Proxy Signal

High judge agreement matters, but cheap proxy grading is both tempting and noisy. That tradeoff is part of the benchmark story, not a side detail.

References

Primary paper, benchmark foils, judge-evaluation context, reproducibility work, and prior AI Post Transformers episodes referenced in the discussion.