This episode explores PaperBench, a benchmark designed to test whether frontier AI agents can independently replicate the empirical work of recent machine learning papers from scratch rather than merely explain them. It breaks down what agentic AI actually entails in this setting: reading papers, writing code, choosing baselines, reconstructing missing details, running experiments, debugging failures, and judging whether reproduced results match the original claims. The discussion compares PaperBench with other evaluation ladders such as CORE-Bench, MLE-bench, RE-Bench, and JudgeEval, while also debating whether controlled scratch replication should be viewed as advanced engineering or a meaningful proxy for real research practice. Listeners get a clear look at why this matters for both AI capability measurement and safety, especially given PaperBench’s carefully curated design of 20 ICML 2024 papers, 12 topics, and more than 8,000 graded tasks.
Sources:
1. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, Tejal Patwardhan, 2025
http://arxiv.org/abs/2504.018482. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research3. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts — Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al., 2024
https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+frontier+AI+R%26D+capabilities+of+language+model+agents+against+human+experts4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark5. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — Qian Huang, Jian Vora, Percy Liang, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=MLAgentBench%3A+Evaluating+Language+Agents+on+Machine+Learning+Experimentation6. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering — Jun Shern Chan et al., 2024
https://scholar.google.com/scholar?q=MLE-bench%3A+Evaluating+Machine+Learning+Agents+on+Machine+Learning+Engineering7. EXP-Bench: Can AI Conduct AI Research Experiments? — Patrick Tser Jern Kon et al., 2025
https://scholar.google.com/scholar?q=EXP-Bench%3A+Can+AI+Conduct+AI+Research+Experiments%3F8. MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research — Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, Bryan Hooi, 2025
https://scholar.google.com/scholar?q=MLR-Bench%3A+Evaluating+AI+Agents+on+Open-Ended+Machine+Learning+Research9. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? — Christine Ye et al., 2025
https://scholar.google.com/scholar?q=ReplicationBench%3A+Can+AI+Agents+Replicate+Astrophysics+Research+Papers%3F10. Can Large Language Models Be an Alternative to Human Evaluations? — Cheng-Han Chiang, Hung-yi Lee, 2023
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Be+an+Alternative+to+Human+Evaluations%3F11. RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following — Tianjun Pan et al., 2026
https://arxiv.org/abs/2603.2513312. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan et al., 2024
https://arxiv.org/abs/2410.1278413. When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation — Abeer Badawi et al., 2025
https://arxiv.org/abs/2510.1903214. A Dataset For Computational Reproducibility — Lazaro Costa, Susana Barbosa, Jacome Cunha, 2025
https://arxiv.org/abs/2504.0868415. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers — Yanzheng Xiang et al., 2025
https://arxiv.org/abs/2504.0025516. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding — Deming Ding et al., 2026
https://arxiv.org/abs/2601.1034317. ContextBench: A Benchmark for Context Retrieval in Coding Agents — Han Li et al., 2026
https://arxiv.org/abs/2602.0589218. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp319. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp320. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp321. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3Interactive Visualization: PaperBench: Can AI Replicate AI Research?