← All episodes ProgramBench: Frontier Agents Fail to Rebuild Software from Scratch

ProgramBench: Frontier Agents Fail to Rebuild Software from Scratch

Sep 19, 2026
This episode examines ProgramBench, a new benchmark testing whether frontier language models can rebuild working software from just a compiled binary and its documentation, with no source code, scaffolding, or prescribed architecture to work from. Across nine frontier models and two hundred tasks spanning CLI tools up to FFmpeg, SQLite, and the PHP interpreter, zero tasks were fully resolved, exposing a stark gap between patching existing code and making the upstream architectural decisions—language choice, module boundaries, data structures, error handling—that real software design requires. The discussion unpacks the benchmark's clever self-hosting trick: an LLM agent fuzzes the reference binary to build a behavioral test suite, enabling black-box grading that judges programs by what they do rather than how closely they mimic the original source. Framing the results against Parnas's classic work on information hiding and modular decomposition, the conversation argues that current agents default to monolithic, unstructured code once nobody hands them a skeleton to fill in. It's a sobering data point for anyone assuming coding agents are close to functioning as autonomous software architects rather than sophisticated patch-writers.
Sources:
1. ProgramBench: Can Language Models Rebuild Programs From Scratch? — John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, Ofir Press, 2026
http://arxiv.org/abs/2605.03546
2. On the Criteria To Be Used in Decomposing Systems into Modules — D. L. Parnas, 1972
https://scholar.google.com/scholar?q=On+the+Criteria+To+Be+Used+in+Decomposing+Systems+into+Modules
3. Measuring Coding Challenge Competence With APPS — Dan Hendrycks, Steven Basart, Saurav Kadavath, et al., 2021
https://scholar.google.com/scholar?q=Measuring+Coding+Challenge+Competence+With+APPS
4. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Carlos E. Jimenez, John Yang, Alexander Wettig, et al., 2024
https://scholar.google.com/scholar?q=SWE-bench%3A+Can+Language+Models+Resolve+Real-World+GitHub+Issues%3F
5. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — John Yang, Carlos E. Jimenez, Alexander Wettig, et al., 2024
https://scholar.google.com/scholar?q=SWE-agent%3A+Agent-Computer+Interfaces+Enable+Automated+Software+Engineering
6. QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs — Koen Claessen, John Hughes, 2000
https://scholar.google.com/scholar?q=QuickCheck%3A+A+Lightweight+Tool+for+Random+Testing+of+Haskell+Programs
7. Finding and Understanding Bugs in C Compilers — Xuejun Yang, Yang Chen, Eric Eide, John Regehr, 2011
https://scholar.google.com/scholar?q=Finding+and+Understanding+Bugs+in+C+Compilers
8. CodeT: Code Generation with Generated Tests — Bei Chen, Fengji Zhang, Anh Nguyen, et al., 2022
https://scholar.google.com/scholar?q=CodeT%3A+Code+Generation+with+Generated+Tests
9. Commit0: Library Generation from Scratch — Wenting Zhao, Nan Jiang, Celine Lee, Justin T Chiu, Claire Cardie, Matthias Gallé, Alexander M Rush, 2024
https://scholar.google.com/scholar?q=Commit0%3A+Library+Generation+from+Scratch
10. DevBench: A Comprehensive Benchmark for Software Development — Bowen Li, Wenhan Wu, Ziwei Tang, et al. (incl. John Yang, Ofir Press), 2024
https://scholar.google.com/scholar?q=DevBench%3A+A+Comprehensive+Benchmark+for+Software+Development
11. NL2Repo-bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents — Jingzhe Ding et al., 2026
https://scholar.google.com/scholar?q=NL2Repo-bench%3A+Towards+Long-Horizon+Repository+Generation+Evaluation+of+Coding+Agents
12. LLM4Decompile: Decompiling Binary Code with Large Language Models — Hanzhuo Tan, Qi Luo, Jing Li, Yuqun Zhang, 2024
https://scholar.google.com/scholar?q=LLM4Decompile%3A+Decompiling+Binary+Code+with+Large+Language+Models
13. Position: Humans Are Missing from AI Coding Agent Research — Zora Zhiruo Wang, John Yang, Kilian Lieret, et al., 2026
https://scholar.google.com/scholar?q=Position%3A+Humans+Are+Missing+from+AI+Coding+Agent+Research
Interactive Visualization: ProgramBench: Frontier Agents Fail to Rebuild Software from Scratch