ProgramBench: Frontier Agents Fail to Rebuild Software from Scratch

Yang, Lieret, Ma, Thakkar, Pedchenko, Sootla, McMilin, Yin, Hou, Synnaeve, D. Yang, Press · Meta FAIR / Meta TBD / Stanford / Harvard · May 2026
arXiv 2605.03546 200 tasks9 frontier models 0 / 200 fully resolved Interactive page
The self-hosting pipeline
Step through: the same trick that builds the tests also grades fairly.
How much structure does each benchmark hand over?
Hover cells. Red = the model must decide it itself.
Qualitative reading of the episode's comparison, not a measured quantity.
Tasks passing at a chosen test-pass threshold
Drag to 100 for "Resolved"; 95 is "Almost". Nine models, 1,000 steps, 6 h, no internet.
Headline numbers (0 Resolved; Opus 4.7 Almost on 6/200) are from the paper. Per-task rates here are illustrative mock data.
Difficulty is model-agnostic
Mean tests passed per task (mock). Small CLIs easy, FFmpeg / php-src / typst hard for everyone.
Opus 4.7 funnel
Tasks at or above each threshold (mock)
What the models build: shape of the code
Rectangle area is proportional to lines of code. Solutions passing ≥75% of tests.
File layouts are illustrative; medians (1,173 vs 3,068 lines; depth 1 vs 2; about one third the files and functions) are from the paper.
Parnas 1972: decompose by hidden decisions, not by steps
Pick a change and count the modules it forces you to edit.
Black-box grading: behavior, not source
Pick a test. Candidate can be any language; only stdout, exit code and file effects count.
All-or-nothing scoring
Click cells to flip pass/fail. Resolved needs every test.
Generated vs native test coverage
Line coverage of the agent-built suites
Linter effect on weak assertions
Each dot = a test that runs against a deliberately broken program (270 shown)
Coverage climbs as the agent probes
Illustrative trajectory ending at the reported 79.7%
Forced implementation language
Only the five models named in the episode
GPT +4.2 each and Python 36%→51% are reported. Claude drop sizes are illustrative.
Internet ablation and the cheating judges
Flags are mostly source-code lookup (repo inferred from --help, shallow clone)
Trajectory length per model (commands)
Log scale. Same 0% Resolved, wildly different paths. Hover a dot.
Medians for GPT 5.4 (17) and Sonnet 4.6 (868) are reported; other medians and run scatter are illustrative.
Caveats worth carrying
Hover for emphasis.
References
  1. ProgramBench: Can Language Models Rebuild Programs From Scratch? — Yang, Lieret, Ma, et al., 2026
  2. On the Criteria To Be Used in Decomposing Systems into Modules — Parnas, 1972
  3. Measuring Coding Challenge Competence With APPS — Hendrycks et al., 2021
  4. SWE-bench — Jimenez et al., 2024
  5. SWE-agent — Yang et al., 2024
  6. QuickCheck — Claessen, Hughes, 2000
  7. Finding and Understanding Bugs in C Compilers — Yang, Chen, Eide, Regehr, 2011
  8. CodeT — Chen et al., 2022
  9. Commit0 — Zhao et al., 2024
  10. DevBench — Li et al., 2024
  11. NL2Repo-bench — Ding et al., 2026
  12. LLM4Decompile — Tan et al., 2024
  13. Position: Humans Are Missing from AI Coding Agent Research — Wang et al., 2026