Can an AI agent tell a good research idea from a dead-on-arrival one before it burns a single GPU-hour? SoundnessBench builds 1,099 ICLR proposals labeled by reviewers' soundness sub-scores — not acceptance — and finds twelve frontier LLMs produce a 74% false-positive rate, repeatedly rating flawed proposals as sound. Aggressive fault-hunting prompts flip the verdict on the same proposals, pointing at framing sensitivity rather than a missing-knowledge problem.
Prior research-agent benchmarks score execution. SoundnessBench tests the step before that — scientific triage, the first-gate call on whether an idea is worth running at all.
Hover the bars. When a proposal was genuinely flawed, models called it "sound" nearly three out of four times — while still catching most genuinely good proposals.
Collection → Labeling → Extraction → Verification → Assembly. The verification step — a retrieval-backed audit of atomic claims — is what most benchmark papers skip.
66.93% of extracted candidate proposals survive the fidelity audit at a 0.7 claim-matching threshold — a third are cut before ever reaching the benchmark.
458 low-soundness vs. 641 high-soundness proposals, thresholded on the reviewer soundness sub-score (≥3 high, ≤2 low, middle scores dropped).
No single subfield's writing style drives the result — coverage spans RL, generative modeling, optimization, vision, and more, across five conference years.
Toggle the metric. Notice that alignment-tuned frontier models (Gemini-3.1-Pro, Claude-Opus-4.6) don't reliably outperform a 70B open checkpoint on catching flawed proposals.
"Aggressive" prompts tell the model to default to low-soundness unless the case is airtight. False positives collapse — but so does recall of genuinely good proposals.
The paper only samples two operating points. The dashed region is the unmeasured trade-off curve — a real sweet spot may exist between them, but was never tested.
Under standard prompting, high-soundness recall climbs with parameter count — but low-soundness recall falls the whole way. Bigger gets more optimistic, not more skeptical.
GPT-5.4 on 100 high-soundness proposals with a blunt hypothesis–experiment mismatch injected. It catches an obvious break — but that's a narrower result than "the model reasons well," since real flaws are subtler.
Human audit, contamination split, identifier stripping, surface-feature baseline — none of these confounders account for the optimism bias.