AI Post Transformers · Episode Companion

SoundnessBench: Exposing AI Reviewers' Blind Spots

arXiv 2605.30329 Sy-Tuyen Ho et al. · University of Maryland, College Park Posted May 28, 2026 1,099 ICLR proposals 12 frontier models tested

Can an AI agent tell a good research idea from a dead-on-arrival one before it burns a single GPU-hour? SoundnessBench builds 1,099 ICLR proposals labeled by reviewers' soundness sub-scores — not acceptance — and finds twelve frontier LLMs produce a 74% false-positive rate, repeatedly rating flawed proposals as sound. Aggressive fault-hunting prompts flip the verdict on the same proposals, pointing at framing sensitivity rather than a missing-knowledge problem.

74%
mean false-positive rate on flawed proposals (standard prompting)
26%
mean low-soundness recall
91.8%
mean high-soundness recall
1,099
verified proposals, ICLR 2022–2026

Where SoundnessBench sits gap vs. prior benchmarks

Prior research-agent benchmarks score execution. SoundnessBench tests the step before that — scientific triage, the first-gate call on whether an idea is worth running at all.

The core finding standard prompting, 12 models

Hover the bars. When a proposal was genuinely flawed, models called it "sound" nearly three out of four times — while still catching most genuinely good proposals.

Five-step construction pipeline

Collection → Labeling → Extraction → Verification → Assembly. The verification step — a retrieval-backed audit of atomic claims — is what most benchmark papers skip.

Verification funnel

66.93% of extracted candidate proposals survive the fidelity audit at a 0.7 claim-matching threshold — a third are cut before ever reaching the benchmark.

Label composition

458 low-soundness vs. 641 high-soundness proposals, thresholded on the reviewer soundness sub-score (≥3 high, ≤2 low, middle scores dropped).

Coverage across 16 ML subfields

No single subfield's writing style drives the result — coverage spans RL, generative modeling, optimization, vision, and more, across five conference years.

Model leaderboard standard prompting

Toggle the metric. Notice that alignment-tuned frontier models (Gemini-3.1-Pro, Claude-Opus-4.6) don't reliably outperform a 70B open checkpoint on catching flawed proposals.

strong (recall > 70%) mid (30–70%) weak (< 30%)

Standard vs. Aggressive prompting same models, same proposals

"Aggressive" prompts tell the model to default to low-soundness unless the case is airtight. False positives collapse — but so does recall of genuinely good proposals.

Low-soundness recall (catches flaws) High-soundness recall (catches good ideas)

Two points, not a curve

The paper only samples two operating points. The dashed region is the unmeasured trade-off curve — a real sweet spot may exist between them, but was never tested.

Scale doesn't rescue it — Qwen3.5 family

Under standard prompting, high-soundness recall climbs with parameter count — but low-soundness recall falls the whole way. Bigger gets more optimistic, not more skeptical.

Adversarial injection test

GPT-5.4 on 100 high-soundness proposals with a blunt hypothesis–experiment mismatch injected. It catches an obvious break — but that's a narrower result than "the model reasons well," since real flaws are subtler.

Robustness checks that didn't explain it away

Human audit, contamination split, identifier stripping, surface-feature baseline — none of these confounders account for the optimism bias.

References