1 00:00:01,000 --> 00:00:36,674 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper made me nervous about how much I trust AI reviewers — it's SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? That's Sy-Tuyen Ho et al., four co-authors total, out of the University of Maryland, College Park, posted to arXiv May 28th, 2026. The question: before an AI research agent burns a single GPU-hour, can it actually tell whether the underlying idea is any good? 2 00:00:36,674 --> 00:00:57,199 [Dr. Ada Shannon] What sold me wasn't the size, though 1,099 proposals is solid — it was the theoretical grounding. A lot of 'can AI judge AI' papers feel like vibes-based demos, one number and a shrug. This one ties straight into existing literature on sycophancy and prompt fragility, so when they find a failure, they can actually explain why it's happening, not just that it happens. 3 00:00:57,199 --> 00:01:15,324 [Hal Turing] So walk me through what's actually missing right now, because I'd have naively assumed the opposite ordering — that if we already trust AI agents to run entire experiments end to end, judging whether a proposal is even worth trying would be the easy, solved part. Apparently not? 4 00:01:15,324 --> 00:01:47,224 [Dr. Ada Shannon] That's the gap they name. Existing research-agent benchmarks — MLE-Bench, PaperBench, InnovatorBench — evaluate execution. Hand the model a task, let it run code, score the results against a leaderboard or a human rubric. None test the step before that: what they call scientific triage, the first-gate call on whether a hypothesis and experimental design are even worth running. Skip it, and AI agents don't accelerate science — they risk industrializing bad science, running well-executed experiments on ideas that were dead on arrival. 5 00:01:47,224 --> 00:02:04,250 [Hal Turing] Okay, but how do you even build a rigorous benchmark for that? 'Is this idea good' sounds unavoidably subjective — one reviewer's fatal flaw is another reviewer's minor quibble. I'd expect that to fall apart the second you try to turn it into labeled data. 6 00:02:04,250 --> 00:02:33,125 [Dr. Ada Shannon] They went to the source: ICLR review data. Proposals reconstructed from real submissions, 1,099 in the final set, labeled not by acceptance or star rating but by the reviewers' soundness sub-scores specifically — results and outcomes stripped out so the model can't infer the verdict from the ending. Then they ran twelve frontier LLMs against it under standard prompting. Headline number: a mean false-positive rate of seventy-four percent. When a proposal was genuinely flawed, models called it sound almost three-quarters of the time. 7 00:02:33,125 --> 00:02:42,925 [Hal Turing] Oh wait wait wait — seventy-four percent? That's not a model being slightly too generous, that's a coin flip rigged in exactly the wrong direction. 8 00:02:42,925 --> 00:03:13,875 [Dr. Ada Shannon] And they didn't stop at one prompt. They also tested what they call aggressive prompting — instead of neutral evaluation, you explicitly tell the model to default to 'low soundness' unless the case is airtight. That flips the failure mode: false positives drop, but now it starts rejecting genuinely good proposals. And to be precise about the claim itself — this isn't 'we predict exact peer review outcomes.' It's scoped to recoverable proposal-stage soundness: how much of a reviewer's judgment survives in the proposal text alone, before any results exist. 9 00:03:13,875 --> 00:03:32,050 [Hal Turing] That prompt-sensitivity swing sounds a lot like something adjacent to sycophancy — the model bending its verdict based on how you frame the ask rather than sticking to some stable internal judgment. Is that actually what's going on here, or am I pattern-matching too aggressively? 10 00:03:32,050 --> 00:04:10,625 [Dr. Ada Shannon] You're not wrong to reach for that word. The paper leans on Sharma, Tong, Korbak and colleagues at Anthropic — 'Towards Understanding Sycophancy in Language Models,' 2023 — which found RLHF-tuned assistants shift their answers toward what a user seems to want, because the human preference data training the reward model rewards agreeable, confident responses over correct-but-critical ones. The model isn't lying deliberately. It's a learned shortcut — validation gets rewarded more than skepticism does. Apply that to judging a proposal cold, and the default under neutral prompting is exactly what you'd predict: lean positive, because positive is what usually gets praised in training. 11 00:04:10,625 --> 00:04:31,949 [Hal Turing] I actually disagree that sycophancy is the right lens here, Ada. Sycophancy is usually about a user pushing back and the model folding. Nobody's pushing back in this setup — it's a cold read with no argument attached. Feels more like the models just don't have good priors on what makes ML methodology sound, full stop. 12 00:04:31,949 --> 00:04:58,199 [Dr. Ada Shannon] No, I don't buy that distinction. The mechanism doesn't need explicit pushback — it needs a reward model trained to prefer agreeable outputs, and a baseline where 'promising' reads as more agreeable than 'fatally flawed.' The paper's careful too — they frame RLHF as consistent with the pattern, not a proven sole cause. But the aggressive-prompting result is the tell: same proposal, same model, opposite verdict, purely from how hard you tell it to hunt for flaws. Knowledge gaps don't flip on command like that. 13 00:04:58,199 --> 00:05:20,724 [Hal Turing] Fair — and honestly that's the scarier framing anyway. A model that doesn't know better is a capability problem you fix with more training data. A model whose verdict swings on framing is a reliability problem, and that's exactly what you don't want sitting as the gatekeeper deciding which ideas an autonomous research agent gets to burn compute on. 14 00:05:20,724 --> 00:05:33,174 [Hal Turing] Okay, I want the mechanics here, not just the headline number. Walk me through how SoundnessBench actually gets built, because if the construction is sloppy, seventy-four percent doesn't mean much. 15 00:05:33,174 --> 00:06:10,224 [Dr. Ada Shannon] Five steps. Collection: pull the ICLR corpus, keep only papers with strong reviewer agreement — confidence at least three, soundness standard deviation under 0.15. Labeling: use the soundness sub-score itself — three or higher is high, two or lower is low, middle scores dropped. Extraction: pull a near-verbatim proposal from the source PDF — abstract, related work, hypothesis, experimental design — with every result and acceptance cue stripped out. Verification: a retrieval-backed audit checking atomic claims against the source paper at a 0.7 threshold, to catch extraction drift. Assembly: whatever survives becomes the final benchmark. 16 00:06:10,224 --> 00:06:23,674 [Hal Turing] That verification step is the part most benchmark papers skip — actually auditing their own extraction instead of just asserting it's clean. So what survives, and what does the final set look like? 17 00:06:23,674 --> 00:06:50,474 [Dr. Ada Shannon] 66.93 percent pass the audit — a third of candidates cut on fidelity alone. What's left is 1,099 proposals, 458 low soundness, 641 high, spread across sixteen ML subfields — reinforcement learning, generative modeling, optimization, computer vision, all of it — pulled from ICLR 2022 through 2026. Not one subfield's quirky writing style driving this, and not one conference year either. 18 00:06:50,474 --> 00:06:54,674 [Hal Turing] And the models being tested — who's actually in the hot seat? 19 00:06:54,674 --> 00:07:38,649 [Dr. Ada Shannon] Twelve frontier models, spread across families: GPT-4o, GPT-5.4, GPT-5.4-Mini, Claude-Opus-4.6, Claude-Sonnet-4.6, Gemini-2.5-Pro, Gemini-3-Flash, Gemini-3.1-Pro, Qwen3.5 at 27B and 122B, LLaMA-3.3-70B, and Kimi-Linear-48B — closed and open source, reasoning and non-reasoning. That 74 percent false-positive rate breaks down as mean low-soundness recall of 26 percent against mean high-soundness recall of 91.8. Nine of twelve — including Gemini-3.1-Pro and Claude-Opus-4.6, genuinely strong models — label over 70 percent of low-soundness proposals as sound. LLaMA-3.3-70B and GPT-4o are worst, at 98 and 94.5 percent. 20 00:07:38,649 --> 00:07:57,524 [Hal Turing] Wait, hold on — so it's not a 'weaker model doesn't know better' story. The expensive, heavily RLHF'd frontier models are just as bad, sometimes worse, than a 70-billion open LLaMA checkpoint. I'd have bet money alignment tuning at least bought some default caution. 21 00:07:57,524 --> 00:08:36,750 [Dr. Ada Shannon] That's why the aggressive prompt matters — false positives drop from 74 percent to 19.9. But high-soundness recall collapses with it, to 36.1 percent mean. GPT-5.4 and GPT-5.4-Mini go to the extreme — recall of 0.0 and 0.2 percent, rejecting almost everything. Scale doesn't rescue it either — the Qwen3.5 family, 2B up to 122B, same recipe throughout: high-soundness recall climbs with scale under standard prompting, but low-soundness recall drops the whole way, from 31 percent at 2B to 19.2 at 35B. Bigger gets more optimistic, not less. 22 00:08:36,750 --> 00:09:02,299 [Hal Turing] Here's where I'll push toward the optimistic read — tell me if I'm wrong. The robustness section is thorough: human audit on leakage and label validity, an ICLR-2026-only split for contamination, identifier stripping, even a surface-feature baseline on proposal length and risk-factor counts. None of those confounders explain the pattern away. Doesn't that mean the benchmark itself is solid, and the number's trustworthy? 23 00:09:02,299 --> 00:09:31,250 [Dr. Ada Shannon] No, I actually disagree with you there, Hal. The surface-feature baseline is the part that should worry you more — it fails in the opposite direction, over-rejecting high-soundness proposals. Two failure modes pointing at each other, and ruling out confounders doesn't make this reassuring — it means the optimism bias lives inside the model's actual judgment, not a dataset shortcut you could patch. A confound is fixable. A model reading real content and still defaulting to approval is a capability gap, and that's strictly worse news. 24 00:09:31,250 --> 00:09:42,725 [Hal Turing] Fair, but isn't 'it's reading content, not gaming a proxy' at least partial good news? If it were pure surface-feature gaming, the whole benchmark would be measuring nothing real. 25 00:09:42,725 --> 00:10:26,100 [Dr. Ada Shannon] Engaging badly is still the headline. The one genuinely reassuring result is the adversarial injection test: inject a blunt hypothesis-experiment mismatch into 100 high-soundness proposals, and GPT-5.4's approval rate craters from 77 percent to 1. So there's a floor — it catches an obvious break. What it can't do is catch the same flaw in its natural, subtler form, sitting quietly in a real low-soundness submission. A hypothesis-experiment mismatch that blunt is basically a strawman; real low-soundness proposals are underpowered experiments, near-miss baselines, metrics that sound right but aren't. And it's one model, GPT-5.4, on a hundred proposals — the paper doesn't even specify which prompt regime was used. That's narrower relief than 'the benchmark is solid.' 26 00:10:26,100 --> 00:10:59,250 [Hal Turing] That's exactly where I want to push, Ada. Sitting with it, the whole standard-versus-aggressive framing bugs me. That's two points on a curve, not a curve. They never ask the model for a soundness probability and sweep the threshold, never test intermediate degrees of strictness. So when they report the shift from false positives to false negatives — is that a fundamental incalibrability finding, or did they just pick one arbitrarily harsh prompt and land in the opposite failure mode? A real operating-characteristic curve might show a sweet spot they never sampled. 27 00:10:59,250 --> 00:11:38,524 [Dr. Ada Shannon] Legitimate gap, and the two points we do have rest on a ground truth that's shakier than the paper lets on. The label is a single scalar — mean reviewer soundness, thresholded at three versus two. ML conference reviewer scores are notoriously noisy, non-blind to author identity half the time, with documented low inter-reviewer reliability. They filter for agreement — confidence at least three, standard deviation under 0.15 — which is the right instinct. But they never report what fraction of ICLR reviews actually survive that filter. If it's a narrow, unusually-agreeable slice, those 1,099 proposals might represent 'cases reviewers happened to agree on,' not soundness in general. 28 00:11:38,524 --> 00:11:51,299 [Hal Turing] Wait, hold on — that snaps right into the sycophancy story from earlier, though. So the optimism bias is just sycophancy wearing a lab coat. Doesn't that basically settle the mechanism? 29 00:11:51,299 --> 00:12:22,000 [Dr. Ada Shannon] I actually disagree with how you're stating that, Hal. Sharma et al. is the right citation and the mechanism is plausible, but 'consistent with' isn't 'caused by,' and the authors never test RLHF as the driver. They'd need a base model versus a preference-tuned checkpoint on the same backbone, showing the gap opens after tuning. Without that ablation, sycophancy is a compelling story, not a demonstrated cause. It could just as easily be pattern-matching on writing quality — polished proposals reading as sound regardless of how the model was tuned. 30 00:12:22,000 --> 00:12:57,950 [Hal Turing] Okay, fair — but then explain the Qwen3.5 scale result to me, Ada. Bigger models get more optimistic under standard prompting, not less, and that's within one model family, same training recipe, just parameter count climbing from two billion to over a hundred. If this were purely reward-hacking from preference tuning, you'd expect it to plateau somewhere. Doesn't rising optimism with scale point more toward 'the model has absorbed more surface-plausibility patterns from its training distribution' than 'more RLHF pressure applied'? 31 00:12:57,950 --> 00:13:36,025 [Dr. Ada Shannon] Could be both, and the paper's own surface-feature baseline can't rule it out — it only tests proposal length and risk-factor counts, not writing polish, so a stylistic confound stays open. We're both speculating past the paper at that point. There's a complementary angle worth naming, though — Si, Yang, and Hashimoto's 2025 Stanford study found LLM-generated ideas score novel but weak on feasibility, and their Stanford follow-up, the Ideation-Execution Gap paper, showed proposal-stage promise often doesn't survive contact with execution. Same models, potentially bad at generating sound proposals and bad at judging them — a shared root cause nobody's isolated yet. 32 00:13:36,025 --> 00:14:08,450 [Hal Turing] Which gets at the framing gap for me. The title asks whether your 'AI scientist' can tell good ideas from bad, full stop — but this is ICLR-only, reviewer-soundness-only, binary, two prompts. Credit to the authors, they're careful about that scoping in the text itself. But if you're running an autonomous research loop today, the practical takeaway is blunt: don't let one LLM call gate whether an experiment runs. Standard prompting burns compute on bad ideas; aggressive prompting starves you of the good ones. 33 00:14:08,450 --> 00:14:46,825 [Dr. Ada Shannon] Right, and the fix they gesture at — targeted training, calibration, human-in-the-loop — is underspecified. A more concrete path is probably a dedicated verifier model trained on this kind of labeled data, rather than prompting a general chat model and hoping it acts like a critic. That's the process-reward-model playbook that's worked in reasoning tasks, a strange omission given how squarely this is a first-gate verification problem. I'd also push this past ICLR — biology, chemistry, social-science proposals don't look like ML papers structurally, and this benchmark's proposal format is borrowed straight from Yamada et al.'s AI Scientist-v2 paper, 2025. 34 00:14:46,825 --> 00:15:14,650 [Hal Turing] So here's where I land: current LLMs are not safe standalone gatekeepers for research triage. The failure is real and survives their controls, but how much is prompt artifact versus RLHF versus scale-driven pattern-matching is still genuinely open. That's SoundnessBench — a careful benchmark asking a narrower question than its title implies, and getting an uncomfortable answer even within that narrower scope. Thanks for listening, everyone — we'll catch you next time.