1 00:00:01,000 --> 00:00:35,625 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And today's paper is NaturalReasoning: Reasoning in the Wild with 2.8 Million Challenging Questions. Lead author Weizhe Yuan, eleven authors total, Weizhe Yuan et al., out of Meta and New York University. It first hit arXiv back in February twenty twenty-five, the version we're reading from is dated November sixth, twenty twenty-five, and it's a NeurIPS twenty twenty-five paper. Ada, you flagged this one for us, what got you hooked? 2 00:00:35,625 --> 00:01:07,650 [Dr. Ada Shannon] Honestly, it wasn't the two point eight million number, everybody claims scale these days. What got me is they actually argue why this should work, not just show a chart where the line goes up. The whole reasoning-data field is bottlenecked by verification right now. If you can't automatically check whether an answer is right, it's hard to build a clean training signal around it, so everyone clusters around math and code, the checkable domains. This paper's pitch is stop starting from benchmarks, start from the text itself. That's a structural argument about where the data comes from, not an engineering flex. 3 00:01:07,650 --> 00:01:39,525 [Hal Turing] Okay, quick background for anyone not neck deep in this. Over the past year we got a new class of deliberation-style models, OpenAI's o1, DeepSeek's R1. They're trained with reinforcement learning where the reward comes from checking the final answer against a known correct answer, not a person rating the response. The model sits with a problem, produces a long chain of reasoning, and gets scored purely on whether it landed on the right answer. It's worked remarkably well. But doesn't that put a hard ceiling on what you can even train on? 4 00:01:39,525 --> 00:02:14,574 [Dr. Ada Shannon] Exactly the trap. It's called Reinforcement Learning from Verifiable Rewards, RLVR, the reward is an automated check against ground truth instead of a learned reward model or a human rater. And here's the consequence before the mechanism: you can only train this way where correctness is checkable by a script, math problems, code with unit tests, chess. The moment you want reasoning training for an economics argument or an open-ended physics derivation with no single numeric answer, there's no verifier to write. The recent reasoning gains are real, but they're math-and-code shaped. 5 00:02:14,574 --> 00:02:47,199 [Hal Turing] So that's the gap this paper tries to plug, and it leans hard on synthetic data generation, having an LLM manufacture the training data instead of paying humans to write or curate it by hand. But the specific trick they use is called question backtranslation, a name borrowed from a totally different corner of NLP. If you've ever touched machine translation, backtranslation already means something to you there. So is this the same idea repurposed, or just a name collision? I'm genuinely curious how far the analogy holds up. 6 00:02:47,199 --> 00:03:21,074 [Dr. Ada Shannon] Same core idea, pointed somewhere new. Sennrich, Haddow, and Birch, out of the University of Edinburgh, introduced backtranslation for machine translation in 2016: you've got plenty of good text in your target language but little parallel data, so you train a target-to-source model and use it to manufacture synthetic source sentences for real target text you already have. The trick generalizes. Instead of collecting question-answer pairs from humans, you start from real, high-quality text, treat that as the answer side, and ask a strong model to work backward and invent the question it would be answering. 7 00:03:21,074 --> 00:03:35,399 [Hal Turing] Oh wait wait wait, so instead of a person sitting down writing 'here's a hard problem,' the model looks at a dense paragraph from, say, a physics text and asks itself what question this paragraph would be a good answer to? 8 00:03:35,399 --> 00:04:14,474 [Dr. Ada Shannon] Exactly. NaturalReasoning does exactly that: Llama-3.3-70B-Instruct sweeps pretraining corpora, DCLM-baseline and FineMath specifically, flags documents with real reasoning content, then backtranslates the strong ones into questions with reference answers grounded in the source passage. It descends from Li et al.'s Self-Alignment with Instruction Backtranslation, out of Meta in 2023 — two of that paper's authors, Xian Li and Jason Weston, are on this one too. That earlier paper backtranslated general web text into instructions; this one hunts specifically for reasoning-rich documents, across math, physics, CS, economics, social science, well beyond the usual benchmark-shaped stuff. 9 00:04:14,474 --> 00:04:29,724 [Hal Turing] So once you've got two point eight million backtranslated questions sitting there, how do you actually turn that into a better model? I'm guessing distillation is where this goes, but is that the whole story, or is there something trickier underneath? 10 00:04:29,724 --> 00:05:16,224 [Dr. Ada Shannon] Distillation's the baseline move: take a strong teacher, here Llama-3.3-70B-Instruct itself, have it generate long chain-of-thought answers, and fine-tune a student to reproduce them. Same supervised fine-tuning loop people have run for years, the teacher's just good at reasoning now. But they also test skipping external reward models entirely with self-rewarding, from Yuan and Weston's own 2024 paper out of Meta and NYU, where one model both generates answers and judges its own answers well enough to build a preference signal, no separate reward model, no extra human labels. So the shape so far: two point eight million questions, one model family doing the annotating, writing, teaching, and judging, end to end. Whether that holds up once you look at how it verifies its own output, that's where it gets interesting. 11 00:05:16,224 --> 00:05:52,799 [Dr. Ada Shannon] verify whether a correct answer can actually be pulled from that same document before it becomes a reference answer. And every one of the 2.8 million questions also gets a separate response from Llama-3-70B-Instruct specifically, that's the teacher target for the distillation you just asked about. So it's really four LLM passes per document: rate it on Problem Completeness, Complexity and Technical Depth, Correctness, and Thinking and Reasoning, synthesize the question if it clears the bar, verify the reference answer, then generate the teacher response. No human touches any of it. 12 00:05:52,799 --> 00:06:29,574 [Hal Turing] Okay so that's the full pipeline, and now I want to poke at it, because something's been nagging me. Llama-3.3-70B-Instruct, or close cousins in the same family, decides which documents have 'high reasoning content,' writes the question, checks whether a reference answer can be pulled out — and then Llama-3-70B-Instruct specifically writes the teacher response you distill into the student. Same family doing the judging, the writing, the grading, and the teaching. How do you rule out that the ninety-three percent 'high quality' number is really just Llama telling you it likes Llama's own taste in questions? 13 00:06:29,574 --> 00:07:01,799 [Dr. Ada Shannon] Fair worry, and credit where due — the quality scoring in Section 3 does pull in three separate models, DeepSeek-R1-Distill-Qwen-32B, Qwen2.5-72B-Instruct, and Llama-3-70B-Instruct, specifically so one family's preferences don't dominate. But the diversity stops there. Document-mining, question synthesis, reference-answer derivation, and the difficulty proxy all stay inside the Llama family. The only fully independent check is that hundred-question human study, and for a paper this size, that's a thin reed to hang the whole quality claim on. 14 00:07:01,799 --> 00:07:35,324 [Hal Turing] And that difficulty number is exactly where I'd want independent verification, because it's just response length — longer chain-of-thought from Llama3.3-70B-Instruct equals harder question, full stop. But length and difficulty aren't the same axis. Ask someone to 'discuss the ethics of gene editing' and you'll get four hundred rambling words that aren't hard to produce, just open-ended. Did they ever check that against pass rate under repeated sampling, or whether independent judges agree the question is actually hard? 15 00:07:35,324 --> 00:08:06,749 [Dr. Ada Shannon] No, and that's the gap — no cross-validation against sampling-based difficulty, no cross-judge agreement check, length is the entire proxy end to end. Which loops back into the human eval, because that's the one place a length-inflation artifact could've been caught, and it's underpowered. A hundred questions per dataset, two annotators, scores averaged, and nowhere do they report inter-rater agreement — no Cohen's kappa, nothing. NaturalReasoning scores six four five, WebInstruct five nine two — is that gap real, or noise at n equals a hundred? We can't tell from what's published. 16 00:08:06,749 --> 00:08:38,099 [Hal Turing] Oh wait wait wait — hold on, that's actually worse than you're giving it credit for. Those same three LLM judges scoring the automatic quality numbers include Llama-3-70B-Instruct, which is literally one of the models used to build the dataset. So the human study is supposed to be independent corroboration of the automatic scores, and the automatic scores themselves aren't fully independent of the pipeline that produced the data. That's two layers of the same circularity stacked on top of each other. 17 00:08:38,099 --> 00:09:20,199 [Dr. Ada Shannon] And I'll grant the ranking direction matches across both evals — at least consistent. But consistency isn't the same as being right; if the whole apparatus shares one blind spot, both measurements agree and still miss it. Set that aside though — there's a comparison worth making: WebInstruct, the closest real competitor, also multi-domain, also mined at web scale. But WebInstruct recalls thirteen million existing question-answer pairs with rule-based extraction — Yue and colleagues, University of Waterloo, the MAmmoTH2 paper, 2024. NaturalReasoning only synthesizes two point eight million, LLM-generated end to end. You're trading scale for precision: fewer questions, but each one is genuinely novel, and it scores higher on quality. 18 00:09:20,199 --> 00:10:07,749 [Hal Turing] That tradeoff is basically the whole synthetic-data debate in miniature — more data that's noisier, or less that's cleaner. And it's not new. This backtranslation approach traces to a 2023 paper out of Meta AI, Self-Alignment with Instruction Backtranslation, Xian Li, Ping Yu, Chunting Zhou and coauthors — and Xian Li is the senior author on today's paper too, so there's a direct throughline. That paper backtranslated instructions from ordinary web text; NaturalReasoning does the same trick pointed at reasoning-annotated documents instead. What it doesn't revisit is that older paper's known failure mode: self-augmented data can drift, because the model doing the backtranslation is also the one deciding what counts as 'good.' 19 00:10:07,749 --> 00:10:55,949 [Dr. Ada Shannon] Same tension we've circled all episode, just a different layer. Practical takeaway though — despite the skepticism, there's real signal here. If you need reasoning data without paying for human annotation, this is a legitimate option now: a hundred thousand randomly sampled questions gets a 70B model most of the way to matching DeepSeek-R1-Distill-Llama-70B on an eighth of the data. On self-training, self-scoring rivals or beats a trained external reward model, which matters if you can't afford to train one. They're upfront about limits too — the actual RLVR experiment, verifiable rewards with their General Verifier, ran just fifty GRPO steps, called preliminary by the authors themselves. Honest, but the abstract's 'effective for RL' framing rides on a thin proof of concept. They also flag that models trained on this data may carry unexamined biases. 20 00:10:55,949 --> 00:11:27,074 [Hal Turing] Fair place to land it. What sticks with me: this is a genuinely clever, cheap way to mine reasoning questions out of text nobody thought to look at that way, and the distillation and self-training results back it up. But nearly every quality and difficulty number here comes from inside the same small family of models, checked by a human study too small to lean on hard. Treat the headline numbers as promising, not settled. That's NaturalReasoning — thanks for listening, everyone, and we'll see you next time.