AI Post Transformers · Episode Companion

NaturalReasoning: Backtranslating Reasoning Questions at Scale

NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions — Weizhe Yuan, Jane Yu, Song Jiang, Karthik Padthe, Yang Li, Ilia Kulikov, Kyunghyun Cho, Dong Wang, Yuandong Tian, Jason E Weston, Xian Li (Meta & NYU)
arXiv:2502.13124 NeurIPS 2025 Meta · NYU 11 authors v2 · Nov 6, 2025
2.8M
backtranslated questions
4
LLM passes / document
6.45
human quality score (vs 5.92)
1/8
data to match distill baseline

From raw text to a training question

Llama-3.3-70B-Instruct sweeps DCLM-baseline and FineMath, flags reasoning-rich passages, then works backward from the passage to the question it would answer — four LLM passes per surviving document.

Per-document rating rubric

Every candidate document is scored on four dimensions before it's allowed into the backtranslation step. Illustrative pass-rate proxy — not exact reported figures.

Backtranslation's lineage: 2016 → 2025

The trick starts in machine translation — train a reverse model to manufacture synthetic inputs for real outputs you already trust — and gets re-pointed at reasoning text nine years later. Hover a node for detail.

Who checks whose work?

Document mining, question synthesis, reference-answer verification and the difficulty proxy all stay inside the Llama family. Only quality scoring pulls in outside judges — and the human study meant to corroborate everything is thin. Hover any cell.
not involved partial / shared authority sole authority
The trap: the same Llama family that mines, writes, verifies and teaches also casts 1 of 3 votes in the one stage meant to diversify judgment. The only fully independent check is a 100-question human study with no reported inter-rater agreement.

Does it hold up against the closest competitor?

WebInstruct trades precision for scale — 13M recalled QA pairs vs 2.8M synthesized ones. Toggle views below.

Two ways to train on this data

Straightforward distillation aside, the paper also tests a reward-model-free loop. Compare it against the RLVR approach it's meant to widen beyond math and code.

References