This episode explores a paper that argues AI can help mathematics most by orchestrating the full research workflow rather than acting as a one-shot chatbot. It discusses why real mathematical work depends on durable memory, branching hypotheses, literature search, proof attempts, computation, and recorded failures, and contrasts that with both ordinary chat interfaces and formal theorem provers such as Lean or Coq. The conversation details the paper’s multi-agent design, where a coordinator delegates parallel tasks like literature review, coding, proof exploration, and claim checking into a living draft document with provenance and uncertainty markers. It also highlights reported results on 100 research-level problems, where the full system outperformed strong single-shot models by using tactics such as SAT reduction, theorem retrieval, and coordinated theory-computation pipelines, making the episode interesting for listeners curious about how AI might become a practical research collaborator instead of just a clever text generator.
Sources:
1. AI Co-Mathematician for Mathematical Research
https://arxiv.org/pdf/2605.066512. A Survey on Deep Learning for Theorem Proving — Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, Xujie Si, 2024
https://arxiv.org/abs/2404.099393. GPT-f: Generative Language Modeling for Automated Theorem Proving — Stanislas Polu, Ilya Sutskever, 2020
https://openai.com/index/generative-language-modeling-for-automated-theorem-proving//4. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs — Albert Q. Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothee Lacroix, Yuhuai Wu, Guillaume Lample, 2023
https://arxiv.org/abs/2210.122835. Solving olympiad geometry without human demonstrations — Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, Thang Luong, 2024
https://www.nature.com/articles/s41586-023-06747-56. Exploration and Explanation in Computational Notebooks — Adam Rule, Aurelien Tabard, James D. Hollan, 2018
https://adamrule.com/files/papers/chi_2018_computational_notebooks_camera_ready.pdf7. What’s Wrong with Computational Notebooks? Pain Points, Needs, and Design Opportunities — Souti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma, Titus Barik, 2020
https://www.microsoft.com/en-us/research/publication/whats-wrong-with-computational-notebooks/8. Principles for data analysis workflows — Sara Stoudt, Valeri N. Vasquez, Ciera C. Martinez, 2021
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.10087709. Towards an AI co-scientist — Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic and many others, 2025
https://arxiv.org/abs/2502.1886410. Towards Autonomous Mathematics Research — T. Feng, T. H. Trinh, G. Bingham, D. Hwang, Y. Chervonyi, J. Jung, J. Lee, C. Pagano, S.-h. Kim, F. Pasqualotto, S. Gukov, J. N. Lee, J. Kim, K. Hou, G. Ghiasi, Y. Tay, Y. Li, C. Kuang, Y. Liu, H. Lin, E. Z. Liu, N. Nayakanti, X. Yang, H.-t. Cheng, D. Hassabis, K. Kavukcuoglu, Q. V. Le, and T. Luong, 2026
https://scholar.google.com/scholar?q=Towards+Autonomous+Mathematics+Research11. Olympiad-level formal mathematical reasoning with reinforcement learning — T. Hubert, R. S. Mehta, L. Sartran, M. Z. Horvath, G. Zuzic, E. Wieser, A. Huang, J. Schrittwieser, Y. Schroecker, H. Masoom, O. Bertolli, T. Zahavy, A. Mandhane, J. Yung, I. Beloshapka, B. Ibarz, V. Veeriah, L. Yu, O. Nash, P. Lezeau, S. Mercuri, C. Sonne, B. Mehta, A. Davies, D. Zheng, F. Pedregosa, Y. Li, I. von Glehn, M. Rowland, S. Albanie, A. Velingker, S. Schmitt, E. Lockhart, E. Hughes, H. Michalewski, N. Sonnerat, D. Hassabis, P. Kohli, and D. Silver, 2025
https://scholar.google.com/scholar?q=Olympiad-level+formal+mathematical+reasoning+with+reinforcement+learning12. FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI — E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Jarviniemi, M. Barnett, R. Sandler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V. Enugandla, and M. Wildon, 2024
https://scholar.google.com/scholar?q=FrontierMath%3A+A+Benchmark+for+Evaluating+Advanced+Mathematical+Reasoning+in+AI13. Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math — S. Pandit, A. Xu, X.-P. Nguyen, Y. Ming, C. Xiong, and S. Joty, 2025
https://scholar.google.com/scholar?q=Hard2Verify%3A+A+Step-Level+Verification+Benchmark+for+Open-Ended+Frontier+Math14. Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning — Wang Yang et al., 2025
https://scholar.google.com/scholar?q=Longer+Context%2C+Deeper+Thinking%3A+Uncovering+the+Role+of+Long-Context+Ability+in+Reasoning15. InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models — Yuchen Yan et al., 2025
https://scholar.google.com/scholar?q=InftyThink%3A+Breaking+the+Length+Limits+of+Long-Context+Reasoning+in+Large+Language+Models16. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models — Andy Zhou et al., 2023
https://scholar.google.com/scholar?q=Language+Agent+Tree+Search+Unifies+Reasoning+Acting+and+Planning+in+Language+Models17. Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training — Xidong Feng et al., 2023
https://scholar.google.com/scholar?q=Alphazero-like+Tree-Search+can+Guide+Large+Language+Model+Decoding+and+Training18. ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving — Zhibin Gou et al., 2023
https://scholar.google.com/scholar?q=ToRA%3A+A+Tool-Integrated+Reasoning+Agent+for+Mathematical+Problem+Solving19. Efficient Tool Use with Chain-of-Abstraction Reasoning — Silin Gao et al., 2024
https://scholar.google.com/scholar?q=Efficient+Tool+Use+with+Chain-of-Abstraction+Reasoning20. MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning — Shuo Yin et al., 2024
https://scholar.google.com/scholar?q=MuMath-Code%3A+Combining+Tool-Use+Large+Language+Models+with+Multi-perspective+Data+Augmentation+for+Mathematical+Reasoning21. HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement — Jilin Hu et al., 2025
https://scholar.google.com/scholar?q=HybridProver%3A+Augmenting+Theorem+Proving+with+LLM-Driven+Proof+Synthesis+and+Refinement22. A Minimal Agent for Automated Theorem Proving — Borja Requena et al., 2026
https://scholar.google.com/scholar?q=A+Minimal+Agent+for+Automated+Theorem+Proving23. Tree Search for Language Model Agents — Jing Yu Koh et al., 2024
https://scholar.google.com/scholar?q=Tree+Search+for+Language+Model+Agents24. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp325. AI Post Transformers: Agentic AI and the Next Intelligence Explosion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-28-agentic-ai-and-the-next-intelligence-exp-d06561.mp326. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3Interactive Visualization: AI Co-Mathematician for Mathematical Research