This episode explores Leo Breiman’s “Statistical Modeling: The Two Cultures” as a sharp argument about a core divide in data science: building explicit probabilistic models to explain the world versus training algorithms that simply predict well on new data. It traces how that split maps onto inference versus prediction, connects Breiman’s critique to earlier ideas from Tukey and Box, and shows how later work such as Shmueli’s formalized the distinction. The discussion also grounds the debate in concrete methods, from linear regression and Cox models to CART, bagging, random forests, and neural networks, highlighting why algorithmic approaches gained ground on messy, high-dimensional problems. Listeners would find it interesting because it explains a foundational argument that still shapes modern machine learning, while also probing where interpretable classical models remain essential in areas like medicine, policy, and reliability.
Sources:
1. Breiman's Two Cultures of Statistical Modeling
https://www2.math.uu.se/~thulin/mm/breiman.pdf2. The Future of Data Analysis — John W. Tukey, 1962
https://scholar.google.com/scholar?q=The+Future+of+Data+Analysis3. Science and Statistics — George E. P. Box, 1976
https://scholar.google.com/scholar?q=Science+and+Statistics4. Regression Models and Life-Tables — D. R. Cox, 1972
https://scholar.google.com/scholar?q=Regression+Models+and+Life-Tables5. Statistical Modeling: The Two Cultures — Leo Breiman, 2001
https://scholar.google.com/scholar?q=Statistical+Modeling%3A+The+Two+Cultures6. Bagging Predictors — Leo Breiman, 1996
https://scholar.google.com/scholar?q=Bagging+Predictors7. Random Forests — Leo Breiman, 2001
https://scholar.google.com/scholar?q=Random+Forests8. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001
https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine9. Clinical versus Actuarial Judgment — Robyn M. Dawes, David Faust, Paul E. Meehl, 1989
https://scholar.google.com/scholar?q=Clinical+versus+Actuarial+Judgment10. To Explain or To Predict? — Galit Shmueli, 2010
https://scholar.google.com/scholar?q=To+Explain+or+To+Predict%3F11. Choosing Prediction Over Explanation in Psychology: Lessons From Machine Learning — Tal Yarkoni, Jacob Westfall, 2017
https://scholar.google.com/scholar?q=Choosing+Prediction+Over+Explanation+in+Psychology%3A+Lessons+From+Machine+Learning12. Classification and Regression Trees — Leo Breiman, Jerome H. Friedman, Richard A. Olshen, Charles J. Stone, 1984
https://scholar.google.com/scholar?q=Classification+and+Regression+Trees13. Induction of Decision Trees — J. R. Quinlan, 1986
https://scholar.google.com/scholar?q=Induction+of+Decision+Trees14. A Random Forest Guided Tour — Gérard Biau, Erwan Scornet, 2016
https://scholar.google.com/scholar?q=A+Random+Forest+Guided+Tour15. Learning Representations by Back-Propagating Errors — David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams, 1986
https://scholar.google.com/scholar?q=Learning+Representations+by+Back-Propagating+Errors16. Multilayer Feedforward Networks Are Universal Approximators — Kurt Hornik, Maxwell Stinchcombe, Halbert White, 1989
https://scholar.google.com/scholar?q=Multilayer+Feedforward+Networks+Are+Universal+Approximators17. ImageNet Classification with Deep Convolutional Neural Networks — Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton, 2012
https://scholar.google.com/scholar?q=ImageNet+Classification+with+Deep+Convolutional+Neural+Networks18. Deep Learning — Yann LeCun, Yoshua Bengio, Geoffrey Hinton, 2015
https://scholar.google.com/scholar?q=Deep+Learning19. Arcing Classifiers — Leo Breiman, 1998
https://scholar.google.com/scholar?q=Arcing+Classifiers20. Generalized Additive Models — Trevor Hastie and Robert Tibshirani, 1990
https://scholar.google.com/scholar?q=Generalized+Additive+Models21. The Elements of Statistical Learning — Trevor Hastie, Robert Tibshirani, and Jerome Friedman, 2001
https://scholar.google.com/scholar?q=The+Elements+of+Statistical+Learning22. Sparse Neural Additive Model: Interpretable Deep Learning with Feature Selection via Group Sparsity — Shiyun Xu, Zhiqi Bu, Pratik Chaudhari, Ian J. Barnett, 2022/2023
https://scholar.google.com/scholar?q=Sparse+Neural+Additive+Model%3A+Interpretable+Deep+Learning+with+Feature+Selection+via+Group+Sparsity23. Neural Additive Models for Location Scale and Shape: A Framework for Interpretable Neural Regression Beyond the Mean — Anton Frederik Thielmann, Rene-Marcel Kruse, Thomas Kneib, Benjamin Safken, 2024
https://scholar.google.com/scholar?q=Neural+Additive+Models+for+Location+Scale+and+Shape%3A+A+Framework+for+Interpretable+Neural+Regression+Beyond+the+Mean24. Conformal Prediction: A Data Perspective — Xiaofan Zhou, Baiting Chen, Yu Gui, Lu Cheng, 2024
https://scholar.google.com/scholar?q=Conformal+Prediction%3A+A+Data+Perspective25. Large language model validity via enhanced conformal prediction methods — John J. Cherian, Isaac Gibbs, Emmanuel J. Candes, 2024
https://scholar.google.com/scholar?q=Large+language+model+validity+via+enhanced+conformal+prediction+methods26. CPSign: conformal prediction for cheminformatics modeling — Staffan Arvidsson McShane, Ulf Norinder, Jonathan Alvarsson, Ernst Ahlberg, Lars Carlsson, Ola Spjuth, 2024
https://scholar.google.com/scholar?q=CPSign%3A+conformal+prediction+for+cheminformatics+modeling27. Open Problems in Mechanistic Interpretability — Lee Sharkey et al., 2025
https://scholar.google.com/scholar?q=Open+Problems+in+Mechanistic+Interpretability28. AI Post Transformers: Evaluating LLM Embeddings for Psychometric Personality Prediction — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/evaluating-llm-embeddings-for-psychometric-personality-prediction/29. AI Post Transformers: Introducing RTEB: Retrieval Embedding Benchmark — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/introducing-rteb-retrieval-embedding-benchmark/30. AI Post Transformers: Information Bottleneck-based Causal Attention for Medical Image Recognition — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/information-bottleneck-based-causal-attention-for-medical-image-recognition/Interactive Visualization: Breiman's Two Cultures of Statistical Modeling