This episode explores Paul Werbos’s 2004 review of reverse differentiation and argues that reverse-mode automatic differentiation, backpropagation, hand-coded adjoints, and adjoint circuits are largely the same core idea expressed in different technical communities. It explains the mechanics of automatic differentiation and reverse mode in clear terms, then traces how these methods diverged historically and why that fragmentation slowed progress in fields like neural networks, control, and scientific computing. The discussion highlights Werbos’s main claim that better integrated, derivative-aware software could make advanced nonlinear modeling and intelligent control far more practical, while also questioning how much evidence supports that agenda beyond synthesis and historical interpretation. Listeners would find it interesting for its sharp distinction between gradients as infrastructure versus models or optimizers, and for its perspective on how today’s differentiable programming ecosystem was once a contested software vision.
Sources:
1. Reverse-Mode Differentiation Across AD and Neural Nets
https://www.werbos.com/AD2004.pdf2. A Simple Automatic Derivative Evaluation Program — R. E. Wengert, 1964
https://scholar.google.com/scholar?q=A+Simple+Automatic+Derivative+Evaluation+Program3. Taylor Expansion of the Accumulated Rounding Error — Seppo Linnainmaa, 1976
https://scholar.google.com/scholar?q=Taylor+Expansion+of+the+Accumulated+Rounding+Error4. The Complexity of Partial Derivatives — Walter Baur and Volker Strassen, 1983
https://scholar.google.com/scholar?q=The+Complexity+of+Partial+Derivatives5. Automatic Differentiation in Machine Learning: a Survey — Atılım Güneş Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, Jeffrey Mark Siskind, 2018
https://scholar.google.com/scholar?q=Automatic+Differentiation+in+Machine+Learning%3A+a+Survey6. Learning Representations by Back-Propagating Errors — David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams, 1986
https://scholar.google.com/scholar?q=Learning+Representations+by+Back-Propagating+Errors7. Backpropagation Through Time: What It Does and How to Do It — Paul J. Werbos, 1990
https://scholar.google.com/scholar?q=Backpropagation+Through+Time%3A+What+It+Does+and+How+to+Do+It8. Backpropagation Applied to Handwritten Zip Code Recognition — Yann LeCun, Bernhard Boser, John S. Denker, Don Henderson, Richard E. Howard, Wayne Hubbard, Lawrence D. Jackel, 1989
https://scholar.google.com/scholar?q=Backpropagation+Applied+to+Handwritten+Zip+Code+Recognition9. Gradient-Based Learning Applied to Document Recognition — Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, 1998
https://scholar.google.com/scholar?q=Gradient-Based+Learning+Applied+to+Document+Recognition10. Neuro-Dynamic Programming: An Overview — Dimitri P. Bertsekas, John N. Tsitsiklis, 1995
https://scholar.google.com/scholar?q=Neuro-Dynamic+Programming%3A+An+Overview11. Neuro-Dynamic Programming — Dimitri P. Bertsekas, John N. Tsitsiklis, 1996
https://scholar.google.com/scholar?q=Neuro-Dynamic+Programming12. Approximate Dynamic Programming and Reinforcement Learning — Lucian Bușoniu, Bart De Schutter, Robert Babuška, 2010
https://scholar.google.com/scholar?q=Approximate+Dynamic+Programming+and+Reinforcement+Learning13. An Approximate Dynamic Programming Algorithm for Large-Scale Fleet Management: A Case Application — Hugo P. Simão, Jeff Day, Abraham P. George, Ted Gifford, John Nienow, Warren B. Powell, 2009
https://scholar.google.com/scholar?q=An+Approximate+Dynamic+Programming+Algorithm+for+Large-Scale+Fleet+Management%3A+A+Case+Application14. Some New Tools for Prediction and Analysis in the Behavioral Sciences — Paul J. Werbos, 1974
https://scholar.google.com/scholar?q=Some+New+Tools+for+Prediction+and+Analysis+in+the+Behavioral+Sciences15. The Difficulty of Learning Long-Term Dependencies with Gradient Descent is Officially Overcome — Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, 2001
https://scholar.google.com/scholar?q=The+Difficulty+of+Learning+Long-Term+Dependencies+with+Gradient+Descent+is+Officially+Overcome16. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation — Andreas Griewank, 2000
https://scholar.google.com/scholar?q=Evaluating+Derivatives%3A+Principles+and+Techniques+of+Algorithmic+Differentiation17. Differential Dynamic Programming — David H. Jacobson, David Q. Mayne, 1970
https://scholar.google.com/scholar?q=Differential+Dynamic+Programming18. Backpropagation-free training of deep physical neural networks — Ali Momeni, Babak Rahmani, Matthieu Mallejac, Philipp Del Hougne, Romain Fleury, 2023
https://scholar.google.com/scholar?q=Backpropagation-free+training+of+deep+physical+neural+networks19. Fully forward mode training for optical neural networks — Zhiwei Xue, Tiankuang Zhou, Zhihao Xu, Shaoliang Yu, Qionghai Dai, Lu Fang, 2024
https://scholar.google.com/scholar?q=Fully+forward+mode+training+for+optical+neural+networks20. Brain-like training of a pre-sensor optical neural network with a backpropagation-free algorithm — Zheng Huang, Conghe Wang, Caihua Zhang, Wanxin Shi, Shukai Wu, Sigang Yang, Hongwei Chen, 2025
https://scholar.google.com/scholar?q=Brain-like+training+of+a+pre-sensor+optical+neural+network+with+a+backpropagation-free+algorithm21. Backpropagation-Free Deep Learning with Recursive Local Representation Alignment — Alexander G. Ororbia, Ankur Mali, Daniel Kifer, C. Lee Giles, 2023
https://scholar.google.com/scholar?q=Backpropagation-Free+Deep+Learning+with+Recursive+Local+Representation+Alignment22. Exploring the Promise and Limits of Real-Time Recurrent Learning — Kazuki Irie, Anand Gopalakrishnan, Jürgen Schmidhuber, 2023
https://scholar.google.com/scholar?q=Exploring+the+Promise+and+Limits+of+Real-Time+Recurrent+Learning23. Real-Time Recurrent Reinforcement Learning — Julian Lemmel, Radu Grosu, 2023/2025
https://scholar.google.com/scholar?q=Real-Time+Recurrent+Reinforcement+Learning24. Second-order forward-mode optimization of recurrent neural networks for neuroscience — Youjing Yu, Rui Xia, Qingxi Ma, Máté Lengyel, Guillaume Hennequin, 2024
https://scholar.google.com/scholar?q=Second-order+forward-mode+optimization+of+recurrent+neural+networks+for+neuroscience25. Dynamic predictive coding: A model of hierarchical sequence learning and prediction in the neocortex — Linxing Preston Jiang, Rajesh P. N. Rao, 2024
https://scholar.google.com/scholar?q=Dynamic+predictive+coding%3A+A+model+of+hierarchical+sequence+learning+and+prediction+in+the+neocortex26. Predictive coding networks for temporal prediction — Beren Millidge, Mufeng Tang, Mahyar Osanlouy, Nicol S. Harper, Rafal Bogacz, 2024
https://scholar.google.com/scholar?q=Predictive+coding+networks+for+temporal+prediction27. Where is the error? Hierarchical predictive coding through dendritic error computation — Fabian A. Mikulasch, Lucas Rudelt, Michael Wibral, Viola Priesemann, 2023
https://scholar.google.com/scholar?q=Where+is+the+error%3F+Hierarchical+predictive+coding+through+dendritic+error+computation28. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp329. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp330. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3Interactive Visualization: Reverse-Mode Differentiation Across AD and Neural Nets