This episode examines "On First-Order Meta-Learning Algorithms" by Alex Nichol, Joshua Achiam, and John Schulman, which challenges the assumption that MAML's expensive second-derivative computation is essential for effective few-shot learning. The discussion traces the lineage from MAML's nested optimization — where an outer loop backpropagates through an inner loop's gradient steps via the Hessian — through First-Order MAML's approximation, to Reptile, a stripped-down algorithm that simply runs SGD on sampled tasks and nudges the initialization toward the result, with no meta-gradient or train-test split required. A central tension drives the conversation: why pulling an initialization toward "wherever SGD landed" produces a genuinely different target than plain joint training across tasks, rather than just averaging into one generic model. The hosts set up a Taylor-expansion argument to explain which gradient terms MAML, FOMAML, and Reptile weight differently, revealing the mathematical reason the cheaper approximation retains nearly all the useful signal. Listeners interested in the mechanics of meta-learning, gradient-based optimization tradeoffs, or the history of few-shot learning approaches will find the paper's practical implications for scaling meta-learning algorithms especially relevant.
Sources:
1. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018
http://arxiv.org/abs/1803.029992. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks3. How to train your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019
https://scholar.google.com/scholar?q=How+to+train+your+MAML4. Meta-Learning with Implicit Gradients — Aravind Rajeswaran, Chelsea Finn, Sham Kakade, Sergey Levine, 2019
https://scholar.google.com/scholar?q=Meta-Learning+with+Implicit+Gradients5. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017
https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning6. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016
https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent7. Using Fast Weights to Deblur Old Memories — Geoffrey E. Hinton, David C. Plaut, 1987
https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Deblur+Old+Memories8. Parallelized Stochastic Gradient Descent — Martin Zinkevich, Markus Weimer, Lihong Li, Alex J. Smola, 2010
https://scholar.google.com/scholar?q=Parallelized+Stochastic+Gradient+Descent