One initialization, no extra parameters, no task-specific architecture — trained so a few steps of ordinary gradient descent adapt it fast to a brand-new task. This page visualizes the inner/outer loop structure, the second-order gradient math, and how MAML compares to prior meta-learners and plain pretrain-then-fine-tune.
MAML replaces a learned optimizer with plain gradient descent, and instead optimizes the starting weights so that descent works fast from there. Explore the N-way K-shot setup and where MAML sits relative to prior meta-learning families.
Omniglot support set — hover a cell for its class
What each approach actually learns
MAML's outer-loop objective is identical across domains — only the loss function changes
The inner loop takes a few gradient steps on K examples from one task. The outer loop scores those adapted weights on fresh query examples from the same task, and pushes the shared initialization toward "easy to adapt."
Solid arrows = data flow · dashed purple = gradient flow (backprop through adaptation)
Toggle to compare objectives
Because θ′ is itself a function of θ, the outer-loop update requires differentiating through the inner step — a Hessian-vector product. In the RL setting, TRPO's own second-order machinery would need a third derivative, so MAML approximates it with finite differences instead.
Hover a node for its role
MiniImageNet, 5-way 1-shot — toggle the variant
Dropping the Hessian term costs ~0.6 points of accuracy here but saves ~33% compute — a comparison the paper runs only on this one classification benchmark, not on regression or RL.
TRPO's meta-update would need a 3rd derivative
Sinusoid extrapolation from five points, classification accuracy against prior meta-learners, and RL adaptation speed against pretraining and random initialization.
Toggle to compare MAML's one-step adaptation against a pretrained-then-fine-tuned baseline
Hover a bar for the exact figure
2D navigation, new goal — MAML vs. pretrained vs. random init
The episode flags real gaps between the "model-agnostic, any problem" framing and what was actually run.
Every "new task" at test time is an in-distribution perturbation: same sine frequency range, same alphabet-agnostic character pool, same MuJoCo robot body. No out-of-distribution morphology or frequency range is ever tested.
A 40-unit MLP for regression, four conv layers for classification, a 100-unit policy for RL. No recurrent architecture and no pixel-based RL are ever run, despite the title's "model-agnostic" framing.