AI Post Transformers · Episode Companion

Model-Agnostic Meta-Learning for Fast Task Adaptation

Finn, Abbeel, Levine · ICML 2017 UC Berkeley / OpenAI 📄 PMLR v70/finn17a Classification · Regression · RL

One initialization, no extra parameters, no task-specific architecture — trained so a few steps of ordinary gradient descent adapt it fast to a brand-new task. This page visualizes the inner/outer loop structure, the second-order gradient math, and how MAML compares to prior meta-learners and plain pretrain-then-fine-tune.

Learn an initialization, not an update rule

MAML replaces a learned optimizer with plain gradient descent, and instead optimizes the starting weights so that descent works fast from there. Explore the N-way K-shot setup and where MAML sits relative to prior meta-learning families.

5-way 1-shot: one image per class

Omniglot support set — hover a cell for its class

support example (K=1) query example (evaluated)

Meta-learner families, 2016-2017

What each approach actually learns

Same recipe, three problem types

MAML's outer-loop objective is identical across domains — only the loss function changes

Ω
Omniglot / MiniImageNet — classification loss
∫
Sinusoid regression — MSE loss
R
MuJoCo / 2D nav — policy gradient (TRPO)

Inner loop adapts, outer loop generalizes

The inner loop takes a few gradient steps on K examples from one task. The outer loop scores those adapted weights on fresh query examples from the same task, and pushes the shared initialization toward "easy to adapt."

MAML training pipeline

Solid arrows = data flow · dashed purple = gradient flow (backprop through adaptation)

Step by step

  1. 1Sample a batch of tasks from the task distribution p(T)
  2. 2For each task, sample K support examples
  3. 3Inner loop: θ′ = θ − α·∇θ L_task(θ)
  4. 4Evaluate θ′ on fresh query examples from the same task
  5. 5Outer loop: θ ← θ − β·∇θ Σ_tasks L_task(θ′)
  6. 6Repeat across the task batch — θ becomes "primed" to adapt

Why this differs from pretrain → fine-tune

Toggle to compare objectives

Gradient through a gradient

Because θ′ is itself a function of θ, the outer-loop update requires differentiating through the inner step — a Hessian-vector product. In the RL setting, TRPO's own second-order machinery would need a third derivative, so MAML approximates it with finite differences instead.

Backprop-through-adaptation graph

Hover a node for its role

First-order vs. full second-order MAML

MiniImageNet, 5-way 1-shot — toggle the variant

Dropping the Hessian term costs ~0.6 points of accuracy here but saves ~33% compute — a comparison the paper runs only on this one classification benchmark, not on regression or RL.

Where exact 2nd-order breaks down

TRPO's meta-update would need a 3rd derivative

Results across three domains

Sinusoid extrapolation from five points, classification accuracy against prior meta-learners, and RL adaptation speed against pretraining and random initialization.

Sinusoid regression: 5 points, one gradient step

Toggle to compare MAML's one-step adaptation against a pretrained-then-fine-tuned baseline

ground truth adapted prediction 5 training points

5-way 1-shot accuracy by method

Hover a bar for the exact figure

RL: reward vs. gradient steps at test time

2D navigation, new goal — MAML vs. pretrained vs. random init

What the paper doesn't test

The episode flags real gaps between the "model-agnostic, any problem" framing and what was actually run.

Claim vs. evidence

Narrow task families

Every "new task" at test time is an in-distribution perturbation: same sine frequency range, same alphabet-agnostic character pool, same MuJoCo robot body. No out-of-distribution morphology or frequency range is ever tested.

Small networks only

A 40-unit MLP for regression, four conv layers for classification, a 100-unit policy for RL. No recurrent architecture and no pixel-based RL are ever run, despite the title's "model-agnostic" framing.