An Algorithmic Perspective on Imitation Learning

Osa, Pajarinen, Neumann, Bagnell, Abbeel, Peters — Foundations and Trends in Robotics, 2018
AI Post Transformers · Hosts: Hal Turing & Dr. Ada Shannon
arXiv:1811.06711 ↗

Two Families, One Fork in the Road

Skip manual programming and reward engineering — learn the policy from demonstrations instead. But "learn how" splits immediately into two philosophies.
Behavioral Cloning path Inverse RL path Rejected alternatives

BC treats imitation as supervised learning: state → action, directly. IRL recovers the expert's implicit reward first, then solves an RL problem to get a policy — better generalization, at the cost of a nested solver.

Compounding Error vs. DAgger's Fix

In supervised learning, test data doesn't depend on the model. In BC, the policy's own actions decide what it sees next.
Expert demonstrated states Policy rollout (drifted) Aggregated correction

One small prediction error pushes the policy into a state the expert never demonstrated — and error compounds from there.

Seventeen Demonstrations and a Helicopter

The introduction claims manufacturing, elder care, the service industry. The cited experiments are lab-scale.
Demonstrations used Implied industrial-scale bar

Dashed reference line = order-of-magnitude scale implied by "manufacturing, elder care, service industry." None of the six results cited come close.

Dynamic Movement Primitives: Structure vs. Learned-Everything

A spring-damper attractor guaranteed to converge, shaped by a learned forcing term along a decaying phase variable.
Phase variable x(t) Resulting trajectory
Basis function activation weight, blue → orange → red

Early in the motion the forcing term (weighted Gaussian basis functions) dominates and shapes the path; late in the motion the attractor takes over and pulls hard toward the goal — convergence is guaranteed by construction, not learned.

The Field Didn't Resolve BC vs. IRL — It Split It

BC absorbed manipulation at scale. IRL's descendants went and ate alignment instead.

Maximum-entropy IRL's KL-regularized reward recovery is the direct mathematical ancestor of RLHF and DPO. Meanwhile DMPs' stability guarantees resurface as safety layers on diffusion policies — the "soon to be superseded" structure gets more valuable at scale, not less.

References