BC treats imitation as supervised learning: state → action, directly. IRL recovers the expert's implicit reward first, then solves an RL problem to get a policy — better generalization, at the cost of a nested solver.
One small prediction error pushes the policy into a state the expert never demonstrated — and error compounds from there.
Dashed reference line = order-of-magnitude scale implied by "manufacturing, elder care, service industry." None of the six results cited come close.
Early in the motion the forcing term (weighted Gaussian basis functions) dominates and shapes the path; late in the motion the attractor takes over and pulls hard toward the goal — convergence is guaranteed by construction, not learned.
Maximum-entropy IRL's KL-regularized reward recovery is the direct mathematical ancestor of RLHF and DPO. Meanwhile DMPs' stability guarantees resurface as safety layers on diffusion policies — the "soon to be superseded" structure gets more valuable at scale, not less.