A trajectory optimizer (DDP) plays teacher to a flexible neural-network student, sidestepping the poor local optima that sink direct policy search — using importance-sampled policy gradients instead of naive imitation, because the teacher's guidance is only locally valid.