Drag conceptually along the spectrum — over-assistance produces clean performance but no learning; under-assistance produces uninstructive failure. L2C hunts for the productive middle.
The learner's skill-conditioned policy is folded into the environment's dynamics, collapsing a two-player game into a single-agent POMDP that PPO can solve.
a = λ·aE + (1−λ)·aH
λ is a per-axis vector, not a scalar — roll and yaw get independent assistance levels.
Value of Independence (VoI)
Counterfactual: how well would the human fly if the AI were removed right now?
θt − θt−1
Proposition 1: upskill-dominance implies VoI-dominance — cheap surrogate stands in for the intractable counterfactual.
Learner reward (task performance) and Coach reward (Value of Independence) can move in opposite directions on the same timestep — more help now can suppress the counterfactual solo-performance signal the coach is scored on.
Twelve gates, two overlapping at the center crossing. Skill θ is never directly observed — the coach maintains a separate Bayesian belief per gate, updated from clear-time relative to a skill-calibrated target. Hover a gate.
Success → upskill (αS) or downskill (βS). Failure → upskill (αF). Each is a sigmoid in θ.
Within-subject, L2C is the only arm clearing significance on both outcomes. Between-group contrasts point the same direction with medium–large effect sizes, but at n=11/arm land at p = 0.09–0.16 — directionally consistent, not conventionally significant.