AI Post Transformers · Episode Companion

AI Coaching That Preserves Human Skill Development

A non-cooperative dynamic game reframes AI coaching: the learner optimizes immediate performance while the coach is trained on Value of Independence — a counterfactual measure of skill retained if the AI vanished. Tested on FPV drone racing against fixed-schedule and safety-filter baselines.
⚡ arXiv:2606.25337 University of Pennsylvania × Johns Hopkins Wang, Gu, Loquercio, Hu, Mangharam · Jun 2026 View Paper ↗

The Assistance Dial: Two Failure Modes

Drag conceptually along the spectrum — over-assistance produces clean performance but no learning; under-assistance produces uninstructive failure. L2C hunts for the productive middle.

Over-assistance — AI flies the drone, nothing sticks Productive failure zone — Kapur 2008 Under-assistance — crash with no signal

Coach ↔ Learner ↔ Environment Loop

The learner's skill-conditioned policy is folded into the environment's dynamics, collapsing a two-player game into a single-agent POMDP that PPO can solve.

Blending Equation

a = λ·aE + (1−λ)·aH

λ is a per-axis vector, not a scalar — roll and yaw get independent assistance levels.

Coach Objective

Value of Independence (VoI)

Counterfactual: how well would the human fly if the AI were removed right now?

Training Surrogate

θt − θt−1

Proposition 1: upskill-dominance implies VoI-dominance — cheap surrogate stands in for the intractable counterfactual.

Cooperative vs. Non-Cooperative Framing

Learner reward (task performance) and Coach reward (Value of Independence) can move in opposite directions on the same timestep — more help now can suppress the counterfactual solo-performance signal the coach is scored on.

Reward Divergence Over a Session

Lineage: Dynamic Game Theory

Figure-Eight Track · Per-Gate Skill Belief

Twelve gates, two overlapping at the center crossing. Skill θ is never directly observed — the coach maintains a separate Bayesian belief per gate, updated from clear-time relative to a skill-calibrated target. Hover a gate.

Low skill belief Mid High skill belief

Per-Axis Assistance (λ)

PFA: Skill State Transitions

Success → upskill (αS) or downskill (βS). Failure → upskill (αF). Each is a sigmoid in θ.

L2C vs. Baselines (N = 33, 11 per arm)

Head-to-Head Contrasts (Welch's t, Holm-corrected)

Within-subject, L2C is the only arm clearing significance on both outcomes. Between-group contrasts point the same direction with medium–large effect sizes, but at n=11/arm land at p = 0.09–0.16 — directionally consistent, not conventionally significant.

References