1 00:00:01,000 --> 00:00:44,189 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "AI Coaching for Accelerating Human Skill Development with Reinforcement Learning," from Wei Wang, Enlin Gu, Antonio Loquercio, Haimin Hu, and Rahul Mangharam — that's Wei Wang et al., five authors total — out of the University of Pennsylvania and Johns Hopkins University, posted to arXiv on June 24th, 2026. And Ada, the setup here genuinely stopped me: they open by saying AI copilots that help you succeed right now might be quietly wrecking your ability to ever succeed without them. 2 00:00:44,189 --> 00:01:21,433 [Dr. Ada Shannon] Right, and that's the sentence that makes this paper worth an episode. Every shared-control system we've talked about on this show has been optimized for one thing: keep the human from crashing, keep the lap time down, keep the score up, right now, in this episode of trying. Nobody was asking what happens to the human's own skill after a hundred sessions of the AI quietly bailing them out. This paper is the first one I've seen that turns 'does the human get better on their own' into the actual objective function the AI is trained against, instead of a side effect nobody measured. 3 00:01:21,433 --> 00:01:33,740 [Hal Turing] So walk me through the dilemma, because on the surface it sounds almost contradictory. If the AI helps you and you do better, isn't that just... good coaching? Why would assistance make you worse? 4 00:01:33,740 --> 00:02:13,678 [Dr. Ada Shannon] Because there are two very different failure modes and they sit on opposite ends of the same dial. Crank assistance up and the AI is basically flying the drone for you — you get a clean lap, but you never had to solve the hard part yourself, so nothing sticks. Crank it down to zero and you just crash into the same gate over and over with no idea why, which is failure with no signal in it. The psychology literature calls the productive version of failure exactly that — Manu Kapur's 2008 work on productive failure showed students who struggled with a problem before getting instruction retained it far better than students who were walked through it. The trick is finding failures worth having. 5 00:02:13,678 --> 00:02:20,691 [Hal Turing] Okay, so this isn't a brand new idea — shared control itself has been around a while, right? Cars, wheelchairs, that kind of thing? 6 00:02:20,691 --> 00:03:00,350 [Dr. Ada Shannon] Decades old. The clean mathematical version comes from Dragan and Srinivasa at Berkeley and CMU in 2013 — they formalized shared control as a blend, your action times some weight plus the robot's action times one minus that weight, with the weight increasing as the robot gets more confident about your intent. Reddy, Dragan, and Levine took that further in 2018, training the blending end-to-end with deep RL instead of hand-coding the intent inference. Both of those, and basically every shared-control system since, are tuned to maximize task performance right now. This paper takes that exact blending mechanism and points it at a completely different target. 7 00:03:00,350 --> 00:03:09,035 [Hal Turing] Which is where the game theory comes in — I saw them call this a 'non-cooperative dynamic game,' which sounded almost adversarial to me at first. 8 00:03:09,035 --> 00:03:49,577 [Dr. Ada Shannon] It's not adversarial, but it is genuinely not cooperative, and that distinction matters more than it sounds. In a cooperative Assistance Game — this is the Laidlaw et al. AssistanceZero line of work from 2025 — the human and the robot share one reward, so it collapses into a solvable single-agent problem. Here the learner wants task performance now, full stop, while the coach is scored on something called the Value of Independence — basically, 'if I vanished right now, how well would this person fly without me?' Two different reward functions, two players, that's a non-cooperative dynamic game by definition, going back to the Basar and Olsder formalism from the 1980s. 9 00:03:49,577 --> 00:03:56,682 [Hal Turing] Wait, hold on — VoI is a counterfactual though, right? The coach is being graded on a version of the world where it isn't even— 10 00:03:56,682 --> 00:04:20,784 [Dr. Ada Shannon] —where it isn't there. Exactly, and that's the hard part they spend the rest of the paper solving, we'll get to it. But sit with the definition for a second: it's not 'how good was the lap,' it's 'how good would the lap have been with the assistance switched off.' That's a genuinely different training signal than anything in the shared-control literature, and I think it's the paper's real contribution, more than the drone racing demo. 11 00:04:20,784 --> 00:04:36,574 [Hal Turing] I'll push back a little here, Ada — calling the learner and coach 'non-cooperative' still feels like a stretch to me. They both want the human to eventually fly well. Isn't that just... a shared goal with different time horizons, not actually opposed incentives? 12 00:04:36,574 --> 00:05:13,772 [Dr. Ada Shannon] I actually disagree with you there, Hal. Different time horizons on the same objective would still make it cooperative — you could just discount the reward differently and stay in one shared function. What's happening here is the two rewards can point in opposite directions on the same timestep: the learner's reward goes up when the coach helps, and the coach's own reward, VoI, can go down at that exact moment because more help now means the counterfactual solo performance stays weak. That's not a horizon mismatch, that's a structural conflict, and it's precisely why they needed genuine game-theoretic machinery instead of just reshaping a single reward. 13 00:05:13,772 --> 00:05:26,218 [Hal Turing] Okay — I'll take that, the conflict is real even if the end goal looks aligned from a distance. So how does this compare to the closest prior coaching work, because I know they didn't invent 'fading assistance' out of nowhere? 14 00:05:26,218 --> 00:06:07,457 [Dr. Ada Shannon] No, and they're upfront about it. Shen et al.'s Cyber Racing Coach, from 2025, already does the productive-failure idea in racing by decaying assistance over time — but on a fixed schedule that doesn't know or care what the learner's actual skill is, or what the track is throwing at them right now. Z-COACH, Srivastava et al., also 2025, uses shared autonomy to figure out which sub-skills sit in the learner's zone of proximal development, but it locks in a fixed assistance level once it's picked the curriculum. This paper's whole pitch is closing both loops at once — skill and physical context — rather than committing to either a fixed decay curve or a fixed curriculum up front. 15 00:06:07,457 --> 00:06:16,466 [Hal Turing] So the promise is a coach that's actually watching you, moment to moment, deciding whether this particular gate, right now, is worth letting you crash into. 16 00:06:16,466 --> 00:06:31,420 [Dr. Ada Shannon] That's the pitch. Whether the machinery underneath actually delivers that in practice is exactly what we should dig into next — because reducing a two-player non-cooperative game with a counterfactual reward down to something you can train with off-the-shelf RL is not a small ask. 17 00:06:31,420 --> 00:06:44,795 [Hal Turing] Okay, so let's get into the actual machinery, because I want to understand how you turn 'non-cooperative game with a counterfactual reward' into something you can bolt PPO onto. Where does the blending actually happen? 18 00:06:44,795 --> 00:07:32,163 [Dr. Ada Shannon] So the executed action is a straight linear combination: a equals lambda times the expert action plus one-minus-lambda times the learner's action, and lambda is itself a vector, not a scalar — one entry per control axis. That per-axis part matters. The coach can crank up assistance on yaw while leaving roll almost entirely to the learner, in the same instant, depending on what it's observing. And crucially, lambda is a policy — pi-C-lambda — that reacts to the coach's observations in real time. It's anchored to a frozen, pretrained expert policy, so the coach never has to learn how to fly the drone from scratch. It only has to learn when to lean on an already-competent pilot. That's a nice separation of concerns — task competence and coaching strategy are decoupled. 19 00:07:32,163 --> 00:07:41,358 [Hal Turing] Right, and that's the trick that makes the POMDP reduction work, isn't it? Because on paper you've still got two agents with different rewards fighting for the same action. 20 00:07:41,358 --> 00:08:39,455 [Dr. Ada Shannon] Exactly, and here's the sleight of hand. The learner isn't treated as a free agent from the coach's perspective — its behavior is fully determined by its skill level theta through a fixed policy class. So instead of solving a two-player game, you fold the learner's skill-conditioned policy straight into the environment dynamics. The learner becomes part of the world the coach is acting in, not a co-player it has to negotiate with. That collapses the whole thing to a single-agent POMDP, which is something PPO can actually chew on. And to be clear about what's generating that learner behavior in training — it's a Boltzmann noisily-rational model. Actions get sampled proportional to the exponential of theta times the expert's Q-function, so a low-theta novice samples almost uniformly across the action set, and a high-theta expert clusters tightly around the optimal action. I want to flag plainly: this is a synthetic learner, not real human trajectories. No human ever generates the training data here. 21 00:08:39,455 --> 00:08:48,000 [Hal Turing] Wait, hold on — so the entire coach is trained against a simulated stand-in for a person, and then just... deployed on actual humans at test time? 22 00:08:48,000 --> 00:08:56,637 [Dr. Ada Shannon] That's precisely the structure, yes. And it's worth sitting with for a second before we move on, because that's a real design choice, not an incidental detail. 23 00:08:56,637 --> 00:09:03,557 [Hal Turing] Okay, so given the learner is fake, how does the coach get any signal that its actions are actually changing this simulated person's skill? 24 00:09:03,557 --> 00:10:12,845 [Dr. Ada Shannon] That's the second piece — the probabilistic finite-state automaton, the PFA. It watches a performance index, basically cumulative task reward over a horizon, and classifies each attempt as success or failure against a threshold. On success, the learner upskills with some probability alpha-S, downskills with beta-S, or stays put. On failure, they upskill with probability alpha-F. Each of those is a sigmoid in theta with a sensible shape — novices gain the most from small wins, skill gets stickier once you're more advanced, and skilled learners extract more from their mistakes than beginners do. Now, the actual coach reward they train with isn't the counterfactual Value of Independence directly — that's intractable on-policy. They use a surrogate: theta-t minus theta-t-minus-one, literally the change in skill level. Proposition 1 is the piece that justifies swapping one for the other — it proves that if a coaching policy induces a higher upskill probability at every skill level, it also dominates on the actual VoI objective. So optimizing the cheap surrogate is provably sufficient for the expensive counterfactual one. 25 00:10:12,845 --> 00:10:18,511 [Hal Turing] So let's ground this in the actual deployment, because FPV drone racing is a pretty specific choice of testbed. 26 00:10:18,511 --> 00:10:59,471 [Dr. Ada Shannon] It is. The human only controls two axes — roll and yaw — while thrust and pitch rate are automated by the expert policy, which makes sense given most participants had never flown a drone before. The track is a figure-eight with twelve gates, two of which spatially overlap at the center crossing. Since theta is never directly observable at deployment, the coach maintains a separate belief over skill for each individual gate on the track, updated via Bayesian inference from how fast the learner clears that specific segment relative to a skill-calibrated target time. So it's not one global skill estimate — it's a per-gate belief that lets the coach scaffold differently at a technical gate versus an easy straightaway. 27 00:10:59,471 --> 00:11:06,437 [Hal Turing] And that's the piece that's supposed to separate this from the two baselines they're comparing against, right? Walk me through RBF and MIA. 28 00:11:06,437 --> 00:11:51,530 [Dr. Ada Shannon] RBF is Rule-based Fading, adapted from Shen et al.'s Cyber Racing Coach — it decays assistance along a fixed curve over the session, with no sensitivity to skill or to what's happening physically at that moment. MIA is Minimally-invasive Assistance, an RL safety-filter copilot in the spirit of the human-centered safety-filter line of work, calibrated so its hands-off behavior matches the same frozen expert L2C uses, for a fair comparison. What L2C claims over RBF is closing the loop on skill and context instead of following a preset curve. What it claims over MIA is that MIA is built purely to keep you safe and task-competent right now, not to cultivate independence later — it has no pedagogical objective at all. 29 00:11:51,530 --> 00:12:06,995 [Hal Turing] Okay, but isn't that comparison a little uneven though? RBF was designed for haptic car-racing feedback, MIA for a totally different safety-filter setup — you're taking two systems built for other domains and dropping them into drone racing to lose to the home team. 30 00:12:06,995 --> 00:12:26,917 [Dr. Ada Shannon] I get the instinct, but I don't think it's as uneven as you're making it sound, Hal. All three conditions run on the identical simulator, identical track, identical expert policy, identical verbal and visual cues — the only thing that changes across arms is the assistance-modulation logic itself. That's about as controlled as a between-subjects human study gets. 31 00:12:26,917 --> 00:12:44,054 [Hal Turing] Sure, but the baselines are still hobbled versions of themselves, adapted out of their native domain, while L2C was purpose-built for exactly this task and this action space. That feels like it stacks the deck before a single participant even sits down. 32 00:12:44,054 --> 00:13:10,571 [Dr. Ada Shannon] I actually disagree with you there. Every RL method in this space gets evaluated on whatever benchmark its authors picked — that's true of RBF's original driving paper too. The fair question isn't 'was this L2C's home turf,' it's 'holding the environment fixed, does closing the loop on skill and context outperform a fixed curve or a safety filter.' We can debate what that tells us about generalization, but let's at least get the numbers on the table first. 33 00:13:10,571 --> 00:13:13,961 [Hal Turing] Fair — let's hear them, then we can argue about what they mean. 34 00:13:13,961 --> 00:14:01,748 [Dr. Ada Shannon] For L2C, lap time dropped 27.9% on average post-coaching, p equals 0.005, with a large effect size. Failure count dropped by 3.52 failures per lap, p less than 0.001, very large effect. RBF showed no reliable lap-time change — 11.3% reduction, p equals 0.21 — but did hit significance on failure count, a 2-failure reduction, roughly a third of L2C's effect size. MIA showed no reliable change on either metric — 6.2% on lap time, 1.11 failures per lap, neither significant. This came out of an N of 33 user study: pre-test, then a 15-lap, roughly 40-minute coached session, then a post-test, split across the three arms. 35 00:14:01,748 --> 00:14:06,020 [Hal Turing] And the head-to-head comparisons, L2C directly against each baseline? 36 00:14:06,020 --> 00:14:35,788 [Dr. Ada Shannon] Those are Welch's t-tests with Holm correction, and with only eleven participants per arm, all four contrasts favor L2C with medium-to-large effect sizes, but the p-values land between 0.09 and 0.16 — the paper itself describes that range as expected given the sample-size limit. So within-subject, L2C is the only method that clears significance on both outcomes. Between-group, the comparisons point the same direction but don't clear a conventional threshold at that sample size. 37 00:14:35,788 --> 00:15:18,188 [Hal Turing] So within-subject, L2C clears significance on both lap time and failure count — that part's real, and it's the only method that does both. But here's what bugs me, Ada: the abstract says 'significant gains in human learning outcomes over state-of-the-art AI coaching baselines.' The actual head-to-head numbers, the ones that would justify 'over the baselines,' sit at p equals point-oh-nine to point-one-six after Holm correction. Is that headline earned, or is it borrowing the word 'significant' from a different comparison than the one it's actually describing? 38 00:15:18,188 --> 00:16:00,077 [Dr. Ada Shannon] Fair gotcha, and credit where due, the paper doesn't hide the number — they write it's 'expected given the sample-size limit' right in the results section. But that caveat doesn't survive into the abstract. What's actually established is within-subject: eleven people per arm, each compared against their own baseline self, and L2C moved the needle where RBF and MIA mostly didn't. The between-group claim — L2C beats MIA, L2C beats RBF head-to-head — is directionally consistent with medium-to-large effect sizes, but not conventionally significant. Those are two different flavors of 'significant,' and only one of them actually cleared the bar. 39 00:16:00,077 --> 00:16:35,510 [Hal Turing] That same generosity with scope shows up again later, too. Everything here — training, deployment, the user study — is one task, FPV drone racing, one simulated figure-eight track with twelve gates, and the human only ever controls two of four axes. Yet the abstract talks about 'human motor-skill development' in general, and the conclusion goes as far as coding agents needing 'explicit incentive for what the human retains.' That's a big leap from twelve gates and a joystick. 40 00:16:35,510 --> 00:17:11,222 [Dr. Ada Shannon] It is, and coding agents make the leap especially strained — there's no blending vector, no shared-control channel, no physical skill dynamics to hang a VoI-style reward on. The entire machinery here assumes a continuous control interface that a code-completion tool just doesn't have. They also cite Kulveit and colleagues' Gradual Disempowerment, the 2025 paper on systemic existential risk from incremental AI development, to frame over-reliance as part of something bigger. That's a legitimate concern in its own right, but citing it doesn't establish that a twelve-gate drone study says anything conclusive about it. 41 00:17:11,222 --> 00:17:25,479 [Hal Turing] Oh, come on, that's not really the paper's fault though — that's just framing the stakes! Every conclusion gestures at where an idea could go. You can't hold one sentence in the discussion section to a rigorous evidentiary standard. 42 00:17:25,479 --> 00:17:59,984 [Dr. Ada Shannon] No, I actually disagree with you there. There's a difference between 'this suggests a direction worth exploring' and what they wrote, which reads as a confident claim about what coding agents currently lack. And it fits a pattern — remember Proposition 1, the formal guarantee that skill dominance implies VoI dominance? That's only proven under beta-S equals zero, no downskilling. The actual PFA they train with allows downskilling. So even the math meant to anchor this in rigor only strictly covers a simplified version of their own system, not the one they ran. 43 00:17:59,984 --> 00:18:16,424 [Hal Turing] Okay, I'll give you that one — a proof that doesn't cover your own experimental setup is a real gap, not a nitpick. So where does L2C actually sit next to its closest competitors, since RBF and MIA aren't strawmen, they're adaptations of real prior systems. 44 00:18:16,424 --> 00:19:05,465 [Dr. Ada Shannon] Right. RBF comes straight from Shen and colleagues' Cyber Racing Coach, a haptic shared-control framework out of the University of Michigan, 2025 — its whole design is a fixed fading curve, decaying assistance over time regardless of how the learner's actually doing. Z-COACH, from Srivastava and colleagues at Stanford, also 2025, uses Zone of Proximal Development to pick which sub-skills to target, but still applies a fixed assistance level once it picks. The paper is honest that Z-COACH isn't superseded here — they call it complementary for curriculum design, since L2C only decides how much to help, not what to practice next. Their own limitations section also flags that the PFA may not hold up in messier real skill domains, and that verbal and visual cues remain fixed-rule, not learned. 45 00:19:05,465 --> 00:19:18,096 [Hal Turing] That complementary framing is honestly the most credible part of the whole comparison — it doesn't oversell. So practically, if I'm building assistive tooling, what do I actually take from this beyond drones? 46 00:19:18,096 --> 00:20:04,629 [Dr. Ada Shannon] The transferable idea is the objective, not the drone-specific machinery: measure and reward independent competence, not just in-the-moment performance. That echoes Macnamara and colleagues' 2024 human-subjects study on AI assistance and skill decay — assistance can erode retention without the user even noticing. Worth noting too, Loquercio, a co-author here, also co-authored Kaufmann et al.'s Champion-Level Drone Racing paper out of Zurich, Nature 2023 — that expert-policy lineage is literally what pi-star descends from in this study. None of that changes the core honesty problem, though: what's demonstrated is narrow, what's claimed is broad, and listeners should weight the idea over the drone numbers. 47 00:20:04,629 --> 00:20:24,366 [Hal Turing] Fair takeaway. So: real within-subject gains, a results section that's honest about its own p-values, an abstract that reaches further than the data supports, and a genuinely interesting reframe of what AI assistance should even be optimizing for. That's it for this one — thanks for listening, and we'll catch you next time.