1 00:00:01,000 --> 00:00:49,994 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "Hierarchical Reinforcement Learning for Air Combat at DARPA's AlphaDogfight Trials," by Adrian P. Pope and Jaime S. Ide, et al. — ten authors total — out of the Lockheed Martin Artificial Intelligence Center and Primordial Labs, with a co-author from the United States Air Force, published in IEEE Transactions on Artificial Intelligence in 2022. And Ada, here's the number that got me: this agent didn't just beat a graduate of the USAF's F-16 Weapons Instructor Course, it beat him five to zero. A clean sweep, best of five, against one of the most qualified fighter pilots on the planet. 2 00:00:49,994 --> 00:01:36,805 [Dr. Ada Shannon] And the context that makes that number land is DARPA's Air Combat Evolution program, ACE, which is ultimately aiming at full-scale piloted aircraft flying with AI assistance. Nobody hands an untested algorithm a live F-16, so DARPA needed a low-risk sandbox first — one-versus-one, within-visual-range dogfighting in a high-fidelity simulator, as a precursor called the AlphaDogfight Trials. It's foundational in the literal sense: pilots master basic maneuvering before they're trusted with escort or suppression-of-air-defense missions. So this was never really about building a killer dogfighting AI for its own sake, it was about building trust incrementally. And the agent in this paper, PHANG-MAN, finished second overall out of eight finalists before that 5-0 human match. 3 00:01:36,805 --> 00:01:50,551 [Hal Turing] Okay, so "hierarchical" is doing a lot of work in that title, and I don't think we've actually covered hierarchical reinforcement learning on the show before. What's the core idea, and what's it buying you here that a plain, flat policy wouldn't? 4 00:01:50,551 --> 00:02:38,709 [Dr. Ada Shannon] Plain deep RL trains one network to map states straight to actions — it has to figure out, at every single timestep, both what tactical situation it's in and what to physically do about it. Hierarchical RL splits that into two layers: a high-level policy makes an infrequent, coarse choice — which specialist handles this — and a low-level policy does the fast, continuous control once that specialist is picked. Think of a manager assigning a project to a team, then getting out of the way of the day-to-day work. This isn't new — Dayan and Hinton proposed that manager-worker framing back in 1993, and DeepMind's FeUdal Networks in 2017 revived it end-to-end. PHANG-MAN's twist is that the "teams" are entire pre-trained specialist policies, not sub-goals. 5 00:02:38,709 --> 00:03:01,186 [Hal Turing] So it's basically the reinforcement-learning cousin of mixture-of-experts routing — a router picking a specialist network, except the routing decision here is temporally extended, you commit to a specialist for a stretch instead of per-token. Which raises the bigger question: what even counts as an "autonomous combat system," and had anyone gotten a learned agent to fly well before this paper? 6 00:03:01,186 --> 00:03:46,605 [Dr. Ada Shannon] Autonomous combat systems is DARPA's umbrella term for AI that pilots or assists piloting combat aircraft, and deep RL just means approximating the policy with neural networks trained by trial and error instead of hand-coded rules. There's real prior art: rule-based systems, game-theoretic utility functions over discrete maneuvers, dynamic programming. The standout was Nick Ernest's genetic fuzzy tree system, ALPHA, out of the University of Cincinnati in 2016, which also beat a retired USAF instructor — but with hand-built decision trees in a lower-fidelity simulator. This paper's environment, JSBSim, is different: it's an open-source F-16 flight dynamics model, six degrees of freedom, taking raw 7 00:03:46,605 --> 00:03:56,868 [Hal Turing] Wait, hold on — raw meaning actual continuous stick-and-rudder inputs? Aileron, elevator, rudder, throttle, not picking from some menu of pre-canned barrel rolls? 8 00:03:56,868 --> 00:04:57,286 [Dr. Ada Shannon] Exactly, fifty times a second — genuinely continuous control, which is a much harder problem than the maneuver-library approach and a real step toward something you could eventually trust near an actual airframe. Two more pieces of vocabulary before we go further: curriculum learning just means training easy-to-hard — scripted, weak opponents first, tougher intelligent ones only once you're winning consistently — because dropping a random policy into a fight against a skilled opponent from day one gives it nothing to learn from. Maximum entropy RL is the training objective PHANG-MAN uses via Soft Actor-Critic: maximize reward plus the randomness of your own actions, so the policy stays exploratory instead of collapsing onto one predictable move. And a "gun snap" — nobody's firing anything in JSBSim — just means hitting the geometry that would let you: inside the Weapons Engagement Zone, five hundred to three thousand feet out, within a two-degree cone off your nose. 9 00:04:57,286 --> 00:05:23,339 [Hal Turing] So that's the vocabulary and the stakes: a two-level hierarchical agent, trained with curriculum learning and max-entropy SAC, trying to out-fly human expert pilots in a simulator faithful enough that DARPA was willing to bet real trust on the result. So let's actually open up PHANG-MAN itself, because hierarchical so far has been an abstraction. What's literally sitting inside this thing? 10 00:05:23,339 --> 00:06:14,562 [Dr. Ada Shannon] Two layers. At the bottom you've got three low-level policies, Control Zone, Aggressive Shooter, and Conservative Shooter, each one a separate SAC network running at fifty hertz, so they're the ones actually moving the stick. Above them sits the policy selector, evaluated at ten hertz, once every five low-level actions, and its whole job is picking which of those three gets control given the current geometry of the fight. Here's the design choice that matters: the low-level policies are trained first, completely independently, and then frozen solid before the selector ever sees a single frame. The paper justifies that by citing Nachum and colleagues' HIRO work out of Google Brain, 2018, saying freezing avoids the non-stationarity problem you'd otherwise get training both levels at once. 11 00:06:14,562 --> 00:06:30,212 [Hal Turing] Wait, but HIRO's whole contribution is an off-policy correction that lets you train the low-level policy while it's still changing underneath the high-level policy, right? That's not a paper about why you should freeze things, that's a paper about how to not have to. 12 00:06:30,212 --> 00:07:15,956 [Dr. Ada Shannon] Right, and they never run that comparison. They cite the non-stationarity problem HIRO solves, then sidestep it entirely rather than adopting the solution, and there's no ablation showing frozen beats joint training here. It's a real gap between the citation and the actual design choice. The other reference point worth naming is Frans and colleagues' Meta Learning Shared Hierarchies, out of OpenAI and Berkeley, 2018, which has the identical selector-plus-sub-policies skeleton. The difference is MLSH meta-learns shared sub-skills across a distribution of tasks so they transfer; PHANG-MAN trains three fixed specialists for one task, dogfighting, and just learns when to hand off between them. Structurally similar, ambition-wise much narrower. 13 00:07:15,956 --> 00:07:26,776 [Hal Turing] So the sub-policies aren't learning general skills, they're basically three flavors of one job. Which brings me to how each one gets its personality, and that's all reward shaping, isn't it— 14 00:07:26,776 --> 00:08:07,690 [Dr. Ada Shannon] Sorry, jump in here because this is the fun part. Every low-level policy shares the same handful of ingredients, track angle to the opponent's nose, adverse angle off their tail, closure rate, gun-snap proximity, and a deck-avoidance penalty for flying too low, but the weighting is completely different per policy. Control Zone rewards a full positional-dominance geometry, nothing the opponent can do about it. The two shooters drop that positional term entirely and reward pure track angle, so they'll take head-on shots Control Zone would never attempt. Conservative Shooter's gun-snap reward stays flat with distance; Aggressive Shooter's spikes up close. Same ingredient list, three completely different fighting styles just from reweighting. 15 00:08:07,690 --> 00:08:14,377 [Hal Turing] And you get all three of those into fighting shape through curriculum learning, which is where the self-play piece comes in, presumably. 16 00:08:14,377 --> 00:08:50,043 [Dr. Ada Shannon] Exactly the mechanism. Training starts against scripted, easy opponents, and only once win rate crosses fifty percent do intelligent opponents get introduced, earlier PHANG-MAN iterations, individual low-level policies, hand-built agents mimicking rival teams. Eventually you're sampling from thirty total opponents, weighted by a win/loss ratio computed over the last hundred matchups, updated after every single game, clipped between point-two and eleven-point-seven percent so nothing gets neglected or over-trained-against. It's an automatic difficulty ladder that keeps pointing the agent at whatever's currently beating it. 17 00:08:50,043 --> 00:08:55,151 [Hal Turing] That's a lot of simultaneous games. What's actually running underneath all that, compute-wise? 18 00:08:55,151 --> 00:09:40,198 [Dr. Ada Shannon] An Ape-X-style setup, distributed actors each running their own JSBSim instance and pushing experience into a central prioritized replay buffer, with one learner consuming it. Wide, shallow MLPs, twelve-thousand-plus neurons in a single hidden layer, which is a deliberate choice, wide-shallow nets had already been shown to outperform deep-narrow ones at equal parameter count. And there's a wonderfully ugly bug they had to fix: identical networks on different GPU architectures, Volta versus Tesla, produced outputs differing at the sixth decimal, and compounded fifty times a second that drift was enough to flip wins into losses. Fix was brute-force truncating states to six decimals and actions to two before they ever hit the network. 19 00:09:40,198 --> 00:09:46,885 [Hal Turing] So after all that scaffolding, did the hierarchy actually earn its keep, or could one good specialist have done just as well? 20 00:09:46,885 --> 00:10:23,712 [Dr. Ada Shannon] The hierarchy earned it. Selector-driven PHANG-MAN matched or beat its own best single low-level policy against every opponent they tested, and often by a wide margin, meaning it's genuinely combining strengths rather than just picking a favorite and sticking with it. You can watch that show up across the agent iterations in their tables too, each version's win-loss delta climbing as they added reward terms, truncation fixes, and better opponent pools. That competition run included a semifinal win and a loss in the championship final, before the clean five-nothing sweep against an actual USAF Weapons Instructor Course graduate. 21 00:10:23,712 --> 00:11:00,214 [Hal Turing] So let's push on this, because it's been nagging me the whole discussion: every number here—the training curve, the agent selection, even that 5-0 sweep against the USAF pilot—happens inside JSBSim with zero sensor noise and complete, noise-free knowledge of the opponent's position, velocity, attitude, and health. That's a God's-eye view, not partial observability with an estimation layer bolted on. So how much of PHANG-MAN's edge is really the hierarchical architecture, versus just being handed information no real cockpit sensor suite would ever give a pilot? 22 00:11:00,214 --> 00:11:42,614 [Dr. Ada Shannon] It's baked in deeper than a footnote. The reward terms depend on ground-truth distance, track angle, and closure rate handed straight from the sim, not estimates—that's the load-bearing substrate of the whole method, not a peripheral simplification. To their credit, the authors don't dodge it. Their own Limitations and Future Work section explicitly lists imperfect information, partial and estimated knowledge, and domain randomization as things they haven't done—future work, not solved problems. So the paper's honest that this is an idealized slice of the problem. That honesty doesn't change the fact that 'outperformed human expert pilots' is doing a lot of work for a result earned entirely inside a noise-free simulator. 23 00:11:42,614 --> 00:12:13,636 [Hal Turing] Oh, hold on—sorry to jump in, but this is the part that actually undermines the causal story even more: the supplementary material's own confession about why they lost the championship final. Buried in there, they admit the loss to Heron Systems was partly caused by a training hack—inflating aircraft health ten-fold and leaving the opponent's health out of the reward and state entirely. That's not a footnote, that's a confound sitting right on top of their headline result. 24 00:12:13,636 --> 00:12:58,079 [Dr. Ada Shannon] Right, and it's a real one. They say the hack biased PHANG-MAN toward repositioning over finishing kills. In the final, they landed seven percent more total shots than Heron's agent, but from farther away, so each gun snap did less damage, and the agent kept disengaging inside 800 feet instead of pressing an advantage a more aggressive policy would've closed. So the explanation for the one loss that mattered most isn't 'the hierarchy failed'—it's a reward-shaping artifact that trained conservative behavior at the worst possible moment. And they never re-ran the final without that hack to isolate whether it was really the cause. The claim that the hierarchy drives the tactical quality sits on top of an acknowledged, unmeasured confound. 25 00:12:58,079 --> 00:13:11,314 [Hal Turing] Given the MLSH parallel we mentioned and the HIRO freezing rationale we already picked apart, is swapping three frozen specialists on one task a meaningfully different contribution, or mostly MLSH's idea applied to a new domain? 26 00:13:11,314 --> 00:13:55,478 [Dr. Ada Shannon] Honestly, mostly the latter, and it should be named plainly. Where it genuinely differs from MLSH is that the sub-policies aren't shared skills reused across tasks—they're full competing strategies for one task, a simpler design, not a harder one. It's also not the first system to claim beating a weapons-school graduate — Ernest's genetic fuzzy tree system we mentioned earlier gets barely a mention here despite the near-identical headline claim. There's also Sun, Piao, Yang, Zhao and colleagues, 2021, who built a hierarchical air-combat agent using genuine multi-agent self-play instead of a fixed curriculum—exactly the population-level training this paper's own discussion flags as unexplored, already existing. 27 00:13:55,478 --> 00:14:30,494 [Hal Turing] Stepping back—who does this actually matter to right now? ACE's whole premise is building trust in autonomy incrementally, and an agent whose behavior you can watch, literally which specialist it's invoking and when, seems genuinely more useful for human-machine teaming than an opaque end-to-end policy. But every ounce of that trust needs the caveat printed right in the paper: single opponent, no missiles, visual-range only, pure simulation, and an AI competitor already beat it in the same tournament. 28 00:14:30,494 --> 00:15:09,921 [Dr. Ada Shannon] And the authors are upfront about what's next: sim2real transfer via domain randomization — Tobin, Fong, Ray, Schneider, Zaremba, and Abbeel, OpenAI, 2017 — and porting this to other airframes, quadcopters, small fixed-wing UAVs, even the F-22. None of that exists yet. So there's a real gap between the impact statement's language about operationalizing autonomy for real-world defense and what's actually demonstrated: one F-16 against scripted and self-play opponents, plus a five-match sample against a single human pilot. The paper is honest about that gap in its own limitations section—it's the framing up front that outruns it a little. 29 00:15:09,921 --> 00:15:48,374 [Hal Turing] So here's where I land. The contribution is real—first hierarchical RL agent to beat expert fighter pilots in this competition, and the policy-selector design is a clean, legible way to combine specialists. But the causal story, that the hierarchy itself produces the good tactics, is confounded on two fronts: an idealized, noise-free observation space baked into every reward term, and an acknowledged training artifact that skewed the one head-to-head loss that mattered most. Good result, overclaimed framing. Thanks for flying along with us today, and we'll see you next time.