AI Post Transformers · Episode Companion

Hierarchical RL Beats an F-16 Instructor Pilot 5-0

Adrian P. Pope, Jaime S. Ide, Daria Micovic, Henry Diaz, David Rosenbluth, Lee Ritholtz, Jason C. Twedt, Thayne T. Walker, Kevin Alcedo, Daniel Javorsek — Lockheed Martin AI Center / Primordial Labs / USAF · IEEE Transactions on AI, 2022
arXiv:2105.00990 ↗ DARPA AlphaDogfight Trials Hierarchical RL Soft Actor-Critic JSBSim F-16
System Overview
PHANG-MAN routes control between three frozen specialist policies via a slower policy selector — a temporally-extended cousin of mixture-of-experts routing, applied to continuous F-16 flight control.

Control Pipeline

Ground-truth combat geometry → 10 Hz policy selector → one of three frozen 50 Hz specialists → continuous stick/rudder into JSBSim.
Sim State position, velocity, opponent geometry Policy Selector 10 Hz 1 choice per 5 low-level actions Control Zone positional dominance SAC · frozen · 50Hz Aggressive Shooter close-range gun snaps SAC · frozen · 50Hz Conservative Shooter track-angle at any range SAC · frozen · 50Hz JSBSim F-16 aileron/elevator/ rudder/throttle 5-0
Hover any box — the selector picks a specialist for a stretch of time (temporally extended), unlike per-token routing in mixture-of-experts.

The Headline Result

PHANG-MAN vs. a USAF F-16 Weapons Instructor Course graduate, best-of-five, simulated within-visual-range dogfight.
Two-Layer Hierarchy & Reward Shaping
All three specialists share the same reward ingredients — only the weighting differs, producing three distinct fighting styles from one recipe.

Reward Term Weighting by Specialist

Hover a cell for the exact mock weight. Same five geometric ingredients, reweighted per policy (illustrative, based on the paper's reward description).
low weight mid weight dominant weight

Update Frequency Mismatch

The selector commits for 5 low-level ticks — a deliberate temporal abstraction.

Freeze-then-Route vs. HIRO

The paper cites HIRO's non-stationarity problem, then sidesteps it by freezing low-level policies entirely — no ablation against joint training is run.
Gap flagged in the episode: HIRO's actual contribution is an off-policy correction that lets you train jointly. PHANG-MAN cites the problem, then avoids solving it.
Curriculum Learning & Opponent Pool
Training starts against weak scripted opponents; tougher intelligent opponents enter only once win rate clears 50%, building an automatic difficulty ladder.

Win Rate Over Training, With Curriculum Events

Mock training curve illustrating the >50% gate that triggers each new tier of opponents.

Opponent Pool Sampling Weight (of 30 total)

Weighted by win/loss ratio over the last 100 matchups, clipped between 0.2% and 11.7% so nothing is neglected or over-trained-against. Top 10 shown.
Does the Hierarchy Earn Its Keep?
Selector-driven PHANG-MAN matched or beat its own best single specialist against every tested opponent — but the championship-final loss reveals a confound.
Win rate (%) by opponent tier — the hierarchical selector never underperforms its best individual specialist.

The Training-Hack Confound

In the tournament final, an inflated-health / no-opponent-health-signal hack biased PHANG-MAN toward repositioning over finishing kills.
PHANG-MAN landed 7% more total shots than Heron's agent — but from farther away, doing less damage per snap, and disengaged inside 800 ft instead of pressing the advantage. The final was never re-run without the hack.

References