Imitative Reinforcement Learning for UCAV Pursuit-Lock-Launch Dogfights

arXiv:2406.11562 Siyuan Li, Rongchang Zuo, Bofei Liu, Yaoyu He, Peng Liu, Yingnan Zhao · Harbin Institute of Technology + collaborators · 2024

TD3-style actor-critic RL blended with a behavior-cloning term from simulator-generated expert trajectories, trained on a three-stage pursuit → lock → launch dogfight task. Reports near-100% success where pure-RL baselines collapse to ~0% — but never benchmarks against DAgger, the standard fix for imitation learning's distribution-shift problem.

The task: a Markov Decision Process with three sequential subtasks and a hard terminal condition — one missile, no do-overs.

Mission Pipeline: Pursuit → Lock → Launch

Own aircraft spawns at a random position beyond lock range; opponent is fixed. Lock requires distance 100–3000 units AND angle <15° held for a full 5 seconds.

Action Space

Continuous flight controls + one irreversible discrete decision.

MDP Loop

State → Action → Reward → Transition, repeated every control step inside Harfang3D.

Two loss terms share one actor: a TD3 critic-driven RL loss, and a behavior-cloning loss toward PID-autopilot expert actions, weighted by λ.

TD3 + Behavior-Cloning Architecture

Twin critics (Q1, Q2) estimate value pessimistically; the actor is pulled by both the critic gradient and the expert BC loss.

λ Weighting Schedule

Loss = (1−λ)·L_RL + α·λ·L_BC. Toggle between the fixed linear decay and the Q-value-driven adaptive scheme.

λ over training episodes

Reward Shaping (5 terms)

Sparse-reward-at-the-end would be brutal with one missile and a 5s hold — so reward is shaped across every step.

Six methods × three fixed (non-adaptive) opponent behaviors. Proposed method wins broadly; pure-RL baselines master pursuit and lock but never learn to pull the trigger correctly.

Success Rate by Opponent

Policy Mastery Heatmap

Each cell: does the method's policy master this stage? TD3 / DSAC-v2 stall at launch.

Launch Efficiency

Single-missile (training condition) vs. unlimited-missile stress test — hit rate per launch attempt.

The paper claims "adaptability to dynamic environments" — but every opponent is a fixed, non-learning script, the expert is a cheap PID autopilot, and the one obvious control experiment is missing.

Distribution Shift: BC vs. DAgger vs. This Paper

Pure BC drifts off the expert's demonstrated states and compounds error. DAgger iteratively re-queries the expert to correct drift. This paper never tests that fix.

The Missing Baseline

Cited in Related Work, never benchmarked — despite sharing the paper's own loss formula.

Reported Variance: Circling Opponent

Return swings roughly ±300 across runs, yet Table 3 reports best-of-4, not mean-of-4.

References