TD3-style actor-critic RL blended with a behavior-cloning term from simulator-generated expert trajectories, trained on a three-stage pursuit → lock → launch dogfight task. Reports near-100% success where pure-RL baselines collapse to ~0% — but never benchmarks against DAgger, the standard fix for imitation learning's distribution-shift problem.
The task: a Markov Decision Process with three sequential subtasks and a hard terminal condition — one missile, no do-overs.
Own aircraft spawns at a random position beyond lock range; opponent is fixed. Lock requires distance 100–3000 units AND angle <15° held for a full 5 seconds.
Continuous flight controls + one irreversible discrete decision.
State → Action → Reward → Transition, repeated every control step inside Harfang3D.
Two loss terms share one actor: a TD3 critic-driven RL loss, and a behavior-cloning loss toward PID-autopilot expert actions, weighted by λ.
Twin critics (Q1, Q2) estimate value pessimistically; the actor is pulled by both the critic gradient and the expert BC loss.
Loss = (1−λ)·L_RL + α·λ·L_BC. Toggle between the fixed linear decay and the Q-value-driven adaptive scheme.
Sparse-reward-at-the-end would be brutal with one missile and a 5s hold — so reward is shaped across every step.
Six methods × three fixed (non-adaptive) opponent behaviors. Proposed method wins broadly; pure-RL baselines master pursuit and lock but never learn to pull the trigger correctly.
Each cell: does the method's policy master this stage? TD3 / DSAC-v2 stall at launch.
Single-missile (training condition) vs. unlimited-missile stress test — hit rate per launch attempt.
The paper claims "adaptability to dynamic environments" — but every opponent is a fixed, non-learning script, the expert is a cheap PID autopilot, and the one obvious control experiment is missing.
Pure BC drifts off the expert's demonstrated states and compounds error. DAgger iteratively re-queries the expert to correct drift. This paper never tests that fix.
Cited in Related Work, never benchmarked — despite sharing the paper's own loss formula.
Return swings roughly ±300 across runs, yet Table 3 reports best-of-4, not mean-of-4.