1 00:00:01,000 --> 00:00:37,455 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're talking drones — but not delivery drones, acrobatic ones. The paper is "Deep Drone Acrobatics," first author Elia Kaufmann et al., six authors total, including Antonio Loquercio, René Ranftl, Matthias Müller, Vladlen Koltun, and Davide Scaramuzza, out of the University of Zurich, ETH Zurich, and Intel's Intelligent Systems Lab. It ran at Robotics: Science and Systems, 2020. 2 00:00:37,455 --> 00:01:05,922 [Dr. Ada Shannon] Here's the number that got me: three g's of acceleration, on a quadrotor flying a Power Loop, a Barrel Roll, and something called a Matty Flip, using nothing but a camera and an IMU strapped to the frame. No Vicon, no OptiTrack, no motion capture rig watching from the ceiling. And the policy that flies it has never once touched a real camera image during training — it learned entirely in simulation and then got bolted onto a physical drone and just... flew the maneuver. 3 00:01:05,922 --> 00:01:39,220 [Hal Turing] Right, and that's the core question the paper's chasing: can a policy trained purely in simulation actually fly extreme acrobatics zero-shot on real hardware, using only what's onboard? Why is this even hard? Because at those accelerations you get motion blur wrecking your vision-based state estimate, and the control margins are so tight that a tiny mistake doesn't mean a wobble, it means a crashed quadrotor. Human pilots train for years to fly a barrel roll safely. This is asking a neural net to do it with no do-overs on the real platform. 4 00:01:39,220 --> 00:02:14,421 [Dr. Ada Shannon] And it's worth being specific about the failure mode. Vision-based state estimation — think visual-inertial odometry, VINS-Mono style pipelines — relies on tracking features frame to frame. Crank up the angular rate and the camera image smears, features get lost, and your position estimate either degrades or blows up entirely. That's a known, cited problem in this space. So previous agile-flight work mostly dodged it by assuming near-perfect external state estimation. This paper doesn't get that luxury — it has to survive on onboard sensing alone, mid-flip. 5 00:02:14,421 --> 00:02:25,010 [Hal Turing] So — and correct me if I'm wrong here — this is basically reinforcement learning, right? The policy tries stuff in sim, gets punished for crashing, and gradually gets better at not dying during a loop? 6 00:02:25,010 --> 00:02:59,747 [Dr. Ada Shannon] No, no, that's not how I read it at all. There's zero reward shaping, zero exploration-and-punish loop here. This is imitation learning — specifically DAgger, dataset aggregation. There's a privileged expert, a model-predictive controller with full ground-truth state access, that already knows how to fly the maneuver perfectly. The student network just watches what the expert would do at states the student itself visits, and imitates it. No trial and error against a reward signal, no policy gradient. That distinction actually matters a lot for how you should read every result later in this paper. 7 00:02:59,747 --> 00:03:11,542 [Hal Turing] Oh wait wait wait — hold on, back up. So the student flies on its own, generates a trajectory, and then the expert basically grades that trajectory after the fact and hands back corrections? That's the DAgger loop? 8 00:03:11,542 --> 00:03:52,967 [Dr. Ada Shannon] Exactly that. Ross, Gordon, and Bagnell laid out DAgger back in 2011 — you fly with your current student policy, log the states you actually visited, have the expert label the *correct* action at each of those states, fold that into the training set, and retrain. Repeat. It fixes the classic imitation-learning problem where a student drifts slightly off the expert's trajectory and then has no idea what to do because it never saw that state during training. Here the privileged expert is an MPC that operates on simplified drone dynamics with full state access — position, velocity, attitude, all of it, ground truth, no noise. 9 00:03:52,967 --> 00:04:25,382 [Hal Turing] Okay, that lands. So the real trick isn't the learning algorithm, it's the gap between what that expert sees — perfect state — and what the deployed student sees, which is just onboard camera and IMU. That's the sim-to-real problem in a nutshell: the observation models don't match, so a policy trained on one won't behave the same on the other. And their fix is what they call input abstraction — instead of feeding raw pixels, which look wildly different in simulation versus a real camera, they feed the network geometry-based feature tracks that look roughly the same in both worlds. 10 00:04:25,382 --> 00:05:43,772 [Dr. Ada Shannon] Right, and it's worth flagging where that idea comes from before we go further, since it's not original to this paper — Zhou, Krähenbühl, and Koltun's "Does Computer Vision Matter for Action?" out of Intel Labs and UT Austin, Science Robotics 2019, made the formal case that abstracting sensory input shrinks the sim-to-real gap. This paper builds two lemmas directly on top of it: Lemma 1 says the sim-to-real performance gap is bounded by how different the simulated and real observation models are, and Lemma 2 says that if you pass both observations through the same abstraction function first, that distance can only shrink — so instead of chasing photorealistic simulation, you strip the signal down to what sim and reality actually share. The direct ancestor for the privileged-learning setup itself is "Learning by Cheating," Chen, Zhou, Koltun, and Krähenbühl, also Intel Labs and UT Austin, CoRL 2019 — a privileged expert driving a car teaches a sensor-limited student to drive. But acrobatics is a nastier version of that problem: driving has a shoulder to pull onto when you're unsure. A quadrotor mid-flip doesn't get to slow down and think about it. 11 00:05:43,772 --> 00:05:53,943 [Hal Turing] So walk me through what the policy is actually chasing, mechanically. It maps observations to actions, sure, but how is the target maneuver itself defined and handed to it? 12 00:05:53,943 --> 00:06:34,763 [Dr. Ada Shannon] The policy pi maps onboard observations to an action — thrust plus body rates — trained to minimize squared error against a reference trajectory at every timestep. That reference is planned in flat-output space, position and yaw, because any smooth trajectory there is guaranteed flyable by the underactuated platform. The acrobatic core is a circular motion primitive at constant tangential velocity, with one sharp constraint: speed has to exceed the free-fall velocity at the top of the loop by a ten-percent margin, or the required thrust orientation becomes undefined. Seventh-order polynomials then stitch in the entry, transitions, and exit segments. 13 00:06:34,763 --> 00:06:52,875 [Hal Turing] And the thing generating training labels off that reference is the MPC expert you mentioned. But here's what I actually want to understand — why feature tracks specifically? This same group did Deep Drone Racing, and that one leaned on domain randomization, throwing wild textures and lighting at the sim. Why not just do that again? 14 00:06:52,875 --> 00:07:34,671 [Dr. Ada Shannon] The expert's an MPC running on a simplified dynamics model with full ground-truth state — no perception problem to solve, just tracking. And on feature tracks versus randomization, they made a deliberate switch. Deep Drone Racing randomized textures and lighting hoping the network would learn invariance through exposure. Here they skip that bet entirely — the visual frontend is lifted from VINS-Mono, Harris corners tracked with Lucas-Kanade, filtered by distance from the epipolar line. Feature tracks encode scene geometry and ego-motion, not surface appearance, so a simulated wall and a real wall produce statistically similar tracks even though they look nothing alike as pixels. 15 00:07:34,671 --> 00:07:39,594 [Hal Turing] Okay — so how does that actually get built into the network itself? What's the architecture look like? 16 00:07:39,594 --> 00:08:19,021 [Dr. Ada Shannon] Architecturally it's three branches running asynchronously, since the camera and IMU update at different rates. Feature tracks go through a small PointNet-style net into a temporal convolution stack, IMU gets its own temporal conv stack, and the reference trajectory history gets a third. All three concatenate into a plain three-layer MLP outputting thrust and body rates at 100 Hz. Training is off-policy DAgger for 150 rollouts total in batches of thirty. They randomize IMU bias and thrust-to-weight ratio by up to ten percent each iteration — but zero randomization of scene geometry or appearance, which matters for what's coming. 17 00:08:19,021 --> 00:08:27,845 [Hal Turing] And does that actually pay off numerically, or is this one of those papers where the ablation table quietly undercuts the headline claim? Give me the real comparison. 18 00:08:27,845 --> 00:09:14,888 [Dr. Ada Shannon] It holds up. Full model: 24 centimeters tracking error, 100 percent success across all four maneuvers, versus 43 centimeters for a strong VIO-plus-MPC baseline — up to 45 percent lower error, gap widening on the hardest combo sequence. Strip feature tracks and keep only IMU, the No-FT ablation, and it's still 28 centimeters and near-perfect on individual maneuvers, but drifts harder on the twenty-second repeated-loop endurance test where full Ours never fails. Swap feature tracks for raw images and it's a different story: 80 percent success in the training scene, then zero the moment you drop in unseen COCO backgrounds. Feature tracks held 100 percent in both novel test scenes. On the physical drone, both Ours and No-FT largely succeed too. 19 00:09:14,888 --> 00:09:34,207 [Hal Turing] Wait — say that physical result again. Ten runs per maneuver, and No-FT is already at ninety to a hundred percent on that same small sample. That's not much room to statistically separate 'abstraction helps' from 'the expert labels were just good and both policies got lucky ten times.' 20 00:09:34,207 --> 00:10:40,895 [Dr. Ada Shannon] I actually disagree with you there, Hal — ten runs isn't nothing when it's a physical quadrotor pulling 3g with zero crashes; that's not a cheap result to fake. But you're right it's thin for any real confidence interval, and there's a second issue stacked on it: all of that flying happened in one fixed real-world space, no geometry or appearance randomization anywhere in training. So is this zero-shot to a genuinely novel scene, or transfer to a space that happens to already resemble the sim? There's a second data point in that same table that bugs me even more: the Combo maneuver — triple Barrel Roll, double Power Loop, then a Matty Flip, no stopping between any of them — is the only test of long-horizon compounding difficulty in the whole paper, and it drops to 95% while both ablations, No FT and Only-Ref, collapse straight to zero. The paper never says why. Is that accumulated drift across the transitions? A specific hand-off where orientation error compounds? One unlucky rollout out of twenty? We don't know, and that's exactly the test that's supposed to be stress-testing compounding failure. 21 00:10:40,895 --> 00:11:03,372 [Hal Turing] That connects to something else that bothered me on the physical side. Real-world 'success' isn't measured by any sensor — it's whichever the safety pilot judges as 'executed correctly,' no ground-truth position error, no stated threshold, no second observer cross-checking the call. So what's actually backing that number? 22 00:11:03,372 --> 00:11:43,682 [Dr. Ada Shannon] Nothing quantitative, and the paper is upfront about why: it literally states the tracking-error metric can't be computed outdoors at all, because it requires exact state estimation that only exists in simulation. So every claim about Ours beating VIO-MPC on accuracy — the 45% lower tracking error — is a simulation-only result. On hardware you get a binary pass or fail from one person's judgment call, nothing that confirms the sim ranking survives contact with reality. It's entirely plausible the real robot is tracking worse than VIO-MPC while still 'succeeding' by the pilot's definition, and we'd never see it in the data. 23 00:11:43,682 --> 00:12:36,577 [Hal Turing] Oh — wait, hold on, that actually loops back to something more basic I keep snagging on. The privileged expert isn't some oracle that can fly anything — it's an MPC tracking a reference trajectory that was already built with polynomial constraints respecting the platform's actual physical limits, back in Eq. 4. So the 'privilege' isn't unlimited state access plus magic; it's state access plus a trajectory a classical controller could already fly given the true state. Doesn't that mean the real contribution here is sim-to-real transfer of state estimation, not the acrobatics themselves? The acrobatic capability was arguably already solved by MPC — the paper's actual trick is getting a student to approximate that MPC without seeing the state. 24 00:12:36,577 --> 00:13:31,840 [Dr. Ada Shannon] That's a fair reframing, and I'd go further: it's the same blind spot as treating VINS-Mono's feature-track frontend as a fixed, untouchable abstraction layer. Meanwhile a whole separate line of research is trying to learn VIO away entirely — end-to-end, from pixels. This paper leans on classical VIO as scaffolding rather than asking if it'll still hold up. And the randomization they do apply — IMU bias and thrust-to-weight, plus or minus 10% — never touches vision at all. What happens to those Harris-corner tracks in real outdoor lighting, motion blur, or a textureless wall that looks nothing like the fisheye renders in Figure 5? We don't get an answer. Side note — Antonio Loquercio, one of the authors here, later co-authored 'AI Coaching for Accelerating Human Skill Development with Reinforcement Learning,' so this abstraction-for-transfer instinct clearly followed him. 25 00:13:31,840 --> 00:13:58,032 [Hal Turing] And zooming out even further — this is one quadrotor, 1.15 kilograms, 4-to-1 thrust-to-weight, with hard-coded constants like gravity and specific loop radii baked into the reference trajectories. The conclusion claims the methodology 'is not limited to autonomous flight and can enable progress in other areas of robotics.' That's a big reach off one platform, one room, and pre-scripted circles. 26 00:13:58,032 --> 00:14:58,451 [Dr. Ada Shannon] I'd push back gently on how hard we read that line — it's a conclusion-section flourish, not the paper's central claim, and the core contribution, sim-to-real transfer via abstraction for a genuinely brutal control problem, stands on its own. But you're right that the generalization case isn't made here. The modular-abstraction idea itself comes from Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem, and Vladlen Koltun's 'Driving Policy Transfer via Modularity and Abstraction' out of Intel Labs and KAUST, CoRL 2018 — validated first in driving with semantic segmentation, not feature tracks, which raises exactly the domain-specificity question you're asking. And the performance bound underpinning all of this, from Yunpeng Pan and colleagues' 'Agile Autonomous Driving Using End-to-End Deep Imitation Learning,' Georgia Tech, RSS 2018, assumes Lipschitz continuity and a fixed smoothness constant — asserted here, never verified for acrobatics specifically. 27 00:14:58,451 --> 00:15:29,473 [Hal Turing] Which does matter practically, though — onboard-only acrobatic flight with no motion-capture rig is genuinely useful for inspection drones, racing, search-and-rescue, anywhere you can't instrument the environment. Where this heads next is whether the same abstraction trick — geometry-driven intermediate features instead of raw pixels — carries over to manipulation or legged locomotion, and whether it survives contact with actual unstructured, novel environments instead of one tuned flying space. 28 00:15:29,473 --> 00:15:55,711 [Dr. Ada Shannon] Which ties it straight back to where this line of work started — Chen, Zhou, Koltun, and Krähenbühl's 'Learning by Cheating' gave it the privileged-teacher framework, and Zhou, Krähenbühl, and Koltun's 'Does Computer Vision Matter for Action?' gave it the abstraction argument. This paper's real achievement is proving that lineage survives the jump from steering a car to pulling three g's in a barrel roll. The open question is whether it survives the next jump too. 29 00:15:55,711 --> 00:16:22,693 [Hal Turing] That's a good place to land it. Bottom line: genuinely first-of-its-kind acrobatic transfer, built on borrowed theory that mostly holds up, with evaluation that's thinner than the headline suggests — small samples, one room, subjective success calls, and a generalization claim the experiments don't back yet. Thanks for flying through this one with me, Ada. 30 00:16:22,693 --> 00:16:26,733 [Dr. Ada Shannon] Anytime, Hal. Thanks for listening, everyone — we'll catch you next time.