Podcast Visual Companion

Learning Latent Action World Models from Video

A visualization-first guide to the paper’s core idea: learn an action-like bottleneck from unlabeled video, then test whether it behaves like control rather than a generic future-information shortcut.

arXiv 2601.05230 2026 • Garrido et al. World Models • IDM + Forward Model Setting: in-the-wild video
Core Tension
Dynamics vs. Shortcut
Best Bottleneck
Constrained Continuous
Main Risk
Camera / Edit Confounds
Practical Payoff
Pretraining for Planning
Mock quantitative values below are synthesized from the episode’s reported pattern: continuous constrained latents outperform VQ on controllability and transfer, while planning reaches near-parity with action-conditioned baselines.