This episode explores Causal-JEPA, a world-modeling approach that masks whole object trajectories rather than image patches to force a model to reason about interactions between entities. It explains how the method combines object-centric representations with JEPA-style latent prediction, asking the model to reconstruct hidden objects from scene context and then predict future dynamics, instead of relying on pixel reconstruction or simple autoregressive rollouts. The discussion highlights the paper’s core argument that this training setup makes counterfactual and causal reasoning more necessary by blocking shortcut strategies like temporal interpolation and self-contained single-object motion prediction. Listeners would find it interesting for its sharp comparison between patch-based scaling and object-centric structure, and for its claim that better world models may come from making interaction reasoning unavoidable rather than merely possible.
Sources:
1. Causal-JEPA: Learning World Models through Object-Level Latent Interventions — Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun, Randall Balestriero, 2026
http://arxiv.org/abs/2602.113892. MONet: Unsupervised Scene Decomposition and Representation — Christopher P. Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=MONet%3A+Unsupervised+Scene+Decomposition+and+Representation3. Multi-Object Representation Learning with Iterative Variational Inference — Klaus Greff, Raphael Lopez Kaufman, Rishabh Kabra, Nick Watters, Chris Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=Multi-Object+Representation+Learning+with+Iterative+Variational+Inference4. Object-Centric Learning with Slot Attention — Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf, 2020
https://scholar.google.com/scholar?q=Object-Centric+Learning+with+Slot+Attention5. Bridging the Gap to Real-World Object-Centric Learning — Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, Francesco Locatello, 2023
https://scholar.google.com/scholar?q=Bridging+the+Gap+to+Real-World+Object-Centric+Learning6. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence7. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas, 2023
https://scholar.google.com/scholar?q=Self-Supervised+Learning+from+Images+with+a+Joint-Embedding+Predictive+Architecture8. Revisiting Feature Prediction for Learning Visual Representations from Video — Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas, 2024
https://scholar.google.com/scholar?q=Revisiting+Feature+Prediction+for+Learning+Visual+Representations+from+Video9. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran, Adrien Bardes, David Fan and many others including Yann LeCun, Michael Rabbat, Nicolas Ballas, 2025
https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning10. CLEVRER: CoLlision Events for Video REpresentation and Reasoning — Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, Joshua B. Tenenbaum, 2020
https://scholar.google.com/scholar?q=CLEVRER%3A+CoLlision+Events+for+Video+REpresentation+and+Reasoning11. Counterfactual VQA: A Cause-Effect Look at Language Bias — Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, Ji-Rong Wen, 2021
https://scholar.google.com/scholar?q=Counterfactual+VQA%3A+A+Cause-Effect+Look+at+Language+Bias12. What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models — Letian Zhang, Xiaotong Zhai, Zhongkai Zhao, Yongshuo Zong, Xin Wen, Bingchen Zhao, 2024
https://scholar.google.com/scholar?q=What+If+the+TV+Was+Off%3F+Examining+Counterfactual+Reasoning+Abilities+of+Multi-modal+Language+Models13. ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos — Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng, 2023
https://scholar.google.com/scholar?q=ACQUIRED%3A+A+Dataset+for+Answering+Counterfactual+Questions+In+Real-Life+Videos14. Towards Causal Representation Learning — Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, Yoshua Bengio, 2021
https://scholar.google.com/scholar?q=Towards+Causal+Representation+Learning15. Interventional Causal Representation Learning — Kartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua Bengio, 2023
https://scholar.google.com/scholar?q=Interventional+Causal+Representation+Learning16. Desiderata for Representation Learning: A Causal Perspective — Yixin Wang, Michael I. Jordan, 2024
https://scholar.google.com/scholar?q=Desiderata+for+Representation+Learning%3A+A+Causal+Perspective17. Provably Learning Object-Centric Representations — Stefan Bauer, Bernhard Schölkopf and collaborators, 2023
https://scholar.google.com/scholar?q=Provably+Learning+Object-Centric+Representations18. SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models — Yuhang Wu, Yueting Zhuang, Francesco Locatello, et al., 2022
https://scholar.google.com/scholar?q=SlotFormer%3A+Unsupervised+Visual+Dynamics+Simulation+with+Object-Centric+Models19. Object-Centric Video Prediction via Decoupling of Object Dynamics and Interactions — Angel Villar-Corrales, Ismail Wahdan, Sven Behnke, 2023
https://scholar.google.com/scholar?q=Object-Centric+Video+Prediction+via+Decoupling+of+Object+Dynamics+and+Interactions20. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning — Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto, 2024
https://scholar.google.com/scholar?q=DINO-WM%3A+World+Models+on+Pre-trained+Visual+Features+enable+Zero-shot+Planning21. Conditional Object-Centric Learning from Video — Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff, 2022
https://scholar.google.com/scholar?q=Conditional+Object-Centric+Learning+from+Video22. Attention over Learned Object Embeddings Enables Complex Visual Reasoning — David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, Matt Botvinick, 2021
https://scholar.google.com/scholar?q=Attention+over+Learned+Object+Embeddings+Enables+Complex+Visual+Reasoning23. Dyn-O: Building Structured World Models with Object-Centric Representations — Zizhao Wang, Kaixin Wang, Li Zhao, Peter Stone, Jiang Bian, 2025
https://scholar.google.com/scholar?q=Dyn-O%3A+Building+Structured+World+Models+with+Object-Centric+Representations24. Learning Interactive World Model for Object-Centric Reinforcement Learning — Fan Feng, Phillip Lippe, Sara Magliacane, 2025
https://scholar.google.com/scholar?q=Learning+Interactive+World+Model+for+Object-Centric+Reinforcement+Learning25. Object-Centric World Model for Language-Guided Manipulation — Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim, 2025
https://scholar.google.com/scholar?q=Object-Centric+World+Model+for+Language-Guided+Manipulation26. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model — Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak, 2026
https://scholar.google.com/scholar?q=Planning+in+8+Tokens%3A+A+Compact+Discrete+Tokenizer+for+Latent+World+Model27. Learning nonparametric latent causal graphs with unknown interventions — Yibo Jiang, Bryon Aragam, 2023
https://scholar.google.com/scholar?q=Learning+nonparametric+latent+causal+graphs+with+unknown+interventions28. Learning Linear Causal Representations from Interventions under General Nonlinear Mixing — Simon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam, Bernhard Schölkopf, Pradeep Ravikumar, 2023
https://scholar.google.com/scholar?q=Learning+Linear+Causal+Representations+from+Interventions+under+General+Nonlinear+Mixing29. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp330. AI Post Transformers: Learning Latent Action World Models from Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-learning-latent-action-world-models-from-1570a4.mp331. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3Interactive Visualization: Causal-JEPA for Object-Level World Models