Object masking changes the training game: instead of reconstructing pixels or rolling forward every slot autoregressively, Causal-JEPA hides whole object trajectories, reconstructs them from scene context, and then predicts future dynamics. The point is not just object structure. The point is to make interaction reasoning harder to avoid.
The visual argument of the paper is simple: if hidden slots cover enough of an object’s past, the predictor cannot survive on cheap single-object extrapolation. It has to read collision traces, occlusions, and relative motion from the rest of the scene.
Rows are training recipes. Columns are failure modes the episode emphasized: temporal interpolation, inertial self-dynamics, and interaction avoidance. Hotter cells mean the shortcut remains available.
Use the stepper to watch the hidden red object disappear from history. What remains visible are the social traces of physics: impact timing, reflected motion, and geometry of nearby entities.
The model’s job is not “guess the next frame.” It is “recover the missing entity from the rest of the movie, then continue the movie.”
Toggle between visible context and inferred dependency strength. Hover cells to see which observed objects carry the hidden object’s recoverable evidence.
The episode separated two claims. The within-family masking ablation supports the reasoning story. The planning story is more about latent compression and systems efficiency than about pure causal metaphysics.
Mock data visualizes the qualitative claim from the episode: when the planner attends over object slots instead of a large patch sea, the compute curve bends down sharply enough to matter.
The paper uses causal language to describe the bias induced by masking. The stronger read is not “causality proven.” It is “the objective makes interaction dependence more necessary than unmasked latent prediction does.”
This positions the method near object-centric JEPA models that apply pressure in the objective rather than hard-coding separate interaction modules.