Visualization Companion arXiv 2604.24764 Interactive Viz URL Short-horizon geometry, not full world simulation

World-R1 Improves 3D Consistency in Text-to-Video

World-R1 wraps reinforcement learning around an existing text-to-video model and scores the generated clips with pretrained 3D critics plus a vision-language judge. The visual question is whether camera motion reveals a stable scene, or whether the world liquefies when the viewpoint shifts.

Base backboneWan2.1-T2V 1.3B / 14B
OptimizationFlow-GRPO post-training
Main signal3D reconstruction + VLM rewards
Extra arXiv IDs in transcriptnone beyond 2604.24764
Core claim
Geometry ↑
Architecture change
None
Reported PSNR gain
+10.23 dB
Failure mode
Reward gaming

References