World-R1 wraps reinforcement learning around an existing text-to-video model and scores the generated clips with pretrained 3D critics plus a vision-language judge. The visual question is whether camera motion reveals a stable scene, or whether the world liquefies when the viewpoint shifts.