A visual-first map of the paper’s claim: move transport into training, let inference collapse into a single forward pass, and treat equilibrium as the point where generated samples no longer need to drift toward the data distribution.
Paper Deng, Li, Li, Du, He · 2026
Core idea 1-NFE generation via training-time drift reduction
Reported FID 1.54 latent · 1.61 pixel on ImageNet 256
Instead of paying for many corrective steps during sampling, the model tries to internalize that correction during optimization. The diagram below tracks where the “motion” lives in drifting versus diffusion-like inference.
The left lane is the one-shot inference path. The right lane is the training loop where residual motion is measured and driven toward zero.
Field Mechanics
The paper frames training as shrinking a field that would otherwise keep nudging generated samples. Toggle between an early, misaligned state and a near-equilibrium state.
The vector field is illustrative: stronger arrows mean samples still need to move. Hover a point or cell to inspect local pressure.
A kernel-style interaction view. Hot cells indicate stronger positive/negative influence between real and generated mini-batch samples.
Fast Sampling vs Quality Correction
The chart contrasts one-step and multi-step families. Values mix paper-reported numbers with labeled mock baselines to emphasize the tradeoff shape rather than claim an exact benchmark table.
Lower is better for FID and latency. The drifting bars show the transcript-reported ImageNet-256 results for latent and pixel generation.
Method Landscape
Each family lands in a different part of the latency-quality-design space. Hover a point to see what it buys and what it gives up.
The diagonal tension is the whole story: low-latency systems want single-pass generation, but iterative systems keep quality and controllability advantages.
References
Compact source map for the companion page. Extracted arXiv IDs from the provided material: 2602.04770.