Latent Reasoning with Normalizing Flows
A frozen VAE compresses verbose written rationales into a short latent sequence, then shallow normalizing flows let one left-to-right transformer predict continuous thought slots and normal answer tokens inside the same causal stream. The point of the page is the shape of that trade: fewer explicit tokens, exact likelihoods, easier sampling, but thinner auditability.
Matched Signal
Reported averages from the episode transcript. Per-benchmark detail below is illustrative unless explicitly labeled otherwise.
Unified Autoregressive Pipeline
Toggle between the training path that learns from visible rationales and the inference path that removes them. The backbone stays causal in both cases.
Flow Warping
Illustrative density view: a simple base distribution gets bent into a richer latent space while keeping tractable likelihoods.
Token Footprint
Illustrative slot budget, not a quoted paper number. The page shows why replacing long written traces with a few latent positions is attractive.
Why This Formulation Matters
The paper’s pitch is not “hide reasoning.” It is “keep reasoning latent without giving up normal autoregressive serving primitives.”
Step-by-Step Latent Thought Construction
Progress through the path from visible teacher rationale to sampled latent slots and answer tokens. The heatmap and latent plane are illustrative diagrams of how structure changes across stages.
Benchmark Lift
Switch between the reported benchmark average and an illustrative per-benchmark decomposition that preserves the transcript’s averages and HumanEval caveat.
Pass@k Shape
Illustrative MBPP+ curves honoring the transcript’s claim that a latent method’s pass@1 can match the base model’s pass@128.
Gain Heatmap
Per-benchmark delta over the base model using the same illustrative matched split as the grouped bars.
Method Tradeoff Matrix
This grid is a qualitative map of what the episode argued, not a paper table. Warmer cells indicate stronger fit on that axis.
Serve vs Audit
Change the weighting and watch the “best” method move. Deployment prefers compact scoreable latents; oversight still prefers visible traces.
What The Evidence Actually Covers
The discussion ended with a narrow reading: strong post-training signal on one 8B code model, much weaker evidence on transfer, latency, teacher robustness, and latent faithfulness.
References
Compact source list from the episode prompt. arXiv links are shown directly when available; other items keep their provided source URLs.