Interactive Visual Companion

Latent Reasoning with Normalizing Flows

A frozen VAE compresses verbose written rationales into a short latent sequence, then shallow normalizing flows let one left-to-right transformer predict continuous thought slots and normal answer tokens inside the same causal stream. The point of the page is the shape of that trade: fewer explicit tokens, exact likelihoods, easier sampling, but thinner auditability.

arXiv 2606.06447 Backbone Qwen3-8B-Base Posted June 4, 2026 Claim latent slots keep exact likelihoods Transcript scan checking IDs…

Matched Signal

Reported averages from the episode transcript. Per-benchmark detail below is illustrative unless explicitly labeled otherwise.

Base
55.8
Diffusion
61.6
NF Unified
68.8
NF + RL
70.1
Training-only compression path
VAE out at inference
Latent path advantage emphasized in the episode
sampleable + scoreable
Main caveat repeated in the discussion
strong on code, narrow elsewhere

Unified Autoregressive Pipeline

Toggle between the training path that learns from visible rationales and the inference path that removes them. The backbone stays causal in both cases.

Flow Warping

Illustrative density view: a simple base distribution gets bent into a richer latent space while keeping tractable likelihoods.

Token Footprint

Illustrative slot budget, not a quoted paper number. The page shows why replacing long written traces with a few latent positions is attractive.

Example budget: prompt 12, visible rationale 48, answer 24 versus prompt 12, begin-thought 1, latent slots 4, answer 24.

Why This Formulation Matters

The paper’s pitch is not “hide reasoning.” It is “keep reasoning latent without giving up normal autoregressive serving primitives.”

Latent thought slots are still sampled left-to-right. Exact conditional densities make RL on latent trajectories easier to define. Auditability remains weaker than explicit text CoT.

Step-by-Step Latent Thought Construction

Progress through the path from visible teacher rationale to sampled latent slots and answer tokens. The heatmap and latent plane are illustrative diagrams of how structure changes across stages.

Benchmark Lift

Switch between the reported benchmark average and an illustrative per-benchmark decomposition that preserves the transcript’s averages and HumanEval caveat.

Pass@k Shape

Illustrative MBPP+ curves honoring the transcript’s claim that a latent method’s pass@1 can match the base model’s pass@128.

Gain Heatmap

Per-benchmark delta over the base model using the same illustrative matched split as the grouped bars.

Method Tradeoff Matrix

This grid is a qualitative map of what the episode argued, not a paper table. Warmer cells indicate stronger fit on that axis.

Serve vs Audit

Change the weighting and watch the “best” method move. Deployment prefers compact scoreable latents; oversight still prefers visible traces.

What The Evidence Actually Covers

The discussion ended with a narrow reading: strong post-training signal on one 8B code model, much weaker evidence on transfer, latency, teacher robustness, and latent faithfulness.

References

Compact source list from the episode prompt. arXiv links are shown directly when available; other items keep their provided source URLs.