AI Post TransformersarXiv:2504.198742025 paperOnline vector quantization
TurboQuant
A visual companion to the episode on data-oblivious online vector compression:
random rotation spreads energy, scalar quantization does the heavy lifting,
and a residual 1-bit correction repairs inner-product bias when retrieval or attention cares more about scores than raw reconstruction.
Online Compression Pipeline
The paper’s argument is structural: mix coordinates with a random rotation, quantize each coordinate simply, and optionally add a tiny residual sketch when the downstream objective is dot products rather than just reconstruction.
From dense vector to score-preserving code
Flow Diagram
Hover boxes for what each stage is optimizing. The second branch exists because low MSE does not guarantee faithful rankings or attention scores.
Bit allocation by objective
Stacked Bars
Mock allocations showing how the same total rate gets split differently when inner-product repair is enabled.
Where TurboQuant sits
Method Map
The page positions TurboQuant between fully learned codebooks and purely per-coordinate scalar quantizers.
Rotation as Geometry Regularizer
Before rotation, energy can be concentrated in a few unstable coordinates. After rotation, the vector looks more balanced coordinate-by-coordinate, which is the condition that makes simple scalar quantization unexpectedly strong in high dimension.
Coordinate covariance heatmap
Heatmap
Synthetic 16×16 covariance slice. The rotated view becomes less blocky and less dominated by a few channels.
low couplingmoderatehigh
Coordinate magnitude profile
Bar Profile
The pre-rotation profile is spiky. The rotated profile is intentionally flatter, which reduces the worst-coordinate problem for fixed-step quantizers.
Quantization bins vs. spread
Step Diagram
A single scalar quantizer works poorly on anisotropic coordinates, then improves once the rotation spreads energy more evenly. Hover points to compare clipping pressure.
Distortion, Recall, and Objective Switching
The transcript’s caution shows up here: the theorem is broad, but the evidence is narrower. These mock curves emphasize the paper’s core tradeoff rather than claim exact benchmark numbers.
Distortion-rate frontier
Line Chart
TurboQuant tracks close to the theoretical floor while staying online and data-oblivious.
ANN recall at equal code size
Bar Chart
A stylized comparison against classical product quantization and other online or scalar-heavy schemes.
Operating regime map
Mode Surface
The residual correction helps most in low-bit score-sensitive regimes. At higher rates, simpler MSE-focused quantization can win on net utility.
Systems Pressure and Practical Limits
KV-cache compression is a serving-time intervention, not a training-time cleanup step. The visuals below focus on bandwidth pressure, salient-token sensitivity, and the missing end-to-end throughput measurements the episode called out.
Decode traffic per token
Flow + Bars
Compression reduces bytes moved more than it reduces arithmetic. That is why the page centers memory flow rather than FLOP counts.
Attention sensitivity heatmap
Matrix
Not all rows or tokens matter equally. Salient regions and sink behavior complicate any purely oblivious compression story.
Production readiness scoreboard
Comparison Grid
The theory is strong, but the empirical column is intentionally mixed: proofs, selective benchmarks, and open systems questions coexist.
References
Compact source list for the visuals, including the main paper, classic retrieval quantization work, and recent KV-cache compression papers discussed in the episode.