AI Post Transformers · Visualization Companion

MELT: Decoupling Compute From Memory

A systems-first reading of looped reasoning: MELT keeps iterative latent compute but swaps growing per-loop KV state for a single gated shared cache per layer. The visual question is not “can it reason?” but “what happens to memory, state reuse, and practical deployment pressure as loops increase?”

arXiv: 2605.07721 Live Viz URL Transcript IDs 2605.07721 Claim constant-memory looping Tension architecture vs cache engineering
Core Shift
Looped compute
Reasoning passes continue to scale, but cached attention state stops growing with every pass.
Memory Story
1 cache / layer
MELT rewrites shared cache slots through gates instead of appending new per-loop KV slices.
Closest Ancestors
UT + LSTM
Universal Transformer style iteration, recurrent-style keep / overwrite / forget behavior.
Practical Use
KV wall relief
Interesting where long context and deeper reasoning are blocked by memory traffic, not parameter count.

Looped Architecture at a Glance

Same iterative ambition, different memory bookkeeping. Toggle between the growing Ouro-style stack and the MELT shared-cache rewrite.

What the Paper Is Really Selling

The page below is not prose-first. Read the stacked bars as a claim allocation meter: where the contribution appears strongest, inherited, or still under-proven.

strong support inherited prior capability evidence gap

Memory Scaling vs Loop Depth

Mocked but mechanically faithful: baseline looped KV grows with sequence length × loop count, while MELT stays close to flat after initial cache allocation.

Compute-Memory Surface

Hover the heatmap. Hot colors mean harder hardware pressure. MELT changes the memory axis more than the compute axis.

Axes: sequence length on x, loop count on y. Cells approximate deployment pain rather than benchmark score.

Shared Cache Update Mechanics

Each loop rewrites one persistent memory page through gates. Hover cells to inspect keep / overwrite intensity across layers and steps.

Adaptation Recipe

The architecture is only part of the story. Step through chunk-wise training, transition interpolation, and teacher-aligned distillation.

Baseline Landscape

This scatter compresses the episode’s framing: standard decoders, cache-efficient variants, recurrent alternatives, Ouro, and MELT occupy different memory / reasoning trade zones.

Open Questions Radar

The stronger the fill, the more unresolved experimental demand remains before “new reasoning paradigm” becomes a safe conclusion.

References

Compact source map. arXiv links appear when the episode supplied a direct paper identifier.