Training Million-Token LLMs Beyond the Memory Barrier
This page treats the paper as a systems-engineering claim under a microscope: activations can be flattened with replay, but the KV path keeps growing, so the real drama moves from autograd storage to page movement, residency, and bandwidth.
Paper: arXiv:2602.02108ICLR 2026Model: Qwen2.5-7BClaim: 4M-token training on one H200Transcript arXiv IDs found: 2602.02108
Headline Slope
~10 MB
Per extra 10K tokens, as discussed in the episode
Core Trick
Replay
Drop activations, recompute during backward
Remaining Wall
KV
Cache residency and transfer scheduling still grow
Interpretation
Systems
Strong feasibility result, not automatic proof of deep long-range learning
Where the Bottleneck Moves
A normal long-context training run explodes because activation storage scales with sequence length. OOMB’s argument is not “memory disappears”; it is “activations stop dominating, so KV state becomes the hard part.”
Chunk-Recurrent Training, Step by Step
This section animates the execution model. The model walks through the giant sequence chunk by chunk, discards per-chunk activations after the forward pass, then replays those chunks during backward so gradients can still be computed.
Paged KV Management Is the Real Second Battle
Once activations are flattened, the surviving growth term is KV state. Here the issue is not only capacity, but page movement and access timing: can the GPU stay busy while host memory serves as overflow?
Feasibility Result vs Product Claim
The most defensible reading is narrow: the memory wall can be bent sharply with recomputation, paging, offload, and optional sparsity. The stronger claim, that this alone proves robust million-token learning, remains open.
References
Compact map of the paper family and adjacent prior episodes mentioned in the discussion.