AI Post Transformers • Visual Companion

Training Million-Token LLMs Beyond the Memory Barrier

This page treats the paper as a systems-engineering claim under a microscope: activations can be flattened with replay, but the KV path keeps growing, so the real drama moves from autograd storage to page movement, residency, and bandwidth.

Paper: arXiv:2602.02108 ICLR 2026 Model: Qwen2.5-7B Claim: 4M-token training on one H200 Transcript arXiv IDs found: 2602.02108
Headline Slope
~10 MB
Per extra 10K tokens, as discussed in the episode
Core Trick
Replay
Drop activations, recompute during backward
Remaining Wall
KV
Cache residency and transfer scheduling still grow
Interpretation
Systems
Strong feasibility result, not automatic proof of deep long-range learning

Where the Bottleneck Moves

A normal long-context training run explodes because activation storage scales with sequence length. OOMB’s argument is not “memory disappears”; it is “activations stop dominating, so KV state becomes the hard part.”

Chunk-Recurrent Training, Step by Step

This section animates the execution model. The model walks through the giant sequence chunk by chunk, discards per-chunk activations after the forward pass, then replays those chunks during backward so gradients can still be computed.

Paged KV Management Is the Real Second Battle

Once activations are flattened, the surviving growth term is KV state. Here the issue is not only capacity, but page movement and access timing: can the GPU stay busy while host memory serves as overflow?

Feasibility Result vs Product Claim

The most defensible reading is narrow: the memory wall can be bent sharply with recomputation, paging, offload, and optional sparsity. The stronger claim, that this alone proves robust million-token learning, remains open.

References

Compact map of the paper family and adjacent prior episodes mentioned in the discussion.

Out of the Memory Barrier
Li et al., 2026 • arXiv:2602.02108
Transformer-XL
Dai et al., 2019 • recurrent context across segments
Recurrent Memory Transformer
Bulatov et al., 2022 • explicit memory tokens
Ring Attention
Liu, Zaharia, Abbeel, 2023 • exact long attention across devices
Infini-attention
Munkhdalai et al., 2024 • long-context compression path
PagedAttention
Kwon et al., 2023 • serving-side paging ancestor
DeepSeek-V4 and Practical Million-Token Context
AI Post Transformers • prior context-window episode
CacheFlow and 3D-Parallel KV Cache Restoration
AI Post Transformers • KV movement and restoration