AI Post Transformers · Visual Companion

AllMem for Efficient Long-Context Modeling

A visual map of the AllMem bet: keep exact attention over a recent token window, compress the distant past into learned online memory, and trade quadratic long-context cost for a hybrid path that still preserves sharp local reasoning.

arXiv 2602.13680 Qwen3 0.6B and 1.7B 4k and 8k recent windows Long-context distillation On-policy post-training
Paper
Ziming Wang, Xiang Wang, Kailong Peng, Lang Qin, Juan Gabriel Kostelec, Christos Sourmpis, Axel Laborieux, Qinghai Guo. Posted February 14, 2026.
Core Design
Recent context stays as exact KV. Evicted context becomes a learned memory state updated online as new tokens arrive.
Bench Anchors
Transcript claims near-lossless 4k-window behavior on LongBench up to 37k and an 8k-window variant beating full attention on InfiniteBench at 128k.
Transcript ID Scan
Scanning transcript for additional arXiv IDs.
Tab 1

Hybrid Map

Tab 2

Attention Plane

Tab 3

Bench Curves

Tab 4

Design Space

References

compact arXiv-linked paper set

Related Episodes

public episode cross-links