A neuroscience-inspired fix for the recall gap in efficient Transformer alternatives: instead of one fixed-size memory state, MoM routes tokens across several independent memory slots plus a shared accumulating memory — borrowing Mixture-of-Experts routing and a hippocampus-style multiplexing analogy.
All six recall benchmarks are truncated to 2K tokens — a window a single well-tuned memory state can often handle. The only test past 2K is perplexity extrapolation to 32K, which shows the model doesn't collapse, but never proves task-level retrieval of a fact planted far back in context. Titans, a peer approach to the same interference problem, is listed in the paper's taxonomy but never run as a baseline.