1 00:00:01,000 --> 00:00:55,288 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called MoM: Linear Sequence Modeling with Mixture-of-Memories. That's Jusen Du plus four co-authors — Weigao Sun, Disen Lan, Jiaxi Hu, and Yu Cheng — five authors total, out of Shanghai AI Laboratory, working with Tsinghua University, Fudan University, Hong Kong University of Science and Technology in Guangzhou, and the Chinese University of Hong Kong. It landed on arXiv back in February 2025, number 2502.13685. And here's what hooked me, Ada — the whole premise is that giving a model more than one memory at once might fix something that's been quietly bottlenecking every efficient alternative to the Transformer we've ever covered. 2 00:00:55,288 --> 00:01:27,842 [Dr. Ada Shannon] It's a real bottleneck, not a hypothetical one. The question this paper is asking is: can you take these linear sequence models — built to be cheap and fast — and give them enough memory capacity, arranged the right way, that they stop losing information they should be able to recall? Right now they can write fluent text and still forget a fact you mentioned six paragraphs ago. That's the gap. And the fix borrows an idea from a completely different corner of the deep learning toolbox and grafts it onto memory itself, which is a more interesting move than it sounds. 3 00:01:27,842 --> 00:01:43,957 [Hal Turing] Back up for a second, because I want everyone tracking what 'forgets a fact from six paragraphs ago' actually means mechanically, not just as a vibe. When we say a Transformer remembers something and a linear model forgets it, what's actually different under the hood? 4 00:01:43,957 --> 00:02:39,452 [Dr. Ada Shannon] It's architecture, squarely. A Transformer keeps a key-value pair for every token it's ever seen — the KV cache — so nothing gets overwritten. A linear sequence model, going back to the 'Transformers are RNNs' paper by Katharopoulos, Vyas, Pappas, and Fleuret in 2020, does something different: it swaps softmax attention for a kernel trick that lets the whole thing be rewritten as a recurrence with one fixed-size memory state — a single matrix updated at every step instead of a cache that keeps growing. That's what gives linear-time training and constant-memory inference. But because it's one fixed-size slot, every new write can partially overwrite what's already there. The paper names two failure modes directly: limited capacity — that one slot just can't hold enough — and memory interference, where new information degrades old information sitting in that same state. 5 00:02:39,452 --> 00:03:01,465 [Hal Turing] So the Transformer's trick is basically never throw anything away, and the price is attention cost that blows up quadratically with sequence length. The linear models flip that trade completely — cheap and constant per step, but lossy. Is that the entire reason this whole line of research exists — dodging that quadratic wall? 6 00:03:01,465 --> 00:03:39,035 [Dr. Ada Shannon] That's exactly it. O(n squared) is fine at a few thousand tokens, brutal at a few hundred thousand — that's why the field went looking for alternatives: state space models like Mamba, linear attention, linear RNNs like RWKV. They all made the same bet, give up some recall precision, get linear scaling back. For plain language-modeling perplexity, that bet mostly paid off. It's specifically recall-heavy tasks — needle-in-a-haystack, multi-hop QA, pulling one exact detail from way back — where the single fixed-size state shows its limits. That's the precise gap MoM is aimed at. 7 00:03:39,035 --> 00:03:47,533 [Hal Turing] So where does the actual fix come from? I saw a mention of neuroscience buried in here, which is not what I expected walking into a memory-architecture paper. 8 00:03:47,533 --> 00:03:51,481 [Dr. Ada Shannon] Right, they lean on a genuinely biological analogy. The hippocamp— 9 00:03:51,481 --> 00:03:59,840 [Hal Turing] Oh wait, wait, wait — hold on, hold on. You're telling me there's actual brain science backing this and it's not just a cute metaphor slapped on the abstract? 10 00:03:59,840 --> 00:04:40,057 [Dr. Ada Shannon] Both, honestly. The hippocampus separates simultaneous memories using theta and gamma oscillations — different phases of brain-wave activity acting like separate channels, so two memories forming close together in time don't blend into mush. Neuroscientist David Eagleman has made the broader point that the brain isn't running one memory trace at a time, it's multiplexing — keeping distinct items distinct by giving each one its own slot in time. The MoM authors use that framing directly: instead of one memory state everything gets crammed into, why not several states, computationally separated, the way the hippocampus separates memories? 11 00:04:40,057 --> 00:04:54,314 [Hal Turing] That's a satisfying parallel, and it also sounds a lot like something else — routing inputs to different specialized components instead of forcing everything through one path. Is that literally Mixture-of-Experts, just relocated? 12 00:04:54,314 --> 00:05:50,181 [Dr. Ada Shannon] Same underlying idea, redirected. Mixture-of-Experts — from Shazeer, Mirhoseini, and colleagues out of Google Brain in 2017 — scales a model's capacity by routing each input to a small subset of specialized sub-networks instead of running everything through one dense block, so you get more total capacity without paying for it on every token. MoM points that exact routing philosophy at memory instead of feed-forward layers. So the shape of the fix: instead of one memory state, maintain several independent ones. A lightweight router decides which memory or memories each incoming token gets written to. There's also a shared memory, always updated alongside the routed ones, so you keep a continuously accumulating global summary too. It's still built on the same linear-recurrence machinery as state space models like Mamba — Gu and Dao's 2023 paper — just multiplied and reorganized. 13 00:05:50,181 --> 00:06:08,200 [Hal Turing] Okay — one memory slot, two failure modes, capacity and interference, and a fix borrowed from Mixture-of-Experts wearing a hippocampus metaphor. What I want to know now is how that router actually decides where a token goes, because 'a lightweight router decides' is doing a lot of work in that sentence. 14 00:06:08,200 --> 00:06:23,850 [Dr. Ada Shannon] It is doing a lot of work, and it's worth being precise about — the routing mechanics are exactly where this either holds together as a real architectural contribution or turns out to be a strong baseline with a good metaphor attached. Let's get into how it actually works. 15 00:06:23,850 --> 00:07:05,182 [Dr. Ada Shannon] Mechanically it's almost disarmingly simple. Every token passes through one small linear layer — just a learned matrix, x times W_g — producing a raw score for each of the memory states. Softmax turns those into a distribution, then you take the top-k, say two out of four states, and renormalize just those k scores so they sum back to one. No separate gating network, no extra hidden layers — one matmul, one softmax, one top-k, one renormalization. The k selected memories get this token's information; the rest sit completely untouched, no decay, no bleed, no interference from tokens that were never routed there. 16 00:07:05,182 --> 00:07:22,132 [Hal Turing] So the decision itself is cheap. Once a token lands in, say, memory two and memory four, what's actually happening inside those states — is it literally the update math we covered with the matrix state and KV projections, and how does that ever turn back into one output vector the next layer can use? 17 00:07:22,132 --> 00:08:13,495 [Dr. Ada Shannon] Same math, duplicated per memory — each activated state gets its own key and value projection, and updates via that outer-product accumulation we already described. But MoM doesn't commit to one update rule; it's explicitly update-rule-agnostic, a wrapper you can drop any linear recurrence into. In their experiments that's Gated DeltaNet — Yang, Kautz and Hatamizadeh out of NVIDIA, 2024 — as the backbone inside every memory slot. Recombining is a weighted sum using those same router scores: multiply the query against each activated memory, then combine the outputs by importance weight. And there's a continuously-active shared memory layered on top, no routing at all, there for every single token, holding global context that shouldn't depend on which slot a token happened to land in. 18 00:08:13,495 --> 00:08:34,114 [Hal Turing] Oh — hold on, wait. Sorting tokens into different memories, running separate little recurrences per memory, then gathering everything back together at the end — that's exactly the scatter-gather pattern that tanks GPU throughput in practice, even when the underlying math is clean on paper. How do they not eat that cost at runtime? 19 00:08:34,114 --> 00:09:13,077 [Dr. Ada Shannon] That's the real engineering problem they had to solve. Instead of processing tokens in arrival order, they physically reorder the sequence by routing assignment first — every memory-one token grouped together, memory-two together, and so on — concatenated into one variable-length sequence. Then they run the existing, already-optimized Triton kernels from prior linear-attention work on that reorganized sequence, one contiguous segment per memory. Outputs get split back out per memory and restored to original token order before the weighted mixing happens. So routing becomes a sequence-reordering problem they solve once, rather than a new class of GPU kernel they'd have to write from scratch. 20 00:09:13,077 --> 00:09:29,796 [Hal Turing] That's a genuinely satisfying trick — reduce a new engineering problem to one that's already solved instead of reinventing kernels from scratch. So does it actually pay off on the benchmarks, Ada, or is this an elegant idea that nets out to a rounding error once you actually run the numbers? 21 00:09:29,796 --> 00:10:13,078 [Dr. Ada Shannon] Real margin, not a rounding error. On the six recall benchmarks — FDA, SWDE, SQuAD, Natural Questions, TriviaQA, Drop, all truncated to 2K tokens — Gated DeltaNet alone averages 24.78 at 380 million parameters and 32.30 at 1.3 billion. MoM, same recurrence, wrapped in routing, hits 28.16 and 36.04. RetNet, GLA, HGRN2, GSA are all down in the 17-to-22 range, well behind even the Gated DeltaNet baseline. On LongBench — summarization, few-shot, synthetic, code — MoM averages 15.64 against Gated DeltaNet's 13.98. 22 00:10:13,078 --> 00:10:36,112 [Hal Turing] Okay, but couldn't that just be more effective capacity spread across more memory slots doing the work, rather than the separation itself? Before I buy that routing is the mechanism, I want to see it isolated against a model with the exact same activated parameter budget but one bigger memory instead of several small ones. 23 00:10:36,112 --> 00:11:50,462 [Dr. Ada Shannon] That's exactly the control they ran. They took a single memory and expanded its dimensionality to match MoM's total activated capacity — same size, no routing — and separate memories still won: 41.97 average on common-sense tasks versus 41.32, 28.16 versus 26.32 on recall. Separation itself is doing work, not just parameter count. On efficiency, latency and GPU memory both scale linearly against Transformer++ out to half a million tokens, where Transformer++ hits out-of-memory. Trained at 2K context and extrapolated to 32K, MoM's perplexity held up better than every other linear baseline. And routing real ARC-easy tokens through a trained model showed genuine specialization — one memory leaning toward basic nouns and verbs, another toward proper nouns and scientific terms, another toward technical adjectives, another toward fragmented phrases — with an auxiliary load-balancing loss keeping activation roughly uniform across memories so none gets starved. Scaling memory count from one to eight at a fixed 0.5 activation ratio, performance climbs the whole way up. 24 00:11:50,462 --> 00:12:30,818 [Hal Turing] Okay, I'm sold that separation itself is pulling real weight, not just extra capacity in disguise. But now I want to push on something that's been nagging me since you first framed this whole paper around long-context recall. Every single one of those six benchmarks — FDA, SWDE, SQuAD, NQ, TriviaQA, Drop — gets truncated to 2K tokens before evaluation. That's not long context, Ada. That's short enough that a single well-tuned memory state should already have plenty of room to hold everything. If the whole motivation is 'linear models forget things from way back,' shouldn't the headline evidence actually test 'way back'? 25 00:12:30,818 --> 00:13:18,651 [Dr. Ada Shannon] That's the sharpest question you've asked all episode, and no, it doesn't fully hold up. The recall suite itself comes from Arora, Eyuboglu, and Zhang's Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff paper — Stanford and Together AI, 2024 — and MoM just inherits that 2K-truncated protocol wholesale. The only evidence that stretches past 2K is Figure 4, and that's perplexity extrapolation out to 32K on the Fineweb dataset, not task-level recall accuracy. Perplexity climbing gracefully tells you the model isn't collapsing at long range. It doesn't tell you whether it can actually retrieve a fact planted at token twenty thousand. Those are different claims, and this paper only demonstrates the weaker one. 26 00:13:18,651 --> 00:13:44,565 [Hal Turing] Wait — hold on, that's actually a bigger problem than it sounds, because if the recall numbers all sit inside a window a single memory can comfortably handle, how much of MoM's 28.16 average is routing doing real interference-reduction work, versus just inheriting a stronger base recurrence? Gated DeltaNet alone already beat every other single-memory baseline in that table. 27 00:13:44,565 --> 00:14:39,271 [Dr. Ada Shannon] Exactly the conflation to watch for. Gated DeltaNet — Yang, Kautz, and Hatamizadeh's Gated Delta Networks: Improving Mamba2 with Delta Rule, NVIDIA, 2024 — scored 24.78 average recall on its own, miles ahead of RetNet, GLA, HGRN2, GSA sitting in the 17-to-22 range. MoM's headline 28.16 rides on that same strong backbone, so part of the gap to weaker baselines is just 'built on the best update rule available,' not routing. The paper does control for this, buried in Appendix G.1: Table 9 matches activated parameters at 400M and compares MoM directly against Gated DeltaNet — 26.51 versus 24.78. That's a real, routing-attributable lift, about 1.7 points, but noticeably smaller than the 3.4-point headline gap implies. The fair number is the quieter one. 28 00:14:39,271 --> 00:15:03,327 [Hal Turing] So there's a genuine effect, it's just more modest once you isolate it properly — that's honest science, even if it undercuts the marketing framing a bit. What about a completely different fix for the same interference problem? That update-rule taxonomy included Titans, learning to memorize at test time instead of routing to parallel states. Did they ever actually run it as a baseline? 29 00:15:03,327 --> 00:16:01,655 [Dr. Ada Shannon] Never. Titans — Behrouz, Zhong, and Mirrokni out of Google Research, 2024 — takes a genuinely different philosophy: instead of routing tokens to separate parallel states, it updates memory at test time using a gradient-based 'surprise' signal, learning on the fly which inputs deserve to overwrite what's already stored. It's listed in their own taxonomy table as a peer approach to this exact interference problem, but MoM never runs it experimentally. So whether routing to multiple states or test-time gradient updates is the more effective fix stays an open question. Practically, MoM's real strength is reusing existing linear-attention Triton kernels through that token-reordering trick — a plausible drop-in for anyone already running a linear-attention stack who wants a recall boost without redesigning serving. And routing at the token-mixing level rather than channel-mixing, like standard MoE does, is a genuinely distinct design point worth remembering on its own. 30 00:16:01,655 --> 00:17:10,340 [Hal Turing] So where does that leave the core question — can separate, routed memories close the recall gap without giving up linear-time efficiency? Partially, and the paper's more useful for what it opens up than what it fully proves. Worth noting: Yu Cheng, one of the corresponding authors here, also worked on the Memory Intelligence Agent paper, and Jiaxi Hu shows up on the Kimi K2.5 visual agentic intelligence work — this memory-and-routing thread clearly keeps resurfacing in their research. There's a real, controlled routing benefit at matched parameters, genuine efficiency wins over Transformers, and an implementation that reuses existing kernels. But the long-context claim motivating the whole paper still rests on perplexity extrapolation, not proven task-level retrieval past 2K, and Titans remains untested. As Eagleman put it, the enemy of memory is other memories — MoM gives linear models a clever way to keep those memories from fighting. Thanks for listening, everyone — catch you next time.