1 00:00:01,000 --> 00:00:41,325 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. So Ada, today we're cracking open a paper called MemGPT: Towards LLMs as Operating Systems. First author is Charles Packer, et al. — seven authors total, including Sarah Wooders, Kevin Lin, Vivian Fang, Shishir Patil, Ion Stoica, and Joseph Gonzalez, all out of UC Berkeley. It first hit arXiv in October 2023, with a revised version posted in February 2024. And the pitch is bold: treat a language model's context window like RAM in a computer. 2 00:00:41,325 --> 00:01:08,525 [Dr. Ada Shannon] Here's what got me: this isn't just another retrieval trick bolted onto a chatbot. The theoretical grounding is what sold me on this one — they're not asking 'how do we cram more tokens into the prompt,' they're asking 'what did operating systems figure out fifty years ago about giving programs the illusion of more memory than they actually have, and can we just steal that wholesale for LLMs?' That's a genuinely different framing, and it comes with a real architecture attached, not just a bigger retrieval index. 3 00:01:08,525 --> 00:01:47,775 [Hal Turing] Right, and to get why that framing matters, you have to understand the problem they're solving. Every transformer has a fixed context window — the maximum number of tokens it can look at in one shot. And you can't just make that window bigger for free, because self-attention cost scales quadratically with sequence length. Double your context, quadruple your compute and memory. That quadratic wall is why Llama 2 tops out around 4k tokens and even GPT-4's original release was capped at 8k — that's maybe 140 messages of actual conversation before you fall off the edge. 4 00:01:47,775 --> 00:02:23,425 [Dr. Ada Shannon] And even if you brute-force past that wall — say with a 100k or 128k window like Claude 2 or GPT-4 Turbo — you run into a second, nastier problem: the model doesn't actually use all that space well. That's the headline finding from the Lost in the Middle paper, Nelson Liu and coauthors out of Stanford, 2023. They showed models are much better at recalling information at the very start or the very end of a long context than anything buried in the middle. So scaling context isn't just expensive — past a certain point it's not even buying you reliable recall— 5 00:02:23,425 --> 00:02:50,625 [Hal Turing] Oh wait wait wait — that's actually the part that made this click for me. Because that's exactly the OS problem, right? A program doesn't need all of RAM loaded at once, it just needs the OS to page in the right data at the right moment and page out what it doesn't need. MemGPT is proposing the LLM equivalent: don't fight the context window, manage it. Treat the window itself as a scarce resource with a paging policy, the same way virtual memory pages data between RAM and disk. 6 00:02:50,625 --> 00:03:20,074 [Dr. Ada Shannon] That's the core move, yes — main context is their word for what's actually in the prompt tokens, analogous to physical RAM, and external context is everything sitting outside the window that has to be explicitly paged in before the model can see it, analogous to disk. But here's where I'd push back a little on how clean that analogy is. A real OS kernel pages memory transparently — the application doesn't decide when a page fault happens. In MemGPT, the LLM itself decides when to page things in and out, by emitting function calls. 7 00:03:20,074 --> 00:03:31,674 [Hal Turing] I mean, isn't that the interesting part though, not a flaw? The model calling functions like a search over its own memory instead of a human-written retrieval pipeline deciding for it? 8 00:03:31,674 --> 00:03:58,749 [Dr. Ada Shannon] I actually disagree with you there, Hal — or at least, I think you're glossing over why that's a genuinely bigger claim than it sounds. Self-directed paging means the system's reliability is now hostage to whether the underlying model calls the right function at the right time. A kernel doesn't have off days. An LLM's function-calling accuracy absolutely does, and we'll see later in this paper that it's a real bottleneck. Calling it 'virtual memory' undersells how much is riding on the model's own judgment. 9 00:03:58,749 --> 00:04:22,500 [Hal Turing] No no no, I hear you on the risk, but I don't think that breaks the analogy — I think it's just a more honest version of it. Every real paging system has a policy for what to evict and when, LRU, clock, whatever. MemGPT's policy just happens to be 'ask the language model.' It's still a memory hierarchy with an eviction policy, it's just implemented differently than in silicon. 10 00:04:22,500 --> 00:04:42,125 [Dr. Ada Shannon] Okay, fair — if we define it at that level of abstraction, a hierarchy with a policy, I'll grant you it holds. I just don't want listeners walking away thinking this is deterministic the way a kernel's page table is. We're getting ahead of ourselves though — let's lay out what's actually inside that hierarchy before we argue about it further. 11 00:04:42,125 --> 00:05:07,024 [Hal Turing] Fair. So inside main context, MemGPT splits the prompt into three pieces. There's working context, a fixed-size read/write scratchpad the model can edit directly — things like key facts about a user or the persona it's playing. Then there's the FIFO queue, which holds the rolling recent conversation history, plus a running recursive summary of anything that's already been evicted out of it. 12 00:05:07,024 --> 00:05:39,049 [Dr. Ada Shannon] And outside the window entirely, you've got two more tiers, both queryable but never automatically visible. Archival storage is a searchable database for arbitrary long-form content — think an uploaded document collection. Recall storage is the searchable log of literally every past message the agent has ever sent or received. Both only enter the model's view when it explicitly calls a search function and pulls results back in — that's the function-calling layer doing the actual paging work we were just debating. 13 00:05:39,049 --> 00:06:05,724 [Hal Turing] Which sits this paper at the intersection of a handful of threads that were all bubbling up around the same time — agent memory systems, this new idea of a 'memory operating system,' the whole context-window-scaling arms race we just talked about, tool-using LLM agents following work like Toolformer, and long-context models generally. MemGPT isn't really competing with any one of those lines, it's stitching them together into one architecture. 14 00:06:05,724 --> 00:06:20,924 [Dr. Ada Shannon] Which is exactly why I wanted to cover it. Next up, we get into how this thing actually runs — the queue manager, the function executor, and some genuinely striking numbers on how much this improves memory recall over just using a bigger model. 15 00:06:20,924 --> 00:06:49,324 [Dr. Ada Shannon] So let's get into the plumbing. There's a component called the queue manager — it appends incoming messages to the FIFO queue, concatenates everything into the prompt, fires the inference call, and writes both the user message and the model's output to recall storage. The interesting part is what happens as that queue fills up. Once prompt tokens cross a warning threshold — 70 percent of the context window — the queue manager doesn't wait for a crash. It injects a system message: a 'memory pressure' alert, right into the conversation the model is seeing. 16 00:06:49,324 --> 00:06:56,550 [Hal Turing] Wait — so the model gets warned about its own memory running out, like a phone buzzing at 20 percent battery? 17 00:06:56,550 --> 00:07:24,050 [Dr. Ada Shannon] Basically, yes. That warning is what triggers the model to proactively call functions — save something to working context, archive it — before eviction happens. If it ignores the warning and hits 100 percent, the queue manager forces the issue: it flushes a chunk of messages and generates a new recursive summary folding in whatever was just evicted plus the prior summary. That summary sits at the front of the queue as the only trace of what's gone. The evicted messages aren't destroyed — they live in recall storage, searchable later. 18 00:07:24,050 --> 00:07:32,600 [Hal Turing] And the function executor is what actually runs those calls the model makes, right? How does that fit with the event-driven piece? 19 00:07:32,600 --> 00:08:03,700 [Dr. Ada Shannon] Right — every function call gets parsed, validated, executed, and the result — including errors, like trying to write to a full context — gets fed straight back to the model as its next input. Here's the clever bit for chaining: functions can carry a flag, request_heartbeat, that hands control back to the processor immediately instead of waiting for the next user message. That's what lets MemGPT do multi-step retrieval — page through results, check a value, decide it needs another lookup — all before ever yielding a turn back to the human. 20 00:08:03,700 --> 00:08:08,300 [Hal Turing] Okay, give me the numbers, because you teased these were striking. 21 00:08:08,300 --> 00:08:42,000 [Dr. Ada Shannon] Deep memory retrieval — the agent's asked something answerable only from a session five conversations back. GPT-3.5 Turbo alone: 38.7 percent. Add MemGPT: 66.9. GPT-4 alone: 32.1 — lower than 3.5, interestingly. Add MemGPT: 92.5. GPT-4 Turbo: 35.3 baseline, 93.4 with MemGPT. That's not incremental, that's nearly tripling GPT-4's accuracy on a task purely about remembering something from outside the context window. 22 00:08:42,000 --> 00:08:56,400 [Hal Turing] GPT-4 alone doing worse than 3.5 on that is wild. But what about the softer metric — the conversation opener task? That one's not just 'did you remember,' it's 'did you say something engaging.' 23 00:08:56,400 --> 00:09:27,325 [Dr. Ada Shannon] Right, and this is the one I find almost more interesting, because MemGPT is measured against a human-written opener as the gold standard, using similarity scores — SIM-1, SIM-3, SIM-H. Humans score 0.800 on SIM-1 and SIM-3 by definition. MemGPT with GPT-4 hits 0.868 and 0.843 — it actually beats the human baseline on those. The paper's own explanation is that MemGPT's openers are more verbose and touch more persona facts than what the human labelers wrote. 24 00:09:27,325 --> 00:09:38,325 [Hal Turing] Hold on, isn't 'beats the human' doing a lot of work there? If it's just cramming in more persona facts, that's not necessarily a better opener, that's a wordier one. 25 00:09:38,325 --> 00:10:06,400 [Dr. Ada Shannon] I'd push back on that a little. SIM-1 and SIM-3 measure overlap with the persona traits, not word count — an opener surfacing more true facts about the user is, by that metric, doing exactly what the task asks. Now, SIM-H — similarity to the actual human opener — MemGPT with GPT-4 only hits 0.773, well below the 1.0 a human gets against itself. So it's not fooling anyone into thinking it wrote the human's exact line. It's optimizing the persona-coverage metric harder than a human labeler bothered to. 26 00:10:06,400 --> 00:10:15,575 [Hal Turing] Fair, I'll take that. So separate from conversation stuff entirely — what about the document QA side, actual long documents? 27 00:10:15,575 --> 00:10:54,650 [Dr. Ada Shannon] That one's a retriever-reader setup on NaturalQuestions-Open — embed a Wikipedia dump, a retriever pulls the top-K passages, and the model answers using only what's retrieved. They sampled 50 questions. Baselines and MemGPT share the same retriever, but baselines get the top-K documents handed directly into their prompt. MemGPT gets archival storage instead and searches it itself, paginating through results. There's a chart, Figure 5, plotting accuracy against documents retrieved — the fixed-context baselines rise and plateau with the retriever's ranking quality, while MemGPT stays essentially flat, since it isn't capped by how many documents fit in one prompt. 28 00:10:54,650 --> 00:11:02,950 [Hal Turing] And then there's the nested key-value task, which sounds like the cleanest stress test of the multi-hop retrieval story. 29 00:11:02,950 --> 00:11:42,625 [Dr. Ada Shannon] It's synthetic on purpose — 140 UUID key-value pairs, about 8,000 tokens, nesting depth from zero to four hops, where the value you retrieve can itself be another key. GPT-3.5 falls to zero accuracy at just one nesting level. Plain GPT-4 collapses to zero by two levels, GPT-4 Turbo holds on a little longer before collapsing by three. MemGPT with GPT-4 stays flat across all four — it's not holding four chained lookups in its head, it's just re-querying archival storage node by node, checking whether the value it just got is itself a key. Which raises a question the paper doesn't really want asked too loudly. 30 00:11:42,625 --> 00:11:45,100 [Hal Turing] Go ahead, ask it loudly. 31 00:11:45,100 --> 00:12:05,700 [Dr. Ada Shannon] Is that multi-hop reasoning, or just a for-loop? A hundred forty UUID pairs, eight thousand tokens, entirely authored and controlled by Packer's team — no ambiguity, no paraphrase, no real-world noise. It's a solid demo of function chaining working mechanically. It's a much weaker demo of reasoning that generalizes to messy real documents. 32 00:12:05,700 --> 00:12:22,650 [Hal Turing] Same problem shows up in the document QA numbers, Figure 5 — fifty sampled questions. Fifty. No confidence intervals, no repeated seeds, and we're meant to read real curve shape into that as retrieved documents go from zero to two hundred. 33 00:12:22,650 --> 00:12:45,425 [Dr. Ada Shannon] Thin for a curve, agreed. And there's a second wrinkle beside it — both DMR and document QA use GPT-4 as judge, and the paper admits MemGPT's answers run 'generally more verbose' than the gold answers. Zheng et al.'s MT-Bench work documents exactly this: LLM judges reward length independent of correctness. Some slice of that 32-to-92.5 jump on GPT-4 DMR could just be MemGPT talking more. 34 00:12:45,425 --> 00:13:07,900 [Hal Turing] Wait, hold on — isn't the real problem upstream of that? The paper flat-out says GPT-3.5's 'degraded performance' on DMR comes from its 'limited function calling.' So the whole architecture is riding on how cleanly the base model emits function calls. Isn't this a GPT-4-tool-use result wearing an OS-paging costume? 35 00:13:07,900 --> 00:13:30,175 [Dr. Ada Shannon] I actually disagree with you there, Hal. That's like calling RAID 'just a result of disk controllers working' — sure, the substrate has to function, but the architecture still does real work on top. GPT-4 alone already has decent function calling and still collapses to zero accuracy by nesting level three. MemGPT with GPT-4 stays flat through all four. Function-calling quality alone doesn't explain that gap. 36 00:13:30,175 --> 00:13:47,150 [Hal Turing] Sure, but MemGPT-with-GPT-3.5 versus MemGPT-with-GPT-4 is 66.9 versus 92.5 — if the architecture were doing the heavy lifting, shouldn't the underlying model matter less, not more? 37 00:13:47,150 --> 00:14:17,975 [Dr. Ada Shannon] Fair — it amplifies whatever tool-use quality it's given, it doesn't replace it. Real architectural value, bottlenecked by reliability, not either-or. Which loops back to Liu et al.'s 'Lost in the Middle,' 2023, Stanford — the paper MemGPT's own QA and KV benchmarks are built directly on. Liu showed models underuse the middle of long contexts; MemGPT's fix is page it out and re-fetch on demand. Legitimate — but it also means this win partly sidesteps a known attention pathology, not proves general long-term memory. 38 00:14:17,975 --> 00:14:46,175 [Hal Turing] Which connects to Park et al.'s Generative Agents, Stanford, 2023 — a completely different philosophy. Instead of OS paging and function calls, their agents build a reflection tree, synthesizing raw observations into higher-level reflections and retrieving by recency, importance, and relevance. MemGPT treats memory like a filesystem with syscalls; Park's agents treat it like a diary they periodically reread. Nobody's benchmarked one against the other. 39 00:14:46,175 --> 00:15:08,500 [Dr. Ada Shannon] And the MSC comparison has its own fairness gap — the baseline gets a lossy summary of prior sessions, MemGPT gets full paginated access to the entire history. Not apples-to-apples; MemGPT structurally has a bigger memory budget by design. Also nobody reports latency or call counts — multi-step paging could mean five to ten times the inference cost of a single pass, never measured. 40 00:15:08,500 --> 00:15:30,450 [Hal Turing] There's a privacy angle too nobody flags — Figure 1's whole running example is storing an ex-boyfriend's name and birthday indefinitely in recall storage. Cute demo, zero discussion of retention limits, consent, or what happens when a stale memory gets pulled back and just overwrites something true. No rollback, no verification. 41 00:15:30,450 --> 00:15:49,600 [Dr. Ada Shannon] And it's GPT-3.5, GPT-4, GPT-4 Turbo only — closed OpenAI models. Self-directed editing depends on precise function-call formatting, so whether this holds up on weaker open-weight models, or only exists behind an expensive frontier API, is still an open question. 42 00:15:49,600 --> 00:15:54,800 [Hal Turing] So practically — building agent memory today, what do you take from this? 43 00:15:54,800 --> 00:16:25,375 [Dr. Ada Shannon] A proof of concept for the paging metaphor, not a finished product. The paper's own future-work list says as much: other unbounded-context domains, real database and cache tiers instead of one vector store, better control flow and eviction policy — right now it's all LLM judgment, zero verification. Given how much context windows have grown since 2023, the more interesting question ahead isn't windows versus paging, it's whether you need either once retrieval is good enough. 44 00:16:25,375 --> 00:16:51,675 [Hal Turing] Good note to land on. MemGPT's OS framing is a genuinely useful lens for agent memory — the DMR and KV numbers are real. But the evaluation leans on small samples, an author-controlled synthetic task, a verbosity-prone judge, and a structurally generous baseline. Treat the headline percentages as promising, not settled. Thanks for listening to AI Post Transformers — we'll catch you next time.