The Memory Hierarchy: RAM vs. Disk, for LLMs
MemGPT splits the world into main context (what's actually in the prompt — analogous to physical RAM) and external context (everything else, paged in on demand — analogous to disk). Click any block below.
System Architecture
Why Context Windows Are Scarce
Even Big Windows Don't Help: Lost in the Middle
Liu et al. (2023) showed recall accuracy is highest for facts near the start or end of context — and sags badly in the middle. The longer the window, the worse the sag. This is the second reason MemGPT pages instead of just stuffing more tokens in.
Recall Accuracy by Position in Context
The Queue Manager: Paging in Action
As the FIFO queue fills with conversation history, the queue manager monitors token usage against two thresholds — a soft warning and a hard eviction point. Step through what happens.
Walk Through the Eviction Cycle
Context Window Fill Meter
Benchmark Results
Four evaluations from the paper: deep memory retrieval, nested key-value lookups across nesting depth, document QA at scale, and conversation-opener quality.
Deep Memory Retrieval (DMR) Accuracy
Nested Key-Value Lookup vs. Nesting Depth
Document QA vs. Documents Retrieved (Fig. 5)
Conversation Opener: Similarity to Human Baseline
Does the OS Analogy Actually Hold?
Hal and Ada spend most of the episode arguing this exact point. A real kernel pages memory transparently; MemGPT's "kernel" is the language model itself, deciding via function calls.
Kernel vs. MemGPT, Dimension by Dimension
| Dimension | Real OS Kernel | MemGPT |
|---|---|---|
| Who decides eviction? | Kernel Deterministic page-replacement algorithm | The LLM Emits its own function calls to save/archive |
| Reliability | Fixed A kernel has no "off days" | Model-dependent Hostage to function-call accuracy |
| Eviction policy | LRU / clock Well-defined algorithm | "Ask the model" Still a hierarchy + policy, just implemented differently |
| Failure mode | Page fault Handled transparently, app unaware | Missed call Wrong/late function call = lost context |
| Verification | Guaranteed Page tables are provably consistent | None yet No rollback, no verification on overwrite |
Pick a Side
How Thin Is the Evidence?
References
- MemGPT: Treating LLM Context Windows Like Virtual Memory —
Packer, Wooders, Lin, Fang, Patil, Stoica, Gonzalez, UC Berkeley, 2023.
arxiv.org/abs/2310.08560 - Lost in the Middle: How Language Models Use Long Contexts —
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni,
Percy Liang, 2023.
scholar.google.com → - Generative Agents: Interactive Simulacra of Human Behavior —
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang,
Michael S. Bernstein, 2023.
scholar.google.com → - Beyond Goldfish Memory: Long-Term Open-Domain Conversation —
Jing Xu, Arthur Szlam, Jason Weston, 2021.
scholar.google.com → - Improving Language Models by Retrieving from Trillions of Tokens (RETRO) —
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, et al., 2022.
scholar.google.com → - ReAct: Synergizing Reasoning and Acting in Language Models —
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022.
scholar.google.com →