AI Post Transformers · Episode Companion

MemGPT: Treating LLM Context Windows Like Virtual Memory

UC Berkeley's MemGPT reframes the fixed context window as scarce RAM and pages facts, summaries, and full history in and out of external storage tiers via the model's own function calls — instead of trying to cram everything into one ever-larger prompt.

arXiv:2310.08560 Packer, Wooders, Lin, Fang, Patil, Stoica, Gonzalez · UC Berkeley · Oct 2023 Hosts: Hal Turing & Dr. Ada Shannon

The Memory Hierarchy: RAM vs. Disk, for LLMs

MemGPT splits the world into main context (what's actually in the prompt — analogous to physical RAM) and external context (everything else, paged in on demand — analogous to disk). Click any block below.

System Architecture

Click a component to see its role. Arrows show the paging path.
Click a box in the diagram above to learn what it does.

Why Context Windows Are Scarce

Self-attention cost scales quadratically — doubling context ≈ 4× compute.
140 messages. That's roughly what an 8K window (GPT-4's original release) holds before conversation falls off the edge — the wall MemGPT is built to work around.

Even Big Windows Don't Help: Lost in the Middle

Liu et al. (2023) showed recall accuracy is highest for facts near the start or end of context — and sags badly in the middle. The longer the window, the worse the sag. This is the second reason MemGPT pages instead of just stuffing more tokens in.

Recall Accuracy by Position in Context

Hover a cell for the exact recall rate at that position.
The connection: MemGPT's benchmarks (deep memory retrieval, nested key-value lookup) are built directly on top of this failure mode — paging a fact back into the front of context on demand sidesteps the middle-of-context blind spot rather than solving it.

The Queue Manager: Paging in Action

As the FIFO queue fills with conversation history, the queue manager monitors token usage against two thresholds — a soft warning and a hard eviction point. Step through what happens.

Walk Through the Eviction Cycle

Context Window Fill Meter

Tracks live against the 70% warning line and 100% eviction line.
0%70% warn100% evict
Evicted messages are never destroyed — they move to recall storage, searchable later. Only a recursive summary remains visible in the queue.

Benchmark Results

Four evaluations from the paper: deep memory retrieval, nested key-value lookups across nesting depth, document QA at scale, and conversation-opener quality.

Deep Memory Retrieval (DMR) Accuracy

Recall something from five conversations back.

Nested Key-Value Lookup vs. Nesting Depth

140 UUID pairs, ~8K tokens. Click a legend swatch to isolate a model.

Document QA vs. Documents Retrieved (Fig. 5)

NaturalQuestions-Open, 50 sampled questions, same retriever for both.

Conversation Opener: Similarity to Human Baseline

SIM-1 / SIM-3 measure persona-fact overlap; SIM-H measures overlap with the actual human-written line.
MemGPT beats the human baseline on SIM-1/SIM-3 (more persona facts surfaced) but trails badly on SIM-H (0.773 vs. 1.0) — it isn't reproducing the human's line, it's out-covering it.

Does the OS Analogy Actually Hold?

Hal and Ada spend most of the episode arguing this exact point. A real kernel pages memory transparently; MemGPT's "kernel" is the language model itself, deciding via function calls.

Kernel vs. MemGPT, Dimension by Dimension

DimensionReal OS KernelMemGPT
Who decides eviction? Kernel Deterministic page-replacement algorithm The LLM Emits its own function calls to save/archive
Reliability Fixed A kernel has no "off days" Model-dependent Hostage to function-call accuracy
Eviction policy LRU / clock Well-defined algorithm "Ask the model" Still a hierarchy + policy, just implemented differently
Failure mode Page fault Handled transparently, app unaware Missed call Wrong/late function call = lost context
Verification Guaranteed Page tables are provably consistent None yet No rollback, no verification on overwrite

Pick a Side

Two hosts, one paper — toggle between their hottest takes.

How Thin Is the Evidence?

Sample size behind the headline document-QA curve, no confidence intervals reported.
Also: both DMR and document QA use GPT-4 as judge, and the paper admits MemGPT's answers run more verbose than gold answers — a known confound (Zheng et al., MT-Bench) for LLM-judged evaluations.

References

  1. MemGPT: Treating LLM Context Windows Like Virtual Memory — Packer, Wooders, Lin, Fang, Patil, Stoica, Gonzalez, UC Berkeley, 2023.
    arxiv.org/abs/2310.08560
  2. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023.
    scholar.google.com →
  3. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2023.
    scholar.google.com →
  4. Beyond Goldfish Memory: Long-Term Open-Domain Conversation — Jing Xu, Arthur Szlam, Jason Weston, 2021.
    scholar.google.com →
  5. Improving Language Models by Retrieving from Trillions of Tokens (RETRO) — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, et al., 2022.
    scholar.google.com →
  6. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022.
    scholar.google.com →