AI Post Transformers • Interactive Visualization Companion

EverMemOS for Long-Horizon Agent Memory

A visualization-first tour of memory as state management: raw traces become MemCells, cells consolidate into MemScenes, and query-time recollection rebuilds only the context that should govern the next decision.

arXiv 2601.02163 Posted Jan 9, 2026 Mode Structured long-horizon agent memory Transcript IDs 2601.02163 Viz URL live page

Memory Lifecycle

The paper’s key move is organizational: do not keep a flat pile of retrieved snippets. Build a memory layer that writes, consolidates, and reconstructs context with explicit state transitions.

Atomic unitMemCell

Episodic traces, facts, and bounded foresight are stored as typed cells rather than raw chat chunks.

Abstraction unitMemScene

Cells merge into semantic scenes that preserve temporal relations and reduce fragment sprawl.

Query-time policyReconstruct

Only the context judged necessary and sufficient is assembled back into prompt space.

Core tensionState > Storage

The challenge is not just remembering more. It is deciding which memory should govern the current turn.

Why Flat Memory Fails

Hover the matrix: the problem is not token volume alone. Long-lived agents accumulate stale preferences, contradictory facts, fragmented traces, and relevance collisions.

Hot cells19

High-conflict or high-staleness regions under the selected memory policy.

Scene coherence0.42

Illustrative score for how well retrieved evidence aligns into one governing situation.

Prompt clutter73%

Share of retrieved context that is present but not decision-relevant.

Recollective Loop

Step through the system’s query-time behavior. The flow below exposes how query rewriting, scene retrieval, cell drill-down, reranking, and selective assembly split the job.

Step 1

Rewrite the user request into memory-relevant facets like persona, temporal status, task state, and safety constraints.

Step 2

Retrieve candidate scenes first, then descend to supporting cells instead of directly ranking isolated snippets.

Step 3

Rerank for sufficiency, not just semantic similarity, so contradictory but obsolete notes do not dominate prompt space.

Step 4

Assemble only the governing evidence slice needed for the next answer or action.

Benchmark Shape

Toggle between absolute scores and relative lift. The two annotated gains come from the podcast discussion; the rest are illustrative mock values used to show the trade space between quality, latency, and memory control.

LoCoMo+9.2%

Relative gain over strongest baseline, as discussed in the episode.

LongMemEval+6.7%

Relative gain in the emphasized GPT-4.1-mini setup from the discussion.

Reading of the paperStack win

The improvement appears to come from better organization plus retrieval policy, not a new base model.

References

EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning — arXiv:2601.02163
MemoryBank: Enhancing Large Language Models with Long-Term Memory — Scholar
MemGPT: Towards LLMs as Operating Systems — Scholar
A Survey on the Memory Mechanism of Large Language Model based Agents — Scholar
MemOS: A Memory OS for AI System — Scholar
Zep: A Temporal Knowledge Graph Architecture for Agent Memory — Scholar
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Scholar
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Scholar
AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Episode
AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Episode
AI Post Transformers: Neural Computers as Learned Latent Runtimes — Episode
AI Post Transformers: How Induction Heads Emerge in Transformers — Episode