Real Context Size and Context Rot

A visual companion on the gap between advertised token capacity and usable context: how retrieval-only demos flatter models, why performance decays unevenly with length, and why a giant prompt buffer is not memory.
Anchor paper arXiv:2404.06654
Benchmark ecosystem arXiv:2308.14508
Theme capacity ≠ reliability
Mode visual-heavy / interactive
RULER
Context Rot
Needle vs multi-hop
Lost in the middle
RAG ≠ memory

Episode frame

Main distinction
Accepted ≠ Usable
Primary benchmark
RULER
Failure lens
Context Rot
Models discussed
18 frontier systems
The visuals below use realistic mock data shaped by the episode’s claims: retrieval holds up longer than tracing/aggregation, performance falls non-uniformly with length, distractors and semantic mismatch accelerate decay, and structured memory beats “just add tokens.”
stable use partial decay severe rot
Hover nodes for notes on the dominant failure style.
Click tabs above to step through the progression from lexical lookup to longer-range integration.
vanilla NIAH RULER broader suite Chroma-style stress
Dashed overlays mark advertised window; solid bars show usable range under the selected scoring rule.
Retrieval Tracing Aggregation Distractor QA
Switch modes to compare a flat token pile with tiered memory and retrieval.
Hover cells to inspect retrieval reliability by position and distractor density.
RULER
What’s the Real Context Size of Your Long-Context Language Models?
arXiv:2404.06654
LongBench
A Bilingual, Multitask Benchmark for Long Context Understanding
arXiv:2308.14508
Lost in the Middle
Position effects in long context use
Scholar link
InfiniteBench
Long-context evaluation beyond 100K tokens
Scholar link
L-Eval
Standardized long-context evaluation
Scholar link
Episode thesis
Long context is an upper bound on buffer size, not a guarantee of memory or reasoning quality.