Rows are task conditions, columns are prompt lengths. Hotter cells indicate higher failure probability.
As prompts grow, models drift from exact retrieval toward priors, copying, omission, and hallucination.
The benchmark moves from simple lookup to state maintenance across long buffers.
Needle-in-a-haystack looks flattering because it often measures literal matching more than robust context use.
Mock benchmark cutoffs at the longest length where performance stays “satisfactory.”
Retrieval stays high longest; tracing, aggregation, and distractor-rich QA bend downward earlier.
A big prompt is a cluttered desk. A memory architecture ranks, stores, retrieves, compresses, and persists.
Unique facts near the edges are often easier to recover than facts buried in the middle under distraction.
Compact anchors from the episode and related benchmark ecosystem.