← All episodes Lost in the Middle and Long Context LLMs

Lost in the Middle and Long Context LLMs

Apr 6, 2026
This episode examines the widening gap between long-context marketing claims and actual downstream performance. Using Lost in the Middle (Liu et al., 2023) as the anchor paper, the hosts explain why the ability to accept 32K, 128K, or even million-token prompts is only an interface claim—not evidence that a model can reliably use information spread across that context. They situate the discussion in the broader evolution of long-context language models, from the Transformer architecture and Transformer-XL to systems advances like FlashAttention, and describe how modern applications have turned the prompt into a packed working memory of retrieved documents, chat history, tool outputs, transcripts, and examples. The conversation focuses on the difference between retrieval and reasoning. The hosts contrast impressive needle-in-a-haystack results and power-law-style retrieval trends reported in long-context evaluations with a growing body of “context rot” findings showing that real task performance often deteriorates as more tokens are added. They explain why locating a planted fact is not the same as summarizing long documents, answering questions across many sources, performing multi-hop reasoning, or learning patterns from buried examples. A central theme is positional bias: Lost in the Middle shows that models often display primacy and recency effects, producing a U-shaped accuracy curve where relevant information placed at the beginning or end of a prompt is used more effectively than information buried in the middle. The episode also connects this paper to a broader research wave. It references Same Task, More Tokens (Levy et al., 2024) on downstream degradation under longer prompts, RULER (Hsieh et al., 2024) on more demanding synthetic long-context evaluations, and Long-Context LLMs Struggle with Long In-Context Learning (Li et al., 2024) on many-shot learning plateau effects. Together, these studies paint a more cautious picture than benchmark headlines suggest: larger context windows can improve coverage, but they do not guarantee better integration, robustness, or reasoning. The discussion gives practitioners a grounded view of what million-token context windows can and cannot be expected to deliver in real systems.
Sources:
1. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023
http://arxiv.org/abs/2307.03172
2. Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models — Mosh Levy, Alon Jacoby, Yoav Goldberg, 2024
http://arxiv.org/abs/2402.14848
3. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024
http://arxiv.org/abs/2404.06654
4. Long-context LLMs Struggle with Long In-context Learning — Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, Wenhu Chen, 2024
http://arxiv.org/abs/2404.02060
5. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
6. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context
7. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
8. How to Train Long-Context Language Models — Shouyuan Chen, Weizhe Yuan, Yifan Chen, et al., 2023
https://scholar.google.com/scholar?q=How+to+Train+Long-Context+Language+Models
9. LongChat: Scaling Up Instruction Tuning for Long Contexts — Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Cai, et al., 2023
https://scholar.google.com/scholar?q=LongChat%3A+Scaling+Up+Instruction+Tuning+for+Long+Contexts
10. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
11. Needle In A Haystack — Greg Kamradt, 2023
https://scholar.google.com/scholar?q=Needle+In+A+Haystack
12. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, et al., 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
13. Mitigating the lost-in-the-middle phenomenon: a context re-ranking strategy for improving retrieval-augmented generation performance — approx. 2024 authors unclear from snippet, 2024
https://scholar.google.com/scholar?q=Mitigating+the+lost-in-the-middle+phenomenon%3A+a+context+re-ranking+strategy+for+improving+retrieval-augmented+generation+performance
14. LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression — Jiang et al. (approx.), 2024
https://scholar.google.com/scholar?q=LongLLMLingua%3A+Accelerating+and+enhancing+LLMs+in+long+context+scenarios+via+prompt+compression
15. Found in the middle: How language models use long contexts better via plug-and-play positional encoding — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Found+in+the+middle%3A+How+language+models+use+long+contexts+better+via+plug-and-play+positional+encoding
16. An efficient recipe for long context extension via middle-focused positional encoding — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=An+efficient+recipe+for+long+context+extension+via+middle-focused+positional+encoding
17. Length extrapolation of transformers: A survey from the perspective of positional encoding — approx. 2024–2025 survey authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Length+extrapolation+of+transformers%3A+A+survey+from+the+perspective+of+positional+encoding
18. Lost in the Middle, and In-Between: Enhancing Language Models' Ability to Reason Over Long Contexts in Multi-Hop QA — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Lost+in+the+Middle%2C+and+In-Between%3A+Enhancing+Language+Models%27+Ability+to+Reason+Over+Long+Contexts+in+Multi-Hop+QA
19. When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=When+to+Memorize+and+When+to+Stop%3A+Gated+Recurrent+Memory+for+Long-Context+Reasoning
20. Mastering long-context multi-task reasoning with transformers and recurrent memory — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Mastering+long-context+multi-task+reasoning+with+transformers+and+recurrent+memory
21. Augmenting language models with long-term memory — approx. 2024–2025 authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Augmenting+language+models+with+long-term+memory
22. AI Post Transformers: Lost in the Middle: How Language Models Use Long Contexts — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/lost-in-the-middle-how-language-models-use-long-contexts/
23. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, Thu,
https://podcast.do-not-panic.com/episodes/rope/
24. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
25. AI Post Transformers: CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-cacheslide-position-aware-kv-cache-reuse-cd59c7.mp3
26. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
27. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
28. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3