2. Merrill & Sabharwal (2024). Theoretical expressive power of reasoning models.
3. Hong et al. (2025). Context Rot.
4. Khattab et al. (2021). Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval.
5. OpenAI (2025). Context compaction work.
6. Smith (2025). Long-context compaction work.
7. Wu et al. (2021). Task-specific long-context methods.
8. Wu et al. (2025). Context compaction / long-context scaffolding.
9. Anthropic (2025). Self-delegation / sub-agent work.
10. Schroeder et al. (2025). Self-delegation work.
11. Sun et al. (2025). Self-delegation work.
12. Chen et al. (2025). Deep research benchmark/work.
13. Bertsch et al. (2025). Information aggregation benchmark/work.
14. Bai et al. (2025). Code repository understanding benchmark/work.
15. Yang et al. (2025). Qwen3 technical report.
16. Core Context Aware Transformers for Long Context Language Modeling (2024).
17. Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis (2024).
18. Recursively Summarizing Enables Long-Term Dialogue Memory in LLMs (2024).
19. Augmenting Language Models with Long-Term Memory (2024).
20. Retrieval Augmented Generation or Long-Context LLMs? A Comparative Study and Hybrid Approach (2024/25).
21. LongRAG: Enhancing Retrieval-Augmented Generation with Long-Context LLMs (2024/25).
22. Let’s (Not) Just Put Things in Context: Test-Time Training for Long-Context LLMs (2025).
23. Z1: Efficient Test-Time Scaling with Code (2025).
24. Survey on Test-Time Scaling in LLMs (2025).