Tests whether plain regex retrieval can outperform vector search inside real agent harnesses like Chronos, Claude Code, and Codex CLI — and finds that inline grep wins every harness–model pairing tested, sometimes by 20+ points, while the same backbone model swings from 93.1% to 76.7% purely by changing which harness runs it.
Every question is dense-primed with a top-15 vector context block before the grep-vs-vector tool loop even starts — so no condition tested here is "purely lexical." Click a path to trace it.
Chronos spans 83.6–93.1% with inline grep vs 62.9–83.6% with inline vector. The narrowest gap is Claude Code + Opus 4.6 (76.7 vs 75.0); the widest is Chronos + Gemini 3.1 Flash-Lite (86.2 vs 62.9).
Programmatic vector beats programmatic grep on 5 of 10 pairs. But the standout is Codex CLI + GPT-5.4: identical regex, identical corpus — forcing the agent to open a file instead of reading inline drops accuracy from 93.1% to 55.2%.
As distractor sessions grow from 39 (s5) to 66 (full haystack), accuracy is not monotone. The winning retriever depends on which harness the model runs inside, not just which model it is.
Chronos, grep-only, full haystack. Single-session categories hit ceiling effects; multi-session aggregation and temporal reasoning show the real variance the aggregate Table 1 numbers hide. Hover a cell.
| 1 | Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sen, Kasturi, Lumer, Gulati, Subbiah, 2026 | 2605.15184 |
| 2 | ReAct: Synergizing Reasoning and Acting in Language Models — Yao et al., 2022 | Scholar |
| 3 | WebGPT: Browser-assisted question-answering with human feedback — Nakano et al. (OpenAI), 2021 | Scholar |
| 4 | Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al. (FAIR), 2020 | Scholar |
| 5 | LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Wu, Wang, Yu et al., 2024 | Scholar |
| 6 | The Probabilistic Relevance Framework: BM25 and Beyond — Robertson, Zaragoza, 2009 | Scholar |
| 7 | Dense Passage Retrieval for Open-Domain Question Answering — Karpukhin, Oğuz, Min et al. (FAIR), 2020 | Scholar |
| 8 | Lost in the Middle: How Language Models Use Long Contexts — Liu, Lin, Hewitt et al., 2023 | Scholar |
| 9 | SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Formal, Piwowarski, Clinchant, 2021 | Scholar |
| 10 | BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of IR Models — Thakur, Reimers, Rücklé, Srivastava, Gurevych, 2021 | Scholar |
| 11 | MemGPT: Towards LLMs as Operating Systems — Packer, Fang, Patil, Lin, Wooders, Gonzalez, 2023 | Scholar |
| 12 | Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval — Sen, Lumer, Gulati, Subbiah, 2026 | Scholar |
Explore the full companion viz: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy