Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

arXiv:2605.15184 Sen, Kasturi, Lumer, Gulati, Subbiah · 2026 PricewaterhouseCoopers

Tests whether plain regex retrieval can outperform vector search inside real agent harnesses like Chronos, Claude Code, and Codex CLI — and finds that inline grep wins every harness–model pairing tested, sometimes by 20+ points, while the same backbone model swings from 93.1% to 76.7% purely by changing which harness runs it.

How a question flows through Chronos

Every question is dense-primed with a top-15 vector context block before the grep-vs-vector tool loop even starts — so no condition tested here is "purely lexical." Click a path to trace it.

Same corpus, same grader — the harness bundles prompt construction, tool ergonomics, and transcript formatting into one variable, which is why retrieval mode alone doesn't explain the paper's biggest swings.

Inline grep beats inline vector — every single pairing

Chronos spans 83.6–93.1% with inline grep vs 62.9–83.6% with inline vector. The narrowest gap is Claude Code + Opus 4.6 (76.7 vs 75.0); the widest is Chronos + Gemini 3.1 Flash-Lite (86.2 vs 62.9).

Inline Grep Inline Vector

Inline vs. programmatic (file-read) delivery

Programmatic vector beats programmatic grep on 5 of 10 pairs. But the standout is Codex CLI + GPT-5.4: identical regex, identical corpus — forcing the agent to open a file instead of reading inline drops accuracy from 93.1% to 55.2%.

Grep Vector
Codex CLI + GPT-5.4 under inline grep: 93.1%. Same regex, same corpus, file-read delivery only: 55.2%.

Session-limit sweep (Experiment 2)

As distractor sessions grow from 39 (s5) to 66 (full haystack), accuracy is not monotone. The winning retriever depends on which harness the model runs inside, not just which model it is.

Grep-only Vector-only

Table 4 — accuracy by LongMemEval category

Chronos, grep-only, full haystack. Single-session categories hit ceiling effects; multi-session aggregation and temporal reasoning show the real variance the aggregate Table 1 numbers hide. Hover a cell.

Lower accuracy Mid Ceiling (~100%)

References

1Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sen, Kasturi, Lumer, Gulati, Subbiah, 20262605.15184
2ReAct: Synergizing Reasoning and Acting in Language Models — Yao et al., 2022Scholar
3WebGPT: Browser-assisted question-answering with human feedback — Nakano et al. (OpenAI), 2021Scholar
4Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al. (FAIR), 2020Scholar
5LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Wu, Wang, Yu et al., 2024Scholar
6The Probabilistic Relevance Framework: BM25 and Beyond — Robertson, Zaragoza, 2009Scholar
7Dense Passage Retrieval for Open-Domain Question Answering — Karpukhin, Oğuz, Min et al. (FAIR), 2020Scholar
8Lost in the Middle: How Language Models Use Long Contexts — Liu, Lin, Hewitt et al., 2023Scholar
9SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Formal, Piwowarski, Clinchant, 2021Scholar
10BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of IR Models — Thakur, Reimers, Rücklé, Srivastava, Gurevych, 2021Scholar
11MemGPT: Towards LLMs as Operating Systems — Packer, Fang, Patil, Lin, Wooders, Gonzalez, 2023Scholar
12Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval — Sen, Lumer, Gulati, Subbiah, 2026Scholar

Interactive Visualization

Explore the full companion viz: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy