1 00:00:01,000 --> 00:00:33,350 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called 'Is Grep All You Need? How Agent Harnesses Reshape Agentic Search,' by Sahil Sen et al., five co-authors total, out of PricewaterhouseCoopers, posted to arXiv on May 14th, 2026. And I want to start with the number that made me sit up: on one of their harnesses, plain old grep beats a full vector index by more than twenty points. Not a rounding error. 2 00:00:33,350 --> 00:00:47,075 [Hal Turing] So let's get into the mechanics, because 'grep beats vector' is a nice headline but the setup is doing a lot of work. Ada, walk me through Chronos itself — what's actually happening when the agent gets a question? 3 00:00:47,075 --> 00:01:21,525 [Dr. Ada Shannon] Right, so Chronos isn't just 'here's a question, go search.' Every episode starts with category-conditioned dynamic prompting — the system instructions literally change depending on whether the pipeline detects a temporal-reasoning question versus a preference-recall question. Then, before the tool-calling loop even begins, the agent gets an initial context block: the top-15 results from vector search, dumped in up front. Only after that does the grep-versus-vector tool loop start. So Chronos is never purely 'lexical' or 'purely dense' — it's dense-primed, then whatever retrieval mode the experiment condition allows kicks in on top of that seed. 4 00:01:21,525 --> 00:01:29,900 [Hal Turing] Wait, so even the grep-only condition gets a vector head start? That seems like it complicates the clean comparison a bit. 5 00:01:29,900 --> 00:02:06,725 [Dr. Ada Shannon] It does, and it's worth flagging now — we'll come back to what that means for interpreting the results. But on the tool side itself, the two retrievers are genuinely different animals. Grep loads the conversation turns and the Chronos-extracted temporal events — dates, intervals, spans — into memory per question, and runs regex matching, scored by match count. Zero embedding cost, all in-process. Vector search is the opposite: everything's embedded into a per-question index at ingestion time, and at query time the tool embeds the natural-language query, does approximate nearest-neighbor search, then reranks before handing the agent its top-k. 6 00:02:06,725 --> 00:02:12,075 [Hal Turing] And on the grading side — GPT-4o is the judge for all of this? 7 00:02:12,075 --> 00:02:41,775 [Dr. Ada Shannon] Sole grader, yeah. For every question it gets the question text, the reference answer, and the agent's hypothesis, and returns a binary judgment — did it get it right or not. But it's not a flat rubric — there's category-conditioned tolerance baked in. Off-by-one temporal counts get some slack, preference items get scored more rubric-style, and there's specific abstention handling for the _abs variants where the correct answer is 'I don't know.' Same grader, same prompt template, same decoding settings held fixed across every single condition in both experiments. 8 00:02:41,775 --> 00:02:51,775 [Hal Turing] Okay, so let's get to the actual numbers, because Table 1 is dense. Inline grep beats inline vector for literally every harness-model pair? 9 00:02:51,775 --> 00:03:33,000 [Dr. Ada Shannon] Every single one. Chronos spans 83.6 to 93.1 percent with inline grep across the five backbones, versus 62.9 to 83.6 with inline vector. Biggest gap is Chronos with Gemini 3.1 Flash-Lite — 86.2 versus 62.9. Narrowest is Claude Code with Opus 4.6 — 76.7 versus 75.0, basically a coin flip. But here's the number that actually stopped me: that same Opus 4.6 backbone hits 93.1 percent on Chronos and only 76.7 on Claude Code. Same model, same corpus, same questions — swapping the harness moves the needle by as much as swapping the retriever does within a fixed harness. 10 00:03:33,000 --> 00:03:54,675 [Hal Turing] Oh wait wait wait — that's the framing you led with off the top, right? The harness matters as much as the retrieval strategy. I actually think that's overstating it a little, Ada — isn't it just as plausible Claude Code's sandboxing is choking something unrelated to retrieval entirely, and they're conflating two different variables under one number? 11 00:03:54,675 --> 00:04:27,225 [Dr. Ada Shannon] No, I don't think that's overstating it — that's the actual finding, and it holds up across multiple backbones, not just one weird outlier. GPT-5.4 ties the best Chronos inline-grep score at 93.1 on Codex CLI too. This isn't cherry-picked. And I'd push back on 'conflating' — the whole point of a harness is that it bundles prompt construction, tool ergonomics, and transcript formatting together. You can't cleanly separate 'retrieval' from 'orchestration' in a real deployed system, so measuring them tangled together is actually more honest than a synthetic ablation would be. 12 00:04:27,225 --> 00:04:38,975 [Hal Turing] Fair — I'll grant the pattern's consistent. I just want us to be careful not to treat 'harness' as one clean variable when it's really a bundle of five things happening at once. 13 00:04:38,975 --> 00:05:17,750 [Dr. Ada Shannon] Agreed on that nuance. And it gets messier once you add programmatic, file-based delivery into the mix. Programmatic vector actually beats programmatic grep on five of the ten harness-model pairs — Chronos and Claude Code with Opus 4.6, Codex with GPT-5.4, Gemini CLI with both Gemini variants. But the standout is Codex with GPT-5.4: 93.1 percent under inline grep, then it falls off a cliff to 55.2 percent under programmatic grep. Same regex, same corpus — just forcing the agent to open a file and read it instead of getting results dumped inline nearly halves its accuracy. 14 00:05:17,750 --> 00:05:26,625 [Hal Turing] That's brutal. So cheap retrieval doesn't mean cheap end-to-end if the read-the-file step is where the agent actually falls apart. 15 00:05:26,625 --> 00:05:59,975 [Dr. Ada Shannon] Exactly the takeaway. Then Experiment 2 stress-tests the noise dimension — they sweep session limits, s5, s10, s20, s30, and full, from 39 up to 66 sessions per question, holding delivery fixed and running paired grep-only and vector-only tables so both retrievers face identical distractor exposure at each step. And the accuracy curves are not monotone at all. Chronos Opus on grep rises to 90.5 at s20, dips to 85.3 at s30, then climbs back to 89.7 at full. It's genuinely bumpy. 16 00:05:59,975 --> 00:06:05,200 [Hal Turing] Any consistent pattern at all, or is it just noise all the way down? 17 00:06:05,200 --> 00:06:45,025 [Dr. Ada Shannon] There's a vendor-level pattern that's pretty sticky. Claude Code favors grep for both Opus and Haiku at every single session-limit configuration they report. Gemini CLI with Gemini 3.1 Pro favors vector at every configuration too, with the gap actually widening to 89.7 versus 78.5 at full haystack. Meanwhile on Chronos, the same Gemini 3.1 Pro backbone crosses over — vector's ahead from s5 through s20, then grep takes it at full, 86.6 versus 84.5. So the retriever that wins depends on which shell you're running the model inside, not just which model it is. 18 00:06:45,025 --> 00:06:54,750 [Hal Turing] And the per-category table — that's the one broken out by knowledge-update, multi-session, temporal reasoning, all six LongMemEval buckets? 19 00:06:54,750 --> 00:07:31,000 [Dr. Ada Shannon] Right, Table 4, grep-only, full haystack, Chronos harness. It's basically ceiling effects on the single-session categories — single-session-assistant and single-session-preference both hit 100 percent for three of the five models. Where it actually gets hard is multi-session aggregation, which sits in the 69 to 84 percent range across the board, and temporal reasoning, which swings wildly — Gemini 3.1 Pro and Flash-Lite both hit 100 percent there, but GPT-5.4 drops to 67.7. So the aggregate accuracy numbers in Table 1 are hiding a lot of category-level variance underneath. 20 00:07:31,000 --> 00:07:49,250 [Dr. Ada Shannon] Right, multi-session sits in the low-to-mid seventies across the board, GPT-5.4 dips to 74.2, Gemini Pro to 69.3. But honestly, Hal, I don't want to keep walking the grid. There's something underneath all these tables that's been nagging me since we opened Table 1. 21 00:07:49,250 --> 00:08:27,750 [Hal Turing] Yeah, same itch here. LongMemEval's own paper says its answers are licensed by literal spans, dates, counts, stated preferences. That's practically the definition of a grep-friendly benchmark. So when the abstract says grep generally yields higher accuracy than vector, and the conclusion says it consistently does, isn't the honest version of that claim just grep wins at recalling verbatim facts from chat transcripts? The Limitations section even says they're not claiming grep beats vector in general, but that's paragraph six of section five, and the confident version is what's in the abstract everyone reads. 22 00:08:27,750 --> 00:09:04,550 [Dr. Ada Shannon] That's exactly the tension, and it's worth naming what this really is: Thakur, Reimers, Rucklé, Srivastava, and Gurevych's BEIR benchmark, out of UKP Lab Darmstadt, 2021, relearned inside a tool-calling loop instead of a static pipeline. BEIR's whole point was that BM25 is an embarrassingly strong zero-shot baseline against dense retrievers, nobody should be shocked lexical wins when ground truth is a date or a number. There's a second wrinkle: Chronos, the harness doing most of the winning, is the authors' own prior system, Sen, Lumer, Gulati, and Subbiah, 2026. They're not a neutral party grading someone else's harness. 23 00:09:04,550 --> 00:09:39,050 [Hal Turing] And that's before we even get to the grader. GPT-4o, alone, binary judgments, category-conditioned tolerances that are never fully spelled out, no inter-rater check, no human-validated subset. On a hundred sixteen questions, one flipped grading call is basically a full percentage point. So when section 4.2.4 builds a crossover story off eighty-five-three versus eighty-four-five, that's a one-question swing dressed up as a pattern. Honestly, Ada, I think a good chunk of this paper's most specific claims might be grader noise. 24 00:09:39,050 --> 00:09:59,550 [Dr. Ada Shannon] I actually disagree with you there, Hal, at least on the headline. Sure, eighty-five-three versus eighty-four-five, I won't defend that to the decimal. But inline grep beats inline vector on ten out of ten harness-model pairs. That's not one shaky measurement, that's ten independent readings all pointing the same direction. If it were pure grader noise you'd expect that to wash out close to half the time, not zero. 25 00:09:59,550 --> 00:10:22,125 [Hal Turing] Sure, but those aren't ten independent draws, Ada, same benchmark, same grader, same literal-span answer format every time. If GPT-4o leans systematically toward verbatim matches, that bias hits every row identically. Ten-for-ten is consistent with a real effect, but it's also consistent with one shared confound. 26 00:10:22,125 --> 00:10:36,575 [Dr. Ada Shannon] Fair, that's a real confound, not just variance averaging out. I'll take that framing, direction's probably real, precision isn't. Grep wins on this task distribution, but don't trust any single margin to the decimal. 27 00:10:36,575 --> 00:10:50,275 [Hal Turing] That I can live with. There's another citation doing similar work without being tested, though, you've name-dropped context rot a couple times as the reason for file-based delivery. Where's that actually from? 28 00:10:50,275 --> 00:11:11,550 [Dr. Ada Shannon] Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, and Liang, Lost in the Middle, Stanford, 2024. That's where context rot comes from, and it's the entire justification for programmatic delivery, get results out of the window before they crowd the model. But this paper never measures position-within-context as a variable. It just imports the concept and assumes file-based routing fixes it— 29 00:11:11,550 --> 00:11:27,825 [Hal Turing] Oh wait wait wait, but we already saw it doesn't fix it! Codex on GPT-5.4 went from ninety-three percent inline grep to fifty-five percent programmatic grep. That's not relief from context rot, that's a brand new failure mode. 30 00:11:27,825 --> 00:11:55,650 [Dr. Ada Shannon] Right, and that failure mode has a name, Packer, Fang, Patil, Lin, Wooders, and Gonzalez's MemGPT, Berkeley, 2023. Paging results to disk instead of context is exactly their OS-paging metaphor. What this paper shows that MemGPT's pitch didn't emphasize is that paging only helps if the read-back step is reliable. They call it brittle read-integrate-retry cycles, but that's a guess, no trace log, no failure taxonomy. Their single biggest, most quotable number is also their least explained one. 31 00:11:55,650 --> 00:12:21,925 [Hal Turing] So if I'm a practitioner reading this, the takeaway isn't switch to grep. It's that harness and delivery path are bigger levers than retriever choice, Opus 4.6 swung from ninety-three to seventy-seven percent just by changing which CLI it ran in, same retriever, same corpus. And file-based delivery isn't a free context-pressure fix, it's a competence test for your agent's tool-use loop. Validate the read-back before you ship it. 32 00:12:21,925 --> 00:12:44,850 [Dr. Ada Shannon] The authors flag reasonable next steps, hybrid retrieval where the agent picks per query, non-chat corpora like code or scientific synthesis where evidence isn't a literal span, broader vendor coverage, finishing the missing Codex rows. My honest read: a solid, narrow empirical result, grep wins on verbatim-recall conversational memory, under these specific harnesses, wrapped in language that reads more general than the evidence supports. 33 00:12:44,850 --> 00:13:07,100 [Hal Turing] Fair place to land. Don't walk away thinking grep beats vector everywhere, walk away thinking retrieval, harness, and delivery are one system. If your task looks like literal fact recall over chat history, give lexical search a real seat at the table instead of defaulting to vector. That's it for this one, thanks for listening, and we'll catch you next time.