AI Post Transformers • Visual Companion

Experimental Comparison of Agentic and Enhanced RAG

A visual-first map of the paper’s core claim: naïve → enhanced → agentic is not a simple quality ladder. Targeted fixed modules often beat tool-using agents on cost and latency, while agents help mainly when iterative recovery is genuinely needed.

Primary Paper
Ferrazzi et al., 2026
Known arXiv ID
Compared Designs
Naïve / Enhanced / Agentic
Decision Axes
Quality × Latency × Token Cost

RAG design space at a glance

The page starts with a structural view: what remains fixed in pipelines, what moves into the model in agentic systems, and where “middle-ground” methods like corrective or self-reflective RAG sit.

System comparison diagram

Hover nodes for module-level notes. Toggle above to switch views.

Cost–adaptivity radar

Mock normalized values illustrating the paper’s practical framing.
Naïve Enhanced Agentic

Interactive control flow

Fixed pipelines expose explicit stages. Agentic systems move branching into the LLM loop. Step through a representative query to see where decisions differ.

Step-by-step execution trace

Current scenario: “What’s our internal refund policy for enterprise accounts?”

Per-step resource profile

Animated bars show why agents often pay extra in tokens and latency.

Experimental tradeoffs

The paper’s core empirical message is conditional advantage: enhanced wins on many well-specified failure modes, while agentic gains appear when retries and evidence repair rescue bad first-pass retrieval.

Scenario comparison bars

Four benchmark dimensions from the discussion: routing, alignment, document adjustment, and model sensitivity.

Operating frontier

Bubble chart: quality vs latency, bubble size = token burn.

Failure-mode heatmaps

These matrices visualize where each design fails: bad routing, misaligned phrasing, noisy retrieval, or over-long evidence. Hover cells to inspect severity.

Query × failure type heatmap

Blue = mild issue, orange/red = severe issue. Mock data shaped by the episode’s qualitative conclusions.

“Lost in the middle” context position map

Relevant evidence near the center of long contexts is used less reliably; reranking and trimming help more than extra agent deliberation alone.

References

Compact references used in the episode and this visualization. Additional arXiv IDs found in transcript: 2601.07711 only.

[1]Ferrazzi, Cvjeticanin, Piraccini, Giannuzzi. Is Agentic RAG worth it? An experimental comparison of RAG approaches (2026). arXiv:2601.07711
[2]Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020).
[3]Asai et al. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (2023).
[4]Shi et al. Corrective Retrieval Augmented Generation (2024).
[5]Gao, Tao, Qi et al. A Survey on Retrieval-Augmented Text Generation for Large Language Models (2024).
[6]Gao, Ma, Lin, Callan. HyDE: Precise Zero-Shot Dense Retrieval without Relevance Labels (2023).
[7]Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (2023).
[8]Schick et al. Toolformer: Language Models Can Teach Themselves to Use Tools (2023).
[9]Karpukhin et al. Dense Passage Retrieval for Open-Domain Question Answering (2020).
[10]Liu et al. Lost in the Middle: How Language Models Use Long Contexts (2024).