RAG and agent prompts push tens of thousands of tokens through the model before generation even starts. Position-independent caching (PIC) reuses that work by splicing per-token KV vectors — a primitive that simply doesn't exist inside compressed-state linear-attention layers.
Prefill cost vs. a single decode step
Two attention families, two caching realities
Naively adding two segments' end-states produces a structurally wrong object — Table 2's RetNet case shows this fails even in the simplest family. The fix: cache each segment's transition operator, then left-multiply.
Composition law, step by step
RetNet: naive add vs. operator compose
Family comparison: scalar → diagonal → dense
3.25x average TTFT reduction, 1.66x more sustainable QPS at a one-second SLO, and a 5.7x cold-prefill speedup at eight workers — traded against a modest, honestly-reported quality gap.
Headline speedups
Segment parallelism: LPT scheduling across workers
Quality cost: Hypic vs. Full Recompute
The 8.92% linear-state drift number is measured with full-attention layers disabled — isolating one failure mode, not the deployed stack. Nobody ran the diagonal (GLA) family end to end. Both gaps are argued, not settled, in the transcript.