AI Post Transformers — Episode Companion

Hypic: Position-Independent KV Caching for Hybrid-Attention LLM Serving

↗ arXiv:2607.01299 Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, et al. Xiaohongshu Inc. · Peking University · SJTU Posted July 12, 2026

Hypic caches independent prompt segments on hybrid-attention models that compress history into a fixed-size state instead of a per-token KV cache. It replaces splice-and-correct with a transition operator that composes cached segments algebraically — provably matching sequential computation.

Why prefill dominates, and where old caching tricks stop working

RAG and agent prompts push tens of thousands of tokens through the model before generation even starts. Position-independent caching (PIC) reuses that work by splicing per-token KV vectors — a primitive that simply doesn't exist inside compressed-state linear-attention layers.

Prefill cost vs. a single decode step

Mock latency scaling with prompt size — dashed line is one decode step for reference.

Two attention families, two caching realities

Full-attention layers expose a per-token list to splice; linear layers only expose one compressed blob.
The transition operator: compose, don't add

Naively adding two segments' end-states produces a structurally wrong object — Table 2's RetNet case shows this fails even in the simplest family. The fix: cache each segment's transition operator, then left-multiply.

Composition law, step by step

Correct path (green) vs. the naive shortcut (red) that Table 2 shows diverges.

RetNet: naive add vs. operator compose

Error vs. ground truth as more segments are chained. Naive add compounds; composition stays flat.

Family comparison: scalar → diagonal → dense

Storage cost, composition complexity, and end-to-end production coverage (0 = never run through the full pipeline).
What it buys in production

3.25x average TTFT reduction, 1.66x more sustainable QPS at a one-second SLO, and a 5.7x cold-prefill speedup at eight workers — traded against a modest, honestly-reported quality gap.

Headline speedups

Average across four hybrid-attention models, five workloads.

Segment parallelism: LPT scheduling across workers

Longest-segment-first assignment avoids straggler workers as parallelism scales.

Quality cost: Hypic vs. Full Recompute

1.71-point average gap against the from-scratch baseline.
Where the evidence runs out

The 8.92% linear-state drift number is measured with full-attention layers disabled — isolating one failure mode, not the deployed stack. Nobody ran the diagonal (GLA) family end to end. Both gaps are argued, not settled, in the transcript.

Isolated linear-state drift (Table 4)

Full-attention layers disabled to isolate the recurrent path only.

Seam-window correction: attention-sink decay

Error concentrates at the seam and fades fast — the paper recomputes a fixed 8-token window.

The undecomposed 1.71-point gap

Linear-state drift and seam residual were never measured together on the full deployed stack.

References

  1. HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching — Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen, 2026. arXiv:2607.01299
  2. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, et al. (Yale University), 2024. Scholar
  3. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, et al. (U Chicago), EuroSys 2025. Scholar
  4. Marconi: Prefix Caching for the Era of Hybrid LLMs — collaborators incl. Princeton / Tri Dao (author list unverified), MLSys 2025. Scholar
  5. RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation — Chao Jin, Zili Zhang, et al. (Peking University / Shanghai AI Lab), 2024. Scholar
  6. EPIC: Efficient Position-Independent Caching for Serving Large Language Models — Junhao Hu, Wenrui Huang, Weidong Wang, et al., 2025. Scholar
  7. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh (NVIDIA), 2025. Scholar
  8. Mooncake: Trading More Storage for Less Computation — Ruoyu Qin, Zheming Li, Weiran He, et al., 2025. Scholar
  9. You Need an Encoder for Native Position-Independent Caching — Shiju Zhao, Junhao Hu, Jiaqi Zheng, Guihai Chen, 2026. Scholar