Prefill-as-a-Service for Cross-Datacenter KV Cache

Visual companion for the episode on whether hybrid-attention models shrink KV cache enough to make prefill → remote transfer → local decode practical over commodity Ethernet. This page emphasizes the systems bottleneck: moving attention state, not routing requests.
arXiv: 2604.15039 Topic: PD Disaggregation Focus: KV Transfer Boundary Model Claim: Hybrid Attention Extra arXiv IDs in transcript: none detected
paper date
2026-04-16
case study
1T internal model
headline result
+54%
naive hetero
+32%
Mock visual data below is designed to illustrate the paper’s mechanism and the hosts’ skepticism: architecture shrinkage can shift the feasible network boundary, but evidence generalization remains limited.

Selective Offloading Pipeline

Requests are triaged by prompt length, cache status, and link conditions. The point is not “send everything remote,” but to remote-prefill only when compute savings exceed KV transport cost.

Local path
Remote prefill path
KV transfer object
Potential bottleneck / hot spot

Request Mix Radar

Only a subset of traffic is worth remote prefill. Hover bars to inspect why some requests stay local.

Where the Serving Stack Breaks

The limiting object is the persistent attention state. Request routing is cheap; shipping per-request KV is not.

KV Cache Growth by Architecture

Dense attention scales persistent state across many layers. Hybrid designs reduce the number of full-attention layers or use bounded-state blocks.

Layer-State Heatmap

Rows are layers, columns are prompt positions. Hybrid models carry “bright” state in fewer places.

Transferable State Composition

Mock stacked bars show how much of the prompt-state footprint comes from full attention versus bounded-state blocks as sequence length grows.

Break-even Heatmap

Cells indicate whether remote prefill wins after adding KV transfer latency. Axes: bandwidth and prompt length.

Latency Budget Waterfall

Compare local prefill against remote prefill with transfer. Use the request selector to see where network dominates.

Cross-Cluster Link Utilization Timeline

Bandwidth-aware scheduling exists because “longer than X” is too brittle; offload decisions should react to congestion and cache unevenness.

Throughput Comparison

Illustrative reconstruction of the reported style of result: homogeneous PD vs naive heterogeneous vs selective offloading.

What Carries the Gain?

Ablation-style mock bars visualize the hosts’ critique: architecture, routing, bandwidth awareness, and hardware mix are entangled.

Decision Surface: Offload or Keep Local?

Interactive matrix over prompt length and cache-hit probability. Hover cells for the recommended policy under the selected architecture/network assumption.

References

Compact source map for the episode and adjacent serving literature.

Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
Qin et al., 2026 · arXiv:2604.15039
Mooncake
KVCache-centric serving architecture · 2024 · Scholar
DistServe
Prefill/decode disaggregation for goodput · 2024 · Scholar
SARATHI
Chunked prefills and efficient inference · 2024 · Scholar
vLLM / SGLang / Dynamo
Paged KV, scheduling, serving stacks · 2023–2025 · vLLM
KV cache reuse & security
LMCache, HotPrefix, KVShare, CacheSolidarity · 2024–2025