Requests are triaged by prompt length, cache status, and link conditions. The point is not “send everything remote,” but to remote-prefill only when compute savings exceed KV transport cost.
Only a subset of traffic is worth remote prefill. Hover bars to inspect why some requests stay local.
The limiting object is the persistent attention state. Request routing is cheap; shipping per-request KV is not.
Dense attention scales persistent state across many layers. Hybrid designs reduce the number of full-attention layers or use bounded-state blocks.
Rows are layers, columns are prompt positions. Hybrid models carry “bright” state in fewer places.
Mock stacked bars show how much of the prompt-state footprint comes from full attention versus bounded-state blocks as sequence length grows.
Cells indicate whether remote prefill wins after adding KV transfer latency. Axes: bandwidth and prompt length.
Compare local prefill against remote prefill with transfer. Use the request selector to see where network dominates.
Bandwidth-aware scheduling exists because “longer than X” is too brittle; offload decisions should react to congestion and cache unevenness.
Illustrative reconstruction of the reported style of result: homogeneous PD vs naive heterogeneous vs selective offloading.
Ablation-style mock bars visualize the hosts’ critique: architecture, routing, bandwidth awareness, and hardware mix are entangled.
Interactive matrix over prompt length and cache-hit probability. Hover cells for the recommended policy under the selected architecture/network assumption.
Compact source map for the episode and adjacent serving literature.