Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

Forys, Wu, Xiao, Nie, Liu, Antonova, Jones, Mullins, Luk, Zhao, Constantinides · Imperial College London & University of Cambridge · submitted 2026-08-04
arXiv:2608.03741

Splitting the inference pipeline: two cuts vs four

Toggle between the now-standard prefill/decode split and this paper's PDAF proposal, which cuts attention away from FFN inside each phase too.

Context length: chatbot vs agentic

Agentic tool-calling loops re-prefill an ever-growing context nearly every turn, pushing average context far beyond chat traffic.

Why this matters

2.06x
peak PDAF gain (NPU space)
100K+
tokens, agentic context spikes
8
nodes, real B200 validation
9
prefill:output ratios swept

The headline number is real but conditional — it only materializes in a specific regime of workload ratio and hardware flexibility, explored in the next tabs.

Compute-bound vs memory-bandwidth-bound, per stage

Prefill reads many tokens in parallel (compute-bound). Decode reads one token against a growing KV cache (memory-bandwidth-bound). Hover a bar for detail.

Why attention and FFN batch differently

FFN applies identical weights to every token — bigger batches mean more reuse. Attention reads a distinct KV cache per request — bigger batches just add memory traffic. The heatmap shows utilization as batch size grows.

PDAF throughput multiplier vs prefill:output ratio

Figure 5's sweep: on the custom NPU space, PDAF is worse than doing nothing at a balanced ratio, then jumps sharply once workloads turn prefill-heavy — exactly where agentic traffic lands.

Why GPUs barely benefit from the extra cut

Table V: when the search picks hardware for each of PDAF's four stages on the real GPU catalog, 31 of 32 stage assignments across all eight models just land on an H100 (one lands on an A100). Compute and bandwidth are bundled on real GPUs, so there's no independent knob for attention vs FFN to exploit beyond what prefill/decode already captured.

KV latent rank vs PDAF gain

Widening the MLA-style compressed KV vector (kv_lora_rank, 128→4096) makes decode-attention more clearly memory-bandwidth-bound — and PDAF's benefit climbs monotonically with it.

Total experts (flat) vs active experts (collapses)

Growing total MoE experts while holding active experts fixed adds only capacity pressure — nearly free. Growing active experts adds real compute per token and collapses PDAF's benefit outright.

Quantization: which stage tolerates 4-bit depends on the workload

Qwen3.5-32B, 8-bit baseline. Dropping one stage to 4-bit at a time — accuracy relative to task. FFN precision matters for short GSM8K-style tasks; attention precision matters for long-context BFCL.

Table II: decoupled compute × memory design space

Real GPUs bundle compute and bandwidth — an H100 is an H100. The synthetic NPU space lets compute range 25–20,000 TFLOPs completely independently of memory, drawn from seven technology classes. Real GPUs (grey) sit as fixed bundled points; the design space (colored) sweeps freely.

HeteroPanacea validated against a real 8-node B200 cluster

Roofline compute predictions vs measured GPU kernels (Fig. 2), and the power model vs datasheet TDP (Fig. 3). The scheduler layer above this roofline — greedy prefill batching, FIFO decode — is not independently validated against a live serving stack.

References