Toggle between the now-standard prefill/decode split and this paper's PDAF proposal, which cuts attention away from FFN inside each phase too.
Agentic tool-calling loops re-prefill an ever-growing context nearly every turn, pushing average context far beyond chat traffic.
The headline number is real but conditional — it only materializes in a specific regime of workload ratio and hardware flexibility, explored in the next tabs.
Prefill reads many tokens in parallel (compute-bound). Decode reads one token against a growing KV cache (memory-bandwidth-bound). Hover a bar for detail.
FFN applies identical weights to every token — bigger batches mean more reuse. Attention reads a distinct KV cache per request — bigger batches just add memory traffic. The heatmap shows utilization as batch size grows.
Figure 5's sweep: on the custom NPU space, PDAF is worse than doing nothing at a balanced ratio, then jumps sharply once workloads turn prefill-heavy — exactly where agentic traffic lands.
Table V: when the search picks hardware for each of PDAF's four stages on the real GPU catalog, 31 of 32 stage assignments across all eight models just land on an H100 (one lands on an A100). Compute and bandwidth are bundled on real GPUs, so there's no independent knob for attention vs FFN to exploit beyond what prefill/decode already captured.
Widening the MLA-style compressed KV vector (kv_lora_rank, 128→4096) makes decode-attention more clearly memory-bandwidth-bound — and PDAF's benefit climbs monotonically with it.
Growing total MoE experts while holding active experts fixed adds only capacity pressure — nearly free. Growing active experts adds real compute per token and collapses PDAF's benefit outright.
Qwen3.5-32B, 8-bit baseline. Dropping one stage to 4-bit at a time — accuracy relative to task. FFN precision matters for short GSM8K-style tasks; attention precision matters for long-context BFCL.
Real GPUs bundle compute and bandwidth — an H100 is an H100. The synthetic NPU space lets compute range 25–20,000 TFLOPs completely independently of memory, drawn from seven technology classes. Real GPUs (grey) sit as fixed bundled points; the design space (colored) sweeps freely.
Roofline compute predictions vs measured GPU kernels (Fig. 2), and the power model vs datasheet TDP (Fig. 3). The scheduler layer above this roofline — greedy prefill batching, FIFO decode — is not independently validated against a live serving stack.