AI Post Transformers · Visual Companion

LAPS for Length-Aware LLM Serving

This page visualizes the paper’s main claim: long prefills and short re-prefills are not just different sizes of the same job. They hit different bottlenecks, so LAPS splits them into separate traffic classes and schedules them differently to protect time to first token.

arXiv 2601.11589 Venue MLSys 2025 Core metric TTFT / prefill latency Extra IDs none beyond transcript’s 2601.11589
What the visuals emphasize
2 queues
short-prefill vs long-prefill
2 modes
temporal and spatial split
3 wins
latency, SLOs, throughput

Where Prefill Stops Looking Compute-Bound

Drag the threshold. The crossover marks the point where short chat turns behave more like memory traffic and long prompts behave more like compute-heavy tensor work.

320 tokens
compute cost KV traffic + scheduler overhead total prefill latency

Traffic Distribution

Mock trace points emulate the paper’s workload intuition: many arrivals cluster in the short re-prefill regime, while a smaller set of long prompts consumes heavy compute slices.

0
short-prefill requests
0
long-prefill requests
0%
share near crossover

Head-of-Line Blocking vs LAPS

Toggle between a single mixed prefill queue and the LAPS dual-queue path. Short follow-ups wait less when they stop sitting behind giant context loads.

Temporal vs Spatial Disaggregation

LAPS has two deployment shapes: alternate service on one prefill instance, or dedicate different instances to different traffic classes when more hardware exists.

Performance Delta

These bars use paper-inspired mock values anchored to the transcript’s directional claims: lower prefill latency and SLO violations, higher throughput under mixed traffic.

User-Perceived Impact

TTFT matters because the first visible response dominates how “fast” a chat product feels. This panel sketches why small queueing changes produce outsized UX differences.

Batch Shape Heatmap

Short-request waiting windows and length-aware bucketing create more regular batch shapes. That regularity makes CUDA Graph capture more useful.

cool = poor graph reuse warm = moderate reuse hot = efficient execution

Short-Prefill Fast Path

A small wait, then smarter grouping, then CUDA Graph-backed execution. The policy stack is visualized as a step-by-step pipeline rather than a single scheduler tweak.

References

  1. LAPS: A Length-Aware-Prefill LLM Serving System
  2. ORCA
  3. PagedAttention
  4. SARATHI
  5. DistServe
  6. BucketServe
  7. DéjàVu
  8. Aggregation or Disaggregation?
  9. FlowPrefill
  10. BurstGPT
  11. SageServe
  12. ServeGen
  13. Fairness in Serving LLMs
  14. FairBatching
  15. Locality-Aware Fair Scheduling
  16. AI Post Transformers: Splitwise
  17. AI Post Transformers: Prefill-as-a-Service
  18. AI Post Transformers: Speculative Decoding
  19. AI Post Transformers: Shared KV Cache
  20. AI Post Transformers: Compute-Bandwidth-Memory Trade-offs