This page visualizes the paper’s main claim: long prefills and short re-prefills are not just different sizes of the same job. They hit different bottlenecks, so LAPS splits them into separate traffic classes and schedules them differently to protect time to first token.
Drag the threshold. The crossover marks the point where short chat turns behave more like memory traffic and long prompts behave more like compute-heavy tensor work.
Mock trace points emulate the paper’s workload intuition: many arrivals cluster in the short re-prefill regime, while a smaller set of long prompts consumes heavy compute slices.
Toggle between a single mixed prefill queue and the LAPS dual-queue path. Short follow-ups wait less when they stop sitting behind giant context loads.
LAPS has two deployment shapes: alternate service on one prefill instance, or dedicate different instances to different traffic classes when more hardware exists.
These bars use paper-inspired mock values anchored to the transcript’s directional claims: lower prefill latency and SLO violations, higher throughput under mixed traffic.
TTFT matters because the first visible response dominates how “fast” a chat product feels. This panel sketches why small queueing changes produce outsized UX differences.
Short-request waiting windows and length-aware bucketing create more regular batch shapes. That regularity makes CUDA Graph capture more useful.
A small wait, then smarter grouping, then CUDA Graph-backed execution. The policy stack is visualized as a step-by-step pipeline rather than a single scheduler tweak.