AI Post Transformers arXiv 2605.01280 Position Paper

Why LLM Serving Needs Mathematical Optimization

A visualization-first companion to the episode on why transformer serving is no longer a simple queue plus cache problem. The visuals below show how prefill/decode asymmetry, KV growth, and cluster-level routing turn serving into an objective-driven control system.

Core mismatch
4
Phase asymmetry, unknown output length, growing memory, and coupled batching.
Decision planes
6
Routing, admission, prefill scheduling, decode scheduling, KV eviction, allocation.
Main tension
TTFT ↔ p99
Local heuristics often improve one surface while destabilizing another.
Transcript IDs
1
Additional arXiv IDs found in transcript text: none beyond the source paper.

What the page tries to draw

Not a paper summary. Each tab renders one control problem: where interference appears, which objective shifts the schedule, how cache affinity fights load balance, and where stronger baselines still leave optimization room.

Serving Is a Coupled Control Loop

The position paper’s strongest claim is structural: each request mutates future memory pressure and schedule quality. The diagram shows why a clean objective beats isolated heuristics once prefill, decode, admission, and eviction begin interfering.

Flow Diagram mode = heuristic
low interference policy coupling hot bottleneck

Prefill and Decode Live on Different Geometry

Prefill burns compute in parallel over the prompt. Decode is lighter per step but stretched over time and dominated by memory bandwidth. Continuous batching couples them one iteration at a time.

Occupancy Heatmap mixed workers
idle / cold busy overloaded

Objective Choice Changes the “Best” Policy

The same serving trace looks different if the product cares most about time-to-first-token, p99 completion, or goodput under an SLO. This chart uses plausible mock measurements to show how policy ranking flips.

Policy Comparison objective = ttft

Cache Affinity Fights Load Balance

Short queue routing smooths load but can destroy reusable KV locality. Pure affinity preserves warm state but can produce hotspots. The useful operating point is usually somewhere in between, and it changes with memory pressure.

Routing Affinity Matrix policy = shortest queue
reuse-rich balanced churn-heavy

Reference Map

The cited work clusters around four themes: scheduling, prefill/decode separation, KV memory management, and cache-aware routing. The position paper sits above them as a call for cleaner objectives and stronger algorithmic framing.

Citation Theme Map papers + episode links