xLLM: Co-Locating Online and Offline LLM Inference

JD.com Production System | Enterprise-Scale GPU Cluster Optimization

arXiv:2510.14686

xLLM Two-Layer Architecture

The system separates high-level scheduling intelligence from low-level hardware optimization through a clean abstraction boundary.

SERVICE LAYER Dynamic PD Disaggregation Prefill/Decode resource allocation EPD Multimodal Routing Encode-Prefill-Decode for vision inputs Online-Offline Co-Location Elastic scheduling with SLO guarantees ENGINE LAYER Multi-Layer Pipeline CPU/GPU overlap Dual-stream xTensor Memory Virtual memory Fragmentation fix Speculative Decoding Draft-verify Multi-token gen MoE Load Balancing Dynamic expert routing (EPLB) abstraction
1.7-2.2x
Throughput vs Baselines
8-GPU
Tested Cluster Size
Ascend
Target Hardware (NPU)

Dynamic Prefill-Decode Disaggregation

Real-time resource reallocation between compute-bound prefill and memory-bound decode phases based on workload characteristics.

Time of Day 00:00 06:00 12:00 18:00 GPU Instances

Resource Profile Comparison

PREFILL (Compute-Bound) Compute Utilization 95% Memory Bandwidth 40% DECODE (Memory-Bound) Compute Utilization 30% Memory Bandwidth 90% Different Resource Bottlenecks → Disaggregation Opportunity Prefill saturates compute, Decode saturates memory bandwidth Running on same instance creates interference and wastes resources

Online-Offline Workload Co-Location

Maximize GPU utilization during traffic valleys by packing offline batch jobs without violating online SLOs.

GPU Cluster Allocation Online Chatbot (Low Latency SLO) Offline Batch (Best Effort) Idle (Wasted Capacity)

Utilization Metrics by Traffic Pattern

100% 50% 0% Workload Scenarios

Performance Analysis

Throughput comparison on Ascend NPUs running Qwen and DeepSeek models at 8-GPU scale.

Throughput Comparison (Requests/Second) 200 150 100 50 0 Framework

TPOT (Time Per Output Token) Heatmap

Latency breakdown across batch sizes and sequence lengths. Lower is better.

TPOT Matrix (ms) BS=1 BS=4 BS=8 BS=16 BS=32 BS=64 512 1024 2048 4096 8192 Low Latency High Latency

xTensor Memory Management

Virtual memory system eliminating fragmentation by decoupling logical and physical tensor layouts.

KV Cache Growth Pattern

Unpredictable memory demand as generation progresses makes traditional allocation wasteful or prone to OOM.

Generated Tokens KV Cache Size (MB) 0 500 1000 1500 2000 0 512 1024 1536 2048 Unpredictable Growth Region
~0%
Fragmentation with xTensor
15-30%
Traditional Fragmentation
<2%
Virtual Memory Overhead

References & Related Work