xLLM Two-Layer Architecture
The system separates high-level scheduling intelligence from low-level hardware optimization through a clean abstraction boundary.
1.7-2.2x
Throughput vs Baselines
8-GPU
Tested Cluster Size
Ascend
Target Hardware (NPU)
Dynamic Prefill-Decode Disaggregation
Real-time resource reallocation between compute-bound prefill and memory-bound decode phases based on workload characteristics.
Resource Profile Comparison
Online-Offline Workload Co-Location
Maximize GPU utilization during traffic valleys by packing offline batch jobs without violating online SLOs.
Utilization Metrics by Traffic Pattern
Performance Analysis
Throughput comparison on Ascend NPUs running Qwen and DeepSeek models at 8-GPU scale.
TPOT (Time Per Output Token) Heatmap
Latency breakdown across batch sizes and sequence lengths. Lower is better.
xTensor Memory Management
Virtual memory system eliminating fragmentation by decoupling logical and physical tensor layouts.
KV Cache Growth Pattern
Unpredictable memory demand as generation progresses makes traditional allocation wasteful or prone to OOM.
~0%
Fragmentation with xTensor
15-30%
Traditional Fragmentation
<2%
Virtual Memory Overhead