SGLang treats LLM applications as executable programs with branches, joins, tool calls, cacheable prefixes, and hard output contracts. This page visualizes where the speedups come from: runtime-aware prefix reuse, grammar-constrained decoding, and scheduling over whole workflows rather than isolated prompt-response calls.
The optimization unit is the LM program, not the single token stream. Once branching and schemas are explicit, the runtime can cache, schedule, and constrain execution with intent.
Most of this page is diagrams: program DAGs, reuse heatmaps, grammar-state flows, and workload-sensitive benchmark views. The text only orients the visuals.
The episode positions SGLang between language systems like Guidance, LMQL, DSPy and serving systems like vLLM, TGI, and TensorRT-LLM.
The workflow below shows the paper’s shift: from one prompt call to a structured LM program with branches, tools, joins, and constrained output.
Python stays in charge of control flow, but the runtime sees the structure instead of opaque SDK spaghetti.
Shared scaffolds dominate many LM programs. Toggle the view to compare repeated prefill work against cached prefix reuse across diverging branches.
The matrix encodes token positions by branch. Dark blue means unique compute; orange-red means repeated work that SGLang can collapse.
JSON is not just a prompt wish. The runtime narrows valid next tokens through grammar states so the model does less wandering and fewer invalid retries.
This micro-chart shows how constrained decoding shifts probability mass from “any token” to “schema-valid next token” over generation steps.
Throughput wins vary with workload structure. Heavy prefix sharing and strict schemas light up the runtime; chaotic low-overlap traffic should compress the gain.
The heatmap estimates where the design should matter most across reuse, grammar rigidity, and branching pressure.