SGLang Inference
From prefix reuse to disaggregated serving
A document assistant receives the same long handbook from hundreds of readers. Which computations can it reuse? Where should the cached state live? And when should the machines reading prompts be separate from the machines generating answers?
This companion follows those questions through SGLang: its radix tree, scheduler, structured generation, and cache hierarchy, then the request handoff between separate prefill and decode pools. Four interactive models let you change the workload and inspect the consequences.
Start here after the basics in LLM Inference. You should recognize attention, tokens, prefill and decode, and the KV cache. No experience operating a GPU cluster is required.
10 chapters · 4 interactive models
Build and evict a radix tree. Overlap CPU and GPU work. Fetch a prefix from a slower cache. Budget a prefill–decode handoff in bytes and milliseconds.
Start reading →~90 minutes, 10 chaptersPublished . Code and deployment examples reference SGLang v0.5.20. The original research and later implementation features are identified separately. Numerical models illustrate tradeoffs; they are not GPU benchmarks.
Contents
Inside the runtime
Disaggregated inference
Putting it to work
Read with the sources
Each chapter links to the relevant documentation, release-pinned code, or paper. The annotated reading guide collects the main sources, including SGLang, DistServe, and Mooncake, and explains what each contributes.