SGLang Inference

From prefix reuse to disaggregated serving

A document assistant receives the same long handbook from hundreds of readers. Which computations can it reuse? Where should the cached state live? And when should the machines reading prompts be separate from the machines generating answers?

This companion follows those questions through SGLang: its radix tree, scheduler, structured generation, and cache hierarchy, then the request handoff between separate prefill and decode pools. Four interactive models let you change the workload and inspect the consequences.

Start here after the basics in LLM Inference. You should recognize attention, tokens, prefill and decode, and the KV cache. No experience operating a GPU cluster is required.

10 chapters · 4 interactive models

Build and evict a radix tree. Overlap CPU and GPU work. Fetch a prefix from a slower cache. Budget a prefill–decode handoff in bytes and milliseconds.

Start reading →~90 minutes, 10 chapters

Published . Code and deployment examples reference SGLang v0.5.20. The original research and later implementation features are identified separately. Numerical models illustrate tradeoffs; they are not GPU benchmarks.

Contents

Inside the runtime

  1. 01What SGLang optimizesFrom language model programs to a serving engine
  2. 02RadixAttentionFind and reuse the shared prefix
  3. 03Scheduling around reuseKeep the CPU and GPU working together
  4. 04Structured outputs and faster decodingConstraints, drafts, and verification
  5. 05HiCacheReuse beyond GPU memory

Disaggregated inference

  1. 06Why separate prefill and decode?Interference, specialization, and goodput
  2. 07Crossing the prefill–decode boundaryFollow one request and its KV cache
  3. 08The cost of moving KVBytes, bandwidth, and topology
  4. 09Routing and balancing the serviceCache affinity meets queue pressure

Putting it to work

  1. 10Putting the system togetherDeploy, measure, and choose

Read with the sources

Each chapter links to the relevant documentation, release-pinned code, or paper. The annotated reading guide collects the main sources, including SGLang, DistServe, and Mooncake, and explains what each contributes.