Section 06

Why separate prefill and decode?

Interference, specialization, and goodput

The assistant is generating a response at a comfortable rhythm. Then somebody uploads a long new handbook. The server must process that prompt, and the ongoing response pauses. Total throughput may still look healthy: many prompt tokens were processed during the pause. The person waiting for the next word experiences something different.

Prefill–decode disaggregationprefill–decode disaggregationExecuting prompt processing and ongoing token generation on separate workers, with a handoff of KV state between them.See in glossary →, usually abbreviated PD, assigns the two phases to separate serving workers. A prefill worker computes the state needed to start an answer. A decode worker receives that state and continues generation. SGLang implements this deployment pattern with explicit worker roles and KV-transfer backends. SGLang PD documentation

Two phases, different shapes of work

During prefill, many known positions can be processed together. During ordinary decode, each active sequence contributes its next position, and that position depends on the previously generated token. These shapes produce different opportunities for parallelism and memory reuse.

The common shorthand is “prefill is compute-bound, decode is memory-bound.” It is a useful starting intuition, not a law. Batch size, context length, model architecture, quantization, and communication can change the limiting resource. A large decode batch can use considerable compute; long-context attention can add substantial KV traffic. Our architectural argument is that the two phases have different operating points, not that one hardware resource is always the sole bottleneck.

A shared engine chooses settings that accommodate both. A large prefill batch may improve prompt-processing efficiency while making ongoing decoders wait. A setting that protects token cadence may restrict the prompt work admitted each round. Separating the workers allows each pool to choose budgets and parallelism suited to its phase, within the implementation’s compatibility constraints.

What chunking already solves

Before adding another deployment boundary, compare against a well-tuned colocated baseline. Chunked prefill can reduce the longest uninterrupted prompt step. Continuous batching can keep resources useful as requests enter and leave. Cache reuse can remove much of the prompt work in the first place.

Consider a synthetic long prefill that takes 160 milliseconds. If it blocks ongoing decode for the full duration, a user may see a conspicuous pause. Splitting it into eight pieces could create more frequent opportunities for other work. The total prefill may take longer because of repeated overhead or less efficient kernels, but the worst interruption can shrink.

Disaggregation takes the next step: prefill no longer competes for the same worker’s execution slots as decode. It adds a new task, however—moving state between the workers. The right comparison includes both the benefit of isolation and the transfer and queueing that isolation introduces.

ArrangementMain opportunityMain cost to examine
Colocated, unchunkedSimple execution and local KVLong prefill interruptions
Colocated, chunkedShorter scheduling intervalsRepeated overhead and shared resource contention
Separate prefill/decodePhase-specific capacity and schedulingKV handoff, extra queues, and pool imbalance

Latency objectives define useful work

A service might promise that an answer begins within one second and continues with short gaps. Time to first token, or TTFT, includes waiting before generation begins. Inter-token latency, or ITL, measures gaps during the stream. Time per output token, or TPOT, often summarizes a request’s post-first-token duration divided by subsequent token count; it does not reveal every individual pause.

A service-level objectiveservice-level objective (SLO)A defined service target, such as a bound on a specified percentile of first-token or inter-token latency.See in glossary → is a specified target, such as a percentile bound or a per-request deadline. State the exact definition. “P99 ITL below 50 milliseconds” and “every token gap below 50 milliseconds for 99% of requests” are different requirements because they aggregate samples differently.

GoodputgoodputThe rate of useful completed work meeting explicitly stated requirements, such as both first-token and streaming-latency objectives.See in glossary → counts useful completed work under the chosen constraints. For an illustrative benchmark, call a request successful only if it completes, its TTFT is at most one second, and every measured token gap is at most 80 milliseconds. If 100 requests finish per second but only 70 meet both latency requirements, the request goodput is 70 per second. Those thresholds are our example, not SGLang defaults.

DistServe develops this goodput-oriented argument for separating prefill and decode, including phase-specific resource allocation and bandwidth-aware placement. It is a related research foundation for understanding the tradeoff. We do not need to claim that every SGLang design choice is directly copied from that paper. DistServe, OSDI 2024

A small timeline explains the appeal

Imagine two decode requests that need a GPU step roughly every 20 milliseconds. A new prompt arrives and occupies the same worker for 100 milliseconds. If it is executed without useful interleaving, both old requests lose several opportunities to advance. Their prompts were already complete, yet someone else’s prompt length changes their token cadence.

Now put that new request on a prefill worker. Existing decode work can continue on its own pool. After prefill, suppose the new request pays a 15-millisecond handoff and joins the decode queue. That request’s first token might arrive later than it would on an otherwise idle colocated worker. The old requests may nevertheless have much better latency. The gain is visible at the service level, not necessarily in a single isolated request.

This example deliberately omits queueing detail. In a real service, the prefill pool might already be busy, or the decode pool might lack space. The isolation of phases removes one source of interference but does not guarantee immediate service at either stage.

Separate pools can be the wrong size

Suppose a deployment devotes most GPUs to decoding because outputs are usually long. Tomorrow, users start uploading much larger documents but asking for short answers. Prefill demand grows while decode demand shrinks. An idle decode GPU does not automatically execute the work queued in a separate prefill process.

The same problem appears during bursts and failures. A failed prefill worker reduces prompt capacity even when the decode pool has spare resources. Scaling a pool involves loading weights, establishing communication, and warming caches; it is not instantaneous. Operational decisions should reflect the timescale on which demand changes.

Additional replicas also consume weight memory. A single worker may have kept one model copy and flexibly alternated phases. Splitting roles can require another model instance and its runtime buffers. If the original workload fits comfortably on one GPU, using two GPUs merely to demonstrate PD does not prove better cost efficiency.

The handoff must pay for itself

Let the old system’s problematic waiting time be the delay caused by competing prefill work. The new system trades some of that delay for transfer time, coordination, and separate-pool queues. Disaggregation is attractive when the removed interference and improved phase efficiency outweigh those new costs under the actual latency objectives.

Large prompts make the decision less obvious than “longer means better.” They increase prefill work, which can strengthen the case for separation. They also create more KV to transfer, which can weaken it. More compact KV representations, stronger interconnects, and effective reuse change the balance.

High prefix reuse similarly has two sides. It reduces the prefill compute that needs isolation, while the decode stage still needs the cached history. If the destination has no compatible state, the payload can remain large despite a cheap prefill. We will quantify this distinction in chapter 8.

Choose the experiment before the architecture

For the handbook service, first record input lengths, output lengths, prefix reuse, arrival bursts, and latency targets. Establish a colocated baseline with sensible chunking. Then compare PD using the same model, workload, precision, and total hardware budget.

Measure each pool’s utilization and queue depth alongside the client-visible stream. A throughput increase without acceptable token cadence may not help the product. A lower TTFT achieved by rejecting most requests is also incomplete unless rejection rate is reported. The goal is to establish how much demand the system serves well, and what resources that requires.

The next chapter follows the actual handoff. Once a request’s KV leaves one worker for another, ownership, allocation, and readiness become as important as arithmetic speed.

Sources and further reading