Section 07

Crossing the prefill–decode boundary

Follow one request and its KV cache

A question enters a disaggregated service through one public endpoint, but two serving workers will execute it. The prefill worker will produce the prompt’s state. The decode worker will consume that state and continue the response. Before either can safely hand off the request, they must agree on which request they mean and where its data belongs.

This chapter follows the logical lifecycle documented in SGLang v0.5.20’s prefill implementation and decode implementation. Backend optimizations can overlap or subdivide operations; the diagram shows dependencies, not a claim that every arrow is a serialized network round trip.

Two kinds of messages

The control planecontrol planeThe coordination path carrying request identities, destinations, allocation metadata, and readiness information.See in glossary → carries request identity, routing choices, endpoint information, and allocation metadata. The data planedata planeThe path that moves or processes bulk payloads, such as the KV tensors transferred between prefill and decode workers.See in glossary → moves the large KV payload. Separating these responsibilities is useful because a successful HTTP request to a worker says little about whether the GPU-to-GPU transfer path is healthy.

For example, an API port may be reachable while a required bootstrap endpoint or transport path is not. The router can accept a request and select healthy-looking workers, yet the request may wait for coordination or state transfer. Observing only the public API hides where the delay occurs.

The question text and sampling configuration also need to agree across the two legs. Model and tokenizer compatibility cannot be repaired by a fast transport. If the workers interpret the history differently, successfully copying bytes does not produce a valid continuation.

1. Select workers and identify the handoff

The gateway selects prefill and decode destinations and supplies routing information that lets the workers coordinate the request. SGLang’s HTTP PD deployment presents a single endpoint while dispatching the two roles internally. Its gateway documentation describes PD routing separately from ordinary routing to interchangeable full-service replicas. Model Gateway

A handoff needs an identity more specific than “send the next request to GPU 1.” Multiple requests can be bootstrapping or transferring simultaneously. The workers must associate destination allocations, source spans, transfer completion, and output metadata with the correct request.

Our handbook question may match a prefix on the selected prefill worker. That reduces the prompt work that worker performs. It does not settle the decode-side allocation question; the destination must still have the state it will read.

2. Bootstrap and reserve destination space

BootstrapPD bootstrapInitial coordination between prefill and decode workers that associates a request with transfer endpoints and destination allocations.See in glossary → is the initial coordination that establishes the handoff. On the decode side, requests enter a preallocation queue. The receiver coordinates with the source and allocates KV space when capacity permits. The prefill side tracks requests whose handshake and preallocation are not yet ready. Decode and prefill lifecycle source

Destination allocation is a form of admission control. A decode worker needs enough memory for the incoming state and must leave room for the generation workload it will run. If it cannot reserve the necessary resources, producing more prefills for it can create a backlog of completed but unusable work.

That backlog consumes resources on both sides. The source may need to retain values until transfer completes, and the destination may hold reserved slots while waiting for data. Counting only requests actively executing GPU kernels misses these in-flight obligations.

3. Perform the missing prefill work

Once a request is eligible, the prefill scheduler admits it according to its budgets. It reuses compatible prefix state where available and computes the missing portion. Chunking and supported transfer optimizations can affect when pieces become available.

The prompt forward pass produces the information needed to choose the first output token. In the ordinary PD flow, the handoff includes output metadata for that boundary token along with the prompt’s KV state. The decode-side code incorporates the committed output token before continuing. It should not rerun the entire prompt merely to recreate what the prefill worker already computed. Decode-side commit handling

Be precise about the cache boundary. The token just sampled is an output of the prompt forward pass. Its own KV is normally produced when that token is fed into a subsequent step. “The first token exists” and “KV for that token exists” are different statements.

4. Move state and wait for readiness

  1. Gateway → both workers: request identity, compatible input, and handoff metadata.
  2. Decode ↔ prefill: bootstrap and destination allocation information.
  3. Prefill GPU: reuse existing state and compute the missing prompt positions.
  4. Transfer backend → decode memory: copy the required KV and associated metadata.
  5. Decode scheduler: observe completion, prepare request metadata, then join a running batch.
  6. Decode → gateway → client: continue the response stream.
Logical dependencies. Bootstrap, queueing, GPU execution, and transfers can overlap across requests; this is not a packet-level protocol trace.

SGLang documents Mooncake and NIXL transfer backends. A transfer engine handles the movement and completion tracking of data through supported infrastructure; it does not, by itself, choose the best model, pool ratio, or application-level latency objective. PD transfer configuration

Mooncake TransferEngine should not be confused with the complete Mooncake serving architecture described in its paper. Similarly, configuring NIXL as the PD transport is distinct from selecting a NIXL-related storage integration. Names can recur at different layers while the ownership and purpose of the bytes remain different.

The destination must not interpret an allocated buffer as ready merely because its address is known. It polls or receives completion information and advances the request only after the required data is usable. The source must also respect transfer lifetime before releasing storage that a pending operation still reads.

5. Admit the continuation into decode

After transfer completion, the request enters the decode-side waiting path. The pinned implementation uses a prebuilt extension path to populate metadata without repeating the prompt forward, then merges the request into ongoing decode work. Decode lifecycle

This final waiting step matters. KV arrival is not the same as an immediate GPU execution slot. Existing batches, running-request limits, and memory pressure can still delay progress. A timeline that stops measuring at “transfer complete” can therefore miss a material part of user-visible latency.

Likewise, the moment the client observes its first token depends on stream handling and gateway behavior. Record client timestamps rather than inferring TTFT solely from a prefill kernel’s end time. Whether a boundary token is forwarded early or after another milestone also affects what a first-token metric hides about the next gap.

Failure is part of the lifecycle

A worker may disappear during bootstrap. A transfer may stall after memory has been reserved. A client may cancel while both legs hold resources. Robust handling must stop or complete in-flight operations safely, release reservations, and avoid sending the request into normal decode with incomplete state.

The PD documentation exposes heartbeat and timeout controls. Longer waiting budgets can accommodate slow work, but they can also retain resources longer after a peer fails. Those settings trade tolerance against recovery speed; increasing every timeout is not a general solution to a congested service. PD timeout settings

A retry is another inference attempt unless the system explicitly provides stronger semantics. It can repeat prefill, consume additional capacity, and—after partial streaming—risk confusing the client if output is silently stitched together. Benchmark and application code should record cancellation, failure, and retry outcomes explicitly.

State travels farther than the HTTP request

For debugging, record milestones: arrival, worker selection, destination allocation, prefill start/end, transfer start/end, decode admission, first client token, and completion. Separate request IDs from worker IDs so that one request can be followed across both processes.

Our handbook service now has several possible bottlenecks that look identical from a spinning browser: no destination capacity, a slow prefill queue, a busy transfer path, or a full decode batch. Locating the delay is the prerequisite for tuning it. The next chapter puts a numerical budget on the largest payload in that handoff.

Sources and further reading