You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When an SGLang server running in decode mode (--disaggregation-mode=decode) it always requires a KV cache transfer from a prefill server. It lacks local prefill support.
When kv cache params are not in the request the SGLang server rejects such requests with an HTTP 400: Invalid request: Disaggregated request received without bootstrap room id.
To summarize: SGLang disaggregation server does not have an option to provide prefill locally like vLLM does. (and it does not have kv cache error policy either)
llm-d sends a request to a decode server alone in the configurations below. In each of them, the request carries the client body and no bootstrap fields:
Path
Configuration that selects decode-only
Requests affected
Coordinator
The conditional-decode step is in the pipeline and prefix-based-pd-decider is in EPP
Requests that the Prefer: if-available gate forwards
Sidecar
disagg-profile-handler with deciders.prefill: prefix-based-pd-decider
Requests for which the decider does not select prefill
Sidecar
prefix-based-pd-decider with nonCachedTokens: 0, which is the default
All requests
Sidecar
disagg-profile-handler with no deciders.prefill
All requests
Sidecar
EPP configuration with no disagg-profile-handler, which is the only plugin that sets x-prefiller-host-port
All requests
The conditional-decode step and prefix-based-pd-decider are optional. The default coordinator configuration file, config/coordinator/coordinator.yaml, has the step as a comment, so the step does not run. A coordinator pipeline without the conditional-decode step, and a sidecar deployment with always-disagg-pd-decider, do not send a decode-only request.
llm-d permits the configurations in the table:
The coordinator accepts pipelines configured with both conditional-decode and kv-sglang.
The sidecar forwards requests lacking a prefill endpoint to the decode server when using the sglang connector.
The prefix-based-pd-decider selects the decode-only path when it cannot read cache state or prompt length. For SGLang, this fallback fails every request.
Proposal
Documentation. Add the limitations:
In docs/disaggregation.md: prefix-based-pd-decider is not supported with SGLang. SGLang uses always-disagg-pd-decider.
In docs/coordinator_architecture.md: the conditional-decode step is not supported with kv-sglang.
Track upstream SGLang support.[DRAFT][PD] Add conditional aggregation to decode workers sgl-project/sglang#34773 (draft) adds a do_local_prefill request field and the decode server flag --disaggregation-decode-enable-conditional-agg, with which a decode server prefills a request locally. The PR does not add the field to the OpenAI routes, which llm-d uses. When SGLang accepts the field on those routes, llm-d can set it on the decode-only path and remove the errors proposed above.
Current SGLang behavior:
When an SGLang server running in decode mode (
--disaggregation-mode=decode) it always requires a KV cache transfer from a prefill server. It lacks local prefill support.When kv cache params are not in the request the SGLang server rejects such requests with an HTTP
400:Invalid request: Disaggregated request received without bootstrap room id.To summarize: SGLang disaggregation server does not have an option to provide prefill locally like vLLM does. (and it does not have kv cache error policy either)
llm-d sends a request to a decode server alone in the configurations below. In each of them, the request carries the client body and no bootstrap fields:
conditional-decodestep is in the pipeline andprefix-based-pd-decideris in EPPPrefer: if-availablegate forwardsdisagg-profile-handlerwithdeciders.prefill: prefix-based-pd-deciderprefix-based-pd-deciderwithnonCachedTokens: 0, which is the defaultdisagg-profile-handlerwith nodeciders.prefilldisagg-profile-handler, which is the only plugin that setsx-prefiller-host-portThe
conditional-decodestep andprefix-based-pd-deciderare optional. The default coordinator configuration file,config/coordinator/coordinator.yaml, has the step as a comment, so the step does not run. A coordinator pipeline without theconditional-decodestep, and a sidecar deployment withalways-disagg-pd-decider, do not send a decode-only request.llm-d permits the configurations in the table:
conditional-decodeandkv-sglang.sglangconnector.prefix-based-pd-deciderselects the decode-only path when it cannot read cache state or prompt length. For SGLang, this fallback fails every request.Proposal
docs/disaggregation.md:prefix-based-pd-decideris not supported with SGLang. SGLang usesalways-disagg-pd-decider.docs/coordinator_architecture.md: theconditional-decodestep is not supported withkv-sglang.do_local_prefillrequest field and the decode server flag--disaggregation-decode-enable-conditional-agg, with which a decode server prefills a request locally. The PR does not add the field to the OpenAI routes, which llm-d uses. When SGLang accepts the field on those routes, llm-d can set it on the decode-only path and remove the errors proposed above.Related: sgl-project/sglang#34773.
/kind feature
/area epp
/area coordinator
/area sglang