Skip to content

SGLang does not support conditional aggregation to decode workers #3075

Description

@revit13

Current SGLang behavior:

When an SGLang server running in decode mode (--disaggregation-mode=decode) it always requires a KV cache transfer from a prefill server. It lacks local prefill support.
When kv cache params are not in the request the SGLang server rejects such requests with an HTTP 400: Invalid request: Disaggregated request received without bootstrap room id.

To summarize: SGLang disaggregation server does not have an option to provide prefill locally like vLLM does. (and it does not have kv cache error policy either)

llm-d sends a request to a decode server alone in the configurations below. In each of them, the request carries the client body and no bootstrap fields:

Path Configuration that selects decode-only Requests affected
Coordinator The conditional-decode step is in the pipeline and prefix-based-pd-decider is in EPP Requests that the Prefer: if-available gate forwards
Sidecar disagg-profile-handler with deciders.prefill: prefix-based-pd-decider Requests for which the decider does not select prefill
Sidecar prefix-based-pd-decider with nonCachedTokens: 0, which is the default All requests
Sidecar disagg-profile-handler with no deciders.prefill All requests
Sidecar EPP configuration with no disagg-profile-handler, which is the only plugin that sets x-prefiller-host-port All requests

The conditional-decode step and prefix-based-pd-decider are optional. The default coordinator configuration file, config/coordinator/coordinator.yaml, has the step as a comment, so the step does not run. A coordinator pipeline without the conditional-decode step, and a sidecar deployment with always-disagg-pd-decider, do not send a decode-only request.

llm-d permits the configurations in the table:

  • The coordinator accepts pipelines configured with both conditional-decode and kv-sglang.
  • The sidecar forwards requests lacking a prefill endpoint to the decode server when using the sglang connector.
  • The prefix-based-pd-decider selects the decode-only path when it cannot read cache state or prompt length. For SGLang, this fallback fails every request.

Proposal

  1. Documentation. Add the limitations:
    • In docs/disaggregation.md: prefix-based-pd-decider is not supported with SGLang. SGLang uses always-disagg-pd-decider.
    • In docs/coordinator_architecture.md: the conditional-decode step is not supported with kv-sglang.
  2. Track upstream SGLang support. [DRAFT][PD] Add conditional aggregation to decode workers sgl-project/sglang#34773 (draft) adds a do_local_prefill request field and the decode server flag --disaggregation-decode-enable-conditional-agg, with which a decode server prefills a request locally. The PR does not add the field to the OpenAI routes, which llm-d uses. When SGLang accepts the field on those routes, llm-d can set it on the decode-only path and remove the errors proposed above.

Related: sgl-project/sglang#34773.

/kind feature
/area epp
/area coordinator
/area sglang

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/coordinatorarea/eppkind/featureCategorizes issue or PR as related to a new feature.needs-triageIndicates an issue or PR lacks a triage label and requires one.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions