[Bugfix] Set breakable graph env before Ray actor import - #53293
Conversation
DSv4 auto-enables breakable CUDA graphs in VllmConfig. MP workers inherit that setting before model import, but Ray creates actors first and copies the driver environment only later in initialize_worker. By then eager_break_during_capture has already returned undecorated attention and indexer functions. Ray therefore enables breakable capture without its required breakpoints, producing an oversized graph and garbled output. Copy only VLLM_USE_BREAKABLE_CUDAGRAPH into actor runtime_env before import; all other environment propagation is unchanged. In two-replica, two-host TP8 GB300 serving, the Ray baseline moved from 0/10 to 10/10 deterministic checks and matched MP's 0.37 GiB graph. Co-authored-by: Codex <codex@openai.com> Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
|
Thanks @alexeldeib this looks reasonable to me but it would be good to get approval from someone more familiar with the ray side, cc @jeffreywang88 |
There was a problem hiding this comment.
@alexeldeib thanks for identifying and fixing the bug! Couple questions:
- I'm wondering if we could make this more general. For example, maintain a list of environment variables that should always be set before actor creation, and add
VLLM_USE_BREAKABLE_CUDAGRAPHto that list. - Have you identified other env vars that need the same treatment?
Note for future reviewers
Both mechanisms from the existing and this PR copy the environment variable from the driver. The differences are:
-
Mechanism 1 (general): The driver passes the environment variables to
initialize_worker, which applies them inside the already-running actor. This is intentionally done late because some values aren't known when the actor is created.CUDA_VISIBLE_DEVICES,local_rank, and the GPU mapping all depend on where Ray places the actor, which is only known after scheduling. -
Mechanism 2 (this PR): The variable is put in
runtime_envwhen the actor is created, so Ray sets it before python imports vLLM.
This timing matters because @eager_break_during_capture reads VLLM_USE_BREAKABLE_CUDAGRAPH at import time. With mechanism 1, the variable is set only after vLLM has already been imported, so it's too late. Mechanism 2 solves that by setting the variable early, while leaving placement-dependent variables to be set later, once Ray knows where the actor is running.
yeah, I initially did it as a frozenset and iterate over the list. but since I only found 1 that mattered, I avoided adding more. the logic could be easily adjusted for sure.
I have not, so didn't add any more per above |
|
@alexeldeib let's go ahead with the frozenset approach to be more general and easier to extend to other env vars going forward. |
Maintain a frozen set of driver environment variables that Ray workers need before importing vLLM. Keep breakable CUDA graphs as the only member and use the shared helper in both Ray executors. Co-authored-by: Codex <codex@openai.com> Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
|
@jeffreywang88 adjusted, lmk how that looks to you |
njhill
left a comment
There was a problem hiding this comment.
Thanks @alexeldeib @jeffreywang88
|
/ci run |
|
✅ @alexeldeib, CI is now available for this PR.
|
|
✅ Triggered Buildkite CI #86040 for commit |
…t#53293) Signed-off-by: Ace Eldeib <aeldeib@coreweave.com> Co-authored-by: Codex <codex@openai.com>
…t#53293) Signed-off-by: Ace Eldeib <aeldeib@coreweave.com> Co-authored-by: Codex <codex@openai.com>
Purpose
DSv4 automatically enables breakable CUDA graphs in
VllmConfig. Nativemultiprocessing workers inherit that setting before importing the model, but
Ray creates its worker actors first and copies the driver environment only later
in
initialize_worker.That ordering matters because
@eager_break_during_capturemakes its decisionwhen the model module imports. Without the setting, the decorator permanently
returns the original attention and indexer functions. Ray later enables
breakable capture, but the breakpoints needed to keep those functions outside
the graph are missing, producing an oversized graph and garbled output.
Copy only
VLLM_USE_BREAKABLE_CUDAGRAPHinto the actorruntime_envbeforeimport, for both Ray executors. All existing general environment propagation,
worker exclusions, and node-local values remain unchanged. This follows the
existing pre-actor environment pattern used for Ray device visibility in
#33308.
No open PR duplicates this Ray import-order fix; #41834 concerns unrelated
SM12x model support. AI assistance was used to investigate and prepare this
patch; I reviewed every changed line and the reported test results.
Test Plan
Run the focused Ray environment tests:
Validate DeepSeek-V4-Pro with Ray V2 using two replicas, each TP8 across two
physical GB300 hosts. Retain production CUDA graphs, expert parallelism, MNNVL,
and custom collectives, and do not explicitly set
VLLM_USE_BREAKABLE_CUDAGRAPHso the test exercises automatic enablement.Send 10 deterministic color ("what is the color of the sky, 1 word answer") and 10 deterministic simple math requests to each backend and through the normal router. Repeat after 128 requests at concurrency 16.
Test Result
Both patched TP8 replicas passed the color and math checks before and after
load. The router also passed both checks, and the concurrency-16 run completed
128/128 requests successfully.
Essential Elements of an Effective PR Description Checklist