Skip to content

[Bugfix] Set breakable graph env before Ray actor import - #53293

Merged
njhill merged 5 commits into
vllm-project:mainfrom
alexeldeib:ace/breakable-cudagraph-ray-rootcause
Aug 28, 2026
Merged

njhill merged 5 commits into
vllm-project:mainfrom
alexeldeib:ace/breakable-cudagraph-ray-rootcause

Conversation

@alexeldeib

@alexeldeib alexeldeib commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

DSv4 automatically enables breakable CUDA graphs in VllmConfig. Native
multiprocessing workers inherit that setting before importing the model, but
Ray creates its worker actors first and copies the driver environment only later
in initialize_worker.

That ordering matters because @eager_break_during_capture makes its decision
when the model module imports. Without the setting, the decorator permanently
returns the original attention and indexer functions. Ray later enables
breakable capture, but the breakpoints needed to keep those functions outside
the graph are missing, producing an oversized graph and garbled output.

Copy only VLLM_USE_BREAKABLE_CUDAGRAPH into the actor runtime_env before
import, for both Ray executors. All existing general environment propagation,
worker exclusions, and node-local values remain unchanged. This follows the
existing pre-actor environment pattern used for Ray device visibility in
#33308.

No open PR duplicates this Ray import-order fix; #41834 concerns unrelated
SM12x model support. AI assistance was used to investigate and prepare this
patch; I reviewed every changed line and the reported test results.

Test Plan

Run the focused Ray environment tests:

.venv/bin/python -m pytest tests/test_ray_env_utils.py -q

Validate DeepSeek-V4-Pro with Ray V2 using two replicas, each TP8 across two
physical GB300 hosts. Retain production CUDA graphs, expert parallelism, MNNVL,
and custom collectives, and do not explicitly set
VLLM_USE_BREAKABLE_CUDAGRAPH so the test exercises automatic enablement.

Send 10 deterministic color ("what is the color of the sky, 1 word answer") and 10 deterministic simple math requests to each backend and through the normal router. Repeat after 128 requests at concurrency 16.

Test Result

9 passed in 0.83s
Variant Captured graph Color Math
Unpatched Ray, automatic enablement 0.69 GiB 0/10 0/10
Ray, explicit setting at actor startup 0.37 GiB 10/10 10/10
Native MP, automatic enablement 0.37 GiB 10/10 10/10
Patched Ray, automatic enablement 0.37 GiB 10/10 10/10

Both patched TP8 replicas passed the color and math checks before and after
load. The router also passed both checks, and the concurrency-16 run completed
128/128 requests successfully.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results.
  • Documentation impact was considered; no user-facing option or API changes require documentation.
DSv4 auto-enables breakable CUDA graphs in VllmConfig. MP workers inherit that
setting before model import, but Ray creates actors first and copies the driver
environment only later in initialize_worker.

By then eager_break_during_capture has already returned undecorated attention
and indexer functions. Ray therefore enables breakable capture without its
required breakpoints, producing an oversized graph and garbled output.

Copy only VLLM_USE_BREAKABLE_CUDAGRAPH into actor runtime_env before import; all
other environment propagation is unchanged. In two-replica, two-host TP8 GB300
serving, the Ray baseline moved from 0/10 to 10/10 deterministic checks and
matched MP's 0.37 GiB graph.

Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
@alexeldeib
alexeldeib requested a review from njhill as a code owner August 21, 2026 16:37

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added bug Something isn't working ray anything related with ray labels Aug 21, 2026
@github-project-automation github-project-automation Bot moved this to Backlog in Ray Aug 21, 2026
@njhill

njhill commented Aug 27, 2026

Copy link
Copy Markdown
Member

Thanks @alexeldeib this looks reasonable to me but it would be good to get approval from someone more familiar with the ray side, cc @jeffreywang88

@jeffreywang88 jeffreywang88 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@alexeldeib thanks for identifying and fixing the bug! Couple questions:

  • I'm wondering if we could make this more general. For example, maintain a list of environment variables that should always be set before actor creation, and add VLLM_USE_BREAKABLE_CUDAGRAPH to that list.
  • Have you identified other env vars that need the same treatment?

Note for future reviewers

Both mechanisms from the existing and this PR copy the environment variable from the driver. The differences are:

  • Mechanism 1 (general): The driver passes the environment variables to initialize_worker, which applies them inside the already-running actor. This is intentionally done late because some values aren't known when the actor is created. CUDA_VISIBLE_DEVICES, local_rank, and the GPU mapping all depend on where Ray places the actor, which is only known after scheduling.

  • Mechanism 2 (this PR): The variable is put in runtime_env when the actor is created, so Ray sets it before python imports vLLM.

This timing matters because @eager_break_during_capture reads VLLM_USE_BREAKABLE_CUDAGRAPH at import time. With mechanism 1, the variable is set only after vLLM has already been imported, so it's too late. Mechanism 2 solves that by setting the variable early, while leaving placement-dependent variables to be set later, once Ray knows where the actor is running.

FYI @eicherseiji @kouroshHakha

@alexeldeib

Copy link
Copy Markdown
Contributor Author

For example, maintain a list of environment variables that should always be set before actor creation, and add VLLM_USE_BREAKABLE_CUDAGRAPH to that list.

yeah, I initially did it as a frozenset and iterate over the list. but since I only found 1 that mattered, I avoided adding more. the logic could be easily adjusted for sure.

Have you identified other env vars that need the same treatment?

I have not, so didn't add any more per above

@jeffreywang88

jeffreywang88 commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

@alexeldeib let's go ahead with the frozenset approach to be more general and easier to extend to other env vars going forward.

Maintain a frozen set of driver environment variables that Ray workers need before importing vLLM. Keep breakable CUDA graphs as the only member and use the shared helper in both Ray executors.

Co-authored-by: Codex <codex@openai.com>

Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
@alexeldeib

Copy link
Copy Markdown
Contributor Author

@jeffreywang88 adjusted, lmk how that looks to you

@jeffreywang88 jeffreywang88 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM thank you!

@njhill njhill left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 28, 2026
@njhill

njhill commented Aug 28, 2026

Copy link
Copy Markdown
Member

/ci run

@njhill
njhill enabled auto-merge (squash) August 28, 2026 21:15
@github-actions

Copy link
Copy Markdown

✅ @alexeldeib, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.
@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #86040 for commit 995ab8fd43a4.

@njhill
njhill merged commit 9662ab0 into vllm-project:main Aug 28, 2026
85 checks passed
@github-project-automation github-project-automation Bot moved this from Backlog to Done in Ray Aug 28, 2026
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
…t#53293)

Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
Co-authored-by: Codex <codex@openai.com>
mylibrar pushed a commit to tanyuqian/vllm that referenced this pull request Sep 3, 2026
…t#53293)

Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
Co-authored-by: Codex <codex@openai.com>
D-G-Dimitrov pushed a commit to D-G-Dimitrov/vllm that referenced this pull request Sep 4, 2026
…t#53293)

Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
Co-authored-by: Codex <codex@openai.com>
(cherry picked from commit 9662ab0)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ray anything related with ray ready ONLY add when PR is ready to merge/full CI is needed

3 participants