fix(ray): RayExecutorV2 actor handles survive multi-node session init - #259
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Out-of-the-box multi-node Ray TP (
--distributed-executor-backend raywith the defaultVLLM_USE_RAY_V2_EXECUTOR_BACKEND=1) crashed at engine teardown with:This makes
RayExecutorV2.shutdown()fail whenever teardown runs after Ray has been re-initialized under a new job (which Ray's auto-init hook does during interpreter teardown / after any init failure).Root cause (job/session lifecycle)
RayExecutorV2._init_executor()creates the worker actors viaray.remote(RayWorkerProc).options(...).remote(...)and stores the handles inself.ray_worker_handles. Ray actor handles are only valid within the Ray session/job that created them.shutdown()unconditionally calledray.kill(handle.actor)on those handles. When shutdown runs after Ray has been torn down and re-initialized under a different job -- e.g. theweakref.finalize(self, self.shutdown)finalizer or a GC pass firing during interpreter teardown, or cleanup after an init-time failure --ray.kill()receives a handle whose owning job (01000000) no longer matches the current job (02000000) and raisesActorHandleNotFoundError.Instrumenting the job id at handle-creation vs. use confirmed this: the init path (
ray.get(init_worker_refs)etc.) runs entirely within a single job and boots cleanly multi-node; the mismatch only appears at theray.kill()site inshutdown(), after aConnecting to existing Ray cluster ... Connected to Ray clusterreconnect bumps the job id.Attribution: newly-defaulted-and-always-broken, NOT merge-introduced
RayExecutorV2was introduced upstream byde5e6c44c("[Feat][Executor] Introduce RayExecutorV2", PR [Feat][Executor] Introduce RayExecutorV2 vllm-project/vllm#36836) withVLLM_USE_RAY_V2_EXECUTOR_BACKENDdefaulting to"1".de5e6c44cis an ancestor of the pre-merge SHA691649ce9e, so the v2 executor and its on-by-default flag already existed before feat/integration was merged.git log --oneline 691649ce9e..d0ce9c29e -- vllm/v1/executor/ray_executor_v2.pyreturns zero commits, andgit diff 691649ce9e..d0ce9c29e -- vllm/v1/executor/ray_executor_v2.pyis empty. The only commit touchingvllm/envs.pyin that range is unrelated; the flag default was not flipped.Conclusion: the feat/integration merge neither introduced the v2 executor nor changed the default. The bug is inherent to the upstream v2 executor's shutdown path and has been present (and default-on) the whole time.
Fix
Guard actor teardown against stale Ray sessions.
_init_executorrecords the job that owns the handles (self._ray_job_id).shutdown()skipsray.kill()when Ray is not initialized, or when the current job differs from the owning job -- in both cases the original actors already died with their session, so killing them is impossible and unnecessary. Same-session shutdown is unchanged: handles are killed exactly as before.This is the smallest change that makes handle lifetime match session lifetime; single-node Ray and non-Ray paths are untouched.
Validation (GPU, 2x RTX 3090, multi-node TP2)
Exact command:
vllm serve google/gemma-3-4b-it --tensor-parallel-size 2 --distributed-executor-backend ray, defaultVLLM_USE_RAY_V2_EXECUTOR_BACKEND=1.The capital of France is->Paris., 0ActorHandleNotFoundError.The capital of Italy is->Rome., 0ActorHandleNotFoundError. This is the second-boot case the original error's "job N vs N+1" pointed at.VLLM_USE_RAY_V2_EXECUTOR_BACKEND=0): boots and generatesParis.(unaffected).python -c "import vllm.v1.executor.ray_executor_v2"OK. Added_worker_actors_killableunit tests (same-job / stale-job / ray-down / never-created / shutdown-skips-kill-on-stale-job); verified the guard skipsray.kill()on a stale job and still kills on the same job.AI assistance was used for this change.