Skip to content

fix(ray): RayExecutorV2 actor handles survive multi-node session init - #259

Merged
RhizoNymph merged 1 commit into
feat/dynamic-steeringfrom
fix/ray-v2-multinode-init
Jul 8, 2026
Merged

RhizoNymph merged 1 commit into
feat/dynamic-steeringfrom
fix/ray-v2-multinode-init

Conversation

@RhizoNymph

Copy link
Copy Markdown
Owner

Summary

Out-of-the-box multi-node Ray TP (--distributed-executor-backend ray with the default VLLM_USE_RAY_V2_EXECUTOR_BACKEND=1) crashed at engine teardown with:

ray.exceptions.ActorHandleNotFoundError: ActorHandle objects are not valid across
Ray sessions. The actor handle was created in job 01000000, but the current job is
02000000. ... after calling ray.shutdown() and ray.init().

This makes RayExecutorV2.shutdown() fail whenever teardown runs after Ray has been re-initialized under a new job (which Ray's auto-init hook does during interpreter teardown / after any init failure).

Root cause (job/session lifecycle)

RayExecutorV2._init_executor() creates the worker actors via ray.remote(RayWorkerProc).options(...).remote(...) and stores the handles in self.ray_worker_handles. Ray actor handles are only valid within the Ray session/job that created them.

shutdown() unconditionally called ray.kill(handle.actor) on those handles. When shutdown runs after Ray has been torn down and re-initialized under a different job -- e.g. the weakref.finalize(self, self.shutdown) finalizer or a GC pass firing during interpreter teardown, or cleanup after an init-time failure -- ray.kill() receives a handle whose owning job (01000000) no longer matches the current job (02000000) and raises ActorHandleNotFoundError.

Instrumenting the job id at handle-creation vs. use confirmed this: the init path (ray.get(init_worker_refs) etc.) runs entirely within a single job and boots cleanly multi-node; the mismatch only appears at the ray.kill() site in shutdown(), after a Connecting to existing Ray cluster ... Connected to Ray cluster reconnect bumps the job id.

Attribution: newly-defaulted-and-always-broken, NOT merge-introduced

  • RayExecutorV2 was introduced upstream by de5e6c44c ("[Feat][Executor] Introduce RayExecutorV2", PR [Feat][Executor] Introduce RayExecutorV2 vllm-project/vllm#36836) with VLLM_USE_RAY_V2_EXECUTOR_BACKEND defaulting to "1".
  • de5e6c44c is an ancestor of the pre-merge SHA 691649ce9e, so the v2 executor and its on-by-default flag already existed before feat/integration was merged.
  • git log --oneline 691649ce9e..d0ce9c29e -- vllm/v1/executor/ray_executor_v2.py returns zero commits, and git diff 691649ce9e..d0ce9c29e -- vllm/v1/executor/ray_executor_v2.py is empty. The only commit touching vllm/envs.py in that range is unrelated; the flag default was not flipped.

Conclusion: the feat/integration merge neither introduced the v2 executor nor changed the default. The bug is inherent to the upstream v2 executor's shutdown path and has been present (and default-on) the whole time.

Fix

Guard actor teardown against stale Ray sessions. _init_executor records the job that owns the handles (self._ray_job_id). shutdown() skips ray.kill() when Ray is not initialized, or when the current job differs from the owning job -- in both cases the original actors already died with their session, so killing them is impossible and unnecessary. Same-session shutdown is unchanged: handles are killed exactly as before.

This is the smallest change that makes handle lifetime match session lifetime; single-node Ray and non-Ray paths are untouched.

Validation (GPU, 2x RTX 3090, multi-node TP2)

Exact command: vllm serve google/gemma-3-4b-it --tensor-parallel-size 2 --distributed-executor-backend ray, default VLLM_USE_RAY_V2_EXECUTOR_BACKEND=1.

  • v2 boot 1 (fresh cluster): "Application startup complete", greedy completion The capital of France is -> Paris., 0 ActorHandleNotFoundError.
  • v2 boot 2 (same ray cluster, job N+1): re-booted on the same cluster after killing boot 1; "Application startup complete", The capital of Italy is -> Rome., 0 ActorHandleNotFoundError. This is the second-boot case the original error's "job N vs N+1" pointed at.
  • classic executor (VLLM_USE_RAY_V2_EXECUTOR_BACKEND=0): boots and generates Paris. (unaffected).
  • CPU: python -c "import vllm.v1.executor.ray_executor_v2" OK. Added _worker_actors_killable unit tests (same-job / stale-job / ray-down / never-created / shutdown-skips-kill-on-stale-job); verified the guard skips ray.kill() on a stale job and still kills on the same job.

AI assistance was used for this change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant