Skip to content

[BugFix] Support online dense model DP without overhead - #30739

Merged
youkaichao merged 7 commits into
vllm-project:mainfrom
njhill:non-moe-dp
Jan 2, 2026
Merged

youkaichao merged 7 commits into
vllm-project:mainfrom
njhill:non-moe-dp

Conversation

@njhill

@njhill njhill commented Dec 16, 2025 •

Copy link
Copy Markdown
Member

Currently, there's unnecessary overhead when running non-MoE models in a data parallel configuration because the steps across the ranks are synchronized with redundant all-reduce ops and coordination is done to ensure "idle" ranks perform dummy forward passes.

This PR changes the parallel config at the worker level to be equivalent to DP=1 for non-MoE models, so each rank operates independently. When internal load-balancing is used, the DP coordinator still runs to propagate stats back from the engines for load balancing purposes, but the step/wave synchronization logic is disabled.

Fixes #24461.
Fixes #30655.

This is supported in the online / AsyncLLM case only.

The offline DP will now fail during startup for non-MoE models (it really makes no sense to use it in that configuration).

Benchmark on 4xH100:

vllm serve Qwen/Qwen3-8B --data-parallel-size 4 --uvicorn-log-level=error
vllm bench serve \
    --backend vllm \
    --model Qwen/Qwen3-8B \
    --dataset-name random \
    --random-input-len 128 \
    --random-output-len 512 \
    --ignore-eos \
    --port 8033 \
    --num-prompts 4000 \
    --max-concurrency 200 \
    --seed 42

Before

============ Serving Benchmark Result ============
Successful requests:                     4000      
Failed requests:                         0         
Maximum request concurrency:             200       
Benchmark duration (s):                  104.41    
Total input tokens:                      512000    
Total generated tokens:                  2048000   
Request throughput (req/s):              38.31     
Output token throughput (tok/s):         19615.26  
Peak output token throughput (tok/s):    21597.00  
Peak concurrent requests:                400.00    
Total token throughput (tok/s):          24519.08  
---------------Time to First Token----------------
Mean TTFT (ms):                          131.78    
Median TTFT (ms):                        124.31    
P99 TTFT (ms):                           404.36    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          9.93      
Median TPOT (ms):                        9.95      
P99 TPOT (ms):                           10.08     
---------------Inter-token Latency----------------
Mean ITL (ms):                           9.93      
Median ITL (ms):                         9.79      
P99 ITL (ms):                            15.93     
==================================================

After

============ Serving Benchmark Result ============
Successful requests:                     4000      
Failed requests:                         0         
Maximum request concurrency:             200       
Benchmark duration (s):                  99.24     
Total input tokens:                      512000    
Total generated tokens:                  2048000   
Request throughput (req/s):              40.31     
Output token throughput (tok/s):         20636.52  
Peak output token throughput (tok/s):    22454.00  
Peak concurrent requests:                400.00    
Total token throughput (tok/s):          25795.66  
---------------Time to First Token----------------
Mean TTFT (ms):                          88.94     
Median TTFT (ms):                        74.67     
P99 TTFT (ms):                           379.48    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          9.50      
Median TPOT (ms):                        9.50      
P99 TPOT (ms):                           9.66      
---------------Inter-token Latency----------------
Mean ITL (ms):                           9.50      
Median ITL (ms):                         9.38      
P99 ITL (ms):                            12.86     
==================================================
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant optimization for running dense (non-MoE) models in a data-parallel configuration by removing unnecessary synchronization overhead. The core idea is to treat each data-parallel rank as an independent worker for dense models, effectively setting their data-parallel size to 1 at the worker level. This avoids redundant all-reduce operations and complex wave synchronization, which are only necessary for MoE models. The DP coordinator's role is intelligently adapted: for dense models with internal load balancing, it continues to run for statistics propagation, but with wave coordination disabled. For external load balancing, it's disabled entirely for dense models. The changes are well-structured, with clear separation of concerns. The introduction of data_parallel_index to preserve the original rank is a clean solution. The related configurations and tests, especially the new test_needs_dp_coordination, are thorough and correctly validate the new logic. Overall, this is a solid improvement that should enhance performance for a common use case.

@njhill njhill mentioned this pull request Dec 16, 2025
1 task done
@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Dec 16, 2025
Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: Nick Hill <nhill@redhat.com>
aipaes pushed a commit to aipaes/vllm-ascend that referenced this pull request Jan 15, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
akh64bit pushed a commit to akh64bit/vllm that referenced this pull request Jan 16, 2026
…#30739)

Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: njhill <nickhill123@gmail.com>
jeffreywang88 pushed a commit to nrghosh/ray that referenced this pull request Jan 26, 2026
- Use a MoE model (Deepseek-V2-Lite) because
vllm-project/vllm#30739 changes how vLLM handles
DP ranks - overrides dp_size=1 and dp_rank=0 if non-MoE model

- Fixes doc/source/llm/doc_code/serve/multi_gpu/dp_basic_example.py and
 doc/source/llm/doc_code/serve/multi_gpu/dp_pd_example.py

- vLLM 0.14.0 commit bd877162e optimizes DP for dense models by making each rank independent and only preserving DP coordination for MoE models where it's needed for expert

- Impact: Ray's DPServer DP coordination (rank assignment, stats addresses) was ignored for dense models like Qwen2.5-0.5B-Instruct, causing cascading assertion failures

- Fix: The tests now use an MoE model where vLLM's DP coordination is preserved. Outside of this test, dense model deployments should use Ray Serve replicas (num_replicas) instead of vLLM's data_parallel_size.

Signed-off-by: Nikhil Ghosh <nikhil@anyscale.com>
maoxx241 pushed a commit to maoxx241/vllm-ascend that referenced this pull request Mar 2, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
LCAIZJ pushed a commit to LCAIZJ/vllm-ascend that referenced this pull request Mar 7, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
mystous pushed a commit to mystous/vllm_hybrid that referenced this pull request May 10, 2026
…#30739)

Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: njhill <nickhill123@gmail.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: nanxing <1014662416@qq.com>
my-other-github-account pushed a commit to my-other-github-account/vllm that referenced this pull request May 15, 2026
…#30739)

Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: njhill <nickhill123@gmail.com>
0826joyce pushed a commit to 0826joyce/vllm-serving-optimization that referenced this pull request May 19, 2026
…#30739)

Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: njhill <nickhill123@gmail.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
xqchen7 pushed a commit to nv-action/vllm-benchmarks that referenced this pull request Jul 15, 2026
### What this PR does / why we need it?

Upgrade vllm commit to 0105 (8be6432bdaf6275664d857b1e5e9bf8ed1ce299e)

1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since
vllm-project/vllm#31517 deleted unused arg

2. Remove dense `Qwen/Qwen3-0.6B` in
`tests/e2e/multicard/test_aclgraph_capture_replay.py` and
`tests/e2e/multicard/test_data_parallel.py` due to
vllm-project/vllm#30739
where offline data parallel mode will not be supported/useful for dense
models

3. Adapt `vllm_ascend/worker/worker.py` due to
vllm-project/vllm#31584

4. Adapt `self.block_size` calling due to
vllm-project/vllm#31540

5. Modify `test_mla_v1.py` due to
vllm-project/vllm#28454 , which refactorred
`get_head_size()`

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.13.0
- vLLM main:
vllm-project/vllm@7157596

Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: xqchen7 <chenxueqing7@huawei.com>
plasticchris pushed a commit to plasticchris/vllm that referenced this pull request Jul 20, 2026
…#30739)

Signed-off-by: Nick Hill <nhill@redhat.com>
Signed-off-by: njhill <nickhill123@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kv-connector ready ONLY add when PR is ready to merge/full CI is needed v1

3 participants