Skip to content

[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation - #46448

Merged
vllm-bot merged 4 commits into
vllm-project:mainfrom
EanWang211123:feat/local-argmax/mrv2
Jun 26, 2026
Merged

vllm-bot merged 4 commits into
vllm-project:mainfrom
EanWang211123:feat/local-argmax/mrv2

Conversation

@EanWang211123

@EanWang211123 EanWang211123 commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

use_local_argmax_reduction is already supported in the V1 speculative decoding path (SpecDecodeBaseProposer), but Model Runner V2 draft sampling in DraftModelSpeculator still always calls compute_logits() and all-gathers full vocabulary logits before argmax. This PR wires the existing config flag into the MR V2 greedy draft path via get_top_tokens(), reducing TP communication from O(vocab_size) to O(2 × tp_size) per token. All MR V2 speculators that inherit DraftModelSpeculator (DFlash, Eagle, MTP, Gemma4 MTP, etc.) benefit from this single change.

Summary of Changes

  • File: vllm/v1/worker/gpu/spec_decode/speculator.py only.

  • Read use_local_argmax_reduction from speculative_config in DraftModelSpeculator.__init__.

  • Add _validate_local_argmax_reduction() (called from load_model):

    • Reject incompatible draft_sample_method='probabilistic'.
    • Require draft model to implement get_top_tokens().
  • Add _greedy_sample_draft() and route greedy sampling through get_top_tokens() when the flag is enabled; otherwise keep compute_logits().argmax().

  • Refactor sample_draft(): probabilistic path unchanged; greedy path uses _greedy_sample_draft().

  • All MR V2 speculators inheriting DraftModelSpeculator (DFlash, Eagle, MTP, Gemma4 MTP, etc.) pick up this change with no subclass edits. Default remains off (no behavior change).

Test

Environment

  • 2× NVIDIA RTX 4090 (TP=2, no NVLink)

Models

  • Target: Qwen3-8B
  • Draft: Qwen3-8B-DFlash-b16 (method=dflash, num_speculative_tokens=15)

Benchmark

  • Dataset: HumanEval (openai_human_eval)
  • Concurrency: 32

Commands

Baseline (default, use_local_argmax_reduction=false):

CUDA_VISIBLE_DEVICES=4,5 vllm serve /nfs_models/Qwen/Qwen3-8B \
  --host 0.0.0.0 --port 8000 --served-model-name Qwen3-8B \
  --tensor-parallel-size 2 --trust-remote-code --dtype auto \
  --max-model-len 8192 --max-num-seqs 64 --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.90 \
  --speculative-config '{"method":"dflash","model":"/nfs_models/z-lab/Qwen3-8B-DFlash-b16/","num_speculative_tokens":15}'

With local argmax:

CUDA_VISIBLE_DEVICES=4,5 vllm serve /nfs_models/Qwen/Qwen3-8B \
  ... \
  --speculative-config '{"use_local_argmax_reduction":true,"method":"dflash","model":"/nfs_models/z-lab/Qwen3-8B-DFlash-b16/","num_speculative_tokens":15}'

Results

Config Output throughput (tok/s) Mean TPOT (ms)
Baseline 2159.3 11.56
use_local_argmax_reduction=true 2327.1 10.60

~+7.8% throughput, ~−8.3% TPOT on 4090×2 TP=2.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

cc @benchislett

Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>

@benchislett benchislett left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@benchislett
benchislett enabled auto-merge (squash) June 25, 2026 00:55
@benchislett benchislett added the ready ONLY add when PR is ready to merge/full CI is needed label Jun 25, 2026
@vllm-bot
vllm-bot merged commit 652d962 into vllm-project:main Jun 26, 2026
72 of 79 checks passed
wincent8 pushed a commit to wincent8/vllm that referenced this pull request Jun 29, 2026
…n generation (vllm-project#46448)

Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
philippesic pushed a commit to philippesic/vllm-semantic-cache that referenced this pull request Jul 19, 2026
…n generation (vllm-project#46448)

Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready ONLY add when PR is ready to merge/full CI is needed v1

4 participants