[Core] Add MRV2 virtual-batch PCP for MLA - #46570
Conversation
8929805 to
664eee9
Compare
|
This pull request has merge conflicts that must be resolved before it can be |
|
Documentation preview: https://vllm--46570.org.readthedocs.build/en/46570/ |
1e072bf to
7b99578
Compare
|
This pull request has merge conflicts that must be resolved before it can be |
795bf7a to
f093f03
Compare
Add modular prefill context parallelism for MLA models, including prefix caching, DCP integration, and expert-parallel MoE dispatch. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
fabac90 to
0268d82
Compare
Replace separate replicated-query and TP-spanning predicates with one DCP TP-size value shared by prefill metadata and decode communication. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Unify decode query gathering, clarify indexer KV gathering, and encapsulate PCP-only DCP chunk metadata and context selection. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Keep DCP groups within TP while PCP gathers cache-write inputs across its own axis. Reuse the regular MLA DCP implementation and remove the specialized PCP-DCP metadata and attention path. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Codex <noreply@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
…date page test PR #32 multiplied build_offloading_config's tokens_per_block by prefill_context_parallel_size (context_parallel_factor = dcp * pcp). That breaks the upstream invariant that KV-cache sharding ignores PCP (parallel.py documents PCP does not increase shard count), which the upstream test test_offloading_spec_kv_sharding_ignores_prefill_context_parallel (added by vllm-project#46570) asserts (tokens_per_block == (16,) at PCP=2, not (32,)). Restore the upstream computation (block_size * decode_context_parallel_size). No-op on GB10 (PCP=1). Also update test_spec_compact_page_size_non_divisible_budget to assert the new budget-rounding page-size behavior instead of the removed gcd behavior. Both found by running PR #32's full suite (not just the compact tests) post-merge. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
|
@LucasWilkinson hi ,Congratulations,but a question, if we can support pcp = 16/32/64 ? |
|
tested failed h20-3e*8 ,config: error: so ,should i change tp8->1 or your need to redesign ,make worldsize (8)= attn_cp_size(8) × attn_tp_size(1) × attn_dp_size(1) @LucasWilkinson |
|
another attempt , if i setting -tp 1 --pcp 8,it can start ,but oom error with not start successful: |
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com> Signed-off-by: HeFangjun <hfj0219@mail.ustc.edu.cn>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: OpenAI Codex <codex@openai.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com> (cherry picked from commit b6ff8a2)
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Just a quick question: was SP enabled in this TP performance benchmark? |
Question about the planned layout contract for PCP + MTP/spec decodeThanks for the MRV2 PCP virtual-batch work. I’m trying to understand the intended design for supporting MTP/speculative decoding on top of PCP. My current understanding is:
If MTP input preparation were performed directly on the PCP virtual rows, a physical request split into multiple segments would need a different lookahead token at every internal segment boundary. A single Is the intended design to: A. Restore the global target hidden states, construct the shifted MTP inputs in global physical-request order, and then repartition both or B. Make the speculator PCP-segment-aware and explicitly construct a per-segment next-token/lookahead value? Option A seems cleaner because it preserves the existing speculator request-level semantics and handles virtual-segment boundaries before PCP partitioning. It would presumably require rebuilding MTP-local attention metadata and slot mappings instead of reusing the target-forward metadata. Also, is my understanding correct that |
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
### What this PR does / why we need it? The PCP/DCP solution in the community is different from that of Ascend. The PCP/DCP PR in the community has already been merged(vllm-project/vllm#46570), while Ascend's PCP/DCP needs to be restructured to remain consistent with the community version. Currently, the restructuring is in progress and cannot be rolled out in the short term. To avoid affecting CI, all PCP/DCP test cases have been taken offline. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>

Summary
Adds an MRV2 virtual-batch implementation for prefill context parallelism (PCP), initially targeting MLA and sparse MLA.
PCPManagerthat partitions the globally scheduled MRV2InputBatchinto rank-local virtual rows while preserving global request state for restore, logits, and postprocessing.PCP algorithm
The scheduler continues to produce one global batch. Immediately before input preparation,
PCPManagerconverts that batch into the rank-local virtual batch containing a portion of each prefill for that PCP rank.Each prefill request is partitioned independently. Decode rows are replicated across PCP ranks so their model and KV-cache state remain synchronized.
DualChunkSwap partitioning
PCP uses a DualChunkSwap-style partition to balance causal-attention work. A request is split into
2 × PCPcontiguous chunks, and rankrreceives chunksrand2 × PCP - 1 - r.For PCP=4:
Earlier chunks attend to shorter prefixes, while later chunks attend to longer prefixes. Pairing one early and one late chunk gives each rank a more balanced amount of attention work than assigning one contiguous quarter of the request.
KandV—or latent KV in the case of MLA—are all-gathered so each rank has the full KV needed for KV-cache insertion and attention computation.After model execution, hidden states are gathered and restored to global scheduled-token order before logits, sampling, and postprocessing.
Orthogonal PCP and DCP groups
Parallel-group behavior now follows RFC #46358: Unify Context Parallelism in vLLM (PCP ⊥ DCP).
PCP and DCP are treated as independent axes:
DCP, notPCP × DCP.This is to make it easier to reason about:
(See https://vllm-dev.slack.com/archives/C09JHEC74PM/p1771054846767689?thread_ts=1769445634.668079&cid=C09JHEC74PM for full motivation.)
For TP=2 and PCP=2:
DCP groups span the PCP axis first and then TP. This supports the RFC’s two combined configurations:
DCP = PCP: ranks already share the same TP-local query heads, so no TP query gather is required.DCP = TP × PCP: queries are gathered across TP before attention, and the combined output is restored to each rank’s TP-local heads.Block sizing, block tables, slot mappings, prefix-cache accounting, and KV capacity scale with DCP ownership only. PCP alone partitions prefill compute and does not reduce per-GPU KV capacity.
Results
Measured with
nvidia/GLM-5.2-NVFP4on 4xB300 GPUs,--max-model-len 131072,--max-num-batched-tokens 32768, eager execution, and FP8 KV cache.The GSM8K column is a 100-sample, 5-shot smoke check only, not the full GSM8K evaluation. The 100k TFTT value is mean TTFT across three sequential requests with exactly 100,000 input tokens and one output token.
Reproduction
Run one deployment at a time on 4xB300. The benchmark host used the following common environment:
Start the server:
Run the three-request 100k TFTT measurement:
.venv/bin/vllm bench serve \ --backend vllm \ --model "$MODEL" \ --dataset-name random \ --random-input-len 100000 \ --random-output-len 1 \ --random-range-ratio 0 \ --num-prompts 3 \ --max-concurrency 1 \ --request-rate inf \ --port 8000 \ --ignore-eos \ --disable-tqdmRun the 100-sample GSM8K smoke check:
Focused regression tests:
AI assistance
AI assistance was used to draft and implement this change. The submitter reviewed the changed lines and ran the tests and model evaluations documented above.