Skip to content

[Core] Add MRV2 virtual-batch PCP for MLA - #46570

Merged
MatthewBonanni merged 20 commits into
vllm-project:mainfrom
neuralmagic:codex/mrv2-pcp-virtual-batch
Jul 19, 2026
Merged

MatthewBonanni merged 20 commits into
vllm-project:mainfrom
neuralmagic:codex/mrv2-pcp-virtual-batch

Conversation

@LucasWilkinson

@LucasWilkinson LucasWilkinson commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds an MRV2 virtual-batch implementation for prefill context parallelism (PCP), initially targeting MLA and sparse MLA.

  • Add a stateful PCPManager that partitions the globally scheduled MRV2 InputBatch into rank-local virtual rows while preserving global request state for restore, logits, and postprocessing.
  • Add MRV2 input-batch plumbing for PCP positions, sequence lengths, logits indices, and hidden-state restoration.
  • Add MLA PCP execution support while keeping PCP and DCP responsibilities orthogonal.
  • Add sparse MLA cache insertion for GLM/DSA-style models.
  • Align context-parallel group construction and KV ownership with the orthogonal PCP/DCP design.
  • Update scheduler/offload tests for DCP-owned KV sharding and add concurrent GSM8K eval configurations for TP2+PCP2+EP and TP1+PCP4+EP.

PCP algorithm

The scheduler continues to produce one global batch. Immediately before input preparation, PCPManager converts that batch into the rank-local virtual batch containing a portion of each prefill for that PCP rank.

Each prefill request is partitioned independently. Decode rows are replicated across PCP ranks so their model and KV-cache state remain synchronized.

DualChunkSwap partitioning

PCP uses a DualChunkSwap-style partition to balance causal-attention work. A request is split into 2 × PCP contiguous chunks, and rank r receives chunks r and 2 × PCP - 1 - r.

For PCP=4:

Global prefill:  | C0 | C1 | C2 | C3 | C4 | C5 | C6 | C7 |

PCP rank 0:        C0                                 C7
PCP rank 1:             C1                       C6
PCP rank 2:                  C2             C5
PCP rank 3:                       C3   C4

Earlier chunks attend to shorter prefixes, while later chunks attend to longer prefixes. Pairing one early and one late chunk gives each rank a more balanced amount of attention work than assigning one contiguous quarter of the request.

K and V—or latent KV in the case of MLA—are all-gathered so each rank has the full KV needed for KV-cache insertion and attention computation.

After model execution, hidden states are gathered and restored to global scheduled-token order before logits, sampling, and postprocessing.

Orthogonal PCP and DCP groups

Parallel-group behavior now follows RFC #46358: Unify Context Parallelism in vLLM (PCP ⊥ DCP).

PCP and DCP are treated as independent axes:

  • TP shards model heads.
  • PCP partitions prefill computation.
  • DCP shards KV ownership.
  • KV shard count is now DCP, not PCP × DCP.

This is to make it easier to reason about:

The reason I like this interface is that it keeps vllm/distributed/parallel_state.py cleaner/simpler, and makes it easier for the user to understand that the number of KV shards is simply dcp_size. This makes it a bit easier to reason about the KV replication factor (just becomes num_gpus / min(tp, num_kv_heads) / dcp, so its easy to know if its MLA and I want no replication I do DCP = num_gpus.

(See https://vllm-dev.slack.com/archives/C09JHEC74PM/p1771054846767689?thread_ts=1769445634.668079&cid=C09JHEC74PM for full motivation.)

For TP=2 and PCP=2:

                 TP0       TP1
              +---------+---------+
PCP0          | rank 0  | rank 1  |
              +---------+---------+
PCP1          | rank 2  | rank 3  |
              +---------+---------+

TP groups:        [0, 1]  [2, 3]
PCP groups:       [0, 2]  [1, 3]
DCP = PCP:        [0, 2]  [1, 3]
DCP = TP × PCP:   [0, 2, 1, 3]

DCP groups span the PCP axis first and then TP. This supports the RFC’s two combined configurations:

  • DCP = PCP: ranks already share the same TP-local query heads, so no TP query gather is required.
  • DCP = TP × PCP: queries are gathered across TP before attention, and the combined output is restored to each rank’s TP-local heads.

Block sizing, block tables, slot mappings, prefix-cache accounting, and KV capacity scale with DCP ownership only. PCP alone partitions prefill compute and does not reduce per-GPU KV capacity.

Results

Measured with nvidia/GLM-5.2-NVFP4 on 4xB300 GPUs, --max-model-len 131072, --max-num-batched-tokens 32768, eager execution, and FP8 KV cache.

The GSM8K column is a 100-sample, 5-shot smoke check only, not the full GSM8K evaluation. The 100k TFTT value is mean TTFT across three sequential requests with exactly 100,000 input tokens and one output token.

Deployment GSM8K (100 samples) 100k TFTT KV/GPU KV token capacity 131k contexts
TP1+PCP4+EP4 93% 2.859s 87.29 GiB 1,964,800 14.99
TP1+PCP4 93% 3.996s 84.37 GiB 1,899,200 14.49
TP2+PCP2+EP4 93% 3.549s 114.10 GiB 2,568,320 19.59
TP2+PCP2 94% 4.773s 115.08 GiB 2,590,528 19.76
TP4+EP4 92% 4.746s 127.40 GiB 2,867,776 21.88
TP4 95% 6.015s 130.48 GiB 2,937,024 22.41

Reproduction

Run one deployment at a time on 4xB300. The benchmark host used the following common environment:

export VLLM_USE_V2_MODEL_RUNNER=1
export VLLM_LOGGING_LEVEL=INFO
export MODEL=nvidia/GLM-5.2-NVFP4
export VLLM_DISABLE_LL_BF16_ROUTER=1
# TP1+PCP4+EP4
DEPLOYMENT_ARGS=(--prefill-context-parallel-size 4 --enable-expert-parallel)

# TP1+PCP4
DEPLOYMENT_ARGS=(--prefill-context-parallel-size 4)

# TP2+PCP2+EP4
DEPLOYMENT_ARGS=(--tensor-parallel-size 2 --prefill-context-parallel-size 2 --enable-expert-parallel)

# TP2+PCP2
DEPLOYMENT_ARGS=(--tensor-parallel-size 2 --prefill-context-parallel-size 2)

# TP4+EP4
DEPLOYMENT_ARGS=(--tensor-parallel-size 4 --enable-expert-parallel)

# TP4
DEPLOYMENT_ARGS=(--tensor-parallel-size 4)

Start the server:

.venv/bin/vllm serve "$MODEL" \
  --enforce-eager \
  --max-model-len 131072 \
  --max-num-batched-tokens 32768 \
  --safetensors-load-strategy prefetch \
  --moe-backend flashinfer_cutlass \
  --kv-cache-dtype fp8 \
  --trust-remote-code \
  --disable-uvicorn-access-log \
  --port 8000 \
  --seed 0 \
  "${DEPLOYMENT_ARGS[@]}"

Run the three-request 100k TFTT measurement:

.venv/bin/vllm bench serve \
  --backend vllm \
  --model "$MODEL" \
  --dataset-name random \
  --random-input-len 100000 \
  --random-output-len 1 \
  --random-range-ratio 0 \
  --num-prompts 3 \
  --max-concurrency 1 \
  --request-rate inf \
  --port 8000 \
  --ignore-eos \
  --disable-tqdm

Run the 100-sample GSM8K smoke check:

.venv/bin/python tests/evals/gsm8k/gsm8k_eval.py \
  --num-questions 100 \
  --num-shots 5 \
  --max-tokens 256 \
  --port 8000 \
  --temperature 0 \
  --seed 42 \
  --max-concurrency 16

Focused regression tests:

.venv/bin/python -m pytest \
  tests/v1/attention/test_attention_splitting.py \
  tests/v1/worker/test_gpu_batch_ordering.py -q

.venv/bin/python -m pytest \
  tests/v1/attention/test_sparse_mla_backends.py::test_split_indexer_prefill_chunks -q

AI assistance

AI assistance was used to draft and implement this change. The submitter reviewed the changed lines and ran the tests and model evaluations documented above.

@mergify mergify Bot added the v1 label Jun 24, 2026
@LucasWilkinson
LucasWilkinson force-pushed the codex/mrv2-pcp-virtual-batch branch from 8929805 to 664eee9 Compare June 24, 2026 04:22
@mergify

mergify Bot commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @LucasWilkinson.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Jun 24, 2026
@mergify

mergify Bot commented Jun 24, 2026

Copy link
Copy Markdown
Contributor
@mergify mergify Bot added the documentation Improvements or additions to documentation label Jun 24, 2026
@LucasWilkinson
LucasWilkinson force-pushed the codex/mrv2-pcp-virtual-batch branch from 1e072bf to 7b99578 Compare June 26, 2026 19:54
@mergify mergify Bot removed the needs-rebase label Jun 26, 2026
@mergify

mergify Bot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @LucasWilkinson.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

Add modular prefill context parallelism for MLA models, including prefix caching, DCP integration, and expert-parallel MoE dispatch.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@LucasWilkinson
LucasWilkinson force-pushed the codex/mrv2-pcp-virtual-batch branch from fabac90 to 0268d82 Compare July 14, 2026 23:36
Replace separate replicated-query and TP-spanning predicates with one DCP TP-size value shared by prefill metadata and decode communication.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@mergify mergify Bot removed the needs-rebase label Jul 15, 2026
LucasWilkinson and others added 4 commits July 15, 2026 02:43
Unify decode query gathering, clarify indexer KV gathering, and encapsulate PCP-only DCP chunk metadata and context selection.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Keep DCP groups within TP while PCP gathers cache-write inputs across its own axis. Reuse the regular MLA DCP implementation and remove the specialized PCP-DCP metadata and attention path.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@mergify mergify Bot added the ci/build label Jul 15, 2026
LucasWilkinson and others added 2 commits July 16, 2026 04:19
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@LucasWilkinson LucasWilkinson changed the title [Draft] Add MRV2 virtual-batch PCP for MLA Jul 16, 2026
@LucasWilkinson
LucasWilkinson marked this pull request as ready for review July 16, 2026 05:21
jasl added a commit to jasl/vllm that referenced this pull request Jul 21, 2026
…date page test

PR #32 multiplied build_offloading_config's tokens_per_block by
prefill_context_parallel_size (context_parallel_factor = dcp * pcp). That breaks the
upstream invariant that KV-cache sharding ignores PCP (parallel.py documents PCP does
not increase shard count), which the upstream test
test_offloading_spec_kv_sharding_ignores_prefill_context_parallel (added by vllm-project#46570)
asserts (tokens_per_block == (16,) at PCP=2, not (32,)). Restore the upstream
computation (block_size * decode_context_parallel_size). No-op on GB10 (PCP=1).

Also update test_spec_compact_page_size_non_divisible_budget to assert the new
budget-rounding page-size behavior instead of the removed gcd behavior.

Both found by running PR #32's full suite (not just the compact tests) post-merge.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
ArjunPakhan pushed a commit to ArjunPakhan/vllm that referenced this pull request Jul 21, 2026
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
@sfeng33 sfeng33 mentioned this pull request Jul 21, 2026
1 of 3 tasks
linfeng-yuan pushed a commit to vllm-project/vllm-ascend that referenced this pull request Jul 21, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@yiminghub2024

Copy link
Copy Markdown

@LucasWilkinson hi ,Congratulations,but a question, if we can support pcp = 16/32/64 ?

@yiminghub2024

Copy link
Copy Markdown

tested failed h20-3e*8 ,config:
image

error:
(APIServer pid=1) Value error, World size (64) is larger than the number of available GPUs (8) in this node. If this is intentional and you are using:

so ,should i change tp8->1 or your need to redesign ,make worldsize (8)= attn_cp_size(8) × attn_tp_size(1) × attn_dp_size(1) @LucasWilkinson

@yiminghub2024

Copy link
Copy Markdown

another attempt , if i setting -tp 1 --pcp 8,it can start ,but oom error with not start successful:
(EngineCore pid=359) RuntimeError: Worker failed with error 'CUDA out of memory. Tried to allocate 21.00 GiB. GPU 0 has a total capacity of 139.80 GiB of which 1.95 GiB is free. Including non-PyTorch memory, this process has 137.84 GiB memory in use. Of the allocated memory 134.10 GiB is allocated by PyTorch, and 2.02 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)', please check the stack trace above for the root cause
(APIServer pid=1) Traceback (most recent call last): @LucasWilkinson

ClockOfDestiny pushed a commit to ClockOfDestiny/vllm-ascend-hybrid-recompute that referenced this pull request Jul 22, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Signed-off-by: HeFangjun <hfj0219@mail.ustc.edu.cn>
li1how pushed a commit to li1how/vllm that referenced this pull request Jul 23, 2026
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
(cherry picked from commit b6ff8a2)
Alex-stack-hub pushed a commit to 0moyi0-2024/vllm-ascend_tp that referenced this pull request Jul 27, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@pisceskkk

Copy link
Copy Markdown
Contributor

P4+EP4 92% 4.746s 127.40 GiB 2,867,776 21.88

Just a quick question: was SP enabled in this TP performance benchmark?

@zwzmzd

zwzmzd commented Aug 11, 2026

Copy link
Copy Markdown

Question about the planned layout contract for PCP + MTP/spec decode

Thanks for the MRV2 PCP virtual-batch work. I’m trying to understand the intended design for supporting MTP/speculative decoding on top of PCP.

My current understanding is:

  1. prepare_inputs() first builds a global physical-request InputBatch, including next_prefill_tokens.
  2. PCPManager partitions it into rank-local DualChunkSwap virtual rows for the target-model forward.
  3. restore_for_sampling() all-gathers the target hidden states and restores the global InputBatch before sampling/speculator execution.
  4. AutoRegressiveSpeculator.prepare_prefill_inputs() currently assumes physical-request layout and inserts one next_prefill_token at the tail of each request.

If MTP input preparation were performed directly on the PCP virtual rows, a physical request split into multiple segments would need a different lookahead token at every internal segment boundary. A single next_prefill_tokens[request] value would not be sufficient.

Is the intended design to:

A. Restore the global target hidden states, construct the shifted MTP inputs in global physical-request order, and then repartition both input_ids and hidden_states using the same PCPBatchLayout before MTP forward;

or

B. Make the speculator PCP-segment-aware and explicitly construct a per-segment next-token/lookahead value?

Option A seems cleaner because it preserves the existing speculator request-level semantics and handles virtual-segment boundaries before PCP partitioning. It would presumably require rebuilding MTP-local attention metadata and slot mappings instead of reusing the target-forward metadata.

Also, is my understanding correct that cp_kv_cache_interleave_size is orthogonal to this problem and only affects DCP KV-cache ownership, while the MTP next-token alignment issue comes from the PCP DualChunkSwap layout?

MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Leetrytry pushed a commit to Leetrytry/vllm-ascend that referenced this pull request Sep 11, 2026
### What this PR does / why we need it?
The PCP/DCP solution in the community is different from that of Ascend.
The PCP/DCP PR in the community has already been
merged(vllm-project/vllm#46570), while Ascend's
PCP/DCP needs to be restructured to remain consistent with the community
version. Currently, the restructuring is in progress and cannot be
rolled out in the short term. To avoid affecting CI, all PCP/DCP test
cases have been taken offline.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/build documentation Improvements or additions to documentation kv-connector ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm v1

7 participants