[Feature][MiniMax-H3] Support diffusion continuous batching - #5810
hsliuustc0106 merged 21 commits into
Conversation
04289b1 to
4b726c3
Compare
6121458 to
743bf0c
Compare
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
MiniMax-H3 ran its whole denoise loop inside one forward(), so it could not participate in the step-wise scheduler: one request occupied the engine end to end. Implement the step-execution contract (prepare_encode / denoise_step / step_scheduler / post_decode) so the scheduler can admit and retire H3 requests between denoise steps. H3's DiT is already a variable-length packed model, so co-batched requests are concatenated into a single sequence whose cu_seqlens keeps one document per request plus that request's alignment-padding tail. Attention never crosses a request boundary and a whole batch costs one DiT forward. Backends that ignore cu_seqlens cannot express that isolation, so they fall back to one forward per request. Request mode and step mode share _prepare_request_inputs(), _build_denoise_inputs(), minimax_h3_prepare_denoise_rows(), and _unpack_denoised_rows(), so the two paths cannot drift. Video rows are the batched tensor the runner slices per request; audio rows have a different width, so they travel through request-private state along with the audio sigma schedule. Validated on MiniMaxAI/MiniMax-H3 FL2VA (2xH100, TP2, 672x384): - request mode vs step mode: bitwise identical video and audio - two concurrent requests co-batch into one 9088-row forward and produce the same output as running each alone Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: princepride <wangzhipeng628@gmail.com>
The step-execution notes claimed a throughput and fairness benefit for co-batching H3. Benchmarking on two H100s (TP2, 672x384, 30 steps, 4 requests at concurrency 4) contradicts that: request mode finishes in 174.8s, --max-num-seqs 1 in 179.0s, and --max-num-seqs 4 in 182.1s with 8% more peak memory and mean latency degrading from 111.5s to 175.7s. An H3 denoise step is a compute-bound dense GEMM over an already long packed sequence, so fusing N requests costs N times the FLOPs and buys no amortization -- unlike LLM decoding, which is memory-bandwidth bound. Replace the unsupported claim with the measured table and steer users to max_num_seqs=1 unless they need step-level scheduling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: princepride <wangzhipeng628@gmail.com>
State the workload precisely (209 frames, 16384 packed rows per request) so the numbers are reproducible, and give the per-step evidence behind the conclusion: going from one request per step to four moves the per-request denoise cost only from 1.323s to 1.291s, so there is nearly nothing for batching to amortize. Also note that quantization does not change the verdict: online int8 runs the same workload in 153.3s at 56.9GB in request mode, and --max-num-seqs 4 remains 5.0% slower than request mode there too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
a409b6c to
f3c8780
Compare
|
Local validation on PR head Environment: Python 3.12.3, PyTorch 2.11.0+cu130, vLLM 0.26.0.
Total: 238 passed. This covers request/step parity, packed request isolation, batched vs independent execution, progress propagation, abort-after-inflight-step, scheduler admission/retirement, worker-death handling, and async output behavior. GPU note: I prepared a 2×B300 real-weight packed-batch parity run, but all 8 B300s on this machine are occupied by a pre-existing long-running DLO soak. I did not interrupt that experiment, so I am not claiming a fresh GPU E2E result in this comment. |
|
@lishunyang12 Seems I can enable DLO and CB in the same time. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
@Semmer2 do we have to add a standalone step_batch.py to support this feature?
Hi, @hsliuustc0106 the code itself is reasonable. I'm fine with it staying in a standalone file or being merged into |
Sure, I can change it to |
3b6df66 to
07ca1ee
Compare
Signed-off-by: princepride <wangzhipeng628@gmail.com>
07ca1ee to
9397578
Compare
|
@Semmer2 @hsliuustc0106 PTAL |
LGTM |
…execution # Conflicts: # vllm_omni/diffusion/models/minimax_h3/pipeline_minimax_h3.py
|
CI failure not related and fixed already |
…ject#5810) Signed-off-by: princepride <wangzhipeng628@gmail.com> Signed-off-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
…ject#5810) Signed-off-by: princepride <wangzhipeng628@gmail.com> Signed-off-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Cursor <cursoragent@cursor.com>
…ject#5810) Signed-off-by: princepride <wangzhipeng628@gmail.com> Signed-off-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Purpose
Closes the Diffusion Continuous Batching item of #5700.
MiniMax-H3 ran its whole denoise loop inside one
forward(), so it could not join the step-wise scheduler: one request occupied the engine end to end. This implements the step-execution contract (prepare_encode/denoise_step/step_scheduler/post_decode).H3's DiT is already a variable-length packed model, so co-batched requests are concatenated into one sequence that keeps one attention document per request (plus that request's 64-row padding tail). Attention never crosses a request boundary and a batch costs one DiT forward.
Notes for reviewers:
_prepare_request_inputs(),_build_denoise_inputs(),minimax_h3_prepare_denoise_rows()and_unpack_denoised_rows(), so the two paths cannot drift.cu_seqlens(--diffusion-attention-backend FLASH_ATTN). Other backends stay correct by falling back to one forward per request._run_packed_attentionnow skips the KV prefix length and the 1-D pad mask whencu_seqlensdescribes more than one document, because neither can express a block-diagonal batch.num_outputs_per_prompt > 1(a request state holds one latent tensor) and distributed layerwise offload (its resident-layer window spans a whole denoise loop).Batching does not speed H3 up. It is implemented for scheduler-level control, and the measurements below are recorded in the recipe so users are not misled.
Test Plan
vLLM Version: 0.26.0
vLLM-Omni Commit:
ce17696a(rebased on9235b0ae)Hardware: 2 x H100 80GB
1. Correctness
pytest -sv tests/diffusion/models/minimax_h3 -m 'core_model and cpu'Plus, on real weights, the same seeded prompt through both paths, and two concurrent requests co-batched:
vllm serve $MODEL_ROOT/FL2VA --omni --port 8091 --trust-remote-code \ --tensor-parallel-size 2 --text-encoder-tp-size 2 --vae-use-tiling --enforce-eager \ --diffusion-attention-backend FLASH_ATTN --step-execution --max-num-seqs 22. Performance
Same workload against three server configs (
request= no--step-execution,step1=--max-num-seqs 1,step4=--max-num-seqs 4):python3 benchmarks/diffusion/diffusion_benchmark_serving.py \ --base-url http://127.0.0.1:8091 --endpoint /v1/videos --task t2v --dataset random \ --model $MODEL_ROOT/FL2VA --num-prompts 4 --max-concurrency 4 --warmup-requests 1 \ --width 672 --height 384 --fps 24 --num-inference-steps 30 --seed 1101 --disable-tqdm3. Operator comparison
Loads only the DiT and profiles 3 separate single-request forwards against 1 fused three-request forward, using the same code paths
denoise_step()takes.profile_batching.py
Test Result
1. Correctness
tests/diffusion/models/minimax_h3 -m 'core_model and cpu'0.000e+00)2. Performance (BF16, TP2, 672x384, 209 frames, 30 steps, 4 requests at concurrency 4)
--step-execution --max-num-seqs 1--step-execution --max-num-seqs 4Same picture with online
int8(153.3 s / 158.4 s / 161.0 s), so quantization does not change the verdict.Where a request's time goes (pipeline profiler, per request):
encode_promptdiffuse(29 steps)decode(VAE)3. Operator comparison (1 GPU, 3 requests x 16384 rows, one denoise step)
Two runs, to show the run-to-run spread:
The fused forward is never slower, but the margin is inside run-to-run noise: fusing the forward is roughly time-neutral, so the end-to-end regression comes from outside it.
The per-kernel split is stable across both runs:
torch.catFusing wins on kernel-launch amortization (attention, and the tiny
M=9token-refiner GEMM) and loses on cache locality: the main GEMM is already compute-bound at 16384 rows, and the bandwidth-bound norm/cat/elementwise kernels get 5-11% worse on the larger working set. The two roughly cancel.Per-step cost is therefore close to linear in batch size, which is why merging requests does not reduce total time:
Unlike LLM decoding, which is memory-bandwidth bound and batches almost for free, one H3 denoise step already has ~16k rows of dense math per request, so N requests cost N times the FLOPs.
4. Open-loop arrival experiment (4× H100)
Configuration: 4× H100 80GB, TP=2, USP=2, text-encoder TP=4, VAE patch parallelism=4, BF16, FlashAttention.
Workload: 10 requests at a fixed 5-second arrival interval; 672×384, 4 seconds / 24 FPS, 20 inference steps.
--step-execution --max-num-seqs 4The step-mode run reached
peak compute overlap = 4, confirming packed continuous batching. It does not improve H3 throughput: meandiffusetime grew from 5.27 s to 16.46 s, so throughput fell 10.5% and mean latency increased. H3's dense DiT cost scales close to linearly with packed request count; step execution is useful for scheduler-level control, not throughput.