Skip to content

[Feature][MiniMax-H3] Support diffusion continuous batching - #5810

Merged
hsliuustc0106 merged 21 commits into
vllm-project:mainfrom
princepride:feat/minimax-h3-step-execution
Aug 24, 2026
Merged

hsliuustc0106 merged 21 commits into
vllm-project:mainfrom
princepride:feat/minimax-h3-step-execution

Conversation

@princepride

@princepride princepride commented Aug 5, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Closes the Diffusion Continuous Batching item of #5700.

MiniMax-H3 ran its whole denoise loop inside one forward(), so it could not join the step-wise scheduler: one request occupied the engine end to end. This implements the step-execution contract (prepare_encode / denoise_step / step_scheduler / post_decode).

H3's DiT is already a variable-length packed model, so co-batched requests are concatenated into one sequence that keeps one attention document per request (plus that request's 64-row padding tail). Attention never crosses a request boundary and a batch costs one DiT forward.

Notes for reviewers:

  • Request mode and step mode share _prepare_request_inputs(), _build_denoise_inputs(), minimax_h3_prepare_denoise_rows() and _unpack_denoised_rows(), so the two paths cannot drift.
  • Video rows are the batched tensor the runner slices per request; audio rows have a different width, so they travel through request-private state along with the audio sigma schedule.
  • Multi-request packing needs a backend that consumes cu_seqlens (--diffusion-attention-backend FLASH_ATTN). Other backends stay correct by falling back to one forward per request.
  • _run_packed_attention now skips the KV prefix length and the 1-D pad mask when cu_seqlens describes more than one document, because neither can express a block-diagonal batch.
  • Step mode rejects two request-mode-only features with explicit errors: num_outputs_per_prompt > 1 (a request state holds one latent tensor) and distributed layerwise offload (its resident-layer window spans a whole denoise loop).

Batching does not speed H3 up. It is implemented for scheduler-level control, and the measurements below are recorded in the recipe so users are not misled.

Test Plan

vLLM Version: 0.26.0
vLLM-Omni Commit: ce17696a (rebased on 9235b0ae)
Hardware: 2 x H100 80GB

1. Correctness

pytest -sv tests/diffusion/models/minimax_h3 -m 'core_model and cpu'

Plus, on real weights, the same seeded prompt through both paths, and two concurrent requests co-batched:

vllm serve $MODEL_ROOT/FL2VA --omni --port 8091 --trust-remote-code \
  --tensor-parallel-size 2 --text-encoder-tp-size 2 --vae-use-tiling --enforce-eager \
  --diffusion-attention-backend FLASH_ATTN --step-execution --max-num-seqs 2

2. Performance

Same workload against three server configs (request = no --step-execution, step1 = --max-num-seqs 1, step4 = --max-num-seqs 4):

python3 benchmarks/diffusion/diffusion_benchmark_serving.py \
  --base-url http://127.0.0.1:8091 --endpoint /v1/videos --task t2v --dataset random \
  --model $MODEL_ROOT/FL2VA --num-prompts 4 --max-concurrency 4 --warmup-requests 1 \
  --width 672 --height 384 --fps 24 --num-inference-steps 30 --seed 1101 --disable-tqdm

3. Operator comparison

Loads only the DiT and profiles 3 separate single-request forwards against 1 fused three-request forward, using the same code paths denoise_step() takes.

profile_batching.py
import json, os
from collections import defaultdict
from pathlib import Path
from types import SimpleNamespace

os.environ.setdefault("DIFFUSION_ATTENTION_BACKEND", "FLASH_ATTN")
os.environ.setdefault("MASTER_ADDR", "127.0.0.1")
os.environ.setdefault("MASTER_PORT", "29591")

import torch
from safetensors.torch import safe_open
from vllm.config import VllmConfig, set_current_vllm_config
from vllm.distributed import init_distributed_environment, initialize_model_parallel

from vllm_omni.diffusion.models.minimax_h3.denoise_loop import MiniMaxH3DenoiseBranch
from vllm_omni.diffusion.models.minimax_h3.minimax_h3_transformer import MiniMaxH3DiTModel
from vllm_omni.diffusion.models.minimax_h3.packed_sequence import minimax_h3_packed_sequence
from vllm_omni.diffusion.models.minimax_h3.step_batch import minimax_h3_batched_forward_kwargs

MODEL = os.environ["MODEL_ROOT"] + "/FL2VA"
REQUESTS, ITERS, WARMUP = 3, 3, 2
# 209 frames at 672x384 -> 16384 packed rows per request.
SHAPE = dict(text_len=9, latent_t=62, latent_h=24, latent_w=42, audio_t=348)

init_distributed_environment(world_size=1, rank=0, local_rank=0, distributed_init_method="env://", backend="nccl")
config = json.loads((Path(MODEL) / "transformer" / "config.json").read_text())
od_config = SimpleNamespace(tf_model_config=config, parallel_config=SimpleNamespace(ulysses_degree=1))
vllm_config = VllmConfig()
with set_current_vllm_config(vllm_config):
    initialize_model_parallel(tensor_model_parallel_size=1)
    torch.set_default_device("cuda")
    model = MiniMaxH3DiTModel(od_config, quant_config=None)

    def weights():
        for shard in sorted((Path(MODEL) / "transformer").glob("*.safetensors")):
            with safe_open(shard, framework="pt", device="cpu") as f:
                for name in f.keys():
                    yield name, f.get_tensor(name)

    model.load_weights(weights())
    model.post_load_weights()
model.eval()
torch.set_default_device("cpu")

branches, videos, audios = [], [], []
for index in range(REQUESTS):
    packed = minimax_h3_packed_sequence(**SHAPE, include_keyframe_cond=False)
    gen = torch.Generator().manual_seed(index)
    branch = MiniMaxH3DenoiseBranch(
        packed=packed,
        text_embeddings=torch.randn(SHAPE["text_len"], 5120, generator=gen, dtype=torch.bfloat16),
        token_tags=packed["token_tags"],
        device=torch.device("cuda"),
    )
    branches.append(branch)
    videos.append(torch.randn(int(branch.img_pos.shape[0]), 96, generator=gen).cuda())
    audios.append(torch.randn(int(branch.audio_pos.shape[0]), 32, generator=gen).cuda())


def separate():
    for i, b in enumerate(branches):
        model(**b.forward_kwargs(video_rows=videos[i], audio_rows=audios[i], t_video=0.6, t_audio=0.5,
                                 imgvid_cond_timestep=0.999, audio_ref_cond_timestep=1.0))


def fused():
    n = len(branches)
    model(**minimax_h3_batched_forward_kwargs(
        branches=branches, video_rows=videos, audio_rows=audios,
        t_video=[0.6] * n, t_audio=[0.5] * n,
        imgvid_cond_timesteps=[0.999] * n, audio_ref_cond_timesteps=[1.0] * n))


def run(fn, label):
    from torch.profiler import ProfilerActivity, profile
    with torch.inference_mode():
        for _ in range(WARMUP):
            fn()
        torch.cuda.synchronize()
        start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
        start.record()
        for _ in range(ITERS):
            fn()
        end.record()
        torch.cuda.synchronize()
        with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
            for _ in range(ITERS):
                fn()
            torch.cuda.synchronize()
    ops, counts = defaultdict(float), defaultdict(int)
    for e in prof.key_averages():
        us = getattr(e, "self_device_time_total", 0.0) or 0.0
        if us > 0 and not e.key.startswith("aten::") and e.key != "Command Buffer Full":
            ops[e.key] += us / ITERS / 1000.0
            counts[e.key] += e.count // ITERS
    print(f"[{label}] wall {start.elapsed_time(end) / ITERS:.1f} ms, kernels {sum(ops.values()):.1f} ms")
    return ops, counts


sep_ops, sep_n = run(separate, "separate")
fus_ops, fus_n = run(fused, "fused")
for key in sorted(set(sep_ops) | set(fus_ops), key=lambda k: fus_ops.get(k, 0) - sep_ops.get(k, 0)):
    a, b = sep_ops.get(key, 0.0), fus_ops.get(key, 0.0)
    if max(a, b) >= 1.0:
        print(f"{b - a:9.1f}{a:10.1f}{b:10.1f}  {sep_n.get(key, 0):>5}->{fus_n.get(key, 0):<5} {key[:56]}")

Test Result

1. Correctness

Check Result
tests/diffusion/models/minimax_h3 -m 'core_model and cpu' 86 passed
Request mode vs step mode, same seed, real weights bitwise identical video and audio (max abs diff 0.000e+00)
Two concurrent requests co-batched into one 65280-row forward; each output identical to running it alone

2. Performance (BF16, TP2, 672x384, 209 frames, 30 steps, 4 requests at concurrency 4)

Config Wall time Mean latency Peak memory
request mode 174.8 s 111.5 s 72.4 GB
--step-execution --max-num-seqs 1 179.0 s 113.8 s 72.4 GB
--step-execution --max-num-seqs 4 182.1 s 175.7 s 78.3 GB

Same picture with online int8 (153.3 s / 158.4 s / 161.0 s), so quantization does not change the verdict.

Where a request's time goes (pipeline profiler, per request):

Stage Time Share
encode_prompt 0.37 s 0.9%
diffuse (29 steps) 38.34 s 90.1%
decode (VAE) 3.83 s 9.0%

3. Operator comparison (1 GPU, 3 requests x 16384 rows, one denoise step)

Two runs, to show the run-to-run spread:

3 separate forwards 1 fused forward Ratio
GPU kernel time 10510.7 / 10517.1 ms 10333.4 / 10456.1 ms 0.983 / 0.994
Wall time 7115.6 / 7122.7 ms 7057.3 / 7113.1 ms 0.992 / 0.999

The fused forward is never slower, but the margin is inside run-to-run noise: fusing the forward is roughly time-neutral, so the end-to-end regression comes from outside it.

The per-kernel split is stable across both runs:

Kernel Separate Fused Ratio Launches
FlashAttention (varlen) 3354.5 ms 3232.3 ms 0.96 156 -> 52
Token-refiner GEMM (9 text rows) 27.4 ms 9.1 ms 0.33 150 -> 50
Main DiT GEMM 2407.6 ms 2443.5 ms 1.01 - 1.03 600 -> 200
RMSNorm 244.7 ms 266.2 ms 1.09 - 1.10 630 -> 210
torch.cat 238.6 ms 259.6 ms 1.09 - 1.11 —

Fusing wins on kernel-launch amortization (attention, and the tiny M=9 token-refiner GEMM) and loses on cache locality: the main GEMM is already compute-bound at 16384 rows, and the bandwidth-bound norm/cat/elementwise kernels get 5-11% worse on the larger working set. The two roughly cancel.

Per-step cost is therefore close to linear in batch size, which is why merging requests does not reduce total time:

Requests per step Time per step vs batch=1
1 1.323 s 1.00x
3 3.945 s 2.98x
4 5.164 s 3.90x

Unlike LLM decoding, which is memory-bandwidth bound and batches almost for free, one H3 denoise step already has ~16k rows of dense math per request, so N requests cost N times the FLOPs.

4. Open-loop arrival experiment (4× H100)

Configuration: 4× H100 80GB, TP=2, USP=2, text-encoder TP=4, VAE patch parallelism=4, BF16, FlashAttention.
Workload: 10 requests at a fixed 5-second arrival interval; 672×384, 4 seconds / 24 FPS, 20 inference steps.

Configuration Wall time Throughput Mean latency Peak memory
request mode 60.8 s 9.88 req/min 11.7 s 56.9 GB
--step-execution --max-num-seqs 4 67.9 s 8.84 req/min 23.6 s 57.9 GB

The step-mode run reached peak compute overlap = 4, confirming packed continuous batching. It does not improve H3 throughput: mean diffuse time grew from 5.27 s to 16.46 s, so throughput fell 10.5% and mean latency increased. H3's dense DiT cost scales close to linearly with packed request count; step execution is useful for scheduler-level control, not throughput.

=== comparison
request  (0 - 31s, each cell 0.5s;  '-' waiting, '=' computing)
#0  |>--==============                                                   |    7.5s
#1  |          >-----==============                                      |    8.4s
#2  |                     >-------=============                          |    9.2s
#3  |                                >--------==============             |   10.1s
#4  |                                           >-----------=============|   11.3s

step4  (0 - 35s, each cell 0.5s;  '-' waiting, '=' computing)
#0  |>-===================                                               |   10.8s
#1  |         >----=============================                         |   16.9s
#2  |                   >-----===========================                |   16.8s
#3  |                             >-------===========================    |   17.7s
#4  |                                      >-------======================|   15.1s

metric               |        request |          step4
------------------------------------------------------
completed / failed   |          5 / 0 |          5 / 0
wall time (s)        |           31.3 |           35.1
throughput (req/min) |           9.59 |           8.55
latency mean (s)     |            9.3 |           15.5
latency p50 (s)      |            9.2 |           16.8
latency max (s)      |           11.3 |           17.7
queue delay mean (s) |            3.5 |            3.0
peak outstanding     |              2 |              4
peak compute overlap |              2 |              3
peak memory (MB)     |          56866 |          57210

mean pipeline stage time (s), server-reported
decode               |           0.52 |           0.52
diffuse              |           5.24 |          11.15
encode_prompt        |           0.10 |           0.08
post_decode          |           0.00 |           0.52
prepare_encode       |           0.00 |           0.15
@princepride
princepride force-pushed the feat/minimax-h3-step-execution branch from 04289b1 to 4b726c3 Compare August 5, 2026 14:20
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models enhancement New feature or request labels Aug 6, 2026
@princepride
princepride force-pushed the feat/minimax-h3-step-execution branch 2 times, most recently from 6121458 to 743bf0c Compare August 7, 2026 02:15
@princepride
princepride marked this pull request as ready for review August 7, 2026 08:10
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

princepride and others added 5 commits August 7, 2026 09:56
MiniMax-H3 ran its whole denoise loop inside one forward(), so it could
not participate in the step-wise scheduler: one request occupied the
engine end to end. Implement the step-execution contract
(prepare_encode / denoise_step / step_scheduler / post_decode) so the
scheduler can admit and retire H3 requests between denoise steps.

H3's DiT is already a variable-length packed model, so co-batched
requests are concatenated into a single sequence whose cu_seqlens keeps
one document per request plus that request's alignment-padding tail.
Attention never crosses a request boundary and a whole batch costs one
DiT forward. Backends that ignore cu_seqlens cannot express that
isolation, so they fall back to one forward per request.

Request mode and step mode share _prepare_request_inputs(),
_build_denoise_inputs(), minimax_h3_prepare_denoise_rows(), and
_unpack_denoised_rows(), so the two paths cannot drift.

Video rows are the batched tensor the runner slices per request; audio
rows have a different width, so they travel through request-private
state along with the audio sigma schedule.

Validated on MiniMaxAI/MiniMax-H3 FL2VA (2xH100, TP2, 672x384):
- request mode vs step mode: bitwise identical video and audio
- two concurrent requests co-batch into one 9088-row forward and produce
  the same output as running each alone

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
The step-execution notes claimed a throughput and fairness benefit for
co-batching H3. Benchmarking on two H100s (TP2, 672x384, 30 steps, 4
requests at concurrency 4) contradicts that: request mode finishes in
174.8s, --max-num-seqs 1 in 179.0s, and --max-num-seqs 4 in 182.1s with
8% more peak memory and mean latency degrading from 111.5s to 175.7s.

An H3 denoise step is a compute-bound dense GEMM over an already long
packed sequence, so fusing N requests costs N times the FLOPs and buys no
amortization -- unlike LLM decoding, which is memory-bandwidth bound.
Replace the unsupported claim with the measured table and steer users to
max_num_seqs=1 unless they need step-level scheduling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Signed-off-by: princepride <wangzhipeng628@gmail.com>
State the workload precisely (209 frames, 16384 packed rows per request)
so the numbers are reproducible, and give the per-step evidence behind
the conclusion: going from one request per step to four moves the
per-request denoise cost only from 1.323s to 1.291s, so there is nearly
nothing for batching to amortize.

Also note that quantization does not change the verdict: online int8
runs the same workload in 153.3s at 56.9GB in request mode, and
--max-num-seqs 4 remains 5.0% slower than request mode there too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
@princepride
princepride force-pushed the feat/minimax-h3-step-execution branch from a409b6c to f3c8780 Compare August 7, 2026 09:56
@princepride

Copy link
Copy Markdown
Collaborator Author
@lishunyang12

Copy link
Copy Markdown
Collaborator

Local validation on PR head f3c87809a9074b7bdc91880dcb41428119c81596:

Environment: Python 3.12.3, PyTorch 2.11.0+cu130, vLLM 0.26.0.

Scope Command Result
MiniMax-H3 contracts + step execution pytest -sv tests/diffusion/models/minimax_h3 -m 'core_model and cpu' 125 passed, 5 deselected
Shared diffusion scheduler + multiprocess concurrency pytest -q tests/diffusion/test_diffusion_scheduler.py tests/diffusion/test_multiproc_engine_concurrency.py 98 passed
Async output + streaming-output integration pytest -q tests/diffusion/test_async_output_worker.py tests/diffusion/test_diffusion_streaming_output.py 15 passed

Total: 238 passed. This covers request/step parity, packed request isolation, batched vs independent execution, progress propagation, abort-after-inflight-step, scheduler admission/retirement, worker-death handling, and async output behavior.

GPU note: I prepared a 2×B300 real-weight packed-batch parity run, but all 8 B300s on this machine are occupied by a pre-existing long-running DLO soak. I did not interrupt that experiment, so I am not claiming a fresh GPU E2E result in this comment.

@princepride

Copy link
Copy Markdown
Collaborator Author

@lishunyang12 Seems I can enable DLO and CB in the same time.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Semmer2 do we have to add a standalone step_batch.py to support this feature?

@Semmer2

Semmer2 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

@Semmer2 do we have to add a standalone step_batch.py to support this feature?

Hi, @hsliuustc0106 the code itself is reasonable. I'm fine with it staying in a standalone file or being merged into denoise_loop.py — either works for me.
Actually, my concern is the name: step_batch.py looks like an implementation related to the engine's STEP_BATCH execution mode, while it actually packs N requests' layouts into one forward. The name like batched_packing.py would be clearer and avoid any possible confusion.

@princepride

Copy link
Copy Markdown
Collaborator Author

Actually, my concern is the name: step_batch.py looks like an implementation related to the engine's STEP_BATCH execution mode, while it actually packs N requests' layouts into one forward. The name like batched_packing.py would be clearer and avoid any possible confusion.

Sure, I can change it to batched_packing.py to avoid ambiguity.

@princepride
princepride force-pushed the feat/minimax-h3-step-execution branch from 3b6df66 to 07ca1ee Compare August 21, 2026 03:48
Signed-off-by: princepride <wangzhipeng628@gmail.com>
@princepride
princepride force-pushed the feat/minimax-h3-step-execution branch from 07ca1ee to 9397578 Compare August 21, 2026 06:19
@princepride princepride added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 21, 2026
@princepride

Copy link
Copy Markdown
Collaborator Author
@Semmer2

Semmer2 commented Aug 21, 2026

Copy link
Copy Markdown
Contributor
…execution

# Conflicts:
#	vllm_omni/diffusion/models/minimax_h3/pipeline_minimax_h3.py
@princepride princepride added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 24, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

CI failure not related and fixed already

@hsliuustc0106
hsliuustc0106 disabled auto-merge August 24, 2026 13:08
@hsliuustc0106
hsliuustc0106 merged commit d150a4f into vllm-project:main Aug 24, 2026
7 of 9 checks passed
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
…ject#5810)

Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
…ject#5810)

Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@Isotr0py Isotr0py mentioned this pull request Sep 18, 2026
9 of 13 tasks
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…ject#5810)

Signed-off-by: princepride <wangzhipeng628@gmail.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models enhancement New feature or request ready label to trigger buildkite CI

8 participants