Skip to content

[Model][Performance] Optimize MiniMax-H3 strict Ulysses boundaries - #6173

Merged
hsliuustc0106 merged 6 commits into
vllm-project:mainfrom
mo-ke-ke:codex/minimax-h3-local-sp
Aug 15, 2026
Merged

hsliuustc0106 merged 6 commits into
vllm-project:mainfrom
mo-ke-ke:codex/minimax-h3-local-sp

Conversation

@mo-ke-ke

@mo-ke-ke mo-ke-ke commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

MiniMax-H3's strict Ulysses path previously built the complete packed
[S, 5376] embedding and RoPE tensors on every rank, then immediately selected
that rank's S / world_size rows. After the transformer blocks it also gathered
the full BF16 hidden state before applying the final 128-channel projection.

This PR keeps those two model boundaries rank-local when, and only when, the
registered sp_input---local_sp_prepare hook proves the strict-Ulysses layout:

  • select rank-owned image/audio latent rows before their patch projections, run
    the token refiner on its required full text context before selecting rank-owned
    text rows, construct only the local embedding rows, and slice matching RoPE
    rows;
  • keep transformer output rows local through the final projection, then gather
    [S, 128] FP32 logits instead of [S, 5376] BF16 hidden states;
  • fall back to the existing full-sequence path for SP=1, Ring SP, AllGather-KV,
    hybrid/advanced layouts, missing hooks, non-forward contexts, or shapes that
    cannot be divided evenly.

For MiniMax-H3, the final all-gather payload is reduced from 10,752 bytes to
512 bytes per packed row (21x smaller). This is a boundary-payload reduction,
not a 21x end-to-end performance claim.

The implementation deliberately does not change Ulysses Q/K/V all-to-all,
the attention backend, packed-prefix handling, scheduler semantics, or model
weights. Focused tests exercise the optimized contract and every fallback above.

8x B300 performance and quality

Hardware and software:

  • 8x NVIDIA B300 SXM6, 275,040 MiB per GPU, compute capability 10.3
  • driver 580.126.09, CUDA 13.2, PyTorch 2.13.0+cu132
  • vLLM 0.27.0, vLLM-Omni current main dbc0dd6d
  • official MiniMax-H3 FL2VA checkpoint, BF16, default TRTLLM_ATTN

Fixed workload:

  • T2VA, 1344x768, 107 frames, 24 fps, 50 denoising steps, seed 42
  • Ulysses 8, Ring 1, AllGather-KV 1, text-encoder TP 8, VAE patch parallel 8
  • regional torch.compile; one excluded warmup and five measured generations
    for each engine; B1 (main) -> candidate -> B2 (main) run order
  • same prompt, checkpoint, environment, and generation parameters for every run
Metric Current main (10 runs, B1+B2) This PR (5 runs) Change
Diffusion-stage median 16.8485 s 16.5584 s -1.72%
Full-generation wall median 18.8572 s 18.6598 s -1.05%
Maximum peak memory per rank 89,654 MiB 88,530 MiB -1,124 MiB (-1.25%)

Main diffusion samples were
[16.7780, 16.8498, 16.8513, 16.8456, 16.8471, 16.8226, 16.8169, 16.8514, 16.8806, 16.8583] seconds. Candidate samples were
[16.5572, 16.5727, 16.5562, 16.5878, 16.5584] seconds. The B1 and B2
medians were 16.8471 s and 16.8514 s, respectively, which guards against a
one-directional thermal or clock drift explanation. This is a batch-size-1
latency optimization; no throughput or concurrency-scaling claim is made.

The numerical order changes because the final projection now precedes the
all-gather, so output hashes are not expected to match the baseline bit for bit.
Every repeat within each build was deterministic. Candidate versus baseline:

  • video mean SSIM 0.99813 (minimum-frame SSIM 0.99642), PSNR 53.33 dB,
    MAE 0.000831;
  • audio STFT cosine similarity 0.89359 and RMS ratio 1.00707;
  • gates: mean SSIM >= 0.97, audio cosine >= 0.80, and RMS ratio in [0.5, 2.0]:
    all passed.

Latest-main compatibility was requalified after main advanced to 596c16a5
(including fused Q/K RMSNorm and RoPE from #5990). The conflict resolution
keeps the fused RoPE table construction after selecting this rank's strict-
Ulysses position rows. On the same 8x B300 allocation, the combined MiniMax-
H3 and fused-QK/RoPE suites passed (152 passed, 3 skipped), and an exact
2-step Ulysses-8 T2VA smoke matched the latest-main video bit for bit
(SSIM 1.0, MAE 0) with audio STFT cosine approximately 1.0 and RMS ratio
1.0000045. Peak memory was 89,268 MiB/rank on latest main versus 88,048
MiB/rank for the candidate. These are compatibility-smoke results, not a
replacement formal performance claim for the B -> candidate -> B results
above.

The benchmark used the same Omni.generate contract as the checked-in
MiniMax-H3 E2E test. This PR does not add a second model-specific benchmark
framework because draft #5852 already owns the generic MiniMax-H3 2/4/8-GPU
benchmark and SM120 work.

Duplicate-work and attribution check

Before opening this draft I inspected issue #5700 and searched open PRs by the
roadmap issue, MiniMax-H3/Ulysses/SP-boundary keywords, the new hook name, and
the affected model file.

The closed, unmerged draft #5750 contained independently reviewable versions
of the local-embedding and compact-gather ideas. This change reimplements them
against current main, tightens activation to the exact registered strict-
Ulysses hook, adds current fallback/contract coverage, and carries
Co-authored-by: david6666666 <530634352@qq.com> in the commit.

Test Plan

pre-commit run --files \
  vllm_omni/diffusion/models/minimax_h3/minimax_h3_transformer.py \
  tests/diffusion/models/minimax_h3/test_minimax_h3_parallel.py

PYTHONSAFEPATH=1 PYTHONPATH=$CANDIDATE_ROOT \
  .venv-ci/bin/python -P -m pytest -q \
  $CANDIDATE_ROOT/tests/diffusion/models/minimax_h3 \
  -m "not e2e and not full_model"

PYTHONSAFEPATH=1 PYTHONPATH=$CANDIDATE_ROOT \
  .venv-ci/bin/python -P -m pytest -q \
  $CANDIDATE_ROOT/tests/diffusion/models/minimax_h3/test_minimax_h3_full_model.py::test_minimax_h3_quantization_quality

PYTHONSAFEPATH=1 PYTHONPATH=$CANDIDATE_ROOT \
  .venv-ci/bin/python -P -m pytest -q \
  $CANDIDATE_ROOT/tests/diffusion/models/minimax_h3/test_minimax_h3_e2e.py::test_minimax_h3_t2va_ulysses8_smoke

vLLM Version: 0.27.0

vLLM-Omni Commit: 092b8c7d (latest-main compatibility base 596c16a5; formal performance base dbc0dd6d)

Test Result

  • changed-file pre-commit: all hooks passed, including the four files touched
    by the [Kernel] Fuse Q/K RMSNorm and RoPE #5990 integration;
  • latest-main MiniMax-H3 plus fused-QK/RoPE suites: 152 passed, 3 skipped;
  • focused MiniMax-H3 suite: 147 passed, 3 skipped, 1 deselected;
  • full-model BF16/quantization quality: 1 passed; LPIPS 0.0806 (< 0.2),
    PSNR 25.5733 dB, MAE 0.028717, audio cosine 0.9801 (> 0.8), audio RMS
    1.0842; BF16 peak 68.84 GiB versus quantized 53.34 GiB;
  • real 8x B300 Ulysses-8 T2VA smoke: 1 passed and produced video/audio output;
  • fixed-workload B -> candidate -> B performance and output-quality gates:
    passed, with the samples and results reported above.

The first full-model invocation exposed a missing lpips package in the fresh
B300 venv. After installing the repository's already-declared lpips==0.1.4
development dependency with uv, the test passed as reported. Remaining
warnings are upstream dependency deprecations and existing Qwen docstring
warnings.

AI assistance and submitter accountability

OpenAI Codex assisted with implementation, tests, B300 benchmark orchestration,
and this PR description. The human submitter has reviewed every changed line,
including the latest-main conflict resolution, understands the change end to
end, and is prepared to defend it in review.

Co-authored-by: david6666666 <530634352@qq.com>

Assisted-by: OpenAI Codex
Signed-off-by: mokashliu <mokashliu@tencent.com>
@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

Latest-main compatibility update: upstream/main advanced from dbc0dd6d to 90a08c0e via the disjoint dots.tts merge #4765. I synthesized 90a08c0e + 45194337 in an isolated worktree; the cherry-pick was conflict-free, changed-file pre-commit passed every hook, both changed Python files compiled, and git diff --check passed. The changed-file SHA-256 values remain ada9835f...1647 (transformer) and 1308c0de...f200 (tests), matching the exact files used for the reported B300 qualification. No public-branch rewrite is needed.

@mo-ke-ke
mo-ke-ke marked this pull request as ready for review August 13, 2026 18:13
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: MinimaxH3.

Model owners: @david6666666

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

@vllm-omni-review-bot please review the strict-Ulysses activation guard, rank-local image/audio/text row reconstruction, RoPE/AdaLN span alignment, compact final-projection-before-gather path, and exact fallback behavior for SP1, Ring, AllGather-KV, advanced UAA, missing hooks, and non-divisible sequence lengths.

@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization labels Aug 14, 2026
Resolve MiniMax-H3 fused QK/RoPE integration while preserving rank-local strict-Ulysses row materialization.

Assisted-by: OpenAI Codex

Signed-off-by: mokashliu <mokashliu@tencent.com>

@david6666666 david6666666 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. add label to pass CI next

@david6666666 david6666666 added ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI labels Aug 14, 2026
@david6666666 david6666666 added ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI and removed ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI labels Aug 14, 2026
Assisted-by: OpenAI Codex
Signed-off-by: mokashliu <mokashliu@tencent.com>
@david6666666 david6666666 added ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI and removed ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI labels Aug 14, 2026
@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

Follow-up for the failing Simple · Diffusion Test job:

  • Root cause: the MiniMax-H3 TeaCache extractor called _embed() without the newly required local_span.
  • Fix: pass the explicit full-sequence TeaCache boundary, local_span=(0, seq_len).
  • Existing TeaCache regression tests reproduced the failure before the fix: 7 failed, 1 passed, 18 deselected.
  • The same targeted tests passed after the fix: 8 passed, 18 deselected.
  • Broader B300 regression suite: 187 passed, 3 skipped.
  • Changed-file pre-commit, Python compile, and git diff --check passed.

The human submitter reviewed and understood the added line before commit/push. AI assistance was used for diagnosis, implementation, and testing.

@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

@david6666666 The branch was updated to the latest main at head 1c550507, and the GitHub build, pre-commit, DCO, and docs checks passed. However, no Buildkite contexts were attached to this new head even though the ready label is still present. Could you please retrigger Buildkite (for example by toggling the ready label), or ask a CI maintainer to do so? We need the new main GPU run to validate the TeaCache compatibility fix on the final head. The recurring Intel failure is a separate pre-existing collection error (ModuleNotFoundError: examples.offline_inference) and does not execute the MiniMax-H3 changes.

@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

Adding the CI CODEOWNERS for help: @yenuo26 @congw729 @NickCao. After the branch update to head 1c550507, the ready label remains present but no CUDA/AMD/Intel/NPU Buildkite contexts were attached. The GitHub build, pre-commit, DCO, and docs checks are green. Could one of you please retrigger the Buildkite pipelines by toggling ready or manually rerunning them? The author has only read permission on the upstream repository and cannot toggle the label.

@david6666666 david6666666 added ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI and removed ready label to trigger buildkite CI diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI labels Aug 14, 2026
@mo-ke-ke

Copy link
Copy Markdown
Contributor Author

CI triage for final head 17a1b45c:

  • CUDA Buildkite #13620 passed, including the Simple · Diffusion Test that previously exposed the MiniMax-H3 TeaCache local_span incompatibility. This confirms the follow-up fix in upstream CI.
  • AMD #10584 failed only in unrelated LTX2 phase-adapter tests because vllm::rocm_unquantized_gemm was invoked with the CPU backend; no MiniMax-H3 test failed.
  • Intel [2/N] Add a minimal temporal chunk callback for MiniMax-H3 #7017 repeated the pre-existing collection errors for examples.offline_inference; no MiniMax-H3 test executed.
  • NPU [Bug]: AsyncOmni.generate ignores lora_request during generation #5369 failed only in Wan2.2 I2V startup because the model snapshot was absent while outgoing traffic was disabled; no MiniMax-H3 test failed.

No PR code change is indicated by these remaining failures. @yenuo26 @congw729 @NickCao could you please retry or classify the unrelated hardware failures so the PR can proceed?

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@hsliuustc0106
hsliuustc0106 merged commit 7b76b64 into vllm-project:main Aug 15, 2026
6 of 9 checks passed
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…llm-project#6173)

Signed-off-by: mokashliu <mokashliu@tencent.com>
Co-authored-by: mokashliu <mokashliu@tencent.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models diffusion-x2v-test label to trigger buildkite x2video series of diffusion models test in nightly CI Kernel optimization Codes related to optimize kernel execution to improve hardware utilization ready label to trigger buildkite CI

4 participants