[Model][Performance] Optimize MiniMax-H3 strict Ulysses boundaries - #6173
Conversation
Co-authored-by: david6666666 <530634352@qq.com> Assisted-by: OpenAI Codex Signed-off-by: mokashliu <mokashliu@tencent.com>
|
Latest-main compatibility update: |
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to be related to model: MinimaxH3. Model owners: @david6666666 Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
@vllm-omni-review-bot please review the strict-Ulysses activation guard, rank-local image/audio/text row reconstruction, RoPE/AdaLN span alignment, compact final-projection-before-gather path, and exact fallback behavior for SP1, Ring, AllGather-KV, advanced UAA, missing hooks, and non-divisible sequence lengths. |
Resolve MiniMax-H3 fused QK/RoPE integration while preserving rank-local strict-Ulysses row materialization. Assisted-by: OpenAI Codex Signed-off-by: mokashliu <mokashliu@tencent.com>
david6666666
left a comment
There was a problem hiding this comment.
LGTM. add label to pass CI next
Assisted-by: OpenAI Codex Signed-off-by: mokashliu <mokashliu@tencent.com>
|
Follow-up for the failing
The human submitter reviewed and understood the added line before commit/push. AI assistance was used for diagnosis, implementation, and testing. |
|
@david6666666 The branch was updated to the latest |
|
Adding the CI CODEOWNERS for help: @yenuo26 @congw729 @NickCao. After the branch update to head |
|
CI triage for final head
No PR code change is indicated by these remaining failures. @yenuo26 @congw729 @NickCao could you please retry or classify the unrelated hardware failures so the PR can proceed? |
…llm-project#6173) Signed-off-by: mokashliu <mokashliu@tencent.com> Co-authored-by: mokashliu <mokashliu@tencent.com> Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Purpose
MiniMax-H3's strict Ulysses path previously built the complete packed
[S, 5376]embedding and RoPE tensors on every rank, then immediately selectedthat rank's
S / world_sizerows. After the transformer blocks it also gatheredthe full BF16 hidden state before applying the final 128-channel projection.
This PR keeps those two model boundaries rank-local when, and only when, the
registered
sp_input---local_sp_preparehook proves the strict-Ulysses layout:the token refiner on its required full text context before selecting rank-owned
text rows, construct only the local embedding rows, and slice matching RoPE
rows;
[S, 128]FP32 logits instead of[S, 5376]BF16 hidden states;hybrid/advanced layouts, missing hooks, non-forward contexts, or shapes that
cannot be divided evenly.
For MiniMax-H3, the final all-gather payload is reduced from 10,752 bytes to
512 bytes per packed row (21x smaller). This is a boundary-payload reduction,
not a 21x end-to-end performance claim.
The implementation deliberately does not change Ulysses Q/K/V all-to-all,
the attention backend, packed-prefix handling, scheduler semantics, or model
weights. Focused tests exercise the optimized contract and every fallback above.
8x B300 performance and quality
Hardware and software:
dbc0dd6dTRTLLM_ATTNFixed workload:
torch.compile; one excluded warmup and five measured generationsfor each engine; B1 (main) -> candidate -> B2 (main) run order
Main diffusion samples were
[16.7780, 16.8498, 16.8513, 16.8456, 16.8471, 16.8226, 16.8169, 16.8514, 16.8806, 16.8583]seconds. Candidate samples were[16.5572, 16.5727, 16.5562, 16.5878, 16.5584]seconds. The B1 and B2medians were 16.8471 s and 16.8514 s, respectively, which guards against a
one-directional thermal or clock drift explanation. This is a batch-size-1
latency optimization; no throughput or concurrency-scaling claim is made.
The numerical order changes because the final projection now precedes the
all-gather, so output hashes are not expected to match the baseline bit for bit.
Every repeat within each build was deterministic. Candidate versus baseline:
MAE 0.000831;
all passed.
Latest-main compatibility was requalified after main advanced to
596c16a5(including fused Q/K RMSNorm and RoPE from #5990). The conflict resolution
keeps the fused RoPE table construction after selecting this rank's strict-
Ulysses position rows. On the same 8x B300 allocation, the combined MiniMax-
H3 and fused-QK/RoPE suites passed (
152 passed, 3 skipped), and an exact2-step Ulysses-8 T2VA smoke matched the latest-main video bit for bit
(SSIM 1.0, MAE 0) with audio STFT cosine approximately 1.0 and RMS ratio
1.0000045. Peak memory was 89,268 MiB/rank on latest main versus 88,048
MiB/rank for the candidate. These are compatibility-smoke results, not a
replacement formal performance claim for the B -> candidate -> B results
above.
The benchmark used the same
Omni.generatecontract as the checked-inMiniMax-H3 E2E test. This PR does not add a second model-specific benchmark
framework because draft #5852 already owns the generic MiniMax-H3 2/4/8-GPU
benchmark and SM120 work.
Duplicate-work and attribution check
Before opening this draft I inspected issue #5700 and searched open PRs by the
roadmap issue, MiniMax-H3/Ulysses/SP-boundary keywords, the new hook name, and
the affected model file.
Q/K normalization/RoPE, respectively.
model-local and preserves the public SP topology/fallback behavior.
I posted the exact boundary scope there and asked for conflict disclosure;
there is no overlapping PR at the time this draft is opened.
The closed, unmerged draft #5750 contained independently reviewable versions
of the local-embedding and compact-gather ideas. This change reimplements them
against current main, tightens activation to the exact registered strict-
Ulysses hook, adds current fallback/contract coverage, and carries
Co-authored-by: david6666666 <530634352@qq.com>in the commit.Test Plan
vLLM Version: 0.27.0
vLLM-Omni Commit:
092b8c7d(latest-main compatibility base596c16a5; formal performance basedbc0dd6d)Test Result
by the [Kernel] Fuse Q/K RMSNorm and RoPE #5990 integration;
PSNR 25.5733 dB, MAE 0.028717, audio cosine 0.9801 (> 0.8), audio RMS
1.0842; BF16 peak 68.84 GiB versus quantized 53.34 GiB;
passed, with the samples and results reported above.
The first full-model invocation exposed a missing
lpipspackage in the freshB300 venv. After installing the repository's already-declared
lpips==0.1.4development dependency with
uv, the test passed as reported. Remainingwarnings are upstream dependency deprecations and existing Qwen docstring
warnings.
AI assistance and submitter accountability
OpenAI Codex assisted with implementation, tests, B300 benchmark orchestration,
and this PR description. The human submitter has reviewed every changed line,
including the latest-main conflict resolution, understands the change end to
end, and is prepared to defend it in review.