MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3 - #315
Conversation
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Contributor coordination checklistScope update (2026-08-31): after the team review, this article now uses 8× B300 only. H200 and RTX PRO 5000 article benchmarks, including the PRO 5000 DLO Pareto study, are no longer requested. Those platforms remain covered by the full H3 recipe and RTX PRO 5000 recipe. 0. Shared workload
1. Claim one B300 result bundlePlease reply with the bundle, owner, exact B300 system, and environment availability.
Every result must follow the stage-timing and parallelism addendum: encoder; denoise total, actual forward count, and per-forward time; video/audio VAE; transport; CPU MP4 wall/process CPU; client E2E/residual; and the explicit devices/parallelism for every stage. Required evidence: exact SHAs and commands, preparation versus readiness time, one excluded full-shape warmup, two measured repetitions per claimed single-request A/B, HBM/host RAM, validated MP4 metadata, same-seed quality evidence, raw logs, and a stable artifact link. Stop before repeated measurement after OOM, accelerator error, invalid media, missing audio, a failed quality gate, or an unexpected fallback. Keep prompt, seed, shape, schedule, topology, readiness, cache state, and server lifecycle fixed within each A/B. 2. AttributionFor data, text, figures, or substantive review integrated into the post:
Source-area contacts: base H3 @Isotr0py; DLO/base integration @lishunyang12 @evanchueng @Gaohan123 @david6666666; encoder, attention, kernels, quantization, VAE, and media @gcanlin @yuanwu2017 @bobboli @fan2956 @mo-ke-ke @mglyn @MosCloud @ultism; FastH3 @princepride; VeRL-Omni @NancyFyong @mengchengTang. Please keep final artifact links in this PR so the benchmark provenance stays auditable. |
Benchmark reporting addendum: stage timing and parallelismPlease treat this as part of the contributor checklist above. A benchmark result is incomplete unless it reports both the timing boundary and the device placement/parallelism for every stage. Required stage timings
Important: denoise per-step time must divide by the actual DiT forward count, not Required stage placement and parallelism
Please include this fill-in manifest with every submitted profile: profile: TBD
hardware: TBD
parallelism:
encoder:
devices: TBD
tp: TBD
dp_or_replicas: TBD
offload: TBD
prefix_cache: TBD
attention_backend: TBD
dit:
devices: TBD
tp: TBD
ulysses: TBD
ring: TBD
dp: TBD
cfg: TBD
pp_or_hsdp: TBD
dlo_mode: TBD
dlo_resident_layers: TBD
attention_backend: TBD
execution: eager_or_compile
video_vae:
devices: TBD
patch_parallel_size: TBD
mode: TBD
tiling: TBD
audio_vae:
devices: TBD
placement: rank_local_replicated_or_sharded
output:
transport: TBD
cpu_affinity: TBD
threads: TBD
timing_ms:
prepare_encoder: TBD
encoder: TBD
encoder_handoff_wait: TBD
denoise_total: TBD
requested_sigma_points: TBD
actual_dit_forwards: TBD
denoise_per_forward: TBD
video_vae: TBD
audio_vae: TBD
output_transport: TBD
cpu_mp4_wall: TBD
cpu_mp4_process_cpu: TBD
client_e2e: TBD
residual: TBDPlease attach the raw profiler/log source for every populated field rather than copying only the final summary. |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
|
@hsliuustc0106 I’m interested in taking the 8× H200 common T2VA baseline, pending the frozen prompt, seed, revisions, quality gates, and artifact destination. I’ll first verify that the published MiniMax H3/vLLM-Omni configuration starts cleanly on the available 8× H200 node. Once confirmed, I can provide the requested topology, stage-timing, output-validation, and raw-artifact bundle. |
thanks for help us provide this DP, we expect this blog to be publish next monday. Not sure whether you can catch up the time:) |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Mechanism diagrams added for technical reviewThe draft now includes four mechanism figures:
Please check the diagrams for semantic or compatibility-boundary errors, especially:
The new assets contain no benchmark claims. Each caption links to its exact PR/RFC/blog/cookbook provenance, and all SVGs are editable under |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
|
Updated the SVDQuant diagram arrow layout in
The revised SVG was raster-previewed and rebuilt with the unpublished post before pushing. |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
|
Updated the H3 encoder-disaggregation figure in
The updated diagram and article text were raster-previewed and rebuilt before pushing. |
|
should add Diffusion continuous batching 与 step-boundary abort this feature vllm-project/vllm-omni#5810 |
|
Cache-DiT dynamic quality grading #5853 |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
|
Added a model-pipeline overview at the top of the draft in The new editable SVG introduces:
Sources were checked against the official MiniMax H3 model card, the vLLM-Omni MiniMax-H3 recipe and packed-sequence implementation, and the Diffusers H3 pipeline description. @Isotr0py, please flag any model-pipeline semantic issue. Existing figure captions were renumbered from 1–4 to 2–5. |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
|
Yes, I can catch up with the Monday timeline. Please tag me once the prompt, seed, frozen revisions, quality gates, and artifact destination are finalized. I’ll start the 8× H200 feasibility run immediately and keep the results and artifact links in this PR. |
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
22a8634 to
86fbf99
Compare
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> Co-authored-by: SYLAR <125541396+lishunyang12@users.noreply.github.com> Co-authored-by: david6666666 <530634352@qq.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
86fbf99 to
4b7e40b
Compare
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
* Clarify the sparse attention narrative Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * Clarify the Skip-Softmax quality tradeoff Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * Simplify the attention results introduction Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * Clarify the attention baseline comparison Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> --------- Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Summary
/2026/09/01/The post is approximately 4.3k words. The latest DLO, online-FP8, and attention contributions are integrated in signed commit
4b7e40b, with GitHub-recognized co-authorship for @lishunyang12, @david6666666, and @bobboli.Newly integrated B300 evidence
These rows use different declared timing boundaries and workloads. Their gains are not added together and are not mixed into the FastH3 headline result.
Evidence boundaries
b81aeb7, fixed official T2VA prompt and seed 0, 50 sigma positions / 49 DiT forwards, complete-MP4 client E2E.86b85c07, fixed prompt and seed 1101, 5 sigma positions / 4 DiT forwards, absolute complete-response latency only.Publication status
The PR remains a draft. The summarized measurements are populated, but publication gates remain.
published: falsewhen maintainers approve publicationValidation
ruby scripts/check-blog-summaries.rbgit diff --checkbundle exec jekyll build --unpublishedT_mediaSigned-off-by: hsliuustc0106 liuhongsheng4@huawei.com