Skip to content

MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3 - #315

Merged
ywang96 merged 40 commits into
vllm-project:mainfrom
hsliuustc0106:codex-minimax-h3-production-blog
Sep 1, 2026
Merged

ywang96 merged 40 commits into
vllm-project:mainfrom
hsliuustc0106:codex-minimax-h3-production-blog

Conversation

@hsliuustc0106

@hsliuustc0106 hsliuustc0106 commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • tells one serving journey: MiniMax H3 system bottlenecks → broadly applicable vLLM-Omni runtime improvements → scalable general-H3 architecture → FastVideo FastH3 four-step acceleration → real-time evidence → production boundaries
  • dates the post September 1, 2026, with the generated URL under /2026/09/01/
  • preserves the 8× B300 base-H3 comparison: Diffusers at 82.239 s versus vLLM-Omni at 56.917 s complete-response E2E, a 30.8% latency reduction and 1.445× speedup
  • adds a measured BF16 DLO latency–HBM Pareto frontier for general H3 serving
  • adds an isolated resident BF16-versus-online-FP8 B300 comparison with stage latency and peak-HBM evidence
  • updates the TRTLLM SAGE and SAGE + Skip-Softmax rows to the measured K16 configuration, with model-execution time, speedup, LPIPS, and output samples
  • keeps SVDQuant, Sol-Attn, and Cache-DiT visible in the main optimization narrative without presenting unqualified latency gains
  • reports the FastH3 5/10/15-second duration sweep; the measured 10.125-second MP4 completes in 8.678–8.710 seconds
  • defines “real-time” as complete-MP4 generation faster than playback, not streaming or time to first frame
  • pins the FastH3 artifact revision, size, SHA-256 checksum, topology, launch command, and request
  • presents supplied 5/10/15-second outputs in one direct-link table and sends other hardware to maintained recipes
  • adds VeRL-Omni, UniRL, and RLinf post-training integration to future work and linked GitHub acknowledgments for the contributor groups
  • adds special thanks to Hongsheng Liu and Roger Wang for general support and blog preparation
  • quantifies the chunkwise output-pipeline roadmap with B300 ceilings/targets and the caveated four-L20X PR #6885 feasibility result

The post is approximately 4.3k words. The latest DLO, online-FP8, and attention contributions are integrated in signed commit 4b7e40b, with GitHub-recognized co-authorship for @lishunyang12, @david6666666, and @bobboli.

Newly integrated B300 evidence

  • DLO Pareto frontier: BF16 FL2VA, 8× B300, 5.175-second 1344×768 output, SP8/Ulysses8/Ring1/DP1/TP1, AllGather, and CUDNN attention. At 35 resident DiT blocks, reported HBM falls by 37.5% for a 5.1% latency cost; zero resident blocks is the minimum-memory endpoint.
  • Online FP8: under a fixed resident 10-second profile, BF16 measures 52.572 s stage / 53.118 s offline E2E / 87.16 GiB per rank, while online FP8 measures 49.769 s / 50.331 s / 53.27 GiB. This is 5.3% lower stage time and 38.9% lower peak HBM.
  • TRTLLM attention: K16 SAGE measures 44.787 s, 1.211×, LPIPS 0.3697; K16 SAGE + conservative Skip-Softmax measures 43.867 s, 1.237×, LPIPS 0.3750. Dense and Skip-Softmax-only controls remain in the table.
  • Kernel provenance: the fused modulation/normalization/residual discussion now cites vLLM-Omni PR #6878 in addition to the earlier kernel work.

These rows use different declared timing boundaries and workloads. Their gains are not added together and are not mixed into the FastH3 headline result.

Evidence boundaries

  • Base-H3 runtime A/B: vLLM-Omni b81aeb7, fixed official T2VA prompt and seed 0, 50 sigma positions / 49 DiT forwards, complete-MP4 client E2E.
  • DLO: BF16 FL2VA Pareto study with one excluded first request and two measured requests per policy.
  • Online FP8: three measured requests after one warmup; offline E2E ends at returned video/audio tensors and excludes MP4 muxing. Distinct seeds establish successful generation and output shape, not pixelwise BF16 equivalence.
  • Approximate attention: model-execution boundary with same-workload dense controls and LPIPS/output samples.
  • FastH3: merged integration 86b85c07, fixed prompt and seed 1101, 5 sigma positions / 4 DiT forwards, absolute complete-response latency only.
  • No base-to-FastH3 speedup is inferred because the valid experiments do not share source SHA, prompt, seed, and artifact.

Publication status

The PR remains a draft. The summarized measurements are populated, but publication gates remain.

  • preserve the Diffusers-versus-vLLM-Omni B300 result
  • add the BF16 DLO latency–memory frontier
  • add the resident BF16-versus-online-FP8 capacity/latency comparison
  • update SAGE and SAGE + Skip-Softmax to the K16 measurements and samples
  • preserve the FastH3 critical-path decomposition and 5/10/15-second sweep
  • include generated-video links and explicit media-validation gates
  • pin the FastH3 artifact and merged vLLM-Omni integration
  • preserve contributor attribution while restoring DCO compliance
  • treat issue #5700 as a potentially lagging tracker rather than authoritative evidence
  • publish stable raw evidence bundles for the populated B300 tables, including samples, logs, environment/topology manifests, and media metadata/hashes
  • complete a matched base-versus-FastH3 multi-seed quality comparison before making a parity claim
  • add linked GitHub acknowledgments for named contributors
  • remove published: false when maintainers approve publication

Validation

  • ruby scripts/check-blog-summaries.rb
  • git diff --check
  • bundle exec jekyll build --unpublished
  • supplied MP4s validated as H.264 video plus stereo 32 kHz AAC
  • five-second RTF/inverse factors recomputed from the displayed T_media
  • rewritten DCO commit has an identical tree to the two integrated contributor commits

Signed-off-by: hsliuustc0106 liuhongsheng4@huawei.com

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106

hsliuustc0106 commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Contributor coordination checklist

Scope update (2026-08-31): after the team review, this article now uses 8× B300 only. H200 and RTX PRO 5000 article benchmarks, including the PRO 5000 DLO Pareto study, are no longer requested. Those platforms remain covered by the full H3 recipe and RTX PRO 5000 recipe.

0. Shared workload

  • T2VA; no reference media; official model-card case-T2VA expanded prompt
  • Prompt SHA-256 98f36b879692095e099ae824c18d9e93e7006a490e082fd474a5f531769dcf06; seed 0
  • 10.0 seconds requested; 1344×768; 243 aligned frames at 24 FPS; H.264 + stereo 32 kHz AAC
  • Base: 50 sigma points / 49 expected DiT forwards; few-step paths must report both sigma points and actual forwards
  • Pin every B300 test to vLLM-Omni 86b85c078bc041e04aee4c4d9167fb10fb1994c7, the merged commit for #6714; it includes #6776 + #6824

1. Claim one B300 result bundle

Please reply with the bundle, owner, exact B300 system, and environment availability.

  • Lossless stack: Diffusers vs vLLM-Omni, plus the selected dense attention backend, regular vs Fast Ulysses, fused DiT, VAE, GPU output transport, and CPU MP4 boundaries
  • FastH3 low-latency serving: the pinned 8×B300 USP8/VAE-PP8 profile now includes its stage decomposition and 5/10/15-second complete-MP4 RTF sweep. Cross-path multi-seed quality evidence and a stable raw-artifact link remain publication gates
  • Optional acceleration: online FP8, SVDQuant, or sparse/skip attention as an isolated A/B with quality evidence; claim only the path you can complete

Every result must follow the stage-timing and parallelism addendum: encoder; denoise total, actual forward count, and per-forward time; video/audio VAE; transport; CPU MP4 wall/process CPU; client E2E/residual; and the explicit devices/parallelism for every stage.

Required evidence: exact SHAs and commands, preparation versus readiness time, one excluded full-shape warmup, two measured repetitions per claimed single-request A/B, HBM/host RAM, validated MP4 metadata, same-seed quality evidence, raw logs, and a stable artifact link.

Stop before repeated measurement after OOM, accelerator error, invalid media, missing audio, a failed quality gate, or an unexpected fallback. Keep prompt, seed, shape, schedule, topology, readiness, cache state, and server lifecycle fixed within each A/B.

2. Attribution

For data, text, figures, or substantive review integrated into the post:

  • provide the preferred public display name
  • provide a GitHub-associated email or GitHub no-reply email
  • confirm consent to a Co-authored-by: trailer
  • include your own DCO Signed-off-by: trailer when pushing a commit directly

Source-area contacts: base H3 @Isotr0py; DLO/base integration @lishunyang12 @evanchueng @Gaohan123 @david6666666; encoder, attention, kernels, quantization, VAE, and media @gcanlin @yuanwu2017 @bobboli @fan2956 @mo-ke-ke @mglyn @MosCloud @ultism; FastH3 @princepride; VeRL-Omni @NancyFyong @mengchengTang.

Please keep final artifact links in this PR so the benchmark provenance stays auditable.

@hsliuustc0106

hsliuustc0106 commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Benchmark reporting addendum: stage timing and parallelism

Please treat this as part of the contributor checklist above. A benchmark result is incomplete unless it reports both the timing boundary and the device placement/parallelism for every stage.

Required stage timings

  • Encoder: preparation time and encoder wall time. For disaggregated encoding, report Stage 0 compute and handoff wait separately.
  • DiT denoise: total denoise wall time, requested sigma points, actual DiT forward count, and denoise_wall_time / actual_dit_forwards.
  • Video VAE: video decode wall time and the maximum/critical-path value across VAE ranks.
  • Audio VAE: audio decode wall time separately when instrumentation permits. If only aggregate VAE time exists, label it as aggregate.
  • Output transport: D2H, worker-to-engine, and inter-stage handoff wall times where applicable.
  • CPU MP4: encode/mux wall time, process CPU time, and peak RSS. Do not fold D2H/IPC into this number unless the boundary explicitly includes it.
  • Client E2E: request submission through receipt of the complete response body.
  • Residual: client_e2e - reported_stage_times, with an explanation for material unaccounted time.

Important: denoise per-step time must divide by the actual DiT forward count, not num_inference_steps. Always report both values, especially for Turbo and FastH3.

Required stage placement and parallelism

  • Encoder: device IDs, TP, DP/replicas, offload state, prefix-cache state, and attention backend.
  • DiT: device IDs and process groups; TP, Ulysses, Ring, DP, CFG, PP/HSDP; DLO mode and resident-layer count; attention backend; eager/compile mode.
  • Video VAE: device IDs, VAE patch-parallel size, parallel mode, tiling, and process group.
  • Audio VAE: device IDs and whether execution is rank-local, replicated, or sharded.
  • Output: source/destination ranks, D2H/SHM/IPC path, payload dtype/size, CPU model, socket/NUMA affinity, thread count, and PyAV/FFmpeg versions.

Please include this fill-in manifest with every submitted profile:

profile: TBD
hardware: TBD

parallelism:
  encoder:
    devices: TBD
    tp: TBD
    dp_or_replicas: TBD
    offload: TBD
    prefix_cache: TBD
    attention_backend: TBD
  dit:
    devices: TBD
    tp: TBD
    ulysses: TBD
    ring: TBD
    dp: TBD
    cfg: TBD
    pp_or_hsdp: TBD
    dlo_mode: TBD
    dlo_resident_layers: TBD
    attention_backend: TBD
    execution: eager_or_compile
  video_vae:
    devices: TBD
    patch_parallel_size: TBD
    mode: TBD
    tiling: TBD
  audio_vae:
    devices: TBD
    placement: rank_local_replicated_or_sharded
  output:
    transport: TBD
    cpu_affinity: TBD
    threads: TBD

timing_ms:
  prepare_encoder: TBD
  encoder: TBD
  encoder_handoff_wait: TBD
  denoise_total: TBD
  requested_sigma_points: TBD
  actual_dit_forwards: TBD
  denoise_per_forward: TBD
  video_vae: TBD
  audio_vae: TBD
  output_transport: TBD
  cpu_mp4_wall: TBD
  cpu_mp4_process_cpu: TBD
  client_e2e: TBD
  residual: TBD

Please attach the raw profiler/log source for every populated field rather than copying only the final summary.

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@mo-ke-ke

Copy link
Copy Markdown

@hsliuustc0106 I’m interested in taking the 8× H200 common T2VA baseline, pending the frozen prompt, seed, revisions, quality gates, and artifact destination. I’ll first verify that the published MiniMax H3/vLLM-Omni configuration starts cleanly on the available 8× H200 node. Once confirmed, I can provide the requested topology, stage-timing, output-validation, and raw-artifact bundle.

@hsliuustc0106

hsliuustc0106 commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 I’m interested in taking the 8× H200 common T2VA baseline, pending the frozen prompt, seed, revisions, quality gates, and artifact destination. I’ll first verify that the published MiniMax H3/vLLM-Omni configuration starts cleanly on the available 8× H200 node. Once confirmed, I can provide the requested topology, stage-timing, output-validation, and raw-artifact bundle.

thanks for help us provide this DP, we expect this blog to be publish next monday. Not sure whether you can catch up the time:)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106

Copy link
Copy Markdown
Contributor Author

Mechanism diagrams added for technical review

The draft now includes four mechanism figures:

  1. Turbo LoRA vs FastH3 — new editable H3-specific SVG based on #6476, #6550, and #6714. Review requested from @mglyn, @princepride, and @lishunyang12.
  2. DLO double-buffer execution — reuses the existing official DLO H2D/AllGather/compute timeline and links back to the published DLO post.
  3. Encoder disaggregation — new editable H3-specific SVG adapted from #5885 and RFC #5707. Review requested from @gcanlin and @yuanwu2017.
  4. Online FP8 vs SVDQuant W4A4 — new editable H3-specific SVG adapted from the source-backed cookbook explainers and #5910/#6162. Review requested from @lishunyang12 and @ultism.

Please check the diagrams for semantic or compatibility-boundary errors, especially:

  • Turbo request-switchability and DLO-resident A/B sidecars
  • FastH3 load-time fusion, dense-T2VA scope, VSA boundary, and offload rejection
  • Online FP8 load-time versus per-forward scaling behavior
  • SVDQuant W4A4 base branch plus BF16 low-rank correction
  • Encoder/DiT ownership, conditioning bridge, original-media bypass, and independent parallelism
  • DLO overlap terminology and slot lifecycle

The new assets contain no benchmark claims. Each caption links to its exact PR/RFC/blog/cookbook provenance, and all SVGs are editable under assets/figures/2026-08-29-minimax-h3-production-serving/.

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106

Copy link
Copy Markdown
Contributor Author

Updated the SVDQuant diagram arrow layout in 84e36a1:

  • split checkpoint tensors into explicit R4 + scales and U/V routes
  • split activation into explicit A4 and BF16 x routes
  • replaced overlapping curves with orthogonal, non-crossing input paths
  • moved the branch merge to a side-mounted sum node
  • routed the 4-bit and BF16 branch outputs into separate merge ports

The revised SVG was raster-previewed and rebuilt with the unpublished post before pushing.

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106

Copy link
Copy Markdown
Contributor Author

Updated the H3 encoder-disaggregation figure in 4f60281:

  • redrew the complete serving chain as a strict left-to-right flow
  • made the original prompt/media bypass a separate left-to-right lane
  • removed upward/crossing control-flow arrows; request overlap is now annotation-only
  • verified the merged #5885 recipe and code path before naming the handoff
  • explicitly documents that the current single-node topology uses an orchestrator payload merge plus InlineStageDiffusionClient
  • explicitly documents that OmniConnector is not configured in the current deployment
  • labels OmniConnector (SHM/RDMA) as the RFC #5707 future/cross-node option, outside the current recipe

The updated diagram and article text were raster-previewed and rebuilt before pushing.

@david6666666

Copy link
Copy Markdown
Contributor

should add Diffusion continuous batching 与 step-boundary abort this feature vllm-project/vllm-omni#5810

@david6666666

Copy link
Copy Markdown
Contributor

Cache-DiT dynamic quality grading #5853

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106

Copy link
Copy Markdown
Contributor Author

Added a model-pipeline overview at the top of the draft in 81b61df.

The new editable SVG introduces:

  • multimodal inputs and task selection
  • shared H3/Qwen3-VL, Visual VAE, and Audio VAE encoding
  • the unified packed sequence with conditioning and noisy target A/V latents
  • FL2VA versus Ref2VA DiT partition selection
  • joint video/audio denoising with dual sigma schedules
  • separate video/audio VAE decode and final CPU MP4 mux
  • combined-serving reuse of the tokenizer, processor, H3 encoder, and VAEs

Sources were checked against the official MiniMax H3 model card, the vLLM-Omni MiniMax-H3 recipe and packed-sequence implementation, and the Diffusers H3 pipeline description. @Isotr0py, please flag any model-pipeline semantic issue. Existing figure captions were renumbered from 1–4 to 2–5.

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@mo-ke-ke

Copy link
Copy Markdown

Yes, I can catch up with the Monday timeline. Please tag me once the prompt, seed, frozen revisions, quality gates, and artifact destination are finalized. I’ll start the 8× H200 feasibility run immediately and keep the results and artifact links in this PR.

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review September 1, 2026 04:27
@hsliuustc0106
hsliuustc0106 force-pushed the codex-minimax-h3-production-blog branch from 22a8634 to 86fbf99 Compare September 1, 2026 04:29
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Co-authored-by: SYLAR <125541396+lishunyang12@users.noreply.github.com>
Co-authored-by: david6666666 <530634352@qq.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Comment thread _posts/2026-09-01-minimax-h3-production-serving.md
Comment thread _posts/2026-09-01-minimax-h3-production-serving.md
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review September 1, 2026 04:37
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Comment thread _posts/2026-09-01-minimax-h3-production-serving.md
* Clarify the sparse attention narrative

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* Clarify the Skip-Softmax quality tradeoff

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* Simplify the attention results introduction

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* Clarify the attention baseline comparison

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

---------

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Comment thread _posts/2026-08-29-minimax-h3-production-serving.md Outdated
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@hsliuustc0106
hsliuustc0106 marked this pull request as draft September 1, 2026 07:40
@hsliuustc0106
hsliuustc0106 marked this pull request as ready for review September 1, 2026 07:42
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
@ywang96
ywang96 merged commit c8b3f20 into vllm-project:main Sep 1, 2026
4 checks passed

This branch was successfully deployed

1 active deployment
Preview — b8546e0d Deployed Sep 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

8 participants