Skip to content

[Diffusion][Quantization] Enable MiniMax-H3 global FP8 with DLO - #5910

Merged
hsliuustc0106 merged 12 commits into
vllm-project:mainfrom
lishunyang12:feat/minimax-h3-encoder-fp8-dlo
Aug 15, 2026
Merged

hsliuustc0106 merged 12 commits into
vllm-project:mainfrom
lishunyang12:feat/minimax-h3-encoder-fp8-dlo

Conversation

@lishunyang12

@lishunyang12 lishunyang12 commented Aug 8, 2026 •

Copy link
Copy Markdown
Collaborator

What users get

This PR adds one user-facing MiniMax-H3 online FP8 mode:

  • --quantization fp8 (or quantization="fp8") quantizes all eligible DiT and Qwen3-VL text-decoder linears.
  • The same default works for T2VA, FL2VA/I2VA, and Ref2VA.
  • It can run fully resident with no offload, with ordinary layerwise offload, or with distributed layerwise offload (DLO) through the no-AllGather path.
  • No calibrated or pre-quantized checkpoint is required; the released BF16 checkpoint is quantized at runtime.

This intentionally exposes one simple default. It does not add a MiniMax-H3 layer-selection API or additional user-facing quantization modes.

Quick start

Single GPU, resident FP8, no offload

Use the FL2VA-only partition for the 96 GB single-GPU capacity test. This avoids loading the second Ref2VA DiT.

Terminal 1:

MODEL_ROOT=/path/to/MiniMax-H3 \
GPU_ID=0 \
PORT=8091 \
RUN_DIR=$PWD/h3_single_gpu_fp8_no_offload \
bash examples/online_serving/minimax_h3/run_server_single_gpu_fp8_no_offload.sh

Terminal 2, after the health check succeeds:

BASE_URL=http://127.0.0.1:8091 \
RUN_DIR=$PWD/h3_single_gpu_fp8_no_offload \
bash examples/online_serving/minimax_h3/run_curl_t2va_5s.sh

The server script intentionally passes none of --enable-cpu-offload, --enable-layerwise-offload, or --enable-distributed-layerwise-offload. It samples whole-device memory from before model initialization. The request script generates a fixed five-second, 1344x768, 50-step T2VA output and verifies H.264 video plus 32 kHz stereo AAC.

After the request succeeds, stop the server with Ctrl-C. The run directory contains:

  • gpu.csv and memory_summary.txt for whole-lifecycle capacity and headroom;
  • server.log for worker peak and encode/diffuse/decode timing;
  • client_time.txt, t2va.mp4, media.json, and validation.txt.

The complete guide is in recipes/MiniMaxAI/MiniMax-H3.md.

Python API

from vllm_omni import Omni

omni = Omni(
    model="/path/to/MiniMax-H3/FL2VA",
    quantization="fp8",
)

Distributed layerwise offload

For lower HBM residency, combine the same FP8 default with DLO's full-weight per-rank path:

vllm serve /path/to/MiniMax-H3/FL2VA \
  --omni \
  --quantization fp8 \
  --enable-distributed-layerwise-offload \
  --dlo-no-use-allgather

Runtime-created FP8 weights do not support the sharded DLO AllGather path. Use --dlo-no-use-allgather when combining online FP8 with DLO.

Precision scope

quantization="fp8" means global FP8 for eligible compute-heavy linears, not every tensor in the pipeline.

Component Precision
DiT attention, MLP, condition, and AdaLN linears Online FP8
Qwen3-VL text-decoder attention and MLP linears Online FP8
Qwen vision tower, embeddings, norms, and RoPE Checkpoint precision
Visual VAE and Audio VAE Checkpoint precision
FP32 patch, timestep, and output projections FP32

This keeps model-specific mixed-precision boundaries while providing the largest practical memory reduction through one default option.

Direct BF16 vs global FP8 videos

Matched five-second T2VA, I2VA, and Ref2VA pairs are available in the side-by-side player. The rankings comparison folder also provides individual MP4 links, per-task fidelity metrics, and provenance.

User-visible results

Single-B300 resident, eager, five-second T2VA at 1344x768 and 50 steps:

Configuration Wall time Worker peak
BF16 162.30 s 133,110 MiB
Default global FP8 151.58 s 92,146 MiB
Change 1.071x faster 40,964 MiB lower (30.8%)

The global-FP8 run's sampled whole-device first-request peak was 92,946 MiB. This is a capacity proxy for a 96 GB RTX PRO 6000, not a target-hardware validation; the scripts above record the real card's reported capacity and headroom.

The repeatable encoder-side benefit is primarily capacity, not latency. Isolated steady-state tests found no stable Qwen encoder speedup, while complete-pipeline resident memory fell from 121.10 GiB to 90.33 GiB at text-encoder TP1.

With DLO no-AllGather, a current-head single-B300 smoke generated valid five-second H.264/AAC output with an 18,062 MiB worker request peak and a 50,310 MiB sampled full-lifecycle peak. In the four-GPU benchmark, DLO startup streaming reduced the measured full-lifecycle peak from 58,696 to 37,938 MiB/GPU (35.4%).

The complete T2VA/I2VA/Ref2VA ablation results and generated videos are available here:

https://github.com/lishunyang12/vllm-omni-rankings/tree/main/scripts/minimax_h3_online_fp8

Across those three five-second, 50-step cases, default global FP8 averaged 1.066x speedup and 29.2% lower request-peak memory.

Implementation notes

  • TP-sharded encoder linears are quantized online one at a time, avoiding overlap between the complete BF16 encoder and its FP8 replacement.
  • Completed online-quantized DiT layers can stream back to CPU during DLO initialization.
  • Layerwise offload preserves the physical stride required by transposed Cutlass FP8 weights.
  • The default keeps non-eligible and explicitly mixed-precision components unchanged.

Validation

  • Focused CPU suite: 142 passed.
  • New serve and curl scripts: Bash syntax and safe invalid-model-path checks passed.
  • Single-B300 global-FP8 + DLO no-AllGather smoke: valid 124-frame H.264 video at 24 FPS plus 32 kHz stereo AAC.
  • Resident 50-step BF16/global-FP8 comparison and three-task ablation results are linked above.
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models quantization Code related to quantization labels Aug 8, 2026
@lishunyang12
lishunyang12 force-pushed the feat/minimax-h3-encoder-fp8-dlo branch from 9f92147 to b18eeff Compare August 8, 2026 14:06
@lishunyang12 lishunyang12 changed the title [Diffusion][Quantization] Support MiniMax-H3 encoder FP8 with DLO Aug 8, 2026
@lishunyang12

Copy link
Copy Markdown
Collaborator Author

Single-B300 encoder FP8 isolation

This isolates the incremental effect of quantizing the Qwen3-VL text decoder after the DiT is already FP8.

Workload

  • 1x NVIDIA B300 SXM6 AC (267.7 GiB)
  • TP1 / DP1 / Ulysses1 / Ring1 / text-encoder TP1 / VAE tile1
  • cuDNN attention, eager execution
  • T2VA, 1344x768, 5 seconds, 124 frames at 24 FPS, 32 kHz stereo audio
  • 50 denoising steps, seed 1101
  • no CPU or layerwise offload
Configuration Encode Denoise Decode End-to-end Worker peak
BF16/FP32 13.660 s 141.221 s 5.477 s 162.303 s 133,110 MiB
DiT FP8 + encoder BF16 12.925 s 132.919 s 5.424 s 153.165 s 101,620 MiB
Default global FP8 8.664 s 134.504 s 6.016 s 151.583 s 92,146 MiB

Incremental encoder result

Comparing default global FP8 with DiT-only FP8:

  • encode latency: 12.925 -> 8.664 s, a 33.0% reduction / 1.492x speedup
  • end-to-end latency: 153.165 -> 151.583 s, a 1.0% reduction / 1.010x speedup
  • first-request worker peak: 101,620 -> 92,146 MiB, saving 9,474 MiB (9.3%)

The small end-to-end speed change is expected: encoding runs once, while the DiT runs for 50 denoising steps.

During denoising, the 200 ms NVML trace showed stable device-memory plateaus of 101,030 MiB for DiT-only FP8 and 79,188 MiB for default global FP8, a difference of 21,842 MiB (21.6%). This plateau is reported as an observation, not as a second-request worker peak: the first global-FP8 request also performs online encoder quantization and its temporary workspace/cache raises the formal request peak.

Default FP8 + DLO compatibility

A separate current-head single-B300 smoke used global FP8 with no-AllGather DLO, 1344x768, 5 seconds, and 2 denoising steps:

  • 200 encoder linears quantized online
  • 208 DiT online-FP8 linears streamed back to CPU during loading
  • worker request peak: 18,062 MiB
  • full-lifecycle 200 ms NVML peak: 50,310 MiB
  • output validated as 5.207-second H.264 video plus 32 kHz stereo AAC

Conclusion: encoder FP8 provides meaningful single-GPU capacity/headroom and encode-stage speedup, but should not be presented as a large end-to-end throughput optimization.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@lishunyang12
lishunyang12 requested a review from ywang96 as a code owner August 8, 2026 15:47
@lishunyang12 lishunyang12 added the ready label to trigger buildkite CI label Aug 9, 2026
@lishunyang12

Copy link
Copy Markdown
Collaborator Author

Merged current upstream main in 2ac53dd and resolved the MiniMax-H3 recipe conflict. The merged limitation now reflects this PR: online FP8 with DLO uses rank-local streaming and does not support the sharded AllGather path. Post-merge validation passes: loader/H3 contracts 98 tests, DLO suites 57 tests, and pre-commit on all textual PR files. The PR is ready for review; platform CI has restarted.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

do we apply fp8 to all components? how do we select?

@lishunyang12

Copy link
Copy Markdown
Collaborator Author

The scope is now explicit and selectable. --quantization fp8 quantizes eligible DiT and Qwen3-VL text-decoder linears. It does not quantize the vision tower, embeddings, norms, RoPE, either VAE, or the model-specific FP32 patch/timestep/output projections. For component selection, --diffusion-quantization-config '{"transformer":{"method":"fp8"}}' is DiT-only and --diffusion-quantization-config '{"text_encoder":{"method":"fp8"}}' is text-decoder-only; both entries can be combined. Commit 2031bbe adds the missing text-encoder component resolver, documents the matrix, and tests global, DiT-only, and encoder-only routing. The full related regression set passes: 156 tests.

Comment thread docs/user_guide/quantization/fp8.md Outdated
@lishunyang12

Copy link
Copy Markdown
Collaborator Author

Architecture conclusion after auditing the existing quantization path:

  • MiniMax-H3 should not own a model-specific FP8 implementation. Its text-decoder TP linears now inherit LinearBase and delegate quant-method selection, weight creation, online processing, and offload handling to the existing vLLM quantization factory and shared loader.
  • A plain/global quantization config has generic semantics: it is passed unchanged to every quantization-aware component constructed by a pipeline. An explicit ComponentQuantizationConfig is the only mechanism that narrows the scope, for example to transformer or text_encoder.
  • This does not mean arbitrary encoder modules are rewritten automatically. Encoders implemented with ordinary torch.nn.Linear remain at checkpoint precision until their model integration exposes quantizable layers and routes their weights through the shared component loader. Therefore, many existing diffusion pipelines are still DiT-only today; Flux2 and MiniMax-H3 are examples that route eligible encoder components through the common contract.
  • For MiniMax-H3 specifically, global online FP8 covers the DiT and Qwen3-VL text decoder. The vision tower, embeddings, norms, RoPE, VAEs, and precision-sensitive FP32 projections remain at checkpoint precision.

The refactor is in 5a61bbf4. Relevant CPU regressions: 203 passed. DCO, Python 3.11/3.12 wheel builds, pre-commit, and local repository hooks pass; NPU CI is pending.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

resolve conflicts

Make quantization="fp8" apply online FP8 to eligible MiniMax-H3 DiT and Qwen3-VL text-decoder linears by default. Preserve mixed-precision input/output components, stream completed online-quantized layers to CPU, and retain FP8 physical strides across layerwise offload.

Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
@lishunyang12
lishunyang12 force-pushed the feat/minimax-h3-encoder-fp8-dlo branch from 495f8de to e6b3676 Compare August 13, 2026 09:11
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Comment thread vllm_omni/diffusion/offloader/distributed_layerwise_backend.py
Signed-off-by: lishunyang12 <lishunyang12@163.com>
@lishunyang12

Copy link
Copy Markdown
Collaborator Author

Addressed both latest review threads in 55fab2b: H3 now strips unsafe pre-quantized ModelOpt configs from its BF16 text encoder while preserving online FP8, and resident DLO tensors are reconstructed with their saved stride. Added quantization and transposed-weight regressions. Validation: 51 passed; ruff and diff checks passed.

@david6666666 @hsliuustc0106 could you please double-check the update?

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Addressed both latest review threads in 55fab2b: H3 now strips unsafe pre-quantized ModelOpt configs from its BF16 text encoder while preserving online FP8, and resident DLO tensors are reconstructed with their saved stride. Added quantization and transposed-weight regressions. Validation: 51 passed; ruff and diff checks passed.

@david6666666 @hsliuustc0106 could you please double-check the update?

for text encoder, we usually keep higher precision

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. and removed ready label to trigger buildkite CI labels Aug 14, 2026

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@hsliuustc0106
hsliuustc0106 merged commit d1e230c into vllm-project:main Aug 15, 2026
6 checks passed
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

online/runtime FP8 + DLO AllGather is explicitly unsupported.

  • The loader raises ValueError for that combination ([source](
    _skip_load = _dist_offload and _use_ag and _supports_mmap and not _has_online_quant
    if _skip_load:
    logger.info("DLO+AllGather active: skipping load_weights (will load via mmap in enable())")
    else:
    if _dist_offload and _use_ag and _has_online_quant:
    raise ValueError(
    "Online quantization is incompatible with DLO+AllGather: "
    "the sharding + AllGather mechanism flattens weights by "
    "dtype, which breaks quantized weight/scale layouts. "
    "Please use --dlo-no-use-allgather or disable online "
    "quantization."
    )).
  • Supported configuration:
--quantization fp8 \
--enable-distributed-layerwise-offload \
--dlo-no-use-allgather

The limitation applies specifically to runtime-created/online FP8 weights. A compatible pre-quantized checkpoint may use the AllGather path. The stride-preservation changes in the PR do not remove the explicit online-FP8 rejection.

jwxnlp added a commit to RadiantFrame/vllm-omni that referenced this pull request Aug 19, 2026
…vice results

- deploy_fp8.sh: FP8 single-card config (cpu-offload + FP8 + Cache-DiT +
  FA3 + expandable_segments). Upstream vllm-project#5910 (stream-offload online quant)
  fixes the load-time GPU accumulation that previously OOM'd 80G cards:
  measured 50.5s steady @480p/5s and 36.3 GiB resident (BF16: 52.8s /
  58.9 GiB). FP8 is now the single-card recommendation.
- deploy_fp8_8svc.sh: 8-way variant (1 GPU each, pinned master ports).
  Measured 55.6s/req steady (+10.1% contention vs single-lane, down from
  BF16's +16.5%) and 8.44 videos/min aggregate (last-round-max metric,
  +11% over the BF16 8svc) -- the 8x1-GPU layout record.
- README: FP8 revival post-mortem (sec 3), FP8-vs-BF16 tables, new sec 5a
  with the 8-way measurements; BF16 8svc demoted to archived comparison.
  (Also carries two generate-script reference updates from the nsvc
  consolidation.)

Co-Authored-By: Claude <noreply@anthropic.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…-project#5910)

Signed-off-by: lishunyang12 <lishunyang12@163.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda-test Used to trigger vllm-omni cuda CI separately. diffusion codes related to diffusion models quantization Code related to quantization ready label to trigger buildkite CI

3 participants