[Diffusion][Quantization] Enable MiniMax-H3 global FP8 with DLO - #5910
hsliuustc0106 merged 12 commits into
Conversation
9f92147 to
b18eeff
Compare
Single-B300 encoder FP8 isolationThis isolates the incremental effect of quantizing the Qwen3-VL text decoder after the DiT is already FP8. Workload
Incremental encoder resultComparing default global FP8 with DiT-only FP8:
The small end-to-end speed change is expected: encoding runs once, while the DiT runs for 50 denoising steps. During denoising, the 200 ms NVML trace showed stable device-memory plateaus of 101,030 MiB for DiT-only FP8 and 79,188 MiB for default global FP8, a difference of 21,842 MiB (21.6%). This plateau is reported as an observation, not as a second-request worker peak: the first global-FP8 request also performs online encoder quantization and its temporary workspace/cache raises the formal request peak. Default FP8 + DLO compatibilityA separate current-head single-B300 smoke used global FP8 with no-AllGather DLO, 1344x768, 5 seconds, and 2 denoising steps:
Conclusion: encoder FP8 provides meaningful single-GPU capacity/headroom and encode-stage speedup, but should not be presented as a large end-to-end throughput optimization. |
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
Merged current upstream main in 2ac53dd and resolved the MiniMax-H3 recipe conflict. The merged limitation now reflects this PR: online FP8 with DLO uses rank-local streaming and does not support the sharded AllGather path. Post-merge validation passes: loader/H3 contracts 98 tests, DLO suites 57 tests, and pre-commit on all textual PR files. The PR is ready for review; platform CI has restarted. |
|
do we apply fp8 to all components? how do we select? |
|
The scope is now explicit and selectable. |
|
Architecture conclusion after auditing the existing quantization path:
The refactor is in |
|
resolve conflicts |
cd422ab to
495f8de
Compare
Make quantization="fp8" apply online FP8 to eligible MiniMax-H3 DiT and Qwen3-VL text-decoder linears by default. Preserve mixed-precision input/output components, stream completed online-quantized layers to CPU, and retain FP8 physical strides across layerwise offload. Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
495f8de to
e6b3676
Compare
Signed-off-by: lishunyang12 <lishunyang12@163.com>
Signed-off-by: lishunyang12 <lishunyang12@163.com>
|
Addressed both latest review threads in 55fab2b: H3 now strips unsafe pre-quantized ModelOpt configs from its BF16 text encoder while preserving online FP8, and resident DLO tensors are reconstructed with their saved stride. Added quantization and transposed-weight regressions. Validation: 51 passed; ruff and diff checks passed. @david6666666 @hsliuustc0106 could you please double-check the update? |
for text encoder, we usually keep higher precision |
|
online/runtime FP8 + DLO AllGather is explicitly unsupported.
--quantization fp8 \
--enable-distributed-layerwise-offload \
--dlo-no-use-allgatherThe limitation applies specifically to runtime-created/online FP8 weights. A compatible pre-quantized checkpoint may use the AllGather path. The stride-preservation changes in the PR do not remove the explicit online-FP8 rejection. |
…vice results - deploy_fp8.sh: FP8 single-card config (cpu-offload + FP8 + Cache-DiT + FA3 + expandable_segments). Upstream vllm-project#5910 (stream-offload online quant) fixes the load-time GPU accumulation that previously OOM'd 80G cards: measured 50.5s steady @480p/5s and 36.3 GiB resident (BF16: 52.8s / 58.9 GiB). FP8 is now the single-card recommendation. - deploy_fp8_8svc.sh: 8-way variant (1 GPU each, pinned master ports). Measured 55.6s/req steady (+10.1% contention vs single-lane, down from BF16's +16.5%) and 8.44 videos/min aggregate (last-round-max metric, +11% over the BF16 8svc) -- the 8x1-GPU layout record. - README: FP8 revival post-mortem (sec 3), FP8-vs-BF16 tables, new sec 5a with the 8-way measurements; BF16 8svc demoted to archived comparison. (Also carries two generate-script reference updates from the nsvc consolidation.) Co-Authored-By: Claude <noreply@anthropic.com>
…-project#5910) Signed-off-by: lishunyang12 <lishunyang12@163.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
What users get
This PR adds one user-facing MiniMax-H3 online FP8 mode:
--quantization fp8(orquantization="fp8") quantizes all eligible DiT and Qwen3-VL text-decoder linears.This intentionally exposes one simple default. It does not add a MiniMax-H3 layer-selection API or additional user-facing quantization modes.
Quick start
Single GPU, resident FP8, no offload
Use the FL2VA-only partition for the 96 GB single-GPU capacity test. This avoids loading the second Ref2VA DiT.
Terminal 1:
MODEL_ROOT=/path/to/MiniMax-H3 \ GPU_ID=0 \ PORT=8091 \ RUN_DIR=$PWD/h3_single_gpu_fp8_no_offload \ bash examples/online_serving/minimax_h3/run_server_single_gpu_fp8_no_offload.shTerminal 2, after the health check succeeds:
BASE_URL=http://127.0.0.1:8091 \ RUN_DIR=$PWD/h3_single_gpu_fp8_no_offload \ bash examples/online_serving/minimax_h3/run_curl_t2va_5s.shThe server script intentionally passes none of
--enable-cpu-offload,--enable-layerwise-offload, or--enable-distributed-layerwise-offload. It samples whole-device memory from before model initialization. The request script generates a fixed five-second, 1344x768, 50-step T2VA output and verifies H.264 video plus 32 kHz stereo AAC.After the request succeeds, stop the server with Ctrl-C. The run directory contains:
gpu.csvandmemory_summary.txtfor whole-lifecycle capacity and headroom;server.logfor worker peak and encode/diffuse/decode timing;client_time.txt,t2va.mp4,media.json, andvalidation.txt.The complete guide is in
recipes/MiniMaxAI/MiniMax-H3.md.Python API
Distributed layerwise offload
For lower HBM residency, combine the same FP8 default with DLO's full-weight per-rank path:
Runtime-created FP8 weights do not support the sharded DLO AllGather path. Use
--dlo-no-use-allgatherwhen combining online FP8 with DLO.Precision scope
quantization="fp8"means global FP8 for eligible compute-heavy linears, not every tensor in the pipeline.This keeps model-specific mixed-precision boundaries while providing the largest practical memory reduction through one default option.
Direct BF16 vs global FP8 videos
Matched five-second T2VA, I2VA, and Ref2VA pairs are available in the side-by-side player. The rankings comparison folder also provides individual MP4 links, per-task fidelity metrics, and provenance.
User-visible results
Single-B300 resident, eager, five-second T2VA at 1344x768 and 50 steps:
The global-FP8 run's sampled whole-device first-request peak was 92,946 MiB. This is a capacity proxy for a 96 GB RTX PRO 6000, not a target-hardware validation; the scripts above record the real card's reported capacity and headroom.
The repeatable encoder-side benefit is primarily capacity, not latency. Isolated steady-state tests found no stable Qwen encoder speedup, while complete-pipeline resident memory fell from 121.10 GiB to 90.33 GiB at text-encoder TP1.
With DLO no-AllGather, a current-head single-B300 smoke generated valid five-second H.264/AAC output with an 18,062 MiB worker request peak and a 50,310 MiB sampled full-lifecycle peak. In the four-GPU benchmark, DLO startup streaming reduced the measured full-lifecycle peak from 58,696 to 37,938 MiB/GPU (35.4%).
The complete T2VA/I2VA/Ref2VA ablation results and generated videos are available here:
https://github.com/lishunyang12/vllm-omni-rankings/tree/main/scripts/minimax_h3_online_fp8
Across those three five-second, 50-step cases, default global FP8 averaged 1.066x speedup and 29.2% lower request-peak memory.
Implementation notes
Validation