Skip to content

[K3 Perf] Fuse MXFP4 top-k finalization into latent-tail, ~5% E2E latency reduction - #53152

Merged
yewentao256 merged 4 commits into
mainfrom
wentao-optimize-moe-do-finalize
Aug 21, 2026
Merged

yewentao256 merged 4 commits into
mainfrom
wentao-optimize-moe-do-finalize

Conversation

@yewentao256

@yewentao256 yewentao256 commented Aug 20, 2026 •

Copy link
Copy Markdown
Member

Purpose

Part of #50587

Fuse MXFP4 top-k finalization into the Kimi K3 latent-tail kernel.

Before

MXFP4 MoE kernel
    │
    ├─ GEMM2
    │
    └─finalize kernel
         ├─ unpermute
         ├─ times router weight
         ├─ top-k reduction
         └─ Write [M, 3584] tensor
                    │
                    ▼
latent tail reads the tensor
    └─ AllReduce + RMSNorm + Up Projection + shared expert

Now

MXFP4 MoE kernel(do_finalize=False)
    │
    └─ return:
         ├─ GEMM2 output
         ├─ router weights
         └─ permutation map
                    │
latent tail reads the tensor
    ├─ times router weight and topk
    ├─ AllReduce + RMSNorm + Up Projection + shared expert

This removes one kernel launch and avoids writing and rereading the finalized intermediate tensor.

Test

vllm serve moonshotai/Kimi-K3 \
  --trust-remote-code \
  -tp 8 \
  --load-format fastsafetensors \
  --enable-prefix-caching \
  --reasoning-parser kimi_k3 \
  --enable-auto-tool-choice \
  --tool-call-parser kimi_k3 \
  --host 0.0.0.0 \
  --port 30000

HF_HUB_CACHE=/data/engine/hub_cache HF_HUB_OFFLINE=1 guidellm run   --backend '{"kind":"openai_http","target":"http://127.0.0.1:30000","model":"moonshotai/Kimi-K3","request_format":"/v1/chat/completions","timeout":100000,"extras":{"body":{"reasoning_effort":"max","temperature":1.0,"top_p":0.95}}}'   --tokenizer '{"kind":"huggingface_auto","model":"moonshotai/Kimi-K3","load_kwargs":{"trust_remote_code":true,"local_files_only":true}}'   --data 'kind=synthetic_text,prompt_tokens=8000,output_tokens=1000'   --profile 'kind=concurrent,warmup=0.1,cooldown=0.1'   --override profile.streams '1,4,16'   --constraint 'kind=max_duration,seconds=150'   --constraint 'kind=max_errors,count=10'   --seed 'kind=static,value=42'   --metrics 'kind=generative,sample_size=20'   --output 'kind=json,path=/home/yewentao256/guidellm-results/output_vllm_kimik3_8k1k.json'

lm_eval --model local-completions --model_args "base_url=http://127.0.0.1:30000/v1/completions,model=$MODEL,num_concurrent=1024" --tasks gsm8k --trust_remote_code

Acc

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value |   |Stderr|
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  |0.9659|±  | 0.005|
|     |       |strict-match    |     5|exact_match|↑  |0.9659|±  | 0.005|

Perf

Concurrency Latency P50 Main Latency P50 Throughput Main Throughput
1 9.67 s (−5.6%) 🚀 10.24 s 95.9 tok/s (+5.1%) 🚀 91.2 tok/s
4 13.76 s (−4.7%) 🚀 14.45 s 290.7 tok/s (+5.0%) 🚀 276.8 tok/s
16 22.94 s (−4.4%) 🚀 23.99 s 690.3 tok/s (+4.7%) 🚀 659.3 tok/s

TTFT remains unchanged

Concurrency TTFT P50 Main TTFT P50
1 434.4 ms (−0.2%) 435.3 ms
4 1449.7 ms (−0.1%) 1450.9 ms
16 2644.8 ms (−0.1%) 2648.5 ms

output_vllm_kimik3_8k1k_0817_no_spec.json
output_vllm_kimik3_8k1k.json

Signed-off-by: yewentao256 <zhyanwentao@126.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@yewentao256 yewentao256 added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 20, 2026
@yewentao256

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #84873 for commit 73064103eed8.

@yewentao256 yewentao256 changed the title [K3 Perf] Fuse MXFP4 top-k finalization into latent-tail, 4.4%~5.6% E2E latency reduction Aug 20, 2026
@mgoin mgoin changed the title [K3 Perf] Fuse MXFP4 top-k finalization into latent-tail, ~5% E2E latency and throughput improvement Aug 20, 2026
@yewentao256 yewentao256 changed the title [K3 Perf] Fuse MXFP4 top-k finalization into latent-tail, ~5% E2E latency Aug 20, 2026

@mgoin mgoin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The optimization looks nice, but could we clean up the output contract before merging?

trtllm_fp4_block_scale_moe now uses both output (the finalized destination) and result (a mode dependent list), which makes the data flow difficult to follow. Please rename these clearly, immediately decompose the deferred outputs, and ideally centralize the FlashInfer return conversion in a shared helper.

Also, UnfinalizedMoEOutput now passes through APIs still typed as tensor-only. Those annotations/guards should be updated, with targeted finalized-vs-deferred parity tests added. The current structure feels too fragile to future changes.

@github-project-automation github-project-automation Bot moved this to In review in NVIDIA Aug 20, 2026
@zyongye
zyongye enabled auto-merge (squash) August 21, 2026 03:13

@zyongye zyongye left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will merge this first and doing cleaning up and some perf tuning in a separate PR. Thanks for the effort.

@gold9450412

Copy link
Copy Markdown

I'm so excited!
So many optimizations for K3!
I've been looking forward to this for so long, and I really hope it can be merged in!
I really want to try it out!

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
@yewentao256

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85008 for commit ddac6d2053df.

@yewentao256

Copy link
Copy Markdown
Member Author

Thanks @zyongye ! Also CC @mgoin , I am thinking if we could land this first, and I will address your comments in a following up PR, what do you think?

@yewentao256

Copy link
Copy Markdown
Member Author

Note: current ci failure buildkite/ci/pr/nvidia-h200-quantized-models is not related, happening in other PRs as well

@github-project-automation github-project-automation Bot moved this from In review to Ready in NVIDIA Aug 21, 2026
@yewentao256

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85035 for commit 930b7903b65f.

@yewentao256
yewentao256 merged commit 7a2fdba into main Aug 21, 2026
127 of 128 checks passed
@yewentao256
yewentao256 deleted the wentao-optimize-moe-do-finalize branch August 21, 2026 17:30
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Aug 21, 2026
@zyongye

zyongye commented Aug 21, 2026

Copy link
Copy Markdown
Member

Accuracy validation of this change on GB300:

config deferred finalize GSM8K OCRBench
TP8, conc 64 on 0.9598 887
TP8, conc 8 on 0.9621 —
TEP8 (tp8+ep) off (ep_size=8) 0.9682 895
DEP16 (tp1×dp16×ep16) off (tp_size=1) 0.9644 879
ref (2P1D+EAGLE3, 08-13) — 0.9697 888

Kimi-K3 MXFP4, GB300 4×GPU/node, ddac6d2053, flashinfer 0.6.17.
GSM8K: 5-shot, /v1/completions, temp 0.6, top_p 0.95, n=1319, strict-match.
OCRBench: temp 1.0, thinking effort high, n=1000, raw score.
0 request errors across 5,957 requests.

AI assistance was used to produce these runs.

zyongye pushed a commit that referenced this pull request Aug 24, 2026
Extend the deferred MoE finalize protocol (#53152) to the modular
prepare/finalize path so the TRTLLM FP8 block-scale and NVFP4 routed
kernels can stop after GEMM2, and add a FlashInfer fused
finalize + all-reduce + RMSNorm consumer for MiniMax M3.

Signed-off-by: Yongye Zhu <yongye@inferact.ai>
zyongye pushed a commit that referenced this pull request Aug 24, 2026
Two places went their own way instead of using what #53152 established:

- The TRTLLM FP8 block-scale and NVFP4 routed kernels hand-built an
  UnfinalizedMoEOutput from the raw FlashInfer return. Route both through
  convert_flashinfer_moe_output, as every other TRTLLM expert does. It
  validates the deferred layout, and it absorbs FlashInfer's coming switch
  of the finalized return from a Tensor to List[Tensor].

- The consumer's workspace capacity travelled on a new MoEOutput field.
  The protocol already has a channel for it -- the consumer declares
  FusedMoEConfig.defer_moe_finalize_max_num_tokens at build time and the
  producer gates on should_defer_moe_finalize -- so use that and leave
  MoEOutput alone. This also moves the workspace build out of the
  forward-time custom op, where it ran a collective without the vLLM
  config, into the consuming layer's constructor: an unsupported
  (tp_size, hidden, top_k, dtype) is now a build-time fallback to
  finalizing in the MoE kernel rather than a forward-time assert.

Signed-off-by: Yongye Zhu <yongye@inferact.ai>
yewentao256 added a commit that referenced this pull request Aug 24, 2026
…53152 (#53310)

Signed-off-by: yewentao256 <zhyanwentao@126.com>
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
…ency reduction (vllm-project#53152)

Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
mikeshawcode pushed a commit to mikeshawcode/vllm that referenced this pull request Sep 1, 2026
…ency reduction (vllm-project#53152)

Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
mikeshawcode pushed a commit to mikeshawcode/vllm that referenced this pull request Sep 1, 2026
…llm-project#53152 (vllm-project#53310)

Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
evezhier pushed a commit to evezhier/vllm that referenced this pull request Sep 19, 2026
Extend the deferred MoE finalize protocol (vllm-project#53152) to the modular
prepare/finalize path so the TRTLLM FP8 block-scale and NVFP4 routed
kernels can stop after GEMM2, and add a FlashInfer fused
finalize + all-reduce + RMSNorm consumer for MiniMax M3.

Signed-off-by: Yongye Zhu <yongye@inferact.ai>
evezhier pushed a commit to evezhier/vllm that referenced this pull request Sep 19, 2026
Two places went their own way instead of using what vllm-project#53152 established:

- The TRTLLM FP8 block-scale and NVFP4 routed kernels hand-built an
  UnfinalizedMoEOutput from the raw FlashInfer return. Route both through
  convert_flashinfer_moe_output, as every other TRTLLM expert does. It
  validates the deferred layout, and it absorbs FlashInfer's coming switch
  of the finalized return from a Tensor to List[Tensor].

- The consumer's workspace capacity travelled on a new MoEOutput field.
  The protocol already has a channel for it -- the consumer declares
  FusedMoEConfig.defer_moe_finalize_max_num_tokens at build time and the
  producer gates on should_defer_moe_finalize -- so use that and leave
  MoEOutput alone. This also moves the workspace build out of the
  forward-time custom op, where it ran a collective without the vLLM
  config, into the consuming layer's constructor: an unsupported
  (tp_size, hidden, top_k, dtype) is now a build-time fallback to
  finalizing in the MoE kernel rather than a forward-time assert.

Signed-off-by: Yongye Zhu <yongye@inferact.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

k3 kimi nvidia performance Performance-related issues quantization ready ONLY add when PR is ready to merge/full CI is needed

4 participants