[Kimi-K3] Add GEMM-RS for sequence parallelism - #52079
Merged
Merged
Conversation
gau-nernst
requested review from
AndreasKaratzas,
DarkLight1337,
Harry-Chen,
WoosukKwon,
khluu,
mgoin,
tlrmchlsmth,
yewentao256,
ywang96 and
zyongye
as code owners
August 13, 2026 01:15
Contributor
Author
|
/ci run |
|
✅ Triggered Buildkite CI #83638 for commit |
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
gau-nernst
force-pushed
the
codex/kimi-k3-gemm-rs
branch
from
August 13, 2026 01:52
4210d0f to
50bc8b7
Compare
Contributor
Author
|
/ci run |
|
✅ Triggered Buildkite CI #83645 for commit |
Contributor
Author
|
/ci run |
|
✅ Triggered Buildkite CI #83667 for commit |
Contributor
Author
|
/ci run |
|
✅ Triggered Buildkite CI #83689 for commit |
vrdn-23
added a commit
to vrdn-23/vllm
that referenced
this pull request
Aug 14, 2026
Resolves the recurring vllm/envs.py structural conflict per docs/superpowers/specs/2026-05-14-envs-merge-conflict-resolution-design.md: main's legacy `if TYPE_CHECKING:` block and `environment_variables` dict are dropped wholesale (superseded by the pydantic BaseSettings tree on this branch), then main's semantic delta is ported field-by-field. 6 main-side commits touched vllm/envs.py since base e644c8c (+55 -0). All 10 new vars have already-merged callers, so every port is mandatory: - vllm-project#51447 VLLM_MAX_STOP_STRINGS (int=4), VLLM_MAX_NUM_BAD_WORDS (int=128), VLLM_MAX_BAD_WORDS_TOTAL_TOKENS (int=1024) -> ServerSettings - vllm-project#49948 VLLM_MAX_AUDIO_DECODE_BYTES (int=268_435_456) -> MediaSettings, carrying compile_factor=False to mirror main's ignore-set addition - vllm-project#50484 VLLM_USE_DIRECT_DCP_A2A / _Q_GATHER / _KV_GATHER (bool|None=None) -> QuantSettings, with one shared `_parse_direct_dcp` before-validator reproducing main's maybe_convert_bool exactly - vllm-project#52079 VLLM_KIMI_K3_GEMM_RS (bool=False) -> QuantSettings - vllm-project#49458 VLLM_USE_HW_AGNOSTIC (bool=False) -> UsageSettings - vllm-project#47808 VLLM_ADAPTIVE_VERIFICATION_PROFILE_CONTEXT_LEN (int=8192) -> QuantSettings No deletions, modifications, or renames this window. Nothing was dropped silently: all 6 commits' envs.py deltas are covered above. Env var set parity after resolution: 292 branch fields vs 293 main runtime entries, sole difference VLLM_TRITON_ATTN_USE_TD -- the known deprecation shim divergence, re-confirmed untouched by this merge window. Verified: 54 tests pass across tests/test_envs.py, tests/test_envs_pydantic.py and tests/docs/test_env_vars_gen.py; `pre-commit run --files vllm/envs.py` clean; tests/test_request_input_bounds.py passes (22 tests). The audio, DCP and end-to-end hw-agnostic consumer suites need a CUDA box plus soundfile / multiprocess and were not run here. AI assistance was used to enumerate the port list and apply the resolution; see Appendix G of the playbook for the full audit trail. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Vinay Damodaran <vrdn@hey.com>
4 tasks
zufangzhu
pushed a commit
to zufangzhu/vllm
that referenced
this pull request
Aug 24, 2026
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg> Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
4 tasks
XiaoSongXS
reviewed
Sep 3, 2026
|
|
||
| ptr = x.iterator.toint(loc=loc, ip=ip).ir_value(loc=loc, ip=ip) | ||
| asm = ( | ||
| "multimem.ld_reduce.relaxed.gpu.global.add.acc::f32" |
There was a problem hiding this comment.
functionally speaking, .gpu scope also work on hw. PTX semantic speaking, .sys scope is better.
1 task done
33 of 46 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Add GEMM-RS kernel for Blackwell, based on https://github.com/NVIDIA/cutlass/blob/dcf215a/examples/python/CuTeDSL/cute/blackwell/kernel/distributed/distributed_gemm_reduce_scatter_blackwell.py (
multimem.ld_reduce)ceil(M / world_size), with the exception of the last rankVLLM_KIMI_K3_GEMM_RS, which enables GEMM-RS for O-proj and shared experts+dense MLP down-projmax_num_batched_tokens x 7168 x 2 bytes= 448 MiB for MNBT=32k. Confirmed in vLLM logs KV memory 39.99 GiB (before) -> 39.78 GiB (after) -> not muchInitialization and runtime logic
maybe_init_gemm_rs(), which also logs the reason if it fails. WhenVLLM_KIMI_K3_GEMM_RS=0, it doesn't do anything__init__(), we callself.run_gemm_rs = get_gemm_rs().can_run(self.down_proj.weight). This is to further validate supported weight shapes and dtypeforward(), we check again withshould_run(), which is the heuristics M>=128. The kernel supports any values of M, but right now the baseline is better for M<128Though technically this can work with any SP in general, this PR only enables GEMM-RS for Kimi-K3. A future extension is to make this into GEMM-AR by adding
multimem.st(all-gather) aftermultimem.ld_reduce(reduce-scatter).Microbenchmark
benchmarks/kernels/benchmark_kimi_k3_gemm_rs.pyin this PR. CUDA graph with rotating buffers. All benchmarks were done with GB300.TP4
Note: K=1536 is shared expert down-proj, K=3072 is O-proj
Component breakdown
TP8
Note: K=768 is shared expert down-proj, K=1536 is O-proj
Component breakdown
E2E prefill-only benchmark
All benchmarks were done with 8xGB300, TP8+EP+SP (DeepGEMM MegaMoE),
--max-num-batched-tokens 32768, 8k input - 1 output requests. Baseline is 7aa248fTest Plan
Unit test (also added to distributed CI)
E2E testing, TP8+EP+SP (DeepGEMM MegaMoE)
Test Result
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)