You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
DeepEP v2 has been released on the epv2-release branch, which uses NCCL GIN for RDMA communication, TMA for data movement, and introduces the new ElasticBuffer.
Are there any plans to support DeepEP v2? PR #24443 looks related but still appears to be a draft.
DeepEPv2Dispatcher is selected through the same factory path as the v1 dispatcher (fused_moe_triton/layer.py, ep_moe/layer.py), and there's DeepGEMM runner integration (moe_runner/deep_gemm.py). The ElasticBuffer facade you referenced is there too, with a process-wide instance stored in runtime resources and a comment that "a key change rebuilds ElasticBuffer collectively on every rank." So the answer to "any plans to support DeepEP v2?" is: support has effectively landed — it's no longer gated on PR #24443 being a draft.
What to be aware of before you flip it on
1. There's a quantization requirement, and it's the thing most likely to bite you.
def_validate_deepep_v2_quant_method(quant_method) ->None:
"""Validate the expert formats the DeepEP v2 adapter can feed."""ifnotget_moe_a2a_backend().is_deepep_v2():
returnifisinstance(quant_method, UnquantizedFusedMoEMethod):
return
...
raiseValueError(
"--moe-a2a-backend deepep_v2 requires 128x128 blockwise FP8 or 1x32 MXFP8 ""experts with dynamic activation scaling or unquantized BF16 "f"experts, but this layer {reason}. Use a compatible checkpoint or ""--moe-a2a-backend deepep."
)
So the accepted expert formats are:
unquantized BF16 (no constraint), or
FP8 with weight_block_size == [128, 128] and activation_scheme == "dynamic", or
MXFP8 with weight_block_size == [1, 32] and dynamic activation scaling.
Explicitly rejected: FP4 experts, non-FP8 quant methods, and any FP8 whose block size isn't one of those two shapes. The error message even suggests the fallback (--moe-a2a-backend deepep). If you're testing on a DeepSeek-style FP8 checkpoint, check its quantization_config.weight_block_size and activation_scheme first — that single check will tell you whether the backend is even eligible, and it's the fastest way to avoid a confusing failure.
2. Importability is checked with a clear message. The module guards the import:
"DeepEP v2 (ElasticBuffer) is not available. Install DeepEP v2 from ..."
so you need a DeepEP build that exposes ElasticBuffer (i.e. the v2 / epv2-release line you mentioned) — v1 won't satisfy it.
3. Expect the "collective rebuild" semantics to matter in your topology. The runtime-resources comment about rebuilding on a key change is worth keeping in mind: the buffer is process-wide and rebuilt collectively, so a key change is a rank-synchronized event. If you hit hangs during startup or after a config change, GIN/ElasticBuffer initialization is the first place I'd look rather than the MoE kernels.
Suggested order of operations
Confirm your checkpoint's quant format is eligible (see the three accepted shapes above). If not, no amount of flag-tweaking will work.
Install a DeepEP build with ElasticBuffer.
Launch with --moe-a2a-backend deepep_v2 and keep your existing expert-parallel flags otherwise unchanged.
A/B against --moe-a2a-backend deepep on the same checkpoint and the same bench settings — since both backends are now selectable, that's a clean single-variable comparison and the most useful thing you can post back.
If it fails, the error string will tell you which of the three gates you hit (import vs quant format vs layer construction), so it's worth pasting verbatim.
Verified against main on 2026-09-25: arg_groups/fields/exec_.py (deepep_v2 in moe_a2a_backend), layers/moe/token_dispatcher/__init__.py and deepep_v2.py (DeepEPv2Dispatcher, ElasticBuffer), layers/moe/fused_moe_triton/layer.py (_validate_deepep_v2_quant_method, get_deepep_v2_dispatcher_output_dtype), layers/moe/ep_moe/layer.py and moe_runner/deep_gemm.py (backend selection).
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
DeepEP v2 has been released on the
epv2-releasebranch, which uses NCCL GIN for RDMA communication, TMA for data movement, and introduces the new ElasticBuffer.Are there any plans to support DeepEP v2? PR #24443 looks related but still appears to be a draft.
Thanks!
All reactions