Replies: 2 comments
|
This is an excellent debugging write-up. The pattern you're seeing—TP within a node works, TP across nodes hangs in NCCL LL protocol wait loops—is a classic symptom of rank synchronization drift in multi-node tensor parallelism, not a fabric-level failure. Root Cause AnalysisThe key evidence:
This pattern typically occurs when:
Diagnostic Approach1. Enable NCCL Collective Tracing with OpCountexport NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=COLL,GRAPH
export NCCL_DEBUG_FILE=/tmp/nccl_rank_%r.logThis will log each collective's 2. Check for CUDA Graph Bucket BoundariesSGLang uses CUDA graphs to batch collective calls. The hang correlation with generation length suggests you're hitting a bucket boundary where:
Test: Disable CUDA graphs for collectives: export SGLANG_DISABLE_CUDA_GRAPH=1If hangs stop, the issue is CUDA graph replay desynchronization. 3. Monitor Per-Rank Scheduler StateAdd logging to SGLang's scheduler to track which batch each rank is processing: # In sglang/srt/managers/scheduler.py, before collective calls:
logger.info(f"Rank {self.tp_rank} processing batch_size={len(batch)}, step={self.step}")If ranks show different batch sizes or step counts before a hang, the scheduler is allowing state divergence. 4. Check MoE Routing (if applicable)For DeepSeek V4 Flash/Pro with MoE layers: export NCCL_ALGO=NVLS # Try NVLink SHARP if available
export CUDA_DEVICE_MAX_CONNECTIONS=8 # Increase connection poolMoE routing can cause different ranks to enter different allgather patterns if expert assignments diverge. Likely Fix: Enforce Collective SynchronizationThe root cause is probably that SGLang's scheduler allows ranks to progress through decode steps independently, and over many iterations, small timing differences accumulate until one rank enters a collective while another is still processing the previous step. Workaround 1: Reduce watchdog timeout # In your launch config
--watchdog-timeout 60 # Force earlier detectionThis won't fix the drift but will fail faster, preventing indefinite hangs. Workaround 2: Force synchronous scheduler steps # Before each collective in the model forward pass:
torch.distributed.barrier(group=self.tp_group)Workaround 3: Use PP instead of TP across nodes python -m sglang.launch_server \
--tp 4 --pp 4 \
--dp 1This keeps tensor parallelism within nodes (where it's fast and reliable) and uses pipeline parallelism across nodes (which is more tolerant of latency). Fabric-Level DiagnosticsIf the above doesn't resolve it, the next step is to capture libfabric traffic: export FI_LOG_PROV=cxi
export FI_LOG_LEVEL=debug
export FI_CXI_DISABLE_HOST_REGISTER=1Then use cxi_dump -t -n 0 # Dump all CXI traffic on node 0Look for:
Questions for Further Debugging
SummaryThis is almost certainly scheduler-level rank desynchronization amplified by CUDA graph replay, not a fabric bug. The LL protocol wait loops are behaving correctly—they're just waiting for data that will never arrive because the peer rank is on a different collective sequence. Start with
Your debugging is spot-on—this is exactly the right approach for narrowing down distributed systems issues. |
Read this before acting on the existing replyThe reply above contains two flags that do not exist in the SGLang codebase. I checked the repo:
The canonical, documented way to disable CUDA graphs is
The suggestion to patch in The direction of that reply (rank divergence / desync, not a fabric failure) is plausible and matches the general shape of these bugs. But its concrete prescriptions mostly won't run, so it's worth not spending a repro cycle on them. What your evidence actually supportsYour coredump evidence is strong and specific. The disassembly you posted is NCCL's LL-protocol wait loop — poll the peer's flag, spin, periodically re-check Your TP-vs-PP observation is the most valuable data point you have:
TP collectives span all nodes; PP only passes point-to-point between stages. So the failure requires a collective whose participant set spans the slow link, which points at timing/participation divergence rather than link quality. The first thing to test: compare opCount across ranksYou asked exactly the right question — how to see which collective each rank was on. Take the last export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=COLL
export NCCL_DEBUG_FILE=/tmp/nccl_rank_%r.logIf the last The repo has an official methodology for exactly this in Keep the process alive so you can look at it
Note that --watchdog-timeout 3600 --soft-watchdog-timeout 300Soft then dumps py-spy traces at 300s without killing anything, and the hard kill is pushed out to an hour, leaving you time to attach and diff ranks. (Setting them the other way round, e.g. soft 600 > hard 300, gets you nothing — the hard watchdog still kills first.) Both arguments are real and defined in This is far more useful than only lowering If you have hierarchical cache enabled, check this specific sync pointOnly relevant if your launch includes def _process_hicache_events(self) -> None:
# The HiCache drain is TP-wide consensus; run it before rank-local
# decisions (_should_defer_prefill) or ranks enter different collectives.
if (self.enable_hierarchical_cache or ... or self.enable_unified_cache_external_linker):
self.tree_cache.check_hicache_events()If you're running HiCache, this is the area I'd look at first, and there's directly relevant precedent. Issue #26921 was: "NCCL deadlock in Worth noting the same family is still open: issue #30760 — "HiCache prefetch all_reduce deadlock with TP=4, no PP — mismatched call count in check_prefetch_progress" is open as of today. So "a rank misses a HiCache collective rendezvous" is a live, recurring class in this code. In current
Because these reduce with MIN, every rank must arrive before any rank proceeds; if one rank is slow to reach the rendezvous, every other rank waits in exactly the LL wait loop your coredump shows. In fact This would also explain your "correlates with longer generations" observation: more decode steps means more rendezvous points, so more chances to catch a rank that's fallen behind. To confirm or eliminate this cheaply: run the same long-generation workload with HiCache disabled. If the hang disappears, the sync points above are where to look; if it still hangs, this is a red herring and the divergence is elsewhere (then the SKILL.md binary-search is the right path). On the topology workaroundSwitching to What would help others reproduce
You offered the coredumps, NCCL logs and repro script — that repro script plus the opCount diff is likely enough for a maintainer to pin this down. Verified against |
Uh oh!
There was an error while loading. Please reload this page.
When running SGLang with tensor parallelism spanning multiple nodes (confirmed on both TP16/4-node and TP8/2-node), long-running generations eventually trigger the scheduler watchdog timeout and hang. CUDA coredump analysis shows the GPU is parked in NCCL's normal LL-protocol wait loop (polling for a peer's data that never arrives), not a kernel fault. The same setup works reliably when TP is confined to a single node (e.g. TP4+PP4 on 16 GPUs / 4 nodes never hangs).
This looks like it could be a fabric-layer issue, an SGLang/NCCL rank-synchronization issue that only manifests over inter-node collectives, or a combination - filing here for guidance on how to narrow it down further and whether this pattern is familiar to anyone.
Environment
libfabricCXI provider)Symptom
SGLang's scheduler watchdog fires after the configured hard timeout:
This sends SIGQUIT to the process, which (with CUDA coredump-on-signal enabled) produces a GPU coredump.
Correlation with generation length: hangs appear to correlate with longer generations (more decode steps). Short queries/responses do not seem to trigger it; the failure rate seems to increase with the number of decode iterations, consistent with a rare per-collective-call event rather than a deterministic trigger - though a structural cause (e.g. CUDA graph bucket crossing, MoE routing imbalance) has not been ruled out.
Coredump analysis
Two independent hangs analyzed via
cuda-gdb target cudacore:Hang 1 -
ncclDevKernel_AllReduce_Sum_bf16_RING_LLActive(not an exception/fault state)Hang 2 -
ncclDevKernel_AllGather_RING_LLDisassembly at both PCs (shared pattern across both hangs)
This is NCCL's standard LL-protocol wait primitive: poll the receive-buffer flag for a peer's data, with a spin-count-gated periodic check of
comm->abortFlag. Seeing the identical wait-loop shape in both AllReduce and AllGather suggests this isn't a bug in either collective's kernel code specifically - both are waiting correctly for data that never arrived within the watchdog window. Same issue is present if I setNCCL_ALGO=Treeso this is not RING specific issue.What's been ruled out / narrowed down
Active, and the disassembly shows a legitimate (if endless) wait loop, not a trap.NCCL_NET_GDR_LEVELmakes no difference; hang reproduces identically with GDR on or off. This argues against a GDR-specific data path issue (memory registration, BAR1 mapping, ODP) and toward something common to both paths.FI_LOG_PROV=cxi FI_LOG_LEVEL=warncame back clean during the actual hang window No libfabric-detected error during the hang itself so far.What we're asking for
RING_LL)? Any known issues withaws-ofi-ncclor CXI provider on this generation of hardware/interconnect?NCCL_DEBUG_SUBSYS=COLLatINFOlevel?Happy to share both coredumps, full NCCL INIT/NET/COLL logs, and our
standalone repro script if useful.
All reactions