feat(validator): AKS H100 NCCL all-reduce performance validation - #1695
Conversation
|
🌿 Preview your docs: https://nvidia-preview-feat-nccl-aks-h100.docs.buildwithfern.com/aicr |
Recipe evidence checkOther affected recipes without evidence yet: 6These recipes are affected by this PR but carry no committed evidence pointer, so there is
This gate is warning-only and never blocks merge. See ADR-007 for the trust model. |
📝 WalkthroughWalkthroughAdds AKS H100 support to NCCL all-reduce bandwidth validation. The change updates documentation and recipe constraints, adds AKS RDMA discovery and runtime template handling, extends validator scheduling and supported combinations, and introduces an AKS H100 TrainingRuntime manifest with accompanying tests. Estimated code review effort: 4 (Complex) | ~45 minutes Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
ad305e5 to
67e5544
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@validators/performance/nccl_aks_utils.go`:
- Around line 45-64: applyAKSTemplateData currently derives RDMA sizing from
config.Nodes[0], which can misconfigure mixed AKS pools. Update this function to
determine RDMA availability across the full target node set, or explicitly
reject mixed RDMA-capable and non-capable nodes before setting
RDMA_RESOURCE_LIMITS, RDMA_RESOURCE_REQUESTS, and MAX_MESSAGE_SIZE. Keep the
existing logging around discoverAKSRdmaCount/buildAKSRdmaResourceLine, but base
the decision on all nodes instead of the first one.
In `@validators/performance/testdata/h100/aks/runtime.yaml`:
- Line 73: The AKS runtime manifest uses a worker/launcher image from
public.ecr.aws, which can fail on locked-down Azure clusters without outbound
access. Update the AKS GPU setup docs to explicitly note the cross-cloud egress
requirement for this image pull, and consider mirroring the image to an
Azure-reachable registry; reference the runtime image entries and the existing
AKS setup guidance when making the change.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: 03965fb3-16fa-40b0-abc4-a2cf32214b19
📒 Files selected for processing (10)
docs/integrator/aks-gpu-setup.mddocs/user/validation.mdexamples/recipes/aks-training.yamlrecipes/overlays/a100-aks-training.yamlrecipes/overlays/h100-aks-training.yamlvalidators/performance/nccl_aks_utils.govalidators/performance/nccl_aks_utils_test.govalidators/performance/nccl_all_reduce_bw_constraint.govalidators/performance/nccl_test.govalidators/performance/testdata/h100/aks/runtime.yaml
xdu31
left a comment
There was a problem hiding this comment.
LGTM — well-motivated, well-tested, live-validated on ND96isr_H100_v5 hardware. Runtime template, RDMA discovery, worker scheduling, and support-table entry all wire up cleanly and mirror the EKS pattern. Provisional floor is honestly labeled with a clear recalibration path.
Minor follow-ups (non-blocking):
- Consider a lightweight test asserting the
rdmaIndentconstant matches the actual column of${RDMA_RESOURCE_LIMITS}intestdata/h100/aks/runtime.yaml— cheap insurance against future reformatting (applies to EKSefaIndenttoo). TestSupportedNCCLCombinationsHaveRuntimeTemplatesonly exercises the populated RDMA line; adding an empty-substitution variant would guard the TCP-fallback rendering path.NCCL_SOCKET_IFNAME=eth0is hardcoded — worth a comment noting the observed interface at the tested SKU in case a future Azure CNI variant changes it.
Wire service=aks accelerator=h100 into the NCCL all-reduce performance validator, mirroring the EKS H100 path: - supportedNCCLCombinations: add AKS+H100 to the default variant. - testdata/h100/aks/runtime.yaml: TrainingRuntime derived from the EKS template with the transport swapped from EFA/Libfabric to NCCL's built-in IB/verbs (NCCL_NET_PLUGIN=none) over the ND-series InfiniBand HCAs; workers request one unit of the rdma-shared-device-plugin pool (rdma/hca_shared_devices_a) plus 8 GPUs and keep IPC_LOCK. - nccl_aks_utils.go: discover the shared RDMA pool from node allocatable and render the worker resource lines; 0 devices falls back to TCP with the reduced message-size cap, mirroring the EKS zero-EFA behavior. - platformWorkerScheduling: AKS shares the OKE shape — pin workers to the common GFD nvidia.com/gpu.product label and tolerate the pool taint. - recipes/overlays/h100-aks-training.yaml: declare the performance phase with a provisional >= 150 GB/s busbw floor (half the calibrated EKS value) pending calibration on an ND96isr_H100_v5 testbed. - docs/user/validation.md + examples/recipes/aks-training.yaml updated. Facts verified live on aicr-test1 (Standard_ND96isr_H100_v5): allocatable nvidia.com/gpu=8 and rdma/hca_shared_devices_a=1k, GFD gpu.product NVIDIA-H100-80GB-HBM3, taint nvidia.com/gpu=present:NoSchedule. Signed-off-by: Yuan Chen <yuanchen97@gmail.com>
67e5544 to
4c13df7
Compare
…DIA#1695) Signed-off-by: Yuan Chen <yuanchen97@gmail.com>
Summary
Adds AKS (Azure) support to the
nccl-all-reduce-bwperformance validator — service=aks + accelerator=h100 now runs the NCCL all-reduce bandwidth benchmark over the ND-series InfiniBand fabric — and declares the performance phase in theh100-aks-trainingoverlay with a provisional busbw floor.Motivation / Context
recipes/overlays/h100-aks-training.yamlcarried an intentional-skip comment: the performance phase was deliberately not declared because the validator had no AKS NCCL runtime template, scheduling path, or support-table entry — a declared check would have silently skipped and reported the phase as passed without measuring anything. This PR wires the full path (runtime template + RDMA discovery + worker scheduling + support-table entry + recipe constraint) and replaces that comment with a real performance block, verified end-to-end on a live AKS H100 cluster.Fixes: N/A
Related: #1337, #1256
Type of Change
Component(s) Affected
cmd/aicr,pkg/cli)cmd/aicrd,pkg/server)pkg/recipe)pkg/bundler,pkg/component/*)pkg/collector,pkg/snapshotter)pkg/validator)pkg/errors,pkg/k8s)docs/,examples/)validators/performance(validator pod binary + testdata),recipes/overlaysImplementation Notes
aks -> [h100]added tosupportedNCCLCombinations[variantDefault]only. EKS H100 exists only in the default variant, so no-net/-nvlsentries (those are GB200-specific).testdata/h100/aks/runtime.yaml): derived from the EKS H100 TrainingRuntime, samepublic.ecr.aws/hpc-cloud/nccl-testsimage (ships sshd, nccl-tests, rdma-core).NCCL_NET_PLUGIN=nonebypasses the bundled aws-ofi-nccl (EFA/Libfabric) plugin so NCCL's built-in IB/verbs transport binds the mlx5 HCAs directly; allFI_*/EFA env dropped;IPC_LOCKretained (ibverbs pinned-buffer registration).NCCL_IB_HCAis deliberately not pinned — NCCL auto-detects HCAs and skips inactive ports.nccl_aks_utils.go): workers request one unit ofrdma/hca_shared_devices_a, the shared pool published by the network-operator rdma-shared-device-plugin (resource name pinned byrecipes/components/network-operator/manifests/nic-cluster-policy-aks.yaml). Shared-pool semantics: one unit mounts every/dev/infinibanddevice, so the request is always"1"regardless of HCA count. Zero allocatable devices falls back to TCP with the reduced 4G message cap, mirroring the EKS zero-EFA behavior.platformWorkerScheduling— pin workers to the common GFDnvidia.com/gpu.productlabel (the same labelnarrowByAcceleratorsized the cohort against) and tolerate the pool taint (nvidia.com/gpu=present:NoSchedule). The AKS-nativekubernetes.azure.com/acceleratorlabel was rejected: its value is justnvidia(verified live), so it cannot pin the H100 cohort on a mixed-accelerator cluster.nccl-all-reduce-bw >= 150— half the calibrated EKS H100 floor (300). Live baseline: 157.24 GB/s busbw measured untuned on 2x Standard_ND96isr_H100_v5 (8x H100 SXM, 8x 400Gb NDR IB per node). 8x NDR line rate is ~400 GB/s, so the floor should be recalibrated upward after IB tuning; the comment in the overlay records this.docker buildfrom a macOS checkout is broken repo-wide — git tracks bothvendor/github.com/NVIDIAandvendor/github.com/nvidia, which collide on case-insensitive filesystems, so the docker build context only ever contains one of them and the in-containergo buildfails on thek8s-launch-kitimports. CI/Linux is unaffected; localgo buildworks (case-insensitive lookup). Workaround: build from agit archivetar context.Testing
make qualify(test-coverage + lint + tuning-check + e2e + scan + license-check) passes, exit 0.make test-coverage: 78.7% total (threshold 75%) — passes.validators/performance: 50.1% -> 50.3% (+0.2% vs origin/main baseline via worktree method). No new exported functions (all new helpers are unexported and covered:discoverAKSRdmaCount100%,buildAKSRdmaResourceLine100%,applyAKSTemplateDatacovered byTestApplyAKSTemplateData).golangci-lint run -c .golangci.yaml ./validators/performance/...— 0 issues.nccl_aks_utils_test.go), plus AKS subtests inTestPlatformWorkerScheduling, an AKS assertion in the support-table test, and the template wiring-guard now parsestestdata/h100/aks/runtime.yaml.aicr recipe --service aks --accelerator h100 --intent training --os ubuntuemits the performance block with the>= 150constraint.h100-aks-ubuntu-training-kubeflow, ran the performance phase with a dev validator image (ghcr.io/yuanchen8911/aicr-validators/performance:aks-nccl-dev, linux/amd64).nccl-all-reduce-bwwas selected and passed: 157.24 GB/s busbw at 16 GiB all-reduce across 2x ND96isr_H100_v5 over IB (mlx5), against the >= 150 floor (10% tolerance). The real image ships via CI on merge; the dev image was only used to test pre-merge.Risk Assessment
Rollout notes: Additive: a new (service, accelerator) tuple in the validator support table and a new performance block in the AKS overlay. Existing services/recipes are untouched; AKS clusters without the RDMA shared device plugin degrade to the TCP fallback rather than failing. Revert = drop the overlay block and the AKS table entry.
Checklist
make testwith-race)make lint)git commit -S) — GPG signing info