Skip to content

feat(validator): AKS H100 NCCL all-reduce performance validation - #1695

Merged
mchmarny merged 1 commit into
NVIDIA:mainfrom
yuanchen8911:feat/nccl-aks-h100
Jul 9, 2026
Merged

mchmarny merged 1 commit into
NVIDIA:mainfrom
yuanchen8911:feat/nccl-aks-h100

Conversation

@yuanchen8911

Copy link
Copy Markdown
Contributor

Summary

Adds AKS (Azure) support to the nccl-all-reduce-bw performance validator — service=aks + accelerator=h100 now runs the NCCL all-reduce bandwidth benchmark over the ND-series InfiniBand fabric — and declares the performance phase in the h100-aks-training overlay with a provisional busbw floor.

Motivation / Context

recipes/overlays/h100-aks-training.yaml carried an intentional-skip comment: the performance phase was deliberately not declared because the validator had no AKS NCCL runtime template, scheduling path, or support-table entry — a declared check would have silently skipped and reported the phase as passed without measuring anything. This PR wires the full path (runtime template + RDMA discovery + worker scheduling + support-table entry + recipe constraint) and replaces that comment with a real performance block, verified end-to-end on a live AKS H100 cluster.

Fixes: N/A
Related: #1337, #1256

Type of Change

  • Bug fix (non-breaking change that fixes an issue)
  • New feature (non-breaking change that adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update
  • Refactoring (no functional changes)
  • Build/CI/tooling

Component(s) Affected

  • CLI (cmd/aicr, pkg/cli)
  • API server (cmd/aicrd, pkg/server)
  • Recipe engine / data (pkg/recipe)
  • Bundlers (pkg/bundler, pkg/component/*)
  • Collectors / snapshotter (pkg/collector, pkg/snapshotter)
  • Validator (pkg/validator)
  • Core libraries (pkg/errors, pkg/k8s)
  • Docs/examples (docs/, examples/)
  • Other: validators/performance (validator pod binary + testdata), recipes/overlays

Implementation Notes

  • Support table: aks -> [h100] added to supportedNCCLCombinations[variantDefault] only. EKS H100 exists only in the default variant, so no -net/-nvls entries (those are GB200-specific).
  • Transport (testdata/h100/aks/runtime.yaml): derived from the EKS H100 TrainingRuntime, same public.ecr.aws/hpc-cloud/nccl-tests image (ships sshd, nccl-tests, rdma-core). NCCL_NET_PLUGIN=none bypasses the bundled aws-ofi-nccl (EFA/Libfabric) plugin so NCCL's built-in IB/verbs transport binds the mlx5 HCAs directly; all FI_*/EFA env dropped; IPC_LOCK retained (ibverbs pinned-buffer registration). NCCL_IB_HCA is deliberately not pinned — NCCL auto-detects HCAs and skips inactive ports.
  • RDMA resource (nccl_aks_utils.go): workers request one unit of rdma/hca_shared_devices_a, the shared pool published by the network-operator rdma-shared-device-plugin (resource name pinned by recipes/components/network-operator/manifests/nic-cluster-policy-aks.yaml). Shared-pool semantics: one unit mounts every /dev/infiniband device, so the request is always "1" regardless of HCA count. Zero allocatable devices falls back to TCP with the reduced 4G message cap, mirroring the EKS zero-EFA behavior.
  • Worker scheduling: AKS joins the OKE branch of platformWorkerScheduling — pin workers to the common GFD nvidia.com/gpu.product label (the same label narrowByAccelerator sized the cohort against) and tolerate the pool taint (nvidia.com/gpu=present:NoSchedule). The AKS-native kubernetes.azure.com/accelerator label was rejected: its value is just nvidia (verified live), so it cannot pin the H100 cohort on a mixed-accelerator cluster.
  • Provisional floor + live baseline: the overlay declares nccl-all-reduce-bw >= 150 — half the calibrated EKS H100 floor (300). Live baseline: 157.24 GB/s busbw measured untuned on 2x Standard_ND96isr_H100_v5 (8x H100 SXM, 8x 400Gb NDR IB per node). 8x NDR line rate is ~400 GB/s, so the floor should be recalibrated upward after IB tuning; the comment in the overlay records this.
  • Known limitation (worth a follow-up issue): building the validator images with docker build from a macOS checkout is broken repo-wide — git tracks both vendor/github.com/NVIDIA and vendor/github.com/nvidia, which collide on case-insensitive filesystems, so the docker build context only ever contains one of them and the in-container go build fails on the k8s-launch-kit imports. CI/Linux is unaffected; local go build works (case-insensitive lookup). Workaround: build from a git archive tar context.

Testing

make qualify
  • Full make qualify (test-coverage + lint + tuning-check + e2e + scan + license-check) passes, exit 0.
  • make test-coverage: 78.7% total (threshold 75%) — passes.
  • Coverage delta validators/performance: 50.1% -> 50.3% (+0.2% vs origin/main baseline via worktree method). No new exported functions (all new helpers are unexported and covered: discoverAKSRdmaCount 100%, buildAKSRdmaResourceLine 100%, applyAKSTemplateData covered by TestApplyAKSTemplateData).
  • golangci-lint run -c .golangci.yaml ./validators/performance/... — 0 issues.
  • New table-driven tests mirror the EKS tests (nccl_aks_utils_test.go), plus AKS subtests in TestPlatformWorkerScheduling, an AKS assertion in the support-table test, and the template wiring-guard now parses testdata/h100/aks/runtime.yaml.
  • Recipe hydration verified: aicr recipe --service aks --accelerator h100 --intent training --os ubuntu emits the performance block with the >= 150 constraint.
  • Live verification (aicr-test1, AKS westus): deployed h100-aks-ubuntu-training-kubeflow, ran the performance phase with a dev validator image (ghcr.io/yuanchen8911/aicr-validators/performance:aks-nccl-dev, linux/amd64). nccl-all-reduce-bw was selected and passed: 157.24 GB/s busbw at 16 GiB all-reduce across 2x ND96isr_H100_v5 over IB (mlx5), against the >= 150 floor (10% tolerance). The real image ships via CI on merge; the dev image was only used to test pre-merge.

Risk Assessment

  • Low — Isolated change, well-tested, easy to revert
  • Medium — Touches multiple components or has broader impact
  • High — Breaking change, affects critical paths, or complex rollout

Rollout notes: Additive: a new (service, accelerator) tuple in the validator support table and a new performance block in the AKS overlay. Existing services/recipes are untouched; AKS clusters without the RDMA shared device plugin degrade to the TCP fallback rather than failing. Revert = drop the overlay block and the AKS table entry.

Checklist

  • Tests pass locally (make test with -race)
  • Linter passes (make lint)
  • I did not skip/disable tests to make CI green
  • I added/updated tests for new functionality
  • I updated docs if user-facing behavior changed
  • Changes follow existing patterns in the codebase
  • Commits are cryptographically signed (git commit -S) — GPG signing info
@yuanchen8911
yuanchen8911 requested review from a team as code owners July 9, 2026 21:10
@yuanchen8911 yuanchen8911 added the theme/validation Constraint evaluation, health checks, and conformance evidence label Jul 9, 2026
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor
@github-actions

github-actions Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

Recipe evidence check

Other affected recipes without evidence yet: 6

These recipes are affected by this PR but carry no committed evidence pointer, so there is
nothing to verify. This is expected — evidence is hardware-gated and added over time.

  • a100-aks-training
  • a100-aks-ubuntu-training-kubeflow
  • a100-aks-ubuntu-training
  • h100-aks-training
  • h100-aks-ubuntu-training-kubeflow
  • h100-aks-ubuntu-training

This gate is warning-only and never blocks merge. See ADR-007 for the trust model.

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

Adds AKS H100 support to NCCL all-reduce bandwidth validation. The change updates documentation and recipe constraints, adds AKS RDMA discovery and runtime template handling, extends validator scheduling and supported combinations, and introduces an AKS H100 TrainingRuntime manifest with accompanying tests.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

  • NVIDIA/aicr#1631: Also changes the nccl-all-reduce-bw validator’s worker-template and scheduling flow, with related platform-specific runtime handling.

Suggested labels: area/tests, area/validator, area/docs

Suggested reviewers: mchmarny, lockwobr

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: adding AKS H100 NCCL all-reduce performance validation.
Description check ✅ Passed The description matches the changeset and explains the AKS NCCL validation path and overlay update.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@yuanchen8911
yuanchen8911 force-pushed the feat/nccl-aks-h100 branch 3 times, most recently from ad305e5 to 67e5544 Compare July 9, 2026 21:29
@yuanchen8911
yuanchen8911 requested review from mchmarny and xdu31 July 9, 2026 21:32
mchmarny
mchmarny previously approved these changes Jul 9, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@validators/performance/nccl_aks_utils.go`:
- Around line 45-64: applyAKSTemplateData currently derives RDMA sizing from
config.Nodes[0], which can misconfigure mixed AKS pools. Update this function to
determine RDMA availability across the full target node set, or explicitly
reject mixed RDMA-capable and non-capable nodes before setting
RDMA_RESOURCE_LIMITS, RDMA_RESOURCE_REQUESTS, and MAX_MESSAGE_SIZE. Keep the
existing logging around discoverAKSRdmaCount/buildAKSRdmaResourceLine, but base
the decision on all nodes instead of the first one.

In `@validators/performance/testdata/h100/aks/runtime.yaml`:
- Line 73: The AKS runtime manifest uses a worker/launcher image from
public.ecr.aws, which can fail on locked-down Azure clusters without outbound
access. Update the AKS GPU setup docs to explicitly note the cross-cloud egress
requirement for this image pull, and consider mirroring the image to an
Azure-reachable registry; reference the runtime image entries and the existing
AKS setup guidance when making the change.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 03965fb3-16fa-40b0-abc4-a2cf32214b19

📥 Commits

Reviewing files that changed from the base of the PR and between ad305e5 and 67e5544.

📒 Files selected for processing (10)
  • docs/integrator/aks-gpu-setup.md
  • docs/user/validation.md
  • examples/recipes/aks-training.yaml
  • recipes/overlays/a100-aks-training.yaml
  • recipes/overlays/h100-aks-training.yaml
  • validators/performance/nccl_aks_utils.go
  • validators/performance/nccl_aks_utils_test.go
  • validators/performance/nccl_all_reduce_bw_constraint.go
  • validators/performance/nccl_test.go
  • validators/performance/testdata/h100/aks/runtime.yaml
Comment thread validators/performance/nccl_aks_utils.go
Comment thread validators/performance/testdata/h100/aks/runtime.yaml
xdu31
xdu31 previously approved these changes Jul 9, 2026

@xdu31 xdu31 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — well-motivated, well-tested, live-validated on ND96isr_H100_v5 hardware. Runtime template, RDMA discovery, worker scheduling, and support-table entry all wire up cleanly and mirror the EKS pattern. Provisional floor is honestly labeled with a clear recalibration path.

Minor follow-ups (non-blocking):

  • Consider a lightweight test asserting the rdmaIndent constant matches the actual column of ${RDMA_RESOURCE_LIMITS} in testdata/h100/aks/runtime.yaml — cheap insurance against future reformatting (applies to EKS efaIndent too).
  • TestSupportedNCCLCombinationsHaveRuntimeTemplates only exercises the populated RDMA line; adding an empty-substitution variant would guard the TCP-fallback rendering path.
  • NCCL_SOCKET_IFNAME=eth0 is hardcoded — worth a comment noting the observed interface at the tested SKU in case a future Azure CNI variant changes it.
Wire service=aks accelerator=h100 into the NCCL all-reduce performance
validator, mirroring the EKS H100 path:

- supportedNCCLCombinations: add AKS+H100 to the default variant.
- testdata/h100/aks/runtime.yaml: TrainingRuntime derived from the EKS
  template with the transport swapped from EFA/Libfabric to NCCL's
  built-in IB/verbs (NCCL_NET_PLUGIN=none) over the ND-series InfiniBand
  HCAs; workers request one unit of the rdma-shared-device-plugin pool
  (rdma/hca_shared_devices_a) plus 8 GPUs and keep IPC_LOCK.
- nccl_aks_utils.go: discover the shared RDMA pool from node allocatable
  and render the worker resource lines; 0 devices falls back to TCP with
  the reduced message-size cap, mirroring the EKS zero-EFA behavior.
- platformWorkerScheduling: AKS shares the OKE shape — pin workers to the
  common GFD nvidia.com/gpu.product label and tolerate the pool taint.
- recipes/overlays/h100-aks-training.yaml: declare the performance phase
  with a provisional >= 150 GB/s busbw floor (half the calibrated EKS
  value) pending calibration on an ND96isr_H100_v5 testbed.
- docs/user/validation.md + examples/recipes/aks-training.yaml updated.

Facts verified live on aicr-test1 (Standard_ND96isr_H100_v5): allocatable
nvidia.com/gpu=8 and rdma/hca_shared_devices_a=1k, GFD gpu.product
NVIDIA-H100-80GB-HBM3, taint nvidia.com/gpu=present:NoSchedule.

Signed-off-by: Yuan Chen <yuanchen97@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/docs area/recipes size/XL theme/validation Constraint evaluation, health checks, and conformance evidence

3 participants