Skip to content

feat(uat): Azure AKS UAT phase runner + intent configs (2/2) - #1722

Merged
mchmarny merged 1 commit into
mainfrom
feat/azure-uat-2
Jul 10, 2026
Merged

mchmarny merged 1 commit into
mainfrom
feat/azure-uat-2

Conversation

@mchmarny

Copy link
Copy Markdown
Member

Summary

Completes the Azure AKS UAT pipeline started in #1709 (2/2): the phase runner (tests/uat/azure/run), the per-intent AICRConfig pair (h100-{training,inference}-config.yaml), and the install step re-widened to 90m now that the readiness gate can keep the federated az session alive mid-phase.

Motivation / Context

PR #1709 shipped the account federation, dispatch surface, and pipeline; the burn-in path (provision → connect → destroy) is green on main as of run 29122546423. This PR adds what a full (non-skip_tests) run needs.

The runner is a port of the AWS/GCP siblings (byte-identical there apart from comments) with one Azure-only addition: a federated az session cannot self-refresh (5-minute OIDC assertion, ~60–75m access tokens — the AADSTS700024 failure diagnosed during burn-in), and the post-install readiness gate can outlive one token. The gate therefore re-logins in-loop every AZ_RELOGIN_INTERVAL_SECONDS (25m): mint a fresh ACTIONS_ID_TOKEN, az login --federated-token, then eagerly warm the AKS-audience token (the assertion is only 5-minute valid, so the mint cannot be deferred to kubelogin's next cache miss). Transient re-login failures warn and retry rather than killing the gate.

The configs codify the manually validated AKS commands (aicr-test5): system components on the cpu-worker pool (the AKS system pool carries CriticalAddonsOnly), GPU toleration nvidia.com/gpu=present:NoSchedule only. Deliberately no value overrides — the toolkit RUNTIME_CONFIG_SOURCE env is baked into values-aks.yaml (#1689), and the empty agentgateway:allowedSourceRanges receives the bundler's RFC1918 injection (#1373), sufficient because phase_serve reaches the endpoint via kubectl port-forward, never the LoadBalancer. The NVIDIA-egress override from the manual validation is a human/VPN concern, documented in the inference config.

Fixes: N/A
Related: #1709, #1275, #1276

Type of Change

  • New feature (non-breaking change that adds functionality)
  • Build/CI/tooling

Component(s) Affected

  • Other: UAT pipeline (tests/uat/azure/, .github/workflows/uat-azure.yaml)

Implementation Notes

  • azure-h100 stays nightly-intents: [] (bring-up, manual dispatch only) until a green manual full run; flipping it is a follow-up one-liner in infra/uat/reservations.yaml.
  • Runner divergence from siblings is limited to the re-login machinery and three cloud-specific comments — kept diffable on purpose.

Testing

bash -n tests/uat/azure/run && shellcheck -x tests/uat/azure/run   # clean
yamllint -c .yamllint.yaml tests/uat/azure/tests/*.yaml .github/workflows/uat-azure.yaml  # clean
actionlint .github/workflows/uat-azure.yaml  # no new findings
  • Local end-to-end smoke of both configs against a freshly built aicr: recipe --config resolves (training: 13 components, inference: 16) and bundle --config renders; the inference bundle's rendered gateway carries the RFC1918 loadBalancerSourceRanges and the gpu-operator values carry RUNTIME_CONFIG_SOURCE — matching the validated manual bundles without overrides.
  • No Go changes; coverage gate N/A.
  • Post-merge: manual uat-run dispatch (full run, no skip_tests) is the acceptance test before nightly enrollment.

Risk Assessment

  • Low — Additive files plus a timeout widening; the burn-in path is unaffected, and a revert restores the 60m cap.

Checklist

  • Tests pass locally (make test with -race) — no Go changes
  • Linter passes (make lint)
  • I did not skip/disable tests to make CI green
  • I added/updated tests for new functionality (N/A — CI tooling; validated by dispatch)
  • I updated docs if user-facing behavior changed (N/A — contributor docs already describe the per-intent config contract)
  • Changes follow existing patterns in the codebase
  • Commits are cryptographically signed (git commit -S)
Completes the Azure UAT pipeline started in #1709:

- tests/uat/azure/run: phase runner ported from the AWS/GCP siblings
  (byte-identical there apart from comments), plus an Azure-only
  addition: the post-install readiness gate re-logins the federated az
  session in-loop (fresh GitHub OIDC assertion + eager AKS-audience
  token mint) every AZ_RELOGIN_INTERVAL_SECONDS, because a federated
  session cannot self-refresh and the gate outlives one access token.
- h100-{training,inference}-config.yaml: AICRConfig pair matching the
  manually validated AKS commands — system components on the cpu-worker
  pool (the AKS system pool is CriticalAddonsOnly), GPU pool toleration
  nvidia.com/gpu only. Deliberately no value overrides: the toolkit
  RUNTIME_CONFIG_SOURCE env is baked into values-aks.yaml (#1689), and
  agentgateway's empty allowedSourceRanges gets the bundler's RFC1918
  injection (#1373) — sufficient since phase_serve port-forwards.
- uat-azure.yaml: install re-widened 60m -> 90m for cold-GPU bring-up,
  now safe under the in-loop re-auth.

azure-h100 stays nightly-intents: [] until a green manual full run.

Signed-off-by: Mark Chmarny <mark@chmarny.com>
@mchmarny mchmarny added the theme/validation Constraint evaluation, health checks, and conformance evidence label Jul 10, 2026
@mchmarny
mchmarny requested review from a team as code owners July 10, 2026 21:17
@mchmarny mchmarny added the theme/validation Constraint evaluation, health checks, and conformance evidence label Jul 10, 2026
@mchmarny mchmarny self-assigned this Jul 10, 2026
@coderabbitai

coderabbitai Bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 66ae0060-144e-4851-b3f5-3f3f4694e9ce

📥 Commits

Reviewing files that changed from the base of the PR and between ddea27d and 3c1365d.

📒 Files selected for processing (4)
  • .github/workflows/uat-azure.yaml
  • tests/uat/azure/run
  • tests/uat/azure/tests/h100-inference-config.yaml
  • tests/uat/azure/tests/h100-training-config.yaml

📝 Walkthrough

Walkthrough

Adds a phase-based Azure UAT runner covering preparation, installation, conformance, training, serving, and evidence verification. Adds H100 inference and training AICR configurations with scheduling, validation, deployment, and attestation settings. Updates the Azure workflow’s installation timeout to 90 minutes and revises related timing commentary while keeping the job timeout unchanged.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

  • NVIDIA/aicr#1665: Both changes extend UAT installation timing and deployment readiness validation.
  • NVIDIA/aicr#1709: This change builds on the Azure UAT workflow and runner introduced there.

Suggested reviewers: njhensley

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main addition: an Azure AKS UAT phase runner with intent configs.
Description check ✅ Passed The description matches the changeset and explains the runner, configs, and timeout update.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/azure-uat-2

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

Coverage Report ✅

Metric Value
Coverage 78.9%
Threshold 75%
Status Pass
Coverage Badge
![Coverage](https://img.shields.io/badge/coverage-78.9%25-green)

No Go source files changed in this PR.

@mchmarny
mchmarny merged commit 4a762d2 into main Jul 10, 2026
37 of 38 checks passed
@mchmarny
mchmarny deleted the feat/azure-uat-2 branch July 10, 2026 21:32
Comment thread tests/uat/azure/run
mohityadav8 pushed a commit to mohityadav8/aicr that referenced this pull request Jul 14, 2026
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under
a shared-budget retry loop (mirroring `install_helmfile`'s discipline:
`timeout ${remaining}` per attempt + budget re-check between attempts
so a stalled apiserver cannot burn the full sync-wait window), and
waits for every Application to reach a terminal-pass state using the
same 4-arm predicate the KWOK chainsaw sync gate encodes.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(GitHub Actions auto-injects GITHUB_ACTOR but not the token) so the
in-cluster repo-creds Secret can be provisioned from it; the helmfile
branch never reads it, so the existing lane is unaffected. The post-
install readiness gate stays deployer-agnostic (it validates deployed
cluster state, not deployment mechanism), so a green argocd cell is
direct evidence the GitOps deploy path converges on the same operator-
managed stack the helmfile lane validates.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop (`timeout ${remaining}` per attempt + budget
re-check between attempts, mirroring install_helmfile), waits for the
root `nvidia-stack` Application to be reified under a bounded root-app
grace (per-invocation `timeout` + sleep capped to remaining budget),
and waits for every Application to reach a terminal-pass state using
the same 4-arm predicate the KWOK chainsaw sync gate encodes. The
failure diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates. Install step `timeout-minutes` sized at 110m to fit the
larger argocd budget (helm 5m + apply/sync 30m + root grace 2m + gate
60m = 97m), keeping the fail-closed `::error::` paths reachable
before GitHub Actions kills the step.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified under a bounded root-app grace, and waits
for every Application to reach a terminal-pass state using the same
4-arm predicate the KWOK chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Fails closed at every step; the whole install step also has
a step-level 110m cap in the workflow to fit the argocd branch's
total budget (helm 5m + apply/sync 30m + root grace 2m + gate 60m =
97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Root-app grace deadline is additionally capped at the shared
argocd_deadline so an upstream step that spent most of the shared
budget can never let the root grace add its full 2m on top. Fails
closed at every step; the whole install step also has a step-level
110m cap in the workflow to fit the argocd branch's total budget
(helm 5m + apply/sync 30m + root grace 2m + gate 60m = 97m) and keep
phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 20, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds a `deployer` dispatch input to uat-run.yaml +
uat-aws.yaml (values `helmfile|argocd`; `helmfile` is the default and
selects the existing `<accelerator>-<intent>-config.yaml` filename,
`argocd` selects `<accelerator>-<intent>-argocd-config.yaml`). The
choice uses explicit `helmfile` rather than an empty string because
actionlint rejects empty strings in a choice's `options:`. The
`Validate inputs` step allowlists the deployer value after
AWS_ACCOUNT_ID export so daytime-down teardown still authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 21, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN` (via `kubectl create --from-literal` +
`kubectl label --local` + `kubectl apply` so the token/actor never
traverse a YAML parser), `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state. The terminal-pass predicate has two premature-
convergence guards before the 4-arm KWOK-parity check: fails when
items is empty ("no Applications yet") and when the root Application
is not Synced ("children may not be reified"), so the root-only race
where the app-of-apps root is OutOfSync+Healthy but its child
Applications have not yet been generated cannot flip `bad=""` and
return 0 before gpu-operator/DRA exist.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`;
every sleep between polls (app-of-apps retry 15s, root-app 5s, sync-
wait 15s) is capped to the remaining budget. Each retry loop re-
checks `SECONDS >= deadline` before starting the next attempt.
Root-app grace deadline is additionally capped at the shared
argocd_deadline; when that cap fires the root-app-not-reified error
reports the effective window rather than the nominal grace so a
budget-starved run doesn't look like a 120s hang. Every fail-closed
path in install_argocd emits `::error::` and closes the ::group::
before exiting (including the `helm repo update argo` and `kubectl
wait crd/applications.argoproj.io` steps that used to abort silently
under `set -euo pipefail`).

Failure-diagnostic path is best-effort throughout (kubectl calls
all `|| true`) so a transient apiserver hiccup can't skip the
describe / repo-server logs a reviewer needs. Sync-wait `bad`
variable is initialized with a sentinel so the timeout diagnostic
reads sensibly even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is passed into the install step's env ONLY when
the argocd branch is selected (`inputs.deployer == 'argocd' &&
github.token || ''`), so helmfile's kubectl/aicr/chart-hook
children don't inherit it; `install_argocd`'s `:?` guard fires on
the empty default if a caller misconfigures the deployer input. The
post-install readiness gate stays deployer-agnostic (it validates
deployed cluster state, not deployment mechanism), so a green argocd
cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds a `deployer` dispatch input to uat-run.yaml
(values `helmfile|argocd`; choice uses explicit `helmfile` rather
than an empty string because actionlint rejects empty strings in a
choice's `options:`). Only `run-aws` forwards the input to its
reusable pipeline. gcp/azure/kind `run-*.if:` skips the reusable
pipeline when `deployer != 'helmfile'` so no cluster is provisioned
for a request that cannot be served, and a new top-level
`unsupported-deployer-for-cloud` job fails the workflow with a clear
`::error::` when the (deployer, cloud) pair is unsupported --
mirroring the `unmapped-cloud` pattern so a dispatch never silently
downgrades. The `Validate inputs` step in uat-aws.yaml allowlists
the deployer value after `AWS_ACCOUNT_ID` export so daytime-down
teardown authentication remains available. Install step
`timeout-minutes` sized at 110 to fit argocd's total budget
(helm 5 + apply/sync 30 + root grace 2 + gate 60 = 97m); the job-
budget comment (uat-aws.yaml:98-121) is re-derived with a
per-(deployer, intent) breakdown and calls out the future
argocd-inference variant that would push the total past the 300m
job cap.

UI: run-name and the AWS Test Summary now surface the deployer
segment when it deviates from `helmfile`, so helmfile cells' titles
are unchanged and argocd cells are trivially distinguishable in the
Actions list without opening the run. The AWS install step's
GITHUB_TOKEN scope is now covered by
`TestCredentialBearingUATStepsDisableXtrace` via a new
`awsTokenBearingStepNames` slice so a future `set -x` regression in
the install path (invisible in an env-only diff) is caught before
the token can leak into log lines.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a
green manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100
(NVIDIA#1843) onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/ci area/tests size/XL theme/validation Constraint evaluation, health checks, and conformance evidence

3 participants