Skip to content

feat(uat): nvkind H100 real-hardware evidence lane (DC5) - #1843

Merged
mchmarny merged 3 commits into
NVIDIA:mainfrom
njhensley:uat/nvkind-h100-evidence-lane
Jul 22, 2026
Merged

mchmarny merged 3 commits into
NVIDIA:mainfrom
njhensley:uat/nvkind-h100-evidence-lane

Conversation

@njhensley

Copy link
Copy Markdown
Member

Summary

Adds a service: kind UAT lane (uat-kind) that drives the h100-kind recipes on real silicon via nvkind on a self-hosted GPU runner, emits a Sigstore-signed ADR-007 evidence bundle, and ingests it onto validation.aicr.run as a first-party corroboration source. Scope is honestly H100 ×1, single-GPU — whatever the runner physically has.

Motivation / Context

DC5 (rescoped): give the dashboard a real-silicon corroboration source for the service: kind recipes. Rather than a bespoke workflow, this lands the nvkind lane as a full sibling of the cloud UAT lanes so it rides the same nightly batch, reservation lease, version matrix, and evidence pipeline.

Fixes: N/A (partial — DC5 launch deliverable)
Related: #1278, #1264

Type of Change

  • New feature (non-breaking change that adds functionality)
  • Build/CI/tooling
  • Documentation update

Component(s) Affected

  • Other: UAT (tests/uat, .github/workflows/uat-*), reservation broker (pkg/uatbroker), recipe evidence allowlist, docs

Implementation Notes

The lane is a sibling of uat-aws/uat-gcp/uat-azure, dispatched through the shared uat-run.yaml:

  • broker: pkg/uatbroker recognizes cloud: kind (a self-hosted GPU-runner lane, not a cloud); infra/uat/reservations.yaml gains a kind-h100 row.
  • dispatch: uat-run.yaml gains a run-kind job (+ unmapped-cloud guard); uat-kind.yaml is the reusable lane and holds the keyless signing identity.
  • shared runner: tests/uat/kind/run is a thin shim over tests/uat/lib/phases.sh (same as the cloud runs); numNodes is parametrized in phases.sh (default 2 for the 2-GPU cloud pools; kind overrides to 1). Configs are tests/uat/kind/tests/h100-{training,inference}-config.yaml.
  • version axis: version-parameterized install like the cloud lanes — main builds from source and side-loads ko.local validator/agent images into kind; release cells install the released aicr and self-resolve released images.
  • provisioning is the only real difference: gpu-cluster-setup / gpu-test-cleanup instead of a github.com/mchmarny/cluster actuator — no cloud credentials, no capacity reservation (the runner is the lease).
  • signer: recipes/evidence/allowlist.yaml + evidence-ingest.yaml's FIRST_PARTY_IDENTITY register uat-kind.yaml.

Kept off the merge gate (nightly cron + manual dispatch via uat-run.yaml, exactly like the cloud lanes).

Testing

Not make qualify — the lane's end-to-end path requires the self-hosted H100 runner (no GPU in CI here). Ran the affected surfaces locally:

go test ./pkg/uatbroker/... ./tools/uat-broker/... ./tests/uat/... ./pkg/evidence/allowlist/...   # all green
actionlint .github/workflows/uat-kind.yaml .github/workflows/uat-run.yaml                          # clean
shellcheck -x tests/uat/kind/run tests/uat/lib/phases.sh                                            # clean
yamllint (changed yaml)                                                                             # clean

Not yet hardware-validated. Dispatch gh workflow run uat-run.yaml -f reservation=kind-h100 on the GPU runner to validate. Known open risks: (1) Fulcio/Rekor egress from the self-hosted runner for keyless signing; (2) --phase all deployment/nodewright checks and helmfile install on a single-node nvkind cluster (the existing nvkind lanes run conformance-only via deploy.sh). If the first nightly is red, flip kind-h100's nightly-intents to [] (bring-up opt-out) or add a nightly-intent-min-versions gate.

Risk Assessment

  • Medium — New CI lane; broker change touches shared reservation parsing.

Rollout notes: Off the merge gate; nightly/dispatch only. The broker change is additive (kind accepted alongside aws|gcp|azure); cloud lanes are unchanged (phases.sh numNodes default stays 2). Revert = drop the kind-h100 row + run-kind job.

Checklist

  • Tests pass locally (affected Go packages, -race via make test targets)
  • Linter passes (actionlint/yamllint/shellcheck/golangci-lint on touched paths)
  • I did not skip/disable tests to make CI green
  • I added/updated tests for new functionality (broker committed-registry coverage)
  • I updated docs if user-facing behavior changed (docs/contributor/{uat,evidence-ingest}.md)
  • Changes follow existing patterns in the codebase
  • Commits are cryptographically signed (git commit -S)
Add a service:kind UAT lane that drives the h100-kind recipes on real
silicon via nvkind on a self-hosted GPU runner, emits a Sigstore-signed
ADR-007 evidence bundle, and ingests it onto validation.aicr.run as a
first-party corroboration source. Scope is honestly H100 x1, single-GPU.

The lane is a full sibling of the cloud UAT lanes: reservation registry
(kind-h100 row, cloud: kind) -> uat-run.yaml (run-kind) -> uat-kind.yaml
-> the shared tests/uat/kind/run shim over tests/uat/lib/phases.sh. It
rides the same nightly batch, reservation lease, version matrix, and
evidence emit -> verify -> ingest as aws/gcp/azure; it differs only in
provisioning (gpu-cluster-setup / gpu-test-cleanup, no cloud actuator).

- broker: recognize cloud: kind (pkg/uatbroker)
- registry: kind-h100 row + committed-registry test coverage
- uat-run.yaml: run-kind dispatch + unmapped-cloud guard
- uat-kind.yaml: version-parameterized install (ko.local for main cells,
  released images for release cells), setup-build-tools, phase-by-phase
- tests/uat/kind: thin run shim + h100-{training,inference} AICRConfigs
- phases.sh: parametrize TrainJob numNodes (default 2; kind overrides 1)
- allowlist + evidence-ingest: register uat-kind first-party signer
- docs: uat.md + evidence-ingest.md

Refs NVIDIA#1278

Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
@njhensley
njhensley requested review from a team as code owners July 21, 2026 22:14
@njhensley njhensley self-assigned this Jul 21, 2026
@njhensley njhensley added the theme/ci-dx CI pipelines, developer experience, and build tooling label Jul 21, 2026
@github-actions

Copy link
Copy Markdown
Contributor
@github-actions

Copy link
Copy Markdown
Contributor

Recipe evidence check

No leaf overlays affected by this PR.

This gate is warning-only and never blocks merge.

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 7e59f5be-bd64-46f3-8c69-4a5ea050bb3c

📥 Commits

Reviewing files that changed from the base of the PR and between c01fc42 and 3d2b679.

📒 Files selected for processing (7)
  • .github/workflows/uat-kind.yaml
  • .github/workflows/uat-run.yaml
  • docs/contributor/uat.md
  • infra/uat/reservations.yaml
  • pkg/uatbroker/registry_test.go
  • tests/uat/kind/tests/h100-inference-config.yaml
  • tests/uat/kind/tests/h100-training-config.yaml
💤 Files with no reviewable changes (2)
  • tests/uat/kind/tests/h100-inference-config.yaml
  • tests/uat/kind/tests/h100-training-config.yaml

📝 Walkthrough

Walkthrough

Adds the kind cloud and kind-h100 reservation, routes it through a reusable uat-kind workflow, provisions a single-node H100 kind cluster, and runs training or inference UAT phases. It adds kind-specific test configurations, configurable training node counts, evidence trust and ingest integration, cleanup/debug handling, and contributor documentation updates.

Estimated code review effort: 4 (Complex) | ~60 minutes

Suggested labels: area/validator, theme/supply-chain

Suggested reviewers: yuanchen8911

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: adding the nvkind H100 real-hardware evidence lane.
Description check ✅ Passed The description is directly related to the changeset and explains the new UAT lane, broker updates, and evidence flow.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/uat-run.yaml:
- Around line 244-250: Update the dispatch-routing comment immediately above the
nvkind job to include run-kind alongside run-aws, run-gcp, and run-azure; change
only the comment so it accurately lists all supported routes.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 80412c14-609a-4423-89c7-f780aa7c0aea

📥 Commits

Reviewing files that changed from the base of the PR and between 9e8e572 and c01fc42.

📒 Files selected for processing (15)
  • .github/workflows/evidence-ingest.yaml
  • .github/workflows/uat-kind.yaml
  • .github/workflows/uat-run.yaml
  • docs/contributor/evidence-ingest.md
  • docs/contributor/uat.md
  • infra/uat/reservations.yaml
  • pkg/uatbroker/model.go
  • pkg/uatbroker/registry.go
  • pkg/uatbroker/registry_test.go
  • recipes/evidence/allowlist.yaml
  • tests/uat/kind/cluster-config.yaml
  • tests/uat/kind/run
  • tests/uat/kind/tests/h100-inference-config.yaml
  • tests/uat/kind/tests/h100-training-config.yaml
  • tests/uat/lib/phases.sh
Comment thread .github/workflows/uat-run.yaml

@mchmarny mchmarny left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new hardware lane has runtime blockers before it can produce and publish trustworthy evidence. Please address the inline findings and complete a green manual H100 acceptance run before nightly enrollment.

Comment thread tests/uat/kind/tests/h100-training-config.yaml Outdated
Comment thread .github/workflows/uat-kind.yaml
Comment thread .github/workflows/uat-run.yaml Outdated
Comment thread infra/uat/reservations.yaml Outdated
Comment thread .github/workflows/uat-kind.yaml Outdated
- reservations: opt kind-h100 OUT of the nightly batch (nightly-intents: [])
  until a green manual H100 acceptance run — no unvalidated multi-hour runs
  scheduled on the shared runner on merge (mchmarny)
- uat-kind config: drop requireGpu on the snapshot + validate agents — the
  agent must not hold the single node's only GPU (starves check/snapshot
  workloads) and GPU detection is NFD/PCI-based; snapshot no longer Pending at
  prep (pre-device-plugin) (mchmarny)
- uat-kind: pin AICR_VALIDATOR_IMAGE_TAG=latest for main cells — the commit
  stamp makes the catalog rewrite :latest -> :sha-<commit>, but aicr-build only
  side-loads :latest, so conformance ImagePullBackOff'd (mchmarny)
- uat-run/uat-kind: grant actions: write on the run-kind caller + ingest job
  so the nested evidence-ingest dashboard-publish dispatch is not denied
  (a reusable workflow cannot elevate the caller token) (mchmarny)
- uat-kind: raise job timeout 180 -> 280 to cover the sequential step budgets
  plus always() cleanup headroom, matching the cloud lanes — a mid-run timeout
  would leak the nvkind cluster (mchmarny)
- uat-run: update the stale dispatch-routing comment to include run-kind
  (coderabbitai)
- docs/test: reflect kind-h100 nightly opt-out

Refs NVIDIA#1278

Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
@njhensley

Copy link
Copy Markdown
Member Author

Thanks @mchmarny — all six findings addressed in 3d2b679. Every one was valid; summary:

# Finding Fix
1 nightly-intents enrolls unvalidated multi-hour runs on merge kind-h100 → nightly-intents: [] (opted OUT during bring-up, mirroring aws-gb200); manual-dispatch only, enrolled after a green H100 acceptance run + verified ingestion. Test/doc updated.
2 requireGpu: true snapshot Pending at prep (pre-device-plugin) Dropped requireGpu on both snapshot and validate agents — the agent must not hold the single node's only nvidia.com/gpu (would starve the check/snapshot workloads), and detection is NFD/PCI-based. Now matches the cloud configs, which omit it.
3 Validator :latest → :sha-<commit> rewrite ImagePullBackOffs Pin AICR_VALIDATOR_IMAGE_TAG=latest for main cells (catalog resolution step 3) so it matches the ko.local/…:latest images aicr-build side-loads. Release cells leave it unset.
4 evidence-ingest dispatch denied (actions: read ceiling) actions: write on the run-kind caller and the ingest-evidence job, matching the cloud lanes.
5 Job cap (180) below sum of step budgets (~250) → cleanup killed, leaked cluster Raised timeout-minutes to 280 (cloud-lane parity) with cleanup headroom.
6 Stale dispatch-routing comment (CodeRabbit) Comment now lists run-kind.

On the green manual H100 acceptance run: agreed it's a prerequisite for nightly enrollment — that's exactly why the row now ships opted-out ([]). The lane stays manually dispatchable (gh workflow run uat-run.yaml -f reservation=kind-h100 -f intent=training); I'll enroll intents only after a green run confirms evidence emit/sign (incl. Fulcio/Rekor egress from the runner) + ingestion.

@njhensley
njhensley requested a review from mchmarny July 21, 2026 23:44
@njhensley

Copy link
Copy Markdown
Member Author

@mchmarny heads-up: this latest review evaluated c01fc42 (the original commit) — 4 of the 5 inline comments are anchored to c01fc42, not the fix commit. All five were already addressed in 3d2b679 (pushed before this pass). Current HEAD = 3d2b679:

  1. requireGpu — removed from both the snapshot and validate agents in tests/uat/kind/tests/h100-{training,inference}-config.yaml (grep: 0 occurrences). The single node's only GPU goes to the check/snapshot workloads, detection is NFD/PCI-based, and it matches the cloud configs.
  2. Validator :sha-<commit> rewrite — uat-kind.yaml:193 now sets AICR_VALIDATOR_IMAGE_TAG=latest for main cells (your recommended fix).
  3. actions: write — set at both layers: uat-run.yaml:256 (run-kind) and uat-kind.yaml:392 (ingest-evidence).
  4. nightly-intents: [] — reservations.yaml:178 is now the explicit opt-out; test guard + docs updated.
  5. Job timeout — uat-kind.yaml:91 raised to 280.

No code change needed for this pass — could you re-review against 3d2b679? (Re-requesting.) If your tooling is pinned to c01fc42, that's why the findings re-posted verbatim.

framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under
a shared-budget retry loop (mirroring `install_helmfile`'s discipline:
`timeout ${remaining}` per attempt + budget re-check between attempts
so a stalled apiserver cannot burn the full sync-wait window), and
waits for every Application to reach a terminal-pass state using the
same 4-arm predicate the KWOK chainsaw sync gate encodes.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(GitHub Actions auto-injects GITHUB_ACTOR but not the token) so the
in-cluster repo-creds Secret can be provisioned from it; the helmfile
branch never reads it, so the existing lane is unaffected. The post-
install readiness gate stays deployer-agnostic (it validates deployed
cluster state, not deployment mechanism), so a green argocd cell is
direct evidence the GitOps deploy path converges on the same operator-
managed stack the helmfile lane validates.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop (`timeout ${remaining}` per attempt + budget
re-check between attempts, mirroring install_helmfile), waits for the
root `nvidia-stack` Application to be reified under a bounded root-app
grace (per-invocation `timeout` + sleep capped to remaining budget),
and waits for every Application to reach a terminal-pass state using
the same 4-arm predicate the KWOK chainsaw sync gate encodes. The
failure diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates. Install step `timeout-minutes` sized at 110m to fit the
larger argocd budget (helm 5m + apply/sync 30m + root grace 2m + gate
60m = 97m), keeping the fail-closed `::error::` paths reachable
before GitHub Actions kills the step.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified under a bounded root-app grace, and waits
for every Application to reach a terminal-pass state using the same
4-arm predicate the KWOK chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Fails closed at every step; the whole install step also has
a step-level 110m cap in the workflow to fit the argocd branch's
total budget (helm 5m + apply/sync 30m + root grace 2m + gate 60m =
97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Root-app grace deadline is additionally capped at the shared
argocd_deadline so an upstream step that spent most of the shared
budget can never let the root grace add its full 2m on top. Fails
closed at every step; the whole install step also has a step-level
110m cap in the workflow to fit the argocd branch's total budget
(helm 5m + apply/sync 30m + root grace 2m + gate 60m = 97m) and keep
phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 18, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 20, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.

Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds a `deployer` dispatch input to uat-run.yaml +
uat-aws.yaml (values `helmfile|argocd`; `helmfile` is the default and
selects the existing `<accelerator>-<intent>-config.yaml` filename,
`argocd` selects `<accelerator>-<intent>-argocd-config.yaml`). The
choice uses explicit `helmfile` rather than an empty string because
actionlint rejects empty strings in a choice's `options:`. The
`Validate inputs` step allowlists the deployer value after
AWS_ACCOUNT_ID export so daytime-down teardown still authenticates.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
framsouza added a commit to framsouza/aicr that referenced this pull request Aug 21, 2026
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN` (via `kubectl create --from-literal` +
`kubectl label --local` + `kubectl apply` so the token/actor never
traverse a YAML parser), `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state. The terminal-pass predicate has two premature-
convergence guards before the 4-arm KWOK-parity check: fails when
items is empty ("no Applications yet") and when the root Application
is not Synced ("children may not be reified"), so the root-only race
where the app-of-apps root is OutOfSync+Healthy but its child
Applications have not yet been generated cannot flip `bad=""` and
return 0 before gpu-operator/DRA exist.

Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`;
every sleep between polls (app-of-apps retry 15s, root-app 5s, sync-
wait 15s) is capped to the remaining budget. Each retry loop re-
checks `SECONDS >= deadline` before starting the next attempt.
Root-app grace deadline is additionally capped at the shared
argocd_deadline; when that cap fires the root-app-not-reified error
reports the effective window rather than the nominal grace so a
budget-starved run doesn't look like a 120s hang. Every fail-closed
path in install_argocd emits `::error::` and closes the ::group::
before exiting (including the `helm repo update argo` and `kubectl
wait crd/applications.argoproj.io` steps that used to abort silently
under `set -euo pipefail`).

Failure-diagnostic path is best-effort throughout (kubectl calls
all `|| true`) so a transient apiserver hiccup can't skip the
describe / repo-server logs a reviewer needs. Sync-wait `bad`
variable is initialized with a sentinel so the timeout diagnostic
reads sensibly even if the loop never executed.

The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is passed into the install step's env ONLY when
the argocd branch is selected (`inputs.deployer == 'argocd' &&
github.token || ''`), so helmfile's kubectl/aicr/chart-hook
children don't inherit it; `install_argocd`'s `:?` guard fires on
the empty default if a caller misconfigures the deployer input. The
post-install readiness gate stays deployer-agnostic (it validates
deployed cluster state, not deployment mechanism), so a green argocd
cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.

Workflow surface: adds a `deployer` dispatch input to uat-run.yaml
(values `helmfile|argocd`; choice uses explicit `helmfile` rather
than an empty string because actionlint rejects empty strings in a
choice's `options:`). Only `run-aws` forwards the input to its
reusable pipeline. gcp/azure/kind `run-*.if:` skips the reusable
pipeline when `deployer != 'helmfile'` so no cluster is provisioned
for a request that cannot be served, and a new top-level
`unsupported-deployer-for-cloud` job fails the workflow with a clear
`::error::` when the (deployer, cloud) pair is unsupported --
mirroring the `unmapped-cloud` pattern so a dispatch never silently
downgrades. The `Validate inputs` step in uat-aws.yaml allowlists
the deployer value after `AWS_ACCOUNT_ID` export so daytime-down
teardown authentication remains available. Install step
`timeout-minutes` sized at 110 to fit argocd's total budget
(helm 5 + apply/sync 30 + root grace 2 + gate 60 = 97m); the job-
budget comment (uat-aws.yaml:98-121) is re-derived with a
per-(deployer, intent) breakdown and calls out the future
argocd-inference variant that would push the total past the 300m
job cap.

UI: run-name and the AWS Test Summary now surface the deployer
segment when it deviates from `helmfile`, so helmfile cells' titles
are unchanged and argocd cells are trivially distinguishable in the
Actions list without opening the run. The AWS install step's
GITHUB_TOKEN scope is now covered by
`TestCredentialBearingUATStepsDisableXtrace` via a new
`awsTokenBearingStepNames` slice so a future `set -x` regression in
the install path (invisible in an env-only diff) is caught before
the token can leak into log lines.

Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.

Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a
green manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100
(NVIDIA#1843) onboarding pattern. `argocd-helm` variant deferred.

Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2 participants