feat(uat): nvkind H100 real-hardware evidence lane (DC5) - #1843
Conversation
Add a service:kind UAT lane that drives the h100-kind recipes on real
silicon via nvkind on a self-hosted GPU runner, emits a Sigstore-signed
ADR-007 evidence bundle, and ingests it onto validation.aicr.run as a
first-party corroboration source. Scope is honestly H100 x1, single-GPU.
The lane is a full sibling of the cloud UAT lanes: reservation registry
(kind-h100 row, cloud: kind) -> uat-run.yaml (run-kind) -> uat-kind.yaml
-> the shared tests/uat/kind/run shim over tests/uat/lib/phases.sh. It
rides the same nightly batch, reservation lease, version matrix, and
evidence emit -> verify -> ingest as aws/gcp/azure; it differs only in
provisioning (gpu-cluster-setup / gpu-test-cleanup, no cloud actuator).
- broker: recognize cloud: kind (pkg/uatbroker)
- registry: kind-h100 row + committed-registry test coverage
- uat-run.yaml: run-kind dispatch + unmapped-cloud guard
- uat-kind.yaml: version-parameterized install (ko.local for main cells,
released images for release cells), setup-build-tools, phase-by-phase
- tests/uat/kind: thin run shim + h100-{training,inference} AICRConfigs
- phases.sh: parametrize TrainJob numNodes (default 2; kind overrides 1)
- allowlist + evidence-ingest: register uat-kind first-party signer
- docs: uat.md + evidence-ingest.md
Refs NVIDIA#1278
Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
|
🌿 Preview your docs: https://nvidia-preview-uat-nvkind-h100-evidence-lane.docs.buildwithfern.com/aicr |
Recipe evidence checkNo leaf overlays affected by this PR. This gate is warning-only and never blocks merge. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Enterprise Run ID: 📒 Files selected for processing (7)
💤 Files with no reviewable changes (2)
📝 WalkthroughWalkthroughAdds the Estimated code review effort: 4 (Complex) | ~60 minutes Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/uat-run.yaml:
- Around line 244-250: Update the dispatch-routing comment immediately above the
nvkind job to include run-kind alongside run-aws, run-gcp, and run-azure; change
only the comment so it accurately lists all supported routes.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: 80412c14-609a-4423-89c7-f780aa7c0aea
📒 Files selected for processing (15)
.github/workflows/evidence-ingest.yaml.github/workflows/uat-kind.yaml.github/workflows/uat-run.yamldocs/contributor/evidence-ingest.mddocs/contributor/uat.mdinfra/uat/reservations.yamlpkg/uatbroker/model.gopkg/uatbroker/registry.gopkg/uatbroker/registry_test.gorecipes/evidence/allowlist.yamltests/uat/kind/cluster-config.yamltests/uat/kind/runtests/uat/kind/tests/h100-inference-config.yamltests/uat/kind/tests/h100-training-config.yamltests/uat/lib/phases.sh
mchmarny
left a comment
There was a problem hiding this comment.
The new hardware lane has runtime blockers before it can produce and publish trustworthy evidence. Please address the inline findings and complete a green manual H100 acceptance run before nightly enrollment.
- reservations: opt kind-h100 OUT of the nightly batch (nightly-intents: []) until a green manual H100 acceptance run — no unvalidated multi-hour runs scheduled on the shared runner on merge (mchmarny) - uat-kind config: drop requireGpu on the snapshot + validate agents — the agent must not hold the single node's only GPU (starves check/snapshot workloads) and GPU detection is NFD/PCI-based; snapshot no longer Pending at prep (pre-device-plugin) (mchmarny) - uat-kind: pin AICR_VALIDATOR_IMAGE_TAG=latest for main cells — the commit stamp makes the catalog rewrite :latest -> :sha-<commit>, but aicr-build only side-loads :latest, so conformance ImagePullBackOff'd (mchmarny) - uat-run/uat-kind: grant actions: write on the run-kind caller + ingest job so the nested evidence-ingest dashboard-publish dispatch is not denied (a reusable workflow cannot elevate the caller token) (mchmarny) - uat-kind: raise job timeout 180 -> 280 to cover the sequential step budgets plus always() cleanup headroom, matching the cloud lanes — a mid-run timeout would leak the nvkind cluster (mchmarny) - uat-run: update the stale dispatch-routing comment to include run-kind (coderabbitai) - docs/test: reflect kind-h100 nightly opt-out Refs NVIDIA#1278 Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
|
Thanks @mchmarny — all six findings addressed in 3d2b679. Every one was valid; summary:
On the green manual H100 acceptance run: agreed it's a prerequisite for nightly enrollment — that's exactly why the row now ships opted-out ( |
|
@mchmarny heads-up: this latest review evaluated
No code change needed for this pass — could you re-review against |
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under
a shared-budget retry loop (mirroring `install_helmfile`'s discipline:
`timeout ${remaining}` per attempt + budget re-check between attempts
so a stalled apiserver cannot burn the full sync-wait window), and
waits for every Application to reach a terminal-pass state using the
same 4-arm predicate the KWOK chainsaw sync gate encodes.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(GitHub Actions auto-injects GITHUB_ACTOR but not the token) so the
in-cluster repo-creds Secret can be provisioned from it; the helmfile
branch never reads it, so the existing lane is unaffected. The post-
install readiness gate stays deployer-agnostic (it validates deployed
cluster state, not deployment mechanism), so a green argocd cell is
direct evidence the GitOps deploy path converges on the same operator-
managed stack the helmfile lane validates.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop (`timeout ${remaining}` per attempt + budget
re-check between attempts, mirroring install_helmfile), waits for the
root `nvidia-stack` Application to be reified under a bounded root-app
grace (per-invocation `timeout` + sleep capped to remaining budget),
and waits for every Application to reach a terminal-pass state using
the same 4-arm predicate the KWOK chainsaw sync gate encodes. The
failure diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates. Install step `timeout-minutes` sized at 110m to fit the
larger argocd budget (helm 5m + apply/sync 30m + root grace 2m + gate
60m = 97m), keeping the fail-closed `::error::` paths reachable
before GitHub Actions kills the step.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified under a bounded root-app grace, and waits
for every Application to reach a terminal-pass state using the same
4-arm predicate the KWOK chainsaw sync gate encodes.
Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Fails closed at every step; the whole install step also has
a step-level 110m cap in the workflow to fit the argocd branch's
total budget (helm 5m + apply/sync 30m + root grace 2m + gate 60m =
97m) and keep phases.sh's own `::error::` paths reachable.
Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.
Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (root-app 5s, sync-wait 15s) is capped to the
remaining budget so a nap near the deadline cannot overrun it. Each
retry loop re-checks `SECONDS >= deadline` before starting the next
attempt. Root-app grace deadline is additionally capped at the shared
argocd_deadline so an upstream step that spent most of the shared
budget can never let the root grace add its full 2m on top. Fails
closed at every step; the whole install step also has a step-level
110m cap in the workflow to fit the argocd branch's total budget
(helm 5m + apply/sync 30m + root grace 2m + gate 60m = 97m) and keep
phases.sh's own `::error::` paths reachable.
Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.
Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.
Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds an optional `deployer` dispatch input to
uat-run.yaml + uat-aws.yaml (empty = existing helmfile behavior;
`argocd` = load `<accelerator>-<intent>-argocd-config.yaml`). The
`Validate inputs` step allowlists the deployer value (empty|argocd)
after AWS_ACCOUNT_ID export so daytime-down teardown still
authenticates.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN`, `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state using the same 4-arm predicate the KWOK
chainsaw sync gate encodes.
Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`; every
sleep between polls (app-of-apps retry 15s, root-app 5s, sync-wait
15s) is capped to the remaining budget so a nap near the deadline
cannot overrun it. Each retry loop re-checks `SECONDS >= deadline`
before starting the next attempt. Root-app grace deadline is
additionally capped at the shared argocd_deadline so an upstream step
that spent most of the shared budget can never let the root grace add
its full 2m on top. Fails closed at every step; the whole install
step also has a step-level 110m cap in the workflow to fit the argocd
branch's total budget (helm 5m + apply/sync 30m + root grace 2m +
gate 60m = 97m) and keep phases.sh's own `::error::` paths reachable.
Failure-diagnostic path is best-effort throughout (kubectl calls all
`|| true`) so a transient apiserver hiccup can't skip the describe /
repo-server logs a reviewer needs. Sync-wait `bad` variable is
initialized with a sentinel so the timeout diagnostic reads sensibly
even if the loop never executed.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is explicitly passed into the install step's env
(Actions auto-injects GITHUB_ACTOR but not the token) so the in-cluster
repo-creds Secret can be provisioned from it; the helmfile branch never
reads it. The post-install readiness gate stays deployer-agnostic (it
validates deployed cluster state, not deployment mechanism), so a green
argocd cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds a `deployer` dispatch input to uat-run.yaml +
uat-aws.yaml (values `helmfile|argocd`; `helmfile` is the default and
selects the existing `<accelerator>-<intent>-config.yaml` filename,
`argocd` selects `<accelerator>-<intent>-argocd-config.yaml`). The
choice uses explicit `helmfile` rather than an empty string because
actionlint rejects empty strings in a choice's `options:`. The
`Validate inputs` step allowlists the deployer value after
AWS_ACCOUNT_ID export so daytime-down teardown still authenticates.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a green
manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100 (NVIDIA#1843)
onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Adds end-to-end UAT coverage for the `--deployer argocd` GitOps path,
which today has unit + KWOK-sync coverage but has never been exercised
on real GPU hardware. Refactors `phase_install` to dispatch on the
config's `spec.bundle.deployment.deployer` (existing helmfile body
moved into `install_helmfile`, byte-equivalent), and adds
`install_argocd`: helm-installs the pinned argo-cd chart from
`.settings.yaml`, provisions a prefix-matched ghcr.io repo-creds
Secret from `GITHUB_TOKEN` (via `kubectl create --from-literal` +
`kubectl label --local` + `kubectl apply` so the token/actor never
traverse a YAML parser), `kubectl apply`s the app-of-apps under a
shared-budget retry loop, waits for the root `nvidia-stack`
Application to be reified (grace capped at the shared budget so the
loop can never outlast it), and waits for every Application to reach
a terminal-pass state. The terminal-pass predicate has two premature-
convergence guards before the 4-arm KWOK-parity check: fails when
items is empty ("no Applications yet") and when the root Application
is not Synced ("children may not be reified"), so the root-only race
where the app-of-apps root is OutOfSync+Healthy but its child
Applications have not yet been generated cannot flip `bad=""` and
return 0 before gpu-operator/DRA exist.
Shared-budget discipline (mirroring install_helmfile): a single
ARGOCD_SYNC_TIMEOUT_SECONDS wall clock spans the whole install path
from repo-creds Secret apply through terminal-pass. Every kubectl
invocation (Secret apply, app-of-apps apply retries, root-app grace
poll, terminal-pass poll) is bounded by `timeout ${remaining}`;
every sleep between polls (app-of-apps retry 15s, root-app 5s, sync-
wait 15s) is capped to the remaining budget. Each retry loop re-
checks `SECONDS >= deadline` before starting the next attempt.
Root-app grace deadline is additionally capped at the shared
argocd_deadline; when that cap fires the root-app-not-reified error
reports the effective window rather than the nominal grace so a
budget-starved run doesn't look like a 120s hang. Every fail-closed
path in install_argocd emits `::error::` and closes the ::group::
before exiting (including the `helm repo update argo` and `kubectl
wait crd/applications.argoproj.io` steps that used to abort silently
under `set -euo pipefail`).
Failure-diagnostic path is best-effort throughout (kubectl calls
all `|| true`) so a transient apiserver hiccup can't skip the
describe / repo-server logs a reviewer needs. Sync-wait `bad`
variable is initialized with a sentinel so the timeout diagnostic
reads sensibly even if the loop never executed.
The bundle reaches Argo CD via `aicr bundle --output oci://... --repo
oci://...` -- pushed to ghcr.io/nvidia/aicr-bundle-scratch under the
existing `packages: write` scope; a per-run tag isolates concurrent
runs. `GITHUB_TOKEN` is passed into the install step's env ONLY when
the argocd branch is selected (`inputs.deployer == 'argocd' &&
github.token || ''`), so helmfile's kubectl/aicr/chart-hook
children don't inherit it; `install_argocd`'s `:?` guard fires on
the empty default if a caller misconfigures the deployer input. The
post-install readiness gate stays deployer-agnostic (it validates
deployed cluster state, not deployment mechanism), so a green argocd
cell is direct evidence the GitOps path converges on the same
operator-managed stack the helmfile lane validates.
Workflow surface: adds a `deployer` dispatch input to uat-run.yaml
(values `helmfile|argocd`; choice uses explicit `helmfile` rather
than an empty string because actionlint rejects empty strings in a
choice's `options:`). Only `run-aws` forwards the input to its
reusable pipeline. gcp/azure/kind `run-*.if:` skips the reusable
pipeline when `deployer != 'helmfile'` so no cluster is provisioned
for a request that cannot be served, and a new top-level
`unsupported-deployer-for-cloud` job fails the workflow with a clear
`::error::` when the (deployer, cloud) pair is unsupported --
mirroring the `unmapped-cloud` pattern so a dispatch never silently
downgrades. The `Validate inputs` step in uat-aws.yaml allowlists
the deployer value after `AWS_ACCOUNT_ID` export so daytime-down
teardown authentication remains available. Install step
`timeout-minutes` sized at 110 to fit argocd's total budget
(helm 5 + apply/sync 30 + root grace 2 + gate 60 = 97m); the job-
budget comment (uat-aws.yaml:98-121) is re-derived with a
per-(deployer, intent) breakdown and calls out the future
argocd-inference variant that would push the total past the 300m
job cap.
UI: run-name and the AWS Test Summary now surface the deployer
segment when it deviates from `helmfile`, so helmfile cells' titles
are unchanged and argocd cells are trivially distinguishable in the
Actions list without opening the run. The AWS install step's
GITHUB_TOKEN scope is now covered by
`TestCredentialBearingUATStepsDisableXtrace` via a new
`awsTokenBearingStepNames` slice so a future `set -x` regression in
the install path (invisible in an env-only diff) is caught before
the token can leak into log lines.
Retention follow-up: run-scoped bundle artifacts under
ghcr.io/nvidia/aicr-bundle-scratch are not cleaned up on the success
path. Deferred while this cell is manual-dispatch-only (low
accumulation); noted inline for nightly-enrollment follow-up via
either a workflow teardown step or an org-level retention policy.
Manual dispatch only on `aws-h100` training for now -- nightly
enrollment and extension to other cells are follow-ups after a
green manual run, mirroring the azure-h100 (NVIDIA#1722) and kind-h100
(NVIDIA#1843) onboarding pattern. `argocd-helm` variant deferred.
Fixes: NVIDIA#2194
Signed-off-by: framsouza <fram.souza14@gmail.com>
Summary
Adds a
service: kindUAT lane (uat-kind) that drives theh100-kindrecipes on real silicon via nvkind on a self-hosted GPU runner, emits a Sigstore-signed ADR-007 evidence bundle, and ingests it onto validation.aicr.run as a first-party corroboration source. Scope is honestly H100 ×1, single-GPU — whatever the runner physically has.Motivation / Context
DC5 (rescoped): give the dashboard a real-silicon corroboration source for the
service: kindrecipes. Rather than a bespoke workflow, this lands the nvkind lane as a full sibling of the cloud UAT lanes so it rides the same nightly batch, reservation lease, version matrix, and evidence pipeline.Fixes: N/A (partial — DC5 launch deliverable)
Related: #1278, #1264
Type of Change
Component(s) Affected
tests/uat,.github/workflows/uat-*), reservation broker (pkg/uatbroker), recipe evidence allowlist, docsImplementation Notes
The lane is a sibling of
uat-aws/uat-gcp/uat-azure, dispatched through the shareduat-run.yaml:pkg/uatbrokerrecognizescloud: kind(a self-hosted GPU-runner lane, not a cloud);infra/uat/reservations.yamlgains akind-h100row.uat-run.yamlgains arun-kindjob (+ unmapped-cloud guard);uat-kind.yamlis the reusable lane and holds the keyless signing identity.tests/uat/kind/runis a thin shim overtests/uat/lib/phases.sh(same as the cloud runs);numNodesis parametrized inphases.sh(default 2 for the 2-GPU cloud pools; kind overrides to 1). Configs aretests/uat/kind/tests/h100-{training,inference}-config.yaml.mainbuilds from source and side-loadsko.localvalidator/agent images into kind; release cells install the releasedaicrand self-resolve released images.gpu-cluster-setup/gpu-test-cleanupinstead of agithub.com/mchmarny/clusteractuator — no cloud credentials, no capacity reservation (the runner is the lease).recipes/evidence/allowlist.yaml+evidence-ingest.yaml'sFIRST_PARTY_IDENTITYregisteruat-kind.yaml.Kept off the merge gate (nightly cron + manual dispatch via
uat-run.yaml, exactly like the cloud lanes).Testing
Not
make qualify— the lane's end-to-end path requires the self-hosted H100 runner (no GPU in CI here). Ran the affected surfaces locally:Not yet hardware-validated. Dispatch
gh workflow run uat-run.yaml -f reservation=kind-h100on the GPU runner to validate. Known open risks: (1) Fulcio/Rekor egress from the self-hosted runner for keyless signing; (2)--phase alldeployment/nodewright checks andhelmfileinstall on a single-node nvkind cluster (the existing nvkind lanes run conformance-only viadeploy.sh). If the first nightly is red, flipkind-h100'snightly-intentsto[](bring-up opt-out) or add anightly-intent-min-versionsgate.Risk Assessment
Rollout notes: Off the merge gate; nightly/dispatch only. The broker change is additive (
kindaccepted alongsideaws|gcp|azure); cloud lanes are unchanged (phases.shnumNodesdefault stays 2). Revert = drop thekind-h100row +run-kindjob.Checklist
-raceviamake testtargets)docs/contributor/{uat,evidence-ingest}.md)git commit -S)