You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Prove the full AICR workflow for k8s-aibom end to end on a real GKE cluster with H100 GPUs, using h100-gke-cos-inference as the target environment: generate the recipe, render the bundle, deploy it, pass validation, and produce a genuine CycloneDX ML-BOM from a real GPU inference workload.
This satisfies ADR-019's Follow-Up Decisions requirement for managed-cluster qualification and measured controller/API-server cost, which is the one requirement in that list that cannot be met by inspection or by a simulated cluster.
Why this is not covered by what already exists
tools/k8s-aibom-test/ runs on Kind. Kind is a useful behavioral proof and it is what qualified the component, but it is a simulation of a managed cluster, not one. The control plane, node images, GPU stack, and API-server behavior all differ.
The operational envelope recorded in ADR-019 (14s convergence at 1,000 workloads, 230.4 mCPU / 47.5 MiB burst) was measured on Kind. GKE's managed control plane is the thing that actually needs qualifying.
Per the ADR-019 amendment, registry presence asserts the component is qualified and offerable; stock recipe presence asserts AICR ships it by default, and that is gated on requalification plus the remaining Follow-Up requirements.
This issue runs from a custom overlay or an unmerged branch. It does not modify recipes/overlays/h100-gke-cos-inference.yaml on main, and no stock recipe changes. That keeps the demonstration honest about what has and has not been committed to.
Scope
On a live GKE cluster with H100 nodes:
aicr recipe for service: gke, accelerator: h100, os: cos, intent: inference, with k8s-aibom declared in a custom overlay.
aicr bundle renders through every supported deployer; confirm the rendered output is what actually deploys.
Deploy the bundle and confirm the component installs cleanly against the managed control plane, including CRD application (see below).
Label a namespace aibom.k8saibom.dev/enabled=true, run a real GPU inference workload, and confirm a correct CycloneDX 1.6 ML-BOM is produced for it. Inline vs. truncated behavior at the 256 KiB threshold should be observed rather than assumed.
Measure controller CPU/memory and API-server impact on GKE, and compare against the Kind-measured envelope in ADR-019. A material divergence is a finding, not a footnote.
The CRD-apply step: the registry now pins chart 1.3.0, whose CRDs differ from 1.2.0. On Helm, Helmfile, and Flux the CRDs must be applied before the bundle upgrade (see the component catalog). Whether that step is needed here depends on whether the cluster is fresh; a fresh install applies crds/ automatically.
Out of scope
Merging k8s-aibom into the stock h100-gke-cos-inference overlay. That is the epic's later step and needs the remaining Follow-Up requirements.
Upgrade, rollback, and uninstall evidence — tracked separately; it gates the overlay merge, not this.
Any AIBOM-as-AICR-evidence integration (ADR-007). Still a separate decision.
Done when
Every step above is executed on a real GKE cluster and the commands and outputs are recorded on this issue.
A real ML-BOM produced from a real inference workload is attached or quoted.
Measured controller and API-server cost is recorded and compared to the ADR-019 envelope.
Any divergence from the Kind-based expectations is written down, including "none observed" if that is the result.
Part of #2271.
Goal
Prove the full AICR workflow for
k8s-aibomend to end on a real GKE cluster with H100 GPUs, usingh100-gke-cos-inferenceas the target environment: generate the recipe, render the bundle, deploy it, pass validation, and produce a genuine CycloneDX ML-BOM from a real GPU inference workload.This satisfies ADR-019's Follow-Up Decisions requirement for managed-cluster qualification and measured controller/API-server cost, which is the one requirement in that list that cannot be met by inspection or by a simulated cluster.
Why this is not covered by what already exists
tools/k8s-aibom-test/runs on Kind. Kind is a useful behavioral proof and it is what qualified the component, but it is a simulation of a managed cluster, not one. The control plane, node images, GPU stack, and API-server behavior all differ.Explicitly not a stock-recipe change
Per the ADR-019 amendment, registry presence asserts the component is qualified and offerable; stock recipe presence asserts AICR ships it by default, and that is gated on requalification plus the remaining Follow-Up requirements.
This issue runs from a custom overlay or an unmerged branch. It does not modify
recipes/overlays/h100-gke-cos-inference.yamlonmain, and no stock recipe changes. That keeps the demonstration honest about what has and has not been committed to.Scope
On a live GKE cluster with H100 nodes:
aicr recipeforservice: gke, accelerator: h100, os: cos, intent: inference, withk8s-aibomdeclared in a custom overlay.aicr bundlerenders through every supported deployer; confirm the rendered output is what actually deploys.aicr validatepasses, including thek8s-aibomhealth check — the Deployment rollout terms, theAIBOMControllerConfigReady condition, and the CRD storage-version assertions added in feat(recipes): assert k8s-aibom CRD storage version in the health check #2305.aibom.k8saibom.dev/enabled=true, run a real GPU inference workload, and confirm a correct CycloneDX 1.6 ML-BOM is produced for it. Inline vs. truncated behavior at the 256 KiB threshold should be observed rather than assumed.Depends on
crds/automatically.Out of scope
k8s-aibominto the stockh100-gke-cos-inferenceoverlay. That is the epic's later step and needs the remaining Follow-Up requirements.Done when