Skip to content

k8s-aibom: validate h100-gke-cos-inference end to end on real GKE #2310

Description

@mchmarny

Part of #2271.

Goal

Prove the full AICR workflow for k8s-aibom end to end on a real GKE cluster with H100 GPUs, using h100-gke-cos-inference as the target environment: generate the recipe, render the bundle, deploy it, pass validation, and produce a genuine CycloneDX ML-BOM from a real GPU inference workload.

This satisfies ADR-019's Follow-Up Decisions requirement for managed-cluster qualification and measured controller/API-server cost, which is the one requirement in that list that cannot be met by inspection or by a simulated cluster.

Why this is not covered by what already exists

  • tools/k8s-aibom-test/ runs on Kind. Kind is a useful behavioral proof and it is what qualified the component, but it is a simulation of a managed cluster, not one. The control plane, node images, GPU stack, and API-server behavior all differ.
  • The operational envelope recorded in ADR-019 (14s convergence at 1,000 workloads, 230.4 mCPU / 47.5 MiB burst) was measured on Kind. GKE's managed control plane is the thing that actually needs qualifying.
  • Registry adoption ([Epic]: k8s-aibom — registry-only component adoption (ADR-019) #2234) deliberately proved the component is safe to offer. It proved nothing about a real cluster running it under a real workload.

Explicitly not a stock-recipe change

Per the ADR-019 amendment, registry presence asserts the component is qualified and offerable; stock recipe presence asserts AICR ships it by default, and that is gated on requalification plus the remaining Follow-Up requirements.

This issue runs from a custom overlay or an unmerged branch. It does not modify recipes/overlays/h100-gke-cos-inference.yaml on main, and no stock recipe changes. That keeps the demonstration honest about what has and has not been committed to.

Scope

On a live GKE cluster with H100 nodes:

  1. aicr recipe for service: gke, accelerator: h100, os: cos, intent: inference, with k8s-aibom declared in a custom overlay.
  2. aicr bundle renders through every supported deployer; confirm the rendered output is what actually deploys.
  3. Deploy the bundle and confirm the component installs cleanly against the managed control plane, including CRD application (see below).
  4. aicr validate passes, including the k8s-aibom health check — the Deployment rollout terms, the AIBOMControllerConfig Ready condition, and the CRD storage-version assertions added in feat(recipes): assert k8s-aibom CRD storage version in the health check #2305.
  5. Label a namespace aibom.k8saibom.dev/enabled=true, run a real GPU inference workload, and confirm a correct CycloneDX 1.6 ML-BOM is produced for it. Inline vs. truncated behavior at the 256 KiB threshold should be observed rather than assumed.
  6. Measure controller CPU/memory and API-server impact on GKE, and compare against the Kind-measured envelope in ADR-019. A material divergence is a finding, not a footnote.

Depends on

  • k8s-aibom: requalify and re-pin against the v1beta1 graduation release #2282 (requalify and re-pin to v1.3.0) — run against the qualified pin, not a stale one.
  • The CRD-apply step: the registry now pins chart 1.3.0, whose CRDs differ from 1.2.0. On Helm, Helmfile, and Flux the CRDs must be applied before the bundle upgrade (see the component catalog). Whether that step is needed here depends on whether the cluster is fresh; a fresh install applies crds/ automatically.

Out of scope

  • Merging k8s-aibom into the stock h100-gke-cos-inference overlay. That is the epic's later step and needs the remaining Follow-Up requirements.
  • Upgrade, rollback, and uninstall evidence — tracked separately; it gates the overlay merge, not this.
  • Any AIBOM-as-AICR-evidence integration (ADR-007). Still a separate decision.

Done when

  • Every step above is executed on a real GKE cluster and the commands and outputs are recorded on this issue.
  • A real ML-BOM produced from a real inference workload is attached or quoted.
  • Measured controller and API-server cost is recorded and compared to the ADR-019 envelope.
  • Any divergence from the Kind-based expectations is written down, including "none observed" if that is the result.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions