k8s-aibom · Apache-2.0 · v1.5.1

Your cluster is already running models nobody registered.

Shadow AI — a pulled image and a --model flag is all it takes to serve a model.

Compliance pull — AI inventory obligations assume you can list what is actually running.

github.com/GoogleCloudPlatform/k8s-aibom
Why runtime

Build-time BOMs record intent. This records what is actually serving.

Build time

Intent

Produced in CI, from the artifact you meant to ship. Frozen at the moment of build.

Blind to a hot-swapped HF_MODEL_ID, an edited Deployment, or a digest that was actually pulled.

Runtime · k8s-aibom

Reality

Generated from live Kubernetes API objects: the args, env, image and digest serving right now.

Every attribute carries evidence pointing at the exact spec field it came from.

github.com/GoogleCloudPlatform/k8s-aibom
What it is

A controller that emits one CycloneDX 1.6 ML-BOM per AI workload.

Controller One unprivileged Deployment in k8s-aibom-system Reads the Kubernetes API. Nothing else.
Opted-in namespaces aibom.k8saibom.dev/enabled=true Inference, agents, RAG, training, eval — vLLM, TGI, Triton, NIM, Dynamo, Ollama, Ray Serve, SGLang, KServe
AIBOM CR per workload ML-BOM stored inline, published sha256 Byte-deterministic → GitOps-diffable
GCSoptional sink
Webhookoptional sink
No DaemonSetNo privileged containerNo kernel / eBPFNo sidecarsNo pod-spec changesNamespace opt-in
github.com/GoogleCloudPlatform/k8s-aibom
Confidence model

Every claim carries checkable evidence.

declared
The customer wrote it into the spec.
Example: a --model container arg, or an HF_MODEL_ID env var.
inferred
Derived by heuristic.
Example: runtime vllm, from matching the image against a pattern.
unresolved
Could not be determined — and the document says so.
Example: the image digest of a pod that hasn't pulled yet.

Each attribute ships with an evidence locator pointing at the exact spec field it came from.

github.com/GoogleCloudPlatform/k8s-aibom
New in v1.5.0

The verified tier — Sigstore-backed, and deliberately calibrated.

Model signature claims are checked against operator-configured trust roots (Sigstore public-good via TUF, self-hosted TUF mirror, or a static trusted-root file for air-gapped clusters) plus Rekor transparency-log inclusion.

Calibration A

verified requires a signer-identity constraint the workload author does not control. A claim can never upgrade itself. Under a public trust root with no identity constraints the verifier reports signature-valid-unconstrained and the tier stays claimed.

Calibration B

verified means a signer the operator trusts signed a statement consistent with this claim. It does not mean the bytes on the accelerator were hashed — the controller has no node access, by design.

Designed in an open review window; substantively amended twice by community review from the model-signing ecosystem.

github.com/GoogleCloudPlatform/k8s-aibom
Demo · kubectl-aibom plugin

See it from the command line.

live output — v1.5.0 release verification, GKE, 2026-09-22
$ kubectl aibom summary -n rc-verify
NAMESPACE  NAME                     WORKLOAD    WORKLOAD-NAME  CATEGORY   RUNTIME  MODELS                  CONFIDENCE  SIGNED    READY
rc-verify  apps-deployment-goodsig  Deployment  goodsig        inference  vllm     pkg:npm/sigstore@1.3.0  inferred    verified  True

$ kubectl aibom verify apps-deployment-goodsig -n rc-verify
published: e15230d35f8561617f94012d8782ca3cd10591d13968e0adf24e5eb555566a99
computed:  e15230d35f8561617f94012d8782ca3cd10591d13968e0adf24e5eb555566a99
OK: document bytes match the published digest
model pkg:npm/sigstore@1.3.0: signature verified
kubectl aibom summaryTable of tracked workloads, incl. per-model SIGNED state.
kubectl aibom view <name>Decoded ML-BOM for one workload.
kubectl aibom verify <name>Recompute sha256 vs the published digest; non-zero exit on mismatch.
github.com/GoogleCloudPlatform/k8s-aibom
Cost

Steady state at 1,001 tracked workloads.

1���2mCPU
~63MiB

Measured on live GKE, dual-sampled (metrics-server and kubelet counter deltas), and independently reproduced by NVIDIA on their own cluster (NVIDIA/aicr issue #2310).

github.com/GoogleCloudPlatform/k8s-aibom
Adoption

Qualified into NVIDIA AI Cluster Runtime — four gates, in public, one thread.

Ships in NVIDIA AI Cluster Runtime as a qualified component since AICR v0.20. Enabled by default in exactly one stock recipe — the GKE H100 inference recipe — and available opt-in on every AICR platform.

GATE 1Supply chain
GATE 2Code review
GATE 3Operational safety
GATE 41,000-workload regression
github.com/GoogleCloudPlatform/k8s-aibom
Where the BOM goes

The controller emits facts. Judgments belong in consumers.

AIBOM CR Inline ML-BOM + published sha256 Optional sinks: GCS, webhook
Relay Documented webhook relay recipe
Consumers decide
  • OWASP Dependency-Track
  • OpenSSF GUAC path
  • Kyverno / Gatekeeper policy cookbook

Outputs are designed as evidence for EU AI Act Articles 12/50 logging & transparency, NIST AI RMF inventory controls, and ISO/IEC 42001 inventory clauses.

github.com/GoogleCloudPlatform/k8s-aibom
Get started

Three commands to inventory.

helm install k8s-aibom oci://ghcr.io/googlecloudplatform/charts/k8s-aibom --version 1.5.1 --namespace k8s-aibom-system --create-namespace
kubectl label namespace <ns> aibom.k8saibom.dev/enabled=true
kubectl krew install aibom

Chart is digest-pinned; releases ship provenance and SBOM attestations.

Roadmap

v1.6 train: CronJob coverage, workload-kind allowlist, NIM-operator/LeaderWorkerSet scrapers under design review.

Designed and reviewed in public — contributions welcome.

github.com/GoogleCloudPlatform/k8s-aibom
← → to navigate · add ?notes for speaker notes