All notable changes to k8s-aibom are documented here. The format follows
Keep a Changelog; versions follow
Semantic Versioning (see VERSIONING.md).
-
kubectl aibom find— the incident command: filter tracked workloads by--runtime(exact),--model(case-sensitive substring),--image(substring of the container reference),--digest(bare hex orsha256:-prefixed; a prefix matches) and--signed(unsigned | claimed | verified); filters AND together. Zero matches prints0 matchesand exits 0 so scripts can tell "none found" from "plugin broke". Rows whose document is in an external sink or truncated are shown — never excluded — under image/digest filters, with a stderr warning. Same client and RBAC assummary; no new CRD, no controller flag. -
Infinity embeddings server runtime pattern (#111).
michaelf34/infinity(tag and digest forms, including0.0.77-cpuand0.0.77-rocm) attributes as runtimeinfinity. The match is publisher-anchored at the image-name boundary, somichaelf34/infinity-extraand other registries stay unmatched. -
Metrics are now scrapable, opt-in (#106). The controller's Prometheus endpoint was registered but bound to loopback with no Service — unreachable by any scraper, which made the chart's "metrics" wording an overclaim and left the shipped
grafana/podmonitoring.yamltargeting a pod port that did not exist.metrics.enabled=truenow binds the endpoint onmetrics.port(8080), names the container portmetrics, and renders a ClusterIP Service;metrics.serviceMonitor.enabled=trueadditionally renders a Prometheus Operator ServiceMonitor. Off by default; nothing changes for existing installs. -
aibom_workload_reconcile_outcomes_total{kind,outcome}— the series that separates "opted-in namespace, nothing recognized" from a broken controller:not_opted_in(namespace selector did not match),unmatched(opted in, no inference signal — conservative detection declined),matched(an AIBOM is produced).
-
Served-schema skew is now detected and reported (#104). If the cluster's
AIBOMControllerConfigCRD is replaced by an older schema than the running controller (a GitOps sync pinned to an older revision, aCreateReplacerollback, a manual CRD re-apply), the API server silently prunes stored fields such asspec.verification— previously leaving signature verification OFF behind a fully green status. The controller now reads the served schema via the OpenAPI v3 endpoint (no new RBAC beyond whatsystem:discoveryalready grants; stated explicitly in the ClusterRole) and compares it against every top-level spec field it was built with. Any missing field yieldsDegraded=Truewith reasonSchemaPredatesController, a Warning Event naming the fields and the remedy, and a log line. The check is generic over the spec struct, so the next additive field cannot reintroduce the failure. Found by downstream adversarial upgrade testing (NVIDIA AICR qualification of v1.5.1). -
An upgrade that sets
config.verificationagainst pre-1.5 CRDs now stops before anything is applied (#105). Helm skipscrds/on upgrade. On Helm 4 the config CR's server-side apply then failed with.spec.verification: field not declared in schemaonly after the Deployment had rolled, leaving the releasefailedand half-applied; on Helm 3 the upgrade succeeded and the API server prunedspec.verification, leaving verification off. Whenconfig.verificationis set, the chart now reads the installedAIBOMControllerConfigCRD withlookupand, if itsv1beta1schema lacksspec.verification, fails at render time with the CRD-apply command. Default values never perform the lookup, andhelm templateand client-side--dry-runare unaffected. Withconfig.verificationset, the Helm identity now needsgeton that CRD; without it the render fails onlookupbefore anything is applied. Found by the same downstream qualification. -
Tenant-controlled document growth is bounded. Container component name/version are truncated on the same rule as every other authored string, and a per-document component cap (256) bounds pathological specs in memory and on the wire, not only in etcd. Truncation is recorded as
aibom.truncation.applied— mirroring the redaction rule, never silent. Untruncated documents are byte-identical to before. -
Transient sink failures now retry until the archive heals (#91). Previously the BOM input hash was persisted even when a configured external sink failed, so the next reconcile took the dedup fast path and returned before re-emitting — a transient 403 or network blip dropped that BOM from the archive until the workload spec changed. Now: the input hash is not persisted while any sink is failing (dedup unaffected on success), the reconcile requeues on a bounded cadence (1 minute) until delivery succeeds, and the four status/comment texts that promised a retry that never happened now describe the real behavior. After a partial failure, sinks that already succeeded are re-emitted on the retry; the default timestamped GCS path template makes that a duplicate archive object, never an overwrite.
-
The bootstrap-race deferral now requeues explicitly. It previously waited for a status-update watch event that the Owns-watch's GenerationChangedPredicate filters out — an AIBOM whose first status write raced the cache could sit unpopulated until an unrelated event arrived, and the empty result could collapse a concurrently scheduled sink retry.
- Pod attribution is now ownership-based, not selector-based (Deployment, StatefulSet, DaemonSet, Job). Previously, two same-namespace workloads with overlapping selectors and a shared container name could cross-contaminate image digests in each other's BOMs, and a principal with pod-create permission in an opted-in namespace could plant a chosen digest in another workload's record. Digests now enter a BOM only from pods tied to the workload through the controller ownerReference chain (Deployment → ReplicaSet → Pod walked explicitly), with a belt-and-braces image-name match on the candidate's container status. Pods with no controller owner never contribute. Found by internal review; regression-tested end to end.
- Webhook sinks reject credentials over cleartext at config load.
An
http://endpoint combined with anyauthconfiguration is now a load-time validation error (all-or-nothing fallback, named LoadError). Plain http without auth remains legal for in-cluster receivers; https with auth is unchanged.
The trust release: the verified confidence tier the README has
promised since v1.0 — cryptographic verification of model signature
claims against configurable Sigstore trust roots with Rekor
transparency-log inclusion — designed in public (Design 002),
substantively amended twice by external review from the model-signing
community, and hardened so a claim can never upgrade itself. Also:
the output sanitization guarantee, the kubectl-aibom plugin, the
non-default-configuration e2e matrix, and the NIM model env vars.
With no signature annotations present, output is byte-identical to
v1.4.0.
-
Chart:
extraVolumes/extraVolumeMountsvalues — required to mount a static trusted-root file forverification.trustRootMode=staticBundle(air-gapped clusters). -
e2e: non-default-configuration matrix (#59) — strict-readiness break/recover, webhook sink with a bearer-token Secret under real RBAC, and signature verification with verified and tampered outcomes against a static trust root. Closes the test-debt class behind the one code defect external qualification found.
-
NIM model-declaration env vars
NIM_MODEL_NAMEandNIM_SERVED_MODEL_NAMEjoin the default model-identity allowlist. Reported by an AICR maintainer during Design 003 review: a NIM container's served model can differ from its image default, and without these names such workloads carried no declared model signal. -
kubectl-aibom:
summarygains a SIGNED column (per-model signature states, deduplicated) andverifyappends the recorded per-model signature facts to its integrity verdict — the verified tier is demonstrable in one command. -
Sigstore/Rekor signature verification — the
verifiedconfidence tier (Design 002, #56):spec.verificationonAIBOMControllerConfigenables cryptographic verification of model signature claims (model.k8saibom.dev/oms-signature, optionalmodel.k8saibom.dev/digest) against configurable trust roots (Sigstore public-good via TUF, self-hosted TUF mirror, or a static trusted-root file).verifiedrequires chain validity, Rekor inclusion, a satisfied signer-identity constraint (a non-public trust root counts), and no contradicted declared binding — digest-over-name precedence. All outcomes are recorded facts on the model component (signature.*properties;ModelSummary.Signednow populates); verification failures never fail a reconcile, and with verification absent, output is byte-identical to v1.4.0. Retires schema-divergences entry D-001. Fixed alongside: the reserved signature annotations are no longer mis-extracted as phantom model identities. -
kubectl-aibomplugin (#58):summary(per-namespace or-Atable of workload, category, runtime, models, confidence, Ready),view(decoded, pretty-printed BOM;--rawfor the canonical bytes the published digest covers), andverify(recomputes sha256 againststatus.bomDocument.sha256, non-zero exit on mismatch — script- and CI-composable). Built viamake kubectl-pluginorgo install .../cmd/kubectl-aibom@latest; verified live against a v1.4.0 install. -
Output sanitization guarantee (#57): every string emitted into a BOM passes a redaction filter at the build boundary — URI userinfo, known credential query parameters (pre-signed URL signatures, SAS tokens), and well-known secret token shapes are replaced before emission, with an
aibom.redaction.appliedproperty recording the redaction class on any affected component or service. The audit behind it (issue #57) confirmed no default extraction path emits credential material; the filter guarantees the residual vectors (URI-shaped identity fields such as KServestorageUriand model annotations, and operator-extended allowlists). Clean documents are byte-identical to v1.4.0 output.
The downstream-coverage release, cut the day after k8s-aibom began shipping in NVIDIA AICR v0.20.0: detection patterns NVIDIA's catalog needs, the chart CR template graduation deferred out of v1.3.0, and the re-baselined performance record — bundled so downstream distributions requalify once.
- Runtime image patterns for NVIDIA NIM (
nvcr.io/nim/*→nim) and NVIDIA Dynamo backend workers (nvcr.io/nvidia/ai-dynamo/{vllm,sglang, tensorrtllm}-runtime→vllm/sglang/tensorrt-llm, nightly variants included). Dynamo infrastructure images (frontend, planner, operator) deliberately do not match. - TGI's GHCR namespace (
ghcr.io/huggingface/text-generation-inference→tgi) — previously a documented deferred false negative; real deployment signal arrived.
- Performance documentation re-baselined on live-GKE measurements of v1.2.0 and v1.3.0 (1,001 workloads, dual-sampled): steady state is 1–2m CPU / ~61Mi, statistically identical across both versions and consistent with NVIDIA/aicr#2310's independent measurement. Every published figure now carries version + environment + sampling method; v1.1.0-era Kind steady-state figures are superseded, with the ~370m convergence burst retained as the upper bound.
- The chart now renders the default
AIBOMControllerConfigataibom.k8saibom.dev/v1beta1, matching the CRD storage version (#49). No behavioral change: the schema is identical under dual serving, andv1alpha1manifests remain valid through 1.x. Tools asserting on the rendered CR'sapiVersionshould follow the guidance in docs/migration-v1beta1.md (assert the CRD storage version, not blanket apiVersion replacement).
The graduation release: the aibom.k8saibom.dev APIs reach v1beta1
(Design 001), satisfying the non-alpha storage requirement for stock
AICR adoption (NVIDIA/aicr ADR-019). Upgrading requires applying the
new CRDs — see docs/migration-v1beta1.md; skipping the step stalls
the rollout loudly and safely, with the previous pod still serving.
- Readiness gating had a start-ordering race (since v1.1.0): the
cache-sync check called
WaitForCacheSyncbefore the manager started, trivially passing against an empty informer set — a pod could report Ready before (or without) its informers syncing. A first fix (a manager Runnable) was defeated by the same class of race: controllers create their informers after plain runnables start. Readiness is now asserted per probe against the load-bearing informers themselves — each readyz evaluation asks the cache for the v1beta1AIBOMandAIBOMControllerConfiginformers and their sync state, failing while the API server cannot serve those versions (e.g. stranded CRDs). With this fix, upgrading to the graduation release without the required CRD apply stalls the rollout with the previous pod still serving. Both defeated implementations were caught by the v1.3.0 release-candidates' stranded-CRD boundary tests on a real cluster. - Chart CRD files (and generated manifests) keep their YAML document
separators: without them,
helm show crdsconcatenates the two CRDs into one invalid stream andkubectl applysilently applies only the first — breaking the documented CRD-upgrade command exactly when it matters. Found by the rc.3 boundary test's recovery step.
v1beta1API forAIBOMandAIBOMControllerConfig— schema-identical tov1alpha1(conversion strategy remainsNone), served alongsidev1alpha1, and the storage version from this release onward.v1alpha1remains served and field-frozen through 1.x; its removal will be a separate, announced release with a documented migration step. Design: docs/design/001-api-graduation-v1beta1.md. The controller operates on thev1beta1types internally; both versions remain registered and served. An integration test proves the dual-serving round-trip (write v1alpha1 → read v1beta1 and vice versa, same UID, identical fields).
- Opt-in strict configuration readiness:
--strict-config-readiness(chart valuereadiness.strictConfig) fails the readiness probe while the activeAIBOMControllerConfigis invalid. Off by default — the controller deliberately stays Ready on last-known-good config so an operator typo cannot take down inventory; distributions requiring configuration-aware readiness (e.g. AICR) enable it via values. An absent CR (defaults-by-choice) is not treated as invalid.
The qualification release: every blocking finding from NVIDIA/AICR's ADR-019 Phase 1 qualification of v1.0.0 (gates 3 and 4), fixed with tests, plus readiness hardening. Details in the sections below and the qualification record on issue #8.
- Chart
config.*values render verbatim into the defaultAIBOMControllerConfig:discovery(incl. namespace selector),bomGeneration,sinks, andloggingare now normal public values — no template patching or post-install CR mutation needed. - Chart default resources are set from measured footprint (requests 50m/128Mi, memory limit 256Mi; no CPU limit by design — see docs/quality-baseline.md).
- Readiness now gates on informer cache sync:
/readyzfails until the controller can observe the cluster, and the chart wires liveness and readiness probes against the health endpoints (previously no probe consulted them).
-
Scrape and BOM-build failures now flip
Ready=False(with reason and message) on the workload's existing AIBOM, so failures are observable in status rather than only in logs; prior document/summary fields are preserved and the failure path never creates AIBOMs. -
Non-conflict status-persistence failures now emit the
aibom_status_persist_failures_totalmetric and anAIBOMStatusPersistFailedwarning Event. -
Truncation reason now distinguishes "no external sink is configured" from "configured sinks all failed this cycle" — the latter previously reported the former's message.
-
Every reconcile now runs under a finite 60s deadline, bounding all Kubernetes API operations (previously unbounded; a stalled API request could consume the reconcile forever). The 30s per-sink deadline nests inside it.
-
GCS writes are capped at 4 attempts (matching the webhook sink's bounded attempt count) within the existing 30s elapsed bound.
-
docs/webhook-sink-protocol.md backoff schedule corrected to match the code (250ms/1s/3s, 4 total attempts).
-
Sink credential Secrets are now read via a direct (uncached) API reader. Previously the first Secret read started a cluster-wide Secret informer, which the namespace-scoped Role correctly forbids — with sinks configured and
rbac.sinkSecretAccess=true, config reload stalled (Ready=Truestale at the prior observedGeneration) with repeatedsecrets is forbiddenlist errors. Found by AICR gate-3 qualification. The Role also narrows togetonly.
First tagged release. Every release publishes a coherent, verifiable artifact set: a multi-arch image (linux/amd64, linux/arm64) on ghcr.io/googlecloudplatform/k8s-aibom carrying Sigstore build-provenance and CycloneDX SBOM attestations; a digest-pinned Helm chart on oci://ghcr.io/googlecloudplatform/charts carrying a build-provenance attestation (the chart has no separate SBOM attestation); and a digest-pinned install.yaml release asset.
- AIBOMControllerConfig is now a regular Helm release resource instead
of a
pre-installhook, so Helm owns install/upgrade/rollback/uninstall deterministically, and the invalidnamespaceon the cluster-scoped object is gone. Migration for existing from-source installs: the hook-created CR carries no Helm ownership metadata; before upgrading an existing release, delete it (kubectl delete aibomcontrollerconfig default) or annotate it for Helm adoption. - Secret access is now opt-in. The namespace-scoped Role granting
Secret reads (used only for sink credentials) is rendered only when
rbac.sinkSecretAccess=true. With no sinks configured (the default), the controller holds no Secret permissions. Set the value if your AIBOMControllerConfig references credential Secrets. - ClusterRole rules deduplicated;
aibomcontrollerconfigsnarrowed to read-only + status (matching the controller's kubebuilder markers).
image.digestchart value: a digest takes precedence over the tag so releases can be pinned immutably without patching the chart.- Controller version is stamped at build time via ldflags
(
main.controllerVersion); local builds reportdev. The image carries OCI identity labels. - Community health files: CODEOWNERS, issue forms, PR template.
VERSIONING.md, this changelog,docs/compatibility.md,docs/release-checklist.md.- Release pipeline: tag-triggered publishing of the signed multi-arch image (with provenance and CycloneDX SBOM attestations), the OCI chart (with a provenance attestation), and digest-pinned install.yaml; a dry-run job exercises the release path on every PR.
- Dependency updates cleared all critical/high Dependabot alerts (golang.org/x/net, golang.org/x/crypto, google.golang.org/grpc).