Evidence for the three items NVIDIA/aicr#2962 lists as not yet
measured after the v1.3.0 → v1.5.1 re-pin (aicr#2976): the API-server
side of the controller's load, a measured rollback, and the shape of
the upgrade path. Run 2026-09-29 on a live regional GKE cluster
(v1.35.8-gke.1036000, 9 nodes), default chart values throughout —
the configuration AICR ships — against 1,002 tracked workloads
(1,001 generated replicas: 0 vLLM Deployments plus one pre-existing).
Attribution is client-side. Every request number below comes from
the controller process's own rest_client_requests_total and
controller_runtime_reconcile_total counters, read from its loopback
metrics endpoint via kubectl port-forward (v1.5.1 binds :8080 on
loopback; v1.6 makes this scrapable with metrics.enabled). They are
monotonic within one process and attributable by construction. This
is the same attribution NVIDIA used for the README's earlier "~2
requests per inventoried workload" figure.
Server-side counter deltas are invalid on regional GKE. kubectl get --raw /metrics is load-balanced across the control plane's
replicas, each with independent lifetime counters. Two snapshots 90
seconds apart landed on different instances (process_start_time_seconds
differed) and produced a nonsensical delta (67,000 pod creates that
never happened). Server-side counters are therefore used only as a
same-instance sanity cross-check, never as the measurement. Anyone
reproducing this on a regional cluster should tag each scrape with
process_start_time_seconds or attribute client-side.
One incidental finding. An ephemeral debug container could not be
attached to the controller pod: its runAsNonRoot security context
rejected an image with a non-numeric user. The hardening applies even
to operator-initiated sidecars.
| UTC | Step | Result |
|---|---|---|
| 17:18 | Baseline snapshot, no controller | 9-node regional GKE; pre-existing pod watchers from system components |
| 17:23 | helm install --version 1.3.0 (default values) |
1.3.0 CRDs installed; image sha256:f8e48d4e… (AICR's prior pin) |
| 17:24 | Apply 1,001 replicas: 0 vLLM Deployments |
1,002 AIBOMs, all Ready, 2 s after the apply finished |
| 17:25 | helm upgrade --version 1.5.1 (default values; Helm leaves CRDs at 1.3.0) |
23 s; image sha256:7b027315…; 1,002 AIBOMs preserved; input hashes unchanged; config Ready=True, Degraded=False |
| 17:26–17:41 | Steady-state window, 935 s | see below |
| 17:43 | helm rollback → rev 1 (v1.3.0), CRDs still 1.3.0 |
17 s; 1,002 AIBOMs preserved; no Warning events |
| 17:46 | Apply v1.5.1 CRDs over the running v1.3.0 controller | verification now served; v1.3.0 controller unaffected, Ready=True |
| 17:47 | helm upgrade --version 1.5.1 again (rev 4) |
24 s; green |
| 17:48 | helm rollback → v1.3.0 (rev 5) with the v1.5.1 CRDs in place |
19 s; 1,002/1,002 AIBOMs with input hash AND BOM hash unchanged; no Warning events; config green |
| 17:52 | Uninstall, remove load, delete CRDs | Cluster restored to as-found |
Controller v1.5.1, default values, 1,002 tracked workloads, after the post-restart walk had completed. Attributed client-side over 935 s:
| Counter | Delta | Rate |
|---|---|---|
rest_client_requests_total GET 200 |
15 | 0.96 requests/minute (0.016/s) |
| any write verb (POST/PUT/PATCH/DELETE) | 0 | 0 |
controller_runtime_reconcile_total (all controllers) |
0 | 0 |
process_cpu_seconds_total |
1.00 cpu-s | 1.07 mCPU average |
| memory | — | 48 Mi (kubectl top) / 93 Mi RSS |
The 15 GETs are informer watch re-establishments — nine informers (Deployment, StatefulSet, DaemonSet, Job, Pod, Namespace, ReplicaSet, AIBOM, AIBOMControllerConfig) whose watches time out and reconnect on controller-runtime's default 5–10 minute jitter. At steady state the controller makes no writes and runs no reconciles; the API server sees roughly one long-poll reconnect per minute for the whole inventory, independent of workload count.
Restart cost (upgrade or rollback), for completeness: each version
change re-walks the inventory once: 11 GETs (one list per informer)
and 1,003 status PUTs (one per AIBOM — the dedup fast path refreshing
LastReconciled, plus the config CR), ~21 cpu-seconds, identical in
shape on v1.3.0 and v1.5.1. This matches the README's published "~2
requests per inventoried workload" for a cold start. v1.5.1 adds one
informer over v1.3.0 (ReplicaSet, for the ownership walk that fixed
#97); its steady-state cost is one reconnect per 5–10 minutes.
Two variants, both from the running v1.5.1 to v1.3.0 via helm rollback, both with 1,002 tracked workloads:
- CRDs never upgraded (AICR's default-values path). 17 s to Ready on the v1.3.0 image. Every AIBOM preserved. No Warning events. The v1.3.0 controller re-walked the inventory with the same 1-PUT-per-AIBOM signature as an upgrade.
- CRDs upgraded to v1.5.1 first (the path the documented manual
CRD step produces), then rollback. 19 s to Ready. The v1.3.0
controller runs correctly under the newer, additive-only schema —
it neither reads nor is disturbed by the
verificationfield. 1,002/1,002 AIBOMs kept bothinputHashandbomHashunchanged across the full v1.3.0 → v1.5.1 → v1.3.0 round trip: no re-emits, no content drift. Config CRReady=True,Degraded=False.
Helm never touches CRDs on upgrade or rollback (both variants confirm the served schema was unchanged by the release operations), so the CRD state is set only by the explicit manual step — which is why variant 2 is the one to qualify against going forward.
From v1.6, the controller detects served-schema skew (#104, PR #107):
the variant-1 upgrade path (v1.5.1+ controller, 1.3.0 CRDs) will
report Degraded=True with reason SchemaPredatesController until
the CRDs are applied, while Ready stays True. That is the intended
signal — it makes "apply the CRDs" visible instead of silent — and a
qualification gate that asserts Degraded=False will correctly
require the CRD step. Nothing about the rollback direction changes.
The snapshot and delta helpers, the 1,001-workload manifest generator,
and the raw counter files from this run are retained by the
maintainer; the commands above are the complete sequence. Attribution
requires reading the controller's metrics endpoint: on v1.5.1 use
kubectl port-forward <pod> 18080:8080 (loopback bind); from v1.6,
metrics.enabled=true exposes it directly.