Skip to content

Latest commit

 

History

History
121 lines (101 loc) · 6.7 KB

File metadata and controls

121 lines (101 loc) · 6.7 KB

v1.5.1 on GKE: upgrade, rollback, and API-server cost — measured

Evidence for the three items NVIDIA/aicr#2962 lists as not yet measured after the v1.3.0 → v1.5.1 re-pin (aicr#2976): the API-server side of the controller's load, a measured rollback, and the shape of the upgrade path. Run 2026-09-29 on a live regional GKE cluster (v1.35.8-gke.1036000, 9 nodes), default chart values throughout — the configuration AICR ships — against 1,002 tracked workloads (1,001 generated replicas: 0 vLLM Deployments plus one pre-existing).

Methodology, and one thing that does not work

Attribution is client-side. Every request number below comes from the controller process's own rest_client_requests_total and controller_runtime_reconcile_total counters, read from its loopback metrics endpoint via kubectl port-forward (v1.5.1 binds :8080 on loopback; v1.6 makes this scrapable with metrics.enabled). They are monotonic within one process and attributable by construction. This is the same attribution NVIDIA used for the README's earlier "~2 requests per inventoried workload" figure.

Server-side counter deltas are invalid on regional GKE. kubectl get --raw /metrics is load-balanced across the control plane's replicas, each with independent lifetime counters. Two snapshots 90 seconds apart landed on different instances (process_start_time_seconds differed) and produced a nonsensical delta (67,000 pod creates that never happened). Server-side counters are therefore used only as a same-instance sanity cross-check, never as the measurement. Anyone reproducing this on a regional cluster should tag each scrape with process_start_time_seconds or attribute client-side.

One incidental finding. An ephemeral debug container could not be attached to the controller pod: its runAsNonRoot security context rejected an image with a non-numeric user. The hardening applies even to operator-initiated sidecars.

Timeline

UTC Step Result
17:18 Baseline snapshot, no controller 9-node regional GKE; pre-existing pod watchers from system components
17:23 helm install --version 1.3.0 (default values) 1.3.0 CRDs installed; image sha256:f8e48d4e… (AICR's prior pin)
17:24 Apply 1,001 replicas: 0 vLLM Deployments 1,002 AIBOMs, all Ready, 2 s after the apply finished
17:25 helm upgrade --version 1.5.1 (default values; Helm leaves CRDs at 1.3.0) 23 s; image sha256:7b027315…; 1,002 AIBOMs preserved; input hashes unchanged; config Ready=True, Degraded=False
17:26–17:41 Steady-state window, 935 s see below
17:43 helm rollback → rev 1 (v1.3.0), CRDs still 1.3.0 17 s; 1,002 AIBOMs preserved; no Warning events
17:46 Apply v1.5.1 CRDs over the running v1.3.0 controller verification now served; v1.3.0 controller unaffected, Ready=True
17:47 helm upgrade --version 1.5.1 again (rev 4) 24 s; green
17:48 helm rollback → v1.3.0 (rev 5) with the v1.5.1 CRDs in place 19 s; 1,002/1,002 AIBOMs with input hash AND BOM hash unchanged; no Warning events; config green
17:52 Uninstall, remove load, delete CRDs Cluster restored to as-found

API-server cost at steady state (the unmeasured item)

Controller v1.5.1, default values, 1,002 tracked workloads, after the post-restart walk had completed. Attributed client-side over 935 s:

Counter Delta Rate
rest_client_requests_total GET 200 15 0.96 requests/minute (0.016/s)
any write verb (POST/PUT/PATCH/DELETE) 0 0
controller_runtime_reconcile_total (all controllers) 0 0
process_cpu_seconds_total 1.00 cpu-s 1.07 mCPU average
memory — 48 Mi (kubectl top) / 93 Mi RSS

The 15 GETs are informer watch re-establishments — nine informers (Deployment, StatefulSet, DaemonSet, Job, Pod, Namespace, ReplicaSet, AIBOM, AIBOMControllerConfig) whose watches time out and reconnect on controller-runtime's default 5–10 minute jitter. At steady state the controller makes no writes and runs no reconciles; the API server sees roughly one long-poll reconnect per minute for the whole inventory, independent of workload count.

Restart cost (upgrade or rollback), for completeness: each version change re-walks the inventory once: 11 GETs (one list per informer) and 1,003 status PUTs (one per AIBOM — the dedup fast path refreshing LastReconciled, plus the config CR), ~21 cpu-seconds, identical in shape on v1.3.0 and v1.5.1. This matches the README's published "~2 requests per inventoried workload" for a cold start. v1.5.1 adds one informer over v1.3.0 (ReplicaSet, for the ownership walk that fixed #97); its steady-state cost is one reconnect per 5–10 minutes.

Rollback (the unmeasured direction)

Two variants, both from the running v1.5.1 to v1.3.0 via helm rollback, both with 1,002 tracked workloads:

  1. CRDs never upgraded (AICR's default-values path). 17 s to Ready on the v1.3.0 image. Every AIBOM preserved. No Warning events. The v1.3.0 controller re-walked the inventory with the same 1-PUT-per-AIBOM signature as an upgrade.
  2. CRDs upgraded to v1.5.1 first (the path the documented manual CRD step produces), then rollback. 19 s to Ready. The v1.3.0 controller runs correctly under the newer, additive-only schema — it neither reads nor is disturbed by the verification field. 1,002/1,002 AIBOMs kept both inputHash and bomHash unchanged across the full v1.3.0 → v1.5.1 → v1.3.0 round trip: no re-emits, no content drift. Config CR Ready=True, Degraded=False.

Helm never touches CRDs on upgrade or rollback (both variants confirm the served schema was unchanged by the release operations), so the CRD state is set only by the explicit manual step — which is why variant 2 is the one to qualify against going forward.

A forward-looking note for qualification

From v1.6, the controller detects served-schema skew (#104, PR #107): the variant-1 upgrade path (v1.5.1+ controller, 1.3.0 CRDs) will report Degraded=True with reason SchemaPredatesController until the CRDs are applied, while Ready stays True. That is the intended signal — it makes "apply the CRDs" visible instead of silent — and a qualification gate that asserts Degraded=False will correctly require the CRD step. Nothing about the rollback direction changes.

Reproduction

The snapshot and delta helpers, the 1,001-workload manifest generator, and the raw counter files from this run are retained by the maintainer; the commands above are the complete sequence. Attribution requires reading the controller's metrics endpoint: on v1.5.1 use kubectl port-forward <pod> 18080:8080 (loopback bind); from v1.6, metrics.enabled=true exposes it directly.