Skip to content

Fix readiness start-ordering race; correct stranded-CRD semantics - #44

Merged
glenmessenger merged 1 commit into
mainfrom
fix/readyz-start-race
Aug 20, 2026
Merged

glenmessenger merged 1 commit into
mainfrom
fix/readyz-start-race

Conversation

@glenmessenger

Copy link
Copy Markdown
Collaborator

The v1.3.0-rc.1 boundary test on GKE falsified our loud-failure hypothesis as stated and exposed a real bug shipping since v1.1.0:

Observed: upgrading v1.2.0 → rc.1 without the CRD apply briefly passed readiness, completed the rollout (killing the healthy old pod), and then failed — pod now NotReady/restarting with no matches for kind ... v1beta1 every 10s, reconciliation dead. Loud eventually, but through a false-Ready window that destroyed the working instance first.

Root cause: the readyz cache-sync gate called WaitForCacheSync before mgr.Start() registered informers — an empty informer set syncs trivially, so the gate passed from t=0 and has been decorative since v1.1.0. (The v1.2.0 strictConfig check is unaffected — it evaluates per-probe against the config store and was genuinely verified.)

Fix: readiness derives from a manager RunnableFunc — non-leader-election runnables execute only after the manager's caches actually sync, so reaching the runnable is the condition. Corrected stranded-CRD behavior (to be re-verified on rc.2): the new pod never reports Ready, the rolling update stalls with the old pod still serving (zero inventory downtime), ProgressDeadlineExceeded signals the operator, and the CRD apply completes the rollout.

Docs (migration guide, Design 001) corrected from the hypothesized behavior to the verified one; changelog entry in the 1.3.0 section.

Full suite green. Next: rc.2 and the boundary test rerun.

…docs

The cache-sync readyz check called WaitForCacheSync before mgr.Start,
trivially passing against an empty informer set — a false-Ready window
that let a stranded-CRD upgrade complete its rollout and kill the
healthy old pod before failing. Readiness now derives from a manager
Runnable (executes only after caches genuinely sync): the stranded
upgrade stalls with the previous pod still serving. Docs corrected to
the verified stalled-rollout behavior. Found by the v1.3.0 rc boundary
test on GKE.
@glenmessenger
glenmessenger merged commit 008bf2c into main Aug 20, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant