fix(recipes): raise ebs-csi node sidecar memory limits for p5 nodes - #1693
Conversation
The aws-ebs-csi-driver chart's default 32Mi limit on the node DaemonSet's livenessProbe and node-driver-registrar sidecars is too tight on high-core instances (e.g. p5.48xlarge, 192 vCPU). The Go binary's resident set alone is ~24Mi; per-P runtime overhead (Go sizes runtime structures by GOMAXPROCS, which defaults to the host core count) pushes the working set to ~31.7Mi during cold start under bringup load, cgroup-OOM-killing the livenessProbe container (observed: CONSTRAINT_MEMCG, RSS 31.7Mi vs 32Mi limit). Killing the livenessProbe removes the /healthz endpoint the ebs-plugin livenessProbe depends on, so the kubelet restarts ebs-plugin -> the pod enters CrashLoopBackOff -> the aws-ebs-csi-driver health check trips -> the expected-resources deployment-validation check fails, failing UAT on p5. Raise both node sidecars to 128Mi, giving headroom above the observed ~31.7Mi peak independent of instance core count. Node-pressure-independent and does not change any image, chart version, or pin (BOM unaffected). Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Enterprise Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughThis change modifies the AWS EBS CSI driver Helm values configuration to add explicit resource limits for two sidecar containers, nodeDriverRegistrar and livenessProbe. Each sidecar receives a 128Mi memory limit along with associated CPU and memory request values, accompanied by comments explaining the rationale related to chart defaults on high-core instances and potential liveness failures from tight limits. Estimated code review effort: 1 (Trivial) | ~3 minutes 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Recipe evidence checkProtected recipesRecipes with committed evidence (
Other affected recipes without evidence yet: 23These recipes are affected by this PR but carry no committed evidence pointer, so there is
How to refresh evidenceRun on a cluster matching the recipe's aicr snapshot -o snapshot.yaml
aicr validate \
-r recipes/overlays/<slug>.yaml \
-s snapshot.yaml \
--emit-attestation ./out \
--push ghcr.io/<your-fork>/aicr-evidence
# Copy to the per-source path printed in the emit 'copyTo' hint:
# recipes/evidence/<slug>/<source>/<bundle-digest>.yamlThis gate is warning-only and never blocks merge. See ADR-007 for the trust model. |
…VIDIA#1693) Signed-off-by: Nathan Hensley <nhensley@nvidia.com>
Summary
Raise the
aws-ebs-csi-drivernode DaemonSet sidecar memory limits (livenessProbe,nodeDriverRegistrar) from the chart-default 32Mi to 128Mi, so they don't OOM on high-core GPU instances (p5.48xlarge, 192 vCPU).Motivation / Context
On p5.48xlarge UAT clusters,
expected-resources(deployment-phase validation) intermittently-to-consistently fails on theaws-ebs-csi-driverhealth check:Root cause (from node dmesg during bringup):
The chart's 32Mi limit on the node sidecars is too tight on high-core nodes. The Go binary's resident set alone is ~24 MiB (
file-rss); per-P runtime overhead (Go sizes runtime structures byGOMAXPROCS, which defaults to the host core count — 192 on p5) pushes the working set to ~31.7 MiB during cold start under bringup load, cgroup-OOM-killing thelivenessProbecontainer. That removes the/healthzendpoint theebs-pluginlivenessProbe depends on → kubelet restartsebs-plugin→ CrashLoopBackOff → theaws-ebs-csi-driverhealth check trips →expected-resourcesfails → UAT fails on p5. The smallerm7i.xlargesystem nodes never hit this (fewer cores → smaller runtime footprint), which is why it is p5-specific.Fix: give both node sidecars headroom above the observed ~31.7 MiB peak (128Mi), independent of instance core count.
Fixes: N/A
Related: N/A
Type of Change
Component(s) Affected
pkg/recipe) — component values (recipes/components/aws-ebs-csi-driver/values.yaml)Implementation Notes
sidecars:block settingnodeDriverRegistrarandlivenessProbelimits to 128Mi (requests unchanged at 40Mi). These sidecar keys apply to both node and controller instances in the chart.GOMAXPROCSpinning was considered but deliberately not used as the fix: forensics show it is a minor contributor (~5 MiB of the working set), and raising the limit is the deterministic, core-count-independent fix. It can be added later as defense-in-depth if desired.Testing
livenessprobeexit 137/OOMKilled at 32Mi;ebs-pluginexit 2 from liveness kill), and the failingaws-ebs-csi-driverchainsaw assert.make qualify) covers recipe/render checks.Risk Assessment
Rollout notes: Takes effect on the next deploy/UAT bringup. Backwards compatible; no migration.
Checklist
make testwith-race) — N/A (no Go changes)make lint) — N/A (values-only)git commit -S)