SCHED-1788: Fix inconsistent behaviour with automatic powering up - #2560
Merged
Merged
Conversation
itechdima
approved these changes
May 29, 2026
itechdima
approved these changes
May 29, 2026
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Automatic power-up of ephemeral (CLOUD) nodes was inconsistent: identical customer configs behaved differently because topology.conf membership was driven by live pods. Whenever an ephemeral pod appeared/disappeared, topology.conf changed and each reapply let Slurm re-evaluate and sometimes power nodes up — so we could neither guarantee nor disable automatic power-ups. Automatic power-downs were also entirely unsupported (SuspendExcStates=CLOUD excluded every ephemeral node).
Solution
Topology is now built in two stages so topology.conf is stable regardless of pod lifecycle:
Stage 1 lists every Slurm node from each NodeSet's full replica range (including powered-down ephemeral nodes) under the unknown switch/block.
Stage 2 overlays real IB switches only for GPU-enabled NodeSets whose pods are scheduled to a K8s node (not necessarily Running), moving them off unknown.
Removed SuspendExcStates=CLOUD so automatic power-up is no longer blocked for CLOUD nodes (Slurm still won't resume DRAINed nodes; static NodeSets stay protected via SuspendExcNodes).
Added a cluster-global slurmConfig.suspendTime (seconds, default -1 = disabled) to make automatic power-down opt-in. The new default preserves today's "no power-down" behavior.
Testing
go build ./..., go vet, and unit tests pass for internal/controller/topologyconfcontroller, internal/render/common, and internal/webhook/....
Updated/added unit tests for the two-stage builders: powered-down nodes present under unknown, scheduled GPU pods placed on IB switches, CPU/unscheduled/unlabeled nodes staying under unknown.
Regenerated CRDs/de
Release Notes
Feature: Ephemeral node topology is now stable across pod lifecycle, making automatic power-ups consistent. Automatic power-downs are now supported but disabled by default — enable per cluster via slurmConfig.suspendTime (idle seconds; -1 disables).