Skip to content

SCHED-1788: Fix inconsistent behaviour with automatic powering up - #2560

Merged
Uburro merged 1 commit into
soperator-release-3.0from
SCHED-1788/0
May 29, 2026
Merged

Uburro merged 1 commit into
soperator-release-3.0from
SCHED-1788/0

Conversation

@Uburro

@Uburro Uburro commented May 29, 2026

Copy link
Copy Markdown
Collaborator

Problem

Automatic power-up of ephemeral (CLOUD) nodes was inconsistent: identical customer configs behaved differently because topology.conf membership was driven by live pods. Whenever an ephemeral pod appeared/disappeared, topology.conf changed and each reapply let Slurm re-evaluate and sometimes power nodes up — so we could neither guarantee nor disable automatic power-ups. Automatic power-downs were also entirely unsupported (SuspendExcStates=CLOUD excluded every ephemeral node).

Solution

Topology is now built in two stages so topology.conf is stable regardless of pod lifecycle:
Stage 1 lists every Slurm node from each NodeSet's full replica range (including powered-down ephemeral nodes) under the unknown switch/block.
Stage 2 overlays real IB switches only for GPU-enabled NodeSets whose pods are scheduled to a K8s node (not necessarily Running), moving them off unknown.
Removed SuspendExcStates=CLOUD so automatic power-up is no longer blocked for CLOUD nodes (Slurm still won't resume DRAINed nodes; static NodeSets stay protected via SuspendExcNodes).
Added a cluster-global slurmConfig.suspendTime (seconds, default -1 = disabled) to make automatic power-down opt-in. The new default preserves today's "no power-down" behavior.

Testing

go build ./..., go vet, and unit tests pass for internal/controller/topologyconfcontroller, internal/render/common, and internal/webhook/....
Updated/added unit tests for the two-stage builders: powered-down nodes present under unknown, scheduled GPU pods placed on IB switches, CPU/unscheduled/unlabeled nodes staying under unknown.
Regenerated CRDs/de

Release Notes

Feature: Ephemeral node topology is now stable across pod lifecycle, making automatic power-ups consistent. Automatic power-downs are now supported but disabled by default — enable per cluster via slurmConfig.suspendTime (idle seconds; -1 disables).

@Uburro Uburro closed this May 29, 2026
@Uburro Uburro reopened this May 29, 2026
@Uburro
Uburro changed the base branch from main to soperator-release-3.0 May 29, 2026 11:31
@Uburro Uburro added the fix label May 29, 2026
@Uburro
Uburro merged commit 478b96f into soperator-release-3.0 May 29, 2026
26 of 41 checks passed
@Uburro
Uburro deleted the SCHED-1788/0 branch May 29, 2026 13:55

This branch had an error being deployed

1 failed deployment
e2e — faf56cb2 Deployed May 29, 2026 by Uburro via e2e-test #6678
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2 participants