Files
actions-runner-controller/controllers
Nikola JokicandCopilot App e2768dddbf Switch the scale set off instead of rebuilding it when runners are outdated
When the runners reject the runner spec they were given, the listener has to
stop acquiring jobs the scale set cannot run. The controller removed the
listener but then deleted the EphemeralRunnerSet, and left the
AutoscalingRunnerSet phase on Running. AutoscalingRunnerSetPhaseOutdated was
declared and read, but never assigned by anything.

Nothing held the scale set switched off as a result. The next reconcile saw a
missing EphemeralRunnerSet, created it, created a listener for it, and the
fresh runners rejected the same spec again, so the scale set churned through
create and teardown cycles against the Actions service instead of resting.

Record the outdated phase and keep the EphemeralRunnerSet, pinned to zero
replicas and patch id. It releases every runner that is not executing a job
while the phase keeps the listener from being rebuilt, and the revision
bookkeeping that decides when the scale set may run again is preserved.

Recovery is driven by the spec update that the phase is waiting for: it moves
the phase back to pending, and the runner spec is then republished to the set
with an advanced revision even when the runner spec itself did not change.
The revision is what tells the EphemeralRunnerSet to stop judging itself by
the runners that failed, so without advancing it a scale set could only be
recovered by editing the pod template, and an edit to anything else would
switch the listener back on against a set parked at zero.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 11:43:46 +02:00
..