mirror of
https://github.com/actions-runner-controller/actions-runner-controller.git
synced 2026-09-30 22:39:01 +02:00
The scale-up suppression read Status.FinishedRunnerCleanupPatchID off the EphemeralRunnerSet that Reconcile fetched through the manager's cached client. That marker is written by the cleanup reconcile, and it is the runner deletions performed by that same reconcile which trigger the next one, so the follow-up reconcile is regularly served from an informer cache that has not yet observed the controller's own status write. The marker read as 0, the guard did not fire, and the controller created a replacement runner for a job that had already finished. That is the exact spurious scale up the guard exists to prevent, so it is a correctness problem in a cluster and not only a flaky test. Make the decision from authoritative state. A cached hit still short circuits, because nothing ever clears the marker, so a hit cannot be a false positive. Only a miss falls through to an uncached Get via a new APIReader, which SetupWithManager fills in from mgr.GetAPIReader() so no construction site can forget it. The extra read is confined to scale-up decisions, where the controller is about to issue creates anyway. The window was transient: updateStatus copies the marker from the in-memory object and patches with MergeFrom, so a stale reconcile produces no diff for that field and cannot clobber the recorded value. The envtest spec that caught this only fails under CI load, so the regression is pinned down directly instead. TestScaleUpServicedByFinished RunnerCleanup drives the decision with a lagging cached client and an API reader that already has the write, and asserts the suppression still holds. Pointing the read back at the cached client fails that case and only that case. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>