mirror of
https://github.com/actions-runner-controller/actions-runner-controller.git
synced 2026-09-30 19:52:52 +02:00
Status.FinishedRunnerCleanupPatchID was only ever set, never cleared, so a marker recorded under one listener incarnation outlived the patch-ID sequence it described. Applying a new actionable revision deletes the idle and pending runners so they are rebuilt from the new spec, and it restarts the listener. A restarted listener numbers its patches from 0 upwards and counts through every integer, so it does not merely risk reusing the leftover value, it passes through it. If that reuse lands on the reconcile that has to refill the pool, the guard suppresses exactly the scale up the revision change asked for. Clearing the marker where the applied revision advances is enough, because the marker only ever means "the gap below Spec.Replicas was made by cleaning up finished runners for this patch ID", and a spec change invalidates that claim outright. That read now bypasses the cache: the patch is a diff against the object that was read, so a cached copy showing 0 while the server held a marker would emit no entry for the field and leave the stale value behind. Clearing the marker also breaks the invariant the cached fast path in scaleUpServicedByFinishedRunnerCleanup rested on. That path was safe only because a marker was never removed, so a cache hit could not be a false positive. Now it can be: a lagging cache can show a marker the server has already cleared, which suppresses the rebuild. The decision is therefore always made against an uncached read. That costs one GET, and only on reconciles that were about to issue creates anyway. One window remains and is documented rather than papered over. A listener that restarts without a spec change keeps the marker and still renumbers from 0, so a collision is still likely. It costs one suppressed reconcile, not an outage: the listener calls back into scaling on every long-poll timeout rather than only on change, and an idle set at its minimum publishes the collapsed patch ID 0, which is never suppressed. This was found while investigating the update-gha-runner-scale-set e2e failure on the tip of the stack. It is a real defect, but it does not explain that failure, whose cause remains open: the suppression here is self-correcting within a long-poll cycle rather than terminal. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>