mirror of
https://github.com/actions-runner-controller/actions-runner-controller.git
synced 2026-09-30 23:44:06 +02:00
The `actions.github.com/integrity-hash` annotation was used as an opaque fingerprint to detect spec drift across AutoscalingRunnerSet, EphemeralRunnerSet and the listener resources. Hashes are brittle: they change whenever unrelated serialization details change, they are invisible to users, and they are not restart-safe. FNV-32a also carries a real collision risk, where the consequence is an update silently never applied. Replace it with explicit, typed state: - `AutoscalingRunnerSetStatus.ObservedGeneration` drives the Pending phase transition via `metadata.generation` instead of an annotation hash. - `EphemeralRunnerSetSpec.ActionableRevision` and `EphemeralRunnerSetStatus.AppliedActionableRevision` form a restart-safe applied marker. The revision is bumped by the AutoscalingRunnerSet controller when `EphemeralRunnerSpec` changes, and only advanced in status after idle/pending runner cleanup succeeds. - `EphemeralRunnerSetStatus.FinishedRunnerCleanupPatchID` records the listener patch ID for which finished runners were reaped, so scale-up is suppressed until the listener publishes a fresh desired state. This prevents creating a replacement runner for a job that already completed. - Listener pod recreation compares pod specs semantically instead of comparing hash annotations. Drift detection uses `apiequality.Semantic`, not `cmp` or `reflect`: - `Semantic.DeepEqual` for the EphemeralRunnerSpec. Most PodSpec collection fields carry `omitempty`, so a template containing an explicitly empty value (`env: []`) is dropped when the EphemeralRunnerSet is written and reads back as nil. A strict comparison reports drift on every reconcile, bumping ActionableRevision each time and deleting every idle and pending runner, forever. Semantic treats nil and empty as equal, understands resource.Quantity, and cannot panic on unexported fields the way cmp can. It is also roughly six times cheaper than cmp.Equal on a realistic spec. - `Semantic.DeepDerivative` for the listener pod, because the live pod carries many fields the desired pod never sets (nodeName, dnsPolicy, default tolerations, the kube-api-access volume, ...). DeepEqual there would spin in a delete/create loop. Container port length is checked separately, since ports come from the --listener-metrics-addr flag rather than from a resource, so disabling metrics would otherwise leave the port on the pod forever. Drift detection is deliberately not short-circuited on metadata.generation. Re-registration changes the runner scale set ID through an annotation, and metadata changes do not bump generation, so a generation-based shortcut would leave the EphemeralRunnerSet pointing at a scale set that no longer exists. The measured saving did not justify the risk. Additionally: - Count deleting runners toward the scale-up total so terminating runners are not double-replaced. - Cleanup of finished runners is no longer deferred; failures now surface as reconcile errors instead of being logged and swallowed. - Status patches for the new fields use `RetryOnConflict` against a freshly read object. - Keep merging EphemeralRunnerSet annotations and labels rather than overwriting them, so metadata applied by admission webhooks or other controllers is preserved. Drift detection compares against the merge result so foreign keys cannot cause a permanent patch loop. - Add unit tests and benchmarks for both drift checks, including a guard that fails if the listener comparison is ever tightened to DeepEqual. - Cover the re-registration path, which previously had no assertion that the new runner scale set ID reaches the EphemeralRunnerSet at all.