Commit Graph
25 Commits
Author SHA1 Message Date
Nikola JokicandCopilot App 41d05325c2 Make the Outdated phase revision-aware
An outdated runner used to flip the whole scale set into the Outdated
phase permanently, which tears down the listener and switches the scale
set off. That verdict outlived the runner spec it was about: a runner
busy with a job survives the revision cleanup that follows a spec
update, and only reports Outdated once the job finishes. The result was
that a freshly applied fix could be discarded by a runner that never ran
it.

Runners are now stamped with the actionable revision they were built
from. An Outdated runner whose revision is behind the applied revision is
considered stale: it is deleted so the scaling logic replaces it with one
built from the current spec, and it no longer contributes to the set's
phase. A runner at the current revision still marks the set Outdated, so
a genuinely bad spec is still surfaced.

Two supporting fixes:

  - patchAppliedActionableRevisionStatus now recomputes the phase in both
    directions. It only ever forced Running, so the early-return path
    could leave a stale Outdated behind.

  - The AutoscalingRunnerSet only tears down on an Outdated set once that
    set has applied its current actionable revision. Otherwise a spec
    update races the EphemeralRunnerSet controller and the teardown fires
    against a phase that predates the update.

Runners created before this change parse to revision 0, which matches
the zero value of AppliedActionableRevision, so they are treated as
current until a revision is actually applied.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-08 17:27:39 +02:00
Nikola Jokic e481f69cff Replace integrity-hash annotations with explicit API state fields
The `actions.github.com/integrity-hash` annotation was used as an opaque
fingerprint to detect spec drift across AutoscalingRunnerSet,
EphemeralRunnerSet and the listener resources. Hashes are brittle: they
change whenever unrelated serialization details change, they are invisible
to users, and they are not restart-safe. FNV-32a also carries a real
collision risk, where the consequence is an update silently never applied.

Replace it with explicit, typed state:

- `AutoscalingRunnerSetStatus.ObservedGeneration` drives the Pending phase
  transition via `metadata.generation` instead of an annotation hash.
- `EphemeralRunnerSetSpec.ActionableRevision` and
  `EphemeralRunnerSetStatus.AppliedActionableRevision` form a restart-safe
  applied marker. The revision is bumped by the AutoscalingRunnerSet
  controller when `EphemeralRunnerSpec` changes, and only advanced in status
  after idle/pending runner cleanup succeeds.
- `EphemeralRunnerSetStatus.FinishedRunnerCleanupPatchID` records the
  listener patch ID for which finished runners were reaped, so scale-up is
  suppressed until the listener publishes a fresh desired state. This
  prevents creating a replacement runner for a job that already completed.
- Listener pod recreation compares pod specs semantically instead of
  comparing hash annotations.

Drift detection uses `apiequality.Semantic`, not `cmp` or `reflect`:

- `Semantic.DeepEqual` for the EphemeralRunnerSpec. Most PodSpec collection
  fields carry `omitempty`, so a template containing an explicitly empty
  value (`env: []`) is dropped when the EphemeralRunnerSet is written and
  reads back as nil. A strict comparison reports drift on every reconcile,
  bumping ActionableRevision each time and deleting every idle and pending
  runner, forever. Semantic treats nil and empty as equal, understands
  resource.Quantity, and cannot panic on unexported fields the way cmp can.
  It is also roughly six times cheaper than cmp.Equal on a realistic spec.
- `Semantic.DeepDerivative` for the listener pod, because the live pod
  carries many fields the desired pod never sets (nodeName, dnsPolicy,
  default tolerations, the kube-api-access volume, ...). DeepEqual there
  would spin in a delete/create loop. Container port length is checked
  separately, since ports come from the --listener-metrics-addr flag rather
  than from a resource, so disabling metrics would otherwise leave the port
  on the pod forever.

Drift detection is deliberately not short-circuited on
metadata.generation. Re-registration changes the runner scale set ID
through an annotation, and metadata changes do not bump generation, so a
generation-based shortcut would leave the EphemeralRunnerSet pointing at a
scale set that no longer exists. The measured saving did not justify the
risk.

Additionally:

- Count deleting runners toward the scale-up total so terminating runners
  are not double-replaced.
- Cleanup of finished runners is no longer deferred; failures now surface as
  reconcile errors instead of being logged and swallowed.
- Status patches for the new fields use `RetryOnConflict` against a freshly
  read object.
- Keep merging EphemeralRunnerSet annotations and labels rather than
  overwriting them, so metadata applied by admission webhooks or other
  controllers is preserved. Drift detection compares against the merge
  result so foreign keys cannot cause a permanent patch loop.
- Add unit tests and benchmarks for both drift checks, including a guard
  that fails if the listener comparison is ever tightened to DeepEqual.
- Cover the re-registration path, which previously had no assertion that
  the new runner scale set ID reaches the EphemeralRunnerSet at all.
2026-09-07 22:08:17 +02:00
Nikola Jokic 54147cfa5e Introduce cache for lookups of desired resource state, to reduce the amount of allocations in the controller (#4568) 2026-09-07 13:49:28 +02:00
Nikola Jokic f6a3d738de Use metrics to display runner statuses instead of status field for EphemeralRunnerSet and AutoscalingRunnerSet (#4557) 2026-07-14 19:31:11 +02:00
Nikola Jokic 767e58e4b1 Upgrade resources in-place, causing 1-1 mapping between autoscaling runner set and ephemeral runner set (#4516) 2026-06-09 13:52:37 +02:00
Nikola JokicandCopilot Autofix powered by AI 9bc1c9e53e Shutdown the scaleset when runner is deprecated (#4404)
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-03-19 13:30:20 +01:00
Nikola Jokic f99c6eda0b Moving to scaleset client for the controller (#4390) 2026-03-13 14:36:41 +01:00
Nikola Jokic 6d07b8d853 Add ephemeral runner finalizer during creation and check finalizer without requeue (#4320) 2025-11-20 23:06:27 +01:00
Nikola Jokic e46c929241 Azure Key Vault integration to resolve secrets (#4090) 2025-06-11 15:53:33 +02:00
Nikola Jokic cae7efa2c6 Create backoff mechanism for failed runners and allow re-creation of failed ephemeral runners (#4059) 2025-05-14 15:38:50 +02:00
Nikola Jokic b349ded2be Increase test timeouts to avoid CI test failures (#3554) 2024-06-21 13:45:48 +02:00
Nikola Jokic 9b51f25800 Rename imports in tests to remove double import and to improve readability (#3455) 2024-05-17 14:37:13 +02:00
Nikola Jokic 963ae48a3f Include self correction on empty batch and avoid removing pending runners when cluster is busy (#3426) 2024-04-16 12:55:25 +02:00
Nikola JokicandFrancesco Renzi 7a643a5107 Fix overscaling when the controller is much faster then the listener (#3371)
Co-authored-by: Francesco Renzi <rentziass@gmail.com>
2024-03-20 15:36:12 +01:00
Nikola Jokic 07bff8aa1e Extend the user agent and fix the build version for the listener app (#2892) 2023-09-14 20:10:49 +02:00
Tingluo Huang 261d4371b5 Update E2E test workflow. (#2395) 2023-03-14 09:00:07 -04:00
Nikola Jokic babbfc77d5 Surface EphemeralRunnerSet stats to AutoscalingRunnerSet (#2382) 2023-03-13 16:16:28 +01:00
c569304271 Add support for self-signed CA certificates (#2268)
Co-authored-by: Bassem Dghaidi <568794+Link-@users.noreply.github.com>
Co-authored-by: Nikola Jokic <jokicnikola07@gmail.com>
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
2023-03-09 17:23:32 +00:00
Francesco Renzi 40c905f25d Simplify the setup of controller tests (#2352) 2023-03-02 18:55:49 +00:00
Francesco Renzi 73e22a1756 Disable metrics serving in proxy tests (#2307) 2023-02-22 16:57:59 +00:00
6b4250ca90 Add support for proxy (#2286)
Co-authored-by: Nikola Jokic <jokicnikola07@gmail.com>
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
Co-authored-by: Ferenc Hammerl <fhammerl@github.com>
2023-02-21 17:33:48 +00:00
dependabot[bot]andYusuke Kuoka 219ba5b477 chore(deps): bump sigs.k8s.io/controller-runtime from 0.13.1 to 0.14.1 (#2132)
Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: Yusuke Kuoka <ykuoka@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Yusuke Kuoka <ykuoka@gmail.com>
2023-01-27 09:23:28 +09:00
Nikola Jokic 882bfab569 Renaming autoScaling to autoscaling in tests matching the convention (#2201) 2023-01-23 17:03:01 +01:00
Tingluo Huang 4932412cd6 Fix L0 test to make it more reliable. (#2178) 2023-01-19 07:33:04 -05:00
622eaa34f8 Introduce new preview auto-scaling mode for ARC. (#2153)
Co-authored-by: Cory Miller <cory-miller@github.com>
Co-authored-by: Nikola Jokic <nikola-jokic@github.com>
Co-authored-by: Ava Stancu <AvaStancu@github.com>
Co-authored-by: Ferenc Hammerl <fhammerl@github.com>
Co-authored-by: Francesco Renzi <rentziass@github.com>
Co-authored-by: Bassem Dghaidi <Link-@github.com>
2023-01-17 12:06:20 -05:00