Commit Graph
50 Commits
Author SHA1 Message Date
Nikola JokicandCopilot App e2768dddbf Switch the scale set off instead of rebuilding it when runners are outdated
When the runners reject the runner spec they were given, the listener has to
stop acquiring jobs the scale set cannot run. The controller removed the
listener but then deleted the EphemeralRunnerSet, and left the
AutoscalingRunnerSet phase on Running. AutoscalingRunnerSetPhaseOutdated was
declared and read, but never assigned by anything.

Nothing held the scale set switched off as a result. The next reconcile saw a
missing EphemeralRunnerSet, created it, created a listener for it, and the
fresh runners rejected the same spec again, so the scale set churned through
create and teardown cycles against the Actions service instead of resting.

Record the outdated phase and keep the EphemeralRunnerSet, pinned to zero
replicas and patch id. It releases every runner that is not executing a job
while the phase keeps the listener from being rebuilt, and the revision
bookkeeping that decides when the scale set may run again is preserved.

Recovery is driven by the spec update that the phase is waiting for: it moves
the phase back to pending, and the runner spec is then republished to the set
with an advanced revision even when the runner spec itself did not change.
The revision is what tells the EphemeralRunnerSet to stop judging itself by
the runners that failed, so without advancing it a scale set could only be
recovered by editing the pod template, and an edit to anything else would
switch the listener back on against a set parked at zero.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 11:43:46 +02:00
Nikola JokicandCopilot App b3e44a1c39 Filter owned-resource events in the workqueue
Every controller in this package woke up on every update of the resources it
owns, including the ones that only touch fields it never reads. An
EphemeralRunner status carries readiness, failure bookkeeping and the job
details written by the listener, a pod reports IPs, its node, a start time and
several conditions, and an EphemeralRunnerSet rewrites
Status.FinishedRunnerCleanupPatchID for every listener patch id. None of that is
input to the owner, yet all of it enqueued a reconcile.

Add an update predicate to each owned watch. A predicate is only allowed to be
an optimisation, so each one is written as the projection of the fields its
reconciler actually reads and drops an update only when all of them are equal:

  - AutoscalingRunnerSet -> EphemeralRunnerSet: object metadata, the whole spec,
    and the Status.Phase and Status.AppliedActionableRevision pair that
    ephemeralRunnerSetOutdatedForAppliedRevision consults.
  - EphemeralRunnerSet -> EphemeralRunner: object metadata, spec, Status.Phase
    and Status.RunnerID.
  - EphemeralRunner -> Pod: UID, deletion timestamp, pod phase, reason and
    message, the container and init container statuses, and the Ready
    condition.

An event of an unexpected type is always delivered, and create, delete and
generic events are untouched. The primary watches keep seeing every update, so
the status patches a reconciler makes to hand work to its own next pass, such
as marking a runner Failed or Outdated before cleaning up its resources, still
re-enqueue.

The Ready condition lookup is extracted into podReady so that the predicate and
updateRunStatusFromPod cannot disagree about what readiness means.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 11:37:23 +02:00
Nikola JokicandCopilot App a7ce96b76c Take the reconcile snapshot lazily, right before the first mutation
Reconcilers deep copied the object they had just fetched on every single
reconcile, purely so a merge patch could be computed on the rare pass
that actually changes something. The copy is a full recursive walk and
allocation of the object, and the overwhelming majority of reconciles
throw it away untouched.

Introduce lazyCopy, which takes the snapshot on the first call to
Mutate and hands back the live object. Because Mutate is the only way
to reach the object, the snapshot cannot be taken after the mutation it
is supposed to be diffed against, which is the way this optimization is
usually gotten wrong.

Apply it to the four Reconcile entry points, and move the two
EphemeralRunnerSet status copies inside the branch that patches, so they
are only paid for when the status really changed.

While here, drop the short circuit in the EphemeralRunner finalizer
block: `addedFinalizers || AddFinalizer(...)` skipped adding the actions
finalizer whenever the first finalizer was added.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 11:32:03 +02:00
Nikola JokicandCopilot App 60a6753d3b Make the Outdated phase revision-aware
An EphemeralRunnerSet goes Outdated when a child runner reports that the
Actions service rejected its runner spec, and the AutoscalingRunnerSet
then tears the listener down so the scale set stops taking jobs. Until
now nothing recorded which runner spec a given Outdated report was about,
so a report from a runner that predates a spec update kept the set
switched off even after the user had already fixed the spec.

Runners now carry the EphemeralRunnerSet's actionable revision as an
annotation, and Outdated runners are split into two groups:

  - stale outdated: created before the currently applied revision. Their
    verdict is about a spec that no longer exists, so they are deleted
    and replaced by runners built from the current spec.
  - outdated: created at or after the applied revision. Their verdict is
    about the current spec, so they hold the set in the Outdated phase.
    They are deliberately not delete-and-replaced: a fresh runner at the
    same revision would report Outdated again, forever.

patchAppliedActionableRevisionStatus now recomputes the phase in both
directions against the revision being applied, because it returns from
Reconcile without reaching updateStatus. The AutoscalingRunnerSet
teardown guard additionally requires the applied revision to have caught
up with the spec revision, so it does not act on an Outdated verdict for
a spec that has already been replaced.

terminated() now allocates a fresh slice instead of appending into the
finished slice, which could alias its backing array.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 22:56:49 +02:00
Nikola JokicandCopilot App ec793dfe9f Track runner spec updates with an actionable revision
When the AutoscalingRunnerSet's runner spec changes, the EphemeralRunnerSet
has to delete its idle and pending runners so they are rebuilt from the new
spec. That was detected by hashing Spec.EphemeralRunnerSpec into the
actions.github.com/integrity-hash annotation and comparing the annotation
against a freshly computed hash.

Replace it with a monotonic revision counter split across spec and status.
The AutoscalingRunnerSet controller bumps Spec.ActionableRevision when it
patches a new runner spec across; the EphemeralRunnerSet controller advances
Status.AppliedActionableRevision only after the cleanup has actually
succeeded. Because the applied marker lives in status and is written last, a
controller that crashes part-way through the cleanup comes back with the
applied revision still behind the spec revision and redoes the work, instead
of skipping runners that are still running the old spec.

The drift check uses apiequality.Semantic.DeepEqual rather than cmp.Equal or
reflect.DeepEqual. Most PodSpec collection fields carry omitempty, so a
template containing an explicitly empty value (`env: []`) is dropped when the
EphemeralRunnerSet is written and reads back as nil. A strict comparison would
report drift on every reconcile, bump the revision each time, and delete every
idle and pending runner forever. helpers_drift_test.go pins that behaviour with
a round-trip-through-JSON test.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 22:56:48 +02:00
Nikola JokicandCopilot App 80a0a64f3f Report the pending phase when the listener is rebuilt
Moving update detection to metadata.generation lost a signal. The desired
listener is built from the AutoscalingRunnerSet's labels and annotations as
well as its spec, and any difference there deletes the listener so it can
be re-created. metadata.generation only moves on spec writes, so a
label-only edit still tore the listener down while the scale set kept
reporting Running. If the rebuild then failed, it reported Running with no
listener indefinitely. Under the old label-inclusive hash the phase did
move, so this was a regression.

Mark the resource pending at the point the listener is deleted rather than
trying to infer it from the generation. That is what the phase already
means: its own doc comment says pending is when the listener is not yet
started. It also covers the case properly, because it keys off the actual
decision to rebuild instead of guessing from the trigger.

This does not change when the listener is rebuilt. A label-only edit
recreates it today and still does; only the reported phase changes. The
earlier claim that labels no longer restart anything was wrong: labels are
propagated to the listener, so editing one is a restart. What
metadata.generation changes is narrower than that, and the test now covers
the listener's identity across a label edit rather than implying it is
untouched.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 22:56:47 +02:00
Nikola JokicandCopilot App b20df164d6 Track AutoscalingRunnerSet updates with metadata.generation
The AutoscalingRunnerSet controller detected changes by hashing the spec
and the labels on every reconcile and comparing the result against the
actions.github.com/integrity-hash annotation it had written on a previous
pass. That is a hand-rolled version of something the API server already
does for us: it bumps metadata.generation on every spec write, and nothing
else. Recording that value in the status as ObservedGeneration gives the
same "has this changed since I last applied it" signal without recomputing
a hash, without an extra non-status write to the object, and with a value a
user can read and reason about.

updateStatus now takes the observed generation and patches it next to the
phase, returning early only when both are already what we want. When the
generation is ahead of the observed generation we move to Pending but
deliberately leave the observed generation where it is; it only catches up
at the end of a reconcile that actually applied the change, so a reconcile
that fails part way through is retried as pending rather than being
mistaken for settled.

The old code returned immediately after stamping the hash. That was needed
because stamping the annotation was itself a write to the object, so
continuing would have worked from a stale copy. Recording the generation
only touches the status subresource, which does not bump generation, so the
rest of the reconcile can run on the same object and apply the change in
the same pass instead of waiting for the next event.

One user-visible difference: the old hash covered Labels as well as Spec,
and metadata.generation does not move when only labels change. Editing just
a label on an AutoscalingRunnerSet no longer forces it through Pending.
Labels still propagate to the EphemeralRunnerSet; they simply no longer
count as an update that has to be applied.

Hash, ListenerSpecHash and RunnerSetSpecHash are removed. The latter two
already had no callers, and the first has none now.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 22:56:47 +02:00
c475023e28 Fix EphemeralRunnerSet metadata drift check comparing against the wrong value (#4634)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-10 11:52:08 +02:00
Nikola Jokic 54147cfa5e Introduce cache for lookups of desired resource state, to reduce the amount of allocations in the controller (#4568) 2026-09-07 13:49:28 +02:00
Diogo Torres 3cfdcf59fe Re-register the runner scale set when it is gone from the Actions service (#4571) 2026-09-03 08:16:06 +01:00
Nikola Jokic f6a3d738de Use metrics to display runner statuses instead of status field for EphemeralRunnerSet and AutoscalingRunnerSet (#4557) 2026-07-14 19:31:11 +02:00
Nikola Jokic 767e58e4b1 Upgrade resources in-place, causing 1-1 mapping between autoscaling runner set and ephemeral runner set (#4516) 2026-06-09 13:52:37 +02:00
Nikola Jokic 30879de182 Fix patch on autoscaling runner set when creating a runner scale set (#4502) 2026-05-22 17:51:47 +02:00
Junya Okabe a401686bd5 Add option to disable workqueue bucket rate limiter (#4451) 2026-04-22 23:26:39 +02:00
Nikola Jokic 802dc28d38 Add multi-label support to scalesets (#4408) 2026-03-19 15:29:40 +01:00
Nikola JokicandCopilot Autofix powered by AI 9bc1c9e53e Shutdown the scaleset when runner is deprecated (#4404)
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-03-19 13:30:20 +01:00
Nikola Jokic f99c6eda0b Moving to scaleset client for the controller (#4390) 2026-03-13 14:36:41 +01:00
Nikola Jokic 95d2107a6a Code style changes on the controller (#4324) 2025-11-21 14:20:44 +01:00
Nikola Jokic e46c929241 Azure Key Vault integration to resolve secrets (#4090) 2025-06-11 15:53:33 +02:00
Nikola JokicandCopilot 389d842a30 Relax version requirements to allow patch version mismatch (#4080)
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-14 21:38:16 +02:00
Ryosei Karaki f832b0b254 upgrade(golangci-lint): v2.1.2 (#4023)
Signed-off-by: karamaru-alpha <mrnk3078@gmail.com>
2025-04-17 16:14:31 +02:00
Nikola Jokic 2dab45c373 Wrap errors in controller helper methods and swap logic in cleanups (#3960) 2025-03-07 11:58:53 +01:00
Nikola Jokic e12a892748 Rename log from target/actual to build/autoscalingRunnerSet version (#3957) 2025-03-04 17:01:34 +01:00
Nikola Jokic a62ca3d853 Exclude label prefix propagation (#3607) 2024-06-21 12:12:14 +02:00
Nikola Jokic 9afd93065f Remove finalizers in one pass to speed up cleanups AutoscalingRunnerSet (#3536) 2024-05-27 09:21:31 +02:00
Nikola Jokic fa7a4f584e Extract single place to set up indexers (#3454) 2024-05-17 14:42:46 +02:00
Nikola Jokic 963ae48a3f Include self correction on empty batch and avoid removing pending runners when cluster is busy (#3426) 2024-04-16 12:55:25 +02:00
Nikola Jokic c9099a5a56 Add annotation with values hash to re-create listener (#3195) 2024-03-19 14:29:49 +01:00
Hidehito Yabuuchi 48706584fd Propagate runner scale set name annotation to EphemeralRunner (#3098) 2024-03-19 12:50:49 +01:00
Nikola Jokic 202a97ab12 Modify user agent format with subsystem and is proxy configured information (#3116) 2023-12-08 13:16:29 +01:00
Nikola Jokic 65fd04540c Bump go version and all direct dependencies to newest for k8s compatibility (#2947) 2023-11-14 16:19:43 +01:00
Nikola Jokic 07bff8aa1e Extend the user agent and fix the build version for the listener app (#2892) 2023-09-14 20:10:49 +02:00
Bassem Dghaidi 8afef51c8b Add DrainJobsMode (aka UpdateStrategy feature) (#2569) 2023-05-23 07:42:30 -04:00
Nikola JokicandBassem Dghaidi 9859bbc7f2 Use build.Version to check if resource version is a mismatch (#2521)
Co-authored-by: Bassem Dghaidi <568794+Link-@users.noreply.github.com>
2023-04-24 10:40:15 +02:00
Nikola Jokic ba1ac0990b Reordering methods and constants so it is easier to look it up (#2501) 2023-04-12 09:50:23 +02:00
Nikola JokicandTingluo Huang 5dea6db412 Fix helm uninstall cleanup by adding finalizers and cleaning them from the controller (#2433)
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
2023-04-03 21:06:12 +02:00
Nikola JokicandTingluo Huang 56e1c62ac2 Add labels to autoscaling runner set subresources to allow easier inspection (#2391)
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
2023-03-27 11:19:34 +02:00
Tingluo Huang 08acb1b831 Get RunnerScaleSet based on both RunnerGroupId and Name. (#2413) 2023-03-15 11:10:09 -04:00
Nikola Jokic babbfc77d5 Surface EphemeralRunnerSet stats to AutoscalingRunnerSet (#2382) 2023-03-13 16:16:28 +01:00
c569304271 Add support for self-signed CA certificates (#2268)
Co-authored-by: Bassem Dghaidi <568794+Link-@users.noreply.github.com>
Co-authored-by: Nikola Jokic <jokicnikola07@gmail.com>
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
2023-03-09 17:23:32 +00:00
Chris PattersonandTingluoHuang 41f2ca3ed9 Adding parameter to configure the runner set name. (#2279)
Co-authored-by: TingluoHuang <TingluoHuang@github.com>
2023-03-03 08:36:14 -05:00
Nikola Jokic 2984de912c Split listener pod label to avoid long names issue (#2341) 2023-03-02 17:25:50 +01:00
6b4250ca90 Add support for proxy (#2286)
Co-authored-by: Nikola Jokic <jokicnikola07@gmail.com>
Co-authored-by: Tingluo Huang <tingluohuang@github.com>
Co-authored-by: Ferenc Hammerl <fhammerl@github.com>
2023-02-21 17:33:48 +00:00
Nikola Jokic 9990243520 Early return if finalizer does not exist to make it more readable (#2262) 2023-02-08 15:21:13 +01:00
Tingluo Huang facae69e0b Remove un-required permissions for the manager-role of the new AutoScalingRunnerSet (#2260) 2023-02-07 12:37:09 -05:00
Nikola Jokic c4297d25bb Avoid deleting scale set if annotation is not parsable or if it does not exist (#2239) 2023-02-03 17:27:31 +01:00
Tingluo Huang 1f4fe4681e Delete RunnerScaleSet on service when AutoScalingRunnerSet is deleted. (#2223) 2023-01-31 15:03:11 -05:00
Tingluo Huang 803818162c Allow update runner group for AutoScalingRunnerSet (#2216) 2023-01-27 09:27:52 -05:00
dependabot[bot]andYusuke Kuoka 219ba5b477 chore(deps): bump sigs.k8s.io/controller-runtime from 0.13.1 to 0.14.1 (#2132)
Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: Yusuke Kuoka <ykuoka@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Yusuke Kuoka <ykuoka@gmail.com>
2023-01-27 09:23:28 +09:00
622eaa34f8 Introduce new preview auto-scaling mode for ARC. (#2153)
Co-authored-by: Cory Miller <cory-miller@github.com>
Co-authored-by: Nikola Jokic <nikola-jokic@github.com>
Co-authored-by: Ava Stancu <AvaStancu@github.com>
Co-authored-by: Ferenc Hammerl <fhammerl@github.com>
Co-authored-by: Francesco Renzi <rentziass@github.com>
Co-authored-by: Bassem Dghaidi <Link-@github.com>
2023-01-17 12:06:20 -05:00