Commit Graph
1882 Commits
Author SHA1 Message Date
Nikola JokicandCopilot App 2877779b72 Patch job started events off the scale decision path
The scaler issued every API call a message asked for before returning, and
the listener acks only once it returns. A full batch of job started events
is 50 of them, two calls each, so the next scale decision sat behind 100
calls of bookkeeping.

Those patches are not what new jobs wait on. Only the desired runner count
creates runners; Status.JobID is a hint the runner set consults when
choosing which idle runner to delete, and the Actions service rejects the
deletion of a runner whose job is still running either way.

Publish the desired count first, hand the job started events to a
background pool, and acquire after. Acquiring last costs the round trip of
the scale patch and saves the whole batch, and the scale decision is
unaffected: it comes from msg.Statistics, a snapshot the service took when
it built the message, so jobs acquired now are reported as assigned in a
later one.

Give the two kinds of traffic their own clients. Sharing one token bucket
is what let the job patches delay the scale patch, so splitting them is
what makes the reordering worth anything; backgrounding alone would just
move the same queue.

Measured against a 5ms API server and a 50ms Actions service, at a full
50 event batch:

  scale patch reaches the API server   1.971s -> 6ms
  listener loop                        2.02s  -> 107ms/msg

The loop no longer spends the rate limit inline, so its cost is now the
two service round trips rather than the batch size.

This does not raise throughput. Job patches still cost two calls per
event, so the job client sustains qps/2 job starts per second, and the
queue is bounded by the real job start rate rather than by how fast the
listener polls: a faster loop polls more often and carries proportionally
fewer events. Measured at qps 40, the queue stays empty through 18
starts/sec and degrades gradually past 20 rather than falling over.

Close drains the queue, since the message these patches came from was
acked long before they run.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-18 15:05:54 +02:00
Nikola JokicandCopilot App 8148c24c37 Publish the desired runner count before the job event patches
Scale used to hold the replica patch back whenever the target dropped, so
that every job started patch had landed before the runner set controller
could act on a lower count. The reasoning was that
deleteIdleEphemeralRunners skips a runner only once it carries a job ID,
so a runner that had just picked up a job could otherwise be deleted.

That guard was unreachable. The controller only deletes idle runners
under Spec.PatchID == 0, and setDesiredWorkerState emits patch ID 0 only
when the target is unchanged and equal to MinRunners, or on the very
first patch, when no previous target exists. Neither can coincide with a
falling target, so a scale down never reaches the deletion path.
Exhaustively walking message sequences over every MinRunners/MaxRunners
pair finds no state where the two occur together.

So the replica patch has no reason to wait, and good reason to go first:
it is the only patch that creates runners, and therefore the one new jobs
wait on, while the job event patches are bookkeeping the controller reads
later. Sending it first also keeps it clear of the client rate limiter,
which a large batch of event patches would otherwise drain ahead of it.

The job events still have to land before Scale returns. The listener acks
the message the moment it does, and nothing other than these patches ever
writes Status.JobID, so a patch dropped after the ack would leave a busy
runner looking idle to the scale down that a later patch ID 0 permits.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-18 15:05:54 +02:00
Nikola JokicandCopilot App 7833438dd7 Handle scale set messages with parallel workers in the listener
The scaleset listener no longer dissects the message it polls. It owns
session management, polling and acking, and hands the whole message to a
single Scale call, so acquiring jobs and recording metrics move to the
only component that still reads them.

That handover is what makes the work parallelisable. The listener used to
replay a message one API call at a time, in a fixed order: every job
started patch, then every job completed, then the scale patch. The job
events touch distinct EphemeralRunners and carry no ordering between
them, so they now run across a bounded worker pool, with the worker that
patches the EphemeralRunnerSet running alongside them.

The one ordering that does matter is kept. deleteIdleEphemeralRunners
skips a runner only once it carries a job request ID, so a patch that
lowers the replica count could offer up a runner that just picked up a
job if it were published while job started patches were still in flight.
The scaling worker therefore waits for the event workers on a scale down,
and only then. A patch that scales up or holds cannot delete anything.

Kubernetes has no bulk write: get, create, update, patch and delete are
single-resource verbs, and deletecollection is the only collection-scoped
mutating verb there is, so N events cannot be collapsed into fewer
requests. Server side apply is a PATCH on one object URI and does not
change that. Issuing the N patches concurrently over the one HTTP/2
connection is the available win; the alternative is writing fewer
objects, which trades N cheap independent writes for one contended,
watch-amplifying, size-bounded write.

The pool size is configurable through listenerConfig.scaler.workers,
alongside the existing qps and burst, and defaults to 10: two calls per
event keeps a full pool well inside the default QPS budget.

Statistics are now cached by the scaler. The listener stopped tracking
them, and a long poll that times out carries no message at all, so
without the cache an idle scale set would stop converging.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-18 15:05:54 +02:00
Nikola JokicandCopilot App e8753bcf57 Cover the outdated runner lifecycle end to end (#4662)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-18 11:18:59 +02:00
Nikola Jokic 6df0e09f84 Revert "Remove legacy e2e tests and promote v2 tests (#4658)" (#4667) 2026-09-18 10:03:18 +02:00
Nikola Jokic 78dacc7a53 Deregister runners from the Actions service in the background (#4664) 2026-09-17 17:11:59 +02:00
dc12c2d417 Keep the outdated phase until the runner spec itself changes (#4660)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-17 17:03:21 +02:00
Nikola JokicandCopilot App be44bb58cf Pin the listener's optimistic lock to API server behavior (#4648)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-17 17:02:45 +02:00
Nikola JokicandCopilot App b9eaf560f8 Switch the scale set off instead of rebuilding it when runners are outdated (#4652)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-15 16:23:37 +02:00
Nikola Jokic e425370456 Remove legacy e2e tests and promote v2 tests (#4658) 2026-09-15 16:12:51 +02:00
Nikola JokicandCopilot App ff8bfdee15 Filter owned-resource events in the workqueue (#4647)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-15 14:14:33 +02:00
9ce3169df3 Let the listener own the EphemeralRunner Running phase transition (#4646)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-15 13:46:40 +02:00
Nikola JokicandCopilot App d386789092 Use lazy copy to patch resources to ensure multiple modifications are applied to the base resource (#4580)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-14 16:39:22 +02:00
Nikola JokicandCopilot App 6e4d85f28f Make the Outdated phase revision-aware (#4644)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-14 14:20:33 +02:00
Nikola JokicandCopilot App 03328aa14f Remove the integrity hash annotation (#4643)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 18:27:55 +02:00
484564e6d6 Defer scale up until the listener publishes a state that accounts for finished runners (#4642)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-11 15:29:02 +02:00
6f89d057c0 Track runner spec updates with an actionable revision (#4638)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-11 12:44:03 +00:00
Nikola JokicandCopilot App 5581cd0c54 Track AutoscalingRunnerSet updates with metadata.generation (#4636)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-11 14:38:30 +02:00
dependabot[bot] db5702bb71 Bump the actions group across 1 directory with 8 updates (#4623) 2026-09-10 23:15:31 +02:00
Nikola JokicandCopilot App 0cfedfbb2c Detect listener pod drift by comparing the pod, not a hash (#4635)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 22:56:43 +02:00
c475023e28 Fix EphemeralRunnerSet metadata drift check comparing against the wrong value (#4634)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-10 11:52:08 +02:00
Nikola JokicandCopilot App 22046a2718 Validate and coerce listener metadata at render time (#4640)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 09:14:48 +00:00
Nikola JokicandCopilot App adf7d0026e Validate label and annotation metadata at chart render time (#4639)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-10 11:12:10 +02:00
b0e69a37bd Render label and annotation values as strings (#4637)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-10 10:22:05 +02:00
KR Ravindra 5e540b6c71 Add default and per-controller max-concurrent-reconciles flags (#4626)
Signed-off-by: KR Ravindra <42912207+KR-Ravindra@users.noreply.github.com>
2026-09-09 09:50:13 +02:00
Nikola Jokic 7c68e1d318 Set listener qps and burst (#4558) 2026-09-08 15:49:58 +02:00
Nikola Jokic bc677d306c Add @actions/actions-runtime to codeowners (#4631) 2026-09-08 15:47:46 +02:00
Nikola Jokic f0d0b1d439 Add guards on timings so we don't publish incorrect metrics (#4621) 2026-09-08 14:32:24 +02:00
Nikola Jokic c3dfb396d3 Fix nil ResourceCache panic in stale scale set tests (#4628) 2026-09-08 12:34:45 +02:00
Nikola Jokic 54147cfa5e Introduce cache for lookups of desired resource state, to reduce the amount of allocations in the controller (#4568) 2026-09-07 13:49:28 +02:00
Junya Okabe 87592be6d3 Remove orphaned testserver package (#4620) 2026-09-07 10:40:28 +02:00
dependabot[bot] a4ac89e2a7 Bump golang.org/x/crypto from 0.51.0 to 0.52.0 (#4562)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 16:45:00 +01:00
Diogo Torres 3cfdcf59fe Re-register the runner scale set when it is gone from the Actions service (#4571) 2026-09-03 08:16:06 +01:00
github-actions[bot] 438440e92a Updates: runner to v2.337.0 (#4613)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-31 16:53:02 +01:00
github-actions[bot] a035c5a393 Updates: runner to v2.336.0 (#4578)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-21 15:51:06 +02:00
dependabot[bot]andNikola Jokic 252eb51766 Bump the actions group across 1 directory with 6 updates (#4559)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Nikola Jokic <jokicnikola07@gmail.com>
2026-07-15 14:44:09 +02:00
Nikola Jokic 03463c051c Fix deprecated calls (#4570) 2026-07-14 19:34:13 +02:00
Nikola Jokic f6a3d738de Use metrics to display runner statuses instead of status field for EphemeralRunnerSet and AutoscalingRunnerSet (#4557) 2026-07-14 19:31:11 +02:00
Nikola Jokic 368e2f28b8 Use Patch instead of Update (#4533) 2026-07-10 12:34:40 +02:00
Nikola Jokic 2fa72b510f Tag fields with optional allowing patches without issues (#4528) 2026-07-10 12:28:10 +02:00
Junya Okabe 68c0a86188 Mark generated files as linguist-generated (#4536) 2026-07-09 14:22:46 +02:00
Junya Okabe e06294f5de Make controller terminationGracePeriodSeconds configurable and align it with graceful shutdown timeout (#4556) 2026-07-07 00:28:57 +02:00
Junya Okabe 2ee6b3b8e8 Increase e2e wait timeouts to reduce flakiness (#4545) 2026-06-30 15:05:58 +02:00
dependabot[bot] 24686a974e Bump the actions group across 1 directory with 2 updates (#4539)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-27 18:30:32 +02:00
dependabot[bot] c7005c3696 Bump the gomod group across 1 directory with 11 updates (#4529)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-19 13:08:17 +02:00
Nikola Jokic 9c50514160 Fix typo and rename status to phase (#4506) 2026-06-15 12:51:13 +02:00
github-actions[bot] 391bc57773 Updates: runner to v2.335.1 (#4523)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-12 16:28:17 +02:00
Nikola Jokic 767e58e4b1 Upgrade resources in-place, causing 1-1 mapping between autoscaling runner set and ephemeral runner set (#4516) 2026-06-09 13:52:37 +02:00
dependabot[bot] 0acef229e2 Bump the actions group across 1 directory with 6 updates (#4518)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-09 10:50:43 +02:00
dependabot[bot]andJiaren Wu 0dc5f8a0c2 Bump the gomod group across 1 directory with 9 updates (#4508)
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Jiaren Wu <jiaren-wu@github.com>
2026-05-30 00:55:53 +02:00