The scaler issued every API call a message asked for before returning, and
the listener acks only once it returns. A full batch of job started events
is 50 of them, two calls each, so the next scale decision sat behind 100
calls of bookkeeping.
Those patches are not what new jobs wait on. Only the desired runner count
creates runners; Status.JobID is a hint the runner set consults when
choosing which idle runner to delete, and the Actions service rejects the
deletion of a runner whose job is still running either way.
Publish the desired count first, hand the job started events to a
background pool, and acquire after. Acquiring last costs the round trip of
the scale patch and saves the whole batch, and the scale decision is
unaffected: it comes from msg.Statistics, a snapshot the service took when
it built the message, so jobs acquired now are reported as assigned in a
later one.
Give the two kinds of traffic their own clients. Sharing one token bucket
is what let the job patches delay the scale patch, so splitting them is
what makes the reordering worth anything; backgrounding alone would just
move the same queue.
Measured against a 5ms API server and a 50ms Actions service, at a full
50 event batch:
scale patch reaches the API server 1.971s -> 6ms
listener loop 2.02s -> 107ms/msg
The loop no longer spends the rate limit inline, so its cost is now the
two service round trips rather than the batch size.
This does not raise throughput. Job patches still cost two calls per
event, so the job client sustains qps/2 job starts per second, and the
queue is bounded by the real job start rate rather than by how fast the
listener polls: a faster loop polls more often and carries proportionally
fewer events. Measured at qps 40, the queue stays empty through 18
starts/sec and degrades gradually past 20 rather than falling over.
Close drains the queue, since the message these patches came from was
acked long before they run.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The scaleset listener no longer dissects the message it polls. It owns
session management, polling and acking, and hands the whole message to a
single Scale call, so acquiring jobs and recording metrics move to the
only component that still reads them.
That handover is what makes the work parallelisable. The listener used to
replay a message one API call at a time, in a fixed order: every job
started patch, then every job completed, then the scale patch. The job
events touch distinct EphemeralRunners and carry no ordering between
them, so they now run across a bounded worker pool, with the worker that
patches the EphemeralRunnerSet running alongside them.
The one ordering that does matter is kept. deleteIdleEphemeralRunners
skips a runner only once it carries a job request ID, so a patch that
lowers the replica count could offer up a runner that just picked up a
job if it were published while job started patches were still in flight.
The scaling worker therefore waits for the event workers on a scale down,
and only then. A patch that scales up or holds cannot delete anything.
Kubernetes has no bulk write: get, create, update, patch and delete are
single-resource verbs, and deletecollection is the only collection-scoped
mutating verb there is, so N events cannot be collapsed into fewer
requests. Server side apply is a PATCH on one object URI and does not
change that. Issuing the N patches concurrently over the one HTTP/2
connection is the available win; the alternative is writing fewer
objects, which trades N cheap independent writes for one contended,
watch-amplifying, size-bounded write.
The pool size is configurable through listenerConfig.scaler.workers,
alongside the existing qps and burst, and defaults to 10: two calls per
event keeps a full pool well inside the default QPS budget.
Statistics are now cached by the scaler. The listener stopped tracking
them, and a long poll that times out carries no message at all, so
without the cache an idle scale set would stop converging.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Changed folder structure to allow multi group registration
* included actions.github.com directory for resources and controllers
* updated go module to actions/actions-runner-controller
* publish arc packages under actions-runner-controller
* Update charts/actions-runner-controller/docs/UPGRADING.md
Co-authored-by: Yusuke Kuoka <ykuoka@gmail.com>