Scale used to hold the replica patch back whenever the target dropped, so
that every job started patch had landed before the runner set controller
could act on a lower count. The reasoning was that
deleteIdleEphemeralRunners skips a runner only once it carries a job ID,
so a runner that had just picked up a job could otherwise be deleted.
That guard was unreachable. The controller only deletes idle runners
under Spec.PatchID == 0, and setDesiredWorkerState emits patch ID 0 only
when the target is unchanged and equal to MinRunners, or on the very
first patch, when no previous target exists. Neither can coincide with a
falling target, so a scale down never reaches the deletion path.
Exhaustively walking message sequences over every MinRunners/MaxRunners
pair finds no state where the two occur together.
So the replica patch has no reason to wait, and good reason to go first:
it is the only patch that creates runners, and therefore the one new jobs
wait on, while the job event patches are bookkeeping the controller reads
later. Sending it first also keeps it clear of the client rate limiter,
which a large batch of event patches would otherwise drain ahead of it.
The job events still have to land before Scale returns. The listener acks
the message the moment it does, and nothing other than these patches ever
writes Status.JobID, so a patch dropped after the ack would leave a busy
runner looking idle to the scale down that a later patch ID 0 permits.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The scaleset listener no longer dissects the message it polls. It owns
session management, polling and acking, and hands the whole message to a
single Scale call, so acquiring jobs and recording metrics move to the
only component that still reads them.
That handover is what makes the work parallelisable. The listener used to
replay a message one API call at a time, in a fixed order: every job
started patch, then every job completed, then the scale patch. The job
events touch distinct EphemeralRunners and carry no ordering between
them, so they now run across a bounded worker pool, with the worker that
patches the EphemeralRunnerSet running alongside them.
The one ordering that does matter is kept. deleteIdleEphemeralRunners
skips a runner only once it carries a job request ID, so a patch that
lowers the replica count could offer up a runner that just picked up a
job if it were published while job started patches were still in flight.
The scaling worker therefore waits for the event workers on a scale down,
and only then. A patch that scales up or holds cannot delete anything.
Kubernetes has no bulk write: get, create, update, patch and delete are
single-resource verbs, and deletecollection is the only collection-scoped
mutating verb there is, so N events cannot be collapsed into fewer
requests. Server side apply is a PATCH on one object URI and does not
change that. Issuing the N patches concurrently over the one HTTP/2
connection is the available win; the alternative is writing fewer
objects, which trades N cheap independent writes for one contended,
watch-amplifying, size-bounded write.
The pool size is configurable through listenerConfig.scaler.workers,
alongside the existing qps and burst, and defaults to 10: two calls per
event keeps a full pool well inside the default QPS budget.
Statistics are now cached by the scaler. The listener stopped tracking
them, and a long poll that times out carries no message at all, so
without the cache an idle scale set would stop converging.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>