Patch job started events off the scale decision path

The scaler issued every API call a message asked for before returning, and
the listener acks only once it returns. A full batch of job started events
is 50 of them, two calls each, so the next scale decision sat behind 100
calls of bookkeeping.

Those patches are not what new jobs wait on. Only the desired runner count
creates runners; Status.JobID is a hint the runner set consults when
choosing which idle runner to delete, and the Actions service rejects the
deletion of a runner whose job is still running either way.

Publish the desired count first, hand the job started events to a
background pool, and acquire after. Acquiring last costs the round trip of
the scale patch and saves the whole batch, and the scale decision is
unaffected: it comes from msg.Statistics, a snapshot the service took when
it built the message, so jobs acquired now are reported as assigned in a
later one.

Give the two kinds of traffic their own clients. Sharing one token bucket
is what let the job patches delay the scale patch, so splitting them is
what makes the reordering worth anything; backgrounding alone would just
move the same queue.

Measured against a 5ms API server and a 50ms Actions service, at a full
50 event batch:

  scale patch reaches the API server   1.971s -> 6ms
  listener loop                        2.02s  -> 107ms/msg

The loop no longer spends the rate limit inline, so its cost is now the
two service round trips rather than the batch size.

This does not raise throughput. Job patches still cost two calls per
event, so the job client sustains qps/2 job starts per second, and the
queue is bounded by the real job start rate rather than by how fast the
listener polls: a faster loop polls more often and carries proportionally
fewer events. Measured at qps 40, the queue stays empty through 18
starts/sec and degrades gradually past 20 rather than falling over.

Close drains the queue, since the message these patches came from was
acked long before they run.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
Nikola Jokic
2026-09-18 15:05:54 +02:00
co-authored by Copilot App
parent 8148c24c37
commit 2877779b72
17 changed files with 928 additions and 155 deletions
@@ -14,19 +14,58 @@ func (c *ListenerConfig) GetScaler() *ScalerConfig {
return c.Scaler
}
// ScalerConfig configures the Kubernetes client used by the ghalistener scaler.
// ScalerConfig configures the Kubernetes clients used by the ghalistener scaler.
//
// The scaler talks to the API server over two independent clients, because the
// two kinds of traffic have very different shapes and only one of them is on the
// critical path:
//
// - The scale client publishes the desired runner count. That is one patch per
// message, and it is the only patch that creates runners, so it must never
// queue behind anything.
// - The job client patches job started events. That is two calls per event and
// can be a whole batch at once, so it is the traffic that actually consumes
// the rate limit.
//
// A single shared client lets a batch of job event patches drain the token
// bucket ahead of the scale patch, which delays the only call new jobs are
// waiting on. Splitting them keeps the scale patch clear of that backlog.
type ScalerConfig struct {
// QPS is the query per second limit of the client that patches job started
// events. This is the bulk of the scaler's API traffic, at up to two calls
// per job started event.
// +optional
// +kubebuilder:validation:Minimum:=1
QPS *int `json:"qps,omitempty"`
// Burst is the burst limit of the client that patches job started events.
// +optional
// +kubebuilder:validation:Minimum:=1
Burst *int `json:"burst,omitempty"`
// Workers is the number of job started and job completed events the scaler
// handles concurrently within a single scale set message. The worker that
// scales the EphemeralRunnerSet runs on top of these.
// ScaleQPS is the query per second limit of the client that publishes the
// desired runner count. The scaler issues at most one such patch per scale
// set message, so this only has to be large enough that the patch never
// waits on a token; it is deliberately a small budget separate from QPS
// rather than a share of it.
// +optional
// +kubebuilder:validation:Minimum:=1
ScaleQPS *int `json:"scaleQPS,omitempty"`
// ScaleBurst is the burst limit of the client that publishes the desired
// runner count.
// +optional
// +kubebuilder:validation:Minimum:=1
ScaleBurst *int `json:"scaleBurst,omitempty"`
// Workers is the number of job started events the scaler patches
// concurrently. The events are drained from a background queue rather than
// being tied to the message they arrived on, so this bounds how many job
// patches are in flight at any moment, not how many a single message may
// carry.
//
// Raising it past what QPS sustains does nothing, since the rate limiter
// rather than the worker count is what bounds throughput.
// +optional
// +kubebuilder:validation:Minimum:=1
Workers *int `json:"workers,omitempty"`
@@ -807,6 +807,16 @@ func (in *ScalerConfig) DeepCopyInto(out *ScalerConfig) {
*out = new(int)
**out = **in
}
if in.ScaleQPS != nil {
in, out := &in.ScaleQPS, &out.ScaleQPS
*out = new(int)
**out = **in
}
if in.ScaleBurst != nil {
in, out := &in.ScaleBurst, &out.ScaleBurst
*out = new(int)
**out = **in
}
if in.Workers != nil {
in, out := &in.Workers, &out.Workers
*out = new(int)