mirror of
https://github.com/actions-runner-controller/actions-runner-controller.git
synced 2026-10-04 03:02:12 +02:00
Patch job started events off the scale decision path
The scaler issued every API call a message asked for before returning, and the listener acks only once it returns. A full batch of job started events is 50 of them, two calls each, so the next scale decision sat behind 100 calls of bookkeeping. Those patches are not what new jobs wait on. Only the desired runner count creates runners; Status.JobID is a hint the runner set consults when choosing which idle runner to delete, and the Actions service rejects the deletion of a runner whose job is still running either way. Publish the desired count first, hand the job started events to a background pool, and acquire after. Acquiring last costs the round trip of the scale patch and saves the whole batch, and the scale decision is unaffected: it comes from msg.Statistics, a snapshot the service took when it built the message, so jobs acquired now are reported as assigned in a later one. Give the two kinds of traffic their own clients. Sharing one token bucket is what let the job patches delay the scale patch, so splitting them is what makes the reordering worth anything; backgrounding alone would just move the same queue. Measured against a 5ms API server and a 50ms Actions service, at a full 50 event batch: scale patch reaches the API server 1.971s -> 6ms listener loop 2.02s -> 107ms/msg The loop no longer spends the rate limit inline, so its cost is now the two service round trips rather than the batch size. This does not raise throughput. Job patches still cost two calls per event, so the job client sustains qps/2 job starts per second, and the queue is bounded by the real job start rate rather than by how fast the listener polls: a faster loop polls more often and carries proportionally fewer events. Measured at qps 40, the queue stays empty through 18 starts/sec and degrades gradually past 20 rather than falling over. Close drains the queue, since the message these patches came from was acked long before they run. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
co-authored by
Copilot App
parent
8148c24c37
commit
2877779b72
@@ -14,19 +14,58 @@ func (c *ListenerConfig) GetScaler() *ScalerConfig {
|
||||
return c.Scaler
|
||||
}
|
||||
|
||||
// ScalerConfig configures the Kubernetes client used by the ghalistener scaler.
|
||||
// ScalerConfig configures the Kubernetes clients used by the ghalistener scaler.
|
||||
//
|
||||
// The scaler talks to the API server over two independent clients, because the
|
||||
// two kinds of traffic have very different shapes and only one of them is on the
|
||||
// critical path:
|
||||
//
|
||||
// - The scale client publishes the desired runner count. That is one patch per
|
||||
// message, and it is the only patch that creates runners, so it must never
|
||||
// queue behind anything.
|
||||
// - The job client patches job started events. That is two calls per event and
|
||||
// can be a whole batch at once, so it is the traffic that actually consumes
|
||||
// the rate limit.
|
||||
//
|
||||
// A single shared client lets a batch of job event patches drain the token
|
||||
// bucket ahead of the scale patch, which delays the only call new jobs are
|
||||
// waiting on. Splitting them keeps the scale patch clear of that backlog.
|
||||
type ScalerConfig struct {
|
||||
// QPS is the query per second limit of the client that patches job started
|
||||
// events. This is the bulk of the scaler's API traffic, at up to two calls
|
||||
// per job started event.
|
||||
// +optional
|
||||
// +kubebuilder:validation:Minimum:=1
|
||||
QPS *int `json:"qps,omitempty"`
|
||||
|
||||
// Burst is the burst limit of the client that patches job started events.
|
||||
// +optional
|
||||
// +kubebuilder:validation:Minimum:=1
|
||||
Burst *int `json:"burst,omitempty"`
|
||||
|
||||
// Workers is the number of job started and job completed events the scaler
|
||||
// handles concurrently within a single scale set message. The worker that
|
||||
// scales the EphemeralRunnerSet runs on top of these.
|
||||
// ScaleQPS is the query per second limit of the client that publishes the
|
||||
// desired runner count. The scaler issues at most one such patch per scale
|
||||
// set message, so this only has to be large enough that the patch never
|
||||
// waits on a token; it is deliberately a small budget separate from QPS
|
||||
// rather than a share of it.
|
||||
// +optional
|
||||
// +kubebuilder:validation:Minimum:=1
|
||||
ScaleQPS *int `json:"scaleQPS,omitempty"`
|
||||
|
||||
// ScaleBurst is the burst limit of the client that publishes the desired
|
||||
// runner count.
|
||||
// +optional
|
||||
// +kubebuilder:validation:Minimum:=1
|
||||
ScaleBurst *int `json:"scaleBurst,omitempty"`
|
||||
|
||||
// Workers is the number of job started events the scaler patches
|
||||
// concurrently. The events are drained from a background queue rather than
|
||||
// being tied to the message they arrived on, so this bounds how many job
|
||||
// patches are in flight at any moment, not how many a single message may
|
||||
// carry.
|
||||
//
|
||||
// Raising it past what QPS sustains does nothing, since the rate limiter
|
||||
// rather than the worker count is what bounds throughput.
|
||||
// +optional
|
||||
// +kubebuilder:validation:Minimum:=1
|
||||
Workers *int `json:"workers,omitempty"`
|
||||
|
||||
@@ -807,6 +807,16 @@ func (in *ScalerConfig) DeepCopyInto(out *ScalerConfig) {
|
||||
*out = new(int)
|
||||
**out = **in
|
||||
}
|
||||
if in.ScaleQPS != nil {
|
||||
in, out := &in.ScaleQPS, &out.ScaleQPS
|
||||
*out = new(int)
|
||||
**out = **in
|
||||
}
|
||||
if in.ScaleBurst != nil {
|
||||
in, out := &in.ScaleBurst, &out.ScaleBurst
|
||||
*out = new(int)
|
||||
**out = **in
|
||||
}
|
||||
if in.Workers != nil {
|
||||
in, out := &in.Workers, &out.Workers
|
||||
*out = new(int)
|
||||
|
||||
Reference in New Issue
Block a user