Six specs construct an EphemeralRunnerSetReconciler literal and call
Reconcile directly instead of going through SetupWithManager, so the
APIReader backfill never runs and the field stays nil.
Nothing is broken today: those specs either publish patch ID 0 or return
via the stale-outdated path, so they never reach the scale-up branch
where the reader is consulted. But the next direct-Reconcile spec that
exercises scale-up with a non-zero patch ID would fail with "APIReader is
not configured" instead of the behaviour it meant to test, and the cause
would not be obvious from the failure.
Set the field the way SetupWithManager does in production, so the specs
exercise the same wiring the real controller has. The helper still errors
on a nil reader; falling back to the cached read is the race that fix
exists to prevent.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Port two specs from the source PR that no layer of the stack had picked
up. Both assert behaviour that is already implemented, they were simply
left without coverage:
- the EphemeralRunnerSet, not just the listener, is re-pointed at a
newly registered runner scale set;
- that re-pointing happens with no spec change on the
AutoscalingRunnerSet at all, which is why drift detection cannot be
keyed on metadata.generation: re-registration writes the scale set
ID as an annotation, and metadata changes do not bump generation.
Ported verbatim and passing unmodified.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
An EphemeralRunnerSet goes Outdated when a child runner reports that the
Actions service rejected its runner spec, and the AutoscalingRunnerSet
then tears the listener down so the scale set stops taking jobs. Until
now nothing recorded which runner spec a given Outdated report was about,
so a report from a runner that predates a spec update kept the set
switched off even after the user had already fixed the spec.
Runners now carry the EphemeralRunnerSet's actionable revision as an
annotation, and Outdated runners are split into two groups:
- stale outdated: created before the currently applied revision. Their
verdict is about a spec that no longer exists, so they are deleted
and replaced by runners built from the current spec.
- outdated: created at or after the applied revision. Their verdict is
about the current spec, so they hold the set in the Outdated phase.
They are deliberately not delete-and-replaced: a fresh runner at the
same revision would report Outdated again, forever.
patchAppliedActionableRevisionStatus now recomputes the phase in both
directions against the revision being applied, because it returns from
Reconcile without reaching updateStatus. The AutoscalingRunnerSet
teardown guard additionally requires the applied revision to have caught
up with the spec revision, so it does not act on an Outdated verdict for
a spec that has already been replaced.
terminated() now allocates a fresh slice instead of appending into the
finished slice, which could alias its backing array.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Code review raised two upgrade consequences of no longer stamping
actions.github.com/integrity-hash. Both are real, and neither is covered
by a test, so cover them.
The AutoscalingRunnerSet reconciler compares a live listener's annotations
against the desired ones exactly, so a listener created by an older
controller is replaced on the first reconcile after an upgrade. That
rollout is one-time rather than a loop, because the replacement is built
by the same path and carries no annotation. It is not additional either:
the listener runs the manager's own image, so a controller upgrade already
forces the same recreation through the Spec comparison.
Reconcilers that update objects in place merge live annotations under the
desired ones, so the annotation survives on objects that already carry it.
Nothing reads it any more, so it is inert metadata; stripping it would mean
patching every managed object on upgrade, which is a separate decision.
Both tests were mutation-checked: re-stamping the annotation fails the
first, and stripping the key in mergeAnnotations fails the second.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Layers 2-4 removed every reader of the actions.github.com/integrity-hash
annotation, so what remained were write-only dead paths. Delete the
annotation constant, all remaining hash helpers that computed it, the
call sites that stamped it onto objects, the integrity-hash fallback in
newResourceCacheObjectRef, and the now-unused AutoscalingListenerSpec.Hash
and EphemeralRunnerSpec.Hash methods.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Status.FinishedRunnerCleanupPatchID was only ever set, never cleared, so
a marker recorded under one listener incarnation outlived the patch-ID
sequence it described. Applying a new actionable revision deletes the
idle and pending runners so they are rebuilt from the new spec, and it
restarts the listener. A restarted listener numbers its patches from 0
upwards and counts through every integer, so it does not merely risk
reusing the leftover value, it passes through it. If that reuse lands on
the reconcile that has to refill the pool, the guard suppresses exactly
the scale up the revision change asked for.
Clearing the marker where the applied revision advances is enough,
because the marker only ever means "the gap below Spec.Replicas was made
by cleaning up finished runners for this patch ID", and a spec change
invalidates that claim outright. That read now bypasses the cache: the
patch is a diff against the object that was read, so a cached copy
showing 0 while the server held a marker would emit no entry for the
field and leave the stale value behind.
Clearing the marker also breaks the invariant the cached fast path in
scaleUpServicedByFinishedRunnerCleanup rested on. That path was safe
only because a marker was never removed, so a cache hit could not be a
false positive. Now it can be: a lagging cache can show a marker the
server has already cleared, which suppresses the rebuild. The decision
is therefore always made against an uncached read. That costs one GET,
and only on reconciles that were about to issue creates anyway.
One window remains and is documented rather than papered over. A
listener that restarts without a spec change keeps the marker and still
renumbers from 0, so a collision is still likely. It costs one
suppressed reconcile, not an outage: the listener calls back into
scaling on every long-poll timeout rather than only on change, and an
idle set at its minimum publishes the collapsed patch ID 0, which is
never suppressed.
This was found while investigating the update-gha-runner-scale-set e2e
failure on the tip of the stack. It is a real defect, but it does not
explain that failure, whose cause remains open: the suppression here is
self-correcting within a long-poll cycle rather than terminal.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The scale-up suppression read Status.FinishedRunnerCleanupPatchID off the
EphemeralRunnerSet that Reconcile fetched through the manager's cached
client. That marker is written by the cleanup reconcile, and it is the
runner deletions performed by that same reconcile which trigger the next
one, so the follow-up reconcile is regularly served from an informer cache
that has not yet observed the controller's own status write. The marker
read as 0, the guard did not fire, and the controller created a
replacement runner for a job that had already finished. That is the exact
spurious scale up the guard exists to prevent, so it is a correctness
problem in a cluster and not only a flaky test.
Make the decision from authoritative state. A cached hit still short
circuits, because nothing ever clears the marker, so a hit cannot be a
false positive. Only a miss falls through to an uncached Get via a new
APIReader, which SetupWithManager fills in from mgr.GetAPIReader() so no
construction site can forget it. The extra read is confined to scale-up
decisions, where the controller is about to issue creates anyway.
The window was transient: updateStatus copies the marker from the
in-memory object and patches with MergeFrom, so a stale reconcile produces
no diff for that field and cannot clobber the recorded value.
The envtest spec that caught this only fails under CI load, so the
regression is pinned down directly instead. TestScaleUpServicedByFinished
RunnerCleanup drives the decision with a lagging cached client and an API
reader that already has the write, and asserts the suppression still
holds. Pointing the read back at the cached client fails that case and
only that case.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Spec.Replicas is the count the listener asked for when it published
Spec.PatchID. When the EphemeralRunnerSet controller cleaned up finished
runners in the same reconcile, it then compared that count against a live
count the cleanup had just reduced, and created runners to replace jobs
that had already completed. Nobody asked for those runners.
Delete the finished runners, record the patch ID the cleanup belongs to in
Status.FinishedRunnerCleanupPatchID, and return, so the scaling decision is
made on the next reconcile against post-cleanup data. Scale up stays
suppressed while Spec.PatchID still equals that recorded patch ID: the gap
below Spec.Replicas is the one the cleanup opened, not new demand. The
listener marks itself dirty on every job completion and publishes a fresh
incrementing patch ID, so the suppression lifts as soon as it reports a
desired state that accounts for the completions.
Runners that are mid-deletion were not counted at all, which let the
controller over-create while deletions were still in flight. Count them
towards the scale-up total.
The cleanup helper is renamed to deleteTerminatedEphemeralRunners because
it is no longer specific to finished runners.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
When the AutoscalingRunnerSet's runner spec changes, the EphemeralRunnerSet
has to delete its idle and pending runners so they are rebuilt from the new
spec. That was detected by hashing Spec.EphemeralRunnerSpec into the
actions.github.com/integrity-hash annotation and comparing the annotation
against a freshly computed hash.
Replace it with a monotonic revision counter split across spec and status.
The AutoscalingRunnerSet controller bumps Spec.ActionableRevision when it
patches a new runner spec across; the EphemeralRunnerSet controller advances
Status.AppliedActionableRevision only after the cleanup has actually
succeeded. Because the applied marker lives in status and is written last, a
controller that crashes part-way through the cleanup comes back with the
applied revision still behind the spec revision and redoes the work, instead
of skipping runners that are still running the old spec.
The drift check uses apiequality.Semantic.DeepEqual rather than cmp.Equal or
reflect.DeepEqual. Most PodSpec collection fields carry omitempty, so a
template containing an explicitly empty value (`env: []`) is dropped when the
EphemeralRunnerSet is written and reads back as nil. A strict comparison would
report drift on every reconcile, bump the revision each time, and delete every
idle and pending runner forever. helpers_drift_test.go pins that behaviour with
a round-trip-through-JSON test.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The specs added with the observed-generation work asserted on end states
only. Both the listener rebuild and a stalled reconcile pass through a
transient phase, and nothing looked at it: the tests waited for the
rebuilt listener and then checked for Running. Removing the pending update
that guards the rebuild left them green, which was pointed out in review
and is true — I checked by deleting that update and re-running, and they
passed.
The window was invisible because it closes on its own between two polls.
Hold it open instead: the AutoscalingListener controller does not run in
this suite, so a finalizer added by the test keeps the deleted listener
around until the test removes it. That turns a race into a state the test
can sit on and assert against.
The label spec now waits for the listener to carry a deletion timestamp
while the scale set reports Pending, holds that to show it is not a blip,
then releases the finalizer and checks the listener comes back with a new
UID and the scale set returns to Running. With the pending update removed
it now fails on the phase, which is the point.
The same trick gives the failure path its first coverage. A max runners
change bumps the generation and forces a listener rebuild, so blocking the
rebuild stalls the reconcile part way through. The new spec pins what the
contract claims: the observed generation stays behind the live generation
and the phase stays Pending for as long as the change cannot be applied,
and the marker only catches up once it can.
Also drops a stale comment that survived a merge and contradicted the code
next to it, still claiming label edits no longer reach the Pending phase.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Moving update detection to metadata.generation lost a signal. The desired
listener is built from the AutoscalingRunnerSet's labels and annotations as
well as its spec, and any difference there deletes the listener so it can
be re-created. metadata.generation only moves on spec writes, so a
label-only edit still tore the listener down while the scale set kept
reporting Running. If the rebuild then failed, it reported Running with no
listener indefinitely. Under the old label-inclusive hash the phase did
move, so this was a regression.
Mark the resource pending at the point the listener is deleted rather than
trying to infer it from the generation. That is what the phase already
means: its own doc comment says pending is when the listener is not yet
started. It also covers the case properly, because it keys off the actual
decision to rebuild instead of guessing from the trigger.
This does not change when the listener is rebuilt. A label-only edit
recreates it today and still does; only the reported phase changes. The
earlier claim that labels no longer restart anything was wrong: labels are
propagated to the listener, so editing one is a restart. What
metadata.generation changes is narrower than that, and the test now covers
the listener's identity across a label edit rather than implying it is
untouched.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The AutoscalingRunnerSet controller detected changes by hashing the spec
and the labels on every reconcile and comparing the result against the
actions.github.com/integrity-hash annotation it had written on a previous
pass. That is a hand-rolled version of something the API server already
does for us: it bumps metadata.generation on every spec write, and nothing
else. Recording that value in the status as ObservedGeneration gives the
same "has this changed since I last applied it" signal without recomputing
a hash, without an extra non-status write to the object, and with a value a
user can read and reason about.
updateStatus now takes the observed generation and patches it next to the
phase, returning early only when both are already what we want. When the
generation is ahead of the observed generation we move to Pending but
deliberately leave the observed generation where it is; it only catches up
at the end of a reconcile that actually applied the change, so a reconcile
that fails part way through is retried as pending rather than being
mistaken for settled.
The old code returned immediately after stamping the hash. That was needed
because stamping the annotation was itself a write to the object, so
continuing would have worked from a stale copy. Recording the generation
only touches the status subresource, which does not bump generation, so the
rest of the reconcile can run on the same object and apply the change in
the same pass instead of waiting for the next event.
One user-visible difference: the old hash covered Labels as well as Spec,
and metadata.generation does not move when only labels change. Editing just
a label on an AutoscalingRunnerSet no longer forces it through Pending.
Labels still propagate to the EphemeralRunnerSet; they simply no longer
count as an update that has to be applied.
Hash, ListenerSpecHash and RunnerSetSpecHash are removed. The latter two
already had no callers, and the first has none now.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>