Every controller in this package woke up on every update of the resources it
owns, including the ones that only touch fields it never reads. An
EphemeralRunner status carries readiness, failure bookkeeping and the job
details written by the listener, a pod reports IPs, its node, a start time and
several conditions, and an EphemeralRunnerSet rewrites
Status.FinishedRunnerCleanupPatchID for every listener patch id. None of that is
input to the owner, yet all of it enqueued a reconcile.
Add an update predicate to each owned watch. A predicate is only allowed to be
an optimisation, so each one is written as the projection of the fields its
reconciler actually reads and drops an update only when all of them are equal:
- AutoscalingRunnerSet -> EphemeralRunnerSet: object metadata, the whole spec,
and the Status.Phase and Status.AppliedActionableRevision pair that
ephemeralRunnerSetOutdatedForAppliedRevision consults.
- EphemeralRunnerSet -> EphemeralRunner: object metadata, spec, Status.Phase
and Status.RunnerID.
- EphemeralRunner -> Pod: UID, deletion timestamp, pod phase, reason and
message, the container and init container statuses, and the Ready
condition.
An event of an unexpected type is always delivered, and create, delete and
generic events are untouched. The primary watches keep seeing every update, so
the status patches a reconciler makes to hand work to its own next pass, such
as marking a runner Failed or Outdated before cleaning up its resources, still
re-enqueue.
The Ready condition lookup is extracted into podReady so that the predicate and
updateRunStatusFromPod cannot disagree about what readiness means.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reconcilers deep copied the object they had just fetched on every single
reconcile, purely so a merge patch could be computed on the rare pass
that actually changes something. The copy is a full recursive walk and
allocation of the object, and the overwhelming majority of reconciles
throw it away untouched.
Introduce lazyCopy, which takes the snapshot on the first call to
Mutate and hands back the live object. Because Mutate is the only way
to reach the object, the snapshot cannot be taken after the mutation it
is supposed to be diffed against, which is the way this optimization is
usually gotten wrong.
Apply it to the four Reconcile entry points, and move the two
EphemeralRunnerSet status copies inside the branch that patches, so they
are only paid for when the status really changed.
While here, drop the short circuit in the EphemeralRunner finalizer
block: `addedFinalizers || AddFinalizer(...)` skipped adding the actions
finalizer whenever the first finalizer was added.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
An EphemeralRunner that exits as Outdated marks its EphemeralRunnerSet
Outdated, and the AutoscalingRunnerSet then stops scaling it. Without a way
to tell which runner spec an Outdated report refers to, a report from a
runner built before the spec was updated keeps the set switched off after
the update that was supposed to fix it.
Stamp each runner with the actionable revision it was built from, and judge
Outdated reports against the revision the set has applied:
- A runner whose revision is older than the applied one is reporting on a
spec that has already been replaced. It is deleted and rebuilt from the
current spec, and it does not hold the set Outdated.
- A runner whose revision is current is reporting on the live spec, so the
set stays Outdated. It is deliberately not delete-and-replaced: a fresh
runner at the same revision would report Outdated again, forever.
Runners without the annotation parse to revision 0, which matches the zero
value of Status.AppliedActionableRevision, so existing runners keep their
current behaviour across an upgrade.
The phase is derived inside patchAppliedActionableRevisionStatus, from a
list read through the same authoritative reader as the set itself, because
the optimistic lock on that patch covers the EphemeralRunnerSet object only
and cannot vouch for a separately-read list. The runners are classified
against the live applied revision rather than the revision this call was
asked to apply: the caller reads the spec from the cache while this function
re-reads the status from the API server, so a lagging reconcile can arrive
with a target behind the live marker, and judging against it would flip a
set that has already moved on back to Outdated.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>