Commit Graph
4 Commits
Author SHA1 Message Date
Nikola JokicandCopilot App f22def86e0 Recreate the listener pod when its config secret changes
Replacing the integrity hash with a pod spec comparison lost the one
signal the spec cannot carry. The listener mounts its config as a secret
volume and parses it once at startup, so a change to the scale set URL,
the TLS certificate, the metrics configuration or the scaler tuning only
reaches the listener after a restart. The pod references the secret by
name, so the spec is byte-identical before and after and the pod was
never recreated.

The desired pod now carries the config secret's resource version as an
annotation, which is free to read and moves exactly when the secret is
written. An empty annotation on the live pod is ignored so that pods
created by an older controller are not all recreated on upgrade.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-09 10:41:31 +02:00
Nikola JokicandCopilot App 41d05325c2 Make the Outdated phase revision-aware
An outdated runner used to flip the whole scale set into the Outdated
phase permanently, which tears down the listener and switches the scale
set off. That verdict outlived the runner spec it was about: a runner
busy with a job survives the revision cleanup that follows a spec
update, and only reports Outdated once the job finishes. The result was
that a freshly applied fix could be discarded by a runner that never ran
it.

Runners are now stamped with the actionable revision they were built
from. An Outdated runner whose revision is behind the applied revision is
considered stale: it is deleted so the scaling logic replaces it with one
built from the current spec, and it no longer contributes to the set's
phase. A runner at the current revision still marks the set Outdated, so
a genuinely bad spec is still surfaced.

Two supporting fixes:

  - patchAppliedActionableRevisionStatus now recomputes the phase in both
    directions. It only ever forced Running, so the early-return path
    could leave a stale Outdated behind.

  - The AutoscalingRunnerSet only tears down on an Outdated set once that
    set has applied its current actionable revision. Otherwise a spec
    update races the EphemeralRunnerSet controller and the teardown fires
    against a phase that predates the update.

Runners created before this change parse to revision 0, which matches
the zero value of AppliedActionableRevision, so they are treated as
current until a revision is actually applied.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-08 17:27:39 +02:00
Nikola JokicandCopilot App 95179fbf54 Remove dead blank-identifier block in helpers.go
Both ephemeralRunnerSetActionableSpecChanged and nextActionableRevision are
referenced from autoscalingrunnerset_controller.go, so the blank assignments
that suppressed unused-function warnings are no longer needed.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-08 13:17:19 +02:00
Nikola Jokic e481f69cff Replace integrity-hash annotations with explicit API state fields
The `actions.github.com/integrity-hash` annotation was used as an opaque
fingerprint to detect spec drift across AutoscalingRunnerSet,
EphemeralRunnerSet and the listener resources. Hashes are brittle: they
change whenever unrelated serialization details change, they are invisible
to users, and they are not restart-safe. FNV-32a also carries a real
collision risk, where the consequence is an update silently never applied.

Replace it with explicit, typed state:

- `AutoscalingRunnerSetStatus.ObservedGeneration` drives the Pending phase
  transition via `metadata.generation` instead of an annotation hash.
- `EphemeralRunnerSetSpec.ActionableRevision` and
  `EphemeralRunnerSetStatus.AppliedActionableRevision` form a restart-safe
  applied marker. The revision is bumped by the AutoscalingRunnerSet
  controller when `EphemeralRunnerSpec` changes, and only advanced in status
  after idle/pending runner cleanup succeeds.
- `EphemeralRunnerSetStatus.FinishedRunnerCleanupPatchID` records the
  listener patch ID for which finished runners were reaped, so scale-up is
  suppressed until the listener publishes a fresh desired state. This
  prevents creating a replacement runner for a job that already completed.
- Listener pod recreation compares pod specs semantically instead of
  comparing hash annotations.

Drift detection uses `apiequality.Semantic`, not `cmp` or `reflect`:

- `Semantic.DeepEqual` for the EphemeralRunnerSpec. Most PodSpec collection
  fields carry `omitempty`, so a template containing an explicitly empty
  value (`env: []`) is dropped when the EphemeralRunnerSet is written and
  reads back as nil. A strict comparison reports drift on every reconcile,
  bumping ActionableRevision each time and deleting every idle and pending
  runner, forever. Semantic treats nil and empty as equal, understands
  resource.Quantity, and cannot panic on unexported fields the way cmp can.
  It is also roughly six times cheaper than cmp.Equal on a realistic spec.
- `Semantic.DeepDerivative` for the listener pod, because the live pod
  carries many fields the desired pod never sets (nodeName, dnsPolicy,
  default tolerations, the kube-api-access volume, ...). DeepEqual there
  would spin in a delete/create loop. Container port length is checked
  separately, since ports come from the --listener-metrics-addr flag rather
  than from a resource, so disabling metrics would otherwise leave the port
  on the pod forever.

Drift detection is deliberately not short-circuited on
metadata.generation. Re-registration changes the runner scale set ID
through an annotation, and metadata changes do not bump generation, so a
generation-based shortcut would leave the EphemeralRunnerSet pointing at a
scale set that no longer exists. The measured saving did not justify the
risk.

Additionally:

- Count deleting runners toward the scale-up total so terminating runners
  are not double-replaced.
- Cleanup of finished runners is no longer deferred; failures now surface as
  reconcile errors instead of being logged and swallowed.
- Status patches for the new fields use `RetryOnConflict` against a freshly
  read object.
- Keep merging EphemeralRunnerSet annotations and labels rather than
  overwriting them, so metadata applied by admission webhooks or other
  controllers is preserved. Drift detection compares against the merge
  result so foreign keys cannot cause a permanent patch loop.
- Add unit tests and benchmarks for both drift checks, including a guard
  that fails if the listener comparison is ever tightened to DeepEqual.
- Cover the re-registration path, which previously had no assertion that
  the new runner scale set ID reaches the EphemeralRunnerSet at all.
2026-09-07 22:08:17 +02:00