The `actions.github.com/integrity-hash` annotation was used as an opaque
fingerprint to detect spec drift across AutoscalingRunnerSet,
EphemeralRunnerSet and the listener resources. Hashes are brittle: they
change whenever unrelated serialization details change, they are invisible
to users, and they are not restart-safe. FNV-32a also carries a real
collision risk, where the consequence is an update silently never applied.
Replace it with explicit, typed state:
- `AutoscalingRunnerSetStatus.ObservedGeneration` drives the Pending phase
transition via `metadata.generation` instead of an annotation hash.
- `EphemeralRunnerSetSpec.ActionableRevision` and
`EphemeralRunnerSetStatus.AppliedActionableRevision` form a restart-safe
applied marker. The revision is bumped by the AutoscalingRunnerSet
controller when `EphemeralRunnerSpec` changes, and only advanced in status
after idle/pending runner cleanup succeeds.
- `EphemeralRunnerSetStatus.FinishedRunnerCleanupPatchID` records the
listener patch ID for which finished runners were reaped, so scale-up is
suppressed until the listener publishes a fresh desired state. This
prevents creating a replacement runner for a job that already completed.
- Listener pod recreation compares pod specs semantically instead of
comparing hash annotations.
Drift detection uses `apiequality.Semantic`, not `cmp` or `reflect`:
- `Semantic.DeepEqual` for the EphemeralRunnerSpec. Most PodSpec collection
fields carry `omitempty`, so a template containing an explicitly empty
value (`env: []`) is dropped when the EphemeralRunnerSet is written and
reads back as nil. A strict comparison reports drift on every reconcile,
bumping ActionableRevision each time and deleting every idle and pending
runner, forever. Semantic treats nil and empty as equal, understands
resource.Quantity, and cannot panic on unexported fields the way cmp can.
It is also roughly six times cheaper than cmp.Equal on a realistic spec.
- `Semantic.DeepDerivative` for the listener pod, because the live pod
carries many fields the desired pod never sets (nodeName, dnsPolicy,
default tolerations, the kube-api-access volume, ...). DeepEqual there
would spin in a delete/create loop. Container port length is checked
separately, since ports come from the --listener-metrics-addr flag rather
than from a resource, so disabling metrics would otherwise leave the port
on the pod forever.
Drift detection is deliberately not short-circuited on
metadata.generation. Re-registration changes the runner scale set ID
through an annotation, and metadata changes do not bump generation, so a
generation-based shortcut would leave the EphemeralRunnerSet pointing at a
scale set that no longer exists. The measured saving did not justify the
risk.
Additionally:
- Count deleting runners toward the scale-up total so terminating runners
are not double-replaced.
- Cleanup of finished runners is no longer deferred; failures now surface as
reconcile errors instead of being logged and swallowed.
- Status patches for the new fields use `RetryOnConflict` against a freshly
read object.
- Keep merging EphemeralRunnerSet annotations and labels rather than
overwriting them, so metadata applied by admission webhooks or other
controllers is preserved. Drift detection compares against the merge
result so foreign keys cannot cause a permanent patch loop.
- Add unit tests and benchmarks for both drift checks, including a guard
that fails if the listener comparison is ever tightened to DeepEqual.
- Cover the re-registration path, which previously had no assertion that
the new runner scale set ID reaches the EphemeralRunnerSet at all.
* feat: allow to discover runner statuses
* fix manifests
* Bump runner version to 2.289.1 which includes the hooks support
* Add feedback from review
* Update reference to newRunnerPod
* Fix TestNewRunnerPodFromRunnerController and make hooks file names job specific
* Fix additional TestNewRunnerPod test
* Cover additional feedback from review
* fix rbac manager role
* Add permissions to service account for container mode if not provided
* Rename flag to runner.statusUpdateHook.enabled and fix needsServiceAccount
Co-authored-by: Yusuke Kuoka <ykuoka@gmail.com>
* added containerMode=kubernetes env variables to the runner
* removed unused logging
* restored configs and charts
* restored makefile cert version and acceptance/run
* added workVolumeClaimTemplate in pod definition, including logic
* added claim template name based on the runner
* Apply suggestions from code review
update errors
* added concurrent cleanup before runner pod is deleted
* update manifests
* added retry after 30s if pod cleanup contains err
* added admission webhook check, made workVolumeClaimTemplate mandatory for k8s
* style changes and added comments
* added izZero timestamp check for deleting runner-linked pods
* changed order of local variable to avoid copy if p is deleted
* removed docker from container mode k8s
* restored charts, config, makefile
* restored forked files back and not the ARC ones
* created PersistentVolume on containerMode k8s
* create pv only if storage class name is local-storage
* removed actions if storage class name is local-storage
* added service account validation if container mode kubernetes
* changed the coding style to match rest of the ARC
* added validation to the runnerdeployment webhook
* specified fields more precisely, added webhook validation to the replicaset as well
* remake manifests
* wraped delete runner-linked-pods in kube mode
* fixed empty line
* fixed import
* makefile changes for hooks
* added cleanup secrets
* create manifests
* docs
* update access modes
* update dockerfile
* nit changes
* fixed dockerfile
* rewrite allowing reuse for runners and runnersets
* deepcopy forgot to stage
* changed privileged
* make manifests
* partly moved to finalizer, still need to apply finalizer first
* finalizer added if env variable used in container mode exists
* bump runner version
* error message moved from Error to Info on cleanup pods/secrets
* removed useless dereferencing, added transformation tests of workVolumeClaimTemplate
* Apply suggestions from code review
* Update controllers/utils_test.go
Co-authored-by: Thomas Boop <52323235+thboop@users.noreply.github.com>
* Update controllers/utils_test.go
Co-authored-by: Thomas Boop <52323235+thboop@users.noreply.github.com>
* add hook version to cli, update to 0.1.2
* Apply suggestions from code review
* Update controllers/utils_test.go
* Update runner/Makefile
* Fix missing secret permission and the error handling
* Fix a runnerpod reconciler finalizer to not trigger unnecessary retry
Co-authored-by: Nikola Jokic <nikola-jokic@github.com>
Co-authored-by: Nikola Jokic <97525037+nikola-jokic@users.noreply.github.com>
Co-authored-by: Yusuke Kuoka <ykuoka@gmail.com>
Without that field, GKE 1.21 refuses to create the CRD
with an error message that conversionReviewVersions is mandatory.
conversionReviewVersions is a required field when creating apiextensions.k8s.io/v1 custom resource definitions.
Webhooks are required to support at least one ConversionReview version understood by the current and previous API server.
See https://kubernetes.io/docs/tasks/extend-kubernetes/custom-resources/_print/#webhook-request-and-response