Let the listener own the EphemeralRunner Running phase transition (#4646)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
This commit is contained in:
Nikola Jokic
2026-09-15 13:46:40 +02:00
committed by GitHub
co-authored by Copilot App Copilot Autofix powered by AI
parent d386789092
commit 9ce3169df3
13 changed files with 610 additions and 67 deletions
+8 -7
View File
@@ -158,13 +158,14 @@ started.
To get a better understanding of health and workings of the cluster
resources, we need to expose the following metrics:
- `pending_ephemeral_runners` - Number of ephemeral runners in a pending state.
This information can show the latency between creating an `EphemeralRunner`
resource, and having an ephemeral runner pod started and ready to receive a
job.
- `running_ephemeral_runners` - Number of ephemeral runners currently running.
This information is helpful to see how many ephemeral runner pods are running
at any given time.
- `pending_ephemeral_runners` - Number of ephemeral runners that have not been
assigned a job yet. This covers both runners whose pod has not finished
starting and runners that are registered and idle, so with a non-zero
`minRunners` it does not drop to zero.
- `running_ephemeral_runners` - Number of ephemeral runners that have been
assigned a job. This information is helpful to see how many ephemeral runners
are executing a workflow job at any given time. It reflects job assignment,
not pod liveness.
- `failed_ephemeral_runners` - Number of ephemeral runners in a `Failed` state.
This information is helpful to catch the faulty image, or some underlying
problem. When the ephemeral runner controller is not able to start the
@@ -66,7 +66,7 @@ The dashboard includes the following metrics:
| Running Jobs | The number of runners that are currently processing jobs. |
| Failed Runners | The total number of ephemeral runners that have failed to properly start. This may require reviewing the custom resource and logs to identify and resolve the root causes. Common causes include resource issues and failure to pull the required image. |
| Listeners | The number of listeners currently running and attempting to manage jobs for the scale set. This should match the number of scale sets deployed. |
| Pending Runners | The total number of ephemeral runners that ARC has requested and is waiting for Kubernetes to provide in a running state. If the Kubernetes API server is responsive, this will typically match the number of runner pods that are in a pending state. This number includes requests for runner pods that have not yet been scheduled. When this number is higher than the number of runner pods in a pending state, it can indicate performance issues. |
| Pending Runners | The total number of ephemeral runners that have not been assigned a job. This covers runners that Kubernetes has not finished starting as well as runners that are registered and idle waiting for work, so with a non-zero `minRunners` it does not fall to zero. On its own it is therefore not a signal of scheduling trouble; compare it against the number of runner pods actually in a pending state, and treat a persistent excess of pending runners over pending pods as the indicator of performance issues. |
| Registered Runners | The total number of ephemeral runners that have been successfully registered. |
| Active Runners | The total number of runners that are active and either available or processing jobs. |
| Out of Memory | The number of containers that have been terminated by the OOMKiller. This can indicate that the requests/ limits for one or more pods on the node were configured improperly, allowing pods to request more memory than the node had available. |