Let the listener own the EphemeralRunner Running phase transition (#4646)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
This commit is contained in:
Nikola Jokic
2026-09-15 13:46:40 +02:00
committed by GitHub
co-authored by Copilot App Copilot Autofix powered by AI
parent d386789092
commit 9ce3169df3
13 changed files with 610 additions and 67 deletions
+8 -7
View File
@@ -158,13 +158,14 @@ started.
To get a better understanding of health and workings of the cluster
resources, we need to expose the following metrics:
- `pending_ephemeral_runners` - Number of ephemeral runners in a pending state.
This information can show the latency between creating an `EphemeralRunner`
resource, and having an ephemeral runner pod started and ready to receive a
job.
- `running_ephemeral_runners` - Number of ephemeral runners currently running.
This information is helpful to see how many ephemeral runner pods are running
at any given time.
- `pending_ephemeral_runners` - Number of ephemeral runners that have not been
assigned a job yet. This covers both runners whose pod has not finished
starting and runners that are registered and idle, so with a non-zero
`minRunners` it does not drop to zero.
- `running_ephemeral_runners` - Number of ephemeral runners that have been
assigned a job. This information is helpful to see how many ephemeral runners
are executing a workflow job at any given time. It reflects job assignment,
not pod liveness.
- `failed_ephemeral_runners` - Number of ephemeral runners in a `Failed` state.
This information is helpful to catch the faulty image, or some underlying
problem. When the ephemeral runner controller is not able to start the