mirror of
https://github.com/zalando/postgres-operator.git
synced 2026-10-03 15:30:07 +02:00
Retry moveMasterPodsOffNode on failure instead of aborting (#3179)
attemptToMoveMasterPodsOffNode's transient errors (e.g. no synced standby available yet) were treated as terminal, so the retry loop gave up after a single ~4-5s attempt instead of retrying over master_pod_move_timeout. This left masters stuck on draining nodes indefinitely once a PodDisruptionBudget started rejecting evictions. Log the failure and return (false, nil) so the retry loop keeps polling every minute until the timeout expires. Fixes #3176
This commit is contained in:
@@ -556,9 +556,10 @@ configuration they are grouped under the `kubernetes` key.
|
||||
* **master_pod_move_timeout**
|
||||
The period of time to wait for the success of migration of master pods from
|
||||
an unschedulable node. The migration includes Patroni switchovers to
|
||||
respective replicas on healthy nodes. The situation where master pods still
|
||||
exist on the old node after this timeout expires has to be fixed manually.
|
||||
The default is 20 minutes.
|
||||
respective replicas on healthy nodes. A failed migration attempt is retried
|
||||
every minute until this timeout expires. The situation where master pods
|
||||
still exist on the old node after this timeout expires has to be fixed
|
||||
manually. The default is 20 minutes.
|
||||
|
||||
* **enable_pod_antiaffinity**
|
||||
toggles [pod anti affinity](https://kubernetes.io/docs/concepts/configuration/assign-pod-node/)
|
||||
|
||||
Reference in New Issue
Block a user