mirror of
https://github.com/zalando/postgres-operator.git
synced 2026-10-09 23:00:18 +02:00
Use maxUnavailable for the critical-op PDB to stop idle alert noise (#3141)
* Use maxUnavailable for the critical-op PDB to stop idle alert noise The critical-op PDB is created with minAvailable equal to numberOfInstances while its selector (critical-operation=true) matches no pods during normal operation. This leaves status.desiredHealthy at N and currentHealthy at 0 permanently, so monitoring stacks fire alerts like kube-prometheus-stack's KubePdbNotEnoughHealthyPods for every idle cluster (#3020). maxUnavailable: 0 provides the same protection while a critical operation is running - no voluntary evictions of labeled pods - but keeps the budget satisfied (desiredHealthy 0) when nothing matches. When PDBs are disabled or there are no instances, the budget relaxes to maxUnavailable 100% instead of minAvailable 0. Fixes #3020 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Update PDB docs for critical-op maxUnavailable semantics Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Clarify why the two PDBs use different budget fields Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Felix Kunde <felix-kunde@gmx.de>
This commit is contained in:
co-authored by
Claude Fable 5
Felix Kunde
parent
d268c589c2
commit
4c1bb1c0ea
+13
-4
@@ -639,9 +639,11 @@ masters in single-node clusters and/or the last remaining running instance in a
|
||||
cluster.
|
||||
|
||||
## PDB for critical operations
|
||||
The `MinAvailable` parameter of this PDB is equal to the `numberOfInstances` set in the
|
||||
cluster manifest, while label selector includes `critical-operation=true` condition. This
|
||||
allows to protect all pods of a cluster, given they are labeled accordingly.
|
||||
The `MaxUnavailable` parameter of this PDB is set to `0`, while label selector includes
|
||||
`critical-operation=true` condition. This blocks voluntary disruptions for all pods of a
|
||||
cluster that are labeled accordingly, without leaving an unsatisfiable budget behind when
|
||||
no pods carry the label (which previously kept monitoring alerts like
|
||||
`KubePdbNotEnoughHealthyPods` firing permanently).
|
||||
For example, Operator labels all Spilo pods with `critical-operation=true` during the major
|
||||
version upgrade run. You may want to protect cluster pods during other critical operations
|
||||
by assigning the label to pods yourself or using other means of automation.
|
||||
@@ -651,7 +653,14 @@ The PDB is only relaxed in two scenarios:
|
||||
* If a cluster is scaled down to `0` instances (e.g. for draining nodes)
|
||||
* If the PDB is disabled in the configuration (`enable_pod_disruption_budget`)
|
||||
|
||||
The PDBs are still in place having `MinAvailable` set to `0`. Disabling PDBs
|
||||
The PDBs are still in place but fully relaxed: the primary PDB with `MinAvailable`
|
||||
set to `0` and the critical operations PDB with `MaxUnavailable` set to `100%`.
|
||||
The two PDBs intentionally use different budget fields matching their purposes:
|
||||
the primary PDB guarantees a minimum count of always-present pods
|
||||
(`MinAvailable`), while the critical operations PDB freezes disruptions for
|
||||
whatever pods currently carry the `critical-operation=true` label - a
|
||||
usually-empty set, which `MaxUnavailable: 0` expresses without producing an
|
||||
unsatisfiable budget while idle. Disabling PDBs
|
||||
helps avoiding blocking Kubernetes upgrades in managed K8s environments at the
|
||||
cost of prolonged DB downtime. See PR [#384](https://github.com/zalando/postgres-operator/pull/384)
|
||||
for the use case.
|
||||
|
||||
Reference in New Issue
Block a user