mirror of
https://github.com/helmfile/helmfile.git
synced 2026-10-09 00:07:10 +02:00
* feat: parallel kubedog tracking with progress printer and safety valves Rework kubedog integration so resource tracking runs in parallel with helm upgrade/install, giving live progress output and recovering from known helm/kubedog wedge conditions. Core: - kubedogTrackingHandle runs tracking in a background goroutine alongside the helm subprocess (startBackgroundKubedogTracking); helm output is buffered and replayed as a single block so it no longer interleaves with progress ticks. - Capture UID+generation baselines before handing off to helm so each tracker waits until the resource actually changes (freshness gate). Progress printer (pkg/kubedog/printer.go): - Styled, auto-sized progress table with a heartbeat flusher, child (pod) status roll-up, pre-ready pod-phase handling, multi-namespace support, and optional color. PreviewBreakdown summarizes kept/filtered resources. Safety valves (verify cluster state via the live API): - Tracker-race valve (always on): when helm succeeds but a dyntracker goroutine is wedged, poll the API and cancel the tracker so wait() returns success instead of blocking until --track-timeout. - Helm-stuck killer (opt-in via helmStuckGrace): if the cluster stays converged while helm v4's hook waiter is wedged, SIGINT the helm subprocess to recover. - Failure watchdog (pkg/kubedog/watchdog.go): surface failing pods that never made it into dyntracker's resource graph. Options: trackFailedLogs, helmStuckGrace, trackTimeout, color (Color/NoColor), and resource filtering (trackKinds/skipKinds/ trackResources). PersistentVolumeClaim support in resource classification. Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com> * test: add unit tests for kubedog tracking Cover the progress printer, resource classification, the failure watchdog, helm-output trimming/dedup, the release hard-timeout helper, and the color/track option plumbing. Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com> * test: update golden logs, e2e snapshots, and values-id fixtures Refresh pkg/app testapply/testdestroy golden logs, e2e template snapshots, and the TestGenerateID values-id golden hashes for the new kubedog progress output and the merged release struct layout. Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com> * refactor: remove dead kubedog display code and dedupe tracker setup Deep-review pass on the parallel kubedog tracking changes: - Remove pkg/kubedog/display.go (308 lines) and display_test.go (453 lines). These rendered progress for the legacy per-kind trackers that this PR replaces; in the merged tree every function is unreferenced outside its own tests. The new progressPrinter (printer.go) supersedes them. TestMain (color.ForceColor for deterministic ANSI in tests) is preserved in a new main_test.go. - Dedupe trackWithKubedog: the post-helm fallback rebuilt the exact same tracker options as buildReleaseTracker. Reuse buildReleaseTracker instead, dropping ~40 lines of duplicated timeout/log/filter/tracker construction. No behavior change; build, go vet, golangci-lint, and the kubedog/state unit tests all pass. Signed-off-by: yxxhero <aiopsclub@163.com> * fix: stop waitForFreshness busy-looping the API after helm finishes Once upstreamDoneCh closes it is always ready, so the select in waitForFreshness stopped blocking on the ticker and re-ran probe() (a live GET) as fast as the round-trip allowed for the whole 3s grace window — hammering the API server once per tracked resource. Track a local view of the channel and nil it out on first delivery so the first hit records the timestamp (one fast retry, as intended) and all subsequent polls are ticker-throttled. Functional behavior is unchanged: return nil when fresh, errUpstreamDoneNoChange after grace. Signed-off-by: yxxhero <aiopsclub@163.com> * test: drop trailing blank line from helm4 OCI pull snapshots The trailing-newline trim in helmexec.info() removes the blank line helm prints after the OCI chart "Digest:" line. The helm3 (output.yaml) snapshots never captured that line, but the helm4 (output-helm4.yaml) snapshots for oci_chart_pull{,_direct,_once,_once2} and issue_473_oci_chart_url_fetch still expected it, so they failed under helm 4. Remove the blank line so the snapshots match the trimmed output. Signed-off-by: yxxhero <aiopsclub@163.com> * test: refresh diff-args integration goldens for styled release headers DisplayAffectedReleases now emits a styled "========== Updated Releases ==========" header (matching the app/e2e goldens already updated by this PR) and helmexec.info() trims the trailing blank after helm's install status. Update the diff-args apply-stderr{,-helm4} and apply-live- stderr{,-helm4} goldens accordingly so they match the actual stderr. Signed-off-by: yxxhero <aiopsclub@163.com> * test: drop trimmed blank line from v1-subhelmfile template golden The trailing-newline trim in helmexec.info() removes the blank line helm prints after '"incubator" has been added to your repositories'. Update the v1-subhelmfile-multi-bases-with-array-values result and result-live goldens so the template stdout comparison matches. Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com> Signed-off-by: yxxhero <aiopsclub@163.com> Co-authored-by: Roman Mykhailiuk <romanm@cybellum.com>
501 lines
18 KiB
Go
501 lines
18 KiB
Go
package kubedog
|
|
|
|
import (
|
|
"bytes"
|
|
"context"
|
|
"strings"
|
|
"sync"
|
|
"testing"
|
|
|
|
"github.com/stretchr/testify/assert"
|
|
"github.com/stretchr/testify/require"
|
|
"github.com/werf/kubedog/pkg/trackers/dyntracker/statestore"
|
|
kdutil "github.com/werf/kubedog/pkg/trackers/dyntracker/util"
|
|
"go.uber.org/zap"
|
|
"go.uber.org/zap/zapcore"
|
|
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
|
|
"k8s.io/apimachinery/pkg/runtime"
|
|
"k8s.io/apimachinery/pkg/runtime/schema"
|
|
"k8s.io/client-go/dynamic/fake"
|
|
|
|
"github.com/helmfile/helmfile/pkg/resource"
|
|
)
|
|
|
|
func TestDetectPodFailureReason(t *testing.T) {
|
|
tests := []struct {
|
|
name string
|
|
obj map[string]any
|
|
want string
|
|
}{
|
|
{
|
|
name: "phase Failed",
|
|
obj: map[string]any{"status": map[string]any{"phase": "Failed"}},
|
|
want: "Failed",
|
|
},
|
|
{
|
|
name: "container waiting CrashLoopBackOff",
|
|
obj: map[string]any{"status": map[string]any{
|
|
"phase": "Running",
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
}},
|
|
want: "CrashLoopBackOff",
|
|
},
|
|
{
|
|
name: "container waiting ImagePullBackOff",
|
|
obj: map[string]any{"status": map[string]any{
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "ImagePullBackOff"},
|
|
}},
|
|
},
|
|
}},
|
|
want: "ImagePullBackOff",
|
|
},
|
|
{
|
|
name: "container terminated OOMKilled",
|
|
obj: map[string]any{"status": map[string]any{
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"terminated": map[string]any{"reason": "OOMKilled"},
|
|
}},
|
|
},
|
|
}},
|
|
want: "OOMKilled",
|
|
},
|
|
{
|
|
name: "init container failing surfaces with init: prefix",
|
|
obj: map[string]any{"status": map[string]any{
|
|
"initContainerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
}},
|
|
want: "init: CrashLoopBackOff",
|
|
},
|
|
{
|
|
name: "running pod with no waiting/terminated reason — healthy",
|
|
obj: map[string]any{"status": map[string]any{
|
|
"phase": "Running",
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"running": map[string]any{},
|
|
}},
|
|
},
|
|
}},
|
|
want: "",
|
|
},
|
|
{
|
|
name: "empty status block",
|
|
obj: map[string]any{},
|
|
want: "",
|
|
},
|
|
}
|
|
for _, tc := range tests {
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
got := detectPodFailureReason(&unstructured.Unstructured{Object: tc.obj})
|
|
assert.Equal(t, tc.want, got)
|
|
})
|
|
}
|
|
}
|
|
|
|
func TestExtractPodSelector(t *testing.T) {
|
|
t.Run("matchLabels round-trips through SelectorFromSet", func(t *testing.T) {
|
|
obj := &unstructured.Unstructured{Object: map[string]any{
|
|
"spec": map[string]any{
|
|
"selector": map[string]any{
|
|
"matchLabels": map[string]any{
|
|
"app": "malware", "component": "worker",
|
|
},
|
|
},
|
|
},
|
|
}}
|
|
got := extractPodSelector(obj)
|
|
// SelectorFromSet sorts keys, so the output is deterministic.
|
|
assert.Equal(t, "app=malware,component=worker", got)
|
|
})
|
|
t.Run("missing selector returns empty string", func(t *testing.T) {
|
|
assert.Equal(t, "", extractPodSelector(&unstructured.Unstructured{Object: map[string]any{}}))
|
|
})
|
|
t.Run("empty matchLabels returns empty string", func(t *testing.T) {
|
|
obj := &unstructured.Unstructured{Object: map[string]any{
|
|
"spec": map[string]any{
|
|
"selector": map[string]any{
|
|
"matchLabels": map[string]any{},
|
|
},
|
|
},
|
|
}}
|
|
assert.Equal(t, "", extractPodSelector(obj))
|
|
})
|
|
}
|
|
|
|
func TestPodIsInTaskStore(t *testing.T) {
|
|
// Build a task store with one Deployment that has one tracked Pod child.
|
|
deployGVK := schema.GroupVersionKind{Group: "apps", Version: "v1", Kind: "Deployment"}
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
ts := statestore.NewReadinessTaskState("app", "ns", deployGVK, statestore.ReadinessTaskStateOptions{})
|
|
ts.AddResourceState("tracked-pod", "ns", watchdogPodGVK)
|
|
ts.AddDependency(ts.Name(), ts.Namespace(), ts.GroupVersionKind(), "tracked-pod", "ns", watchdogPodGVK)
|
|
taskStore.RWTransaction(func(s *statestore.TaskStore) {
|
|
s.AddReadinessTaskState(kdutil.NewConcurrent(ts))
|
|
})
|
|
|
|
assert.True(t, podIsInTaskStore(taskStore, "tracked-pod", "ns"),
|
|
"tracked pod must be found in the task store")
|
|
assert.False(t, podIsInTaskStore(taskStore, "untracked-pod", "ns"),
|
|
"pod the watchdog will need to surface must NOT be found in the task store")
|
|
// Different namespace must not match a same-named pod.
|
|
assert.False(t, podIsInTaskStore(taskStore, "tracked-pod", "other-ns"),
|
|
"pod lookup must be namespace-scoped")
|
|
}
|
|
|
|
// newCaptureLogger returns a SugaredLogger whose Warn-level writes are
|
|
// captured into a buffer — sufficient to verify the watchdog actually emits
|
|
// the expected warning text.
|
|
func newCaptureLogger(t *testing.T) (*zap.SugaredLogger, *bytes.Buffer) {
|
|
t.Helper()
|
|
buf := &bytes.Buffer{}
|
|
var cfg zapcore.EncoderConfig
|
|
cfg.MessageKey = "message"
|
|
core := zapcore.NewCore(zapcore.NewConsoleEncoder(cfg), zapcore.AddSync(&captureWriter{buf: buf, mu: &sync.Mutex{}}), zapcore.WarnLevel)
|
|
return zap.New(core).Sugar(), buf
|
|
}
|
|
|
|
type captureWriter struct {
|
|
buf *bytes.Buffer
|
|
mu *sync.Mutex
|
|
}
|
|
|
|
func (w *captureWriter) Write(p []byte) (int, error) {
|
|
w.mu.Lock()
|
|
defer w.mu.Unlock()
|
|
return w.buf.Write(p)
|
|
}
|
|
|
|
// newWatchdogFakeClient registers the list kinds the watchdog actually
|
|
// needs (Pods, Jobs, Deployments, ReplicaSets). The fake dynamic client
|
|
// panics on List for unregistered list kinds, so this helper is shared by
|
|
// the integration-style tests below.
|
|
func newWatchdogFakeClient(t *testing.T, objs ...runtime.Object) *fake.FakeDynamicClient {
|
|
t.Helper()
|
|
scheme := runtime.NewScheme()
|
|
for _, kind := range []schema.GroupVersionKind{
|
|
{Group: "", Version: "v1", Kind: "PodList"},
|
|
{Group: "batch", Version: "v1", Kind: "JobList"},
|
|
{Group: "apps", Version: "v1", Kind: "DeploymentList"},
|
|
{Group: "apps", Version: "v1", Kind: "ReplicaSetList"},
|
|
} {
|
|
scheme.AddKnownTypeWithName(kind, &unstructured.UnstructuredList{})
|
|
}
|
|
return fake.NewSimpleDynamicClient(scheme, objs...)
|
|
}
|
|
|
|
func TestScanForMissedFailures_WarnsForUntrackedFailingPod(t *testing.T) {
|
|
// Cluster contains a Job and one failing pod owned by it (the one
|
|
// dyntracker missed linking). Task store has the Job but no Pod
|
|
// children — simulating the linkage race the watchdog is built to
|
|
// mitigate. Using a Job keeps the ownership chain direct (Pod → Job).
|
|
jobGVK := schema.GroupVersionKind{Group: "batch", Version: "v1", Kind: "Job"}
|
|
jobGVR := schema.GroupVersionResource{Group: "batch", Version: "v1", Resource: "jobs"}
|
|
const jobUID = "11111111-1111-1111-1111-111111111111"
|
|
|
|
jobObj := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "batch/v1",
|
|
"kind": "Job",
|
|
"metadata": map[string]any{"name": "malware", "namespace": "ns", "uid": jobUID},
|
|
"spec": map[string]any{
|
|
"selector": map[string]any{"matchLabels": map[string]any{"app": "malware"}},
|
|
},
|
|
}}
|
|
failingPod := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "v1",
|
|
"kind": "Pod",
|
|
"metadata": map[string]any{
|
|
"name": "malware-pod-untracked",
|
|
"namespace": "ns",
|
|
"labels": map[string]any{"app": "malware"},
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "batch/v1", "kind": "Job", "name": "malware", "uid": jobUID, "controller": true},
|
|
},
|
|
},
|
|
"status": map[string]any{
|
|
"phase": "Running",
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
},
|
|
}}
|
|
|
|
logger, buf := newCaptureLogger(t)
|
|
tr := &Tracker{
|
|
logger: logger,
|
|
dynamicClient: newWatchdogFakeClient(t, jobObj, failingPod),
|
|
mapper: &staticRESTMapper{mappings: map[schema.GroupVersionKind]schema.GroupVersionResource{
|
|
jobGVK: jobGVR,
|
|
watchdogPodGVK: watchdogPodGVR,
|
|
}},
|
|
trackOptions: &TrackOptions{},
|
|
}
|
|
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
warned := map[string]struct{}{}
|
|
tr.scanForMissedFailures(context.Background(), taskStore, []watchdogWorkload{
|
|
{kind: "job", gvk: jobGVK, name: "malware", namespace: "ns"},
|
|
}, warned)
|
|
|
|
out := buf.String()
|
|
require.Contains(t, out, "malware-pod-untracked", "watchdog must name the missing failing pod")
|
|
require.Contains(t, out, "CrashLoopBackOff", "watchdog must name the failure reason")
|
|
require.Contains(t, out, "kubectl", "watchdog must give the operator an actionable inspection command")
|
|
}
|
|
|
|
func TestScanForMissedFailures_IgnoresStalePodFromPreviousInstall(t *testing.T) {
|
|
// Reproduces the "stale failing pod" false positive: the current Job
|
|
// has its own UID, but a leftover failing pod from the previous install
|
|
// (with a different owner UID) still matches the label selector. The
|
|
// watchdog must skip it — that's not "our" pod and helm will clean it
|
|
// up shortly anyway.
|
|
jobGVK := schema.GroupVersionKind{Group: "batch", Version: "v1", Kind: "Job"}
|
|
jobGVR := schema.GroupVersionResource{Group: "batch", Version: "v1", Resource: "jobs"}
|
|
const currentJobUID = "22222222-2222-2222-2222-222222222222"
|
|
const previousJobUID = "33333333-3333-3333-3333-333333333333"
|
|
|
|
currentJob := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "batch/v1",
|
|
"kind": "Job",
|
|
"metadata": map[string]any{"name": "feeds-db-insert-init", "namespace": "ns", "uid": currentJobUID},
|
|
"spec": map[string]any{
|
|
"selector": map[string]any{"matchLabels": map[string]any{"app": "feeds-db-insert-init"}},
|
|
},
|
|
}}
|
|
stalePod := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "v1",
|
|
"kind": "Pod",
|
|
"metadata": map[string]any{
|
|
"name": "feeds-db-insert-init-cbxvc",
|
|
"namespace": "ns",
|
|
"labels": map[string]any{"app": "feeds-db-insert-init"},
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "batch/v1", "kind": "Job", "name": "feeds-db-insert-init", "uid": previousJobUID, "controller": true},
|
|
},
|
|
},
|
|
"status": map[string]any{"phase": "Failed"},
|
|
}}
|
|
|
|
logger, buf := newCaptureLogger(t)
|
|
tr := &Tracker{
|
|
logger: logger,
|
|
dynamicClient: newWatchdogFakeClient(t, currentJob, stalePod),
|
|
mapper: &staticRESTMapper{mappings: map[schema.GroupVersionKind]schema.GroupVersionResource{
|
|
jobGVK: jobGVR,
|
|
watchdogPodGVK: watchdogPodGVR,
|
|
}},
|
|
trackOptions: &TrackOptions{},
|
|
}
|
|
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
warned := map[string]struct{}{}
|
|
tr.scanForMissedFailures(context.Background(), taskStore, []watchdogWorkload{
|
|
{kind: "job", gvk: jobGVK, name: "feeds-db-insert-init", namespace: "ns"},
|
|
}, warned)
|
|
|
|
assert.NotContains(t, buf.String(), "feeds-db-insert-init-cbxvc",
|
|
"watchdog must NOT warn about a failing pod left over from a previous install (different owner UID)")
|
|
}
|
|
|
|
func TestScanForMissedFailures_StaysQuietWhenDyntrackerAlreadyHasPod(t *testing.T) {
|
|
// Task store DOES contain the failing pod, so dyntracker is already
|
|
// surfacing it via the normal pipeline. Watchdog stays silent.
|
|
jobGVK := schema.GroupVersionKind{Group: "batch", Version: "v1", Kind: "Job"}
|
|
jobGVR := schema.GroupVersionResource{Group: "batch", Version: "v1", Resource: "jobs"}
|
|
const jobUID = "44444444-4444-4444-4444-444444444444"
|
|
|
|
jobObj := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "batch/v1",
|
|
"kind": "Job",
|
|
"metadata": map[string]any{"name": "app", "namespace": "ns", "uid": jobUID},
|
|
"spec": map[string]any{"selector": map[string]any{"matchLabels": map[string]any{"app": "app"}}},
|
|
}}
|
|
failingPod := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "v1",
|
|
"kind": "Pod",
|
|
"metadata": map[string]any{
|
|
"name": "app-pod", "namespace": "ns",
|
|
"labels": map[string]any{"app": "app"},
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "batch/v1", "kind": "Job", "name": "app", "uid": jobUID, "controller": true},
|
|
},
|
|
},
|
|
"status": map[string]any{
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
},
|
|
}}
|
|
|
|
logger, buf := newCaptureLogger(t)
|
|
tr := &Tracker{
|
|
logger: logger,
|
|
dynamicClient: newWatchdogFakeClient(t, jobObj, failingPod),
|
|
mapper: &staticRESTMapper{mappings: map[schema.GroupVersionKind]schema.GroupVersionResource{
|
|
jobGVK: jobGVR,
|
|
watchdogPodGVK: watchdogPodGVR,
|
|
}},
|
|
trackOptions: &TrackOptions{},
|
|
}
|
|
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
ts := statestore.NewReadinessTaskState("app", "ns", jobGVK, statestore.ReadinessTaskStateOptions{})
|
|
ts.AddResourceState("app-pod", "ns", watchdogPodGVK)
|
|
ts.AddDependency(ts.Name(), ts.Namespace(), ts.GroupVersionKind(), "app-pod", "ns", watchdogPodGVK)
|
|
taskStore.RWTransaction(func(s *statestore.TaskStore) {
|
|
s.AddReadinessTaskState(kdutil.NewConcurrent(ts))
|
|
})
|
|
|
|
warned := map[string]struct{}{}
|
|
tr.scanForMissedFailures(context.Background(), taskStore, []watchdogWorkload{
|
|
{kind: "job", gvk: jobGVK, name: "app", namespace: "ns"},
|
|
}, warned)
|
|
|
|
assert.NotContains(t, buf.String(), "watchdog", "watchdog must not warn when dyntracker is already tracking the failing pod")
|
|
}
|
|
|
|
func TestScanForMissedFailures_DoesNotRepeatWarningsAcrossScans(t *testing.T) {
|
|
jobGVK := schema.GroupVersionKind{Group: "batch", Version: "v1", Kind: "Job"}
|
|
jobGVR := schema.GroupVersionResource{Group: "batch", Version: "v1", Resource: "jobs"}
|
|
const jobUID = "55555555-5555-5555-5555-555555555555"
|
|
|
|
jobObj := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "batch/v1",
|
|
"kind": "Job",
|
|
"metadata": map[string]any{"name": "malware", "namespace": "ns", "uid": jobUID},
|
|
"spec": map[string]any{"selector": map[string]any{"matchLabels": map[string]any{"app": "malware"}}},
|
|
}}
|
|
failingPod := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "v1",
|
|
"kind": "Pod",
|
|
"metadata": map[string]any{
|
|
"name": "malware-pod-untracked", "namespace": "ns",
|
|
"labels": map[string]any{"app": "malware"},
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "batch/v1", "kind": "Job", "name": "malware", "uid": jobUID, "controller": true},
|
|
},
|
|
},
|
|
"status": map[string]any{
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
},
|
|
}}
|
|
|
|
logger, buf := newCaptureLogger(t)
|
|
tr := &Tracker{
|
|
logger: logger,
|
|
dynamicClient: newWatchdogFakeClient(t, jobObj, failingPod),
|
|
mapper: &staticRESTMapper{mappings: map[schema.GroupVersionKind]schema.GroupVersionResource{
|
|
jobGVK: jobGVR,
|
|
watchdogPodGVK: watchdogPodGVR,
|
|
}},
|
|
trackOptions: &TrackOptions{},
|
|
}
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
warned := map[string]struct{}{}
|
|
workloads := []watchdogWorkload{{kind: "job", gvk: jobGVK, name: "malware", namespace: "ns"}}
|
|
tr.scanForMissedFailures(context.Background(), taskStore, workloads, warned)
|
|
tr.scanForMissedFailures(context.Background(), taskStore, workloads, warned)
|
|
|
|
count := strings.Count(buf.String(), "[watchdog]")
|
|
assert.Equal(t, 1, count, "watchdog must warn at most once per pod across consecutive scans")
|
|
}
|
|
|
|
func TestScanForMissedFailures_DeploymentPodMatchesViaReplicaSet(t *testing.T) {
|
|
// Deployment pods are not directly owned by the workload; the watchdog
|
|
// must walk Pod → ReplicaSet → Deployment to recognize a failing pod
|
|
// as "ours".
|
|
deployGVK := schema.GroupVersionKind{Group: "apps", Version: "v1", Kind: "Deployment"}
|
|
deployGVR := schema.GroupVersionResource{Group: "apps", Version: "v1", Resource: "deployments"}
|
|
const deployUID = "66666666-6666-6666-6666-666666666666"
|
|
const rsUID = "77777777-7777-7777-7777-777777777777"
|
|
|
|
deploy := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "apps/v1",
|
|
"kind": "Deployment",
|
|
"metadata": map[string]any{"name": "app", "namespace": "ns", "uid": deployUID},
|
|
"spec": map[string]any{
|
|
"selector": map[string]any{"matchLabels": map[string]any{"app": "app"}},
|
|
"replicas": int64(1),
|
|
},
|
|
}}
|
|
rs := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "apps/v1",
|
|
"kind": "ReplicaSet",
|
|
"metadata": map[string]any{
|
|
"name": "app-abc", "namespace": "ns", "uid": rsUID,
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "apps/v1", "kind": "Deployment", "name": "app", "uid": deployUID, "controller": true},
|
|
},
|
|
},
|
|
}}
|
|
failingPod := &unstructured.Unstructured{Object: map[string]any{
|
|
"apiVersion": "v1",
|
|
"kind": "Pod",
|
|
"metadata": map[string]any{
|
|
"name": "app-abc-xyz", "namespace": "ns",
|
|
"labels": map[string]any{"app": "app"},
|
|
"ownerReferences": []any{
|
|
map[string]any{"apiVersion": "apps/v1", "kind": "ReplicaSet", "name": "app-abc", "uid": rsUID, "controller": true},
|
|
},
|
|
},
|
|
"status": map[string]any{
|
|
"containerStatuses": []any{
|
|
map[string]any{"state": map[string]any{
|
|
"waiting": map[string]any{"reason": "CrashLoopBackOff"},
|
|
}},
|
|
},
|
|
},
|
|
}}
|
|
|
|
logger, buf := newCaptureLogger(t)
|
|
tr := &Tracker{
|
|
logger: logger,
|
|
dynamicClient: newWatchdogFakeClient(t, deploy, rs, failingPod),
|
|
mapper: &staticRESTMapper{mappings: map[schema.GroupVersionKind]schema.GroupVersionResource{
|
|
deployGVK: deployGVR,
|
|
watchdogPodGVK: watchdogPodGVR,
|
|
}},
|
|
trackOptions: &TrackOptions{},
|
|
}
|
|
taskStore := kdutil.NewConcurrent(statestore.NewTaskStore())
|
|
warned := map[string]struct{}{}
|
|
tr.scanForMissedFailures(context.Background(), taskStore, []watchdogWorkload{
|
|
{kind: "deploy", gvk: deployGVK, name: "app", namespace: "ns"},
|
|
}, warned)
|
|
|
|
assert.Contains(t, buf.String(), "app-abc-xyz",
|
|
"watchdog must recognize a failing Deployment pod via the Pod → RS → Deployment chain")
|
|
}
|
|
|
|
func TestWatchdogWorkloads_FiltersUntrackableKinds(t *testing.T) {
|
|
in := []*resource.Resource{
|
|
{Kind: "Deployment", Name: "d", Namespace: "ns"},
|
|
{Kind: "ConfigMap", Name: "cm", Namespace: "ns"}, // not tracked at all
|
|
{Kind: "PersistentVolumeClaim", Name: "p", Namespace: "ns"}, // tracked but no pods
|
|
{Kind: "Job", Name: "j", Namespace: "ns"},
|
|
{Kind: "Canary", Name: "c", Namespace: "ns"}, // tracked but complex pod set; skip
|
|
}
|
|
got := watchdogWorkloads(in)
|
|
require.Len(t, got, 2)
|
|
assert.Equal(t, "deploy", got[0].kind)
|
|
assert.Equal(t, "job", got[1].kind)
|
|
}
|