Bump github.com/helmfile/chartify to v0.28.2, which fixes the empty-render
crash upstream (helmfile/chartify#206, fixed in helmfile/chartify#207):
when a chart renders zero resources, chartify now treats it as a no-op
success itself - removing the chart's content dirs, cleaning Chart.yaml
dependencies and the lock file, and skipping the kustomize step - instead
of failing with:
assertion failed: unexpected dir entry "" it must be the abs path to the output directory
That makes the helmfile-side string-matching workaround from #2724 dead
code: the error it matched can no longer be produced. Remove
isChartifyEmptyRenderOutputError and the special-cased no-op branch in
processChartification, so the empty-render case simply flows through the
regular chartify path.
Also bump github.com/helmfile/vals to v0.46.0 and refresh transitive
dependencies.
The regression test for #1757 is updated to assert the new direct
behavior: processChartification succeeds, returns a chartified chart
whose Chart.yaml no longer declares dependencies (so a subsequent
"helm dep build"/"helm template" cannot fail with "found in Chart.yaml,
but missing in charts/ directory"), renders empty output, and is cleaned
up by CleanupChartifyTempDirs.
Fixes#1757 (follow-up to #2724)
Signed-off-by: yxxhero <aiopsclub@163.com>
The '========== Updated Releases ==========' header was a fixed 38
chars regardless of the table width below it. Now the divider line
is extended to match the table's visual width, with the title text
centered within the '=' borders.
- Add TableVisualWidth() to measure table width via runewidth
- Add HeaderDividerCentered() and HeaderDividerCenteredStyled()
for centered dividers with optional bold+blue ANSI styling
- Refactor DisplayAffectedReleases to build the table first, then
compute its width before logging the header
- Update all test snapshots and integration test output files
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: add `helmfile doctor` command for AI-assisted diff analysis
`helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to
summarize the changes and flag risks (data loss, security exposure, breaking
changes, downtime, performance, best-practice issues).
Key design decisions:
- When no LLM is configured, doctor is equivalent to `helmfile diff` with
one exception: --show-secrets is always forced off (secrets never reach
stdout, even without an LLM).
- Secrets are ALWAYS redacted via two layers: (1) ShowSecrets() forced to
false so helm-diff emits <REDACTED> placeholders; (2) a defense-in-depth
text redactor strips residual secret-looking content (Secret YAML blocks,
sensitive key/value lines, base64 blobs, JWT tokens) before LLM transmission.
- LLM configuration precedence: env (HELMFILE_LLM_*) < helmfile.yaml (llm:)
< CLI flags (--llm-*).
- Supports any OpenAI-compatible backend (OpenAI, Azure, One-API, LiteLLM,
Ollama, etc.) with automatic response_format fallback for backends that
don't support JSON mode.
- Prompt injection defense: release names and environment values are
JSON-encoded before insertion into the LLM prompt.
- Exit codes: 0 (success/low-risk), 2 (high-risk gate, bypass with --force),
1 (other errors). Helm-diff's 'detected changes' exit-2 is swallowed.
New packages:
- pkg/agent/llm: OpenAI-compatible client with JSON response parsing, mock
client for testing, prompt builder with injection defense.
- pkg/agent/doctor: secret redactor (state machine + regex), report renderer
(markdown + JSON), config resolver (env < yaml < flag merge).
Testing: 70+ unit tests covering redaction patterns, prompt injection,
response_format fallback, JSON parsing, yaml roundtrip, concurrency safety,
panic recovery, and error propagation. go test -race passes.
Documentation: full doctor section in docs/cli.md, llm: block reference in
docs/configuration.md, updated skills/helmfile for AI agents.
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: fix doctor equivalence wording per PR review
Per review feedback (PR #2660): the docs claimed doctor is 'equivalent to
helmfile diff — same flags, same output, same exit codes' in the unconfigured
path, but this over-promises because:
1. doctor --output is the report format (not helm-diff's output format)
2. helm-diff's --output is exposed as --diff-output in doctor
3. --show-secrets is silently ignored
Updated all three locations (cli.md, cmd/doctor.go Long + godoc, pkg/app/doctor.go
godoc) to say 'falls back to helmfile diff with --show-secrets forced off' and
explicitly note the --output / --diff-output flag difference.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: parallel kubedog tracking with progress printer and safety valves
Rework kubedog integration so resource tracking runs in parallel with
helm upgrade/install, giving live progress output and recovering from
known helm/kubedog wedge conditions.
Core:
- kubedogTrackingHandle runs tracking in a background goroutine alongside
the helm subprocess (startBackgroundKubedogTracking); helm output is
buffered and replayed as a single block so it no longer interleaves
with progress ticks.
- Capture UID+generation baselines before handing off to helm so each
tracker waits until the resource actually changes (freshness gate).
Progress printer (pkg/kubedog/printer.go):
- Styled, auto-sized progress table with a heartbeat flusher, child (pod)
status roll-up, pre-ready pod-phase handling, multi-namespace support,
and optional color. PreviewBreakdown summarizes kept/filtered resources.
Safety valves (verify cluster state via the live API):
- Tracker-race valve (always on): when helm succeeds but a dyntracker
goroutine is wedged, poll the API and cancel the tracker so wait()
returns success instead of blocking until --track-timeout.
- Helm-stuck killer (opt-in via helmStuckGrace): if the cluster stays
converged while helm v4's hook waiter is wedged, SIGINT the helm
subprocess to recover.
- Failure watchdog (pkg/kubedog/watchdog.go): surface failing pods that
never made it into dyntracker's resource graph.
Options: trackFailedLogs, helmStuckGrace, trackTimeout, color
(Color/NoColor), and resource filtering (trackKinds/skipKinds/
trackResources). PersistentVolumeClaim support in resource classification.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* test: add unit tests for kubedog tracking
Cover the progress printer, resource classification, the failure
watchdog, helm-output trimming/dedup, the release hard-timeout helper,
and the color/track option plumbing.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* test: update golden logs, e2e snapshots, and values-id fixtures
Refresh pkg/app testapply/testdestroy golden logs, e2e template
snapshots, and the TestGenerateID values-id golden hashes for the new
kubedog progress output and the merged release struct layout.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* refactor: remove dead kubedog display code and dedupe tracker setup
Deep-review pass on the parallel kubedog tracking changes:
- Remove pkg/kubedog/display.go (308 lines) and display_test.go (453 lines).
These rendered progress for the legacy per-kind trackers that this PR
replaces; in the merged tree every function is unreferenced outside its
own tests. The new progressPrinter (printer.go) supersedes them.
TestMain (color.ForceColor for deterministic ANSI in tests) is preserved
in a new main_test.go.
- Dedupe trackWithKubedog: the post-helm fallback rebuilt the exact same
tracker options as buildReleaseTracker. Reuse buildReleaseTracker instead,
dropping ~40 lines of duplicated timeout/log/filter/tracker construction.
No behavior change; build, go vet, golangci-lint, and the kubedog/state
unit tests all pass.
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: stop waitForFreshness busy-looping the API after helm finishes
Once upstreamDoneCh closes it is always ready, so the select in
waitForFreshness stopped blocking on the ticker and re-ran probe() (a
live GET) as fast as the round-trip allowed for the whole 3s grace
window — hammering the API server once per tracked resource.
Track a local view of the channel and nil it out on first delivery so
the first hit records the timestamp (one fast retry, as intended) and
all subsequent polls are ticker-throttled. Functional behavior is
unchanged: return nil when fresh, errUpstreamDoneNoChange after grace.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: drop trailing blank line from helm4 OCI pull snapshots
The trailing-newline trim in helmexec.info() removes the blank line helm
prints after the OCI chart "Digest:" line. The helm3 (output.yaml)
snapshots never captured that line, but the helm4 (output-helm4.yaml)
snapshots for oci_chart_pull{,_direct,_once,_once2} and
issue_473_oci_chart_url_fetch still expected it, so they failed under
helm 4. Remove the blank line so the snapshots match the trimmed output.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: refresh diff-args integration goldens for styled release headers
DisplayAffectedReleases now emits a styled "========== Updated Releases
==========" header (matching the app/e2e goldens already updated by this
PR) and helmexec.info() trims the trailing blank after helm's install
status. Update the diff-args apply-stderr{,-helm4} and apply-live-
stderr{,-helm4} goldens accordingly so they match the actual stderr.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: drop trimmed blank line from v1-subhelmfile template golden
The trailing-newline trim in helmexec.info() removes the blank line helm
prints after '"incubator" has been added to your repositories'. Update
the v1-subhelmfile-multi-bases-with-array-values result and result-live
goldens so the template stdout comparison matches.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: Roman Mykhailiuk <romanm@cybellum.com>
* chore: bump helm release pins
* chore: align helm module metadata
* chore: finalize helm patch bumps
* fix: add --plain-http flag for Helm 3.21+ OCI push in tests
Helm 3.21.1 introduced stricter security checks that reject HTTP
scheme downgrades when pushing to OCI registries, with the error:
"blob upload Location downgrades scheme from https"
Previously only Helm 4 required --plain-http for HTTP-only OCI
registries. Now Helm 3.21+ also requires this flag.
Add a new requiresPlainHTTPForOCI() helper that returns true for
both Helm 4.x and Helm 3.21+, and use it in execHelmPush() instead
of isHelm4().
* fix: safe fallback in requiresPlainHTTPForOCI when version detection fails
Default to true (require --plain-http) when helm version detection
fails, since any Helm version that supports helm push also supports
the --plain-http flag. This avoids the inconsistent HELMFILE_HELM4
env var fallback which only covered Helm 4.
* fix: update snapshot tests for Helm 4.2.1 OCI pull output
Helm 4.2.1 now outputs additional 'Pulled:' and 'Digest: sha256:...'
lines after each OCI chart pull. The SHA256 digest is non-deterministic
because helm packages include build timestamps, so normalize it with
a regex placeholder.
- Add ociDigestRegex to normalize non-deterministic OCI digest values
- Create output-helm4.yaml for 5 tests that lacked Helm 4 snapshots
- Update output-helm4.yaml for oci_need and postrenderer to include
the new Pulled/Digest lines from Helm dependency pull operations
* fix: update ociDigestRegex to match empty digest in Helm 4.2.1 OCI pull output
Helm 4.2.1 outputs "Digest: sha256:" (empty hash) when pulling OCI charts.
The regex required at least one hex char ([0-9a-f]+), so it did not match
and the digest was not normalized to $DIGEST in snapshot tests.
Also fix the replacement string: Go regex ReplaceAllString interprets $DIGEST
as a capture group reference (resolving to empty). Use $$DIGEST to produce
a literal $DIGEST in the output.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: yxxhero <aiopsclub@163.com>