fix: ensure OCI charts are prepared for needed releases with --include-needs (#923)
When using --include-needs with a selector, releases included via needs
must have their charts prepared (pulled/exported) before diff/sync/apply
can process them. The core fix (ChartPrepareOptions.IncludeTransitiveNeeds
= c.IncludeNeeds()) was already in place, but ForEachState calls in
Diff/Template/Lint/Unittest/Sync/Apply still passed c.IncludeTransitiveNeeds()
instead of c.IncludeNeeds(), creating an inconsistency that would resurface
if SetFilter(true) were ever added.
Changes:
- Use c.IncludeNeeds() in ForEachState for all commands supporting
--include-needs (Diff, Template, Lint, Unittest, Sync, Apply, Doctor)
- Doctor is the only command with SetFilter(true), so this fixes a real
bug: helmfile doctor --include-needs was silently ignored for filtering
- Add explanatory doc comment on ForEachState parameter semantics
- Enhance exectest.Helm.ChartPull to create minimal chart files and track
pulls, enabling OCI chart testing
- Add resetChartCacheForTest() for test isolation from global chart cache
- Add regression tests for issue #923
Signed-off-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
Signed-off-by: yxxhero <aiopsclub@163.com>
In nix/devbox environments, helm plugin directories are typically
symlinks into the Nix store. GetPluginVersion used entry.IsDir()
which does not follow symlinks, causing the plugin to be reported
as not installed. Follow symlinks with os.Stat before skipping
non-directory entries.
Signed-off-by: Shane Starcher <shane.starcher@gmail.com>
Co-authored-by: Shane Starcher <shane.starcher@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add `helmfile doctor` command for AI-assisted diff analysis
`helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to
summarize the changes and flag risks (data loss, security exposure, breaking
changes, downtime, performance, best-practice issues).
Key design decisions:
- When no LLM is configured, doctor is equivalent to `helmfile diff` with
one exception: --show-secrets is always forced off (secrets never reach
stdout, even without an LLM).
- Secrets are ALWAYS redacted via two layers: (1) ShowSecrets() forced to
false so helm-diff emits <REDACTED> placeholders; (2) a defense-in-depth
text redactor strips residual secret-looking content (Secret YAML blocks,
sensitive key/value lines, base64 blobs, JWT tokens) before LLM transmission.
- LLM configuration precedence: env (HELMFILE_LLM_*) < helmfile.yaml (llm:)
< CLI flags (--llm-*).
- Supports any OpenAI-compatible backend (OpenAI, Azure, One-API, LiteLLM,
Ollama, etc.) with automatic response_format fallback for backends that
don't support JSON mode.
- Prompt injection defense: release names and environment values are
JSON-encoded before insertion into the LLM prompt.
- Exit codes: 0 (success/low-risk), 2 (high-risk gate, bypass with --force),
1 (other errors). Helm-diff's 'detected changes' exit-2 is swallowed.
New packages:
- pkg/agent/llm: OpenAI-compatible client with JSON response parsing, mock
client for testing, prompt builder with injection defense.
- pkg/agent/doctor: secret redactor (state machine + regex), report renderer
(markdown + JSON), config resolver (env < yaml < flag merge).
Testing: 70+ unit tests covering redaction patterns, prompt injection,
response_format fallback, JSON parsing, yaml roundtrip, concurrency safety,
panic recovery, and error propagation. go test -race passes.
Documentation: full doctor section in docs/cli.md, llm: block reference in
docs/configuration.md, updated skills/helmfile for AI agents.
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: fix doctor equivalence wording per PR review
Per review feedback (PR #2660): the docs claimed doctor is 'equivalent to
helmfile diff — same flags, same output, same exit codes' in the unconfigured
path, but this over-promises because:
1. doctor --output is the report format (not helm-diff's output format)
2. helm-diff's --output is exposed as --diff-output in doctor
3. --show-secrets is silently ignored
Updated all three locations (cli.md, cmd/doctor.go Long + godoc, pkg/app/doctor.go
godoc) to say 'falls back to helmfile diff with --show-secrets forced off' and
explicitly note the --output / --diff-output flag difference.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: parallel kubedog tracking with progress printer and safety valves
Rework kubedog integration so resource tracking runs in parallel with
helm upgrade/install, giving live progress output and recovering from
known helm/kubedog wedge conditions.
Core:
- kubedogTrackingHandle runs tracking in a background goroutine alongside
the helm subprocess (startBackgroundKubedogTracking); helm output is
buffered and replayed as a single block so it no longer interleaves
with progress ticks.
- Capture UID+generation baselines before handing off to helm so each
tracker waits until the resource actually changes (freshness gate).
Progress printer (pkg/kubedog/printer.go):
- Styled, auto-sized progress table with a heartbeat flusher, child (pod)
status roll-up, pre-ready pod-phase handling, multi-namespace support,
and optional color. PreviewBreakdown summarizes kept/filtered resources.
Safety valves (verify cluster state via the live API):
- Tracker-race valve (always on): when helm succeeds but a dyntracker
goroutine is wedged, poll the API and cancel the tracker so wait()
returns success instead of blocking until --track-timeout.
- Helm-stuck killer (opt-in via helmStuckGrace): if the cluster stays
converged while helm v4's hook waiter is wedged, SIGINT the helm
subprocess to recover.
- Failure watchdog (pkg/kubedog/watchdog.go): surface failing pods that
never made it into dyntracker's resource graph.
Options: trackFailedLogs, helmStuckGrace, trackTimeout, color
(Color/NoColor), and resource filtering (trackKinds/skipKinds/
trackResources). PersistentVolumeClaim support in resource classification.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* test: add unit tests for kubedog tracking
Cover the progress printer, resource classification, the failure
watchdog, helm-output trimming/dedup, the release hard-timeout helper,
and the color/track option plumbing.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* test: update golden logs, e2e snapshots, and values-id fixtures
Refresh pkg/app testapply/testdestroy golden logs, e2e template
snapshots, and the TestGenerateID values-id golden hashes for the new
kubedog progress output and the merged release struct layout.
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
* refactor: remove dead kubedog display code and dedupe tracker setup
Deep-review pass on the parallel kubedog tracking changes:
- Remove pkg/kubedog/display.go (308 lines) and display_test.go (453 lines).
These rendered progress for the legacy per-kind trackers that this PR
replaces; in the merged tree every function is unreferenced outside its
own tests. The new progressPrinter (printer.go) supersedes them.
TestMain (color.ForceColor for deterministic ANSI in tests) is preserved
in a new main_test.go.
- Dedupe trackWithKubedog: the post-helm fallback rebuilt the exact same
tracker options as buildReleaseTracker. Reuse buildReleaseTracker instead,
dropping ~40 lines of duplicated timeout/log/filter/tracker construction.
No behavior change; build, go vet, golangci-lint, and the kubedog/state
unit tests all pass.
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: stop waitForFreshness busy-looping the API after helm finishes
Once upstreamDoneCh closes it is always ready, so the select in
waitForFreshness stopped blocking on the ticker and re-ran probe() (a
live GET) as fast as the round-trip allowed for the whole 3s grace
window — hammering the API server once per tracked resource.
Track a local view of the channel and nil it out on first delivery so
the first hit records the timestamp (one fast retry, as intended) and
all subsequent polls are ticker-throttled. Functional behavior is
unchanged: return nil when fresh, errUpstreamDoneNoChange after grace.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: drop trailing blank line from helm4 OCI pull snapshots
The trailing-newline trim in helmexec.info() removes the blank line helm
prints after the OCI chart "Digest:" line. The helm3 (output.yaml)
snapshots never captured that line, but the helm4 (output-helm4.yaml)
snapshots for oci_chart_pull{,_direct,_once,_once2} and
issue_473_oci_chart_url_fetch still expected it, so they failed under
helm 4. Remove the blank line so the snapshots match the trimmed output.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: refresh diff-args integration goldens for styled release headers
DisplayAffectedReleases now emits a styled "========== Updated Releases
==========" header (matching the app/e2e goldens already updated by this
PR) and helmexec.info() trims the trailing blank after helm's install
status. Update the diff-args apply-stderr{,-helm4} and apply-live-
stderr{,-helm4} goldens accordingly so they match the actual stderr.
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: drop trimmed blank line from v1-subhelmfile template golden
The trailing-newline trim in helmexec.info() removes the blank line helm
prints after '"incubator" has been added to your repositories'. Update
the v1-subhelmfile-multi-bases-with-array-values result and result-live
goldens so the template stdout comparison matches.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: Roman Mykhailiuk <romanm@cybellum.com>
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: Roman Mykhailiuk <romanm@cybellum.com>
fix: helmfile deps broken for OCI charts with underscores in path (#954)
For OCI charts with multi-segment paths (e.g., myrepo/path_with_underscores/example),
helmfile was putting the full path as the dependency name in the generated Chart.yaml.
Helm then reconstructed the OCI reference using this name, and underscores in the
path caused issues with helm's OCI reference handling during dependency update.
Fix: move the chart path prefix into the repository URL and use only the chart
basename as the dependency name, matching Helm's recommended Chart.yaml format
for OCI dependencies:
# Before (broken):
dependencies:
- name: path_with_underscores/example
repository: oci://harbor.custom.com
# After (fixed):
dependencies:
- name: example
repository: oci://harbor.custom.com/path_with_underscores
The resulting OCI reference is identical, but the dependency name is now clean.
Includes backward-compatibility fallback for old lock files that used the full
path as the dependency name.
Signed-off-by: yxxhero <aiopsclub@163.com>
Add appendServerSideFlagsForDiff to validate helm-diff plugin version
(v3.15.10+) before passing the --server-side flag to helm diff upgrade.
Extract resolveServerSideValue helper to share precedence logic between
upgrade and diff paths. Bump helm-diff recommended version to v3.15.10
across Dockerfiles, CI, and integration scripts.
Signed-off-by: yxxhero <aiopsclub@163.com>
go-getter v2 (used since v1.4) removed its built-in S3 getter, so helmfile's own AWS-SDK-v2 S3Getter was added to compensate. However the routing only handled the s3://bucket/key form (Getter==normal, Scheme==s3); the go-getter forced-getter vhost form s3::https://bucket.s3.region.amazonaws.com/key fell through to go-getter v2, which can no longer download S3 at all, producing 'error downloading'.
This restores 1.2.x behavior by:
- routing u.Getter==s3 URLs to the built-in S3Getter
- extending ParseS3Url to parse vhost/path-style amazonaws.com URLs (region/bucket/key), modeled on go-getter v1
- stripping the helmfile @<file> selector before deriving the S3 key
- auto-decompressing archive objects (tar.gz/zip/...) via go-getter v2 decompressors so the @<file> selector resolves inside a tarball, as go-getter v1 did
- cleaning up the cache dir on download/decompress failure (matching the GoGetter branch) and avoiding a nil-response panic in GetObject error handling
Fixes#2643
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: add support for helm 4 --server-side upgrade flag
Add support for the helm 4 upgrade flag --server-side which accepts
"true", "false", or "auto" (default "auto"). This allows users to
explicitly control server-side apply behavior, which is needed for
releases originally installed with Helm 3 and being managed with Helm 4.
The flag can be configured via:
- CLI: --server-side flag on sync, apply, and diff commands
- helmDefaults.serverSide in helmfile.yaml
- releases[].serverSide per-release override
Precedence: release-level > CLI flag > helmDefaults.
Errors are returned when serverSide is set but running Helm 3, or when
an invalid value is provided.
Closes#2640
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: update TestGenerateID expected hashes for new ServerSide field
Adding ServerSide *string to ReleaseSpec changes spew's %#v output and
shifts the FNV hash used by generateValuesID. Update the hard-coded want
values to the new deterministic hashes.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
Update helm-diff plugin version from v3.15.8 to v3.15.9 across
Dockerfiles, recommended version constant, CI matrix, and
integration test default.
Signed-off-by: yxxhero <aiopsclub@163.com>
* chore: bump helm release pins
* chore: align helm module metadata
* chore: finalize helm patch bumps
* fix: add --plain-http flag for Helm 3.21+ OCI push in tests
Helm 3.21.1 introduced stricter security checks that reject HTTP
scheme downgrades when pushing to OCI registries, with the error:
"blob upload Location downgrades scheme from https"
Previously only Helm 4 required --plain-http for HTTP-only OCI
registries. Now Helm 3.21+ also requires this flag.
Add a new requiresPlainHTTPForOCI() helper that returns true for
both Helm 4.x and Helm 3.21+, and use it in execHelmPush() instead
of isHelm4().
* fix: safe fallback in requiresPlainHTTPForOCI when version detection fails
Default to true (require --plain-http) when helm version detection
fails, since any Helm version that supports helm push also supports
the --plain-http flag. This avoids the inconsistent HELMFILE_HELM4
env var fallback which only covered Helm 4.
* fix: update snapshot tests for Helm 4.2.1 OCI pull output
Helm 4.2.1 now outputs additional 'Pulled:' and 'Digest: sha256:...'
lines after each OCI chart pull. The SHA256 digest is non-deterministic
because helm packages include build timestamps, so normalize it with
a regex placeholder.
- Add ociDigestRegex to normalize non-deterministic OCI digest values
- Create output-helm4.yaml for 5 tests that lacked Helm 4 snapshots
- Update output-helm4.yaml for oci_need and postrenderer to include
the new Pulled/Digest lines from Helm dependency pull operations
* fix: update ociDigestRegex to match empty digest in Helm 4.2.1 OCI pull output
Helm 4.2.1 outputs "Digest: sha256:" (empty hash) when pulling OCI charts.
The regex required at least one hex char ([0-9a-f]+), so it did not match
and the digest was not normalized to $DIGEST in snapshot tests.
Also fix the replacement string: Go regex ReplaceAllString interprets $DIGEST
as a capture group reference (resolving to empty). Use $$DIGEST to produce
a literal $DIGEST in the output.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: yxxhero <aiopsclub@163.com>
* fix: Fix broken `trackLogs` functionality in Kubedog tracker
When using helmfile with the following setup
```
releases:
- name: my-app
chart: ./mychart
trackMode: kubedog
trackLogs: true
trackFailOnError: true
```
We don't actually see any logs from the container being printed.
This is because when building the options for the Kubedog tracker, we never
specify `SaveLogsOnlyForNumberOfReplicas` which means this defaults to 0.
Looking at the logic in `pkg/tracker/deployment/tracker.go` we see
```
ignoreLogs := job.ignoreLogs || job.savingLogsReplicas >= job.SaveLogsOnlyForNumberOfReplicas
```
With job.SaveLogsOnlyForNumberOfReplicas always defaulting to 0, this will always ignore logs
This change sets it to a reasonable default of tracking logs from up to 10 pods.
Signed-off-by: Graeme Gillies <ggillies@gitlab.com>
* fix formatting
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: Graeme Gillies <ggillies@gitlab.com>
* feat: Refactor out building tracker opts into it's own function, add tests
Signed-off-by: Graeme Gillies <ggillies@gitlab.com>
---------
Signed-off-by: Graeme Gillies <ggillies@gitlab.com>
Co-authored-by: Graeme Gillies <ggillies@gitlab.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Changed SetValue.Values type from []string to []any to allow passing
maps (not just strings) in the values field of set/setTemplate.
Previously, YAML like:
setTemplate:
- name: source.helm.parameters
values:
- name: demo
- version: v2
would fail with 'cannot unmarshal !!map into string'. Map values are
now serialized to JSON when generating --set flags.
Fixes#1021
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: support more HELMFILE_* env vars as flag fallbacks
Adds env-var fallbacks for global flags, mirroring the existing
HELMFILE_ENVIRONMENT / HELMFILE_KUBE_CONTEXT pattern:
* --helm-binary -> HELMFILE_HELM_BINARY
* --kustomize-binary -> HELMFILE_KUSTOMIZE_BINARY
* --log-level -> HELMFILE_LOG_LEVEL
* --debug -> HELMFILE_DEBUG (expecting "true" lower case)
* --quiet -> HELMFILE_QUIET (expecting "true" lower case)
* --no-color -> HELMFILE_NO_COLOR (expecting "true" lower case),
additionally honors NO_COLOR per no-color.org
(any non-empty value disables color)
Flag values still take precedence; env vars are consulted only when the
flag is unset. The string-flag default values ("helm", "kustomize",
"info") move into the accessor methods so the env-var fallback can
actually trigger when no flag is passed.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* docs: mention new HELMFILE_* env vars in cli.md and templating.md
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* fix: make Color/NoColor/env interaction consistent
Two issues with the env-aware NoColor() introduced together with
HELMFILE_NO_COLOR / NO_COLOR support:
1. Color() consulted the raw GlobalOptions.NoColor field instead of
NoColor(), so in a TTY with only the env set, Color() fell through
to terminal autodetect and ValidateConfig() spuriously errored with
"--color and --no-color cannot be specified at the same time".
2. NoColor() returned true via env even when --color was explicitly
passed, so `helmfile --color` with NO_COLOR (or HELMFILE_NO_COLOR=true)
in the environment hit the same ValidateConfig() error. A flag should
always win over an env var.
Fix both by routing Color() through NoColor() and giving NoColor() an
explicit --color short-circuit. Regression tests added for both paths.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
---------
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>