* feat(telemetry): add opt-in OpenTelemetry tracing (PR 1: lifecycle + root span)
Implements the first increment of docs/proposals/otel-tracing.md (#2767):
- pkg/telemetry: SDK setup from standard OTEL_* env vars (autoexport for
exporter selection, env-driven sampler/propagators, OTEL_SDK_DISABLED),
command-span lifecycle, no-op-by-default accessors
- --otel-tracing flag / HELMFILE_OTEL_TRACING env switch
- root span "helmfile <command>" with file/environment/selectors/exit_code
attributes; TRACEPARENT-based remote-parent extraction for CI correlation
- shutdown flush on both normal-exit and signal paths (nil-safe, 5s bound)
- app.New derives its context from telemetry.CommandContext()
(Background-identical when tracing is disabled)
- docs: otel.md user guide, experimental-features entry, design proposal
- tests: hermetic unit tests, app context-contract pinning, flag registration
Telemetry problems never fail a run: exporter misconfiguration and export
errors degrade to disabled with a warning. When disabled, behavior and
performance are identical to before (no-op tracer, no goroutines, no
network).
Refs: #2767, #2758
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): trace every external process + trace-context bridges (PR 2)
Implements the second increment of docs/proposals/otel-tracing.md (#2767):
- pkg/helmexec/span.go: one span per external process started by helmfile
(helm invocations, hooks, plugin execs) at the ShellRunner choke point —
helm.exec (with helm.subcommand) vs os.exec, with redacted exec.args,
exec.exit_code, and error status on failure
- pkg/helmexec/redact.go: shared argument redaction with two profiles;
legacy is byte-identical to the historical exit-error behavior (existing
goldens unchanged), strict (spans) additionally covers --set=k=v and
credential flags; exit_error.go now uses the shared helper
- orphan-trace bridges with bit-identical cancellation semantics
(context.WithoutCancel of the command context): both kubedog call sites
(state.go) and hook execution (event.Bus gains an optional Ctx consumed
by its default runner; state.go sets it, nil falls back to TODO as before)
- OTLP end-to-end test (in-process httptest receiver, no external
collector): span export, error status/exit code, redaction, and
parent-linkage to the command span
- docs/otel.md updated to the now-traced surface
Verified end-to-end with the console exporter: helmfile template on a
local chart yields the command span plus helm.exec spans for helm
version/dependency/template, all nested under it.
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): state-loading and hook spans (PR 3a)
Implements the third increment of docs/proposals/otel-tracing.md (#2767):
- helmfile.discover_states around findDesiredStateFiles and helmfile.load
around loadDesiredStateFromYamlWithBaseDir; both cover all callers
(incl. nested helmfiles) with no signature changes
- helmfile.render / helmfile.parse children per document part, parented
through a traceCtx field on the unexported desiredStateLoader struct
(set once at its single construction site)
- helmfile.hook span per hook execution: Trigger's per-hook body extracted
into runHook (readability win on its own), the hook's subprocess span
nests under it via a per-hook ctx-swapped ShellRunner clone
(cancellation unchanged — Bus.Ctx never carries cancellation by contract)
- pkg/telemetry/otlptest: shared in-process OTLP/HTTP receiver harness,
now used by helmexec, event, and app span tests
- golden span-tree test at the app layer (root -> discover -> load ->
render/parse, via the exectest fake helm) and a hook-span nesting test
- nil-ctx guard for App literals built directly by tests (App.spanParentCtx)
Verified end-to-end with the console exporter: a template run over a
gotmpl state file with a prepare hook yields the full tree with the hook's
os.exec nested under helmfile.hook.
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): per-release spans nested under the load span (PR 3b)
Implements the per-release increment of docs/proposals/otel-tracing.md
(#2767) — spans nest command -> load -> release -> helm exec:
- helmexec.HelmContext gains an optional Ctx carrying the per-release span
context; the execer's new execWithContext funnel consumes it via a
per-call runner clone (runnerWithCtx) so the shared, cached execer is
never mutated across concurrent workers. The seven Interface methods
that take a HelmContext (Sync/Diff/ReleaseStatus/List/DecryptSecret/
Delete/Test) route through it; nil Ctx behaves exactly as before.
exec() lost its always-nil override parameter on the way (unparam).
- pkg/state/span.go: SetTraceContext + startReleaseSpan/endReleaseSpan
helpers (release/namespace/chart/labels attributes, sorted for stable
output); a typed-nil guard (releaseErrAsError) avoids the classic
nil-pointer-in-interface trap on *ReleaseError.
- release spans in the worker loops: SyncReleases, DiffReleases,
DeleteReleasesForSync, PrepareCharts, and iterateOnReleases (status/
delete/test via a new verb parameter); their HelmContext is stamped with
the release span context where one is built.
- pkg/app sets st.SetTraceContext(loadCtx) right after loading a state
file, rooting all per-release spans under helmfile.load.
- bridged one more detached tracking call found on the way
(trackReleaseIfEnabled's context.Background in the sync worker).
- golden test: release span present, nested under load, correct
attributes; unit test for the runnerWithCtx clone semantics.
Verified with the console exporter: helmfile template yields
release.prepare(demo) under load with full attributes.
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): nest status/delete/test execs under their release spans
Completes the per-release exec nesting for the iterateOnReleases-based
loops (docs/proposals/otel-tracing.md §4.4 phase 2): the do closures now
receive the release span context and stamp it into their HelmContext, so
helm status/delete/test subprocess spans nest under
helmfile.release.<verb> like sync/diff already did.
- scatterGatherReleases/iterateOnReleases/doWithReleaseSpan: do gains a
context parameter (the release span context)
- ReleaseStatuses/DeleteReleases/TestReleases closures stamp
HelmContext.Ctx from it
- integration test with a real execer (version-probe shim binary): the
release's status subprocess nests under helmfile.release.status, same
trace, with helm.subcommand=status
This also makes the otel.md claim ("upgrade, diff, delete, status, test
nested under the release span") fully accurate.
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): OTel metrics — helm exec duration and release results (PR 4)
Implements the metrics increment of docs/proposals/otel-tracing.md
(#2767) on the same provider, switch, and resource as traces:
- pkg/telemetry/metrics.go: helmfile.helm.exec.duration histogram
(subcommand, success) and helmfile.release.count counter (verb,
result). Instruments come from the otel global meter, so recording at
call sites is branch-free no-op when telemetry is disabled.
- Setup builds the resource once and installs both providers; reader
selection delegates to autoexport (OTEL_METRICS_EXPORTER: otlp |
console | prometheus | none), the OTLP reader's interval honors
OTEL_METRIC_EXPORT_INTERVAL (read by the SDK). Shutdown flushes both
providers (errors.Join). StartCommandSpan now carries the meter
provider across state transitions (fixes a nil-shutdown panic).
- helmexec: finishExecSpan records exec duration for helm binaries;
state: endReleaseSpan counts release outcomes for sync/diff/delete/
status/test/prepare (diff counted as success when no hard error).
- otlptest: recorder routes by OTLP path (/v1/traces vs /v1/metrics)
and decodes metrics; new FindMetric helper.
- tests: metrics recorded as no-op when disabled, provider enabled with
the none exporter, degradation on an invalid metrics exporter, and an
integration assertion (status exec duration datapoint + one successful
release.count) in the shim-based state test.
Verified with the console exporter: helmfile template emits
helmfile.helm.exec.duration per subcommand (version/dependency/
template) and helmfile.release.count{verb=prepare,result=success}=1.
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: complete OTel documentation coverage
- docs/cli.md: --otel-tracing in the CLI reference help block (verbatim
from the cobra output)
- CHANGELOG.md: [Unreleased] Added entry for tracing + metrics
- docs/index.md: Observability highlight linking docs/otel.md
- docs/proposals/otel-tracing.md: add OTEL_METRICS_EXPORTER /
OTEL_METRIC_EXPORT_INTERVAL rows to the env-var table and note the
periodic reader + bounded metric cardinality in §7
Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: drop unused id parameter from parsePart (unparam)
The id parameter was never used inside the span wrapper; the caller's id
variable is still used for the render calls and error messages.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): address review — redaction gaps, kubedog valve, phantom metrics
Addresses all Copilot review comments on #2769:
Security (span payloads):
- exec.args: positional arguments are additionally passed through
helmexec.RedactedURL, so credentials embedded in chart/repository URLs
(AddRepo, RegistryLogin, OCI refs) are masked exactly like log output
- release spans sanitize helmfile.chart the same way
- error statuses no longer embed raw errors (which contain rendered
commands, arguments, and subprocess output): the command span, release
spans, hook spans, and exec spans now use generic descriptions; the
concrete exit code remains an attribute, and RecordError on the root
span is dropped
Correctness:
- kubedog safety valve restored: execWithContext now attaches the
per-release span into the runner's own context instead of replacing it,
so trackHandle.Cancel() can interrupt a wedged helm again and app
cancellation semantics stay exactly as before the PR
- diff release spans/metrics: real failures are recorded (exit code 2
"changes detected" still counts as success); previously every diff was
exported as successful
- skipped releases no longer emit phantom spans and inflate
helmfile.release.count: iterateOnReleases callers pass a skip predicate
(skipUndesired for status/test; delete deletes undesired releases and
passes nil)
- Setup shuts down the already-constructed tracer provider (bounded) when
the metrics provider fails, instead of abandoning its batch goroutine
Tests: URL redaction cases (masked/untouched), spanAttachedContext
preserves the runner cancellation chain while attaching the caller's
span, skipUndesired, and a failing-hook span asserting the generic
message.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: lint — restore nolint placement and avoid nil context literal
- the skipUndesired insertion had displaced the // nolint: unparam
directive off iterateOnReleases (helm param is intentionally unused
there); also fixes a skipDesired/skipUndesired comment typo
- use a typed nil in TestSpanAttachedContext (staticcheck SA1012)
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): address review round 2 — remote-ref redaction, wrapper helm binaries, hook release attribution
Addresses all 6 new review comments on #2769:
Security (remote references):
- new helmexec.RedactedRef sanitizes go-getter style references for
telemetry: forced-form prefixes (git::, s3::) preserved, whole URL
userinfo masked (usernames carry tokens too), credential-bearing query
parameters masked using pkg/remote's heuristic (token/password/secret/
key/signature). Applied to helmfile.file (command span), helmfile.path
(discover_states), helmfile.chart (release spans), and exec.args —
log-time RedactedURL is untouched so log output is unchanged
Correctness:
- wrapper helm binaries (--helm-binary custom names) are now classified
as helm operations by an explicit context marker stamped in the execer
funnel, instead of the executable-basename heuristic; the same
classification gates helmfile.helm.exec.duration, so the metric no
longer misses wrapper invocations (classifyExec)
- release-scoped hooks (presync/postsync/preuninstall/postuninstall/
cleanup in the sync/delete/diff workers) now attach their helmfile.hook
spans to the active helmfile.release.* span via a variadic parent on
the trigger functions; global hooks keep the command context and all
29 existing call sites compile unchanged; hook cancellation stays
detached (WithoutCancel) as before
- signal-terminated runs (Shutdown with exitCode 130/143 and nil error)
now mark the command span with error status, consistent with their
nonzero exit code
Tests: RedactedRef table (forced forms, userinfo, s3/token query params,
untouched cases), classifyExec marker case, hookTraceContext parent
attribution + non-cancellability + fallback.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): address review round 3 — redaction corner cases, value runners
Addresses 5 of the 6 new review comments on #2769 (the sixth — an
unused strings import in exit_error.go — is a false positive: Indent
still uses strings.Split/Builder and the package compiles):
- RedactArgs read the previous token from the progressively redacted
output, so {--set, --set-string, secret} leaked the secret
(the masked value hid the following flag). Read the previous token
from the original input, restoring the legacy contract for adjacent
secret flags
- RedactedRef fails closed for malformed references: URL-like refs with
invalid percent escapes export a fully redacted value, and an
unparseable query is dropped entirely instead of exported verbatim
- ShellRunner has value receivers, so a ShellRunner VALUE satisfies the
Runner API; the helm marker stamping and the per-release span
attachment now handle both value and pointer forms (matching
WithContext), so value-runner callers keep release nesting and the
helm.exec classification/metric
Regression tests: adjacent secret flags (legacy + strict), malformed
URL-like ref, malformed query, value-runner marker + span attachment.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): stamp the helm marker on the stdin funnel too
execStdIn (registry login, repo add) called the runner directly, so
wrapper --helm-binary names were misclassified as os.exec and omitted
from helmfile.helm.exec.duration on that path. The marking now goes
through a shared markHelmRunner helper (value and pointer ShellRunner
forms) used by both execution funnels.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): redact helm's --kube-token in strict profile
Helm's global --kube-token carries a bearer token; both the
two-argument and inline forms are now masked in span exec.args
(legacy exit-error output is untouched, matching its historical
behavior).
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix(telemetry): OTel metrics best-practice alignment
- helmfile.helm.exec.duration now declares explicit bucket boundaries
tuned for seconds-scale helm invocations (5ms…600s); the SDK defaults
are millisecond-oriented and lumped every sub-5s invocation — the
common case — into the first bucket, defeating the histogram
- instruments are re-created under the installed provider with the
instrumentation scope version stamped (Setup-time, race-free)
- helmfile.release.count declares the {release} curly-annotation unit
per the metrics naming conventions
Tested end-to-end via the OTLP integration test: exported bounds are
the tuned set, units are asserted, and the scope carries the version.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat(telemetry): per-release duration metrics behind an opt-in switch
New helmfile.release.duration histogram (seconds, same tuned buckets)
with bounded dimensions by default (verb, result). Setting
HELMFILE_OTEL_METRICS_PER_RELEASE=true adds helmfile.release and
helmfile.namespace, answering "which release is slow" from dashboards:
- well-suited to bounded CI runs; long-lived centralized collection
needs a backend capacity/TTL story (documented in docs/otel.md)
- per-release timing remains available in traces without the flag
- env read per call (release operations are low-frequency, and tests
toggle it)
endReleaseSpan now takes the release and the operation start time; the
five worker-loop call sites pass them (doWithReleaseSpan, SyncReleases,
DeleteReleasesForSync, PrepareCharts, DiffReleases).
Verified end-to-end with the console exporter (default dims vs
per-release) and OTLP integration tests pinning both modes.
Refs: #2767, #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor(telemetry): maintainability pass over the runner/metric plumbing
- StartCommandSpan copies the tracingState struct instead of enumerating
fields by hand — that pattern dropped the meter provider once already
- the two value/pointer ShellRunner switches (helm marker, span
attachment) are unified into one withRunnerCtx helper; the duplication
caused two review rounds of value-form misses
- classifyExec derives the helm classification from the span name
(helmExecSpanName constant) instead of returning a third parallel bool
- metrics: shared outcomeAttrs for the verb/result dimensions, and the
bucket slice renamed to durationBuckets with a comment covering both
histograms that use it
No behavior change; full -race suite green, lint clean.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor(telemetry): consolidate test env lists, trace bridges, and hook prep; sync the design doc
Maintainability:
- HermeticEnvVars is now exported from pkg/telemetry (the owner of the
env surface) and used by both telemetry tests and otlptest — the two
copies had already drifted once (HELMFILE_OTEL_METRICS_PER_RELEASE
needed updating in both)
- kubedogTraceContext and hookTraceContext were the same concept written
twice; unified into traceOnlyContext(parent...) in span.go
Readability:
- runHook's nested kubectl rewrite extracted into prepareKubectlHook
with guard-clause structure
Accuracy (docs ↔ code, drifted over five review rounds):
- §4.4 now describes the implemented mechanism: the release span is
INJECTED into the runner's own context (preserving the kubedog safety
valve) rather than the runner context being replaced, and helm
classification is marker-based for wrapper binaries
- §5 exec span rows list the actual attributes incl. URL/query masking
- §6 strict profile documents RedactedRef, --kube-token, and the
adjacent-token guarantee
No behavior change; full -race suite green, lint clean.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor(telemetry): drop the dead noop state, relocate skipUndesired, sync user-facing accuracy
- tracingState.noop was dead weight in the enabled state and a
copy-surface in every transition; a single package-level
noopTracerProvider now backs Tracer while disabled
- skipUndesired moved next to doWithReleaseSpan in span.go, its only
conceptual home (span/metric suppression, not run plumbing)
- accuracy: the package doc, --otel-tracing flag help,
experimental-features entry, and CHANGELOG now all say tracing AND
metrics and list the third instrument (helmfile.release.duration with
the HELMFILE_OTEL_METRICS_PER_RELEASE opt-in) — these had drifted
when the metric was added; the PR description's metric table is
updated to match as well
No behavior change; full -race suite green (except the pre-existing
network-dependent TestStorage_resolveFile flake), lint clean.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor(telemetry): flatten Setup, name the prefix bound, dedupe test fake; fix instrument-count drift
Readability/maintainability:
- Setup drops from 56 to 39 lines: provider construction (including the
shutdown-tracer-on-meter-failure recovery) moves to newProviders in
exporter.go next to the constructors it composes
- refredact's magic 16 becomes maxForcedFormPrefix with a comment
- span_test's hand-rolled fakeRunner removed in favor of the existing
mockRunner (same package)
Accuracy:
- "Two instruments" wording survived in docs/otel.md and the design
proposal §7 after helmfile.release.duration was added; both now say
three and mention the per-release opt-in
No behavior change; full -race suite green (except the pre-existing
network flake), lint clean.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor(telemetry): co-locate span machinery, drop a dead export, fix docs nits
- the span plumbing helpers (markHelmExec, withRunnerCtx,
markHelmRunner, spanAttachedContext) move from exec.go to span.go,
next to the marker type and classifiers they serve — exec.go keeps
only the funnel call sites
- otlptest.SpanNames was never used outside the package; unexported
- isHelmBinary's comment now states it is the FALLBACK classifier
(funnel invocations are marker-classified), replacing the outdated
"cosmetic distinction" framing from before the marker existed
- docs/otel.md: release-scoped hooks nest under their release span
(added in review round 2, never documented)
No behavior change; full -race suite green, lint clean.
Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: skip subhelmfiles when selectors conflict with CLI selectors (#2544)
When CLI selectors are provided (e.g. -l name=b), subhelmfiles whose
explicit selectors are provably incompatible are now skipped entirely,
avoiding unnecessary YAML loading and template rendering.
Two selector sets are incompatible when every pair has a positive label
conflict: same key with different values (e.g. name=b vs name=a).
Negative labels are not compared.
Fixes#2544
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: address PR review comments - use CLI selectors, fix doc comment, add malformed selector test
Agent-Logs-Url: https://github.com/helmfile/helmfile/sessions/1f1c33ce-e50d-4781-85b8-d606b5d4ca54
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* fix: add debug logging, unit tests, docs, and fix integration test for subhelmfile selector skip
- Add debug log when skipping subhelmfile due to selector conflict
- Add TestSubhelmfileSelectorsConflict with 11 cases for direct unit coverage
- Document the selector-based subhelmfile skip optimization in docs/index.md
- Fix integration test: use 'app' label key instead of reserved 'name' key
(GetReleasesWithLabels overwrites labels["name"] with the release name)
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor: avoid map allocation in positiveLabelsCompatibleWith
Compare positive label slices directly instead of allocating a map per
comparison, as label counts are typically small (1-3 entries).
Addresses Copilot review comment on PR #2545.
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: clarify subhelmfile selector docs per Copilot review feedback
Reword the first two bullets to avoid the contradiction between
'CLI selectors are ignored' and the new skip optimization.
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: address Copilot review comments round 3
- Log parse errors from SelectorsAreCompatible at debug level instead of
silently discarding them
- Hoist regex compilation to package-level vars in ParseLabels to avoid
repeated compilation per selector
- Replace EXIT traps with explicit cleanup calls in integration test to
avoid interfering with the parent runner's trap
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
* fix: error on missing secret key when using vals
Add HELMFILE_VALS_FAIL_ON_MISSING_KEY_IN_MAP environment variable
to control whether vals should fail when a referenced key does not
exist in the secret map.
Previously, when a secret reference like ref+vault://path#/nonexistent-key
pointed to a non-existent key, vals would silently return an empty string
without error. This could lead to deployments with missing configuration.
Default behavior remains backward compatible (returns empty string).
Set HELMFILE_VALS_FAIL_ON_MISSING_KEY_IN_MAP=true to enable strict mode.
Fixes#1563
Signed-off-by: yxxhero <aiopsclub@163.com>
* refactor: extract buildValsOptions helper and improve tests
- Extract buildValsOptions() to make vals configuration testable
- Use t.Setenv instead of manual env save/restore in tests
- Test actual vals.Options output including FailOnMissingKeyInMap
Addresses PR review comments on #2496
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: use strconv.ParseBool and make tests hermetic
- Use strconv.ParseBool for FailOnMissingKeyInMap parsing to support
common boolean values like 'TRUE', '1', '0', etc.
- Always set env vars explicitly in tests (even to empty string) to
prevent flaky tests when env vars are set externally
- Add test cases for various boolean formats
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: add documentation for vals-related environment variables
Add documentation for:
- HELMFILE_AWS_SDK_LOG_LEVEL: configure AWS SDK logging for vals
- HELMFILE_VALS_FAIL_ON_MISSING_KEY_IN_MAP: enable strict mode for secret refs
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: improve error handling and case-insensitive comparison
- buildValsOptions now returns error for invalid boolean values
instead of silently defaulting to false
- Use strings.EqualFold for case-insensitive 'off' comparison
to handle OFF, Off, etc.
- Add test cases for invalid boolean and uppercase OFF
- Update docs to mention case-insensitive and error behavior
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: normalize log level and improve singleton initialization
- Normalize AWS log level 'off' to lowercase for true case-insensitivity
- Replace sync.Once with mutex to allow recovery from config errors
- Update tests to expect normalized 'off' value
- Update docs to clarify when error is raised
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
This commit makes the apply logic exit early in the event there are no
changes to releases. I believe this effectively reverts helmfile#522.
Updates relevant snapshots
Clarify conditions under which preapply hooks are triggered to include that they will no longer fire if there is a no-op.
Docs as requested by the maintainer from a copilot request made by them.
Fixes: helmfile#679
Signed-off-by: Thomas Arrow <thomas.arrow@wikimedia.de>
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
Add new documentation explaining how Helmfile merges values from various sources:
- Core architecture and data flow
- Values sources and precedence order
- Deep merge behavior for maps and arrays
- Environment-specific value handling
- Secret management and priorities
- Common patterns and troubleshooting
This guide helps users understand the foundational concepts needed for
writing effective helmfiles, especially regarding value overrides and
merge strategies.
Signed-off-by: yxxhero <aiopsclub@163.com>
* Add IP Network to supported HCL Functions
This patch adds CIDR functions from the `go-cty-funcs` package to
supported HCL functions
Signed-off-by: Oleh Neichev <oleg.neichev@gmail.com>
* Test HCL CIDR Functions
Signed-off-by: Oleh Neichev <oleg.neichev@gmail.com>
---------
Signed-off-by: Oleh Neichev <oleg.neichev@gmail.com>
* feat: allow for HCL values override
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: ensure overriden HCL expression uses range from latest defined block vars
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: implement HCL cty values override tests
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* docs: better describe new behavior
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: add extra parenthesis for better readability
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: implement variable override in decodeGraph() function, AFTER interpolation, providing back access to hv.* and local.* accessors
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: implement better HCL test to override values using local.* and hv.* accessors and pre-processing function calls
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: remove deprecated hclParseError() function (and test)
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: don't let HCL override with null value win
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: proper test condition on HCL map type merge (and tests)
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: more accurate HCL test error statement
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: ensure HCL DAG graph collects dependencies from ALL definitions to ensure proper evaluation order even if only earlier definitions have dependencies
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: verify HCL mixed-types merges are correctly supported
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* docs: improved environment values precedence section with HCL override support
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: HCL test spell-check, linter failure
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: implement HCL override e2e tests
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: correct hcl_loader test error message
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: ensure correct cty type is returned in case of object/map hcl merge
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: ensure hcl locals from a previous definition/file do not leak into this evaluation when merging
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* fix: correct e2e hcl_override test; missing line in output string comparison
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* docs: spell-check on HCL doc
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
* chore: update comment for accuracy in HCL read routine
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
---------
Signed-off-by: Benjamin Zores <benjamin.zores@gmail.com>
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* fix: eliminate os.Chdir in sequential helmfiles to fix relative path resolution
The sequential code path used within() → os.Chdir() to change the
process-wide working directory when processing helmfile.d files.
This broke relative environment variable paths (e.g. KUBECONFIG=kubeconfig.yaml)
because they resolved from the wrong directory after chdir.
Replace the chdir-based approach with the same baseDir parameter pattern
used by the parallel code path, passing explicit directory context through
loadDesiredStateFromYamlWithBaseDir() instead of mutating global process state.
Closes#2409
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* fix: restore within() for single-file sequential to preserve chart path format
The previous approach used baseDir for all sequential processing, which
changed chart path format in output (e.g. from "../../../../charts/raw"
to "test/integration/charts/raw"). This broke integration tests that
compare chart paths in expected output.
Now the sequential branch uses two strategies:
- Single file: use os.Chdir via within() to preserve backward-compatible
relative chart paths in output
- Multiple files with --sequential-helmfiles: use baseDir parameter to
avoid os.Chdir, fixing relative env var paths like KUBECONFIG (#2409)
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* fix: revert e2e snapshot outputs to match within() behavior
The previous commit restored within() for single-file sequential
processing, which produces relative chart paths (e.g. ../../charts/raw)
and filename-only FilePath. Revert the e2e snapshot expected outputs
to match main branch since single-file behavior is now identical.
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* fix: restructure integration test for multi-file sequential processing
- Point -f at helmfile.d/ directly (not parent dir) so findDesiredStateFiles
discovers the yaml files
- Add second helmfile to trigger baseDir path (len > 1)
- Inline environment config to avoid base file relative path issues
- Verify both releases appear in output instead of comparing with parallel
(which may differ in ordering)
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* fix: reduce cognitive complexity and improve accuracy of sequential helmfiles
Replace inline visitSubHelmfiles closure with calls to the existing
processNestedHelmfiles() method, matching the parallel path. This
eliminates duplicated nested logic and reduces gocognit complexity
below the CI threshold of 110. Also fixes help text and docs to
accurately describe that single-file processing still uses within(),
and adds kubeContext verification to the integration test.
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* test: validate kubeContext resolution in sequential helmfiles integration test
Restructure the integration test to replicate the exact user scenario
from issue #2409:
- Multiple files in helmfile.d/ using bases: with relative paths
(../bases/) for environments and defaults
- Environment values set kubeContext via .Environment.Values
- helmDefaults.kubeContext rendered from gotmpl
- Local chart references (../../../../charts/raw) from helmfile.d/
- Run diff against the minikube cluster to exercise kubeContext
resolution, which would fail with "context does not exist" if
os.Chdir() broke relative path resolution
- Also verify template output for both releases and relative values
file (values/common.yaml) resolution
Fix normalizeChart() in util.go to be idempotent — skip re-prefixing
when the chart path already starts with basePath. This prevents
double-prefixing of local chart paths (e.g. helmfile.d/test/.../raw)
when normalizeChart is called multiple times (once during chart
preparation and again during diff/sync).
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
---------
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
Adds a new `helmfile unittest` command that integrates the helm-unittest
plugin, allowing users to define unit test paths per release and run them
via helmfile.
Closes#2376
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* fix: skip cache refresh for shared cache paths to prevent race conditions
When multiple helmfile processes run in parallel (e.g., as ArgoCD plugin),
they share the same OCI chart cache in ~/.cache/helmfile. One process could
delete and re-download (refresh) a cached chart while another process was
still using it, causing "path not found" errors.
This fix:
- Adds isSharedCachePath() helper to detect shared cache paths
- Skips chart deletion/refresh for paths in the shared cache directory
- Users can force refresh by running `helmfile cache cleanup` first
Fixes#2387
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: document OCI chart caching behavior and multi-process safety
Add documentation for:
- OCI chart cache location and behavior
- How to force cache refresh with `helmfile cache cleanup`
- Multi-process safety when using shared cache
- Cache management commands (info, cleanup)
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: address review comments for shared cache handling
- Return error instead of chartActionDownload for corrupted shared cache
- Change refresh skip log from Debugf to Infof for user visibility
- Add t.Helper() to createTestLogger test helper
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: handle symlinks and add debug logging in isSharedCachePath
- Use filepath.EvalSymlinks to resolve symlinks before path comparison
- Add debug logging when filepath.Abs fails
- Add test case for symlink to shared cache directory
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: address copilot review comments
- Include underlying error in corrupted cache error message
- Add cleanup for test directories created in shared cache
- Clarify --skip-refresh flag documentation
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: handle edge case when chartPath equals sharedCacheDir
- isSharedCachePath now returns true for exact match with cache dir
- Add test case for exact match with shared cache directory
Signed-off-by: yxxhero <aiopsclub@163.com>
* test: add integration test for acquireChartLock shared cache behavior
Add TestAcquireChartLockSharedCacheSkipRefresh to verify that
acquireChartLock returns chartActionUseCached instead of
chartActionRefresh when the chart exists in the shared cache,
even when refresh is requested. This tests the core fix for
the race condition issue #2387.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* Initial plan
* Fix helmfile.d parallel processing to respect alphabetical order
Changed multiple helmfile.d file processing from parallel to sequential
to match the documented behavior. Files are now processed in alphabetical
order as stated in the documentation, ensuring predictable deployment order.
This fixes issues where deployments would break because parallel processing
did not respect the intended ordering (e.g., databases before applications).
- Modified visitStatesWithContext to process files sequentially
- Removed unused processStateFileParallel function
- Updated test snapshots to reflect deterministic ordering
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Add --sequential-helmfiles flag for opt-in sequential processing
Based on feedback, added a flag to control processing mode instead of
forcing sequential behavior. Parallel processing is now the default for
performance, with an opt-in --sequential-helmfiles flag for when order matters.
Changes:
- Added SequentialHelmfiles field to GlobalOptions and App
- Added --sequential-helmfiles CLI flag
- Modified visitStatesWithContext to check flag and choose parallel vs sequential
- Updated documentation to clarify parallel is default, sequential is opt-in
- Restored original parallel processing as default behavior
This gives users control over the behavior based on their needs.
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Add documentation and improve code readability
- Added documentation to processStateFileParallel function
- Extracted shouldProcessInParallel variable for clarity
- All tests pass, security scan clean
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Fix lint issue: remove trailing whitespace
Removed trailing whitespace from line 983 in pkg/app/app.go to fix
formatting lint error.
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
This commit adds comprehensive support for Helm 4 while maintaining
full backward compatibility with Helm 3. The implementation includes:
- Updated helm version detection to support both Helm 3 and Helm 4
- Added HELMFILE_HELM4 environment variable to control Helm version
- Modified helm execution paths to handle version-specific binaries
- Updated helm plugin installation to support split architecture
- Helm 4: Uses split plugin architecture (3 separate .tgz files)
- helm-secrets.tgz
- helm-secrets-getter.tgz
- helm-secrets-post-renderer.tgz
- Helm 3: Continues using single plugin installation
- Updated Dockerfiles, CI workflows, and core installation code
- Helm 4 requires post-renderers to be plugins, not executable scripts
- Created Helm plugin structure for integration tests
- Updated helmfile.yaml templates to dynamically select renderer type
- Added test plugins: add-cm, add-cm1, add-cm2
- Updated integration tests for Helm 3/4 compatibility
- Created Helm 4 variant expected output files
- Fixed test determinism issues (repo cleanup between iterations)
- Added version-specific output filtering for warnings/messages
- Updated workflows to test both Helm 3 and Helm 4
- Matrix testing across Helm versions
- Updated helm-diff to v3.14.0 for compatibility
- Updated README and docs with Helm 4 information
- Added migration guidance
- Updated version requirements
All changes are backward compatible - existing Helm 3 users will
see no behavior changes.
fix: update Helm 4 lint expected output to match filtered output
The grep filter removes the semver warning, so the expected output
should not include it. Updated lint-helm4 files to match the filtered
output (warning removed, no extra blank line).
Signed-off-by: Aditya Menon <amenon@canarytechnologies.com>
* feat: Add updateStrategy option in the state file with 'reinstall'/'reinstallIfForbidden' choices to uninstall and apply the specific release(s) (if forbidden to update)
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Fix unit tests related to the new updateStrategy feature
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Fix unit tests related to the new updateStrategy feature
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Resolve linter issue due to cognitive complexity
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Updated index.md to describe the possible values of updateStrategy
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Add validation of updateStrategy parameter and unit test
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Updated unit test
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Removed 'reinstall' update strategy option to only have reinstallIfForbidden, cleanup of pre-sync changes, adapted unit tests
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Display affected releases that were reinstalled
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Make sure to add --wait when deleting a release to be reinstalled due to reinstallIfForbidden
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Apply suggestions from Copilot code review
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
---------
Signed-off-by: Simon Bouchard <sbouchard@rbbn.com>
* Initial plan
* Implement --timeout flag for helmfile sync command
- Add Timeout field to SyncOptions struct in pkg/config/sync.go
- Add --timeout flag to sync command in cmd/sync.go
- Add Timeout field to SyncOpts struct in pkg/state/state.go
- Modify timeoutFlags() function to prioritize CLI timeout over release and default configs
- Add test case to verify CLI timeout overrides other timeout settings
- Follow same pattern as existing --wait and --wait-for-jobs flags
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Fix lint issues: format test struct fields properly
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Update docs: Add --timeout flag documentation for helmfile sync command
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
* Allow caching of remote files to be disabled
Make it possible to automatically update the cache of remote
resources by disabling the caching of those resources using a query
string parameter (`cache=false`).
Signed-off-by: Jess <jess@ros.io>
* Fix test that broke
Because query parameters are being re-encoded, = is being encoded to %3D.
Signed-off-by: Jess <jess@ros.io>
* Add test for disabling caching of remote resources
Signed-off-by: Jess <jess@ros.io>
* Include example usage in docs
Signed-off-by: Jess <jess@ros.io>
---------
Signed-off-by: Jess <jess@ros.io>
* docs: update status section with May 2025 release information
Signed-off-by: yxxhero <aiopsclub@163.com>
* docs: Update Helmfile v1 proposal link in index.md
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* Feat: reuseValues in release
Adding properties to set reuseValues flag on release-level.
Signed-off-by: Adam Blasko <adam.blasko1@gmail.com>
* feat: fixing tests
Most of the tests had issues with flag order, which changed due to moving the value control flags out of the "common flags" for diff
Signed-off-by: Adam Blasko <adam.blasko1@gmail.com>
* fix: fixing lint issue
Signed-off-by: Adam Blasko <adam.blasko1@gmail.com>
---------
Signed-off-by: Adam Blasko <adam.blasko1@gmail.com>
* docs: add skipSchemaValidation to index.md and update related structs
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: add SkipSchemaValidation to config and state handling
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: add HELMFILE_INTERACTIVE env var to enable interactive mode
This commit adds the `HELMFILE_INTERACTIVE` environment variable
to enable the interactive mode by default.
Anything other than `true` will disable the interactive mode.
The precedence has the `--interactive` flag.
Signed-off-by: Krzysztof Łuczak <krzysztof.luczak.pro@gmail.com>
* Trim trailing whitespaces
Signed-off-by: Krzysztof Łuczak <krzysztof.luczak.pro@gmail.com>
---------
Signed-off-by: Krzysztof Łuczak <krzysztof.luczak.pro@gmail.com>
* eat: adding ability for for charts to be pulled without HTTPS
accomplished by:
- Adding PlainHttp attribute to RepositorySpec., HelmDefault, ReleaseSpec
- Adding UnitTests for getOCIChart Flags.
- Adding funciton and unitTests for getChartDownload
- Changing and refactoring how Flags are added to getOCIChart.
Resolves#1224
Signed-off-by: Peter Halliday <peter.halliday@servicenow.com>
* Pass PlainHttp to OCI repo options, fix unit test
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Fix doc
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Use repository fields in non-OCI chart download options
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Update hashes in TestGenerateID
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Make sure repo exists when using its options
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Do not add TLS options if PlainHttp is set, adapt unit tests
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Fix doc
Signed-off-by: Pascal Rivard <privard@rbbn.com>
* Remove 'else if' from appendChartDownloadFlags
Signed-off-by: Pascal Rivard <privard@rbbn.com>
---------
Signed-off-by: Peter Halliday <peter.halliday@servicenow.com>
Signed-off-by: Pascal Rivard <privard@rbbn.com>
Co-authored-by: Peter Halliday <peter.halliday@servicenow.com>
Co-authored-by: Pascal Rivard <privard@rbbn.com>
* Add "PostRendererArgs" option to be passed to helm
This allows using PowerShell scripts on Windows as Post Renderer.
Signed-off-by: Maarten Boekhold <maarten.boekhold@finastra.com>