Files
yxxhero 16259008d5 feat: opt-in OpenTelemetry tracing and metrics (experimental) (#2769)
* feat(telemetry): add opt-in OpenTelemetry tracing (PR 1: lifecycle + root span)

Implements the first increment of docs/proposals/otel-tracing.md (#2767):

- pkg/telemetry: SDK setup from standard OTEL_* env vars (autoexport for
  exporter selection, env-driven sampler/propagators, OTEL_SDK_DISABLED),
  command-span lifecycle, no-op-by-default accessors
- --otel-tracing flag / HELMFILE_OTEL_TRACING env switch
- root span "helmfile <command>" with file/environment/selectors/exit_code
  attributes; TRACEPARENT-based remote-parent extraction for CI correlation
- shutdown flush on both normal-exit and signal paths (nil-safe, 5s bound)
- app.New derives its context from telemetry.CommandContext()
  (Background-identical when tracing is disabled)
- docs: otel.md user guide, experimental-features entry, design proposal
- tests: hermetic unit tests, app context-contract pinning, flag registration

Telemetry problems never fail a run: exporter misconfiguration and export
errors degrade to disabled with a warning. When disabled, behavior and
performance are identical to before (no-op tracer, no goroutines, no
network).

Refs: #2767, #2758
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): trace every external process + trace-context bridges (PR 2)

Implements the second increment of docs/proposals/otel-tracing.md (#2767):

- pkg/helmexec/span.go: one span per external process started by helmfile
  (helm invocations, hooks, plugin execs) at the ShellRunner choke point —
  helm.exec (with helm.subcommand) vs os.exec, with redacted exec.args,
  exec.exit_code, and error status on failure
- pkg/helmexec/redact.go: shared argument redaction with two profiles;
  legacy is byte-identical to the historical exit-error behavior (existing
  goldens unchanged), strict (spans) additionally covers --set=k=v and
  credential flags; exit_error.go now uses the shared helper
- orphan-trace bridges with bit-identical cancellation semantics
  (context.WithoutCancel of the command context): both kubedog call sites
  (state.go) and hook execution (event.Bus gains an optional Ctx consumed
  by its default runner; state.go sets it, nil falls back to TODO as before)
- OTLP end-to-end test (in-process httptest receiver, no external
  collector): span export, error status/exit code, redaction, and
  parent-linkage to the command span
- docs/otel.md updated to the now-traced surface

Verified end-to-end with the console exporter: helmfile template on a
local chart yields the command span plus helm.exec spans for helm
version/dependency/template, all nested under it.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): state-loading and hook spans (PR 3a)

Implements the third increment of docs/proposals/otel-tracing.md (#2767):

- helmfile.discover_states around findDesiredStateFiles and helmfile.load
  around loadDesiredStateFromYamlWithBaseDir; both cover all callers
  (incl. nested helmfiles) with no signature changes
- helmfile.render / helmfile.parse children per document part, parented
  through a traceCtx field on the unexported desiredStateLoader struct
  (set once at its single construction site)
- helmfile.hook span per hook execution: Trigger's per-hook body extracted
  into runHook (readability win on its own), the hook's subprocess span
  nests under it via a per-hook ctx-swapped ShellRunner clone
  (cancellation unchanged — Bus.Ctx never carries cancellation by contract)
- pkg/telemetry/otlptest: shared in-process OTLP/HTTP receiver harness,
  now used by helmexec, event, and app span tests
- golden span-tree test at the app layer (root -> discover -> load ->
  render/parse, via the exectest fake helm) and a hook-span nesting test
- nil-ctx guard for App literals built directly by tests (App.spanParentCtx)

Verified end-to-end with the console exporter: a template run over a
gotmpl state file with a prepare hook yields the full tree with the hook's
os.exec nested under helmfile.hook.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): per-release spans nested under the load span (PR 3b)

Implements the per-release increment of docs/proposals/otel-tracing.md
(#2767) — spans nest command -> load -> release -> helm exec:

- helmexec.HelmContext gains an optional Ctx carrying the per-release span
  context; the execer's new execWithContext funnel consumes it via a
  per-call runner clone (runnerWithCtx) so the shared, cached execer is
  never mutated across concurrent workers. The seven Interface methods
  that take a HelmContext (Sync/Diff/ReleaseStatus/List/DecryptSecret/
  Delete/Test) route through it; nil Ctx behaves exactly as before.
  exec() lost its always-nil override parameter on the way (unparam).
- pkg/state/span.go: SetTraceContext + startReleaseSpan/endReleaseSpan
  helpers (release/namespace/chart/labels attributes, sorted for stable
  output); a typed-nil guard (releaseErrAsError) avoids the classic
  nil-pointer-in-interface trap on *ReleaseError.
- release spans in the worker loops: SyncReleases, DiffReleases,
  DeleteReleasesForSync, PrepareCharts, and iterateOnReleases (status/
  delete/test via a new verb parameter); their HelmContext is stamped with
  the release span context where one is built.
- pkg/app sets st.SetTraceContext(loadCtx) right after loading a state
  file, rooting all per-release spans under helmfile.load.
- bridged one more detached tracking call found on the way
  (trackReleaseIfEnabled's context.Background in the sync worker).
- golden test: release span present, nested under load, correct
  attributes; unit test for the runnerWithCtx clone semantics.

Verified with the console exporter: helmfile template yields
release.prepare(demo) under load with full attributes.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): nest status/delete/test execs under their release spans

Completes the per-release exec nesting for the iterateOnReleases-based
loops (docs/proposals/otel-tracing.md §4.4 phase 2): the do closures now
receive the release span context and stamp it into their HelmContext, so
helm status/delete/test subprocess spans nest under
helmfile.release.<verb> like sync/diff already did.

- scatterGatherReleases/iterateOnReleases/doWithReleaseSpan: do gains a
  context parameter (the release span context)
- ReleaseStatuses/DeleteReleases/TestReleases closures stamp
  HelmContext.Ctx from it
- integration test with a real execer (version-probe shim binary): the
  release's status subprocess nests under helmfile.release.status, same
  trace, with helm.subcommand=status

This also makes the otel.md claim ("upgrade, diff, delete, status, test
nested under the release span") fully accurate.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): OTel metrics — helm exec duration and release results (PR 4)

Implements the metrics increment of docs/proposals/otel-tracing.md
(#2767) on the same provider, switch, and resource as traces:

- pkg/telemetry/metrics.go: helmfile.helm.exec.duration histogram
  (subcommand, success) and helmfile.release.count counter (verb,
  result). Instruments come from the otel global meter, so recording at
  call sites is branch-free no-op when telemetry is disabled.
- Setup builds the resource once and installs both providers; reader
  selection delegates to autoexport (OTEL_METRICS_EXPORTER: otlp |
  console | prometheus | none), the OTLP reader's interval honors
  OTEL_METRIC_EXPORT_INTERVAL (read by the SDK). Shutdown flushes both
  providers (errors.Join). StartCommandSpan now carries the meter
  provider across state transitions (fixes a nil-shutdown panic).
- helmexec: finishExecSpan records exec duration for helm binaries;
  state: endReleaseSpan counts release outcomes for sync/diff/delete/
  status/test/prepare (diff counted as success when no hard error).
- otlptest: recorder routes by OTLP path (/v1/traces vs /v1/metrics)
  and decodes metrics; new FindMetric helper.
- tests: metrics recorded as no-op when disabled, provider enabled with
  the none exporter, degradation on an invalid metrics exporter, and an
  integration assertion (status exec duration datapoint + one successful
  release.count) in the shim-based state test.

Verified with the console exporter: helmfile template emits
helmfile.helm.exec.duration per subcommand (version/dependency/
template) and helmfile.release.count{verb=prepare,result=success}=1.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: complete OTel documentation coverage

- docs/cli.md: --otel-tracing in the CLI reference help block (verbatim
  from the cobra output)
- CHANGELOG.md: [Unreleased] Added entry for tracing + metrics
- docs/index.md: Observability highlight linking docs/otel.md
- docs/proposals/otel-tracing.md: add OTEL_METRICS_EXPORTER /
  OTEL_METRIC_EXPORT_INTERVAL rows to the env-var table and note the
  periodic reader + bounded metric cardinality in §7

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: drop unused id parameter from parsePart (unparam)

The id parameter was never used inside the span wrapper; the caller's id
variable is still used for the render calls and error messages.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review — redaction gaps, kubedog valve, phantom metrics

Addresses all Copilot review comments on #2769:

Security (span payloads):
- exec.args: positional arguments are additionally passed through
  helmexec.RedactedURL, so credentials embedded in chart/repository URLs
  (AddRepo, RegistryLogin, OCI refs) are masked exactly like log output
- release spans sanitize helmfile.chart the same way
- error statuses no longer embed raw errors (which contain rendered
  commands, arguments, and subprocess output): the command span, release
  spans, hook spans, and exec spans now use generic descriptions; the
  concrete exit code remains an attribute, and RecordError on the root
  span is dropped

Correctness:
- kubedog safety valve restored: execWithContext now attaches the
  per-release span into the runner's own context instead of replacing it,
  so trackHandle.Cancel() can interrupt a wedged helm again and app
  cancellation semantics stay exactly as before the PR
- diff release spans/metrics: real failures are recorded (exit code 2
  "changes detected" still counts as success); previously every diff was
  exported as successful
- skipped releases no longer emit phantom spans and inflate
  helmfile.release.count: iterateOnReleases callers pass a skip predicate
  (skipUndesired for status/test; delete deletes undesired releases and
  passes nil)
- Setup shuts down the already-constructed tracer provider (bounded) when
  the metrics provider fails, instead of abandoning its batch goroutine

Tests: URL redaction cases (masked/untouched), spanAttachedContext
preserves the runner cancellation chain while attaching the caller's
span, skipUndesired, and a failing-hook span asserting the generic
message.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: lint — restore nolint placement and avoid nil context literal

- the skipUndesired insertion had displaced the // nolint: unparam
  directive off iterateOnReleases (helm param is intentionally unused
  there); also fixes a skipDesired/skipUndesired comment typo
- use a typed nil in TestSpanAttachedContext (staticcheck SA1012)

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review round 2 — remote-ref redaction, wrapper helm binaries, hook release attribution

Addresses all 6 new review comments on #2769:

Security (remote references):
- new helmexec.RedactedRef sanitizes go-getter style references for
  telemetry: forced-form prefixes (git::, s3::) preserved, whole URL
  userinfo masked (usernames carry tokens too), credential-bearing query
  parameters masked using pkg/remote's heuristic (token/password/secret/
  key/signature). Applied to helmfile.file (command span), helmfile.path
  (discover_states), helmfile.chart (release spans), and exec.args —
  log-time RedactedURL is untouched so log output is unchanged

Correctness:
- wrapper helm binaries (--helm-binary custom names) are now classified
  as helm operations by an explicit context marker stamped in the execer
  funnel, instead of the executable-basename heuristic; the same
  classification gates helmfile.helm.exec.duration, so the metric no
  longer misses wrapper invocations (classifyExec)
- release-scoped hooks (presync/postsync/preuninstall/postuninstall/
  cleanup in the sync/delete/diff workers) now attach their helmfile.hook
  spans to the active helmfile.release.* span via a variadic parent on
  the trigger functions; global hooks keep the command context and all
  29 existing call sites compile unchanged; hook cancellation stays
  detached (WithoutCancel) as before
- signal-terminated runs (Shutdown with exitCode 130/143 and nil error)
  now mark the command span with error status, consistent with their
  nonzero exit code

Tests: RedactedRef table (forced forms, userinfo, s3/token query params,
untouched cases), classifyExec marker case, hookTraceContext parent
attribution + non-cancellability + fallback.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review round 3 — redaction corner cases, value runners

Addresses 5 of the 6 new review comments on #2769 (the sixth — an
unused strings import in exit_error.go — is a false positive: Indent
still uses strings.Split/Builder and the package compiles):

- RedactArgs read the previous token from the progressively redacted
  output, so {--set, --set-string, secret} leaked the secret
  (the masked value hid the following flag). Read the previous token
  from the original input, restoring the legacy contract for adjacent
  secret flags
- RedactedRef fails closed for malformed references: URL-like refs with
  invalid percent escapes export a fully redacted value, and an
  unparseable query is dropped entirely instead of exported verbatim
- ShellRunner has value receivers, so a ShellRunner VALUE satisfies the
  Runner API; the helm marker stamping and the per-release span
  attachment now handle both value and pointer forms (matching
  WithContext), so value-runner callers keep release nesting and the
  helm.exec classification/metric

Regression tests: adjacent secret flags (legacy + strict), malformed
URL-like ref, malformed query, value-runner marker + span attachment.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): stamp the helm marker on the stdin funnel too

execStdIn (registry login, repo add) called the runner directly, so
wrapper --helm-binary names were misclassified as os.exec and omitted
from helmfile.helm.exec.duration on that path. The marking now goes
through a shared markHelmRunner helper (value and pointer ShellRunner
forms) used by both execution funnels.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): redact helm's --kube-token in strict profile

Helm's global --kube-token carries a bearer token; both the
two-argument and inline forms are now masked in span exec.args
(legacy exit-error output is untouched, matching its historical
behavior).

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): OTel metrics best-practice alignment

- helmfile.helm.exec.duration now declares explicit bucket boundaries
  tuned for seconds-scale helm invocations (5ms…600s); the SDK defaults
  are millisecond-oriented and lumped every sub-5s invocation — the
  common case — into the first bucket, defeating the histogram
- instruments are re-created under the installed provider with the
  instrumentation scope version stamped (Setup-time, race-free)
- helmfile.release.count declares the {release} curly-annotation unit
  per the metrics naming conventions

Tested end-to-end via the OTLP integration test: exported bounds are
the tuned set, units are asserted, and the scope carries the version.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): per-release duration metrics behind an opt-in switch

New helmfile.release.duration histogram (seconds, same tuned buckets)
with bounded dimensions by default (verb, result). Setting
HELMFILE_OTEL_METRICS_PER_RELEASE=true adds helmfile.release and
helmfile.namespace, answering "which release is slow" from dashboards:

- well-suited to bounded CI runs; long-lived centralized collection
  needs a backend capacity/TTL story (documented in docs/otel.md)
- per-release timing remains available in traces without the flag
- env read per call (release operations are low-frequency, and tests
  toggle it)

endReleaseSpan now takes the release and the operation start time; the
five worker-loop call sites pass them (doWithReleaseSpan, SyncReleases,
DeleteReleasesForSync, PrepareCharts, DiffReleases).

Verified end-to-end with the console exporter (default dims vs
per-release) and OTLP integration tests pinning both modes.

Refs: #2767, #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): maintainability pass over the runner/metric plumbing

- StartCommandSpan copies the tracingState struct instead of enumerating
  fields by hand — that pattern dropped the meter provider once already
- the two value/pointer ShellRunner switches (helm marker, span
  attachment) are unified into one withRunnerCtx helper; the duplication
  caused two review rounds of value-form misses
- classifyExec derives the helm classification from the span name
  (helmExecSpanName constant) instead of returning a third parallel bool
- metrics: shared outcomeAttrs for the verb/result dimensions, and the
  bucket slice renamed to durationBuckets with a comment covering both
  histograms that use it

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): consolidate test env lists, trace bridges, and hook prep; sync the design doc

Maintainability:
- HermeticEnvVars is now exported from pkg/telemetry (the owner of the
  env surface) and used by both telemetry tests and otlptest — the two
  copies had already drifted once (HELMFILE_OTEL_METRICS_PER_RELEASE
  needed updating in both)
- kubedogTraceContext and hookTraceContext were the same concept written
  twice; unified into traceOnlyContext(parent...) in span.go

Readability:
- runHook's nested kubectl rewrite extracted into prepareKubectlHook
  with guard-clause structure

Accuracy (docs ↔ code, drifted over five review rounds):
- §4.4 now describes the implemented mechanism: the release span is
  INJECTED into the runner's own context (preserving the kubedog safety
  valve) rather than the runner context being replaced, and helm
  classification is marker-based for wrapper binaries
- §5 exec span rows list the actual attributes incl. URL/query masking
- §6 strict profile documents RedactedRef, --kube-token, and the
  adjacent-token guarantee

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): drop the dead noop state, relocate skipUndesired, sync user-facing accuracy

- tracingState.noop was dead weight in the enabled state and a
  copy-surface in every transition; a single package-level
  noopTracerProvider now backs Tracer while disabled
- skipUndesired moved next to doWithReleaseSpan in span.go, its only
  conceptual home (span/metric suppression, not run plumbing)
- accuracy: the package doc, --otel-tracing flag help,
  experimental-features entry, and CHANGELOG now all say tracing AND
  metrics and list the third instrument (helmfile.release.duration with
  the HELMFILE_OTEL_METRICS_PER_RELEASE opt-in) — these had drifted
  when the metric was added; the PR description's metric table is
  updated to match as well

No behavior change; full -race suite green (except the pre-existing
network-dependent TestStorage_resolveFile flake), lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): flatten Setup, name the prefix bound, dedupe test fake; fix instrument-count drift

Readability/maintainability:
- Setup drops from 56 to 39 lines: provider construction (including the
  shutdown-tracer-on-meter-failure recovery) moves to newProviders in
  exporter.go next to the constructors it composes
- refredact's magic 16 becomes maxForcedFormPrefix with a comment
- span_test's hand-rolled fakeRunner removed in favor of the existing
  mockRunner (same package)

Accuracy:
- "Two instruments" wording survived in docs/otel.md and the design
  proposal §7 after helmfile.release.duration was added; both now say
  three and mention the per-release opt-in

No behavior change; full -race suite green (except the pre-existing
network flake), lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): co-locate span machinery, drop a dead export, fix docs nits

- the span plumbing helpers (markHelmExec, withRunnerCtx,
  markHelmRunner, spanAttachedContext) move from exec.go to span.go,
  next to the marker type and classifiers they serve — exec.go keeps
  only the funnel call sites
- otlptest.SpanNames was never used outside the package; unexported
- isHelmBinary's comment now states it is the FALLBACK classifier
  (funnel invocations are marker-classified), replacing the outdated
  "cosmetic distinction" framing from before the marker existed
- docs/otel.md: release-scoped hooks nest under their release span
  (added in review round 2, never documented)

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-09-07 20:34:10 +08:00

624 lines
33 KiB
Markdown

# CLI Reference
## CLI Reference
```
Declaratively deploy your Kubernetes manifests, Kustomize configs, and Charts as Helm releases in one shot
V1 mode = false
YAML library = go.yaml.in/yaml/v3
Usage:
helmfile [command]
Available Commands:
apply Apply all resources from state file only when there are changes
build Build all resources from state file
create Create a helmfile deployment project scaffold
cache Cache management
charts DEPRECATED: sync releases from state file (helm upgrade --install)
completion Generate the autocompletion script for the specified shell
delete DEPRECATED: delete releases from state file (helm delete)
deps Update charts based on their requirements
destroy Destroys and then purges releases
diff Diff releases defined in state file
fetch Fetch charts from state file
help Help about any command
init Initialize the helmfile, includes version checking and installation of helm and plug-ins
lint Lint charts from state file (helm lint)
list List releases defined in state file
repos Add chart repositories defined in state file
show-dag It prints a table with 3 columns, GROUP, RELEASE, and DEPENDENCIES. GROUP is the unsigned, monotonically increasing integer starting from 1. All the releases with the same GROUP are deployed concurrently. Everything in GROUP 2 starts being deployed only after everything in GROUP 1 got successfully deployed. RELEASE is the release that belongs to the GROUP. DEPENDENCIES is the list of releases that the RELEASE depends on. It should always be empty for releases in GROUP 1. DEPENDENCIES for a release in GROUP 2 should have some or all dependencies appeared in GROUP 1. It can be "some" because Helmfile simplifies the DAGs of releases into a DAG of groups, so that Helmfile always produce a single DAG for everything written in helmfile.yaml, even when there are technically two or more independent DAGs of releases in it.
status Retrieve status of releases in state file
sync Sync releases defined in state file
template Template releases defined in state file
test Test charts from state file (helm test)
unittest Unit test charts from state file using helm-unittest plugin
version Print the CLI version
write-values Write values files for releases. Similar to `helmfile template`, write values files instead of manifests.
Flags:
--allow-no-matching-release Do not exit with an error code if the provided selector has no matching releases.
-c, --chart string Set chart. Uses the chart set in release by default, and is available in template as {{ .Chart }}
--color Output with color
--debug Enable verbose output for Helm and set log-level to debug, this disables --quiet/-q effect. Overrides "HELMFILE_DEBUG" OS environment variable when specified
--disable-force-update do not force helm repos to update when executing "helm repo add"
--enable-live-output Show live output from the Helm binary Stdout/Stderr into Helmfile own Stdout/Stderr.
It only applies for the Helm CLI commands, Stdout/Stderr for Hooks are still displayed only when it's execution finishes.
-e, --environment string specify the environment name. Overrides "HELMFILE_ENVIRONMENT" OS environment variable when specified. defaults to "default"
-f, --file helmfile.yaml load config from file or directory. defaults to "helmfile.yaml" or "helmfile.yaml.gotmpl" or "helmfile.d" (means "helmfile.d/*.yaml" or "helmfile.d/*.yaml.gotmpl") in this preference. Specify - to load the config from the standard input.
-b, --helm-binary string Path to the helm binary. Overrides "HELMFILE_HELM_BINARY" OS environment variable when specified (default "helm")
-h, --help help for helmfile
-i, --interactive Request confirmation before attempting to modify clusters
--kube-context string Set kubectl context. Overrides "HELMFILE_KUBE_CONTEXT" OS environment variable when specified. Uses current kubectl context by default
-k, --kustomize-binary string Path to the kustomize binary. Overrides "HELMFILE_KUSTOMIZE_BINARY" OS environment variable when specified (default "kustomize")
--log-level string Set log level. Overrides "HELMFILE_LOG_LEVEL" OS environment variable when specified (default "info")
-n, --namespace string Set namespace. Overrides "HELMFILE_NAMESPACE" OS environment variable when specified. Uses the namespace set in the context by default, and is available in templates as {{ .Namespace }}
--no-color Output without color. Overrides "HELMFILE_NO_COLOR" and "NO_COLOR" OS environment variables when specified
--otel-tracing Enable OpenTelemetry tracing (experimental).
Configure the exporter with standard OTEL_* environment variables (e.g. OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_TRACES_EXPORTER).
Overrides "HELMFILE_OTEL_TRACING" OS environment variable when specified. See docs/otel.md
-q, --quiet Silence output. Equivalent to log-level warn. Overrides "HELMFILE_QUIET" OS environment variable when specified
--repo-retries int Number of times to retry "helm repo add/update" and "helm registry login" on failure, with exponential backoff (1s, 2s, 4s, ..., capped at 30s). Set to 0 to disable retries. Overrides "HELMFILE_REPO_RETRIES" OS environment variable when specified
-l, --selector stringArray Only run using the releases that match labels. Labels can take the form of foo=bar or foo!=bar.
A release must match all labels in a group in order to be used. Multiple groups can be specified at once.
"--selector tier=frontend,tier!=proxy --selector tier=backend" will match all frontend, non-proxy releases AND all backend releases.
The name of a release can be used as a label: "--selector name=myrelease"
--skip-deps skip running "helm repo update" and "helm dependency build"
--state-values-file stringArray specify state values in a YAML file. Used to override .Values within the helmfile template (not values template).
--state-values-set stringArray set state values on the command line (can specify multiple or separate values with commas: key1=val1,key2=val2). Used to override .Values within the helmfile template (not values template).
--state-values-set-string stringArray set state STRING values on the command line (can specify multiple or separate values with commas: key1=val1,key2=val2). Used to override .Values within the helmfile template (not values template).
--sequential-helmfiles Process helmfile.d files sequentially in alphabetical order instead of in parallel
--strip-args-values-on-exit-error Strip the potential secret values of the helm command args contained in a helmfile error message (default true)
-v, --version version for helmfile
Use "helmfile [command] --help" for more information about a command.
```
**Note:** Each command has its own specific flags. Use `helmfile [command] --help` to see command-specific options. For example, `helmfile sync --help` shows operational flags like `--timeout`, `--wait`, and `--wait-for-jobs`.
### init
The `helmfile init` sub-command checks the dependencies required for helmfile operation, such as `helm`, `helm diff plugin`, `helm secrets plugin`, `helm helm-git plugin`, `helm s3 plugin`. When it does not exist or the version is too low, it can be installed automatically.
### cache
The `helmfile cache` sub-command is designed for cache management. Go-getter-backed remote file system are cached by `helmfile`. There is no TTL implemented, if you need to update the cached files or directories, you need to clean individually or run a full cleanup with `helmfile cache cleanup`
#### OCI Chart Cache
OCI charts are cached in the shared cache directory (`~/.cache/helmfile` by default, or `$HELMFILE_CACHE_HOME`). This cache is shared across all helmfile processes.
**Cache Behavior:**
- When a chart exists in the shared cache and is valid, it is reused without re-downloading
- The `--skip-refresh` flag can be used to skip checking for updates to cached charts stored in process-specific temporary directories (it does not affect charts already present in the shared cache)
- When running multiple helmfile processes in parallel (e.g., as an ArgoCD plugin), charts in the shared cache are not refreshed/deleted to prevent race conditions
**Forcing a Cache Refresh:**
To force a refresh of cached OCI charts, run:
```bash
helmfile cache cleanup
```
This will clear the shared cache, allowing the next helmfile command to re-download charts.
#### cache info
Display information about the cache directory.
#### cache cleanup
Remove all cached files from the cache directory.
### sync
The `helmfile sync` sub-command sync your cluster state as described in your `helmfile`. The default helmfile is `helmfile.yaml`, but any YAML file can be passed by specifying a `--file path/to/your/yaml/file` flag.
Under the covers, Helmfile executes `helm upgrade --install` for each `release` declared in the manifest, by optionally decrypting [secrets](#secrets) to be consumed as helm chart values. It also updates specified chart repositories and updates the
dependencies of any referenced local charts.
#### Common sync flags
* `--timeout SECONDS` - Override the default timeout for all releases in this sync operation. This takes precedence over `helmDefaults.timeout` and per-release `timeout` settings.
* `--wait` - Override the default wait behavior for all releases
* `--wait-for-jobs` - Override the default wait-for-jobs behavior for all releases
Examples:
```bash
# Override timeout for all releases to 10 minutes
helmfile sync --timeout 600
# Combine timeout with wait flags
helmfile sync --timeout 900 --wait --wait-for-jobs
# Target specific releases with custom timeout
helmfile sync --selector tier=backend --timeout 1200
```
For Helm 2.9+ you can use a username and password to authenticate to a remote repository.
### deps
The `helmfile deps` sub-command locks your helmfile state and local charts dependencies.
It basically runs `helm dependency update` on your helmfile state file and all the referenced local charts, so that you get a "lock" file per each helmfile state or local chart.
All the other `helmfile` sub-commands like `sync` use chart versions recorded in the lock files, so that e.g. untested chart versions won't suddenly get deployed to the production environment.
For example, the lock file for a helmfile state file named `helmfile.1.yaml` will be `helmfile.1.lock`. The lock file for a local chart would be `requirements.lock`, which is the same as `helm`.
The lock file can be changed using `lockFilePath` in helm state, which makes it possible to for example have a different lock file per environment via templating.
It is recommended to version-control all the lock files, so that they can be used in the production deployment pipeline for extra reproducibility.
To bring in chart updates systematically, it would also be a good idea to run `helmfile deps` regularly, test it, and then update the lock files in the version-control system.
### diff
The `helmfile diff` sub-command executes the [helm-diff](https://github.com/databus23/helm-diff) plugin across all of
the charts/releases defined in the manifest.
To supply the diff functionality Helmfile needs the [helm-diff](https://github.com/databus23/helm-diff) plugin v2.9.0+1 or greater installed. For Helm 2.3+
you should be able to simply execute `helm plugin install https://github.com/databus23/helm-diff`. For more details
please look at their [documentation](https://github.com/databus23/helm-diff#helm-diff-plugin).
#### Notable diff flags
* `--skip-diff-on-install` — skip running `helm diff` entirely for releases that are not yet installed. The release is treated as changed and will be synced on `apply` without showing a diff.
* `--skip-diff-validation-on-install` — for releases that are not yet installed, pass `--disable-validation` to `helm diff` so the diff is shown without K8s API server validation. Useful when a chart bundles CRDs and CRs together: the CRs would fail API validation before the CRDs are installed. This is the CLI-flag equivalent of the per-release `disableValidationOnInstall` field.
### doctor
`helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to summarize the changes and flag risks
(such as data loss, security exposure, breaking changes, downtime, performance, and best-practice issues).
**When no LLM is configured, `doctor` falls back to running `helmfile diff`**
with `--show-secrets` forced off. Most diff flags are accepted for
compatibility, but note two differences: `--output` is reserved for the doctor
report format (use `--diff-output` for helm-diff's plugin output format), and
`--show-secrets` is silently ignored (secrets are always redacted). This makes
it safe to swap into existing CI jobs: the worst case is you get the same diff
output you already had.
#### Configuration
The LLM endpoint speaks the OpenAI Chat Completions protocol (`/v1/chat/completions`). This means it works
with any compatible gateway:
- Direct providers: OpenAI, DeepSeek, Mistral, Together, Groq.
- Unified gateways: One-API, LiteLLM, Azure OpenAI proxy, Cloudflare AI Gateway.
- Local servers: Ollama (with OpenAI compatibility), vLLM, LocalAI.
Configuration precedence (low to high):
1. **Environment variables**: `HELMFILE_LLM_BASE_URL`, `HELMFILE_LLM_API_KEY`, `HELMFILE_LLM_MODEL`,
`HELMFILE_LLM_TIMEOUT` (Go duration, e.g. `90s`), `HELMFILE_LLM_MAX_TOKENS`.
2. **`helmfile.yaml` top-level `llm:` block**:
```yaml
llm:
baseURL: https://one-api.internal/v1
model: gpt-4o
apiKey: {{ env "HELMFILE_LLM_API_KEY" }}
timeout: 60s
maxTokens: 4096
```
3. **CLI flags** (highest precedence): `--llm-base-url`, `--llm-api-key`, `--llm-model`,
`--llm-timeout`, `--llm-max-tokens`.
A layer's non-zero fields override lower layers; empty fields fall through.
> **Note on zero values**: because merge treats "field == zero value" as "not set",
> `--llm-max-tokens 0` does NOT reset MaxTokens to 0 — it is treated as "flag not set"
> and the env/yaml value wins. This matches helmfile's existing flag-override convention
> (same as `--concurrency`, `--context`). Use the yaml block to explicitly request a zero
> value if you really need it.
`baseURL` is **optional**: if you use OpenAI's official endpoint, omit `baseURL` and
the client defaults to `https://api.openai.com/v1`. Set `baseURL` only when targeting
a gateway or non-OpenAI provider. Only `apiKey` and `model` are required to enable
LLM analysis.
#### Output
- **Default (markdown / text)**: a human-readable report with summary, risks sorted by severity
(🔴 high → 🟡 medium → 🟢 low → ⚪ unknown), affected resources, and a footer showing
model, duration, and secrets-redacted count.
- **`--output json`**: structured JSON including the model's analysis, the diff
(always post-redaction), and metadata. Suitable for CI pipelines that want to
post-process (e.g. comment on a pull request). The JSON shape is:
```json
{
"summary": "...",
"risks": [...],
"diff": "...",
"secrets_redacted": 3,
"model": "gpt-4o",
"duration": "8.2s",
"timestamp": "2026-..."
}
```
Key field semantics:
- `risks`: always an array when an analysis ran (even empty: `[]`, never `null`).
When doctor is unconfigured (no LLM), the field is omitted entirely. This lets
CI distinguish "LLM said no risks" from "analysis never happened".
- `diff`: always post-redaction. Doctor never exposes the raw pre-redaction diff
through stdout/JSON. If you need to debug helm-diff itself, run `helmfile diff` directly.
- `secrets_redacted`: always present, even when 0. Lets you confirm the redactor ran.
#### Exit codes
| Code | Meaning |
|---|---|
| 0 | success, or only low/medium risks, or LLM call failed (degraded to plain diff) |
| 2 | at least one high-severity risk and `--force` not passed (CI gate) |
| 1 | other error (state load failure, helm-diff runtime failure, etc.) |
The "detected changes" exit-2 signal from `helm diff --detailed-exitcode` is intentionally swallowed —
doctor's whole job is to react to changes, so reporting them via exit code would be noise.
#### Backend compatibility
Doctor uses `response_format: {type: "json_object"}` (JSON mode) by default for
reliable structured output. If the backend doesn't support JSON mode (common on
early One-API versions, some LiteLLM configs, or Ollama's OpenAI shim), doctor
automatically detects the 400 error, retries without `response_format`, and falls
back to extracting JSON from the model's response via a robust parser that handles
markdown fences and surrounding prose. No user action needed.
#### Safety
- `doctor` never invokes `apply` / `sync` / `destroy`. It is a read-only command.
- LLM calls inherit the global helmfile context timeout, plus a per-request timeout (default 60s).
- If the LLM call fails for any reason, doctor prints the redacted diff with a warning banner and exits 0,
so AI outages never block deployments.
- Large diffs are defensively capped at 32 KB before being sent to the LLM.
Doctor defaults `--context` to 3 (vs diff's 0) so the LLM sees enough surrounding
YAML to ground its analysis. Adjust with `--context N`.
#### Secret handling (always redacted)
Secrets are **always** redacted before any byte leaves the process. This is a
hard safety contract, not a default you can disable.
Two layers enforce it:
1. **`--show-secrets` is silently ignored.** doctor wraps the diff config so
`ShowSecrets()` always reports `false` to helm-diff, which causes
helm-diff itself to substitute `<REDACTED>` for secret values. If you
already pass `--show-secrets` in a wrapper script, doctor still redacts.
2. **Defense-in-depth text redactor.** A second pass strips any residual
secret-looking content from the captured diff before it is sent to the
LLM. It catches:
- `kind: Secret` resource blocks (full YAML mode),
- sensitive key/value lines (`password`, `token`, `apiKey`, `client_secret`,
`bearer`, …) regardless of casing or separator (`api_key` = `apiKey` =
`API_KEY`),
- free-form base64 runs of 40+ chars (encoded keys/certs),
- JWT-shaped tokens (`eyJ...`).
The redaction count is always surfaced — even when 0 — in the report footer
and JSON output, so you can confirm the redactor ran:
```markdown
---
Model: gpt-4o | Duration: 8.2s | Secrets redacted: 3
```
In `--output json` mode the field is `secrets_redacted` (always present).
If you want to drop the *structure* of Secret resources entirely (not just
their values), pass `--suppress-secrets`:
```bash
helmfile doctor --suppress-secrets
```
#### Prompt injection defense
Release names and environment values from `helmfile.yaml` are embedded in the
LLM prompt as context. Since helmfile.yaml can come from untrusted sources
(e.g. a GitOps pull request), doctor JSON-encodes these fields via
`encoding/json` before insertion. This ensures that a malicious release name
like `"foo\n\nIgnore previous instructions"` is treated as opaque string data
by the model, not as a directive.
#### Known limitations
- **Double state load**: doctor loads the helmfile state twice — once to
peek at the `llm:` block and release names, and again inside `helmfile diff`.
For large helmfiles with remote bases or heavy templating this can add
noticeable latency. Caching across loads would require touching core code,
which doctor intentionally avoids. In the unconfigured path (no LLM),
the cost equals plain `helmfile diff`.
- **`--log-output stdout`**: if you point helmfile's logger at stdout (via
`--log-output stdout`), log lines will be captured alongside the diff and
sent to the LLM. Doctor warns when it detects this configuration. Prefer
the default (stderr) when running doctor.
#### Examples
Quick local analysis (OpenAI official — no baseURL needed):
```bash
export HELMFILE_LLM_API_KEY=sk-...
export HELMFILE_LLM_MODEL=gpt-4o
helmfile doctor
```
CI gate that fails the build on high-severity risk:
```bash
helmfile --environment prod doctor --output json > doctor-report.json
# exit code 2 stops CI; --force bypasses when a human has approved the change
```
Pin the gateway per-environment via `helmfile.yaml`:
```yaml
environments:
prod:
values:
- llmGateway: https://prod-llm-gateway.internal/v1
llm:
baseURL: {{ .Values.llmGateway | quote }}
model: gpt-4o
apiKey: {{ env "HELMFILE_LLM_API_KEY" }}
```
### apply
The `helmfile apply` sub-command begins by executing `diff`. If `diff` finds that there is any changes, `sync` is executed. Adding `--interactive` instructs Helmfile to request your confirmation before `sync`.
An expected use-case of `apply` is to schedule it to run periodically, so that you can auto-fix skews between the desired and the current state of your apps running on Kubernetes clusters.
### destroy
The `helmfile destroy` sub-command uninstalls and purges all the releases defined in the manifests.
`helmfile --interactive destroy` instructs Helmfile to request your confirmation before actually deleting releases.
`destroy` basically runs `helm uninstall --purge` on all the targeted releases. If you don't want purging, use `helmfile delete` instead.
If `--skip-charts` flag is not set, destroy would prepare all releases, by fetching charts and templating them.
### delete (DEPRECATED)
The `helmfile delete` sub-command deletes all the releases defined in the manifests.
`helmfile --interactive delete` instructs Helmfile to request your confirmation before actually deleting releases.
Note that `delete` doesn't purge releases. So `helmfile delete && helmfile sync` results in sync failed due to that releases names are not deleted but preserved for future references. If you really want to remove releases for reuse, add `--purge` flag to run it like `helmfile delete --purge`.
If `--skip-charts` flag is not set, destroy would prepare all releases, by fetching charts and templating them.
### secrets
The `secrets` parameter in a `helmfile.yaml` causes the [helm-secrets](https://github.com/jkroepke/helm-secrets) plugin to be executed to decrypt the file.
To supply the secret functionality Helmfile needs the `helm secrets` plugin installed. For Helm 2.3+
you should be able to simply execute `helm plugin install https://github.com/jkroepke/helm-secrets
`.
### test
The `helmfile test` sub-command runs a `helm test` against specified releases in the manifest, default to all
Use `--cleanup` to delete pods upon completion.
### lint
The `helmfile lint` sub-command runs a `helm lint` across all of the charts/releases defined in the manifest. Non local charts will be fetched into a temporary folder which will be deleted once the task is completed.
### unittest
The `helmfile unittest` sub-command runs `helm unittest` (from the [helm-unittest plugin](https://github.com/helm-unittest/helm-unittest)) on releases that have `unitTests` defined. It automatically generates the final merged values files for each release and passes them to `helm unittest`.
This requires the `helm-unittest` plugin to be installed. You can install it with:
```bash
helm plugin install https://github.com/helm-unittest/helm-unittest
```
Releases without `unitTests` defined are skipped. Non-local charts will be fetched into a temporary folder which will be deleted once the task is completed.
Example helmfile configuration:
```yaml
releases:
- name: my-app
chart: ./charts/my-app
values:
- values.yaml
unitTests:
- tests
```
The `unitTests` paths are relative to the chart directory and follow helm-unittest conventions.
If a path does not contain glob characters, it is treated as a directory and `/*_test.yaml` is appended automatically.
You can also specify explicit glob patterns (e.g., `tests/**/*_test.yaml`).
Running `helmfile unittest` will:
1. Merge all values files defined for the release
2. Run `helm unittest ./charts/my-app --values <merged-values> --file tests/*_test.yaml`
You can pass additional flags:
```bash
# Run with additional values
helmfile unittest --values extra-values.yaml
# Run with --set overrides
helmfile unittest --set key=value
# Target specific releases
helmfile unittest --selector name=my-app
# Fail fast on first test failure
helmfile unittest --fail-fast
# Enable colored output (Helm 3 only; ignored on Helm 4 due to flag parsing issues)
helmfile unittest --color
# Enable verbose plugin output
helmfile unittest --debug-plugin
# Pass extra arguments to helm unittest
helmfile unittest --args "--strict"
```
### create
The `helmfile create` sub-command generates a helmfile deployment project scaffold with best-practice directory structure.
```bash
# Create a project in a new directory
helmfile create my-project
# Create a project in the current directory
helmfile create
# Specify a custom output directory
helmfile create my-project --output-dir /path/to/project
# Overwrite existing scaffold files
helmfile create my-project --force
```
This generates:
* `helmfile.yaml` — Main configuration with commented examples for repositories, environments, and releases
* `environments/default.yaml` — Default environment values file
* `values/.gitkeep` — Placeholder for release-specific value files
**Flags:**
| Flag | Default | Description |
|------|---------|-------------|
| `-o`, `--output-dir` | `""` | Output directory (defaults to NAME or current directory) |
| `--force` | false | Overwrite existing scaffold files |
The command validates the project name (no path separators, `.`, `..`, or whitespace-only names). Without `--force`, it atomically checks all target paths before writing to avoid partial scaffolds.
### fetch
The `helmfile fetch` sub-command downloads or copies local charts to a local directory for debug purpose. The local directory
must be specified with `--output-dir`.
### list
The `helmfile list` sub-command lists releases defined in the manifest. Optional `--output` flag accepts `json` to output releases in JSON format.
If `--skip-charts` flag is not set, list would prepare all releases, by fetching charts and templating them.
### version
The `helmfile version` sub-command prints the version of Helmfile.Optional `-o` flag accepts `json` `yaml` `short` to output version in JSON, YAML or short format.
default it will check for the latest version of Helmfile and print a tip if the current version is not the latest. To disable this behavior, set environment variable `HELMFILE_UPGRADE_NOTICE_DISABLED` to any non-empty value.
### show-dag
It prints a table with 3 columns, GROUP, RELEASE, and DEPENDENCIES.
GROUP is the unsigned, monotonically increasing integer starting from 1. All the releases with the same GROUP are deployed concurrently. Everything in GROUP 2 starts being deployed only after everything in GROUP 1 got successfully deployed.
RELEASE is the release that belongs to the GROUP.
DEPENDENCIES is the list of releases that the RELEASE depends on. It should always be empty for releases in GROUP 1. DEPENDENCIES for a release in GROUP 2 should have some or all dependencies appeared in GROUP 1. It can be "some" because Helmfile simplifies the DAGs of releases into a DAG of groups, so that Helmfile always produce a single DAG for everything written in helmfile.yaml, even when there are technically two or more independent DAGs of releases in it.
### print-env
The `helmfile print-env` sub-command prints the parsed environment configuration including merged values (with decrypted secrets). This is useful for debugging environment configuration.
```bash
# Print environment in YAML format (default)
helmfile print-env
# Print environment in JSON format
helmfile print-env --output json
# Print a specific environment
helmfile print-env -e production
```
### status
The `helmfile status` sub-command retrieves the status of releases in the state file by running `helm status` for each release.
### Additional CLI Flags
The following global flags are also available but not shown in the main help output:
| Flag | Default | Description |
|------|---------|-------------|
| `--kubeconfig` | `""` | Use a particular kubeconfig file |
| `--allow-failed-releases` | false | Continue preparing charts for other releases when chart preparation fails for a release; failed releases are skipped and all failures are reported at the end |
| `--skip-refresh` | false | Skip running `helm repo update` (lighter than `--skip-deps` which also skips dependency build) |
| `--enforce-plugin-verification` | false | Fail plugin installation if verification is not supported |
| `--oci-plain-http` | false | Use plain HTTP for OCI registries (required for local/insecure registries in Helm 4) |
| `--repo-retries` | `0` | Number of times to retry `helm repo add/update` and `helm registry login` on failure, with exponential backoff (1s, 2s, 4s, ..., capped at 30s). Set to 0 to disable retries. Overrides `HELMFILE_REPO_RETRIES` |
#### fetch flags
| Flag | Default | Description |
|------|---------|-------------|
| `--output-dir` | temp dir | Directory to store charts. If not set, a temporary directory is used and deleted when the command terminates |
| `--output-dir-template` | (default template) | Go text template for generating the output directory. Available fields: `{{ .OutputDir }}`, `{{ .ChartName }}`, `{{ .Release.* }}`, `{{ .Environment.Name }}`, `{{ .Environment.KubeContext }}`, `{{ .Environment.Values.* }}` |
| `--write-output` | false | Write a helmfile.yaml to stdout with chart references updated to point to the downloaded local chart paths. Requires `--output-dir` |
| `--concurrency` | 0 | Maximum number of concurrent helm processes to run, 0 is unlimited |
This is useful for air-gapped environments: download charts with `--output-dir` and `--write-output`, then transfer the output directory and the generated helmfile.yaml to the air-gapped environment.
#### template-args (template / apply / sync / diff)
| Flag | Default | Available on | Description |
|------|---------|--------------|-------------|
| `--template-args` | `""` | `template`, `apply`, `sync`, `diff` | Extra args appended to the helm rendering invocation. Reaches the final `helm template` for `template`; reaches both the `helm diff` rendering and chartify's internal `helm template` pre-render for `apply`/`sync`/`diff`. |
The most common use case is enabling Helm's [`lookup`](https://helm.sh/docs/chart_template_guide/functions_and_pipelines/#using-the-lookup-function) function, which queries the live cluster during rendering. `lookup` requires a server-side connection, so pass `--dry-run=server`:
```bash
# Render manifests with cluster access so lookup() resolves live values
helmfile template --template-args="--dry-run=server"
# Enable lookup() during the helm-diff phase of apply (and during diff/sync)
helmfile apply --template-args="--dry-run=server"
helmfile diff --template-args="--dry-run=server"
```
Notes:
- For `template`, the args reach the final `helm template` output.
- For `apply` and `diff`, the args reach the `helm diff upgrade` rendering so `lookup()` resolves live values during the diff phase (helm-diff supports `--dry-run=server`, which enables the lookup template function).
- For `sync`, the real `helm upgrade` already connects to the cluster, so `lookup()` works without this flag; `--template-args` is only needed there to pass additional flags to chartify's pre-render step.
- Charts that use `lookup()` should always guard against the empty result (e.g. with `default dict`), because helm renders client-side whenever it has no server connection.
To avoid passing the flag on every invocation, set it permanently under `helmDefaults`:
```yaml
helmDefaults:
templateArgs:
- --dry-run=server
```
The CLI `--template-args` flag overrides `helmDefaults.templateArgs` on a per-invocation basis (it does not merge with it), mirroring the precedence of `diffArgs`/`syncArgs`.
#### destroy flags
| Flag | Default | Description |
|------|---------|-------------|
| `--skip-charts` | false | Don't prepare charts when destroying releases |
| `--deleteWait` | false | Override helmDefaults.wait, sets `helm uninstall --wait` |
| `--deleteTimeout` | 300 | Time in seconds to wait for helm uninstall |
| `--cascade` | background | Pass cascade to helm exec |
| `--concurrency` | 0 | Maximum number of concurrent helm processes to run, 0 is unlimited |
#### list flags
| Flag | Default | Description |
|------|---------|-------------|
| `--skip-charts` | false | Don't prepare charts when listing releases |
| `--keep-temp-dir` | false | Keep temporary directory after listing |
| `--output` | `""` | Output format: `json` for JSON output |