Commit Graph
10 Commits
Author SHA1 Message Date
yxxhero 16259008d5 feat: opt-in OpenTelemetry tracing and metrics (experimental) (#2769)
* feat(telemetry): add opt-in OpenTelemetry tracing (PR 1: lifecycle + root span)

Implements the first increment of docs/proposals/otel-tracing.md (#2767):

- pkg/telemetry: SDK setup from standard OTEL_* env vars (autoexport for
  exporter selection, env-driven sampler/propagators, OTEL_SDK_DISABLED),
  command-span lifecycle, no-op-by-default accessors
- --otel-tracing flag / HELMFILE_OTEL_TRACING env switch
- root span "helmfile <command>" with file/environment/selectors/exit_code
  attributes; TRACEPARENT-based remote-parent extraction for CI correlation
- shutdown flush on both normal-exit and signal paths (nil-safe, 5s bound)
- app.New derives its context from telemetry.CommandContext()
  (Background-identical when tracing is disabled)
- docs: otel.md user guide, experimental-features entry, design proposal
- tests: hermetic unit tests, app context-contract pinning, flag registration

Telemetry problems never fail a run: exporter misconfiguration and export
errors degrade to disabled with a warning. When disabled, behavior and
performance are identical to before (no-op tracer, no goroutines, no
network).

Refs: #2767, #2758
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): trace every external process + trace-context bridges (PR 2)

Implements the second increment of docs/proposals/otel-tracing.md (#2767):

- pkg/helmexec/span.go: one span per external process started by helmfile
  (helm invocations, hooks, plugin execs) at the ShellRunner choke point —
  helm.exec (with helm.subcommand) vs os.exec, with redacted exec.args,
  exec.exit_code, and error status on failure
- pkg/helmexec/redact.go: shared argument redaction with two profiles;
  legacy is byte-identical to the historical exit-error behavior (existing
  goldens unchanged), strict (spans) additionally covers --set=k=v and
  credential flags; exit_error.go now uses the shared helper
- orphan-trace bridges with bit-identical cancellation semantics
  (context.WithoutCancel of the command context): both kubedog call sites
  (state.go) and hook execution (event.Bus gains an optional Ctx consumed
  by its default runner; state.go sets it, nil falls back to TODO as before)
- OTLP end-to-end test (in-process httptest receiver, no external
  collector): span export, error status/exit code, redaction, and
  parent-linkage to the command span
- docs/otel.md updated to the now-traced surface

Verified end-to-end with the console exporter: helmfile template on a
local chart yields the command span plus helm.exec spans for helm
version/dependency/template, all nested under it.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): state-loading and hook spans (PR 3a)

Implements the third increment of docs/proposals/otel-tracing.md (#2767):

- helmfile.discover_states around findDesiredStateFiles and helmfile.load
  around loadDesiredStateFromYamlWithBaseDir; both cover all callers
  (incl. nested helmfiles) with no signature changes
- helmfile.render / helmfile.parse children per document part, parented
  through a traceCtx field on the unexported desiredStateLoader struct
  (set once at its single construction site)
- helmfile.hook span per hook execution: Trigger's per-hook body extracted
  into runHook (readability win on its own), the hook's subprocess span
  nests under it via a per-hook ctx-swapped ShellRunner clone
  (cancellation unchanged — Bus.Ctx never carries cancellation by contract)
- pkg/telemetry/otlptest: shared in-process OTLP/HTTP receiver harness,
  now used by helmexec, event, and app span tests
- golden span-tree test at the app layer (root -> discover -> load ->
  render/parse, via the exectest fake helm) and a hook-span nesting test
- nil-ctx guard for App literals built directly by tests (App.spanParentCtx)

Verified end-to-end with the console exporter: a template run over a
gotmpl state file with a prepare hook yields the full tree with the hook's
os.exec nested under helmfile.hook.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): per-release spans nested under the load span (PR 3b)

Implements the per-release increment of docs/proposals/otel-tracing.md
(#2767) — spans nest command -> load -> release -> helm exec:

- helmexec.HelmContext gains an optional Ctx carrying the per-release span
  context; the execer's new execWithContext funnel consumes it via a
  per-call runner clone (runnerWithCtx) so the shared, cached execer is
  never mutated across concurrent workers. The seven Interface methods
  that take a HelmContext (Sync/Diff/ReleaseStatus/List/DecryptSecret/
  Delete/Test) route through it; nil Ctx behaves exactly as before.
  exec() lost its always-nil override parameter on the way (unparam).
- pkg/state/span.go: SetTraceContext + startReleaseSpan/endReleaseSpan
  helpers (release/namespace/chart/labels attributes, sorted for stable
  output); a typed-nil guard (releaseErrAsError) avoids the classic
  nil-pointer-in-interface trap on *ReleaseError.
- release spans in the worker loops: SyncReleases, DiffReleases,
  DeleteReleasesForSync, PrepareCharts, and iterateOnReleases (status/
  delete/test via a new verb parameter); their HelmContext is stamped with
  the release span context where one is built.
- pkg/app sets st.SetTraceContext(loadCtx) right after loading a state
  file, rooting all per-release spans under helmfile.load.
- bridged one more detached tracking call found on the way
  (trackReleaseIfEnabled's context.Background in the sync worker).
- golden test: release span present, nested under load, correct
  attributes; unit test for the runnerWithCtx clone semantics.

Verified with the console exporter: helmfile template yields
release.prepare(demo) under load with full attributes.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): nest status/delete/test execs under their release spans

Completes the per-release exec nesting for the iterateOnReleases-based
loops (docs/proposals/otel-tracing.md §4.4 phase 2): the do closures now
receive the release span context and stamp it into their HelmContext, so
helm status/delete/test subprocess spans nest under
helmfile.release.<verb> like sync/diff already did.

- scatterGatherReleases/iterateOnReleases/doWithReleaseSpan: do gains a
  context parameter (the release span context)
- ReleaseStatuses/DeleteReleases/TestReleases closures stamp
  HelmContext.Ctx from it
- integration test with a real execer (version-probe shim binary): the
  release's status subprocess nests under helmfile.release.status, same
  trace, with helm.subcommand=status

This also makes the otel.md claim ("upgrade, diff, delete, status, test
nested under the release span") fully accurate.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): OTel metrics — helm exec duration and release results (PR 4)

Implements the metrics increment of docs/proposals/otel-tracing.md
(#2767) on the same provider, switch, and resource as traces:

- pkg/telemetry/metrics.go: helmfile.helm.exec.duration histogram
  (subcommand, success) and helmfile.release.count counter (verb,
  result). Instruments come from the otel global meter, so recording at
  call sites is branch-free no-op when telemetry is disabled.
- Setup builds the resource once and installs both providers; reader
  selection delegates to autoexport (OTEL_METRICS_EXPORTER: otlp |
  console | prometheus | none), the OTLP reader's interval honors
  OTEL_METRIC_EXPORT_INTERVAL (read by the SDK). Shutdown flushes both
  providers (errors.Join). StartCommandSpan now carries the meter
  provider across state transitions (fixes a nil-shutdown panic).
- helmexec: finishExecSpan records exec duration for helm binaries;
  state: endReleaseSpan counts release outcomes for sync/diff/delete/
  status/test/prepare (diff counted as success when no hard error).
- otlptest: recorder routes by OTLP path (/v1/traces vs /v1/metrics)
  and decodes metrics; new FindMetric helper.
- tests: metrics recorded as no-op when disabled, provider enabled with
  the none exporter, degradation on an invalid metrics exporter, and an
  integration assertion (status exec duration datapoint + one successful
  release.count) in the shim-based state test.

Verified with the console exporter: helmfile template emits
helmfile.helm.exec.duration per subcommand (version/dependency/
template) and helmfile.release.count{verb=prepare,result=success}=1.

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: complete OTel documentation coverage

- docs/cli.md: --otel-tracing in the CLI reference help block (verbatim
  from the cobra output)
- CHANGELOG.md: [Unreleased] Added entry for tracing + metrics
- docs/index.md: Observability highlight linking docs/otel.md
- docs/proposals/otel-tracing.md: add OTEL_METRICS_EXPORTER /
  OTEL_METRIC_EXPORT_INTERVAL rows to the env-var table and note the
  periodic reader + bounded metric cardinality in §7

Refs: #2767
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: drop unused id parameter from parsePart (unparam)

The id parameter was never used inside the span wrapper; the caller's id
variable is still used for the render calls and error messages.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review — redaction gaps, kubedog valve, phantom metrics

Addresses all Copilot review comments on #2769:

Security (span payloads):
- exec.args: positional arguments are additionally passed through
  helmexec.RedactedURL, so credentials embedded in chart/repository URLs
  (AddRepo, RegistryLogin, OCI refs) are masked exactly like log output
- release spans sanitize helmfile.chart the same way
- error statuses no longer embed raw errors (which contain rendered
  commands, arguments, and subprocess output): the command span, release
  spans, hook spans, and exec spans now use generic descriptions; the
  concrete exit code remains an attribute, and RecordError on the root
  span is dropped

Correctness:
- kubedog safety valve restored: execWithContext now attaches the
  per-release span into the runner's own context instead of replacing it,
  so trackHandle.Cancel() can interrupt a wedged helm again and app
  cancellation semantics stay exactly as before the PR
- diff release spans/metrics: real failures are recorded (exit code 2
  "changes detected" still counts as success); previously every diff was
  exported as successful
- skipped releases no longer emit phantom spans and inflate
  helmfile.release.count: iterateOnReleases callers pass a skip predicate
  (skipUndesired for status/test; delete deletes undesired releases and
  passes nil)
- Setup shuts down the already-constructed tracer provider (bounded) when
  the metrics provider fails, instead of abandoning its batch goroutine

Tests: URL redaction cases (masked/untouched), spanAttachedContext
preserves the runner cancellation chain while attaching the caller's
span, skipUndesired, and a failing-hook span asserting the generic
message.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: lint — restore nolint placement and avoid nil context literal

- the skipUndesired insertion had displaced the // nolint: unparam
  directive off iterateOnReleases (helm param is intentionally unused
  there); also fixes a skipDesired/skipUndesired comment typo
- use a typed nil in TestSpanAttachedContext (staticcheck SA1012)

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review round 2 — remote-ref redaction, wrapper helm binaries, hook release attribution

Addresses all 6 new review comments on #2769:

Security (remote references):
- new helmexec.RedactedRef sanitizes go-getter style references for
  telemetry: forced-form prefixes (git::, s3::) preserved, whole URL
  userinfo masked (usernames carry tokens too), credential-bearing query
  parameters masked using pkg/remote's heuristic (token/password/secret/
  key/signature). Applied to helmfile.file (command span), helmfile.path
  (discover_states), helmfile.chart (release spans), and exec.args —
  log-time RedactedURL is untouched so log output is unchanged

Correctness:
- wrapper helm binaries (--helm-binary custom names) are now classified
  as helm operations by an explicit context marker stamped in the execer
  funnel, instead of the executable-basename heuristic; the same
  classification gates helmfile.helm.exec.duration, so the metric no
  longer misses wrapper invocations (classifyExec)
- release-scoped hooks (presync/postsync/preuninstall/postuninstall/
  cleanup in the sync/delete/diff workers) now attach their helmfile.hook
  spans to the active helmfile.release.* span via a variadic parent on
  the trigger functions; global hooks keep the command context and all
  29 existing call sites compile unchanged; hook cancellation stays
  detached (WithoutCancel) as before
- signal-terminated runs (Shutdown with exitCode 130/143 and nil error)
  now mark the command span with error status, consistent with their
  nonzero exit code

Tests: RedactedRef table (forced forms, userinfo, s3/token query params,
untouched cases), classifyExec marker case, hookTraceContext parent
attribution + non-cancellability + fallback.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): address review round 3 — redaction corner cases, value runners

Addresses 5 of the 6 new review comments on #2769 (the sixth — an
unused strings import in exit_error.go — is a false positive: Indent
still uses strings.Split/Builder and the package compiles):

- RedactArgs read the previous token from the progressively redacted
  output, so {--set, --set-string, secret} leaked the secret
  (the masked value hid the following flag). Read the previous token
  from the original input, restoring the legacy contract for adjacent
  secret flags
- RedactedRef fails closed for malformed references: URL-like refs with
  invalid percent escapes export a fully redacted value, and an
  unparseable query is dropped entirely instead of exported verbatim
- ShellRunner has value receivers, so a ShellRunner VALUE satisfies the
  Runner API; the helm marker stamping and the per-release span
  attachment now handle both value and pointer forms (matching
  WithContext), so value-runner callers keep release nesting and the
  helm.exec classification/metric

Regression tests: adjacent secret flags (legacy + strict), malformed
URL-like ref, malformed query, value-runner marker + span attachment.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): stamp the helm marker on the stdin funnel too

execStdIn (registry login, repo add) called the runner directly, so
wrapper --helm-binary names were misclassified as os.exec and omitted
from helmfile.helm.exec.duration on that path. The marking now goes
through a shared markHelmRunner helper (value and pointer ShellRunner
forms) used by both execution funnels.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): redact helm's --kube-token in strict profile

Helm's global --kube-token carries a bearer token; both the
two-argument and inline forms are now masked in span exec.args
(legacy exit-error output is untouched, matching its historical
behavior).

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* fix(telemetry): OTel metrics best-practice alignment

- helmfile.helm.exec.duration now declares explicit bucket boundaries
  tuned for seconds-scale helm invocations (5ms…600s); the SDK defaults
  are millisecond-oriented and lumped every sub-5s invocation — the
  common case — into the first bucket, defeating the histogram
- instruments are re-created under the installed provider with the
  instrumentation scope version stamped (Setup-time, race-free)
- helmfile.release.count declares the {release} curly-annotation unit
  per the metrics naming conventions

Tested end-to-end via the OTLP integration test: exported bounds are
the tuned set, units are asserted, and the scope carries the version.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* feat(telemetry): per-release duration metrics behind an opt-in switch

New helmfile.release.duration histogram (seconds, same tuned buckets)
with bounded dimensions by default (verb, result). Setting
HELMFILE_OTEL_METRICS_PER_RELEASE=true adds helmfile.release and
helmfile.namespace, answering "which release is slow" from dashboards:

- well-suited to bounded CI runs; long-lived centralized collection
  needs a backend capacity/TTL story (documented in docs/otel.md)
- per-release timing remains available in traces without the flag
- env read per call (release operations are low-frequency, and tests
  toggle it)

endReleaseSpan now takes the release and the operation start time; the
five worker-loop call sites pass them (doWithReleaseSpan, SyncReleases,
DeleteReleasesForSync, PrepareCharts, DiffReleases).

Verified end-to-end with the console exporter (default dims vs
per-release) and OTLP integration tests pinning both modes.

Refs: #2767, #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): maintainability pass over the runner/metric plumbing

- StartCommandSpan copies the tracingState struct instead of enumerating
  fields by hand — that pattern dropped the meter provider once already
- the two value/pointer ShellRunner switches (helm marker, span
  attachment) are unified into one withRunnerCtx helper; the duplication
  caused two review rounds of value-form misses
- classifyExec derives the helm classification from the span name
  (helmExecSpanName constant) instead of returning a third parallel bool
- metrics: shared outcomeAttrs for the verb/result dimensions, and the
  bucket slice renamed to durationBuckets with a comment covering both
  histograms that use it

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): consolidate test env lists, trace bridges, and hook prep; sync the design doc

Maintainability:
- HermeticEnvVars is now exported from pkg/telemetry (the owner of the
  env surface) and used by both telemetry tests and otlptest — the two
  copies had already drifted once (HELMFILE_OTEL_METRICS_PER_RELEASE
  needed updating in both)
- kubedogTraceContext and hookTraceContext were the same concept written
  twice; unified into traceOnlyContext(parent...) in span.go

Readability:
- runHook's nested kubectl rewrite extracted into prepareKubectlHook
  with guard-clause structure

Accuracy (docs ↔ code, drifted over five review rounds):
- §4.4 now describes the implemented mechanism: the release span is
  INJECTED into the runner's own context (preserving the kubedog safety
  valve) rather than the runner context being replaced, and helm
  classification is marker-based for wrapper binaries
- §5 exec span rows list the actual attributes incl. URL/query masking
- §6 strict profile documents RedactedRef, --kube-token, and the
  adjacent-token guarantee

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): drop the dead noop state, relocate skipUndesired, sync user-facing accuracy

- tracingState.noop was dead weight in the enabled state and a
  copy-surface in every transition; a single package-level
  noopTracerProvider now backs Tracer while disabled
- skipUndesired moved next to doWithReleaseSpan in span.go, its only
  conceptual home (span/metric suppression, not run plumbing)
- accuracy: the package doc, --otel-tracing flag help,
  experimental-features entry, and CHANGELOG now all say tracing AND
  metrics and list the third instrument (helmfile.release.duration with
  the HELMFILE_OTEL_METRICS_PER_RELEASE opt-in) — these had drifted
  when the metric was added; the PR description's metric table is
  updated to match as well

No behavior change; full -race suite green (except the pre-existing
network-dependent TestStorage_resolveFile flake), lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): flatten Setup, name the prefix bound, dedupe test fake; fix instrument-count drift

Readability/maintainability:
- Setup drops from 56 to 39 lines: provider construction (including the
  shutdown-tracer-on-meter-failure recovery) moves to newProviders in
  exporter.go next to the constructors it composes
- refredact's magic 16 becomes maxForcedFormPrefix with a comment
- span_test's hand-rolled fakeRunner removed in favor of the existing
  mockRunner (same package)

Accuracy:
- "Two instruments" wording survived in docs/otel.md and the design
  proposal §7 after helmfile.release.duration was added; both now say
  three and mention the per-release opt-in

No behavior change; full -race suite green (except the pre-existing
network flake), lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor(telemetry): co-locate span machinery, drop a dead export, fix docs nits

- the span plumbing helpers (markHelmExec, withRunnerCtx,
  markHelmRunner, spanAttachedContext) move from exec.go to span.go,
  next to the marker type and classifiers they serve — exec.go keeps
  only the funnel call sites
- otlptest.SpanNames was never used outside the package; unexported
- isHelmBinary's comment now states it is the FALLBACK classifier
  (funnel invocations are marker-classified), replacing the outdated
  "cosmetic distinction" framing from before the marker existed
- docs/otel.md: release-scoped hooks nest under their release span
  (added in review round 2, never documented)

No behavior change; full -race suite green, lint clean.

Refs: #2769
Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-09-07 20:34:10 +08:00
34ead21107 feat: allow helmfile to continue on failed releases (#2616)
* feat: allow helmfile to continue on failed releases

Co-authored-by: Peter Honeder <peter.honeder@unwired.at>
Signed-off-by: Niklas Ott <niklas.ott@unwired.at>

* fix: skip failed-prep releases, complete flag wiring, add tests and docs (#64)

Review follow-ups for --allow-failed-releases (#2616):

- Track per-release chart preparation failures in PrepareCharts (returned
  as a map keyed by release) and remove those releases from the state in
  Run.WithPreparedCharts when --allow-failed-releases is set, so a failed
  release is never executed against its original, un-prepared chart
  reference (which could either fail again with a duplicate error or, for
  charts requiring chartify, bypass patches/dependency modifications and
  produce an unintended result). All failures are still reported at the
  end via the aggregated MultiError.
- Complete the release identity on error results from
  prepareChartForRelease so failures are attributed to the correct
  release.
- With --allow-failed-releases, continue building dependencies of the
  remaining charts when 'helm dep build' fails for one release, and skip
  the affected releases during execution.
- Wire --allow-failed-releases into 'helmfile unittest' and 'helmfile
  status'; remove the dead flag wiring for write-values and list (both
  never prepare charts, see commandsSkipChartPrep).
- Simplify control flow (guard clauses, errors.As, drop dead code and
  redundant else branches).
- Add end-to-end coverage in pkg/app/issue_2616_test.go and extend the
  state-level tests; document the flag in docs/cli.md and CHANGELOG.md.

Signed-off-by: yxxhero <11087727+yxxhero@users.noreply.github.com>

---------

Signed-off-by: Niklas Ott <niklas.ott@unwired.at>
Signed-off-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
Co-authored-by: Peter Honeder <peter.honeder@unwired.at>
Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com>
2026-09-03 06:26:46 +08:00
henrichter-sap 7cf4ff9a25 feat: add --skip-diff-validation-on-install CLI flag (#2728)
Signed-off-by: Richter <h.richter@sap.com>
2026-08-03 21:11:45 +08:00
yxxhero 7bebfab71a feat: add --repo-retries for helm repo and registry login commands (#2683)
* feat: add --repo-retries for retrying helm repo and registry login commands

Add a configurable retry mechanism for chart repository operations to
handle unstable networks (corporate proxies, slow internal registries).

Closes #1894

- New --repo-retries N flag and HELMFILE_REPO_RETRIES env var (default 0
  = opt-in, backward compatible)
- Retry applies to helm repo add, helm repo update (incl. ACR), and
  helm registry login with exponential backoff (1s, 2s, 4s, ..., capped 30s)
- Single retryRepoOp helper; per-attempt args/buffer are local to avoid
  state leaking across retries
- Tests cover succeed-after-retry, exhausted-retries, disabled-by-default,
  and regression guards for password-buffer and args-accumulation

Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: address PR review (overflow guard, cancellable sleep, flag-override, docs)

Address Copilot review feedback on #2683:

- Cap backoff shift exponent at 5 to prevent time.Duration overflow on
  large --repo-retries values
- Make retry sleep context-aware (sleepCtx) so Ctrl+C aborts the retry
  loop promptly via the ShellRunner context
- Log a concise exit status instead of the verbose ExitError dump, and
  clarify the retry-counter wording ('retry N/M')
- Use -1 sentinel as the CLI default so --repo-retries=0 can explicitly
  disable retries even when HELMFILE_REPO_RETRIES is set
- Align help text and docs: retry applies 'on failure' (not just
  transient errors), document the 0-disables behavior
- Add tests for overflow guard, cancellable sleep, and flag-zero-disables

Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: abort retries on canceled context, hide sentinel default, align comment

Address follow-up Copilot review on #2683:

- Fix tight-loop bug: sleepCtx now returns whether it completed vs was
  interrupted by context cancellation, and retryRepoOp aborts the retry
  loop on interruption so Ctrl+C no longer spins into rapid helm calls
- Hide the -1 sentinel from --help by overriding the displayed default
  to 0 (pflag DefValue), matching the documented default while keeping
  the flag-override semantics
- Correct HelmExecOptions.RepoRetry comment: 'on failure' not 'transient
  network errors', matching the actual retry behavior
- Add Test_Retry_AbortsOnCanceledContext covering the no-tight-loop path

Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: copy args per retry in RegistryLogin, make cancel test deterministic

Address follow-up Copilot review on #2683:

- RegistryLogin: pass a per-attempt copy of args to execStdIn so its
  internal append (for helm.extra) can't alias the shared slice across
  retries
- Test_Retry_AbortsOnCanceledContext: cancel the context deterministically
  inside the op closure after the first attempt, replacing the flaky
  time.Sleep(20ms) goroutine

Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: return error on unknown managed repo type instead of silent skip

Address Copilot review on #2683: AddRepo logged an error for an unknown
managed type but returned nil, silently succeeding while skipping the
repo add. Now returns an error so misconfigurations fail loudly.

Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-07-23 18:10:04 +08:00
yxxhero e3f757d5ed feat: add --template-args flag to template/apply/sync for helm lookup() support (#2666)
* feat: add --template-args to enable helm lookup() during template/apply/sync (#1833)

Add a --template-args flag to the template, apply, and sync subcommands so
extra args (most notably --dry-run=server) can be passed to the helm template
invocation, enabling Helm's lookup() function to resolve live cluster values.

- template: --template-args reaches both chartify's pre-render helm template
  and the final helm template output (flagsForTemplate).
- apply/sync: --template-args reaches chartify's pre-render helm template.
  apply/sync already inject --dry-run=server automatically for cluster
  operations; the flag is an explicit opt-in for the template subcommand or
  for passing additional flags.
- When --dry-run is present in template args, kube-context/kubeconfig are also
  injected into chartify so lookup() can actually reach the cluster.
- Resolves the long-stale PR #1833 rebased onto current main, which already
  contains the cluster-connectivity infrastructure (issues #2271, #2309,
  #2355, #2444).
- Includes integration test (lookup.sh) covering both chartify and
  non-chartify scenarios.

Signed-off-by: yxxhero <aiopsclub@163.com>

* test: make lookup template nil-safe to fix integration CI

The lookup() function returns an empty map when the chart is rendered
without a cluster connection (notably the helm-diff phase of `helmfile
apply`). The original fixture chained `index` over the lookup result,
panicking with "index of untyped nil" during apply's diff rendering.

Guard every index with `default dict` so the template falls back to
"overwritten" when lookup is empty, while still resolving to the live
value ("init") when cluster access is available (--dry-run=server via
--template-args, or a real helm upgrade).

Signed-off-by: yxxhero <aiopsclub@163.com>

* feat: enable lookup() during apply/diff via --template-args in helm-diff

Thread --template-args into the helm-diff rendering path so that
`helmfile apply`/`diff --template-args="--dry-run=server"` resolves
Helm's lookup() function during the diff phase too. helm-diff supports
`--dry-run=server`, which explicitly "enables the cluster access ...
and the lookup template function".

Previously --template-args only reached chartify's pre-render (which is a
no-op for plain charts due to chartify's early-return when there is no
forceNamespace/patches/injections) and the final `helm template` of the
`template` subcommand. As a result `helmfile apply` on a lookup chart
rendered client-side during the diff phase.

Changes:
- pkg/state: add TemplateArgs to DiffOpts; append it in appendExtraDiffFlags
  (reaches every helm-diff invocation: apply, standalone diff, interactive
  sync), mirroring the existing flagsForTemplate handling.
- pkg/config + cmd: add --template-args to the diff/doctor commands and to
  DiffConfigProvider, so lookup works for `helmfile diff` as well.
- pkg/app: populate DiffOpts.TemplateArgs from apply/diff/sync-interactive.
- docs/cli.md: correct the previous overpromising wording and document diff
  support plus the nil-safe lookup guidance.
- tests: unit-test the TemplateArgs handling in appendExtraDiffFlags and
  flagsForTemplate; integration lookup.sh now exercises apply with
  --template-args="--dry-run=server".

Signed-off-by: yxxhero <aiopsclub@163.com>

* refactor: de-duplicate chartify template-args logic, add helmDefaults.templateArgs

Address review feedback on #2666:

1. Eliminate stale duplicated test helpers (issue_2444_test.go, issue_2355_test.go).
   Both files intentionally copied the processChartification flag-building logic
   with explicit SYNC WARNING comments, then drifted out of sync when #2666
   refactored the production code (needsKubeConnection gate, user-args merge).
   Extract the real logic into pure, unit-tested helpers
   (buildChartifyTemplateArgs, commandRequiresCluster) and delete the copies.

2. Add unit coverage for the new chartify merge path: template +
   --template-args=--dry-run=server now triggers kubeconfig/kube-context
   injection (TestTemplateArgsDryRunTriggersKubeInjection,
   TestTemplateArgsMergedBeforeInjection) — previously only covered by the
   cluster-dependent integration test.

3. Add a negative integration case (lookup.sh assert_template_fallback)
   verifying lookup() falls back to the default value WITHOUT --template-args,
   guarding against a regression that silently always connects to the cluster.

4. Add helmDefaults.templateArgs for parity with diffArgs/syncArgs, so users
   can enable lookup() support permanently instead of passing the flag on every
   invocation. CLI --template-args overrides (does not merge with) the default.
   Resolved via effectiveTemplateArgs, wired into the chartify, flagsForTemplate,
   and appendExtraDiffFlags paths.

5. Minor: capitalize --template-args help text to match surrounding flags;
   document helmDefaults.templateArgs precedence in docs/cli.md.

Signed-off-by: yxxhero <aiopsclub@163.com>

* test: cover helmDefaults->chartify composition; fix helm helm-diff typo

Address remaining review nits on #2666:

- Add TestHelmDefaultsTemplateArgsReachesChartify, a belt-and-suspenders test
  for the processChartification composition (effectiveTemplateArgs ->
  buildChartifyTemplateArgs), closing the last unit-level coverage gap for
  helmDefaults.templateArgs reaching the chartify path.

- Fix pre-existing typo in cmd/bind_diff_flags.go: 'pass args to helm helm-diff'
  -> 'Pass args to helm-diff' (doubled 'helm', lowercase).

Signed-off-by: yxxhero <aiopsclub@163.com>

* fix: correct 'helm helm-diff' typo in apply --diff-args help text

Sibling of the bind_diff_flags.go fix; the same doubled-'helm' typo and
lowercase help existed in cmd/apply.go's --diff-args registration, leaving
the apply and diff/doctor help strings inconsistent.

Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: add helmDefaults.templateArgs to configuration reference

The complete helmfile.yaml schema in docs/configuration.md documents
diffArgs and syncArgs under helmDefaults but was missing the new
templateArgs field added in #2666. Add it beside syncArgs for
discoverability, noting the --template-args CLI override.

Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-06-28 12:19:33 +08:00
yxxhero 9b943adc9e feat: add helmfile doctor command for AI-assisted diff analysis (#2660)
* feat: add `helmfile doctor` command for AI-assisted diff analysis

`helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to
summarize the changes and flag risks (data loss, security exposure, breaking
changes, downtime, performance, best-practice issues).

Key design decisions:
- When no LLM is configured, doctor is equivalent to `helmfile diff` with
  one exception: --show-secrets is always forced off (secrets never reach
  stdout, even without an LLM).
- Secrets are ALWAYS redacted via two layers: (1) ShowSecrets() forced to
  false so helm-diff emits <REDACTED> placeholders; (2) a defense-in-depth
  text redactor strips residual secret-looking content (Secret YAML blocks,
  sensitive key/value lines, base64 blobs, JWT tokens) before LLM transmission.
- LLM configuration precedence: env (HELMFILE_LLM_*) < helmfile.yaml (llm:)
  < CLI flags (--llm-*).
- Supports any OpenAI-compatible backend (OpenAI, Azure, One-API, LiteLLM,
  Ollama, etc.) with automatic response_format fallback for backends that
  don't support JSON mode.
- Prompt injection defense: release names and environment values are
  JSON-encoded before insertion into the LLM prompt.
- Exit codes: 0 (success/low-risk), 2 (high-risk gate, bypass with --force),
  1 (other errors). Helm-diff's 'detected changes' exit-2 is swallowed.

New packages:
- pkg/agent/llm: OpenAI-compatible client with JSON response parsing, mock
  client for testing, prompt builder with injection defense.
- pkg/agent/doctor: secret redactor (state machine + regex), report renderer
  (markdown + JSON), config resolver (env < yaml < flag merge).

Testing: 70+ unit tests covering redaction patterns, prompt injection,
response_format fallback, JSON parsing, yaml roundtrip, concurrency safety,
panic recovery, and error propagation. go test -race passes.

Documentation: full doctor section in docs/cli.md, llm: block reference in
docs/configuration.md, updated skills/helmfile for AI agents.

Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: fix doctor equivalence wording per PR review

Per review feedback (PR #2660): the docs claimed doctor is 'equivalent to
helmfile diff — same flags, same output, same exit codes' in the unconfigured
path, but this over-promises because:

  1. doctor --output is the report format (not helm-diff's output format)
  2. helm-diff's --output is exposed as --diff-output in doctor
  3. --show-secrets is silently ignored

Updated all three locations (cli.md, cmd/doctor.go Long + godoc, pkg/app/doctor.go
godoc) to say 'falls back to helmfile diff with --show-secrets forced off' and
explicitly note the --output / --diff-output flag difference.

Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-06-22 16:52:35 +08:00
Dominik Schmidt 33eadc993e feat: support HELMFILE_* env vars for more global flags (#2606)
* feat: support more HELMFILE_* env vars as flag fallbacks

Adds env-var fallbacks for global flags, mirroring the existing
HELMFILE_ENVIRONMENT / HELMFILE_KUBE_CONTEXT pattern:

* --helm-binary       -> HELMFILE_HELM_BINARY
* --kustomize-binary  -> HELMFILE_KUSTOMIZE_BINARY
* --log-level         -> HELMFILE_LOG_LEVEL
* --debug             -> HELMFILE_DEBUG       (expecting "true" lower case)
* --quiet             -> HELMFILE_QUIET       (expecting "true" lower case)
* --no-color          -> HELMFILE_NO_COLOR    (expecting "true" lower case),
                         additionally honors NO_COLOR per no-color.org
                         (any non-empty value disables color)

Flag values still take precedence; env vars are consulted only when the
flag is unset. The string-flag default values ("helm", "kustomize",
"info") move into the accessor methods so the env-var fallback can
actually trigger when no flag is passed.

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

* docs: mention new HELMFILE_* env vars in cli.md and templating.md

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

* fix: make Color/NoColor/env interaction consistent

Two issues with the env-aware NoColor() introduced together with
HELMFILE_NO_COLOR / NO_COLOR support:

1. Color() consulted the raw GlobalOptions.NoColor field instead of
   NoColor(), so in a TTY with only the env set, Color() fell through
   to terminal autodetect and ValidateConfig() spuriously errored with
   "--color and --no-color cannot be specified at the same time".

2. NoColor() returned true via env even when --color was explicitly
   passed, so `helmfile --color` with NO_COLOR (or HELMFILE_NO_COLOR=true)
   in the environment hit the same ValidateConfig() error. A flag should
   always win over an env var.

Fix both by routing Color() through NoColor() and giving NoColor() an
explicit --color short-circuit. Regression tests added for both paths.

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

---------

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
2026-05-22 09:16:52 +08:00
Dominik Schmidt 31ac918512 feat: support HELMFILE_NAMESPACE env var for default namespace (#2592)
* feat: support HELMFILE_NAMESPACE env var for default namespace

Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --namespace
CLI flag takes precedence, falling back to HELMFILE_NAMESPACE when
unset.

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

* docs: mention HELMFILE_NAMESPACE in cli.md and templating.md

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

---------

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
2026-05-19 21:43:11 +08:00
Dominik Schmidt c15cbb096a feat: support HELMFILE_KUBE_CONTEXT env var for default kube context (#2593)
* feat: support HELMFILE_KUBE_CONTEXT env var for default kube context

Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --kube-context
CLI flag takes precedence, falling back to HELMFILE_KUBE_CONTEXT when
unset.

Refs #1213.

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

* docs: mention HELMFILE_KUBE_CONTEXT in cli.md and templating.md

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>

---------

Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
2026-05-19 20:43:28 +08:00
yxxhero e703b15075 docs: restructure documentation and improve newcomer experience (#2573)
* feat: add --write-output flag to helmfile fetch for air-gapped environments

Add --write-output flag to helmfile fetch that outputs a modified
helmfile.yaml with chart references updated to point to downloaded
local chart paths. Combined with --output-dir, this enables preparing
all charts for deployment in air-gapped environments.

Usage:
  helmfile fetch --output-dir ./charts --write-output > helmfile-airgapped.yaml

Fixes #2571

Signed-off-by: yxxhero <yxxhero@users.noreply.github.com>
Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: restructure documentation and improve newcomer experience

Split the monolithic index.md (1990 lines) into focused topic pages,
update mkdocs.yml navigation, and add missing documentation for
undocumented code features.

Structure changes:
- Extract configuration.md (helmfile.yaml reference)
- Extract cli.md (CLI commands and flags)
- Extract templating.md (template syntax and env vars)
- Extract environments.md (environment configuration)
- Extract releases.md (DAG, needs, selectors)
- Extract hooks.md (lifecycle hooks)
- Extract integrations.md (ArgoCD, Azure ACR, OCI)
- Slim index.md to ~270 line landing page with step-by-step tutorial

Newcomer improvements:
- Add 5-step Getting Started tutorial with explanations
- Reorganize nav: Getting Started now shows core learning path
  (Writing Helmfile → Values → Environments → Releases)
- Add Quick Reference table to configuration.md
- Simplify writing-helmfile.md title

Code-vs-docs gap fixes:
- Document 23 undocumented release fields (valuesTemplate,
  setTemplate, forceNamespace, adopt, trackMode, etc.)
- Document 6 undocumented helmDefaults fields (enableDNS,
  forceConflicts, skipRefresh, takeOwnership, etc.)
- Document print-env command and missing CLI flags
- Document kubectlApply hook field
- Document environment defaults field and merge order
- Document kubedogQPS/kubedogBurst advanced settings
- Document template partials (_*.tpl) auto-loading

Cleanup:
- Fix Docker image version from v0.156.0 to v1.1.0
- Fix heading nesting in advanced-features.md
- Update experimental-features.md with current features
- Fix broken cross-references and anchor links

Signed-off-by: yxxhero <aiopsclub@163.com>

* Revert changes to pkg/app from docs/restructure-and-improve branch

Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: add create subcommand to README and CLI reference

Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <yxxhero@users.noreply.github.com>
Signed-off-by: yxxhero <aiopsclub@163.com>
2026-05-03 19:33:33 +08:00