mirror of
https://github.com/helmfile/helmfile.git
synced 2026-09-30 19:24:03 +02:00
ad81e4231e40fea9798ccc6dc93e2c0363e95e7d
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
16259008d5 |
feat: opt-in OpenTelemetry tracing and metrics (experimental) (#2769)
* feat(telemetry): add opt-in OpenTelemetry tracing (PR 1: lifecycle + root span) Implements the first increment of docs/proposals/otel-tracing.md (#2767): - pkg/telemetry: SDK setup from standard OTEL_* env vars (autoexport for exporter selection, env-driven sampler/propagators, OTEL_SDK_DISABLED), command-span lifecycle, no-op-by-default accessors - --otel-tracing flag / HELMFILE_OTEL_TRACING env switch - root span "helmfile <command>" with file/environment/selectors/exit_code attributes; TRACEPARENT-based remote-parent extraction for CI correlation - shutdown flush on both normal-exit and signal paths (nil-safe, 5s bound) - app.New derives its context from telemetry.CommandContext() (Background-identical when tracing is disabled) - docs: otel.md user guide, experimental-features entry, design proposal - tests: hermetic unit tests, app context-contract pinning, flag registration Telemetry problems never fail a run: exporter misconfiguration and export errors degrade to disabled with a warning. When disabled, behavior and performance are identical to before (no-op tracer, no goroutines, no network). Refs: #2767, #2758 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): trace every external process + trace-context bridges (PR 2) Implements the second increment of docs/proposals/otel-tracing.md (#2767): - pkg/helmexec/span.go: one span per external process started by helmfile (helm invocations, hooks, plugin execs) at the ShellRunner choke point — helm.exec (with helm.subcommand) vs os.exec, with redacted exec.args, exec.exit_code, and error status on failure - pkg/helmexec/redact.go: shared argument redaction with two profiles; legacy is byte-identical to the historical exit-error behavior (existing goldens unchanged), strict (spans) additionally covers --set=k=v and credential flags; exit_error.go now uses the shared helper - orphan-trace bridges with bit-identical cancellation semantics (context.WithoutCancel of the command context): both kubedog call sites (state.go) and hook execution (event.Bus gains an optional Ctx consumed by its default runner; state.go sets it, nil falls back to TODO as before) - OTLP end-to-end test (in-process httptest receiver, no external collector): span export, error status/exit code, redaction, and parent-linkage to the command span - docs/otel.md updated to the now-traced surface Verified end-to-end with the console exporter: helmfile template on a local chart yields the command span plus helm.exec spans for helm version/dependency/template, all nested under it. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): state-loading and hook spans (PR 3a) Implements the third increment of docs/proposals/otel-tracing.md (#2767): - helmfile.discover_states around findDesiredStateFiles and helmfile.load around loadDesiredStateFromYamlWithBaseDir; both cover all callers (incl. nested helmfiles) with no signature changes - helmfile.render / helmfile.parse children per document part, parented through a traceCtx field on the unexported desiredStateLoader struct (set once at its single construction site) - helmfile.hook span per hook execution: Trigger's per-hook body extracted into runHook (readability win on its own), the hook's subprocess span nests under it via a per-hook ctx-swapped ShellRunner clone (cancellation unchanged — Bus.Ctx never carries cancellation by contract) - pkg/telemetry/otlptest: shared in-process OTLP/HTTP receiver harness, now used by helmexec, event, and app span tests - golden span-tree test at the app layer (root -> discover -> load -> render/parse, via the exectest fake helm) and a hook-span nesting test - nil-ctx guard for App literals built directly by tests (App.spanParentCtx) Verified end-to-end with the console exporter: a template run over a gotmpl state file with a prepare hook yields the full tree with the hook's os.exec nested under helmfile.hook. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): per-release spans nested under the load span (PR 3b) Implements the per-release increment of docs/proposals/otel-tracing.md (#2767) — spans nest command -> load -> release -> helm exec: - helmexec.HelmContext gains an optional Ctx carrying the per-release span context; the execer's new execWithContext funnel consumes it via a per-call runner clone (runnerWithCtx) so the shared, cached execer is never mutated across concurrent workers. The seven Interface methods that take a HelmContext (Sync/Diff/ReleaseStatus/List/DecryptSecret/ Delete/Test) route through it; nil Ctx behaves exactly as before. exec() lost its always-nil override parameter on the way (unparam). - pkg/state/span.go: SetTraceContext + startReleaseSpan/endReleaseSpan helpers (release/namespace/chart/labels attributes, sorted for stable output); a typed-nil guard (releaseErrAsError) avoids the classic nil-pointer-in-interface trap on *ReleaseError. - release spans in the worker loops: SyncReleases, DiffReleases, DeleteReleasesForSync, PrepareCharts, and iterateOnReleases (status/ delete/test via a new verb parameter); their HelmContext is stamped with the release span context where one is built. - pkg/app sets st.SetTraceContext(loadCtx) right after loading a state file, rooting all per-release spans under helmfile.load. - bridged one more detached tracking call found on the way (trackReleaseIfEnabled's context.Background in the sync worker). - golden test: release span present, nested under load, correct attributes; unit test for the runnerWithCtx clone semantics. Verified with the console exporter: helmfile template yields release.prepare(demo) under load with full attributes. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): nest status/delete/test execs under their release spans Completes the per-release exec nesting for the iterateOnReleases-based loops (docs/proposals/otel-tracing.md §4.4 phase 2): the do closures now receive the release span context and stamp it into their HelmContext, so helm status/delete/test subprocess spans nest under helmfile.release.<verb> like sync/diff already did. - scatterGatherReleases/iterateOnReleases/doWithReleaseSpan: do gains a context parameter (the release span context) - ReleaseStatuses/DeleteReleases/TestReleases closures stamp HelmContext.Ctx from it - integration test with a real execer (version-probe shim binary): the release's status subprocess nests under helmfile.release.status, same trace, with helm.subcommand=status This also makes the otel.md claim ("upgrade, diff, delete, status, test nested under the release span") fully accurate. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): OTel metrics — helm exec duration and release results (PR 4) Implements the metrics increment of docs/proposals/otel-tracing.md (#2767) on the same provider, switch, and resource as traces: - pkg/telemetry/metrics.go: helmfile.helm.exec.duration histogram (subcommand, success) and helmfile.release.count counter (verb, result). Instruments come from the otel global meter, so recording at call sites is branch-free no-op when telemetry is disabled. - Setup builds the resource once and installs both providers; reader selection delegates to autoexport (OTEL_METRICS_EXPORTER: otlp | console | prometheus | none), the OTLP reader's interval honors OTEL_METRIC_EXPORT_INTERVAL (read by the SDK). Shutdown flushes both providers (errors.Join). StartCommandSpan now carries the meter provider across state transitions (fixes a nil-shutdown panic). - helmexec: finishExecSpan records exec duration for helm binaries; state: endReleaseSpan counts release outcomes for sync/diff/delete/ status/test/prepare (diff counted as success when no hard error). - otlptest: recorder routes by OTLP path (/v1/traces vs /v1/metrics) and decodes metrics; new FindMetric helper. - tests: metrics recorded as no-op when disabled, provider enabled with the none exporter, degradation on an invalid metrics exporter, and an integration assertion (status exec duration datapoint + one successful release.count) in the shim-based state test. Verified with the console exporter: helmfile template emits helmfile.helm.exec.duration per subcommand (version/dependency/ template) and helmfile.release.count{verb=prepare,result=success}=1. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * docs: complete OTel documentation coverage - docs/cli.md: --otel-tracing in the CLI reference help block (verbatim from the cobra output) - CHANGELOG.md: [Unreleased] Added entry for tracing + metrics - docs/index.md: Observability highlight linking docs/otel.md - docs/proposals/otel-tracing.md: add OTEL_METRICS_EXPORTER / OTEL_METRIC_EXPORT_INTERVAL rows to the env-var table and note the periodic reader + bounded metric cardinality in §7 Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * fix: drop unused id parameter from parsePart (unparam) The id parameter was never used inside the span wrapper; the caller's id variable is still used for the render calls and error messages. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review — redaction gaps, kubedog valve, phantom metrics Addresses all Copilot review comments on #2769: Security (span payloads): - exec.args: positional arguments are additionally passed through helmexec.RedactedURL, so credentials embedded in chart/repository URLs (AddRepo, RegistryLogin, OCI refs) are masked exactly like log output - release spans sanitize helmfile.chart the same way - error statuses no longer embed raw errors (which contain rendered commands, arguments, and subprocess output): the command span, release spans, hook spans, and exec spans now use generic descriptions; the concrete exit code remains an attribute, and RecordError on the root span is dropped Correctness: - kubedog safety valve restored: execWithContext now attaches the per-release span into the runner's own context instead of replacing it, so trackHandle.Cancel() can interrupt a wedged helm again and app cancellation semantics stay exactly as before the PR - diff release spans/metrics: real failures are recorded (exit code 2 "changes detected" still counts as success); previously every diff was exported as successful - skipped releases no longer emit phantom spans and inflate helmfile.release.count: iterateOnReleases callers pass a skip predicate (skipUndesired for status/test; delete deletes undesired releases and passes nil) - Setup shuts down the already-constructed tracer provider (bounded) when the metrics provider fails, instead of abandoning its batch goroutine Tests: URL redaction cases (masked/untouched), spanAttachedContext preserves the runner cancellation chain while attaching the caller's span, skipUndesired, and a failing-hook span asserting the generic message. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix: lint — restore nolint placement and avoid nil context literal - the skipUndesired insertion had displaced the // nolint: unparam directive off iterateOnReleases (helm param is intentionally unused there); also fixes a skipDesired/skipUndesired comment typo - use a typed nil in TestSpanAttachedContext (staticcheck SA1012) Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review round 2 — remote-ref redaction, wrapper helm binaries, hook release attribution Addresses all 6 new review comments on #2769: Security (remote references): - new helmexec.RedactedRef sanitizes go-getter style references for telemetry: forced-form prefixes (git::, s3::) preserved, whole URL userinfo masked (usernames carry tokens too), credential-bearing query parameters masked using pkg/remote's heuristic (token/password/secret/ key/signature). Applied to helmfile.file (command span), helmfile.path (discover_states), helmfile.chart (release spans), and exec.args — log-time RedactedURL is untouched so log output is unchanged Correctness: - wrapper helm binaries (--helm-binary custom names) are now classified as helm operations by an explicit context marker stamped in the execer funnel, instead of the executable-basename heuristic; the same classification gates helmfile.helm.exec.duration, so the metric no longer misses wrapper invocations (classifyExec) - release-scoped hooks (presync/postsync/preuninstall/postuninstall/ cleanup in the sync/delete/diff workers) now attach their helmfile.hook spans to the active helmfile.release.* span via a variadic parent on the trigger functions; global hooks keep the command context and all 29 existing call sites compile unchanged; hook cancellation stays detached (WithoutCancel) as before - signal-terminated runs (Shutdown with exitCode 130/143 and nil error) now mark the command span with error status, consistent with their nonzero exit code Tests: RedactedRef table (forced forms, userinfo, s3/token query params, untouched cases), classifyExec marker case, hookTraceContext parent attribution + non-cancellability + fallback. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review round 3 — redaction corner cases, value runners Addresses 5 of the 6 new review comments on #2769 (the sixth — an unused strings import in exit_error.go — is a false positive: Indent still uses strings.Split/Builder and the package compiles): - RedactArgs read the previous token from the progressively redacted output, so {--set, --set-string, secret} leaked the secret (the masked value hid the following flag). Read the previous token from the original input, restoring the legacy contract for adjacent secret flags - RedactedRef fails closed for malformed references: URL-like refs with invalid percent escapes export a fully redacted value, and an unparseable query is dropped entirely instead of exported verbatim - ShellRunner has value receivers, so a ShellRunner VALUE satisfies the Runner API; the helm marker stamping and the per-release span attachment now handle both value and pointer forms (matching WithContext), so value-runner callers keep release nesting and the helm.exec classification/metric Regression tests: adjacent secret flags (legacy + strict), malformed URL-like ref, malformed query, value-runner marker + span attachment. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): stamp the helm marker on the stdin funnel too execStdIn (registry login, repo add) called the runner directly, so wrapper --helm-binary names were misclassified as os.exec and omitted from helmfile.helm.exec.duration on that path. The marking now goes through a shared markHelmRunner helper (value and pointer ShellRunner forms) used by both execution funnels. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): redact helm's --kube-token in strict profile Helm's global --kube-token carries a bearer token; both the two-argument and inline forms are now masked in span exec.args (legacy exit-error output is untouched, matching its historical behavior). Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): OTel metrics best-practice alignment - helmfile.helm.exec.duration now declares explicit bucket boundaries tuned for seconds-scale helm invocations (5ms…600s); the SDK defaults are millisecond-oriented and lumped every sub-5s invocation — the common case — into the first bucket, defeating the histogram - instruments are re-created under the installed provider with the instrumentation scope version stamped (Setup-time, race-free) - helmfile.release.count declares the {release} curly-annotation unit per the metrics naming conventions Tested end-to-end via the OTLP integration test: exported bounds are the tuned set, units are asserted, and the scope carries the version. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): per-release duration metrics behind an opt-in switch New helmfile.release.duration histogram (seconds, same tuned buckets) with bounded dimensions by default (verb, result). Setting HELMFILE_OTEL_METRICS_PER_RELEASE=true adds helmfile.release and helmfile.namespace, answering "which release is slow" from dashboards: - well-suited to bounded CI runs; long-lived centralized collection needs a backend capacity/TTL story (documented in docs/otel.md) - per-release timing remains available in traces without the flag - env read per call (release operations are low-frequency, and tests toggle it) endReleaseSpan now takes the release and the operation start time; the five worker-loop call sites pass them (doWithReleaseSpan, SyncReleases, DeleteReleasesForSync, PrepareCharts, DiffReleases). Verified end-to-end with the console exporter (default dims vs per-release) and OTLP integration tests pinning both modes. Refs: #2767, #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): maintainability pass over the runner/metric plumbing - StartCommandSpan copies the tracingState struct instead of enumerating fields by hand — that pattern dropped the meter provider once already - the two value/pointer ShellRunner switches (helm marker, span attachment) are unified into one withRunnerCtx helper; the duplication caused two review rounds of value-form misses - classifyExec derives the helm classification from the span name (helmExecSpanName constant) instead of returning a third parallel bool - metrics: shared outcomeAttrs for the verb/result dimensions, and the bucket slice renamed to durationBuckets with a comment covering both histograms that use it No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): consolidate test env lists, trace bridges, and hook prep; sync the design doc Maintainability: - HermeticEnvVars is now exported from pkg/telemetry (the owner of the env surface) and used by both telemetry tests and otlptest — the two copies had already drifted once (HELMFILE_OTEL_METRICS_PER_RELEASE needed updating in both) - kubedogTraceContext and hookTraceContext were the same concept written twice; unified into traceOnlyContext(parent...) in span.go Readability: - runHook's nested kubectl rewrite extracted into prepareKubectlHook with guard-clause structure Accuracy (docs ↔ code, drifted over five review rounds): - §4.4 now describes the implemented mechanism: the release span is INJECTED into the runner's own context (preserving the kubedog safety valve) rather than the runner context being replaced, and helm classification is marker-based for wrapper binaries - §5 exec span rows list the actual attributes incl. URL/query masking - §6 strict profile documents RedactedRef, --kube-token, and the adjacent-token guarantee No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): drop the dead noop state, relocate skipUndesired, sync user-facing accuracy - tracingState.noop was dead weight in the enabled state and a copy-surface in every transition; a single package-level noopTracerProvider now backs Tracer while disabled - skipUndesired moved next to doWithReleaseSpan in span.go, its only conceptual home (span/metric suppression, not run plumbing) - accuracy: the package doc, --otel-tracing flag help, experimental-features entry, and CHANGELOG now all say tracing AND metrics and list the third instrument (helmfile.release.duration with the HELMFILE_OTEL_METRICS_PER_RELEASE opt-in) — these had drifted when the metric was added; the PR description's metric table is updated to match as well No behavior change; full -race suite green (except the pre-existing network-dependent TestStorage_resolveFile flake), lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): flatten Setup, name the prefix bound, dedupe test fake; fix instrument-count drift Readability/maintainability: - Setup drops from 56 to 39 lines: provider construction (including the shutdown-tracer-on-meter-failure recovery) moves to newProviders in exporter.go next to the constructors it composes - refredact's magic 16 becomes maxForcedFormPrefix with a comment - span_test's hand-rolled fakeRunner removed in favor of the existing mockRunner (same package) Accuracy: - "Two instruments" wording survived in docs/otel.md and the design proposal §7 after helmfile.release.duration was added; both now say three and mention the per-release opt-in No behavior change; full -race suite green (except the pre-existing network flake), lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): co-locate span machinery, drop a dead export, fix docs nits - the span plumbing helpers (markHelmExec, withRunnerCtx, markHelmRunner, spanAttachedContext) move from exec.go to span.go, next to the marker type and classifiers they serve — exec.go keeps only the funnel call sites - otlptest.SpanNames was never used outside the package; unexported - isHelmBinary's comment now states it is the FALLBACK classifier (funnel invocations are marker-classified), replacing the outdated "cosmetic distinction" framing from before the marker existed - docs/otel.md: release-scoped hooks nest under their release span (added in review round 2, never documented) No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com> |
||
|
|
34ead21107 |
feat: allow helmfile to continue on failed releases (#2616)
* feat: allow helmfile to continue on failed releases Co-authored-by: Peter Honeder <peter.honeder@unwired.at> Signed-off-by: Niklas Ott <niklas.ott@unwired.at> * fix: skip failed-prep releases, complete flag wiring, add tests and docs (#64) Review follow-ups for --allow-failed-releases (#2616): - Track per-release chart preparation failures in PrepareCharts (returned as a map keyed by release) and remove those releases from the state in Run.WithPreparedCharts when --allow-failed-releases is set, so a failed release is never executed against its original, un-prepared chart reference (which could either fail again with a duplicate error or, for charts requiring chartify, bypass patches/dependency modifications and produce an unintended result). All failures are still reported at the end via the aggregated MultiError. - Complete the release identity on error results from prepareChartForRelease so failures are attributed to the correct release. - With --allow-failed-releases, continue building dependencies of the remaining charts when 'helm dep build' fails for one release, and skip the affected releases during execution. - Wire --allow-failed-releases into 'helmfile unittest' and 'helmfile status'; remove the dead flag wiring for write-values and list (both never prepare charts, see commandsSkipChartPrep). - Simplify control flow (guard clauses, errors.As, drop dead code and redundant else branches). - Add end-to-end coverage in pkg/app/issue_2616_test.go and extend the state-level tests; document the flag in docs/cli.md and CHANGELOG.md. Signed-off-by: yxxhero <11087727+yxxhero@users.noreply.github.com> --------- Signed-off-by: Niklas Ott <niklas.ott@unwired.at> Signed-off-by: yxxhero <11087727+yxxhero@users.noreply.github.com> Co-authored-by: Peter Honeder <peter.honeder@unwired.at> Co-authored-by: yxxhero <11087727+yxxhero@users.noreply.github.com> |
||
|
|
7cf4ff9a25 |
feat: add --skip-diff-validation-on-install CLI flag (#2728)
Signed-off-by: Richter <h.richter@sap.com> |
||
|
|
7bebfab71a |
feat: add --repo-retries for helm repo and registry login commands (#2683)
* feat: add --repo-retries for retrying helm repo and registry login commands Add a configurable retry mechanism for chart repository operations to handle unstable networks (corporate proxies, slow internal registries). Closes #1894 - New --repo-retries N flag and HELMFILE_REPO_RETRIES env var (default 0 = opt-in, backward compatible) - Retry applies to helm repo add, helm repo update (incl. ACR), and helm registry login with exponential backoff (1s, 2s, 4s, ..., capped 30s) - Single retryRepoOp helper; per-attempt args/buffer are local to avoid state leaking across retries - Tests cover succeed-after-retry, exhausted-retries, disabled-by-default, and regression guards for password-buffer and args-accumulation Signed-off-by: yxxhero <aiopsclub@163.com> * fix: address PR review (overflow guard, cancellable sleep, flag-override, docs) Address Copilot review feedback on #2683: - Cap backoff shift exponent at 5 to prevent time.Duration overflow on large --repo-retries values - Make retry sleep context-aware (sleepCtx) so Ctrl+C aborts the retry loop promptly via the ShellRunner context - Log a concise exit status instead of the verbose ExitError dump, and clarify the retry-counter wording ('retry N/M') - Use -1 sentinel as the CLI default so --repo-retries=0 can explicitly disable retries even when HELMFILE_REPO_RETRIES is set - Align help text and docs: retry applies 'on failure' (not just transient errors), document the 0-disables behavior - Add tests for overflow guard, cancellable sleep, and flag-zero-disables Signed-off-by: yxxhero <aiopsclub@163.com> * fix: abort retries on canceled context, hide sentinel default, align comment Address follow-up Copilot review on #2683: - Fix tight-loop bug: sleepCtx now returns whether it completed vs was interrupted by context cancellation, and retryRepoOp aborts the retry loop on interruption so Ctrl+C no longer spins into rapid helm calls - Hide the -1 sentinel from --help by overriding the displayed default to 0 (pflag DefValue), matching the documented default while keeping the flag-override semantics - Correct HelmExecOptions.RepoRetry comment: 'on failure' not 'transient network errors', matching the actual retry behavior - Add Test_Retry_AbortsOnCanceledContext covering the no-tight-loop path Signed-off-by: yxxhero <aiopsclub@163.com> * fix: copy args per retry in RegistryLogin, make cancel test deterministic Address follow-up Copilot review on #2683: - RegistryLogin: pass a per-attempt copy of args to execStdIn so its internal append (for helm.extra) can't alias the shared slice across retries - Test_Retry_AbortsOnCanceledContext: cancel the context deterministically inside the op closure after the first attempt, replacing the flaky time.Sleep(20ms) goroutine Signed-off-by: yxxhero <aiopsclub@163.com> * fix: return error on unknown managed repo type instead of silent skip Address Copilot review on #2683: AddRepo logged an error for an unknown managed type but returned nil, silently succeeding while skipping the repo add. Now returns an error so misconfigurations fail loudly. Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com> |
||
|
|
e3f757d5ed |
feat: add --template-args flag to template/apply/sync for helm lookup() support (#2666)
* feat: add --template-args to enable helm lookup() during template/apply/sync (#1833) Add a --template-args flag to the template, apply, and sync subcommands so extra args (most notably --dry-run=server) can be passed to the helm template invocation, enabling Helm's lookup() function to resolve live cluster values. - template: --template-args reaches both chartify's pre-render helm template and the final helm template output (flagsForTemplate). - apply/sync: --template-args reaches chartify's pre-render helm template. apply/sync already inject --dry-run=server automatically for cluster operations; the flag is an explicit opt-in for the template subcommand or for passing additional flags. - When --dry-run is present in template args, kube-context/kubeconfig are also injected into chartify so lookup() can actually reach the cluster. - Resolves the long-stale PR #1833 rebased onto current main, which already contains the cluster-connectivity infrastructure (issues #2271, #2309, #2355, #2444). - Includes integration test (lookup.sh) covering both chartify and non-chartify scenarios. Signed-off-by: yxxhero <aiopsclub@163.com> * test: make lookup template nil-safe to fix integration CI The lookup() function returns an empty map when the chart is rendered without a cluster connection (notably the helm-diff phase of `helmfile apply`). The original fixture chained `index` over the lookup result, panicking with "index of untyped nil" during apply's diff rendering. Guard every index with `default dict` so the template falls back to "overwritten" when lookup is empty, while still resolving to the live value ("init") when cluster access is available (--dry-run=server via --template-args, or a real helm upgrade). Signed-off-by: yxxhero <aiopsclub@163.com> * feat: enable lookup() during apply/diff via --template-args in helm-diff Thread --template-args into the helm-diff rendering path so that `helmfile apply`/`diff --template-args="--dry-run=server"` resolves Helm's lookup() function during the diff phase too. helm-diff supports `--dry-run=server`, which explicitly "enables the cluster access ... and the lookup template function". Previously --template-args only reached chartify's pre-render (which is a no-op for plain charts due to chartify's early-return when there is no forceNamespace/patches/injections) and the final `helm template` of the `template` subcommand. As a result `helmfile apply` on a lookup chart rendered client-side during the diff phase. Changes: - pkg/state: add TemplateArgs to DiffOpts; append it in appendExtraDiffFlags (reaches every helm-diff invocation: apply, standalone diff, interactive sync), mirroring the existing flagsForTemplate handling. - pkg/config + cmd: add --template-args to the diff/doctor commands and to DiffConfigProvider, so lookup works for `helmfile diff` as well. - pkg/app: populate DiffOpts.TemplateArgs from apply/diff/sync-interactive. - docs/cli.md: correct the previous overpromising wording and document diff support plus the nil-safe lookup guidance. - tests: unit-test the TemplateArgs handling in appendExtraDiffFlags and flagsForTemplate; integration lookup.sh now exercises apply with --template-args="--dry-run=server". Signed-off-by: yxxhero <aiopsclub@163.com> * refactor: de-duplicate chartify template-args logic, add helmDefaults.templateArgs Address review feedback on #2666: 1. Eliminate stale duplicated test helpers (issue_2444_test.go, issue_2355_test.go). Both files intentionally copied the processChartification flag-building logic with explicit SYNC WARNING comments, then drifted out of sync when #2666 refactored the production code (needsKubeConnection gate, user-args merge). Extract the real logic into pure, unit-tested helpers (buildChartifyTemplateArgs, commandRequiresCluster) and delete the copies. 2. Add unit coverage for the new chartify merge path: template + --template-args=--dry-run=server now triggers kubeconfig/kube-context injection (TestTemplateArgsDryRunTriggersKubeInjection, TestTemplateArgsMergedBeforeInjection) — previously only covered by the cluster-dependent integration test. 3. Add a negative integration case (lookup.sh assert_template_fallback) verifying lookup() falls back to the default value WITHOUT --template-args, guarding against a regression that silently always connects to the cluster. 4. Add helmDefaults.templateArgs for parity with diffArgs/syncArgs, so users can enable lookup() support permanently instead of passing the flag on every invocation. CLI --template-args overrides (does not merge with) the default. Resolved via effectiveTemplateArgs, wired into the chartify, flagsForTemplate, and appendExtraDiffFlags paths. 5. Minor: capitalize --template-args help text to match surrounding flags; document helmDefaults.templateArgs precedence in docs/cli.md. Signed-off-by: yxxhero <aiopsclub@163.com> * test: cover helmDefaults->chartify composition; fix helm helm-diff typo Address remaining review nits on #2666: - Add TestHelmDefaultsTemplateArgsReachesChartify, a belt-and-suspenders test for the processChartification composition (effectiveTemplateArgs -> buildChartifyTemplateArgs), closing the last unit-level coverage gap for helmDefaults.templateArgs reaching the chartify path. - Fix pre-existing typo in cmd/bind_diff_flags.go: 'pass args to helm helm-diff' -> 'Pass args to helm-diff' (doubled 'helm', lowercase). Signed-off-by: yxxhero <aiopsclub@163.com> * fix: correct 'helm helm-diff' typo in apply --diff-args help text Sibling of the bind_diff_flags.go fix; the same doubled-'helm' typo and lowercase help existed in cmd/apply.go's --diff-args registration, leaving the apply and diff/doctor help strings inconsistent. Signed-off-by: yxxhero <aiopsclub@163.com> * docs: add helmDefaults.templateArgs to configuration reference The complete helmfile.yaml schema in docs/configuration.md documents diffArgs and syncArgs under helmDefaults but was missing the new templateArgs field added in #2666. Add it beside syncArgs for discoverability, noting the --template-args CLI override. Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com> |
||
|
|
9b943adc9e |
feat: add helmfile doctor command for AI-assisted diff analysis (#2660)
* feat: add `helmfile doctor` command for AI-assisted diff analysis `helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to summarize the changes and flag risks (data loss, security exposure, breaking changes, downtime, performance, best-practice issues). Key design decisions: - When no LLM is configured, doctor is equivalent to `helmfile diff` with one exception: --show-secrets is always forced off (secrets never reach stdout, even without an LLM). - Secrets are ALWAYS redacted via two layers: (1) ShowSecrets() forced to false so helm-diff emits <REDACTED> placeholders; (2) a defense-in-depth text redactor strips residual secret-looking content (Secret YAML blocks, sensitive key/value lines, base64 blobs, JWT tokens) before LLM transmission. - LLM configuration precedence: env (HELMFILE_LLM_*) < helmfile.yaml (llm:) < CLI flags (--llm-*). - Supports any OpenAI-compatible backend (OpenAI, Azure, One-API, LiteLLM, Ollama, etc.) with automatic response_format fallback for backends that don't support JSON mode. - Prompt injection defense: release names and environment values are JSON-encoded before insertion into the LLM prompt. - Exit codes: 0 (success/low-risk), 2 (high-risk gate, bypass with --force), 1 (other errors). Helm-diff's 'detected changes' exit-2 is swallowed. New packages: - pkg/agent/llm: OpenAI-compatible client with JSON response parsing, mock client for testing, prompt builder with injection defense. - pkg/agent/doctor: secret redactor (state machine + regex), report renderer (markdown + JSON), config resolver (env < yaml < flag merge). Testing: 70+ unit tests covering redaction patterns, prompt injection, response_format fallback, JSON parsing, yaml roundtrip, concurrency safety, panic recovery, and error propagation. go test -race passes. Documentation: full doctor section in docs/cli.md, llm: block reference in docs/configuration.md, updated skills/helmfile for AI agents. Signed-off-by: yxxhero <aiopsclub@163.com> * docs: fix doctor equivalence wording per PR review Per review feedback (PR #2660): the docs claimed doctor is 'equivalent to helmfile diff — same flags, same output, same exit codes' in the unconfigured path, but this over-promises because: 1. doctor --output is the report format (not helm-diff's output format) 2. helm-diff's --output is exposed as --diff-output in doctor 3. --show-secrets is silently ignored Updated all three locations (cli.md, cmd/doctor.go Long + godoc, pkg/app/doctor.go godoc) to say 'falls back to helmfile diff with --show-secrets forced off' and explicitly note the --output / --diff-output flag difference. Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com> |
||
|
|
33eadc993e |
feat: support HELMFILE_* env vars for more global flags (#2606)
* feat: support more HELMFILE_* env vars as flag fallbacks
Adds env-var fallbacks for global flags, mirroring the existing
HELMFILE_ENVIRONMENT / HELMFILE_KUBE_CONTEXT pattern:
* --helm-binary -> HELMFILE_HELM_BINARY
* --kustomize-binary -> HELMFILE_KUSTOMIZE_BINARY
* --log-level -> HELMFILE_LOG_LEVEL
* --debug -> HELMFILE_DEBUG (expecting "true" lower case)
* --quiet -> HELMFILE_QUIET (expecting "true" lower case)
* --no-color -> HELMFILE_NO_COLOR (expecting "true" lower case),
additionally honors NO_COLOR per no-color.org
(any non-empty value disables color)
Flag values still take precedence; env vars are consulted only when the
flag is unset. The string-flag default values ("helm", "kustomize",
"info") move into the accessor methods so the env-var fallback can
actually trigger when no flag is passed.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* docs: mention new HELMFILE_* env vars in cli.md and templating.md
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* fix: make Color/NoColor/env interaction consistent
Two issues with the env-aware NoColor() introduced together with
HELMFILE_NO_COLOR / NO_COLOR support:
1. Color() consulted the raw GlobalOptions.NoColor field instead of
NoColor(), so in a TTY with only the env set, Color() fell through
to terminal autodetect and ValidateConfig() spuriously errored with
"--color and --no-color cannot be specified at the same time".
2. NoColor() returned true via env even when --color was explicitly
passed, so `helmfile --color` with NO_COLOR (or HELMFILE_NO_COLOR=true)
in the environment hit the same ValidateConfig() error. A flag should
always win over an env var.
Fix both by routing Color() through NoColor() and giving NoColor() an
explicit --color short-circuit. Regression tests added for both paths.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
---------
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
|
||
|
|
31ac918512 |
feat: support HELMFILE_NAMESPACE env var for default namespace (#2592)
* feat: support HELMFILE_NAMESPACE env var for default namespace Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --namespace CLI flag takes precedence, falling back to HELMFILE_NAMESPACE when unset. Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> * docs: mention HELMFILE_NAMESPACE in cli.md and templating.md Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> --------- Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> |
||
|
|
c15cbb096a |
feat: support HELMFILE_KUBE_CONTEXT env var for default kube context (#2593)
* feat: support HELMFILE_KUBE_CONTEXT env var for default kube context Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --kube-context CLI flag takes precedence, falling back to HELMFILE_KUBE_CONTEXT when unset. Refs #1213. Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> * docs: mention HELMFILE_KUBE_CONTEXT in cli.md and templating.md Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> --------- Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de> |
||
|
|
e703b15075 |
docs: restructure documentation and improve newcomer experience (#2573)
* feat: add --write-output flag to helmfile fetch for air-gapped environments Add --write-output flag to helmfile fetch that outputs a modified helmfile.yaml with chart references updated to point to downloaded local chart paths. Combined with --output-dir, this enables preparing all charts for deployment in air-gapped environments. Usage: helmfile fetch --output-dir ./charts --write-output > helmfile-airgapped.yaml Fixes #2571 Signed-off-by: yxxhero <yxxhero@users.noreply.github.com> Signed-off-by: yxxhero <aiopsclub@163.com> * docs: restructure documentation and improve newcomer experience Split the monolithic index.md (1990 lines) into focused topic pages, update mkdocs.yml navigation, and add missing documentation for undocumented code features. Structure changes: - Extract configuration.md (helmfile.yaml reference) - Extract cli.md (CLI commands and flags) - Extract templating.md (template syntax and env vars) - Extract environments.md (environment configuration) - Extract releases.md (DAG, needs, selectors) - Extract hooks.md (lifecycle hooks) - Extract integrations.md (ArgoCD, Azure ACR, OCI) - Slim index.md to ~270 line landing page with step-by-step tutorial Newcomer improvements: - Add 5-step Getting Started tutorial with explanations - Reorganize nav: Getting Started now shows core learning path (Writing Helmfile → Values → Environments → Releases) - Add Quick Reference table to configuration.md - Simplify writing-helmfile.md title Code-vs-docs gap fixes: - Document 23 undocumented release fields (valuesTemplate, setTemplate, forceNamespace, adopt, trackMode, etc.) - Document 6 undocumented helmDefaults fields (enableDNS, forceConflicts, skipRefresh, takeOwnership, etc.) - Document print-env command and missing CLI flags - Document kubectlApply hook field - Document environment defaults field and merge order - Document kubedogQPS/kubedogBurst advanced settings - Document template partials (_*.tpl) auto-loading Cleanup: - Fix Docker image version from v0.156.0 to v1.1.0 - Fix heading nesting in advanced-features.md - Update experimental-features.md with current features - Fix broken cross-references and anchor links Signed-off-by: yxxhero <aiopsclub@163.com> * Revert changes to pkg/app from docs/restructure-and-improve branch Signed-off-by: yxxhero <aiopsclub@163.com> * docs: add create subcommand to README and CLI reference Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <yxxhero@users.noreply.github.com> Signed-off-by: yxxhero <aiopsclub@163.com> |