* feat(telemetry): add opt-in OpenTelemetry tracing (PR 1: lifecycle + root span) Implements the first increment of docs/proposals/otel-tracing.md (#2767): - pkg/telemetry: SDK setup from standard OTEL_* env vars (autoexport for exporter selection, env-driven sampler/propagators, OTEL_SDK_DISABLED), command-span lifecycle, no-op-by-default accessors - --otel-tracing flag / HELMFILE_OTEL_TRACING env switch - root span "helmfile <command>" with file/environment/selectors/exit_code attributes; TRACEPARENT-based remote-parent extraction for CI correlation - shutdown flush on both normal-exit and signal paths (nil-safe, 5s bound) - app.New derives its context from telemetry.CommandContext() (Background-identical when tracing is disabled) - docs: otel.md user guide, experimental-features entry, design proposal - tests: hermetic unit tests, app context-contract pinning, flag registration Telemetry problems never fail a run: exporter misconfiguration and export errors degrade to disabled with a warning. When disabled, behavior and performance are identical to before (no-op tracer, no goroutines, no network). Refs: #2767, #2758 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): trace every external process + trace-context bridges (PR 2) Implements the second increment of docs/proposals/otel-tracing.md (#2767): - pkg/helmexec/span.go: one span per external process started by helmfile (helm invocations, hooks, plugin execs) at the ShellRunner choke point — helm.exec (with helm.subcommand) vs os.exec, with redacted exec.args, exec.exit_code, and error status on failure - pkg/helmexec/redact.go: shared argument redaction with two profiles; legacy is byte-identical to the historical exit-error behavior (existing goldens unchanged), strict (spans) additionally covers --set=k=v and credential flags; exit_error.go now uses the shared helper - orphan-trace bridges with bit-identical cancellation semantics (context.WithoutCancel of the command context): both kubedog call sites (state.go) and hook execution (event.Bus gains an optional Ctx consumed by its default runner; state.go sets it, nil falls back to TODO as before) - OTLP end-to-end test (in-process httptest receiver, no external collector): span export, error status/exit code, redaction, and parent-linkage to the command span - docs/otel.md updated to the now-traced surface Verified end-to-end with the console exporter: helmfile template on a local chart yields the command span plus helm.exec spans for helm version/dependency/template, all nested under it. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): state-loading and hook spans (PR 3a) Implements the third increment of docs/proposals/otel-tracing.md (#2767): - helmfile.discover_states around findDesiredStateFiles and helmfile.load around loadDesiredStateFromYamlWithBaseDir; both cover all callers (incl. nested helmfiles) with no signature changes - helmfile.render / helmfile.parse children per document part, parented through a traceCtx field on the unexported desiredStateLoader struct (set once at its single construction site) - helmfile.hook span per hook execution: Trigger's per-hook body extracted into runHook (readability win on its own), the hook's subprocess span nests under it via a per-hook ctx-swapped ShellRunner clone (cancellation unchanged — Bus.Ctx never carries cancellation by contract) - pkg/telemetry/otlptest: shared in-process OTLP/HTTP receiver harness, now used by helmexec, event, and app span tests - golden span-tree test at the app layer (root -> discover -> load -> render/parse, via the exectest fake helm) and a hook-span nesting test - nil-ctx guard for App literals built directly by tests (App.spanParentCtx) Verified end-to-end with the console exporter: a template run over a gotmpl state file with a prepare hook yields the full tree with the hook's os.exec nested under helmfile.hook. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): per-release spans nested under the load span (PR 3b) Implements the per-release increment of docs/proposals/otel-tracing.md (#2767) — spans nest command -> load -> release -> helm exec: - helmexec.HelmContext gains an optional Ctx carrying the per-release span context; the execer's new execWithContext funnel consumes it via a per-call runner clone (runnerWithCtx) so the shared, cached execer is never mutated across concurrent workers. The seven Interface methods that take a HelmContext (Sync/Diff/ReleaseStatus/List/DecryptSecret/ Delete/Test) route through it; nil Ctx behaves exactly as before. exec() lost its always-nil override parameter on the way (unparam). - pkg/state/span.go: SetTraceContext + startReleaseSpan/endReleaseSpan helpers (release/namespace/chart/labels attributes, sorted for stable output); a typed-nil guard (releaseErrAsError) avoids the classic nil-pointer-in-interface trap on *ReleaseError. - release spans in the worker loops: SyncReleases, DiffReleases, DeleteReleasesForSync, PrepareCharts, and iterateOnReleases (status/ delete/test via a new verb parameter); their HelmContext is stamped with the release span context where one is built. - pkg/app sets st.SetTraceContext(loadCtx) right after loading a state file, rooting all per-release spans under helmfile.load. - bridged one more detached tracking call found on the way (trackReleaseIfEnabled's context.Background in the sync worker). - golden test: release span present, nested under load, correct attributes; unit test for the runnerWithCtx clone semantics. Verified with the console exporter: helmfile template yields release.prepare(demo) under load with full attributes. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): nest status/delete/test execs under their release spans Completes the per-release exec nesting for the iterateOnReleases-based loops (docs/proposals/otel-tracing.md §4.4 phase 2): the do closures now receive the release span context and stamp it into their HelmContext, so helm status/delete/test subprocess spans nest under helmfile.release.<verb> like sync/diff already did. - scatterGatherReleases/iterateOnReleases/doWithReleaseSpan: do gains a context parameter (the release span context) - ReleaseStatuses/DeleteReleases/TestReleases closures stamp HelmContext.Ctx from it - integration test with a real execer (version-probe shim binary): the release's status subprocess nests under helmfile.release.status, same trace, with helm.subcommand=status This also makes the otel.md claim ("upgrade, diff, delete, status, test nested under the release span") fully accurate. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): OTel metrics — helm exec duration and release results (PR 4) Implements the metrics increment of docs/proposals/otel-tracing.md (#2767) on the same provider, switch, and resource as traces: - pkg/telemetry/metrics.go: helmfile.helm.exec.duration histogram (subcommand, success) and helmfile.release.count counter (verb, result). Instruments come from the otel global meter, so recording at call sites is branch-free no-op when telemetry is disabled. - Setup builds the resource once and installs both providers; reader selection delegates to autoexport (OTEL_METRICS_EXPORTER: otlp | console | prometheus | none), the OTLP reader's interval honors OTEL_METRIC_EXPORT_INTERVAL (read by the SDK). Shutdown flushes both providers (errors.Join). StartCommandSpan now carries the meter provider across state transitions (fixes a nil-shutdown panic). - helmexec: finishExecSpan records exec duration for helm binaries; state: endReleaseSpan counts release outcomes for sync/diff/delete/ status/test/prepare (diff counted as success when no hard error). - otlptest: recorder routes by OTLP path (/v1/traces vs /v1/metrics) and decodes metrics; new FindMetric helper. - tests: metrics recorded as no-op when disabled, provider enabled with the none exporter, degradation on an invalid metrics exporter, and an integration assertion (status exec duration datapoint + one successful release.count) in the shim-based state test. Verified with the console exporter: helmfile template emits helmfile.helm.exec.duration per subcommand (version/dependency/ template) and helmfile.release.count{verb=prepare,result=success}=1. Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * docs: complete OTel documentation coverage - docs/cli.md: --otel-tracing in the CLI reference help block (verbatim from the cobra output) - CHANGELOG.md: [Unreleased] Added entry for tracing + metrics - docs/index.md: Observability highlight linking docs/otel.md - docs/proposals/otel-tracing.md: add OTEL_METRICS_EXPORTER / OTEL_METRIC_EXPORT_INTERVAL rows to the env-var table and note the periodic reader + bounded metric cardinality in §7 Refs: #2767 Signed-off-by: yxxhero <aiopsclub@163.com> * fix: drop unused id parameter from parsePart (unparam) The id parameter was never used inside the span wrapper; the caller's id variable is still used for the render calls and error messages. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review — redaction gaps, kubedog valve, phantom metrics Addresses all Copilot review comments on #2769: Security (span payloads): - exec.args: positional arguments are additionally passed through helmexec.RedactedURL, so credentials embedded in chart/repository URLs (AddRepo, RegistryLogin, OCI refs) are masked exactly like log output - release spans sanitize helmfile.chart the same way - error statuses no longer embed raw errors (which contain rendered commands, arguments, and subprocess output): the command span, release spans, hook spans, and exec spans now use generic descriptions; the concrete exit code remains an attribute, and RecordError on the root span is dropped Correctness: - kubedog safety valve restored: execWithContext now attaches the per-release span into the runner's own context instead of replacing it, so trackHandle.Cancel() can interrupt a wedged helm again and app cancellation semantics stay exactly as before the PR - diff release spans/metrics: real failures are recorded (exit code 2 "changes detected" still counts as success); previously every diff was exported as successful - skipped releases no longer emit phantom spans and inflate helmfile.release.count: iterateOnReleases callers pass a skip predicate (skipUndesired for status/test; delete deletes undesired releases and passes nil) - Setup shuts down the already-constructed tracer provider (bounded) when the metrics provider fails, instead of abandoning its batch goroutine Tests: URL redaction cases (masked/untouched), spanAttachedContext preserves the runner cancellation chain while attaching the caller's span, skipUndesired, and a failing-hook span asserting the generic message. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix: lint — restore nolint placement and avoid nil context literal - the skipUndesired insertion had displaced the // nolint: unparam directive off iterateOnReleases (helm param is intentionally unused there); also fixes a skipDesired/skipUndesired comment typo - use a typed nil in TestSpanAttachedContext (staticcheck SA1012) Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review round 2 — remote-ref redaction, wrapper helm binaries, hook release attribution Addresses all 6 new review comments on #2769: Security (remote references): - new helmexec.RedactedRef sanitizes go-getter style references for telemetry: forced-form prefixes (git::, s3::) preserved, whole URL userinfo masked (usernames carry tokens too), credential-bearing query parameters masked using pkg/remote's heuristic (token/password/secret/ key/signature). Applied to helmfile.file (command span), helmfile.path (discover_states), helmfile.chart (release spans), and exec.args — log-time RedactedURL is untouched so log output is unchanged Correctness: - wrapper helm binaries (--helm-binary custom names) are now classified as helm operations by an explicit context marker stamped in the execer funnel, instead of the executable-basename heuristic; the same classification gates helmfile.helm.exec.duration, so the metric no longer misses wrapper invocations (classifyExec) - release-scoped hooks (presync/postsync/preuninstall/postuninstall/ cleanup in the sync/delete/diff workers) now attach their helmfile.hook spans to the active helmfile.release.* span via a variadic parent on the trigger functions; global hooks keep the command context and all 29 existing call sites compile unchanged; hook cancellation stays detached (WithoutCancel) as before - signal-terminated runs (Shutdown with exitCode 130/143 and nil error) now mark the command span with error status, consistent with their nonzero exit code Tests: RedactedRef table (forced forms, userinfo, s3/token query params, untouched cases), classifyExec marker case, hookTraceContext parent attribution + non-cancellability + fallback. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): address review round 3 — redaction corner cases, value runners Addresses 5 of the 6 new review comments on #2769 (the sixth — an unused strings import in exit_error.go — is a false positive: Indent still uses strings.Split/Builder and the package compiles): - RedactArgs read the previous token from the progressively redacted output, so {--set, --set-string, secret} leaked the secret (the masked value hid the following flag). Read the previous token from the original input, restoring the legacy contract for adjacent secret flags - RedactedRef fails closed for malformed references: URL-like refs with invalid percent escapes export a fully redacted value, and an unparseable query is dropped entirely instead of exported verbatim - ShellRunner has value receivers, so a ShellRunner VALUE satisfies the Runner API; the helm marker stamping and the per-release span attachment now handle both value and pointer forms (matching WithContext), so value-runner callers keep release nesting and the helm.exec classification/metric Regression tests: adjacent secret flags (legacy + strict), malformed URL-like ref, malformed query, value-runner marker + span attachment. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): stamp the helm marker on the stdin funnel too execStdIn (registry login, repo add) called the runner directly, so wrapper --helm-binary names were misclassified as os.exec and omitted from helmfile.helm.exec.duration on that path. The marking now goes through a shared markHelmRunner helper (value and pointer ShellRunner forms) used by both execution funnels. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): redact helm's --kube-token in strict profile Helm's global --kube-token carries a bearer token; both the two-argument and inline forms are now masked in span exec.args (legacy exit-error output is untouched, matching its historical behavior). Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * fix(telemetry): OTel metrics best-practice alignment - helmfile.helm.exec.duration now declares explicit bucket boundaries tuned for seconds-scale helm invocations (5ms…600s); the SDK defaults are millisecond-oriented and lumped every sub-5s invocation — the common case — into the first bucket, defeating the histogram - instruments are re-created under the installed provider with the instrumentation scope version stamped (Setup-time, race-free) - helmfile.release.count declares the {release} curly-annotation unit per the metrics naming conventions Tested end-to-end via the OTLP integration test: exported bounds are the tuned set, units are asserted, and the scope carries the version. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * feat(telemetry): per-release duration metrics behind an opt-in switch New helmfile.release.duration histogram (seconds, same tuned buckets) with bounded dimensions by default (verb, result). Setting HELMFILE_OTEL_METRICS_PER_RELEASE=true adds helmfile.release and helmfile.namespace, answering "which release is slow" from dashboards: - well-suited to bounded CI runs; long-lived centralized collection needs a backend capacity/TTL story (documented in docs/otel.md) - per-release timing remains available in traces without the flag - env read per call (release operations are low-frequency, and tests toggle it) endReleaseSpan now takes the release and the operation start time; the five worker-loop call sites pass them (doWithReleaseSpan, SyncReleases, DeleteReleasesForSync, PrepareCharts, DiffReleases). Verified end-to-end with the console exporter (default dims vs per-release) and OTLP integration tests pinning both modes. Refs: #2767, #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): maintainability pass over the runner/metric plumbing - StartCommandSpan copies the tracingState struct instead of enumerating fields by hand — that pattern dropped the meter provider once already - the two value/pointer ShellRunner switches (helm marker, span attachment) are unified into one withRunnerCtx helper; the duplication caused two review rounds of value-form misses - classifyExec derives the helm classification from the span name (helmExecSpanName constant) instead of returning a third parallel bool - metrics: shared outcomeAttrs for the verb/result dimensions, and the bucket slice renamed to durationBuckets with a comment covering both histograms that use it No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): consolidate test env lists, trace bridges, and hook prep; sync the design doc Maintainability: - HermeticEnvVars is now exported from pkg/telemetry (the owner of the env surface) and used by both telemetry tests and otlptest — the two copies had already drifted once (HELMFILE_OTEL_METRICS_PER_RELEASE needed updating in both) - kubedogTraceContext and hookTraceContext were the same concept written twice; unified into traceOnlyContext(parent...) in span.go Readability: - runHook's nested kubectl rewrite extracted into prepareKubectlHook with guard-clause structure Accuracy (docs ↔ code, drifted over five review rounds): - §4.4 now describes the implemented mechanism: the release span is INJECTED into the runner's own context (preserving the kubedog safety valve) rather than the runner context being replaced, and helm classification is marker-based for wrapper binaries - §5 exec span rows list the actual attributes incl. URL/query masking - §6 strict profile documents RedactedRef, --kube-token, and the adjacent-token guarantee No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): drop the dead noop state, relocate skipUndesired, sync user-facing accuracy - tracingState.noop was dead weight in the enabled state and a copy-surface in every transition; a single package-level noopTracerProvider now backs Tracer while disabled - skipUndesired moved next to doWithReleaseSpan in span.go, its only conceptual home (span/metric suppression, not run plumbing) - accuracy: the package doc, --otel-tracing flag help, experimental-features entry, and CHANGELOG now all say tracing AND metrics and list the third instrument (helmfile.release.duration with the HELMFILE_OTEL_METRICS_PER_RELEASE opt-in) — these had drifted when the metric was added; the PR description's metric table is updated to match as well No behavior change; full -race suite green (except the pre-existing network-dependent TestStorage_resolveFile flake), lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): flatten Setup, name the prefix bound, dedupe test fake; fix instrument-count drift Readability/maintainability: - Setup drops from 56 to 39 lines: provider construction (including the shutdown-tracer-on-meter-failure recovery) moves to newProviders in exporter.go next to the constructors it composes - refredact's magic 16 becomes maxForcedFormPrefix with a comment - span_test's hand-rolled fakeRunner removed in favor of the existing mockRunner (same package) Accuracy: - "Two instruments" wording survived in docs/otel.md and the design proposal §7 after helmfile.release.duration was added; both now say three and mention the per-release opt-in No behavior change; full -race suite green (except the pre-existing network flake), lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> * refactor(telemetry): co-locate span machinery, drop a dead export, fix docs nits - the span plumbing helpers (markHelmExec, withRunnerCtx, markHelmRunner, spanAttachedContext) move from exec.go to span.go, next to the marker type and classifiers they serve — exec.go keeps only the funnel call sites - otlptest.SpanNames was never used outside the package; unexported - isHelmBinary's comment now states it is the FALLBACK classifier (funnel invocations are marker-classified), replacing the outdated "cosmetic distinction" framing from before the marker existed - docs/otel.md: release-scoped hooks nest under their release span (added in review round 2, never documented) No behavior change; full -race suite green, lint clean. Refs: #2769 Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com>
38 KiB
Proposal: OpenTelemetry Tracing Support
- Issue: #2767 — feat-request: OpenTelemetry (tracing) support
- Status: Draft
- Origin discussion: #2758
1. Summary
Add opt-in OpenTelemetry (OTel) distributed tracing to helmfile. When enabled, a helmfile run produces a trace whose spans cover the full execution timeline — state-file discovery, template rendering, per-release operations, hooks, and every helm subprocess helmfile itself starts — exported via OTLP to any OTel-compatible backend (Jaeger, Tempo, Zipkin, Honeycomb, Datadog, Grafana Cloud, cloud-vendor collectors, ...).
This directly serves the motivating use case from CI/CD: see where time is spent, what runs most often (label selectors), and what to target for improvement — the same capability Terragrunt ships today (Terragrunt OTel docs).
When disabled (the default), behavior and performance are identical to today: a no-op tracer provider, no goroutines, no network, no exported spans (explicit guarantees in §7).
Example resulting trace
helmfile sync file=helmfile.yaml env=production (root, ~2m)
├─ helmfile.discover_states (~5ms)
├─ helmfile.load file=helmfile.yaml (~800ms)
│ ├─ helmfile.render pass=values (~400ms)
│ └─ helmfile.parse (~50ms)
├─ helmfile.repos.update (~4s)
│ └─ helm.exec subcommand="repo update" (~4s)
├─ helmfile.release.prepare release=gateway chart=gateway (~9s)
│ └─ helm.exec subcommand="dependency build" (~8s)
├─ helmfile.release.sync release=envoy ns=ingress (~21s)
│ ├─ helmfile.hook event=presync (~2s)
│ │ └─ os.exec command="./migrate.sh" (~2s)
│ ├─ helm.exec subcommand="upgrade --install envoy ..." (~18s)
│ └─ helmfile.hook event=postsync (~1s)
├─ helmfile.release.sync release=api ns=apps (~35s)
└─ helmfile.wait release=api (kubedog) (~30s)
2. Goals and non-goals
Goals
- Opt-in, zero-cost-when-off: tracing disabled by default; no measurable overhead or behavior change when not requested.
- Standard OTLP export configured through the OTel environment-variable specification — no helmfile-specific duplicate knobs for endpoints/protocols/headers/sampling.
- Meaningful hierarchy: spans nest (command → state file → release → hook/subprocess) so the critical path of a sync/apply is visible at a glance — including the kubedog and hook paths, which today run on detached contexts (§4.3).
- Safe by default: span attributes never contain secret-bearing values
(
--setvalues, registry credentials, secret refs). - CI-friendly: W3C
tracecontextpropagation so a CI system that injectsTRACEPARENTgets spans correlated into its own trace.
Non-goals (for the initial implementation)
- OTel metrics and log export (the SDK setup leaves room; see §12 for follow-ups).
- Tracing inside the helm binary or cluster-side (helm/chart hooks are separate processes; context injection into them is an open question, §14).
- Fixing the pre-existing cancellation gaps discovered during design (kubedog path rooted at
context.Background(), hooks atcontext.TODO()— §4.3). This proposal bridges them for trace context only, with byte-identical cancellation semantics; semantic fixes are reported separately. - Any helmfile-config-file surface (
helmfile.yaml) for telemetry — CLI flags + env vars only, matching how observability tooling is usually injected by the platform, not the state author.
3. User-facing design
3.1 Activation
| Mechanism | Default | Description |
|---|---|---|
--otel-tracing (persistent flag) |
false |
Enable tracing for this run |
HELMFILE_OTEL_TRACING=true (env) |
false |
Same, via environment (CI-friendly) |
The feature ships as experimental initially and is listed in
docs/experimental-features.md (see §11 Rollout). We do not require
HELMFILE_EXPERIMENTAL=otel-tracing: tracing is purely additive, invisible when off, and
gating it twice would complicate CI adoption. The experimental label sets stability
expectations only. (Note: this differs from the current entries in that document, which are
gated by HELMFILE_EXPERIMENTAL; the entry will say so explicitly.)
3.2 Exporter & SDK configuration — standard OTel env vars only
When --otel-tracing is set, helmfile initializes the OTel SDK honoring the standard
environment variables. Exporter selection is delegated to
go.opentelemetry.io/contrib/exporters/autoexport (already in the module graph, §8) rather
than hand-rolled, so helmfile maintains no exporter-construction code of its own:
| Variable | Default in helmfile | Read by | Notes |
|---|---|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
http://localhost:4318 |
otlp client (SDK) | default follows the protocol: 4318 with http/protobuf (below), 4317 with grpc |
OTEL_EXPORTER_OTLP_PROTOCOL / ..._TRACES_PROTOCOL |
http/protobuf |
autoexport | grpc, http/protobuf (traces-specific var wins; http/json is not supported by autoexport traces in v0.67.0) |
OTEL_EXPORTER_OTLP_HEADERS / ..._TRACES_HEADERS |
— | otlp client (SDK) | e.g. collector auth tokens |
OTEL_EXPORTER_OTLP_TIMEOUT / ..._TRACES_TIMEOUT |
10s |
otlp client (SDK) | per-export timeout |
OTEL_EXPORTER_OTLP_INSECURE |
per endpoint scheme | otlp client (SDK) | plaintext export for local collectors |
OTEL_TRACES_EXPORTER |
otlp |
autoexport | otlp | console | none (verified against autoexport v0.67.0; more values possible via RegisterSpanExporter) |
OTEL_TRACES_SAMPLER (+..._ARG) |
parentbased_always_on |
our wrapper | standard samplers (the Go SDK core does not parse this env itself) |
OTEL_SERVICE_NAME |
helmfile |
SDK resource | service identity (resource WithFromEnv; verified in sdk v1.44.0 resource/env.go) |
OTEL_RESOURCE_ATTRIBUTES |
— | SDK resource | e.g. deployment.environment=ci,cicd.pipeline=release (same verified reader) |
OTEL_PROPAGATORS |
tracecontext,baggage |
our wrapper | extract parent from CI-injected TRACEPARENT (SDK core does not parse this env either) |
OTEL_METRICS_EXPORTER |
otlp |
autoexport (metrics) | otlp | console | prometheus | none — delivered in PR 4 |
OTEL_METRIC_EXPORT_INTERVAL |
60000 (ms) |
SDK metric reader | periodic export interval; final flush happens at exit |
OTEL_SDK_DISABLED |
false |
our wrapper | standard kill switch (the Go SDK does not read this one itself) |
console (JSON spans on stdout, via stdouttrace) exists for local debugging without a
collector — note stdout, so it does not interleave with helmfile's stderr logs;
none produces no export (used by tests and for pure-propagation setups).
HTTP proxies (HTTPS_PROXY/HTTP_PROXY) are honored automatically by the Go HTTP/gRPC
stacks — no helmfile-specific proxy configuration.
Minimal CI example:
export OTEL_EXPORTER_OTLP_ENDPOINT="https://otel.example.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer ${OTEL_TOKEN}"
export OTEL_SERVICE_NAME="helmfile-ci"
export OTEL_RESOURCE_ATTRIBUTES="cicd.pipeline=deploy,cicd.run_id=4821"
helmfile --otel-tracing -e production -l tier=backend apply
4. Architecture
4.1 New package pkg/telemetry
Single, dependency-light façade over the OTel SDK. No other package imports OTel SDK exporters directly.
pkg/telemetry/
├── telemetry.go // Setup/Shutdown lifecycle, CommandContext(), Tracer(name)
├── exporter.go // thin autoexport wiring + propagators (no exporter construction)
└── telemetry_test.go
Public surface sketch:
package telemetry
type Options struct {
Enabled bool
Version string // helmfile version, recorded as service.version
Logger *zap.SugaredLogger // for one-line diagnostics, never span data
}
// Setup initializes the global provider. Idempotent; a no-op when
// opts.Enabled is false. Exporter/sampler misconfiguration and
// OTEL_SDK_DISABLED=true degrade to disabled with a warning — telemetry
// problems never fail a helmfile run.
func Setup(ctx context.Context, opts Options)
// StartCommandSpan starts the root span for one command invocation, joins a
// remote parent from TRACEPARENT/TRACESTATE/BAGGAGE env vars, and makes its
// context the one returned by CommandContext. No-op when disabled.
func StartCommandSpan(command string, attrs ...attribute.KeyValue)
// CommandContext returns the root command span's context, or
// context.Background() when telemetry is disabled. Single source of truth
// for deriving App.ctx.
func CommandContext() context.Context
// Tracer returns a tracer for the given instrumentation scope. Never nil:
// returns the OTel no-op tracer when telemetry is disabled, so callers need
// no `if enabled` branches. Sanctioned scopes: ScopeHelmfile, ScopeHelm.
func Tracer(name string) trace.Tracer
// Shutdown ends the command span (recording runErr and exitCode), flushes
// buffered spans bounded by ctx, and reverts to disabled. Idempotent and
// nil-safe (safe before Setup, twice, or on a signal that raced Setup).
// ShutdownTimeout is the recommended flush bound.
func Shutdown(ctx context.Context, runErr error, exitCode int) error
Design constraints for maintainability:
- No domain knowledge in this package. Span-attribute redaction rules live with the
existing redaction code in
pkg/helmexec(§6); telemetry only consumes the result. A telemetry package that knows about--setsemantics would be a layering violation. - Sanctioned instrumentation scopes are only
"helmfile"(app/state layer) and"helm"(helmexec layer), documented intelemetry.go.
4.2 Lifecycle
main.go cmd/root.go
────── ───────────
rootCmd.Execute() ─────────────► PersistentPreRunE:
logger setup (existing)
telemetry.Setup(ctx, opts) ← provider + exporters
start root span "helmfile <sub>"
RunE / subcommand execution … (child spans attach via context)
errChan <- Execute()
┌────────────────────────────────┐
shutdown: │ end root span (status=error on │
rootSpan.End() │ failure, exit code attr) │
provider.Shutdown(ctx 5s) ◄─────┘ │
PersistentPreRunE(cmd/root.go) is the natural init point: it already centralizes logger construction fromGlobalOptions, and receives the*cobra.Command, soc.Name()gives the root span name (helmfile sync, ...).App.ctxderivation requires no signature changes.app.New(conf)is called from every subcommand (cmd/*.go, 22 call sites) and currently roots its context atcontext.Background()(pkg/app/app.go). Instead of threading a parameter through all callers,app.Newreplaces that singlecontext.Background()withtelemetry.CommandContext()— Background-identical when tracing is off, span-rooted when on. Cancellation is unaffected (context.WithCancelon either parent behaves the same for SIGINT handling inmain.go).- Shutdown must run even on failure and on signals. The end of
rootCmd.Execute()inmain.go(both theerrChanand theSIGINT/SIGTERMpaths; on the error path beforeerrors.HandleExitCoder, which terminates viaOsExiter, and on the signal path afterapp.CleanWaitGroup.Wait()and beforeos.Exit(130/143)) calls the returned shutdown func with a 5s-timeout context so buffered spans are flushed. The stored shutdown must be nil-safe — a signal arriving beforePersistentPreRunEcompleted (i.e. beforeSetup) must be a no-op, never a panic. This is the one behavior that is easy to get wrong and is explicitly tested (§10). - Setup failures (e.g. unusable exporter configuration) are logged as a warning and do not fail the run — telemetry must never break deployments. Export errors after startup surface only through OTel's own error handler, wired to the zap logger.
4.3 Verified context-reality map (as of this writing)
All claims below were checked against the code; line numbers are anchors for reviewers:
| Path | Today | Consequence for tracing |
|---|---|---|
| cmd → app | app.New roots at context.Background(); ctx, Cancel = WithCancel(ctx) (pkg/app/app.go, New) |
span must be injected here (§4.2) |
| app → all helm execs | getHelm() constructs the ShellRunner with Ctx: a.ctx (pkg/app/app.go:1008); the resulting execer is cached per (helm binary, kube-context) in a.helms and shared by all releases and workers (pkg/app/app.go:982–1021) |
once App.ctx is span-rooted, every non-kubedog helm call nests automatically, with zero changes to getHelm. The shared-instance cache is also why per-release contexts must ride per-call parameters, never mutation of the shared execer |
| kubedog path (sync with tracking) | startBackgroundKubedogTracking(gocontext.Background(), …) (pkg/state/state.go:1294) → bufferHelmOutput derives releaseCtx := context.WithCancel(ctx) and swaps it in via execer.WithContext(releaseCtx) (pkg/state/helmx.go:363–365) |
helm execs on this path run on a Background-rooted context; runner-level spans would become orphan traces. Bridged in §4.4 |
| hooks | both event.Bus constructions (triggerGlobalReleaseEvent, triggerReleaseEvent, pkg/state/state.go:3666, 3703) duplicate the same literal and pass no Runner, so the default kicks in: ShellRunner{Dir: bus.BasePath, Logger: bus.Logger, Ctx: goContext.TODO()} with an inline comment acknowledging it should be app.Ctx (pkg/event/bus.go:61–71) |
hook execs are detached; spans would be orphans. Bridged in §4.4 |
| non-kubedog release workers | release loops (SyncReleases etc., pkg/state/state.go:1212 ff.) call the shared helmexec.Interface with a HelmContext (pkg/helmexec/context.go) that carries no go-context |
per-release spans need the §4.4 mechanism |
| subprocess funnel | exactly three ShellRunner construction sites exist (verified exhaustive): pkg/app/app.go:129 (Init, Ctx: a.ctx), pkg/app/app.go:1006 (getHelm, Ctx: a.ctx), and the hooks default (pkg/event/bus.go:62, Ctx: TODO — §4.4 bridge). Every external process helmfile itself starts goes through Execute/ExecuteStdIn (pkg/helmexec/runner.go); helm commands additionally funnel through execer.exec() (pkg/helmexec/exec.go:1207). Exception: kustomize executes inside the chartify library, outside this funnel (§12) |
one instrumentation point covers everything except chartify-internal execs; spans nest wherever the runner's Ctx carries a span |
Two pre-existing gaps surfaced by this analysis — kubedog tracking not being cancellable via
App.ctx, and hooks likewise — are out of scope for this proposal beyond trace-context
bridging (§4.4), because fixing their cancellation semantics would be a behavior change.
They should be reported as separate issues.
4.4 Context plan
Phase 1 (no state-package changes):
- Root span ctx via
telemetry.CommandContext()→app.New(§4.2). Everya.ctx-rooted exec nests. Discover/load/render spans inpkg/appalso need no signature changes: thehelmfile.loadspan starts insideloadDesiredStateFromYamlWithBaseDir(pkg/app/app.go:932) witha.ctxas parent — its two call sites (app.go:1045 and the nested-helmfile path at app.go:1280) are thereby both covered — andhelmfile.renderchildren parent through an unexportedctxfield on thedesiredStateLoaderstruct (pkg/app/desired_state_file_loader.go:29), set once at its single construction site (pkg/app/app.go:938). Zero method-signature changes, zero exported API. - Runner-level spans in
ShellRunner.Execute/ExecuteStdInnest for all non-kubedog, non-hook execs automatically. - Orphan-bridge for kubedog and hooks, with identical cancellation semantics:
pkg/state/state.go:1294: passcontext.WithoutCancel(telemetry.CommandContext())instead ofgocontext.Background().WithoutCancelpreserves values (the span) while dropping cancellation — andBackgroundnever carried cancellation anyway, so SIGINT/timeout behavior is bit-for-bit unchanged; only trace context is added. (Phase 1 uses the root span fromtelemetry.CommandContext(), which requires no new plumbing inpkg/state; phase 2 re-parents under the per-releasest.traceCtx.)pkg/event/bus.go: add an optionalCtx context.Contextfield toBus; the default runner construction usesbus.Ctxwhen set,TODOwhen nil (so behavior is unchanged for any nil-Ctx caller). The two construction sites inpkg/state/state.go:3666, 3703set it from the same span-rooted, cancel-stripped context.Dir/Loggerwiring of the default runner is untouched.
Phase 2 (per-release spans, still zero changes to helmexec.Interface signatures):
- Start
helmfile.release.<verb>spans in the seven release-worker loops — sixscatterGathersites inpkg/state/state.go(prepareSyncReleases:894,DeleteReleasesForSync:1135,SyncReleases:1241,PrepareCharts:2328,prepareDiffReleases:3068,DiffReleases:3266) plusiterateOnReleases(pkg/state/state_run.go:59, the shared loop behind test/lint/unittest-style iteration) — all following the samescatterGathershape. (An eighthscatterGathersites,scatterGatherEnvSecretFilesatpkg/state/create.go:486, decrypts environment secrets rather than processing releases; an optionalhelmfile.env_secretsspan there is a follow-up in the spirit of §4.5 item 6.) - Release spans are rooted via a new unexported
traceCtxfield onHelmState, exposed through a purely additive exported setter called bypkg/appright after state creation. (Why a setter and not a constructor parameter:st.loggeris injected throughstate.NewCreator— a 9-positional-parameter exported function,pkg/state/create.go:80— so threading a context through it would churn a public signature used by the app loader; an additive setter touches no existing signature and defaults to nil, i.e. current behavior.) - Carry the span context per call by adding an optional
Ctx context.Contextfield toHelmContext(pkg/helmexec/context.go), stamped insidecreateHelmContext(pkg/state/state.go:3168) so all eight call sites (state.go:1007, 1026, 1147, 1261, 3034, 3184, 3358, 3374) inherit it without per-call-site edits. - Funnel it inside
helmexec: theexecermethods that take aHelmContextpass itsCtxtoexecWithContext, which injects the release span into the runner's own context (trace.ContextWithSpan) rather than replacing it — cancellation authority stays with the runner (including the kubedog safety valve installed viaWithContext), which a naive context replacement would have overridden. A nilCtxbehaves exactly like the plain funnel. Helm-vs-other classification comes from an explicit marker stamped in the execer funnels (withRunnerCtx/markHelmRunner, value and pointer runner forms), not from the executable basename — so wrapper--helm-binarynames classify correctly. Theexectest.Helmfake is unaffected: app tests pre-seedApp.helmswith it (pkg/app/app_template_test.go:115–116), so it replaces the wholehelmexec.Interfaceand bypassesexecerinternals entirely. Why per-call threading rather than the existingWithContextclone:WithContext(pkg/helmexec/exec.go:251) is suited to whole-execution substitution (kubedog path), but release workers share one cached execer across concurrent workers (§4.3), so a per-release context must travel with the per-releaseHelmContextparameter.
4.5 Instrumentation points (in priority order)
helmexec.ShellRunner.Execute/ExecuteStdIn(pkg/helmexec/runner.go) — the single choke point for every external process started by helmfile itself: helm invocations, hooks, and helmfile plugin execs. One span per subprocess: namehelm.execwhencmdis the helm binary, elseos.exec. The one exception is kustomize, which runs inside thegithub.com/helmfile/chartifylibrary (see §5/§12). This instrumentation alone delivers most of the requested value (where does time go).- Root command span —
cmd/root.go(see §4.2). - State loading — span
helmfile.loadstarted insideloadDesiredStateFromYamlWithBaseDir(pkg/app/app.go:932, covering both callers incl. nested helmfiles), withhelmfile.render/helmfile.parsechildren intwo_pass_renderer.goparented via the loader-structctxfield (§4.4 step 1) — rendering is frequently the hidden time sink. - Per-release operations — release loops in
pkg/state/state.go(phase 2, §4.4). - Hooks —
pkg/event/bus.goTrigger(pkg/event/bus.go:56): spanhelmfile.hookwithhook.event,hook.nameattributes; naturally parents theos.execspan of the hook command once the §4.4 bridge is in place. - (Phase 2) kubedog wait spans, vals/remote secret resolution,
helm reporetry loops.
5. Span taxonomy
| Span name | Attributes (beyond standard otel.*) |
Notes |
|---|---|---|
helmfile <subcommand> (root) |
helmfile.command, helmfile.file, helmfile.environment, helmfile.selectors, helmfile.exit_code |
error status + recorded error on failure. Service identity (service.name, service.version) lives on the OTel resource, not on spans |
helmfile.discover_states |
helmfile.path |
findDesiredStateFiles (pkg/app/app.go:1642) |
helmfile.load |
helmfile.state_file |
one per file in helmfile.d, nested helmfiles |
helmfile.render |
helmfile.state_file, helmfile.pass=values|main |
two-pass rendering (pkg/app/two_pass_renderer.go) |
helmfile.repos.update |
— | wraps helm repo update |
helmfile.release.prepare |
helmfile.release, helmfile.namespace, helmfile.chart, helmfile.chart_version |
chart pull/build/registry login |
helmfile.release.sync / .diff / .template / .delete / .test / .lint / .unittest |
same as above + helmfile.labels |
one per selected release (phase 2) |
helmfile.hook |
hook.event (presync/…), hook.name |
|
helm.exec / os.exec |
exec.command, exec.args (strict-redacted, plus URL userinfo/query masking via RedactedRef), exec.redacted, exec.exit_code (on failure), helm.subcommand (helm only) |
one per external process; helm classification is marker-based (wrapper binaries included); release identity comes from the parent release span. Hooks, kustomize (invisible, §12), plugin execs land here too |
| Attribute values are strings/ints only; no structured payloads, no output capture in spans | ||
| (output already flows through logs). |
6. Security and redaction
Traces leave the machine they run on. Ground rules, checked against what exists today:
- The existing redaction is not sufficient on its own. The exit-error path
(
pkg/helmexec/exit_error.go:8–20) redacts only the argument following a flag whose name starts with--set(two-argument form). It does not cover the single-argument--set=key=valueform, nor credential-bearing flags such as--usernameor--password. (Registry passwords themselves already travel via stdin —--password-stdin,pkg/helmexec/exec.goRegistryLogin— but usernames appear in args.) Therefore: - One shared redaction implementation, two profiles — spans get the strict superset
without touching existing error output. Extract the exit-error redaction into an
exported helper in
pkg/helmexec(keeping it in the domain that owns flag semantics) with two profiles:legacy: byte-identical to today's exit-error behavior — the current goldens inpkg/helmexec/exit_error_test.go(which assert the exact--set/*** STRIP ***shape) keep passing unchanged. The exit-error path switches to this profile, so error messages are unchanged.strict: the superset required for spans — all--set*forms including--set=k=v, plus--username,--password,--key-file,--kube-token, ...; positional arguments are additionally passed throughhelmexec.RedactedRef, which masks go-getter forced forms (git::,s3::), whole URL userinfo, and credential-bearing query parameters, failing closed for malformed references. The previous token is always read from the original input so adjacent secret flags cannot leak. Both profiles are the same code path, so the span view is guaranteed at least as redacted as the error view. Unifying the two profiles (i.e. tightening exit-error messages too) would change observable output and is deliberately deferred to a separate follow-up PR with its own test updates — this proposal changes no existing message content.
- Never record:
vals://-resolved values, environment variables, exporter headers.OTEL_EXPORTER_OTLP_HEADERSis the only place collector credentials live; helmfile never logs it. - Chart/repo URLs go through the existing
redactedURL(pkg/helmexec/exec.go:184) — credentials embedded in URLs are masked. - Release names, namespaces, chart names, and label selectors are assumed non-secret (consistent with existing helmfile log output).
- Note for reviewers:
execer.exectoday logs the full unredacted command line at Debug level (pkg/helmexec/exec.go:1218). Span attributes deliberately do not mirror that log line; §10 pins this with a redaction test.
7. Performance and zero-impact guarantees
When disabled (the default):
telemetry.Tracerreturns the OTel no-op tracer. No-op span start/end is a few ns and allocation-free; provider setup, exporters, and the batch worker goroutine never start.telemetry.CommandContext()returnscontext.Background();app.Newbehaves exactly as today. No flag checks appear in hot loops.
When enabled:
- Standard SDK BatchSpanProcessor (5s interval / 512-span batches); metrics
use a periodic reader (default 60s,
OTEL_METRIC_EXPORT_INTERVAL) with a final flush at exit. Three instruments exist —helmfile.helm.exec.duration,helmfile.release.duration, andhelmfile.release.count— so metric cardinality stays tiny by default (bounded by subcommands and verbs, never by release names);HELMFILE_OTEL_METRICS_PER_RELEASEopts into name/namespace dimensions for bounded CI runs. Async --concurrency=16run produces at most one span per helm invocation plus one per release — hundreds, not tens of thousands. Export happens off the critical path; the only synchronous cost is the ≤5s shutdown flush, paid only when tracing is on. - SDK spans and exporters are goroutine-safe; helmfile's parallel release workers need no extra locking.
Functional-impact checklist (the review criteria for every PR in §11):
- No
helmexec.Interfacesignature changes in any phase;getHelm()unchanged. app.Newchanges one line (Background()→telemetry.CommandContext()), no call-site churn; cancellation semantics identical.- Kubedog and hook bridging uses
context.WithoutCancel, which drops cancellation and keeps values — the swapped-out parents (Background/TODO) never propagated cancellation either, so SIGINT/timeout behavior is unchanged. - Telemetry setup or export failures never fail or slow the run (warning log only, export off the critical path).
- All pre-existing behavior, including the known cancellation gaps of §4.3 and the exact content of exit-error messages (legacy redaction profile, §6.2), is preserved bit-for-bit; gap fixes and redaction unification are out of scope and filed separately.
8. Dependency impact
The OTel libraries are already in the module graph as indirect dependencies, required
transitively by existing direct dependencies (helm.sh/helm/v4 v4.2.4 requires
go.opentelemetry.io/otel v1.44.0; helm v3 and helmfile/vals also carry otel modules).
Every module this proposal would import — otel, otel/trace, otel/sdk,
otel/exporters/otlp/otlptrace/otlptracegrpc, .../otlptracehttp,
otel/exporters/stdout/stdouttrace, and contrib/exporters/autoexport v0.67.0 — is
already pinned in go.sum (verified), so promoting them to direct requires brings
zero new modules to download. No conflict with existing OTel usage in the process:
none of helmfile's direct dependencies registers OTel globals in library code — verified
no SetTracerProvider call sites under helm v3/v4 pkg/, vals, or chartify (helm's own
OTel wiring, where present, lives in its CLI layer, not the libraries helmfile imports).
Choosing autoexport (§3.2) also means
helmfile maintains no exporter-construction code; its RegisterSpanExporter hook covers
any future backend (e.g. zipkin) without helmfile changes. Note the resulting defaults:
http/protobuf on localhost:4318 unless the user overrides the protocol/endpoint.
9. Code layout of the change
cmd/root.go // --otel-tracing flag, Setup call, root span, shutdown handoff
main.go // shutdown on both exit and signal paths
pkg/config/global.go // OtelTracing option (+ accessor on GlobalImpl)
pkg/envvar/const.go // OtelTracing = "HELMFILE_OTEL_TRACING"
pkg/telemetry/… // new package (§4.1)
pkg/app/app.go // app.New: Background() → telemetry.CommandContext(); helmfile.load span; loader ctx wiring
pkg/app/desired_state_file_loader.go // ctx field on the unexported desiredStateLoader struct (render-span parent)
pkg/app/two_pass_renderer.go // helmfile.render/helmfile.parse spans
pkg/helmexec/runner.go // helm.exec / os.exec spans (single choke point)
pkg/helmexec/redact.go // shared args redaction, legacy+strict profiles (extracted from exit_error.go)
pkg/helmexec/exit_error.go // calls shared helper with legacy profile (output byte-identical)
pkg/helmexec/context.go // HelmContext.Ctx field (phase 2)
pkg/helmexec/exec.go // execCtx funnel beside exec/execStdIn (phase 2)
pkg/state/state.go // WithoutCancel bridges (kubedog call site + both event.Bus constructions); release spans (phase 2)
pkg/state/helmx.go // (no change — bridge happens at its caller)
pkg/event/bus.go // optional Ctx field consumed by the default runner
docs/experimental-features.md // feature entry → promoted out when stable
docs/otel.md // user guide (config, backends, CI recipes, sample trace)
10. Testing strategy
- Unit (
pkg/telemetry): table-driven tests for enabled/disabled no-op guarantees,CommandContext()identity, default resource attributes (service.name=helmfile,service.versionfrompkg/app/version), and propagator extraction ofTRACEPARENT. Global state reset viaexport_test.goso tests stay isolated. - Redaction (
pkg/helmexec): table-driven tests for both profiles —legacypinned byte-identical by the existingexit_error_test.gogoldens;strictcovering--set vand--set=k=v,--set-string/--set-file/--set-json,--username/--password, benign flags untouched. Pins §6.6. - Span-hierarchy golden tests: drive
Appwith the existing fake helm (pkg/exectest/helm.go, pre-seeded intoApp.helmsthe way current app tests do) and an in-memory exporter; assert the span tree (names, parent links, order) fortemplate/syncover a small fixture — catches context-plumbing regressions. Includes an orphan-span regression test for the kubedog and hook paths (§4.3): every exported span must have the root span as an ancestor. - OTLP end-to-end: an
httptestserver speaking OTLP/HTTP+protobuf, pointed at byOTEL_EXPORTER_OTLP_ENDPOINT; decode exported payloads (go.opentelemetry.io/proto/otlp, already in the graph) and assert count/attributes. Runs in unit-test context — no external collector needed in CI. - Lifecycle tests: shutdown flushes on command failure and on SIGINT, so spans never
vanish on failing deploys — the case users care about most. (Requires extracting
main.go's signal/select/exit logic into a small pure function; the extraction itself is behavior-preserving and covered by the same tests.) Also covers the nil-shutdown race noted in §4.2. - Regression suite: existing
make testmust pass unchanged with tracing compiled in but disabled. Note the existingpkg/apptests construct&App{...}literals directly and never callapp.New, so they do not guard the §4.2 one-liner — PR 1 adds a targeted test assertingapp.Newroots cancellation (and the span) exactly as before.
11. Documentation & rollout
docs/otel.md— user guide: enabling, env vars, collector recipes (Jaeger all-in-one, Grafana Tempo, vendor SaaS), CI correlation viaTRACEPARENT, sample trace reading. Linked fromdocs/index.md.- Feature listed under experimental in
docs/experimental-features.mdfor one or two minor releases (feedback on span taxonomy is the main thing that may change), then promoted to stable with a CHANGELOG entry. - PR sequence (each independently shippable and revertible, each measured against the
§7 functional-impact checklist):
- PR 1:
pkg/telemetry+ flag/env + root span + lifecycle +app.Newone-liner + docs + tests (§4.1–4.2). - PR 2:
ShellRunnerinstrumentation + shared redaction extraction/extension (§4.4 phase-1 steps 2–3, §6) — the core value. - PR 3: load/render spans + hook bridging + per-release spans via
HelmContext.Ctx/execCtx(§4.4 phase 2). - PR 4 (post-stabilization): metrics (e.g.
helmfile.helm.exec.durationhistogram,helmfile.release.countby result), kubedog/vals spans.
- PR 1:
12. Future work (explicitly out of scope for v1)
Metrics pipeline on the same provider (— delivered:otel/sdk/metricwith the same env-var config)helmfile.helm.exec.duration,helmfile.release.count, andhelmfile.release.durationviaautoexport.NewMetricReader(OTEL_METRICS_EXPORTER); per-release name/namespace attributes on the duration histogram are opt-in viaHELMFILE_OTEL_METRICS_PER_RELEASE(bounded dimensions by default).TRACEPARENTinjection into helm subprocess env so chart-test hooks / plugins can extend the helmfile trace.- Subprocesses started inside
github.com/helmfile/chartify(kustomize, and any helm calls chartify makes) are invisible toShellRunnerinstrumentation; options are a wrapper span around each chartify call inpkg/state, or upstream OTel support in chartify. - Log-to-trace correlation (zap OTel appender).
- Remaining per-release instrumentation: the flag-preparation loops
(
prepareSyncReleases/prepareDiffReleases) and the exec nesting forhelmexec.Interfacemethods that take noHelmContext(TemplateRelease/Lint/Unittest/Fetch) — their exec spans currently sit flat under the load span while the release span covers the surrounding work. - Fixing the kubedog/hook cancellation gaps (§4.3) — separate issues filed from this design.
13. Alternatives considered
| Alternative | Why rejected |
|---|---|
| Structured logs + collector-side parsing | No hierarchy/timing guarantees; every backend needs custom parsing; poor UX. |
| Prometheus metrics only | Shows counts/durations but not critical paths or nesting; the request is explicitly about tracing where time goes. |
| Hand-rolled exporter selection in helmfile | autoexport already exists in the dependency graph and implements the spec env vars (verified: v0.67.0 supports otlp/console/none, protocol dispatch, none detection); hand-rolled code is pure maintenance burden and would drift from the spec. |
| Helmfile-specific env vars for endpoint/headers etc. | Duplicates the OTel spec; standard vars are already what platform teams configure. |
Full context.Context refactor of pkg/state first |
Large, risky churn unrelated to the feature; the HelmContext.Ctx/execCtx design achieves nesting without it. |
Using the existing WithContext clone for per-release spans |
The cached, shared execer (a.helms) serves concurrent workers; per-release contexts must travel with the per-call HelmContext parameter, not via instance substitution. |
Config-file (helmfile.yaml) telemetry settings |
Telemetry is an operational/platform concern, not state authoring; flag+env matches how it's injected in CI. |
14. Open questions
- Propagation into helm subprocesses: should helmfile inject
TRACEPARENTinto the child process env (opt-in) so hooks/plugins can continue the trace? (§12) - Span-name taxonomy stability: do we commit to the §5 names as stable API for dashboard authors during the experimental window, or reserve the right to rename? Proposal: rename freely while experimental, freeze on promotion.
helmfile.dparallel mode: state-file spans are siblings under the root span — is a synthetichelmfile.parallelgrouping span wanted, or does flat suffice?- Trace output when
--log-level=debug: duplicate a compact span tree to stderr at shutdown for quick local triage without a collector? - Root span noise for trivial commands:
helmfile version/helpalso runPersistentPreRunE— export a root span for them, or skip? (Proposal: skip via a short denylist; cosmetic.)