* feat: add --repo-retries for retrying helm repo and registry login commands
Add a configurable retry mechanism for chart repository operations to
handle unstable networks (corporate proxies, slow internal registries).
Closes#1894
- New --repo-retries N flag and HELMFILE_REPO_RETRIES env var (default 0
= opt-in, backward compatible)
- Retry applies to helm repo add, helm repo update (incl. ACR), and
helm registry login with exponential backoff (1s, 2s, 4s, ..., capped 30s)
- Single retryRepoOp helper; per-attempt args/buffer are local to avoid
state leaking across retries
- Tests cover succeed-after-retry, exhausted-retries, disabled-by-default,
and regression guards for password-buffer and args-accumulation
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: address PR review (overflow guard, cancellable sleep, flag-override, docs)
Address Copilot review feedback on #2683:
- Cap backoff shift exponent at 5 to prevent time.Duration overflow on
large --repo-retries values
- Make retry sleep context-aware (sleepCtx) so Ctrl+C aborts the retry
loop promptly via the ShellRunner context
- Log a concise exit status instead of the verbose ExitError dump, and
clarify the retry-counter wording ('retry N/M')
- Use -1 sentinel as the CLI default so --repo-retries=0 can explicitly
disable retries even when HELMFILE_REPO_RETRIES is set
- Align help text and docs: retry applies 'on failure' (not just
transient errors), document the 0-disables behavior
- Add tests for overflow guard, cancellable sleep, and flag-zero-disables
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: abort retries on canceled context, hide sentinel default, align comment
Address follow-up Copilot review on #2683:
- Fix tight-loop bug: sleepCtx now returns whether it completed vs was
interrupted by context cancellation, and retryRepoOp aborts the retry
loop on interruption so Ctrl+C no longer spins into rapid helm calls
- Hide the -1 sentinel from --help by overriding the displayed default
to 0 (pflag DefValue), matching the documented default while keeping
the flag-override semantics
- Correct HelmExecOptions.RepoRetry comment: 'on failure' not 'transient
network errors', matching the actual retry behavior
- Add Test_Retry_AbortsOnCanceledContext covering the no-tight-loop path
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: copy args per retry in RegistryLogin, make cancel test deterministic
Address follow-up Copilot review on #2683:
- RegistryLogin: pass a per-attempt copy of args to execStdIn so its
internal append (for helm.extra) can't alias the shared slice across
retries
- Test_Retry_AbortsOnCanceledContext: cancel the context deterministically
inside the op closure after the first attempt, replacing the flaky
time.Sleep(20ms) goroutine
Signed-off-by: yxxhero <aiopsclub@163.com>
* fix: return error on unknown managed repo type instead of silent skip
Address Copilot review on #2683: AddRepo logged an error for an unknown
managed type but returned nil, silently succeeding while skipping the
repo add. Now returns an error so misconfigurations fail loudly.
Signed-off-by: yxxhero <aiopsclub@163.com>
---------
Signed-off-by: yxxhero <aiopsclub@163.com>
* feat: support more HELMFILE_* env vars as flag fallbacks
Adds env-var fallbacks for global flags, mirroring the existing
HELMFILE_ENVIRONMENT / HELMFILE_KUBE_CONTEXT pattern:
* --helm-binary -> HELMFILE_HELM_BINARY
* --kustomize-binary -> HELMFILE_KUSTOMIZE_BINARY
* --log-level -> HELMFILE_LOG_LEVEL
* --debug -> HELMFILE_DEBUG (expecting "true" lower case)
* --quiet -> HELMFILE_QUIET (expecting "true" lower case)
* --no-color -> HELMFILE_NO_COLOR (expecting "true" lower case),
additionally honors NO_COLOR per no-color.org
(any non-empty value disables color)
Flag values still take precedence; env vars are consulted only when the
flag is unset. The string-flag default values ("helm", "kustomize",
"info") move into the accessor methods so the env-var fallback can
actually trigger when no flag is passed.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* docs: mention new HELMFILE_* env vars in cli.md and templating.md
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* fix: make Color/NoColor/env interaction consistent
Two issues with the env-aware NoColor() introduced together with
HELMFILE_NO_COLOR / NO_COLOR support:
1. Color() consulted the raw GlobalOptions.NoColor field instead of
NoColor(), so in a TTY with only the env set, Color() fell through
to terminal autodetect and ValidateConfig() spuriously errored with
"--color and --no-color cannot be specified at the same time".
2. NoColor() returned true via env even when --color was explicitly
passed, so `helmfile --color` with NO_COLOR (or HELMFILE_NO_COLOR=true)
in the environment hit the same ValidateConfig() error. A flag should
always win over an env var.
Fix both by routing Color() through NoColor() and giving NoColor() an
explicit --color short-circuit. Regression tests added for both paths.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
---------
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* feat: support HELMFILE_NAMESPACE env var for default namespace
Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --namespace
CLI flag takes precedence, falling back to HELMFILE_NAMESPACE when
unset.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* docs: mention HELMFILE_NAMESPACE in cli.md and templating.md
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
---------
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* feat: support HELMFILE_KUBE_CONTEXT env var for default kube context
Mirrors the existing HELMFILE_ENVIRONMENT pattern: the --kube-context
CLI flag takes precedence, falling back to HELMFILE_KUBE_CONTEXT when
unset.
Refs #1213.
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
* docs: mention HELMFILE_KUBE_CONTEXT in cli.md and templating.md
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>
---------
Signed-off-by: Dominik Schmidt <dev@dominik-schmidt.de>