mirror of
https://github.com/helmfile/helmfile.git
synced 2026-09-30 03:25:36 +02:00
feat: add --repo-retries for helm repo and registry login commands (#2683)
* feat: add --repo-retries for retrying helm repo and registry login commands Add a configurable retry mechanism for chart repository operations to handle unstable networks (corporate proxies, slow internal registries). Closes #1894 - New --repo-retries N flag and HELMFILE_REPO_RETRIES env var (default 0 = opt-in, backward compatible) - Retry applies to helm repo add, helm repo update (incl. ACR), and helm registry login with exponential backoff (1s, 2s, 4s, ..., capped 30s) - Single retryRepoOp helper; per-attempt args/buffer are local to avoid state leaking across retries - Tests cover succeed-after-retry, exhausted-retries, disabled-by-default, and regression guards for password-buffer and args-accumulation Signed-off-by: yxxhero <aiopsclub@163.com> * fix: address PR review (overflow guard, cancellable sleep, flag-override, docs) Address Copilot review feedback on #2683: - Cap backoff shift exponent at 5 to prevent time.Duration overflow on large --repo-retries values - Make retry sleep context-aware (sleepCtx) so Ctrl+C aborts the retry loop promptly via the ShellRunner context - Log a concise exit status instead of the verbose ExitError dump, and clarify the retry-counter wording ('retry N/M') - Use -1 sentinel as the CLI default so --repo-retries=0 can explicitly disable retries even when HELMFILE_REPO_RETRIES is set - Align help text and docs: retry applies 'on failure' (not just transient errors), document the 0-disables behavior - Add tests for overflow guard, cancellable sleep, and flag-zero-disables Signed-off-by: yxxhero <aiopsclub@163.com> * fix: abort retries on canceled context, hide sentinel default, align comment Address follow-up Copilot review on #2683: - Fix tight-loop bug: sleepCtx now returns whether it completed vs was interrupted by context cancellation, and retryRepoOp aborts the retry loop on interruption so Ctrl+C no longer spins into rapid helm calls - Hide the -1 sentinel from --help by overriding the displayed default to 0 (pflag DefValue), matching the documented default while keeping the flag-override semantics - Correct HelmExecOptions.RepoRetry comment: 'on failure' not 'transient network errors', matching the actual retry behavior - Add Test_Retry_AbortsOnCanceledContext covering the no-tight-loop path Signed-off-by: yxxhero <aiopsclub@163.com> * fix: copy args per retry in RegistryLogin, make cancel test deterministic Address follow-up Copilot review on #2683: - RegistryLogin: pass a per-attempt copy of args to execStdIn so its internal append (for helm.extra) can't alias the shared slice across retries - Test_Retry_AbortsOnCanceledContext: cancel the context deterministically inside the op closure after the first attempt, replacing the flaky time.Sleep(20ms) goroutine Signed-off-by: yxxhero <aiopsclub@163.com> * fix: return error on unknown managed repo type instead of silent skip Address Copilot review on #2683: AddRepo logged an error for an unknown managed type but returned nil, silently succeeding while skipping the repo add. Now returns an error so misconfigurations fail loudly. Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: yxxhero <aiopsclub@163.com>
This commit is contained in:
@@ -135,6 +135,13 @@ func setGlobalOptionsForRootCmd(fs *pflag.FlagSet, globalOptions *config.GlobalO
|
||||
fs.BoolVar(&globalOptions.DisableForceUpdate, "disable-force-update", false, `do not force helm repos to update when executing "helm repo add" (Helm 3 only)`)
|
||||
fs.BoolVar(&globalOptions.EnforcePluginVerification, "enforce-plugin-verification", false, `fail plugin installation if verification is not supported (for security purposes)`)
|
||||
fs.BoolVar(&globalOptions.HelmOCIPlainHTTP, "oci-plain-http", false, `use plain HTTP for OCI registries (required for local/insecure registries in Helm 4)`)
|
||||
fs.IntVar(&globalOptions.RepoRetry, "repo-retries", -1, `Number of times to retry "helm repo add/update" and "helm registry login" on failure, with exponential backoff (1s, 2s, 4s, ..., capped at 30s). Set to 0 to disable retries. Overrides "HELMFILE_REPO_RETRIES" OS environment variable when specified`)
|
||||
// The actual default is -1 (a sentinel meaning "flag not set, fall back to
|
||||
// the env var"); display "0" in --help to match the documented default and
|
||||
// avoid confusing users with a negative value.
|
||||
if f := fs.Lookup("repo-retries"); f != nil {
|
||||
f.DefValue = "0"
|
||||
}
|
||||
fs.BoolVarP(&globalOptions.Quiet, "quiet", "q", false, `Silence output. Equivalent to log-level warn. Overrides "HELMFILE_QUIET" OS environment variable when specified`)
|
||||
fs.StringVar(&globalOptions.Kubeconfig, "kubeconfig", "", "Use a particular kubeconfig file")
|
||||
fs.StringVar(&globalOptions.KubeContext, "kube-context", "", `Set kubectl context. Overrides "HELMFILE_KUBE_CONTEXT" OS environment variable when specified. Uses current kubectl context by default`)
|
||||
|
||||
Reference in New Issue
Block a user