Files
helmfile/docs/releases.md
T
Tsukoyachiandyxxhero ca58090af4 feat: add per-release continue-on-error support (#2804)
* feat: add per-release continue-on-error support

Work in progress for discussion #2799.

This introduces the initial release-level configuration needed to
continue processing independent releases after a deployment failure.
The DAG execution and failure propagation behavior are still being
implemented.

Refs #2799

Signed-off-by: Axel Delille <Axel.delille31@gmail.com>

* fix: address review findings for continueOnError

Review fixes for the per-release continueOnError feature:

1. Regenerate values-file ID hashes in pkg/state/temp_test.go — adding
   ContinueOnError to ReleaseSpec shifts generateValuesID hashes, which
   broke TestGenerateID (the test file notes these must be regenerated
   whenever ReleaseSpec changes).

2. Remove the dead skippedErrors variable in withBatches and instead log
   each skipped release with logger.Warnf at decision time, so users see
   why a release never ran instead of only learning from the final error
   list.

3. Aggregate all errors from a state file instead of returning only
   errs[0] in visitStatesWithContext/processStateFileParallel. With
   continueOnError, multiple releases can fail or be skipped in one run;
   reporting only the first error hid the skip errors (and other
   failures), contradicting the feature's contract. Single-error
   rendering is unchanged.

4. Gate tolerated errors on ReleaseErrorCodeFailure so that non-failure
   release errors (e.g. helm-diff's "changes detected" exit code 2 can
   never enable continuation or block dependents.

5. Report skipped-release messages with the dependency's plain release
   name instead of its kubeContext/namespace-qualified needs id
   (e.g. dependency database instead of default/default/database),
   matching the documented message format.

6. Restructure the withBatches helpers into guard-clause style
   (filterBlockedReleases, toleratesBatchErrors) and drop the test-only
   logger nil-guards in favor of a nop logger in tests.

7. Add end-to-end coverage through App.Sync with the exectest fake helm:
   independent releases continue after a tolerated failure, dependents
   are skipped with an explicit error, the exit code stays non-zero, and
   fail-fast remains the default without continueOnError. Also cover the
   non-failure error code case at the withBatches level.

8. Use new(true) instead of a boolPtr helper (CI-enforced check-modernize)
   and document the failure-handling interaction in docs/releases.md.

Signed-off-by: yxxhero <aiopsclub@163.com>
EOF
)
Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: Axel Delille <Axel.delille31@gmail.com>
Signed-off-by: yxxhero <aiopsclub@163.com>
Co-authored-by: yxxhero <aiopsclub@163.com>
2026-09-21 07:48:50 +08:00

9.1 KiB

Releases & DAG

DAG-aware installation/deletion ordering with needs

needs controls the order of the installation/deletion of the release:

releases:
- name: somerelease
  needs:
  - [[KUBECONTEXT/]NAMESPACE/]anotherelease

Be aware that you have to specify the kubecontext and namespace name if you configured one for the release(s).

All the releases listed under needs are installed before(or deleted after) the release itself.

For the following example, helmfile [sync|apply] installs releases in this order:

  1. logging
  2. servicemesh
  3. myapp1 and myapp2
  - name: myapp1
    chart: charts/myapp
    needs:
    - servicemesh
    - logging
  - name: myapp2
    chart: charts/myapp
    needs:
    - servicemesh
    - logging
  - name: servicemesh
    chart: charts/istio
    needs:
    - logging
  - name: logging
    chart: charts/fluentd

Note that all the releases in a same group is installed concurrently. That is, myapp1 and myapp2 are installed concurrently.

On helmfile [delete|destroy], deletions happen in the reverse order.

That is, myapp1 and myapp2 are deleted first, then servicemesh, and finally logging.

Failure handling and continueOnError

By default, helmfile [sync|apply] stops on the first release failure (fail-fast). Set continueOnError: true on a release to keep processing independent releases after it fails:

  - name: myapp1
    chart: charts/myapp
    continueOnError: true
    needs:
    - servicemesh

When a release with continueOnError fails, releases in other branches of the DAG are still processed, while releases that (transitively) depend on the failed release are skipped with an error like release "myapp2" was skipped because dependency "myapp1" failed. The command still exits non-zero, and Helm's own rollback behavior (atomic, rollbackOnFailure, ...) is unaffected. See Continue On Error for details.

Selectors and needs

When using selectors/labels, needs are ignored by default. This behaviour can be overruled with a few parameters:

Parameter default Description
--skip-needs true needs are ignored (default behavior).
--include-needs false The direct needs of the selected release(s) will be included.
--include-transitive-needs false The direct and transitive needs of the selected release(s) will be included.

Let's look at an example to illustrate how the different parameters work:

releases:
- name: serviceA
  chart: my/chart
  needs:
  - serviceB
- name: serviceB
  chart: your/chart
  needs:
  - serviceC
- name: serviceC
  chart: her/chart
- name: serviceD
  chart: his/chart
Command Included Releases Order Explanation
helmfile -l name=serviceA sync - serviceA By default no needs are included.
helmfile -l name=serviceA sync --include-needs - serviceB
- serviceA
serviceB is now part of the release as it is a direct need of serviceA.
helmfile -l name=serviceA sync --include-transitive-needs - serviceC
- serviceB
- serviceA
serviceC is now also part of the release as it is a direct need of serviceB and therefore a transitive need of serviceA.

Note that --include-transitive-needs will override any potential exclusions done by selectors or conditions. So even if you explicitly exclude a release via a selector it will still be part of the deployment in case it is a direct or transitive need of any of the specified releases.

Separating helmfile.yaml into multiple independent files

Once your helmfile.yaml got to contain too many releases, split it into multiple yaml files.

Recommended granularity of helmfile.yaml files is "per microservice" or "per team". And there are two ways to organize your files.

  • Single directory
  • Glob patterns

Single directory

helmfile -f path/to/directory loads and runs all the yaml files under the specified directory, each file as an independent helmfile.yaml. The default helmfile directory is helmfile.d, that is, in case helmfile is unable to locate helmfile.yaml, it tries to locate helmfile.d/*.yaml.

By default, multiple files in helmfile.d are processed in parallel for better performance. If you need files to be processed sequentially in alphabetical order (e.g., for dependency ordering where databases must be deployed before applications), use the --sequential-helmfiles flag.

For example, you can use a <two digit number>-<microservice>.yaml naming convention to control the sync order when using --sequential-helmfiles:

  • helmfile.d/
    • 00-database.yaml
    • 01-backend.yaml
    • 02-frontend.yaml
# Process files sequentially in alphabetical order
helmfile --sequential-helmfiles sync

Note: When processing multiple helmfile.d files, both parallel and sequential modes resolve paths without changing the process working directory, so relative environment variables like KUBECONFIG work correctly.

Glob patterns

In case you want more control over how multiple helmfile.yaml files are organized, use helmfiles: configuration key in the helmfile.yaml:

Suppose you have multiple microservices organized in a Git repository that looks like:

  • myteam/ (sometimes it is equivalent to a k8s ns, that is kube-system for clusterops team)
    • apps/
      • filebeat/
        • helmfile.yaml (no charts/ exists because it depends on the stable/filebeat chart hosted on the official helm charts repository)
        • README.md (each app managed by my team has a dedicated README maintained by the owners of the app)
      • metricbeat/
        • helmfile.yaml
        • README.md
      • elastalert-operator/
        • helmfile.yaml
        • README.md
        • charts/
          • elastalert-operator/
            • <the content of the local helm chart>

The benefits of this structure is that you can run git diff to locate in which directory=microservice a git commit has changes. It allows your CI system to run a workflow for the changed microservice only.

A downside of this is that you don't have an obvious way to sync all microservices at once. That is, you have to run:

for d in apps/*; do helmfile -f $d diff; if [ $? -eq 2 ]; then helmfile -f $d sync; fi; done

At this point, you'll start writing a Makefile under myteam/ so that make sync-all will do the job.

It does work, but you can rely on the Helmfile feature instead.

Put myteam/helmfile.yaml that looks like:

helmfiles:
- apps/*/helmfile.yaml

So that you can get rid of the Makefile and the bash snippet. Just run helmfile sync inside myteam/, and you are done.

All the files are sorted alphabetically per group = array item inside helmfiles:, so that you have granular control over ordering, too.

selectors

When composing helmfiles you can use selectors from the command line as well as explicit selectors inside the parent helmfile to filter the releases to be used.

helmfiles:
- apps/*/helmfile.yaml
- path: apps/a-helmfile.yaml
  selectors:          # list of selectors
  - name=prometheus
  - tier=frontend
- path: apps/b-helmfile.yaml # no selector, so all releases are used
  selectors: []
- path: apps/c-helmfile.yaml # parent selector to be used or cli selector for the initial helmfile
  selectorsInherited: true
  • When a subhelmfile has explicit selectors, those selectors determine which releases from that subhelmfile are considered; parent and CLI selectors are not combined with them for release filtering.
  • When CLI selectors are provided (e.g. helmfile -l name=b sync) and a subhelmfile has explicit selectors that are provably incompatible with them (same key, different value), that subhelmfile may be skipped entirely without loading or rendering it. For example, with -l name=b, a subhelmfile with selectors: [name=a] will be skipped since no release could match both. This optimization does not apply when selectorsInherited: true is set or when no CLI selectors are provided. Use --debug to see log messages about skipped subhelmfiles.
  • When not selector is specified there are 2 modes for the selector inheritance because we would like to change the current inheritance behavior (see issue #344 ).
    • Legacy mode, sub-helmfiles without selectors inherit selectors from their parent helmfile. The initial helmfiles inherit from the command line selectors.
    • explicit mode, sub-helmfile without selectors do not inherit from their parent or the CLI selector. If you want them to inherit from their parent selector then use selectorsInherited: true. To enable this explicit mode you need to set the following environment variable HELMFILE_EXPERIMENTAL=explicit-selector-inheritance (see experimental).
  • Using selector: [] will select all releases regardless of the parent selector or cli for the initial helmfile
  • using selectorsInherited: true make the sub-helmfile selects releases with the parent selector or the cli for the initial helmfile. You cannot specify an explicit selector while using selectorsInherited: true