* feat: add per-release continue-on-error support Work in progress for discussion #2799. This introduces the initial release-level configuration needed to continue processing independent releases after a deployment failure. The DAG execution and failure propagation behavior are still being implemented. Refs #2799 Signed-off-by: Axel Delille <Axel.delille31@gmail.com> * fix: address review findings for continueOnError Review fixes for the per-release continueOnError feature: 1. Regenerate values-file ID hashes in pkg/state/temp_test.go — adding ContinueOnError to ReleaseSpec shifts generateValuesID hashes, which broke TestGenerateID (the test file notes these must be regenerated whenever ReleaseSpec changes). 2. Remove the dead skippedErrors variable in withBatches and instead log each skipped release with logger.Warnf at decision time, so users see why a release never ran instead of only learning from the final error list. 3. Aggregate all errors from a state file instead of returning only errs[0] in visitStatesWithContext/processStateFileParallel. With continueOnError, multiple releases can fail or be skipped in one run; reporting only the first error hid the skip errors (and other failures), contradicting the feature's contract. Single-error rendering is unchanged. 4. Gate tolerated errors on ReleaseErrorCodeFailure so that non-failure release errors (e.g. helm-diff's "changes detected" exit code 2 can never enable continuation or block dependents. 5. Report skipped-release messages with the dependency's plain release name instead of its kubeContext/namespace-qualified needs id (e.g. dependency database instead of default/default/database), matching the documented message format. 6. Restructure the withBatches helpers into guard-clause style (filterBlockedReleases, toleratesBatchErrors) and drop the test-only logger nil-guards in favor of a nop logger in tests. 7. Add end-to-end coverage through App.Sync with the exectest fake helm: independent releases continue after a tolerated failure, dependents are skipped with an explicit error, the exit code stays non-zero, and fail-fast remains the default without continueOnError. Also cover the non-failure error code case at the withBatches level. 8. Use new(true) instead of a boolPtr helper (CI-enforced check-modernize) and document the failure-handling interaction in docs/releases.md. Signed-off-by: yxxhero <aiopsclub@163.com> EOF ) Signed-off-by: yxxhero <aiopsclub@163.com> --------- Signed-off-by: Axel Delille <Axel.delille31@gmail.com> Signed-off-by: yxxhero <aiopsclub@163.com> Co-authored-by: yxxhero <aiopsclub@163.com>
9.1 KiB
Releases & DAG
DAG-aware installation/deletion ordering with needs
needs controls the order of the installation/deletion of the release:
releases:
- name: somerelease
needs:
- [[KUBECONTEXT/]NAMESPACE/]anotherelease
Be aware that you have to specify the kubecontext and namespace name if you configured one for the release(s).
All the releases listed under needs are installed before(or deleted after) the release itself.
For the following example, helmfile [sync|apply] installs releases in this order:
- logging
- servicemesh
- myapp1 and myapp2
- name: myapp1
chart: charts/myapp
needs:
- servicemesh
- logging
- name: myapp2
chart: charts/myapp
needs:
- servicemesh
- logging
- name: servicemesh
chart: charts/istio
needs:
- logging
- name: logging
chart: charts/fluentd
Note that all the releases in a same group is installed concurrently. That is, myapp1 and myapp2 are installed concurrently.
On helmfile [delete|destroy], deletions happen in the reverse order.
That is, myapp1 and myapp2 are deleted first, then servicemesh, and finally logging.
Failure handling and continueOnError
By default, helmfile [sync|apply] stops on the first release failure (fail-fast). Set continueOnError: true on a release to keep processing independent releases after it fails:
- name: myapp1
chart: charts/myapp
continueOnError: true
needs:
- servicemesh
When a release with continueOnError fails, releases in other branches of the DAG are still processed, while releases that (transitively) depend on the failed release are skipped with an error like release "myapp2" was skipped because dependency "myapp1" failed. The command still exits non-zero, and Helm's own rollback behavior (atomic, rollbackOnFailure, ...) is unaffected. See Continue On Error for details.
Selectors and needs
When using selectors/labels, needs are ignored by default. This behaviour can be overruled with a few parameters:
| Parameter | default | Description |
|---|---|---|
--skip-needs |
true |
needs are ignored (default behavior). |
--include-needs |
false |
The direct needs of the selected release(s) will be included. |
--include-transitive-needs |
false |
The direct and transitive needs of the selected release(s) will be included. |
Let's look at an example to illustrate how the different parameters work:
releases:
- name: serviceA
chart: my/chart
needs:
- serviceB
- name: serviceB
chart: your/chart
needs:
- serviceC
- name: serviceC
chart: her/chart
- name: serviceD
chart: his/chart
| Command | Included Releases Order | Explanation |
|---|---|---|
helmfile -l name=serviceA sync |
- serviceA |
By default no needs are included. |
helmfile -l name=serviceA sync --include-needs |
- serviceB- serviceA |
serviceB is now part of the release as it is a direct need of serviceA. |
helmfile -l name=serviceA sync --include-transitive-needs |
- serviceC- serviceB- serviceA |
serviceC is now also part of the release as it is a direct need of serviceB and therefore a transitive need of serviceA. |
Note that --include-transitive-needs will override any potential exclusions done by selectors or conditions. So even if you explicitly exclude a release via a selector it will still be part of the deployment in case it is a direct or transitive need of any of the specified releases.
Separating helmfile.yaml into multiple independent files
Once your helmfile.yaml got to contain too many releases,
split it into multiple yaml files.
Recommended granularity of helmfile.yaml files is "per microservice" or "per team". And there are two ways to organize your files.
- Single directory
- Glob patterns
Single directory
helmfile -f path/to/directory loads and runs all the yaml files under the specified directory, each file as an independent helmfile.yaml.
The default helmfile directory is helmfile.d, that is,
in case helmfile is unable to locate helmfile.yaml, it tries to locate helmfile.d/*.yaml.
By default, multiple files in helmfile.d are processed in parallel for better performance. If you need files to be processed sequentially in alphabetical order (e.g., for dependency ordering where databases must be deployed before applications), use the --sequential-helmfiles flag.
For example, you can use a <two digit number>-<microservice>.yaml naming convention to control the sync order when using --sequential-helmfiles:
helmfile.d/00-database.yaml01-backend.yaml02-frontend.yaml
# Process files sequentially in alphabetical order
helmfile --sequential-helmfiles sync
Note: When processing multiple helmfile.d files, both parallel and sequential modes resolve paths without changing the process working directory, so relative environment variables like
KUBECONFIGwork correctly.
Glob patterns
In case you want more control over how multiple helmfile.yaml files are organized, use helmfiles: configuration key in the helmfile.yaml:
Suppose you have multiple microservices organized in a Git repository that looks like:
myteam/(sometimes it is equivalent to a k8s ns, that iskube-systemforclusteropsteam)apps/filebeat/helmfile.yaml(nocharts/exists because it depends on the stable/filebeat chart hosted on the official helm charts repository)README.md(each app managed by my team has a dedicated README maintained by the owners of the app)
metricbeat/helmfile.yamlREADME.md
elastalert-operator/helmfile.yamlREADME.mdcharts/elastalert-operator/<the content of the local helm chart>
The benefits of this structure is that you can run git diff to locate in which directory=microservice a git commit has changes.
It allows your CI system to run a workflow for the changed microservice only.
A downside of this is that you don't have an obvious way to sync all microservices at once. That is, you have to run:
for d in apps/*; do helmfile -f $d diff; if [ $? -eq 2 ]; then helmfile -f $d sync; fi; done
At this point, you'll start writing a Makefile under myteam/ so that make sync-all will do the job.
It does work, but you can rely on the Helmfile feature instead.
Put myteam/helmfile.yaml that looks like:
helmfiles:
- apps/*/helmfile.yaml
So that you can get rid of the Makefile and the bash snippet.
Just run helmfile sync inside myteam/, and you are done.
All the files are sorted alphabetically per group = array item inside helmfiles:, so that you have granular control over ordering, too.
selectors
When composing helmfiles you can use selectors from the command line as well as explicit selectors inside the parent helmfile to filter the releases to be used.
helmfiles:
- apps/*/helmfile.yaml
- path: apps/a-helmfile.yaml
selectors: # list of selectors
- name=prometheus
- tier=frontend
- path: apps/b-helmfile.yaml # no selector, so all releases are used
selectors: []
- path: apps/c-helmfile.yaml # parent selector to be used or cli selector for the initial helmfile
selectorsInherited: true
- When a subhelmfile has explicit
selectors, those selectors determine which releases from that subhelmfile are considered; parent and CLI selectors are not combined with them for release filtering. - When CLI selectors are provided (e.g.
helmfile -l name=b sync) and a subhelmfile has explicit selectors that are provably incompatible with them (same key, different value), that subhelmfile may be skipped entirely without loading or rendering it. For example, with-l name=b, a subhelmfile withselectors: [name=a]will be skipped since no release could match both. This optimization does not apply whenselectorsInherited: trueis set or when no CLI selectors are provided. Use--debugto see log messages about skipped subhelmfiles. - When not selector is specified there are 2 modes for the selector inheritance because we would like to change the current inheritance behavior (see issue #344 ).
- Legacy mode, sub-helmfiles without selectors inherit selectors from their parent helmfile. The initial helmfiles inherit from the command line selectors.
- explicit mode, sub-helmfile without selectors do not inherit from their parent or the CLI selector. If you want them to inherit from their parent selector then use
selectorsInherited: true. To enable this explicit mode you need to set the following environment variableHELMFILE_EXPERIMENTAL=explicit-selector-inheritance(see experimental).
- Using
selector: []will select all releases regardless of the parent selector or cli for the initial helmfile - using
selectorsInherited: truemake the sub-helmfile selects releases with the parent selector or the cli for the initial helmfile. You cannot specify an explicit selector while usingselectorsInherited: true