Files
helmfile/pkg/agent/doctor/analyze.go
T
yxxhero 9b943adc9e feat: add helmfile doctor command for AI-assisted diff analysis (#2660)
* feat: add `helmfile doctor` command for AI-assisted diff analysis

`helmfile doctor` runs `helmfile diff` and asks an OpenAI-compatible LLM to
summarize the changes and flag risks (data loss, security exposure, breaking
changes, downtime, performance, best-practice issues).

Key design decisions:
- When no LLM is configured, doctor is equivalent to `helmfile diff` with
  one exception: --show-secrets is always forced off (secrets never reach
  stdout, even without an LLM).
- Secrets are ALWAYS redacted via two layers: (1) ShowSecrets() forced to
  false so helm-diff emits <REDACTED> placeholders; (2) a defense-in-depth
  text redactor strips residual secret-looking content (Secret YAML blocks,
  sensitive key/value lines, base64 blobs, JWT tokens) before LLM transmission.
- LLM configuration precedence: env (HELMFILE_LLM_*) < helmfile.yaml (llm:)
  < CLI flags (--llm-*).
- Supports any OpenAI-compatible backend (OpenAI, Azure, One-API, LiteLLM,
  Ollama, etc.) with automatic response_format fallback for backends that
  don't support JSON mode.
- Prompt injection defense: release names and environment values are
  JSON-encoded before insertion into the LLM prompt.
- Exit codes: 0 (success/low-risk), 2 (high-risk gate, bypass with --force),
  1 (other errors). Helm-diff's 'detected changes' exit-2 is swallowed.

New packages:
- pkg/agent/llm: OpenAI-compatible client with JSON response parsing, mock
  client for testing, prompt builder with injection defense.
- pkg/agent/doctor: secret redactor (state machine + regex), report renderer
  (markdown + JSON), config resolver (env < yaml < flag merge).

Testing: 70+ unit tests covering redaction patterns, prompt injection,
response_format fallback, JSON parsing, yaml roundtrip, concurrency safety,
panic recovery, and error propagation. go test -race passes.

Documentation: full doctor section in docs/cli.md, llm: block reference in
docs/configuration.md, updated skills/helmfile for AI agents.

Signed-off-by: yxxhero <aiopsclub@163.com>

* docs: fix doctor equivalence wording per PR review

Per review feedback (PR #2660): the docs claimed doctor is 'equivalent to
helmfile diff — same flags, same output, same exit codes' in the unconfigured
path, but this over-promises because:

  1. doctor --output is the report format (not helm-diff's output format)
  2. helm-diff's --output is exposed as --diff-output in doctor
  3. --show-secrets is silently ignored

Updated all three locations (cli.md, cmd/doctor.go Long + godoc, pkg/app/doctor.go
godoc) to say 'falls back to helmfile diff with --show-secrets forced off' and
explicitly note the --output / --diff-output flag difference.

Signed-off-by: yxxhero <aiopsclub@163.com>

---------

Signed-off-by: yxxhero <aiopsclub@163.com>
2026-06-22 16:52:35 +08:00

110 lines
3.6 KiB
Go

// Package doctor orchestrates the `helmfile doctor` flow: capture helm diff
// output, hand it to an LLM via the OpenAI-compatible Chat Completions
// protocol, then render a structured risk report.
//
// When no LLM is configured (APIKey or Model missing), doctor degrades to
// plain `helmfile diff` with zero behavior change.
package doctor
import (
goContext "context"
"time"
"github.com/helmfile/helmfile/pkg/agent/llm"
)
// Result bundles everything `helmfile doctor` needs to render its output and
// decide on an exit code.
type Result struct {
// Analysis is the structured LLM output. Nil when the LLM was not called
// (diff empty) or when LLMCallFailed is set.
Analysis *llm.Analysis
// RawDiff is the helm diff text AFTER secret redaction. Doctor never
// exposes unredacted secret content through Result.
RawDiff string
// SecretsRedacted counts secret-looking values stripped from RawDiff.
// Surfaced in the report footer so users can spot unexpected leaks.
SecretsRedacted int
// LLMCallFailed indicates the LLM was configured but the call failed;
// the caller should degrade to printing RawDiff with a warning.
LLMCallFailed bool
// LLMError is the underlying LLM error when LLMCallFailed is true.
LLMError error
// Model is the model identifier used (for the report footer).
Model string
// Duration is how long the LLM call took.
Duration time.Duration
}
// HasHighRisk delegates to Analysis.HasHighRisk when an analysis exists.
func (r Result) HasHighRisk() bool {
return r.Analysis != nil && r.Analysis.HasHighRisk()
}
// Options controls a single Analyze run.
type Options struct {
// Client is the LLM client. When nil, Analyze returns a Result that
// signals "no LLM" so the caller can degrade to plain diff.
Client llm.Client
// Environment is the --environment value (may be empty).
Environment string
// Releases is the list of release names that appear in the diff (may be nil).
Releases []string
// Model is the model identifier; surfaced in the report footer.
Model string
// Redactor controls secret redaction. A zero-value SecretRedactor is
// fine: its Redact method falls back to "<REDACTED>". Redaction is
// ALWAYS applied — there is no opt-out at this layer.
Redactor SecretRedactor
}
// Analyze runs the full doctor pipeline against the given diff text.
//
// Behavior:
// - Empty diff → empty Result (caller should print nothing).
// - diff is ALWAYS redacted via opts.Redactor first. RawDiff in the
// returned Result is the redacted text; the original is discarded.
// - nil Client → Result with RawDiff only (the "unconfigured" path).
// Redaction still happens — pipes downstream shouldn't see raw secrets.
// - Non-nil Client → calls LLM; on error returns LLMCallFailed=true.
func Analyze(ctx goContext.Context, diff string, opts Options) Result {
if diff == "" {
return Result{}
}
redacted, redactionCount := opts.Redactor.Redact(diff)
if opts.Client == nil {
return Result{
RawDiff: redacted,
SecretsRedacted: redactionCount,
}
}
start := time.Now()
a, err := opts.Client.Analyze(ctx, redacted, llm.AnalyzeInput{
Environment: opts.Environment,
Releases: opts.Releases,
})
duration := time.Since(start)
if err != nil {
return Result{
RawDiff: redacted,
SecretsRedacted: redactionCount,
LLMCallFailed: true,
LLMError: err,
Model: opts.Model,
Duration: duration,
}
}
return Result{
Analysis: &a,
RawDiff: redacted,
SecretsRedacted: redactionCount,
Model: opts.Model,
Duration: duration,
}
}