Commit Graph
281 Commits
Author SHA1 Message Date
Cody LeeandCursor 912ea91ac0 feat(protect): export air quality readings when present
Prefer a known stats channel and fall back to airQuality for temperature
and humidity. Omit channels whose status is unknown, and add the
particulate and gas series only when the controller includes them.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-29 09:46:13 -05:00
Kevin 2152c020f9 Update uap.go
add band label to per-VAP and per-radio metrics
2026-09-23 13:30:25 +02:00
Jonathan Delgado 160b46d616 fix(otelunifi): callback memory leak 2026-09-17 13:44:48 -07:00
Iain 4da7cd5936 fix(promunifi): accept a bracketed IPv6 http_listen in the health check 2026-09-12 00:14:17 +01:00
Cody LeeandCursor 122226bc4b fix: honor scrape-cache disable and export locate-mode devices
interval=0 now turns off the Prometheus scrape cache as PR #1014
documented, and sub-15s intervals warn instead of clamping (#1083).
Adopted devices stay in Prometheus, Influx, OTel, and Datadog exports
while locate/identify is on (#1075).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 08:51:32 -05:00
Iain e2384b8bb8 fix(promunifi): default an unset http_listen instead of failing validation 2026-09-03 01:12:20 +01:00
Prototype0645 6c08d43b4c fix(inputunifi): apply default_site_name_override to log entries
default_site_name_override is applied in augmentMetrics, which only the
metrics path goes through. Log entries leave by collectControllerEvents,
which reads its sites straight from getFilteredSites — where the override
is deliberately not applied — so events, alarms, IDS records and system-log
entries ship with the controller's stock site name.

Four of the five site-scoped log collectors were affected; collectAnomalies
already did this inline, which is what makes the omission visible.

The consequence is worst for a poller watching several UniFi OS consoles:
each one calls its only site "default", so their log entries are
indistinguishable downstream. In Loki every stream from those consoles lands
under site_name="Default (default)" no matter which console it came from,
and no relabeling downstream can separate them again — the information is
gone by then. That is precisely the case the option was added for, and it
works for the metrics from those same consoles.

Extract the check collectAnomalies was doing into overrideSiteName and call
it from all five, so the two paths agree.

Note that only Site.Name reaches an API path; Site.SiteName is a display
name throughout the library. The override is therefore safe here, which is
what applySiteNameOverride's own comment already says ("keeping default for
API calls").
2026-09-01 09:08:13 +02:00
Prototype0645 ca64173b35 feat(wan): label WAN interface state with its source
unpoller_wan_interface_state ships with site_name and nothing else to
identify where it came from. That is not enough: every UniFi OS console
names its only site "default", so site_name is "Default (default)" on all
of them. An instance polling several controllers emits WAN status that is
indistinguishable, and downstream attribution has to be guessed.

Observed in production before the fix: a scrape config guessing from
site_name filed a UDM's WAN state under the wrong customer. No error, no
missing series — just wrong data under someone else's name.

unifi/v6.1.0 (#244) added SourceName to WANStatus, stamped from the
controller URL in all three read paths. go.mod is already on v6.1.0, so
this only has to read the field.

Two lines, same shape as #1071 which did this for unpoller_wan_*.

Tests use the fakeReport already present in the package: every emitted
metric carries both site_name and source, the raw state is kept as a
label, and a nil status still produces nothing rather than panicking a
poll. TestExportWANStatusIsAttributed fails on master and passes here.
2026-08-31 23:42:30 +02:00
Cody LeeandCursor 8a78355e12 fix: use external test package for remote site helpers
golangci-lint testpackage requires tests in inputunifi_test.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:01:11 -04:00
Cody LeeandCursor e4c2934bdc fix: poll remote multi-site controllers by internalReference
Remote discovery stored Integration display names, so extra sites were dropped when checkSites compared them to legacy Site.Name. Use unifi v6.1.0 InternalReference instead.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 13:54:21 -04:00
Cooper Ry LeesandClaude Opus 5 a73a2e0ab2 fix: don't send unused credentials to a Protect-only console
Found by running this against a real UNVR4 rather than only the fake one.

setDefaults fills an unset user with the placeholder "unifipoller". For a
Protect-only console that placeholder was reaching Login(), the console answered
403, and the controller died during initialisation -- the same symptom #1066 set
out to fix, one layer further in. A Protect API key and no local account is the
config an operator actually writes for a UNVR, so this was the common case, not
an edge one.

Nothing on such a console uses a session unless Protect logs are wanted: the
Integration API authenticates with the key alone. So withhold the credentials
entirely in that case, and say so in the startup summary rather than naming a
username that is never sent.

Also pins the go.mod bump to unifi v6.0.1, which is the release that carries
NewProtectClient (unpoller/unifi#240).

Verified end to end against a UNVR4 (UniFi OS 5.1.31, Protect 7.2.105): the
controller comes up, logs "Auth: Protect API key only (no session needed)",
makes no login request at all, and exports 37 unpoller_protect_* series across
9 cameras and 2 bridges with unpoller_controller_up = 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVreutpEATmBjm6PBw9RjQ
2026-08-31 13:08:34 +00:00
Cooper Ry LeesandClaude Opus 5 31cc5a0f31 feat: poll Protect-only consoles via disable_network (closes #1066)
A UNVR or UNVR Pro runs UniFi Protect with no Network application installed.
UnPoller could not poll one at all: NewUnifi ends with GetServerData(), a GET of
/proxy/network/status, which such a console answers with its UniFi OS SPA HTML.
The controller entry died during initialisation and re-failed every interval,
never even printing a config summary -- while the Protect Integration API on the
same host answered every endpoint with the same key.

Set disable_network = true on that controller. It defaults to false, so nothing
about an existing config changes.

The Protect collectors were already complete and already not site-scoped; three
things stood between them and a Protect-only console:

  - getUnifi now calls unifi.NewProtectClient, which skips the Network probe and
    validates the Protect Integration API instead (unpoller/unifi#240).

  - pollController aborted on getFilteredSites long before reaching
    collectProtect, and collectControllerEvents did the same before
    collectProtectLogs. The Network pass is extracted into pollNetwork and
    skipped wholesale; the event collector list reduces to collectProtectLogs,
    the only site-independent one.

  - Metrics counted a poll successful only if it produced devices or clients. A
    Protect-only console produces neither, so a filtered scrape of one -- the
    Prometheus per-target path -- fell through to the dynamic-controller branch
    and reported ErrDynamicLookupsDisabled despite a successful collection.
    ProtectDevices now counts too.

Two smaller things worth calling out for reviewers:

  - extractDevices dereferenced metrics.Devices unguarded. That was already a
    latent panic; skipping the Network pass makes it reachable, so it is fixed
    here rather than left for the first person to hit it.

  - RawMetrics answers the raw-path kind for these consoles and rejects the
    site-scoped kinds with ErrNetworkDisabled. Returning an empty result would
    read as "this console has no devices" rather than "wrong question".

warnProtectOnly logs an error, without failing the controller, for the two
configurations that can never collect anything: disable_network with neither
Protect save flag, and save_protect_devices with no key to authenticate with.
Silently collecting nothing is the failure mode hardest to spot in a log.

pkg/inputunifi had no tests before this. input_test.go follows inputunas'
input_test.go: an httptest fake UNVR serving the console's SPA HTML for
everything but the Protect paths and the login, covering initialisation,
metrics, events, the filtered scrape, RawMetrics, both warnings, config binding
across toml/json/yaml/env, and that the shipped examples leave Network enabled.
TestProtectOnlyControllerFailsWithoutFlag pins the original bug against that
same console, so the flag is demonstrably what makes the difference.

Requires github.com/unpoller/unifi/v6 with NewProtectClient (unpoller/unifi#240).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVreutpEATmBjm6PBw9RjQ
2026-08-31 13:08:27 +00:00
Prototype0645 24329e278e feat(wan): label WAN metrics with their site and source
Every unpoller_wan_* series shipped with site_name="" and source="". The
exporter said so itself:

    cfg.WANLoadBalanceType,
    "", // site_name - will be set by caller if available
    "", // source - will be set by caller if available

The caller had nothing to set them from: WANEnrichedConfiguration carried
no identity. unifi/v6.0.3 fixes that upstream — GetWANEnrichedConfiguration
now stamps SiteName and SourceName from the site it fetched, the same way
GetSiteDPI does.

This bumps to v6.0.3 and fills the labels in. Two slices needed it, not
one: the base label set and the provider label set built further down for
the isp_name/isp_city descriptors. The test caught the second, which I had
missed.

Why it matters: an instance polling several controllers emitted WAN
metrics that were indistinguishable from one another, since wan_id is the
only other distinguishing label. Attributing them downstream meant
hardcoding a mapping in the scrape config and hoping no second controller
ever gained a gateway — when one does, its metrics are silently filed
under the wrong customer. No error, no missing series, just wrong data.

Tests use the fakeReport already present in the package. They assert every
emitted metric carries both labels, and that a nil configuration still
produces nothing rather than panicking a poll.
TestExportWANIsAttributed fails on master and passes with this change.
2026-08-31 12:06:58 +02:00
Cody LeeandCursor ce4cfe3971 fix: respect configured default_site_name_override in remote mode
Remote API discovery was unconditionally overwriting default_site_name_override
with the console name for Cloud Gateways. Only apply the console name fallback
when the user has not already configured an override.

Fixes #1057

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 16:57:41 -05:00
Cody LeeandCursor 9bc7f2c5bf Complete InfluxDB v3 rollout: tests, docs, and docker example.
Add v3 integration and version tests, migration notes, InfluxDB 3 docker-compose stack, and README updates to finish the remaining plan phases.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 16:24:09 -05:00
Cody LeeandCursor 5e31d2f252 Add InfluxDB v3 support and fix tag/field schema collisions.
Introduce explicit version selection with the influxdb3-go client alongside existing v1 and v2 paths, and resolve overlapping tag/field keys required for InfluxDB 3 write validation.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 15:49:59 -05:00
Cody Lee cb639e0c5c Merge branch 'master' into feat/protect-device-metrics 2026-08-27 09:58:11 -05:00
Iain 12ac4ed7f7 fix: influxdb and datadog reported tx packets under stat_rx_packets 2026-08-23 00:19:29 +01:00
Cody Lee 65ad210f20 feat: add UniFi Protect device metrics (closes #1015)
Collects Protect device data (sensors, cameras, lights, bridges, link
stations, NVR) via the official Integration API and exports it through
Prometheus, InfluxDB, and DataDog. Opt-in via save_protect_devices,
gated by protect_api_key.

Bumps github.com/unpoller/unifi to v6, which added the Protect API
client (breaking change: FlexInt/FlexBool/FlexFloat replace nullable
pointers).
2026-08-19 15:14:43 -05:00
Cody LeeandClaude Opus 5 01ac7ca2b7 refactor(unas): replace disable flag with enable, defaulting to off
`disable = false` is a double negative, and a bool named disable cannot
express opt-in anyway: it zero-values to false, so the flag was inert and
opt-in rested entirely on the device list being empty.

`enable` defaults to false and is now the real gate -- Initialize, Metrics
and DebugInput all return early unless it is set. The two existing guards
remain: an empty device list is still a no-op, and no default URL is ever
synthesized.

Configuring devices while enable is false is always a mistake, so that
combination logs one error instead of silently collecting nothing.

Adds binding tests for the flag across toml, json, yaml and UP_UNAS_ENABLE
(the env name derives from the xml tag, not the json one), plus a test that
all three shipped examples default to off. Both were verified by mutation:
breaking a struct tag or flipping an example fails the suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 09:34:02 -05:00
Cody LeeandClaude Opus 5 3d178bed8d fix: AppendMetrics silently dropped SpeedTests and the batch timestamp
AppendMetrics merges every slice field on Metrics except SpeedTests, so
speed test results were discarded on the way from the input to the
outputs. Both call sites start from a non-nil &Metrics{}, so this hit
every user on every poll: inputunifi collected the results, and the
export code in promunifi, influxunifi and datadogunifi was dead.

TS was dropped the same way. The aggregate starts bare and nothing
restored the timestamp, so it stayed zero -- and influxunifi's collect()
stamps any point that carries no timestamp of its own with the
aggregate's, which meant the zero time. Both Influx clients omit a zero
timestamp and let the server assign one, so the damage was limited to
points being stamped on arrival rather than at poll time, but it made
the fallback path meaningless. First writer wins: the earliest input's
timestamp is the one that describes the batch.

The failure mode here is what makes it worth guarding rather than just
patching. A new metric family needs a field on Metrics and an append
line in AppendMetrics, and omitting the second loses every metric in
that family with no error, no log line, and a passing build.
TestAppendMetricsCoversEverySliceField walks the struct by reflection
and fails naming any slice field that is not merged, so a new field is
covered the moment it is declared rather than when someone notices the
graphs are empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 07:59:04 -05:00
Cody LeeandClaude Opus 5 d5dfc771d5 feat: add opt-in UNAS Pro support (closes #785)
Adds a new `unas` input plugin that polls UNAS Pro storage consoles and
exports console health, storage pools, disks and shares to Prometheus,
InfluxDB and DataDog.

UNAS is a separate plugin rather than a device type inside inputunifi
because a storage-only console has no Network application: it cannot
answer /status, has no sites, and shares none of the UniFi device schema.

The plugin is opt-in and inert until an operator names a console. Opt-in
is expressed as "no devices configured" rather than a `disable` flag,
because a bool named `disable` zero-values to false and so cannot make a
plugin default-off. Initialize returns silently on an empty device list
and, unlike inputunifi, nothing synthesizes a default URL.

Two behaviours are worth calling out for reviewers:

  - Metrics returns (metrics, nil) whenever any console was collected.
    poller.collectMetrics uses `if err != nil {} else if metric != nil`,
    so returning both would discard every healthy console because one
    failed. Only a total failure returns an error.

  - Re-auth fires on total failure, not on a 401. A mid-session 401 from
    GetData surfaces as ErrInvalidStatusCode, not ErrAuthenticationFailed,
    so there is no sentinel to match on. Session expiry fails all four
    endpoints at once, which is exactly the total-failure case.

Prometheus metrics use the `unifi_unas_` prefix, which diverges from the
`unas_` prefix used by the reference implementation; dashboards built
against that will need query edits.

Requires unifi/v5 v5.31.0 for the UNAS client and structs.

Credit to alexgreenbank/unaspoller for mapping the endpoints.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 07:52:37 -05:00
Cody Lee 3dee271de0 inputunifi: skip alarms endpoint on 400 api.err.InvalidObject
Some Network 10.x+ controllers (UniFi OS, Network 10.5.67 confirmed)
return HTTP 400 api.err.InvalidObject from list/alarm instead of a 404
when the endpoint is gone, so collectAlarms fell through the existing
ErrEndpointNotFound skip and logged a real ERROR on every poll.

Fixes #1050
2026-08-18 09:41:03 -05:00
Cody LeeandClaude Sonnet 5 58294420f8 promunifi: add state label to wan_interface_state metric
BACKUP and DISCONNECTED WAN interfaces both reported as 0, making
them indistinguishable in Prometheus (fixes #1045). InfluxDB and
DataDog already expose the raw state; this brings Prometheus in
line by adding a state label alongside the existing 1/0 gauge.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-18 09:32:14 -05:00
Cody Lee 12073ba523 Fix nil-pointer panic in collectEvents for disabled input plugins
InputUnifi.Events returns (nil, nil) when disabled, but collectEvents
dereferenced e.Logs unconditionally after a successful (err == nil)
call, crashing the poller. Same crash class as #1030, found while
verifying the recover-based fix.
2026-08-18 08:36:50 -05:00
Cody LeeandClaude Sonnet 5 8063f06777 Extract recover-and-convert-panic helpers for input plugin calls
Replaces the sent-bool guard duplicated across three goroutines with a
small per-call helper (recoverInitialize/recoverEvents/recoverMetrics)
that wraps the plugin call and converts a panic into a returned error.
Each goroutine then always sends its result exactly once, so there's
no risk of a double-send deadlocking the collector.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-18 08:31:34 -05:00
Cody LeeandClaude Sonnet 5 cc60dc9239 Recover input plugin panics in poller goroutines (fixes #1030)
Metrics/Events/Initialize each fan out to input plugins in their own
goroutines with no recover(), so a panic there (e.g. a UniFi
controller returning an unexpected Site Speed Test aggregated-dashboard
payload) crashes the whole process with exit code 2. Because the
panic occurs in a child goroutine, promunifi's existing safeRefresh
recover() in the caller's goroutine never sees it, which is why the
crash survived the earlier robustness work. This converts a panicking
input into a logged/returned error so polling continues instead of
crashing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-17 15:36:37 -05:00
Sebastian PetersandClaude Opus 5 04f0010fce feat(inputunifi): add save_speedtest toggle for the WAN speed test poll
The Site Speed Test poll against the controller's aggregated-dashboard
endpoint has always run unconditionally, so operators whose controllers
misbehave on that endpoint have no way to skip the request. Every other
optional collection already has a save_* flag; this adds the missing one.

Defaults to true in all three default paths (local defaults,
per-controller defaults and remote discovery) so existing setups keep
behaving exactly as before.

Refs #1030

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 23:11:45 +02:00
Paul Sütterlin 9aeaa0bbb7 feat: implement 1x_identity support 2026-07-12 13:17:12 -05:00
Miguel Deus df0baba630 feat(promunifi): add band label to per-VAP and per-radio metrics
unifi_device_vap_* and unifi_device_radio_* expose the radio identifier
(radio=ng/na/6e) but not the frequency band, so downstream consumers must
hard-code the ng/na/6e -> 2.4/5/6 GHz mapping to group or join by band.

UniFi does not carry band as its own field on the radio/VAP structs, so this
derives it from the radio identifier via a small radioBand() helper and adds
it as a "band" label (GHz: 2.4/5/6) on every series exportVAPtable and
exportRADtable emit. This mirrors the band label already present on the
RogueAP metrics, giving the per-radio/per-VAP series the same dimension.

Derived internally from p.Radio / v.Radio, so the exportVAPtable /
exportRADtable signatures are unchanged and the UAP/UDM/UDB/UBB call sites
need no edits.
2026-06-17 14:38:30 -05:00
Miguel Deus 4ec6fa2010 feat(promunifi): add mac label to per-VAP and per-radio metrics
unifi_device_vap_* and unifi_device_radio_* previously exposed only
(site_name, name, source) as per-device identifiers, none of which is a
stable primary key (name is user-editable, source is the controller
hostname). The device MAC is already available on the parent UAP/UDM/UDB/UBB
struct, so this threads d.Mac through exportVAPtable / exportRADtable and
adds it as a label on every series those functions emit.

This lets downstream consumers join VAP/radio series directly against
unifi_device_info (which already carries mac) without going through the
(site_name, name, source) intersection.

Applied at all four call sites: UAP, UDM, UDB, UBB.
2026-06-09 08:26:06 -05:00
ek-docker-images 6d1b1361fc feat(promunifi): export DOCSIS CI state metrics for UCI devices
Adds two new Prometheus metrics for UCI (UniFi Cable Internet) devices:

- unpoller_device_ci_state_operational (gauge): 1 if DOCSIS CI state ==
  "Operational", else 0. Suitable for alerting on cable link health.

- unpoller_device_ci_state_info (gauge=1, info-style): exposes the full
  DOCSIS CI state table as labels (ci_state, ci_sw_dl_status, ci_mac,
  ci_version, ci_mode) for diagnostic dashboards.

The controller does expose a top-level `internet` boolean on UCI
devices, but it is not a reliable WAN-reachability signal — it stays
false even when the cable link is fully operational and the upstream
WAN is up. The UCI is a cable bridge with no independent internet-
reachability probe; real WAN health lives on the upstream gateway
(e.g. UDM wan1.up).

The ci_state field from ci_state_table IS reported reliably and is the
correct source-of-truth for DOCSIS link health.

Verified on a UCI in Operational state (ci_mode=D3.1).
2026-05-24 15:58:25 -07:00
Cody LeeandClaude Opus 4.7 dabfeffe66 fix(prometheus): serve scrapes from cached background poll (#1013)
Decouples Prometheus scrape cadence from upstream UniFi API calls so a
429 backoff loop on the controller side no longer stalls /metrics. The
output plugin now owns a 60s background poller (configurable) whose
result is served from an in-memory cache. Concurrent /scrape requests
for the same target are coalesced via singleflight to prevent a noisy
scraper from multiplying upstream load.

Adds two new metrics so operators can detect cache staleness and
refresh failures independently:
- unpoller_prometheus_cache_age_seconds
- unpoller_prometheus_refresh_failures_total

Background goroutine recovers from panics so a malformed input payload
no longer silently kills refreshes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 17:43:44 -05:00
Cody LeeandClaude Opus 4.7 fef3ae74f2 test: update integration expectations for new UAP uplink fields
Influx and Datadog integration tests assert that the captured field/gauge
sets exactly match the YAML. Add the new uap uplink_* entries so the
TestInfluxV1Integration and Datadog integration tests stay green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 09:27:30 -05:00
Cody LeeandClaude Opus 4.7 b1a8d60460 feat: add UAP uplink metrics and Prometheus parity for USW/UBB/UDB (closes #988)
Exposes the uplink medium (wire vs wireless) and link speed for UniFi access
points so users can detect when an AP downgrades from gigabit to fast ethernet,
which was the original ask in #988. UAPs previously had zero uplink coverage
in any output plugin; now influxunifi, datadogunifi, and promunifi all report
uplink_type, uplink_speed, uplink_max_speed, and related fields.

Also brings Prometheus to parity with Influx/Datadog by emitting uplink
metrics for USW, UBB, and UDB devices (previously only USG/UDM/UXG had them
in promunifi). A new exportDeviceUplink helper in promunifi/usg.go reuses
the existing unpoller_device_uplink_* descriptors to avoid descriptor
collision (per c48b9917).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 09:23:41 -05:00
Cody LeeandClaude Opus 4.7 511c524e6e feat(influxunifi): add global tags applied to every measurement
Closes #1001. Mirrors the DataDog plugin's global tags feature for
InfluxDB. Per-metric tags take precedence on key collision so
site/device identifiers can never be overwritten by a misconfigured
global. Configurable via TOML/JSON/YAML under influxdb.tags.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 08:59:16 -05:00
Jim Strang c48b9917b0 fix(promunifi): avoid descriptor collision on unpoller_device_uptime_seconds
descIntegrationDevice was registered with namespace prefix
"unpoller_device_", producing "unpoller_device_uptime_seconds" with
labels {device_id} and a different help string than the existing
descDevice() metric of the same FQDN (labels {type, site_name, name,
source, tag}, help "Device Uptime"). Prometheus MustRegister panics on
inconsistent descriptors for the same fully-qualified name, causing
v3.0.0 to crashloop on startup whenever the Prometheus output was
enabled.

Move the Integration/v1 device metrics under a dedicated
"integration_device_" name prefix, matching the convention used by the
other Integration/v1 collectors added in the same release (e.g.
wifi_broadcast_*, acl_rule_*, mclag_domain_*, pending_device_*), where
the bare namespace prefix is passed in and the type prefix is baked
into each metric name string.

Affected metric renames:
  unpoller_device_uptime_seconds              -> unpoller_integration_device_uptime_seconds
  unpoller_device_cpu_utilization_pct         -> unpoller_integration_device_cpu_utilization_pct
  unpoller_device_memory_utilization_pct      -> unpoller_integration_device_memory_utilization_pct
  unpoller_device_load_average_{1,5,15}min    -> unpoller_integration_device_load_average_{1,5,15}min
  unpoller_device_radio_tx_retries_pct        -> unpoller_integration_device_radio_tx_retries_pct
  unpoller_device_uplink_{rx,tx}_rate_bps     -> unpoller_integration_device_uplink_{rx,tx}_rate_bps

Fixes #1002
Fixes #1004
2026-05-09 08:19:44 -04:00
Cody LeeandClaude Opus 4.7 d2948b8bd0 feat: upgrade unifi to v5.26.0 and add Integration/v1 + new legacy metrics
Adds 21 new data types from unifi v5.26.0 across all metric output plugins
(InfluxDB, Prometheus, DataDog). Per-site Integration/v1 calls are gated on
API key configuration and only run for user-configured sites; ErrEndpointNotFound
is handled gracefully so older firmware continues to work without log spam.

Also migrates events collection (collectAlarms, collectAnomalies, collectEvents,
collectIDs, collectProtectLogs) to handle Network 10.x+ endpoint removals via
ErrEndpointNotFound, with debug-level logging to avoid per-poll noise.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 16:45:47 -05:00
Cody LeeandClaude Sonnet 4.6 c596e82cf2 fix: use v2 traffic API as DPI fallback for Network 9.1+ firmware (#985)
The legacy /stat/stadpi and /stat/sitedpi endpoints return empty data
on UniFi Network 9.1+ (issue #834). The v2 /traffic endpoint already
existed in the unifi library and in the collector, but was only called
when both SaveTraffic and SaveDPI were enabled — most users only set
SaveDPI=true and never saw any data.

- Remove the SaveTraffic gate on GetClientTraffic; call it whenever
  SaveDPI is enabled, treating it as a DPI data source
- Downgrade GetClientTraffic errors to debug-log so old firmware that
  lacks the v2 endpoint continues to use the legacy API without error
- Add convertToSiteDPI to aggregate per-client v2 data into per-site
  DPITable entries, filling SitesDPI when the legacy endpoint is empty
- Legacy API results are preserved; v2 data only supplements sites not
  already covered, so old-firmware users are unaffected

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-28 09:42:35 -05:00
Cody LeeandClaude Sonnet 4.6 2f1e28c7d3 chore: apply linter auto-fixes (wsl_v5, nlreturn, tagalign) (#984)
golangci-lint auto-fixes across multiple packages:
- wsl_v5: blank lines between logical blocks
- nlreturn: newlines before return statements
- tagalign: struct field tag alignment

No logic changes.

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 21:10:52 -05:00
Cody LeeandClaude Sonnet 4.6 18c6e66a8e feat: add Site Magic site-to-site VPN metrics (closes #926) (#983)
* feat: add Site Magic site-to-site VPN metrics (closes #926)

Bump github.com/unpoller/unifi/v5 to v5.25.0 which adds:
- GetMagicSiteToSiteVPN / GetMagicSiteToSiteVPNSite API methods
- MagicSiteToSiteVPN types with mesh, connection, device, and status structs
- Missing VPN health fields on Site.Health (SiteToSiteNumActive/Inactive,
  SiteToSiteRxBytes/TxBytes/RxPackets/TxPackets)

Implement VPN metrics collection across all output plugins:
- Collect Site Magic VPN mesh data per-site in inputunifi pollController
- Propagate VPNMeshes through poller.Metrics / AppendMetrics
- Apply DefaultSiteNameOverride for VPN meshes in augmentMetrics /
  applySiteNameOverride
- influxunifi: vpn_mesh, vpn_mesh_connection, vpn_mesh_status tables
- promunifi: vpn_mesh_*, vpn_tunnel_*, vpn_mesh_status_* gauges
- datadogunifi: unifi.vpn_mesh.*, unifi.vpn_tunnel.*, unifi.vpn_mesh_status.*

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(otelunifi): add Site Magic VPN metrics to OpenTelemetry output

Adds exportVPNMeshes to the otel output plugin, emitting the same
unifi_vpn_mesh_*, unifi_vpn_tunnel_*, and unifi_vpn_mesh_status_*
gauges as the other output plugins.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 21:08:09 -05:00
Cody LeeandClaude Sonnet 4.6 a81a6e6e16 feat: add port anomaly metrics (closes #929) (#982)
Collect port anomalies from the UniFi v2 API endpoint
/proxy/network/v2/api/site/{site}/ports/port-anomalies and export
them to all output plugins (Prometheus, InfluxDB, DataDog, OpenTelemetry).

Metrics exported per port:
- port_anomaly_count     – number of anomaly events
- port_anomaly_last_seen – unix timestamp of last event

Labels: site_name, source, device_mac, port_idx, anomaly_type

Bumps github.com/unpoller/unifi/v5 to v5.24.0 which adds GetPortAnomalies.

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 18:56:37 -05:00
Cody LeeandClaude Sonnet 4.6 643c108674 feat: add network topology metrics (closes #931) (#981)
Bumps github.com/unpoller/unifi/v5 to v5.23.0 which adds
GetTopology() fetching vertices (devices/clients) and edges
(wired/wireless connections) from /proxy/network/v2/api/site/{site}/topology.

Changes across the stack:
- poller.Metrics: add Topologies []any field + AppendMetrics support
- inputunifi: collect topology per-site (non-fatal on older controllers),
  pass through augmentMetrics with site name override support
- promunifi: new topology.go with summary, connection-type, link-quality,
  and band-distribution gauges
- influxunifi: new topology.go with topology_summary and topology_edge
  measurements
- datadogunifi: new topology.go with equivalent Datadog gauges
- otelunifi: new topology.go with OpenTelemetry gauge observations

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 18:44:51 -05:00
Cody LeeandClaude Sonnet 4.6 6b33b6b97b feat: firewall policy metrics across all output plugins (closes #928) (#979)
* feat(promunifi): add firewall policy metrics (closes #928)

Bump unifi client to v5.22.0 and wire up firewall policy data end-to-end:

- poller.Metrics: add FirewallPolicies []any slice
- inputunifi: collect GetFirewallPolicies() per poll cycle; apply
  DefaultSiteNameOverride; augment into poller.Metrics
- promunifi: export per-rule (rule_enabled, rule_index) and per-site
  aggregate metrics (rules_total, rules_enabled, rules_disabled,
  rules_by_action, rules_predefined, rules_custom, rules_logging_enabled)

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat: export firewall policies to influx, datadog, and otel outputs

Extends firewall policy support (PR #979) to all remaining output plugins:

- influxunifi: batchFirewallPolicy() writes measurement "firewall_policy"
  with tags (rule_name, action, protocol, ip_version, source/dest zone,
  site_name, source) and fields (enabled, index, predefined, logging)
- datadogunifi: batchFirewallPolicy() emits the same data as Datadog gauges
  under the "firewall_policy.*" namespace
- otelunifi: exportFirewallPolicies() emits per-rule gauges
  (unifi_firewall_rule_enabled, unifi_firewall_rule_index) and per-site
  aggregates (rules_total, rules_enabled, rules_disabled, rules_by_action,
  rules_predefined, rules_custom, rules_logging_enabled)

Also rebases onto master to pick up the otelunifi plugin (PR #978).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 18:26:27 -05:00
Cody LeeandClaude Sonnet 4.6 521c2f88bc feat(otelunifi): add OpenTelemetry output plugin (#978)
* feat(otelunifi): add OpenTelemetry output plugin

Adds a new push-based output plugin that exports UniFi metrics to any
OTLP-compatible backend (Grafana Alloy/Mimir, Honeycomb, Datadog via
OTel, New Relic, etc.) using the Go OpenTelemetry SDK v1.42.

Config (default disabled):
  [otel]
  url      = "http://localhost:4318"
  protocol = "http"   # or "grpc"
  interval = "30s"
  timeout  = "10s"
  disable  = false
  api_key  = ""       # optional Bearer auth

Env var prefix: UP_OTEL_*

Exported metrics:
- Sites:   user/guest/IoT counts, AP/GW/SW counts, latency, uptime,
           tx/rx rates per subsystem
- Clients: uptime, rx/tx bytes & rates; signal/noise/RSSI for wireless
- UAP:     up, uptime, CPU/mem, load, per-radio channel/power,
           per-VAP station count/satisfaction/bytes
- USW:     up, uptime, CPU/mem, load, aggregate rx/tx bytes,
           per-port up/speed/bytes/packets/errors/dropped/PoE
- USG:     up, uptime, CPU/mem, load, per-WAN rx/tx bytes/packets/errors
- UDM/UXG: up, uptime, CPU/mem, load averages

Closes #933

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(otelunifi): rename unused ctx parameter to _ in recordGauge

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(otelunifi): replace Disable with Enable (default false)

Plugin is opt-in: set enable=true / UP_OTEL_ENABLE=true to activate.
Closes part of #933.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 18:19:18 -05:00
Cody LeeandClaude Sonnet 4.6 4c34180047 feat(clients): add MIMO spatial stream metrics for WiFi clients (#977)
* feat(clients): add MIMO spatial stream metrics for WiFi clients

Add tx_nss, rx_nss (spatial stream count) and tx_mcs, rx_mcs (MCS
index) metrics for WiFi clients, sourced from UniFi controller API
fields. These fields are only populated for wireless clients.

- promunifi: adds unifi_client_radio_transmit_spatial_streams,
  unifi_client_radio_receive_spatial_streams,
  unifi_client_radio_transmit_mcs_index, and
  unifi_client_radio_receive_mcs_index gauges
- influxunifi: adds tx_nss, rx_nss, tx_mcs, rx_mcs fields to the
  clients measurement
- go.mod: replace directive to use local unifi library with new fields

Closes #535

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix: use published unifi commit for MIMO fields instead of local replace

Remove the local path replace directive for github.com/unpoller/unifi/v5
and pin to the published pseudo-version at commit f363f61cdbe3a863db5fb3176ef1c0fc282c5674
which contains the RxMcs, RxNSS, TxMcs, TxNSS MIMO fields.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 17:56:16 -05:00
Cody LeeandClaude Sonnet 4.6 cedc52fc89 feat(lokiunifi): add richer low-cardinality stream labels (#932) (#975)
- Add job=unpoller to every Loki stream (alarm, anomaly, event, ids,
  system_log, protect_log, protect_thumbnail) for standard Grafana/Loki
  source filtering with {job="unpoller"}
- Add event_type and inner_alert_action labels to IDS streams using
  EventType and InnerAlertAction fields
- Add event_type and inner_alert_action labels to Alarm streams using
  Key and InnerAlertAction fields
- Skip severity/category on Anomaly: the unifi.Anomaly struct has no
  such fields

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 15:41:20 -05:00
Cody LeeandClaude Sonnet 4.6 117392dd8c feat: export site_to_site_enabled VPN metric (#926) (#976)
Add the site_to_site_enabled FlexBool field from the vpn subsystem
health entry to both InfluxDB and Prometheus outputs. The field was
present in the unifi.Health struct but never exported.

- influxunifi: add site_to_site_enabled to subsystems fields map
- promunifi: add SiteToSiteEnabled gauge descriptor and emit it in
  the vpn case of exportSite
- Update integration_test_expectations.yaml to include the new field

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 15:39:19 -05:00
Cody LeeandClaude Sonnet 4.6 a95804743d feat(lokiunifi): add extra_labels config for custom Loki stream labels (#691) (#973)
Add an ExtraLabels map[string]string field to the Loki Config struct so
users can define static key=value labels that are merged into the stream
labels of every log line sent to Loki. This allows users to distinguish
streams (e.g., by environment or datacenter) without hardcoding values.

Built-in dynamic labels (application, site_name, source, etc.) always
take precedence over extra labels to preserve existing behavior.

Example config (TOML):
  [loki.extra_labels]
  environment = "production"
  datacenter  = "us-east-1"

Closes #691

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 15:25:32 -05:00
Cody LeeandClaude Sonnet 4.6 6c5ff5482d feat(promunifi): add unifi_controller_up gauge metric (closes #356) (#974)
Add a per-controller `<namespace>_controller_up` Prometheus GaugeVec with
a `source` label (controller URL or configured ID). The gauge is set to 1
after each successful poll and 0 on failure, giving operators a standard
metric to alert on controller connectivity issues.

Changes:
- pkg/poller/config.go: add ControllerStatus type and ControllerStatuses
  field to Metrics so any output plugin can consume per-controller health.
- pkg/poller/inputs.go: merge ControllerStatuses when AppendMetrics is
  called (multiple input sources).
- pkg/inputunifi/interface.go: populate ControllerStatuses with Up=true
  on success and Up=false (while still continuing) on per-controller error.
- pkg/promunifi/collector.go: declare and register a prometheus.GaugeVec
  `<namespace>_controller_up`; set the gauge for each controller status
  after every Collect cycle.

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 15:25:00 -05:00