default_site_name_override is applied in augmentMetrics, which only the
metrics path goes through. Log entries leave by collectControllerEvents,
which reads its sites straight from getFilteredSites — where the override
is deliberately not applied — so events, alarms, IDS records and system-log
entries ship with the controller's stock site name.
Four of the five site-scoped log collectors were affected; collectAnomalies
already did this inline, which is what makes the omission visible.
The consequence is worst for a poller watching several UniFi OS consoles:
each one calls its only site "default", so their log entries are
indistinguishable downstream. In Loki every stream from those consoles lands
under site_name="Default (default)" no matter which console it came from,
and no relabeling downstream can separate them again — the information is
gone by then. That is precisely the case the option was added for, and it
works for the metrics from those same consoles.
Extract the check collectAnomalies was doing into overrideSiteName and call
it from all five, so the two paths agree.
Note that only Site.Name reaches an API path; Site.SiteName is a display
name throughout the library. The override is therefore safe here, which is
what applySiteNameOverride's own comment already says ("keeping default for
API calls").
unpoller_wan_interface_state ships with site_name and nothing else to
identify where it came from. That is not enough: every UniFi OS console
names its only site "default", so site_name is "Default (default)" on all
of them. An instance polling several controllers emits WAN status that is
indistinguishable, and downstream attribution has to be guessed.
Observed in production before the fix: a scrape config guessing from
site_name filed a UDM's WAN state under the wrong customer. No error, no
missing series — just wrong data under someone else's name.
unifi/v6.1.0 (#244) added SourceName to WANStatus, stamped from the
controller URL in all three read paths. go.mod is already on v6.1.0, so
this only has to read the field.
Two lines, same shape as #1071 which did this for unpoller_wan_*.
Tests use the fakeReport already present in the package: every emitted
metric carries both site_name and source, the raw state is kept as a
label, and a nil status still produces nothing rather than panicking a
poll. TestExportWANStatusIsAttributed fails on master and passes here.
Remote discovery stored Integration display names, so extra sites were dropped when checkSites compared them to legacy Site.Name. Use unifi v6.1.0 InternalReference instead.
Co-authored-by: Cursor <cursoragent@cursor.com>
Found by running this against a real UNVR4 rather than only the fake one.
setDefaults fills an unset user with the placeholder "unifipoller". For a
Protect-only console that placeholder was reaching Login(), the console answered
403, and the controller died during initialisation -- the same symptom #1066 set
out to fix, one layer further in. A Protect API key and no local account is the
config an operator actually writes for a UNVR, so this was the common case, not
an edge one.
Nothing on such a console uses a session unless Protect logs are wanted: the
Integration API authenticates with the key alone. So withhold the credentials
entirely in that case, and say so in the startup summary rather than naming a
username that is never sent.
Also pins the go.mod bump to unifi v6.0.1, which is the release that carries
NewProtectClient (unpoller/unifi#240).
Verified end to end against a UNVR4 (UniFi OS 5.1.31, Protect 7.2.105): the
controller comes up, logs "Auth: Protect API key only (no session needed)",
makes no login request at all, and exports 37 unpoller_protect_* series across
9 cameras and 2 bridges with unpoller_controller_up = 1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVreutpEATmBjm6PBw9RjQ
A UNVR or UNVR Pro runs UniFi Protect with no Network application installed.
UnPoller could not poll one at all: NewUnifi ends with GetServerData(), a GET of
/proxy/network/status, which such a console answers with its UniFi OS SPA HTML.
The controller entry died during initialisation and re-failed every interval,
never even printing a config summary -- while the Protect Integration API on the
same host answered every endpoint with the same key.
Set disable_network = true on that controller. It defaults to false, so nothing
about an existing config changes.
The Protect collectors were already complete and already not site-scoped; three
things stood between them and a Protect-only console:
- getUnifi now calls unifi.NewProtectClient, which skips the Network probe and
validates the Protect Integration API instead (unpoller/unifi#240).
- pollController aborted on getFilteredSites long before reaching
collectProtect, and collectControllerEvents did the same before
collectProtectLogs. The Network pass is extracted into pollNetwork and
skipped wholesale; the event collector list reduces to collectProtectLogs,
the only site-independent one.
- Metrics counted a poll successful only if it produced devices or clients. A
Protect-only console produces neither, so a filtered scrape of one -- the
Prometheus per-target path -- fell through to the dynamic-controller branch
and reported ErrDynamicLookupsDisabled despite a successful collection.
ProtectDevices now counts too.
Two smaller things worth calling out for reviewers:
- extractDevices dereferenced metrics.Devices unguarded. That was already a
latent panic; skipping the Network pass makes it reachable, so it is fixed
here rather than left for the first person to hit it.
- RawMetrics answers the raw-path kind for these consoles and rejects the
site-scoped kinds with ErrNetworkDisabled. Returning an empty result would
read as "this console has no devices" rather than "wrong question".
warnProtectOnly logs an error, without failing the controller, for the two
configurations that can never collect anything: disable_network with neither
Protect save flag, and save_protect_devices with no key to authenticate with.
Silently collecting nothing is the failure mode hardest to spot in a log.
pkg/inputunifi had no tests before this. input_test.go follows inputunas'
input_test.go: an httptest fake UNVR serving the console's SPA HTML for
everything but the Protect paths and the login, covering initialisation,
metrics, events, the filtered scrape, RawMetrics, both warnings, config binding
across toml/json/yaml/env, and that the shipped examples leave Network enabled.
TestProtectOnlyControllerFailsWithoutFlag pins the original bug against that
same console, so the flag is demonstrably what makes the difference.
Requires github.com/unpoller/unifi/v6 with NewProtectClient (unpoller/unifi#240).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVreutpEATmBjm6PBw9RjQ
Every unpoller_wan_* series shipped with site_name="" and source="". The
exporter said so itself:
cfg.WANLoadBalanceType,
"", // site_name - will be set by caller if available
"", // source - will be set by caller if available
The caller had nothing to set them from: WANEnrichedConfiguration carried
no identity. unifi/v6.0.3 fixes that upstream — GetWANEnrichedConfiguration
now stamps SiteName and SourceName from the site it fetched, the same way
GetSiteDPI does.
This bumps to v6.0.3 and fills the labels in. Two slices needed it, not
one: the base label set and the provider label set built further down for
the isp_name/isp_city descriptors. The test caught the second, which I had
missed.
Why it matters: an instance polling several controllers emitted WAN
metrics that were indistinguishable from one another, since wan_id is the
only other distinguishing label. Attributing them downstream meant
hardcoding a mapping in the scrape config and hoping no second controller
ever gained a gateway — when one does, its metrics are silently filed
under the wrong customer. No error, no missing series, just wrong data.
Tests use the fakeReport already present in the package. They assert every
emitted metric carries both labels, and that a nil configuration still
produces nothing rather than panicking a poll.
TestExportWANIsAttributed fails on master and passes with this change.
Release tags now fail closed without signing secrets, ship a signed Windows PE via house signerd, and replace per-arch Darwin tarballs with one stapled universal archive.
Co-authored-by: Cursor <cursoragent@cursor.com>
Remote API discovery was unconditionally overwriting default_site_name_override
with the console name for Cloud Gateways. Only apply the console name fallback
when the user has not already configured an override.
Fixes#1057
Co-authored-by: Cursor <cursoragent@cursor.com>
Add v3 integration and version tests, migration notes, InfluxDB 3 docker-compose stack, and README updates to finish the remaining plan phases.
Co-authored-by: Cursor <cursoragent@cursor.com>
Introduce explicit version selection with the influxdb3-go client alongside existing v1 and v2 paths, and resolve overlapping tag/field keys required for InfluxDB 3 write validation.
Co-authored-by: Cursor <cursoragent@cursor.com>
GoReleaser's make man hook failed because go get no longer works
outside a module. Replace deprecated go get with go install for
md2roff and rsrc.
Co-authored-by: Cursor <cursoragent@cursor.com>
golangci-lint v2.9 is built with Go 1.26 and panics when type-checking
dependencies that include go1.27-only source files.
Co-authored-by: Cursor <cursoragent@cursor.com>
Collects Protect device data (sensors, cameras, lights, bridges, link
stations, NVR) via the official Integration API and exports it through
Prometheus, InfluxDB, and DataDog. Opt-in via save_protect_devices,
gated by protect_api_key.
Bumps github.com/unpoller/unifi to v6, which added the Protect API
client (breaking change: FlexInt/FlexBool/FlexFloat replace nullable
pointers).
`disable = false` is a double negative, and a bool named disable cannot
express opt-in anyway: it zero-values to false, so the flag was inert and
opt-in rested entirely on the device list being empty.
`enable` defaults to false and is now the real gate -- Initialize, Metrics
and DebugInput all return early unless it is set. The two existing guards
remain: an empty device list is still a no-op, and no default URL is ever
synthesized.
Configuring devices while enable is false is always a mistake, so that
combination logs one error instead of silently collecting nothing.
Adds binding tests for the flag across toml, json, yaml and UP_UNAS_ENABLE
(the env name derives from the xml tag, not the json one), plus a test that
all three shipped examples default to off. Both were verified by mutation:
breaking a struct tag or flipping an example fails the suite.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
AppendMetrics merges every slice field on Metrics except SpeedTests, so
speed test results were discarded on the way from the input to the
outputs. Both call sites start from a non-nil &Metrics{}, so this hit
every user on every poll: inputunifi collected the results, and the
export code in promunifi, influxunifi and datadogunifi was dead.
TS was dropped the same way. The aggregate starts bare and nothing
restored the timestamp, so it stayed zero -- and influxunifi's collect()
stamps any point that carries no timestamp of its own with the
aggregate's, which meant the zero time. Both Influx clients omit a zero
timestamp and let the server assign one, so the damage was limited to
points being stamped on arrival rather than at poll time, but it made
the fallback path meaningless. First writer wins: the earliest input's
timestamp is the one that describes the batch.
The failure mode here is what makes it worth guarding rather than just
patching. A new metric family needs a field on Metrics and an append
line in AppendMetrics, and omitting the second loses every metric in
that family with no error, no log line, and a passing build.
TestAppendMetricsCoversEverySliceField walks the struct by reflection
and fails naming any slice field that is not merged, so a new field is
covered the moment it is declared rather than when someone notices the
graphs are empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds a new `unas` input plugin that polls UNAS Pro storage consoles and
exports console health, storage pools, disks and shares to Prometheus,
InfluxDB and DataDog.
UNAS is a separate plugin rather than a device type inside inputunifi
because a storage-only console has no Network application: it cannot
answer /status, has no sites, and shares none of the UniFi device schema.
The plugin is opt-in and inert until an operator names a console. Opt-in
is expressed as "no devices configured" rather than a `disable` flag,
because a bool named `disable` zero-values to false and so cannot make a
plugin default-off. Initialize returns silently on an empty device list
and, unlike inputunifi, nothing synthesizes a default URL.
Two behaviours are worth calling out for reviewers:
- Metrics returns (metrics, nil) whenever any console was collected.
poller.collectMetrics uses `if err != nil {} else if metric != nil`,
so returning both would discard every healthy console because one
failed. Only a total failure returns an error.
- Re-auth fires on total failure, not on a 401. A mid-session 401 from
GetData surfaces as ErrInvalidStatusCode, not ErrAuthenticationFailed,
so there is no sentinel to match on. Session expiry fails all four
endpoints at once, which is exactly the total-failure case.
Prometheus metrics use the `unifi_unas_` prefix, which diverges from the
`unas_` prefix used by the reference implementation; dashboards built
against that will need query edits.
Requires unifi/v5 v5.31.0 for the UNAS client and structs.
Credit to alexgreenbank/unaspoller for mapping the endpoints.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Some Network 10.x+ controllers (UniFi OS, Network 10.5.67 confirmed)
return HTTP 400 api.err.InvalidObject from list/alarm instead of a 404
when the endpoint is gone, so collectAlarms fell through the existing
ErrEndpointNotFound skip and logged a real ERROR on every poll.
Fixes#1050
BACKUP and DISCONNECTED WAN interfaces both reported as 0, making
them indistinguishable in Prometheus (fixes#1045). InfluxDB and
DataDog already expose the raw state; this brings Prometheus in
line by adding a state label alongside the existing 1/0 gauge.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
InputUnifi.Events returns (nil, nil) when disabled, but collectEvents
dereferenced e.Logs unconditionally after a successful (err == nil)
call, crashing the poller. Same crash class as #1030, found while
verifying the recover-based fix.