The printer card's AI badge collapsed every class that was not Warning or
Failure into green Safe, and the service reported `safe` whenever it had no
verdict. The state entry is created when a monitored print is first seen --
before the first snapshot, let alone the first inference -- so a rejected ML
API token, an unreachable ML API, a failed capture and an unset External URL
all rendered as a healthy watched print: green Safe at score 0.000.
For a safety feature that is the worst failure mode available: it asserts the
print is being watched exactly when it is not. The reporter read that badge and
concluded the loop had never started. It had been calling the ML API every ten
seconds and being turned away with a 401 -- invisible because Obico's auth
layer rejects a bad token before its request log sees it, and because
successful checks log nothing there either.
Add two honest states. Not checking (amber) when the last poll produced no
result, carrying the reason; Starting while a monitored print waits for its
first result. Score and frame count are withheld while not checking, since
0.000 beside "Not checking" reads as a measurement rather than its absence.
The reason is per printer, so a card names its own problem rather than
whichever printer failed most recently, and stays behind settings:read because
it can quote configured URLs -- the badge state does not, because whether a
print is watched is not configuration. An unrecognised class now falls back to
Starting, not Safe.
Test Connection saves the form before probing, so a green result describes the
configuration the loop actually runs with rather than what is typed in the
boxes.
Obico's ml_api container takes an optional ML_API_TOKEN environment variable.
With it set, ml_api/auth.py answers a bare 401 to any request whose
Authorization header isn't "Bearer <token>"; with it unset it ignores the
header entirely. Bambuddy never sent one, so pointing it at a protected server
meant deleting the token there — which the reporter had set for their Home
Assistant integration and did not want to undo.
Settings -> Failure Detection gains an ML API Token field. When it is empty no
header is sent, so an unconfigured install's request stays byte-identical to
what shipped before the setting existed.
This failed in the worst possible way, and that is the more important half of
the change. Obico decorates /p/ with token_required but leaves /hc/ open. Test
Connection pinged /hc/, so it reported success against a server that was
rejecting every real detection call, the settings looked right, and detection
silently never ran. The only symptom was a generic "ML API call failed" buried
in the status card.
So the test now proves what it claims. After health passes it probes GET /p/
with no img parameter: the auth decorator runs before the handler, so 401 means
the token was rejected and 422 ("Invalid request params") means it was
accepted. No inference work is done either way. A probe that itself errors
reports the token as unknown rather than as working — the UI says it could not
be checked instead of claiming success.
The detection loop checks for 401 before raise_for_status, so a rejected token
is reported as a rejected token, naming the setting and the environment
variable, instead of surfacing "401 Unauthorized" with no hint of what to do.
The message never contains the token; a test pins that.
The setting name carries "token", so the support bundle's keyword redactor
masks it with no new rule. Resolving "field omitted" to the saved token is the
route's job, keeping test_connection a pure outbound call with no database
access.
Second fix, same issue: support bundles misreported which printers Obico
watches. The bundle split obico_enabled_printers on commas and read an empty
value as "no printers". The settings UI writes a JSON array, and empty means
*all* printers — the default — so a working Obico setup showed obico_enabled
false against every printer in its own bundle. That is the reporter's bundle
exactly, and it points anyone reading it at the wrong subsystem. The bundle now
parses the setting the way ObicoDetectionService does, keeps a comma fallback
for any install that stored the legacy shape, and factors in the global switch.
The Status panel's Low / High thresholds readout was stuck at 0.38 /
0.78 regardless of the Sensitivity dropdown, so the setting looked
dead. Detection itself was correct -- classify() always used the real
sensitivity -- but get_status() computed the displayed thresholds with
a hardcoded thresholds("medium").
get_status() now takes an optional sensitivity argument and the
/obico/status route passes settings["sensitivity"] (it already loads
settings fresh). The readout updates as soon as the change is saved.
Obico polling could freeze the live camera stream within seconds of opening
the viewer. Cause: when the buffer-reuse path in obico_detection._capture_frame
saw an empty _last_frames[printer_id] entry (stream startup before the first
JPEG lands, or upstream mid-reconnect after a 30s read timeout), it fell
through to capture_camera_frame_bytes() and opened a second RTSP socket. On
firmwares that allow only one camera connection, that second socket forced
the printer to drop the live fan-out connection - the viewer's ffmpeg then
hit its own 30s timeout, looped through 30 reconnects at 0.2s, all racing
the next Obico poll, and the broadcaster pump exited.
Widen the gate from "do we have a buffered frame?" to "is any fan-out stream
registered for this printer?". New is_stream_active() helper checks
_active_streams / _active_chamber_streams independently of buffer state.
_capture_frame consults it first: if a viewer is attached, it returns the
buffered frame when available or None (skip this poll cycle) when not. Never
opens a competing socket while a viewer is connected.
Cost: at most one missed Obico detection cycle per viewer-attach (~10s lag).
Benefit: zero competing-socket events while any viewer is connected.
try_get_active_buffered_frame() refactored to delegate to is_stream_active()
so the two helpers stay in lockstep. The /camera/snapshot caller is unchanged
behaviorally (snapshot is a user-initiated single-shot; falling through to
fresh capture on an empty buffer is the desired behavior there).
The MJPEG fan-out broadcaster from #1089 only solved viewer-side
concurrency. Obico polling (every 5s) and the manual /camera/snapshot
endpoint kept opening their own fresh RTSP sockets, which X1/H2/P2
firmwares tolerated but X2D firmware 01.01.00.00 enforces strict
single-connection on — every poll kicked the live stream.
Add try_get_active_buffered_frame(printer_id): returns the broadcaster's
last buffered frame when a viewer is connected, None otherwise. Obico
and /camera/snapshot consult it before opening a fresh socket. When no
viewer is active they fall through to the existing fresh-capture path.
plate_detection and layer_timelapse intentionally not converted.
The 0.2.3b4 #1003 "fix" POSTed JPEG bytes as multipart form data,
but Obico's /p/ endpoint is declared methods=['GET'] upstream and
reads ?img=URL from the query string. Every POST was 405'd by
Flask's router before any handler ran, which is why the Obico
container logs were silent while Bambuddy kept reporting
"ML API call failed for printer N:" with a blank suffix —
raise_for_status() on the 405 produced an exception whose str()
rendered empty.
Restored the pre-#1003 nonce-URL approach (commit 3e434458):
capture locally with a 20s timeout we control, stash the JPEG
under a single-use 32-byte nonce, hand Obico a
GET /api/v1/obico/cached-frame/{nonce} URL that resolves in
<50ms so its hardcoded 5s read timeout never races RTSP.
Also guards against future silent exceptions: the error format
now falls back to type(exc).__name__ when str(exc) is empty.
Detection also early-returns with an explicit error if
external_url is unset instead of handing Obico a URL it can't
resolve.
The #1003 reverse-proxy scenario (Authelia/Authentik/CF Access
in front of Bambuddy) is addressed by documenting that the
/api/v1/obico/cached-frame/ path must be whitelisted from
external auth at the proxy layer — it is already public on
Bambuddy's side.
Backend: services/obico_detection.py, api/routes/obico.py,
main.py (PUBLIC_API_PATTERNS).
Frontend: FailureDetectionSettings banner + client.ts type +
all 7 locales restored.
Tests: 15 unit + 5 integration tests pass.
The ML API previously called back into Bambuddy to fetch snapshots,
which failed behind reverse proxies with external auth (Authelia, etc.).
Now the detection loop captures the JPEG locally and POSTs it directly
as multipart form data — no callback URL, no nonce cache, no
external_url dependency.
Obico's ML API has a hardcoded 5s read timeout on the URL it fetches, which
our /camera/snapshot regularly exceeds on cold calls (TLS proxy + ffmpeg +
RTSP keyframe wait). The detection loop now captures the JPEG locally with
a 20s timeout we control, stashes the bytes under a single-use 32-byte
nonce, and hands Obico a new /api/v1/obico/cached-frame/{nonce} URL that
returns the cached bytes instantly. The 5s ceiling is no longer a factor.
The nonce is the credential (URL-safe, 256 bits of entropy, single-use,
30s TTL) so the endpoint can be unauthenticated without widening the
camera access surface. Replaces the previous camera-stream-token snapshot
URL approach, which remained vulnerable to the upstream 5s timeout even
when auth was disabled.
Thanks to @fblix for the detailed reproducer with timeout numbers.
surfaced it as "Failed to get image", which Bambuddy reported back as a 400.
Fix: the snapshot endpoint already accepts a reusable camera-stream token (the same mechanism used by <img>-based camera consumers, since browsers can't send auth headers
on image loads). The detection service now appends that token to the URL it gives the ML API. The token is cached on the service, refreshed 5 min before its 60-min expiry,
and is simply ignored when Bambuddy auth is disabled — so no behavior change for users without auth.
Refs #172
Adds a Failure Detection tab under Settings that wires Bambuddy to a
self-hosted Obico ml_api container — no cloud, no account, no WebSocket.
While a print is running, the detection service periodically hands the
printer's camera snapshot URL to the ML API and smooths scores over
time (30-frame warmup + EWM, alpha=2/13, short/long rolling means) so
one noisy frame can't trigger an action. When the smoothed score
crosses HIGH, the configured action fires exactly once per print:
notify, pause, or pause-and-cut-power (via linked smart plugs).
- Backend: new obico_detection + obico_smoothing + obico_actions
services, /obico/status and /obico/test-connection routes
(SETTINGS_READ / SETTINGS_UPDATE), six obico_* AppSettings fields
with validators for sensitivity/action/enabled_printers.
- Frontend: FailureDetectionSettings component (enable, ML URL + test,
sensitivity, action, poll interval, per-printer monitor list, live
status + detection history), new sidebar tab with service-active
bullet, toast on save.
- Tests: 17 detection unit tests + 15 smoothing unit tests + 4
frontend component tests.
- Docs: README bullet, CHANGELOG entry, wiki page under Analytics,
website features.html entry.