bambuddy.log on Windows fills with
Exception in callback _ProactorBasePipeTransport._call_connection_lost()
ConnectionResetError: [WinError 10054] An existing connection was
forcibly closed by the remote host
every time a printer / MQTT broker / camera RSTs a TCP socket instead
of FINing it. The application-layer reconnect (paho-mqtt, httpx)
handles the actual disconnect fine; the traceback is asyncio
bookkeeping. Reported by @cadtoolbox who runs 9 printers including 5
offline X1Es, so the log filled multiple times per minute.
New backend/app/core/asyncio_handlers.py installs a custom
loop.set_exception_handler on Windows that pattern-matches three
signals together (platform == win32, exception is
ConnectionResetError, asyncio message contains
_call_connection_lost) and demotes the entry to DEBUG. Genuine
ConnectionResetErrors raised inside application coroutines have a
different message string and still surface; BrokenPipeError /
ConnectionAbortedError on the same cleanup path also still surface.
Wired from lifespan startup before any task can spawn that might
trip it. Linux / macOS use the Selector loop, so install is an
explicit no-op there with a False return.
9 unit tests in test_asyncio_handlers.py covering signature match,
rejection of unrelated resets, platform gate, suppress vs.
pass-through to default handler.
Auto-Print G-code Injection had two reviewer-reported bugs from the initial
ship:
1. Start snippets were prepended to the entire plate_X.gcode, landing
before the printer's own bed-heat / homing / nozzle-prime sequence —
so a Swapmod start snippet that assumed nozzle-at-temp ran on a cold
printer (pleite). Anchor injection at "; MACHINE_START_GCODE_END" so
snippets land where a slicer-side custom-start-gcode would. Files
without the marker keep prepend behaviour as a fallback with a
warning log.
2. Placeholders like "G1 Z{max_layer_z} F600" were written verbatim;
firmware parsed them as Z1 and crashed the head into the print on
tall models — real safety bug (DevScarabyte). Added a header parser
for the 3MF "; HEADER_BLOCK_START..END" block (lowercased keys,
[units] suffix stripped, spaces -> underscores) and a Prusa-style
{name} substitution pass over both start and end snippets before
injection. Supported placeholders: {max_layer_z} / {max_print_height},
{total_layer_number} / {total_layers}, {total_filament_weight},
{total_filament_length}, plus any other normalised header key.
Unknown placeholders are left verbatim with a warning — a typo never
silently expands to an empty string.
16 new regression tests across 4 new classes in test_gcode_injection.py
(anchored injection + missing-marker fallback, placeholder substitution
including alias resolution + unknown-pass-through, direct unit tests
for each new helper). All 2195 backend unit tests pass.
Wiki print-queue page updated with the supported placeholder list and a
{max_layer_z} safety callout for park moves.
Moderate-severity advisory: PostCSS < 8.5.10 has an XSS via an
unescaped </style> sequence in its CSS Stringify output. Caret range
in package.json already accepts 8.5.12, so this is a lockfile-only
bump (npm audit fix). Build verified clean.
Vite, autoprefixer, and @tailwindcss/postcss all dedupe onto the same
8.5.12 — no nested copies left in node_modules.
Note: Bambuddy doesn't pass user-controlled CSS through PostCSS at
runtime (PostCSS is build-time-only), so the practical impact even on
older versions was nil. This is hygiene + clearing the npm audit
warning.
Adds a finger-friendly amber pill row under the printer status badges
on the SpoolBuddy kiosk dashboard. When any printer reports
awaiting_plate_clear=true, a compact pill appears showing the printer
name plus a "Clear" action; tapping it calls POST /printers/{id}/clear-plate
and optimistically removes the pill before the WebSocket round-trip
lands. Multiple pending printers wrap inline via flex-wrap so the
dashboard stays compact when several finish at once. Pill dimensions
match the existing online/offline printer badges (px-2.5 py-1, text-xs).
The kiosk auth path (X-API-Key) already passes the printers:clear_plate
permission gate via the existing _APIKEY_DENIED_PERMISSIONS denylist
(the permission is intentionally not denied — clear-plate is an
inventory-flow operation, not an admin one), so no auth wiring changes
were needed.
Two adjacent fixes shipped in the same change:
- SpoolBuddyLayout suppresses the global toast viewport while mounted.
The global ToastProvider in App.tsx wraps both the main app and the
kiosk routes, which meant the background-dispatch progress overlay
was rendering on the kiosk display alongside any in-flight prints.
Added setViewportSuppressed(bool) on the toast context; the layout
flips it via useEffect and restores on unmount. State machine and
dispatch-event subscription are untouched — only the visible
viewport is hidden.
- Dispatch toast no longer reads as "frozen at 100%" for fast uploads.
Small files complete FTP in <500ms but the bar would sit at 100%
while the printer's MQTT confirmation landed. When uploadProgressPct
>= 99.9 and status is still 'processing', the byte counter is
replaced with "Awaiting printer..." and the bar gets animate-pulse.
CameraPage hard-coded fps=15 in the stream URL and never read the
URL query string, so /camera/1?fps=5 and similar diagnostic URLs
were silent no-ops. Sibling StreamOverlayPage already honored ?fps=;
this brings CameraPage to parity.
- Read fps via useSearchParams, default 15, clamp 1-30, fallback to
15 on non-numeric input (mirrors StreamOverlayPage's parser)
- Thread the parsed value into the stream URL builder
- 5 new tests in CameraPage.test.tsx pinning default, honored value,
clamp-above-30, clamp-below-1, and non-numeric fallback
Surfaced while triaging #1131 (H2D camera freeze). Independent of
the underlying freeze investigation; restores the diagnostic knob
the issue thread was relying on.
Reprinting from archives sometimes failed immediately with a MicroSD R/W
exception, with the printer's MQTT push referencing a 3MF from a
different unrelated archive. Once it started, every subsequent reprint
hit the same error until the container was restarted.
Root cause from @smandon's support package: paho-mqtt's client-side QoS
1 queue. When the printer's command channel goes half-broken (telemetry
flowing, publishes silently dropped — same #887/#936 pattern),
background_dispatch.py:993 hits its 15s deadline and calls
force_reconnect_stale_session(). That function was force-closing the
underlying socket so paho's auto-reconnect would kick in, but the same
mqtt.Client instance, same client_id, and same in-process QoS 1 queue
stayed alive across the reconnect. Any unacked publish from the broken
session — typically the just-sent project_file for the new archive —
got replayed verbatim on the new connection. The queue accumulates
across multiple stuck dispatches in one Python process, so by the
second or third stuck reprint there were several stale
project_file/resume/stop/clean_print_error commands queued together;
the printer latched onto whichever stale path it processed last,
couldn't find the file on its SD card, and emitted 0500_4003. Container
restart was the only thing that wiped paho's in-process queue.
Replaced socket-close with a context-aware reconnect via a new
_reset_client_for_reconnect() router:
Async-context callers (dispatch deadline, FastAPI handlers via
check_staleness) → hard-reset: client.disconnect() (broker drops
session, clean_session=True), client.loop_stop() (kills paho's
network thread and its queue), null _client, fresh connect() with
incremented client_id. New connection is genuinely empty, no replay.
Paho-network-thread callers (dev-mode probe + ams_filament_setting
zombie detection inside _update_state) → socket-close fallback.
loop_stop() from inside the network thread would self-join and
deadlock, so the safe pattern there is "close the socket and let
paho's loop detect it and auto-reconnect on the same client".
Routing decision uses asyncio.get_running_loop() — paho's callback
thread has no loop, every legitimate hard-reset caller does.
7 regression tests:
- TestForceReconnectRouting (3): sync-context → socket-close fallback,
async-context → hard-reset with disconnect()+loop_stop()+null,
state-disconnected broadcast fires once on either path
- TestHardResetClientDirect (3): helper directly — old client gets
disconnect()+loop_stop(), _client cleared, failing disconnect()
doesn't propagate so background_dispatch's await chain can't break
- TestZombieSessionDetection / TestDeveloperModeProbeTimeout (updated):
paho-thread context still goes through socket-close, preserving the
legacy contract for those paths
Three bugs that surfaced together while debugging an H2D cancel:
1. Cancelling a print stamped failure_reason="Layer shift" in archives
AND left the printer card stuck on "1 problem" forever. Four causes:
(a) POST /printers/{id}/print/stop never set the user-stopped flag, so
on_print_complete couldn't override "failed" -> "cancelled".
(b) HMS-derived failure_reason heuristic mapped any module-0x0C HMS to
"Layer shift". Module 0x0C is "Motion Controller" broadly (includes
cameras, markers, AND the cancel-sequence echo 0C00_001B). Real
layer-shift codes live in module 0x03. Same false-positive class
existed for "Filament runout" (any 0x07) and "Clogged nozzle" (any
0x05). Replaced with a 23-code curated short-code map; unknowns
leave failure_reason=None.
(c) Cancel-echo HMS codes (0300_400C "The task was canceled.",
0500_400E "Printing was cancelled.") were polluting state.hms_errors
via both the hms[] and print_error parse paths. Filter them at
parse time so the frontend never sees them.
(d) Frontend bucketed gcode_state="FAILED" as a problem unconditionally.
Real failures attach an HMS error; user-cancels don't — so FAILED-
without-HMS now buckets as "finished" and only escalates to "error"
when there's an active known HMS.
2. logs/bambuddy.log was silently dropping records from named child
loggers. TraceIDFilter was attached to root_logger, but Python's
logging only invokes a Logger's filters on records originating at that
logger — propagated child-logger records skipped it, formatter raised
KeyError, handler.handleError dropped the record. Moved the filter
from root_logger.addFilter() to handler.addFilter() on each handler,
matching the filter's own docstring guidance.
derive_failure_reason() extracted as a pure function for testability.
status="cancelled" now symmetrically yields "User cancelled" alongside
"aborted".
20 regression tests across:
- backend/tests/unit/test_failure_reason_derivation.py (11)
- backend/tests/unit/services/test_bambu_mqtt.py::TestHMSUserActionFiltering (4)
- backend/tests/unit/test_trace.py::TestFilterMustBeAttachedToHandlerNotLogger (1)
- frontend/src/__tests__/pages/PrintersPageBucketing.test.ts (5; includes
the H2D-cancel-echo "FAILED + only unknown HMS" case)
The delete-API-key mutation called queryClient.invalidateQueries on the
['api-keys'] cache, which in react-query v5 should also refetch active
queries — but in practice the deleted row remained visible until the
user reloaded the page. Switched onSuccess to setQueryData so the
deleted key is filtered out of the cache synchronously the moment the
API confirms; no refetch round-trip required, no invalidation→refetch
race possible. Create-path is unchanged (invalidateQueries was working
there).
Pins the contract with a new SettingsPage test that mocks the GET/DELETE
endpoints, clicks delete, confirms, and asserts the row is gone without
any reload.
Reproduced live during the #1133 rollout: the SpoolBuddy display kept
serving the pre-fix picker for hours after every cache-clear,
chromium-restart, and pkill attempt because a chain of stale state
across HTTP cache + Service Worker + persistent profile prevented
fresh code from reaching the running tab.
Three independent changes — any one of them sufficient on a clean
profile, but all three needed to escape an already-corrupted one:
(1) backend/app/main.py — index.html now served with
Cache-Control: no-cache, must-revalidate on both / and the SPA
catch-all. Vite emits content-hashed JS/CSS bundle filenames so the
assets themselves are safe to cache forever, but the HTML wrapping
them is the only file that knows which hash is current. Without
explicit cache directives Chromium falls back to heuristic caching
(typically 10% of time since Last-Modified) and on long-running
kiosks happily serves stale HTML across browser restarts. That stale
HTML references an old bundle hash which is also still in disk
cache, so the kiosk runs pre-deploy JS forever without ever knowing
why.
(2) frontend/public/sw.js — CACHE_NAME bumped from bambuddy-v25 to
bambuddy-v26 so any client that fetches the new sw.js drops its old
CacheStorage. The SW does network-first for HTML/JS/CSS but
intercepts and falls back to cache, and cache-control on HTTP
responses doesn't reach into the SW's own cache layer.
(3) spoolbuddy/install/install.sh — generated kiosk launcher now uses
--user-data-dir=/tmp/spoolbuddy-kiosk-userdata with a pre-launch
rm -rf, so every kiosk restart starts from a clean slate (no HTTP
cache, no SW registration, no IndexedDB). Trade-off is a slightly
slower first paint and zero offline support; neither matters for a
single-purpose kiosk facing a backend on the same LAN, and the
guarantee that next-deploy-just-works is worth far more.
4 new tests in test_static_html_cache_headers.py: index.html on /
and SPA catch-all paths emit Cache-Control: no-cache,
must-revalidate; API routes are unaffected (no leak of HTML cache
directive onto endpoints we want React Query to cache aggressively).
For existing kiosks already trapped by an old persistent profile,
operator runs once: rm -rf ~/.config/chromium && systemctl restart
getty@tty1.service. The new launcher then picks up automatically.
The picker that opens from <FilamentHoverCard> / SpoolBuddy's slot-action
sheet had two stacked filters that together blocked a real workflow:
1. AssignSpoolModal only listed spools whose tag_uid AND tray_uuid were
both null, hiding any Bambu Lab spool that had been auto-created from
RFID or scanned via SpoolBuddy NFC.
2. FilamentHoverCard rendered its inventory section only when the slot's
vendor was not 'Bambu Lab', so even with #1 fixed the assign button
wasn't visible on a BL slot.
Both filters blocked the same use case: a user has a Bambu Lab spool in
inventory but doesn't want to scan via SpoolBuddy NFC every time and
just wants to pick it from the list.
Both gates removed. Modal lists every spool that isn't already taken by
another (printer/ams_id/tray_id) tuple. Hover-card inventory section
renders for every vendor including Bambu Lab. The AMS-vs-external
special-case in the modal collapsed too — external slots used to be the
only path that allowed picking a tagged spool, that distinction is gone.
Empty slots lost their assign affordance entirely. A physically empty
slot has no spool to attach an inventory record to, and offering the
action there only led to users assigning the wrong spool to a slot the
printer hadn't actually loaded yet. Bambuddy: EmptySlotHoverCard's
inventory prop removed; PrintersPage drops the matching inventory
props on three sites. SpoolBuddy: slot-action picker gates the
assign/unassign block on slotActionPicker.tray !== null.
Modal also gets defensive hardening from the rollout investigation:
its own dedicated cache key (['inventory-spools', 'assign-modal']) so
it can't be poisoned by other components calling getSpools() with
different includeArchived args; getSpools(true) + client-side
!archived_at filter so the picker sees the full inventory regardless
of cache priming order; "Show all spools" toggle now bypasses BOTH
filters (was only bypassing material/profile, label was a lie); a
small "X fetched · Y archived · Z assigned" counter in the empty state
so future "missing spool" reports are debuggable from screenshots.
i18n.inventory.noManualSpools renamed to inventory.noAvailableSpools
with new copy ("No spools available. Add a spool to your inventory or
unassign one from another slot first.") since the empty-state premise
changed. Localised across all 8 languages.
15 frontend tests: assign/unassign render for vendor: 'Bambu Lab',
non-BL vendors unchanged, EmptySlotHoverCard renders no assign
affordance, configure button still works on empty slots, picker lists
BL spools alongside manual ones, picker drops spools assigned
elsewhere unless toggle is on, picker drops archived spools always,
toggle escape-hatch shows everything, empty-state copy update.
SpoolBuddy unassign now invalidates both ['spool-assignments'] and
['spool-assignments', printerId] so the modal's cache stays fresh
(dual cache-key consolidation deferred to its own PR).
awaiting_plate_clear is a Bambuddy-side flag, not a printer-side one,
so toggling it does not produce an MQTT push from the printer. Commit
4e86e8c added the flag to the printer_status payload so MQTT-driven
broadcasts (e.g. when a print finishes and on_print_complete sets the
flag to True alongside a state transition to FINISH) carry it. The
reverse transition didn't: POST /printers/{id}/clear-plate mutated
PrinterManager._awaiting_plate_clear and persisted to the DB, but
emitted no printer_status WebSocket update — and the in-main.py
status-change broadcaster's status_key dedup intentionally excludes
Bambuddy-side flags, so even a coincidentally-arriving MQTT push
wouldn't reflect the change.
The "Mark plate as cleared" button on the printer card disappeared
"immediately" after a click only because the React Query cache was
being optimistically updated client-side; clearing the flag through
any other route (an admin script, a second tab, an automation hitting
the endpoint directly, the scheduler at print_scheduler.py:1844 when
dispatching the next queued print) silently left every UI subscriber
but the originating tab stale until a coincidental status refresh.
Centralised the broadcast in PrinterManager.set_awaiting_plate_clear
itself rather than at each call site, so every current AND future
caller is covered without remembering to wire it up: a new
_broadcast_status_change(printer_id) private coroutine is scheduled
alongside the existing _persist_awaiting_plate_clear whenever the flag
flips under a running event loop. Lazy-imports ws_manager to keep
printer_manager.py clean of application-layer infra at module-import
time, short-circuits when get_status returns None (printer
disconnected — the next reconnect produces a fresh push anyway), and
swallows ws_manager.send_printer_status failures so the persistence
path can complete even if the WS layer is temporarily unavailable.
The same hook is now in place for any other Bambuddy-side flag that
gets added to printer_state_to_dict later — they'll all need to
broadcast their own changes for the same reason.
8 new regression tests in test_printer_manager_status_broadcast.py:
schedules-on-True/False/loop-running/no-loop/loop-stopped contracts,
_broadcast_status_change happy path with payload assertion,
skip-when-no-state, swallow-WS-errors, and an end-to-end live-loop
test that fires set_awaiting_plate_clear(False) and asserts a
broadcast lands with awaiting_plate_clear: false in the payload.
Existing 24 tests in test_scheduler_clear_plate.py continue to pass
unchanged because they instantiate PrinterManager() without
attaching a loop (sync unit-test path) — the new _schedule_async
call short-circuits on the same loop check the existing persistence
call already used.
Builds on the recent uvicorn-access-log-into-bambuddy.log change.
Until now the access line told us who called an endpoint, but there
was no way to tie that line to the application records emitted on the
server side while handling that request. The rogue stop_print mystery
on 2026-04-26 left exactly that gap: even with access logs piped in,
correlating "this POST landed" with "this MQTT publish went out 6 ms
later" required eyeball-matching timestamps across different loggers.
A new ContextVar + middleware + logging filter wire a trace ID through
every record:
* trace_id_middleware mints an 8-char hex ID per request (or honours
a sane inbound X-Trace-Id for cross-system correlation), stores it
in trace_id_var (ContextVar), echoes it on the response as
X-Trace-Id, and resets the var in finally.
* TraceIDFilter, attached to root + uvicorn.access, copies the
current trace_id_var value onto every LogRecord so the format
string [%(trace_id)s] resolves to the right ID per record.
* Records emitted outside any request scope (startup, MQTT
callbacks, scheduler) get a stable "-" placeholder so the column
stays visually aligned and grep stays simple.
ContextVars are the right plumbing because asyncio copies the current
context into every asyncio.create_task, so background work spawned
from inside a request inherits the same ID without explicit threading.
request.state can't make that hop. The logging filter also has no
access to the FastAPI request object — it runs synchronously inside
the stdlib logging machinery — and the ContextVar is the only
mechanism that bridges async request scope to sync log emission.
Inbound X-Trace-Id is hard-validated against [A-Za-z0-9_-]+ (max 64
chars) before being honoured — a hostile/buggy caller cannot smuggle
log-injection payloads (newlines, control chars, megabyte blobs) into
bambuddy.log via the trace ID column; values that fail the gate
silently trigger a freshly minted server-side ID rather than failing
the request.
Middleware is decorated AFTER auth_middleware on purpose: Starlette
stacks @app.middleware decorators LIFO so the last-decorated runs
first inbound, making trace stamp the OUTERMOST layer — auth log
lines and every record emitted on the way down to and back from the
route handler all carry the same ID.
Output now correlates as:
2026-04-26 09:51:39,152 INFO [uvicorn.access] [a4f3b1e7] - "POST
/api/v1/printers/1/print/stop HTTP/1.1" 200
2026-04-26 09:51:39,158 INFO [bambu_mqtt] [a4f3b1e7] [SERIAL] Sent
stop print command
One grep a4f3b1e7 returns the full causality chain.
30 new tests: 22 unit (ContextVar placeholder, filter copies value,
asyncio task propagation, concurrent-request isolation, hex generator
uniqueness, hostile-payload validator, max-length boundary, all four
write verbs survive, GET/HEAD/OPTIONS dropped, URL-substring false-
match guards, edge cases) and 8 integration (X-Trace-Id round-trips,
body matches header, hostile inbound replaced, overlong inbound
replaced, ContextVar resets after request, generator format stable,
each request gets unique ID).
Bambu MQTT can deliver two ams_data push frames for the same printer
~30 ms apart (observed on H2D + dual AMS at K-profile-load / RFID-read
boundaries). Each frame triggers on_ams_change in main.py, whose
auto-assign block reads (printer_id, ams_id, tray_id), decides "no
existing assignment", and INSERTs via auto_assign_spool — and the two
callbacks raced in their respective sessions, both deciding to insert,
with the second commit losing on:
asyncpg.exceptions.UniqueViolationError: duplicate key value
violates unique constraint
"spool_assignment_printer_id_ams_id_tray_id_key"
DETAIL: Key (printer_id, ams_id, tray_id)=(1, 0, 0) already exists.
SQLite's WAL serial-write semantics had been silently swallowing the
race for ~7 weeks since the spool-assignment feature shipped (latent in
ec82092b "Sync", 2026-02-12). When optional Postgres support landed in
610431d6 (2026-04-03) and asyncpg started allowing true concurrent
transactions, it surfaced. Net impact: log noise + one assignment cycle
skipped, retried on the next on_ams_change.
Adds a per-printer asyncio.Lock (_ams_assignment_locks keyed by
printer_id) wrapping the auto-assign critical section. By the time the
second callback's session runs the SELECT, the first's commit is
visible and the early-return "existing assignment" branch fires instead
of a duplicate INSERT.
The Spoolman sync block further down in on_ams_change intentionally
stays OUTSIDE the lock — it's network-bound and idempotent, so
serialising it would block subsequent AMS callbacks for the duration of
a remote roundtrip. Per-printer scope keeps unrelated printers fully
parallel. The auto-unlink block above isn't wrapped because its
DELETE/UPDATE operations don't have the same constraint surface.
5 new regression tests in test_ams_assignment_lock.py: same-printer-
same-lock identity, different-printers-different-lock isolation, second
acquirer waits for first (proves serialisation), different printers run
truly in parallel under a held lock (proves per-printer scope), and an
autouse fixture that resets the module-level dict between tests so
cross-test loop affinity bugs can't surface.
The bidirectional forwarders inside create_tls_proxy._handle catch
(ConnectionError, OSError, asyncio.CancelledError) on writes, but
uvloop's UVStream.write raises a plain RuntimeError from
UVHandle._ensure_alive when the underlying handle is already closed.
asyncio's default selector loop reports the same situation as
ConnectionResetError, so the bug only surfaced on uvloop — and only at
the moment ffmpeg (or a snapshot-capture subprocess) dropped its socket
while the proxy was mid-flush.
The RuntimeError slipped past the except tuple, escaped the forwarder
coroutine, and asyncio's client_connected_cb task-exception handler
logged a noisy multi-line traceback ending in:
RuntimeError: unable to perform operation on
<TCPTransport closed=True ...>; the handler is closed
Adds RuntimeError to the except tuple in both _fwd_to_server and
_fwd_to_client (the latter is the actual frame from the bug report —
server→client is where buffered TLS chunks land after the client has
gone). The forwarders are intentionally fire-and-forget on tear-down;
the existing dst.close() in the finally block already handles cleanup.
No functional regression possible — the connection is already dead by
the time the exception fires; this only changes whether asyncio logs an
"Unhandled exception" trace for it.
2 new regression contract tests in test_camera_tls_proxy.py use
inspect.getsource to assert both forwarder closures' except clauses
include RuntimeError. Source-level rather than a runtime test because
the forwarders are nested closures inside _handle and extracting them
just for testability would require a pure-cosmetic refactor.
Latent since 0feed83c (Fix P2S camera TLS compatibility via OpenSSL
proxy, #661, 2026-03-15) — only commit that ever touched these
forwarders.
Follow-up to #1042. The post-dispatch watchdog _verify_print_response was
fire-and-forget — it correctly detected when the printer never transitioned
(HMS error pending, half-broken MQTT session, plate-clear gate, SD card
fault) and force-reconnected the MQTT session, but the dispatch job had
already been marked successful on the optimistic MQTT-publish-acknowledged
path. The UI carried on showing "Print started successfully" while the
printer sat idle.
The watchdog now returns bool and is awaited inline by both call sites in
_run_reprint_archive and _run_print_library_file. On False the call sites
raise a RuntimeError carrying a user-actionable message ("Printer did not
acknowledge print command — state still {pre_state}. Check the printer for
a pending error...") which routes through the existing _run_active_job →
_mark_job_finished(failed=True) → background_dispatch WS broadcast path.
Library-file flow rolls back the freshly-created archive on timeout so no
phantom row is left behind for a print that never started.
The watchdog now also accepts subtask_id advancing past pre_subtask_id as a
definitive "command landed" signal — same as the queue-side watchdog at
print_scheduler.py:1992 — so slow H2D FINISH→PREPARE transitions (~50 s
observed) don't false-fail when the printer has clearly accepted the
project_file but is still in FINISH. Default timeout raised from 15 s to
90 s to match the queue-side watchdog and give the same headroom on both
dispatch paths. Brief mid-window MQTT disconnects keep polling instead of
immediately failing — matches what the queue watchdog already does and
avoids false-failing on transient telemetry gaps.
11 new tests in test_background_dispatch_watchdog.py: state-change pickup,
subtask_id-change pickup with state still FINISH, neither-changed timeout
plus force_reconnect_stale_session call, pre_subtask_id=None backwards-
compat, post-dispatch subtask_id=None not counting as a change, brief
disconnect not short-circuiting the window, persistent disconnect for the
full window returning False, default-timeout=90s contract, _run_reprint_archive
raises RuntimeError with the captured pre-state args on watchdog False,
_run_reprint_archive happy path doesn't rollback, _run_active_job marks
the job failed with the message when _process_job raises RuntimeError.
feat(oidc): add Azure Entra ID support with configurable email claim resolution
Adds two new OIDC provider fields: email_claim and require_email_verified.
Two new optional fields on Spool: free-text `category` (max 50) and
`low_stock_threshold_pct` (1-99). Powers the "differentiate critical
spools from prototype spools and alert at different thresholds" use
case from #729 without taking on the full multi-tag taxonomy + auto-
apply rules + per-tag alert system the ticket originally proposed.
Form gains:
- Category input with datalist autocomplete sourced from categories
already in use, so casing/spelling stays consistent.
- Per-spool low-stock threshold input. Empty = global default; the
global value renders as the placeholder.
Inventory page:
- New category filter chip (hidden until at least one spool carries
a category — keeps the chip row uncluttered).
- Stat-card "Low Stock" count and the "Low Stock" filter both honour
the per-spool override.
Plus: rename "Delete Tag" button to "Clear RFID Tag" (the original
ticket reporter mistook it for a taxonomy-tag delete; the button
actually clears the RFID UID/UUID off the spool record). Toast key
renamed from `tagDeleted` to `rfidCleared`.
i18n: full translations across all 8 locales.
Tests: 9 new backend schema tests (defaults, partial-update, range
rejection, max-length); 2 new frontend tests (per-spool threshold
pulls extra spools into low-stock count, filter chip hidden when no
categories exist).
`find_matching_untagged_spool` is supposed to attach an incoming Bambu
RFID UUID to a pre-existing manually-logged spool of the same
material/color so users who log inventory before scanning don't end up
with duplicate rows. Two bugs meant it almost never worked for the
actual reporting workflow:
1. Subtype filter was strict. AMS reports `tray_sub_brands="PLA Basic"`
→ matcher required `Spool.subtype = 'Basic'` exactly. The form's
Quick-Add mode only requires `material`, so bulk-logged rows have
`subtype=NULL` and were always excluded → duplicate on first AMS
read.
2. Brand wasn't filtered. The docstring claimed brand was matched but
the WHERE clause didn't include it, so a same-color Polymaker (or
any non-Bambu) untagged row could acquire a Bambu UUID — silent
data corruption.
Fix in the same query: subtype prefers exact match but accepts NULL as
fallback (CASE in ORDER BY ensures exact wins when both exist); brand
restricted to NULL or LOWER(brand) LIKE '%bambu%' (covers 'Bambu',
'Bambu Lab', 'BambuLab', 'bambu lab' — the spellings users actually
type).
6 regression tests added in test_spool_tag_matcher.py.
ntfy supports a Priority header (1=min, 2=low, 3=default, 4=high, 5=urgent)
that controls escalation on the receiving device, but every event was being
sent at the server default — so a "50% complete" ping looked identical to
"print failed" or "printer offline". Add a per-event priority dropdown
section in the Add/Edit Notification modal (visible only for ntfy, listing
only enabled events); the backend reads config.event_priorities and emits
the matching Priority header on POST and PUT (image-attachment) paths.
Unmapped events fall through to the ntfy server default. Out-of-range
and non-numeric values are dropped, not clamped, so a misconfigured value
never silently sends at the wrong urgency. Test sends omit the header by
design so the test path can't accidentally page someone at urgent priority.
Backward compatible: existing providers without event_priorities behave
exactly as before. NtfyConfig.event_priorities is optional; the route
stores config as a JSON blob so no migration is needed.
i18n: full translations across all 8 locales (en/de/fr/it/ja/pt-BR/zh-CN/
zh-TW). README, CHANGELOG, and the wiki notifications page updated.
Tests: 6 backend (Priority set on mapped, omitted on unmapped/missing/
no-priorities, ignored for bad values, propagated through attachment
path), 6 frontend (section visible only for ntfy, lists only enabled
events, save round-trip, edit pre-fill, toggle drops row, non-ntfy
never writes the key).
The global CSP set script-src 'self', so FastAPI's /docs page rendered
blank: the inline boot <script> and the cdn.jsdelivr.net swagger-ui
bundle/CSS were both blocked. /redoc and /docs/oauth2-redirect had the
same problem.
Branch the security_headers_middleware to emit a docs-scoped CSP for
those three paths that allows cdn.jsdelivr.net (scripts + styles), the
FastAPI/Redoc favicon hosts (images), and 'unsafe-inline' for the
inline boot script. Every other route keeps the stricter SPA policy
unchanged.
#1108 — Long-lived camera-stream tokens for HA / Frigate / kiosks. Camera-only
V1, hard 365-day cap (no infinite tokens), pbkdf2 hashed at rest, plaintext
shown to user exactly once on creation. New "Camera API Tokens" panel under
Settings → API Keys with self-service create/revoke, styled confirm modal,
admin "All users" view for leak triage. Auth path: /camera/stream tries the
existing 60-min ephemeral table first, falls through to the long-lived path.
Indexed lookup_prefix keeps verify O(1) per token.
Permission audit: gated the existing API-keys-CRUD + Webhook docs + API
Browser content behind api_keys:read so non-admins with camera:view land on
the API Keys tab and see only the Camera Tokens panel they actually have
permission to use. Grid layout collapses to single column for non-admins.
Tests: 29 new backend (15 service + 14 integration covering create/list/
revoke ownership rules, the auth fall-through, scope enforcement, prefix
collisions) + 6 new frontend tests for the section UI including the new
modal flow. All 77 backend tests + 21 frontend camera tests pass. Ruff
clean (lint + format).
Docs: README updated with fan-out + long-lived-token bullets. Wiki gets a
new "Long-Lived Camera Tokens" section under features/camera.md (HA YAML
example, security model, permission requirements, revoke flow). Website
features.html gets the bullet under Camera Streaming.
Also includes #1089 follow-up tweaks already merged in this branch:
_stream_start_times.setdefault for accurate stream_uptime, subscribe()
RuntimeError retry to close the grace-vs-subscribe race, atomic
unsubscribe count via the iter_subscriber on_unsubscribe callback.
Most Bambu Lab printers only allow one concurrent camera connection, but
GET /printers/{id}/camera/stream opened a fresh upstream per viewer.
Two browser tabs → second viewer fails or kicks the first off.
New MjpegBroadcaster (services/camera_fanout.py) owns one upstream per
printer and fans MJPEG chunks out to N subscribers. 5 s grace window
absorbs tab refreshes without reconnecting. Bounded subscriber queues
drop frames for slow viewers rather than blocking the broadcaster.
Audit-pass fixes:
- _stream_start_times set with setdefault() so stream_uptime reflects
the shared upstream's age, not the most-recent viewer's
- subscribe() retried once on RuntimeError to close a tiny grace race
- unsubscribe() returns post-removal count atomically so the detach log
no longer races with concurrent leavers
Permission gates unchanged; broadcaster has no FastAPI surface.
Tests: 13 broadcaster unit tests + 2 integration tests on /camera/stop.
External-camera path untouched.
The dual-nozzle active-extruder card was the only tile in the printer
status row without a theme icon, making the row look uneven on H2D /
H2S / H2C. Adds a schematic nozzle icon (filament body + heater block
+ tip) matching the SVG @m4rtini2 contributed, sized and coloured to
match the adjacent Nozzle/Bed/Chamber temperature cards.
POST /library/files only rejected the read-only external branch and
then unconditionally wrote to get_library_files_dir() with a UUID
filename. The resulting LibraryFile row pointed at the external folder
via folder_id, so the file showed up in Bambuddy's UI, but the bytes
physically lived in archive/library/files/ and never touched the mount
-- invisible from any other machine accessing the NAS/SMB share.
Writable external uploads now write through to <external_path>/<filename>
with the original filename preserved, and the DB row matches what scan
produces (is_external=True, file_path=<absolute mount path>). Collisions
return 409 instead of silently overwriting; inaccessible or non-writable
mount returns 400; path-traversal filenames are rejected via resolve +
relative_to.
Extract-zip is now rejected against any external folder (not just
read-only) with a clear "extract on the mount and run Scan" message --
the nested-subfolder creation path would need mkdir on the mount plus
matching is_external LibraryFolder rows, which is a separate design.
Scan already handles that shape.
When a file sliced for the wrong nozzle size is dispatched, the printer
goes IDLE -> PREPARE -> FAILED without ever entering RUNNING. Completion
detection required prev=RUNNING or _was_running=True, so on_print_complete
never fired and the queue item stayed at "printing" forever -- blocking
every subsequent pending item for that printer (check_queue seeds
busy_printers from any row in 'printing').
Fire completion on FAILED from PREPARE or SLICING too. Restricted to
those two pre-print states so a stale FAILED on first connection
(prev=None) still can't accidentally advance an unrelated queue item.
Also populate PrintQueueItem.error_message from the current HMS error
list via the existing hms_errors.py lookup, so users see e.g.
"[0500_4038] The nozzle diameter in sliced file is not consistent
with the current nozzle setting" instead of a blank failure reason.
The SSRF guard added in this PR rejected all RFC-1918 private and loopback
addresses, which breaks Bambuddy's primary deployment topology — Spoolman
running on the same LAN as Bambuddy (192.168.x.x, 10.x.x.x, 127.0.0.1).
Users hit "Spoolman URL must not point to a private, loopback, link-local,
multicast, or unspecified address" on legitimate setups.
Rescope the guard to block what's actually dangerous in this context:
cloud metadata endpoints (AWS/Alibaba IMDS), multicast, unspecified,
non-http(s) schemes, and numeric-encoded IP bypasses. Loopback and
RFC-1918 ranges are now explicitly permitted.
Tests:
- test_ssrf_blocked_schemes_and_addresses updated with refined block list
- test_ssrf_allows_lan_spoolman_topologies (new) asserts loopback +
RFC-1918 are accepted so this regression cannot recur silently
- TestSpoolmanInventorySSRFSpoolBuddyPath parametrize lists trimmed
feat(inventory): replace Spoolman iframe with internal inventory UI
When Spoolman is enabled, the Inventory page now uses the same internal
UI (spool list, create/edit modal, archive, delete, weight sync) backed
by a new proxy layer instead of opening an iframe.
1. `_cancel_restart_task` self-await guard (manager.py:389-413).
stop_server() / stop_proxy() are called from inside
_restart_for_cert_renewal, which runs AS _cert_restart_task.
Cancelling+awaiting self flagged a CancelledError on the next
`await` in stop_server, tearing down old listeners but never
letting start_server run — the VP sat on the expired cert
until the process was manually restarted, silently defeating
auto-renewal. Skip when `task is asyncio.current_task()` and
just clear the reference.
2. Clipboard fallback textarea leak (VirtualPrinterCard.tsx:66-81).
The HTTP fallback created a hidden textarea, called
select() + execCommand('copy'), then removed the textarea.
If select() or execCommand threw, removal never ran and the
textarea leaked into the DOM. Move the removal into `finally`
so it happens regardless of the inner block's outcome.
Regression tests in test_tailscale.py::TestCancelRestartTaskSelfAwait
cover both the self-cancel path (must NOT cancel self) and the
outside-cancel path (must still cancel and await).
Add the Tailscale CLI to the production image and document how to
enable Let's Encrypt cert provisioning for virtual printers from a
Docker-deployed Bambuddy.
- Dockerfile installs `tailscale` from the official Debian repo. Only
the CLI is used at runtime; tailscaled itself stays on the host.
The binary is harmless if the socket isn't mounted — the code logs
an actionable hint and falls back to self-signed certs.
- docker-compose.yml adds a commented-out volume mount for
/var/run/tailscale/tailscaled.sock with inline setup instructions.
- tailscale.py's docker-socket hint now also fires when the binary is
present but the daemon socket is unreachable (i.e. the new Docker
pattern), not just when the binary is missing, so users get the
actionable "mount the socket" message instead of opaque CLI stderr.
Enabling the integration on a Docker host:
1. `curl -fsSL https://tailscale.com/install.sh | sh` on host
2. `sudo tailscale up`
3. `sudo tailscale set --operator=<user>` for the container PUID
4. Uncomment the tailscaled.sock mount in docker-compose.yml
5. `docker compose up -d --force-recreate`
6. Flip the Tailscale toggle on the VP card
The Tailscale FQDN copy button used only `navigator.clipboard.writeText`,
which browsers block when `window.isSecureContext === false` — i.e. when
Bambuddy is reached over HTTP on a LAN / tailnet IP, which is the
common case. My catch block swallowed the error and the generic
"Failed to update settings" toast fired instead of actually copying.
Add a legacy `document.execCommand('copy')` fallback via a hidden
textarea for non-secure contexts. New i18n key
`virtualPrinter.toast.copyFailed` added to all 8 locales for the
(rare) both-paths-fail case.
Legacy SQLite installs created the `settings` table without a UNIQUE
constraint on `key`. The seed loop's `INSERT OR IGNORE` silently
degraded to a plain INSERT, so every `systemctl restart` added another
row of `advanced_auth_enabled` / `smtp_auth_enabled`. After a handful
of restarts, `scalar_one_or_none()` in is_advanced_auth_enabled() and
similar sites blew up with `MultipleResultsFound`, 500'ing the login
flow.
Run-migrations now deletes dup rows (keeping MIN(id) per key) and
creates the missing `ix_settings_key` unique index before the seed
loop. Both ops are idempotent — fresh installs and Postgres already
have the index, so they no-op.
Bambu started shipping H2C units with a new serial prefix (`31B8B…`
observed on a January 2026 unit) instead of the legacy `094…` shared by
the H2D/H2C/H2S family. Two serial-prefix-driven paths — the K-profile
edit branch in `kprofiles.py` and the delete-K-profile MQTT command in
`bambu_mqtt.py::delete_kprofile` — were silently routing the new units
through the single-nozzle format.
Match on 5 chars (`31B8B`): covers the 3-char model code plus the two
revision bytes, leaving the revision-letter slot free to iterate. This
mirrors the X2D precedent of using a longer-than-3-char prefix when a
single data point can't confirm family reuse.
Runtime dual-nozzle detection via `device.extruder.info` count and
model-string branches (`self.model in ("H2C", "H2D", …)`) are already
prefix-agnostic — no change needed there.
- backend/app/api/routes/kprofiles.py: add "31B8B" to is_h2d tuple
- backend/app/services/bambu_mqtt.py: same in delete_kprofile
- backend/tests/unit/services/test_bambu_mqtt.py: regression test
`test_h2c_new_prefix_uses_dual_nozzle_format`
Fix a silent correctness bug: archive purge used `created_at` which is
pinned to the first print, so reprinting a two-year-old archive yesterday
would still make it eligible for a 365-day purge. The preview and purge
queries now age each archive by `COALESCE(completed_at, started_at,
created_at)` — reprints refresh the clock.
Also flesh out both purge modals (File Manager + Archives) with an
explicit "What happens when you click Purge" effects list so users see
upfront that library files go to Trash (reversible) while archives are
hard-deleted (irreversible), plus what disk artefacts get removed.
Backend:
- services/archive_purge.py: `_last_activity_expr()` helper used by
preview, purge, and sample query
- tests/integration/test_archive_purge_api.py: new test covering the
reprinted-archive case
Frontend:
- PurgeOldFilesModal / PurgeArchivesModal: new effects bullet list
- i18n: reprint-aware ageLabel/description/warning and effects bullets
across all 8 locales (en/de fully translated, rest English fallback)
Docs:
- wiki/features/archiving.md: "How old is measured" note + effects list
- wiki/features/file-manager.md: "What happens when you click Purge"
section + explicit age-rule breakdown
- CHANGELOG: archive auto-purge entry rewritten to mention reprint
semantics, `archives:purge` permission backfill, and updated test count
Adds an archive counterpart to the library trash sweeper shipped in the
previous commit. Unlike the library flow, archives are hard-deleted —
print history is a decaying timeline, so there is no trash intermediate;
download or favourite anything you want to keep first.
Backend
- New ArchivePurgeService (backend/app/services/archive_purge.py) with
its own 15-minute scheduler loop and a 24h throttle on actual purge
runs. Delegates every delete to the existing safety-checked
ArchiveService.delete_archive so the 3MF, thumbnail, timelapse, source
3MF, F3D, and photo folder all get cleaned up together with the DB
row. Per-row session via async_session() avoids commit-per-row churn
on any caller-passed session.
- New /archives/purge/{preview,settings} + POST /archives/purge routes
gated on a dedicated archives:purge permission (not archives:delete_all)
so admins can delegate bulk-delete to a role without granting
per-archive delete on other users' rows.
- seed_default_groups() now backfills both library:purge and
archives:purge on the Administrators group for upgraded installs —
the original library:purge was added after Administrators was first
seeded so the "create if not exists" path skipped existing DBs and
left admins without the permission.
- 8 new integration tests (defaults, settings roundtrip, bound
validation, preview, manual purge, auto-purge enabled path, 24h
throttle, disabled skip).
Frontend
- Settings → Archives card gains an auto-purge toggle + age input (7d
floor, 10y ceiling, 365d default), with a save-toast on every change.
The bulk "Purge old" button lives on the Archives page header
(rightmost, after Upload 3MF) to match the File Manager pattern —
configuration in Settings, one-shot action on the page.
- New PurgeArchivesModal mirrors PurgeOldFilesModal: live preview (count
+ total size freed + sample filenames) debounced at 300ms, amber
"hard-delete, no undo" warning.
- Admin-only UI gates on archives:purge via the standard hasPermission
hook; Permission TS union updated.
- i18n blocks across all 8 locales (en/de full, other 6 English
fallback per project convention).
Docs
- CHANGELOG entry under 0.2.4b1 following the existing library-trash
entry.
- bambuddy-wiki archiving.md gains a new "Auto-Purge" section.
- bambuddy-website features.html gets a matching bullet.
Verification: python -m ruff check backend/app/ clean; 25 integration
tests pass (8 archive_purge + 17 library_trash regression); npm run
build clean.