Commit Graph
286 Commits
Author SHA1 Message Date
maziggy 12d17bfbe7 fix(photo): source finish photo from forced timelapse + cleanup (#1397)
Bambu's end-gcode lowers the bed at gcode_state=FINISH. Bambuddy's
  live-camera grab captured the bed already dropped, ruining the photo
  framing. Source the photo from a brief Bambu timelapse instead —
  firmware stops timelapse recording AFTER toolhead parks but BEFORE
  bed-drop runs, so the last frame frames the finished print correctly.

  When capture_finish_photo is on AND the user did not opt in to
  timelapse for this print, force timelapse=True at dispatch + mark the
  new PrintArchive.bambuddy_forced_timelapse column. After extraction
  (success or failure), cleanup deletes the locally-attached file,
  clears archive.timelapse_path, and walks the four scanner directories
  (/timelapse, /timelapse/video, /record, /recording) trying FTP DELE
  against the original filename. User-opted-in timelapses pass through
  unchanged.

  Resolver lives at services/background_dispatch.py::resolve_effective_timelapse
  (module-level so the print queue can reuse it). Both dispatch paths
  wired: background_dispatch.py (Print Now / Reprint) AND
  print_scheduler.py:_start_print (the queue). Field testing caught the
  scheduler gap on the first round — AST regression test now asserts
  start_print(timelapse=...) references effective_timelapse, not the raw
  item.timelapse, so a future refactor can't silently drop it.

  Extractor: ffmpeg -i input.mp4 -update 1 -q:v 2 out.jpg. Decoded
  frames overwrite the same output file, so the file left on disk is the
  literal last frame regardless of duration. Bambu records one frame per
  layer-change, so a 16-layer cube produces a 0.6 s timelapse — the
  original -sseof -1.0 approach seeked before the start of the file and
  returned frame 0 (empty bed). Decoding every frame is fine; Bambu
  timelapses are short by construction even on hours-long prints.

  Migration adds bambuddy_forced_timelapse branched on is_sqlite()
  (DEFAULT 0 / DEFAULT FALSE — PG rejects DEFAULT 0 for BOOLEAN).
  Verified live on postgres:16-alpine.

  Photo-task wait_for budget extends 45s -> 75s when timelapse_was_active
  so the notification carries the bed-up photo instead of falling back
  to the live-cam grab on slow links.

  Scope limit, documented in the camera wiki: prints started directly
  on the printer touchscreen / Bambu Handy / Bambu Studio Send bypass
  both dispatch paths, so the override doesn't fire there. Future
  option: mid-print M981 S1 P20000 MQTT toggle in on_print_start.

  Setting description rewritten in all 11 locales to drop the "only
  works when timelapse enabled" caveat (Bambuddy now forces it) and
  explain the kept-or-deleted behaviour.
2026-06-05 13:53:26 +02:00
maziggy 18d534c945 feat(orca-cloud): integrate Orca Cloud profile sync across UI, slicer and SpoolBuddy
Reads, lists, and slices with profiles from a user's Orca Cloud account
  (OrcaSlicer 2.4.0-alpha's Supabase-backed sync) alongside the existing
  Bambu Cloud integration. Four sign-in providers (Google / Apple / GitHub /
  email+password); password defaults. Paste-flow PKCE because Orca's
  Supabase project only allowlists localhost redirect_to — open feature
  request at OrcaSlicer/OrcaSlicer#14028.

  Surfaces:
  - Profiles tab: new "Orca Cloud" tab next to "Bambu Cloud" with the same
    rich layout (search + 5 filter dropdowns + 3-column grouped grid +
    read-only detail modal)
  - SliceModal: 4-tier preset picker (orca_cloud > local > bambu cloud >
    standard); separate status banner per cloud; metadata-aware pre-pick
    scores Orca filaments above local (Orca's sync_pull returns full
    content inline so filament_type / filament_colour come for free, no
    per-setting fetch rate-limit dance)
  - ConfigureAmsSlotModal: orca_cloud as a new preset source (prefixed
    orca_<UUID> to match local_/builtin_); generic Bambu filament-ID
    derivation from parsed material (printer firmware can't grok Orca
    UUIDs); slot mapping persists preset_source='orca_cloud'
  - SpoolForm / SpoolBuddyWriteTagPage: Orca filaments merge into the
    cloud preset list via Promise.allSettled (OrcaProfileMeta is
    structurally identical to SlicerSetting)

  Backend:
  - services/orca_cloud.py: OrcaCloudService with PKCE / token exchange /
    single-use refresh rotation / get_user_info / list_profiles via the
    bare /sync/pull bootstrap path
  - routes/orca_cloud.py: 7 endpoints (auth/start, auth/finish,
    auth/password, status, logout, profiles, profiles/{id}); router-level
    _cloud_api_key_gate + per-route cloud_caller() so API-keyed callers
    (SpoolBuddy kiosk) properly resolve their owner User; just-in-time
    refresh with atomic persist-before-API-call
  - routes/slicer_presets.py: _fetch_orca_cloud_presets mirrors the Bambu
    Cloud fetcher (status vocabulary, 5min cache, permission shortcut);
    _dedupe_by_name extended to 4 tiers; UnifiedPresetsResponse gains
    orca_cloud + orca_cloud_status
  - services/preset_resolver.py: PresetRef.source extended with
    "orca_cloud"; _resolve_orca_cloud walks list + filters
  - 8 columns on users table for tokens (5 persistent) + transient PKCE
    handshake state with 10-min TTL (3); dialect-branched DATETIME /
    TIMESTAMP; auth-disabled mode falls back to Settings table
  - orca_cloud:auth permission folded into can_access_cloud API-key scope
    (same trust dimension)
2026-06-04 15:50:44 +02:00
maziggy 51730a7bf1 feat(ams): gate humidity/temperature alarms on AMS-has-filament (#1619)
The hourly AMS sensor recorder dispatched humidity and temperature alarms
  for every unit above threshold without checking whether the unit was
  actually loaded. Empty AMS units still report ambient readings, so users
  with one loaded + one empty AMS got useful alarms for the loaded one and
  hourly noise for the empty one. Disabling the whole alarm category killed
  both — not a real choice.

  New _ams_has_filament helper inspects tray_exist_bits (hex bitmap, "0" =
  empty) with fallback to the tray array's tray_type strings for shapes
  where the bitmap is missing. The recorder gates the alarm dispatch on
  this check per-AMS-unit, so a multi-AMS printer with one loaded + one
  empty still alarms on the loaded one.

  Sensor history still records regardless of the gate so the System page
  humidity charts stay continuous — only the outbound notification is
  suppressed. 9 unit tests cover the bitmap-zero case, bitmap-missing
  fallback, garbage/blank/int bitmap edges, and defensive malformed tray.
2026-06-04 11:30:26 +02:00
maziggy e10462678f fix(security): GHSA-6mf4-q26m-47pv — fail-closed on auth-probe DB errors (CVSS 9.8)
is_auth_enabled() and auth_middleware both caught every exception during
  the auth-state probe and returned the "allow" answer instead of denying
  the request. Reporter's PoC floods /api/v1/auth/login to exhaust file
  descriptors, forcing the next SQLite connect to raise, then hits a
  protected endpoint during the fail-open window with no token — granting
  unauthenticated access to admin-account creation, API-key creation, DB
  backup download, and printer control. CWE-636 / CWE-755. Affects >= 0.1.6.

  Fix:
  - is_auth_enabled (backend/app/core/auth.py): only returns False for the
    legitimate "settings row absent" case; any actual exception propagates
    so the caller can deny the request.
  - auth_middleware (backend/app/main.py): returns 503 on any probe failure
    instead of await call_next(request).

  4 new regression tests in test_auth_fail_closed.py pin the contract
  (propagates DB exceptions, returns False for no-row, True for "true",
  False for "false"). 1 existing security test renamed and updated to
  accept either 500 or 503 (both fail-closed) and to verify the
  SQLAlchemy detail does not leak in the body.

  Codebase grep confirmed no other auth-decision predicate has the same
  fail-open shape: _validate_api_key returns None on catch (→ 401 fail-
  closed downstream), is_advanced_auth_enabled propagates correctly,
  permissions.py has no catch-alls.

  Reported by @wondercrash via private advisory.
2026-05-30 13:49:09 +02:00
maziggy ed232718f5 fix(prints): connected-edge reconciliation closes the missed-PRINT-COMPLETE loop behind smart-plug ghost prints (#1542 follow-up)
Reporter ran a fresh trace after the doubled-extension fix landed and
  found a distinct second cause behind his ghost prints, hitting 4-of-4
  of his A1s. Timeline:

    22:50 PRINT START
    ...print runs all night...
    23:13 / 00:47 / 09:35 MQTT disconnects (A1 keepalives are unstable)
    print finishes during one of those disconnect windows → PRINT COMPLETE
      is never observed
    smart plug detects idle → cuts power
    power resumes for the next scheduled print → firmware auto-replays the
      leftover .3mf from the SD card
    09:46 Bambuddy reconnects to a fresh PRINT START for the ghost

  The existing IDLE-after-RUNNING completion check at
  backend/app/services/bambu_mqtt.py:3022 was meant to catch the simple
  disconnect-then-finish case via `_previous_gcode_state` preserved across
  reconnects, but with multiple disconnect/reconnect cycles + a smart-plug
  power-off that Bambuddy can't distinguish from any other transient drop,
  the IDLE window that branch needs simply never reaches it. The SD .3mf
  lingers, the firmware ghost-replays every power cycle, and the loop
  repeats.

  Fix: new connected-edge reconciliation pass.

  * `_is_active_archive_stale(archive, state)` — pure decision function
    with three triggers:
      (1) printer state is terminal (IDLE / FINISH / FAILED)
      (2) printer running with a different `subtask_id` than the archive —
          Bambu firmware mints a fresh subtask_id for each print including
          the ghost replay, so a mismatch is unambiguous
      (3) printer running but `subtask_name` is empty — printer doesn't
          know what it's running, archive reference is broken
    Conservative on PAUSE / PREPARE / SLICING and on RUNNING with matching
    subtask. False-positive cost = one misreported "aborted" status that
    the next real PRINT COMPLETE would have overwritten anyway. False-
    negative cost = the ghost-print loop.

  * `reconcile_stale_active_prints(printer_id)` — queries archives in
    `status="printing"` for the printer, runs the decision function, and
    synthesises `on_print_complete(status="aborted", _reconciled=True)`
    for each stale match. Reuses the existing PRINT COMPLETE chain (SD
    cleanup, status update, usage tracker, notifications) — no reimplem-
    entation. Per-archive try/except so one failure doesn't block the
    rest. Returns 0 when status is None / disconnected — the connected
    edge is the only legitimate trigger.

  * `on_printer_status_change` now runs a connected-edge check at the
    start. New `_printer_reconciled_since_connect: dict[int, bool]`
    tracker flips False → True on the first connected status update for
    this connection and back to False on disconnect, so reconciliation
    fires exactly once per (re)connection. The flag is set BEFORE the
    task is spawned so concurrent status updates within the same
    connection don't re-trigger it (no await between check and set,
    asyncio guarantees atomicity). The reconciliation runs as
    `asyncio.create_task` so the hot WebSocket dedup / broadcast path
    isn't blocked.

  * Single handler covers both startup and reconnect — when the first
    MQTT connection completes after startup the printer pushes status,
    the connected edge fires, reconciliation runs. No separate startup
    hook needed.

  Idempotency: when the existing #3022 branch DOES fire on a clean
  disconnect-then-IDLE-on-reconnect, it lands `archive.status` to terminal
  synchronously. The async-scheduled reconciliation then queries
  `WHERE status="printing"` and finds 0 rows → no-op. The narrow race
  where reconcile's query lands between a real on_print_complete's
  archive update and its archive lookup produces at most a duplicate
  notification (no double SD-cleanup since FTP delete on a missing file
  is a no-op 550).

  Ghost-print collateral worth being explicit about: if the ghost is
  already running when reconciliation fires, the synthesised SD-cleanup
  hits 550-file-locked (same root cause as the #1542 first case). The
  cleanup retries 3× then logs "lingering". The ghost runs to completion,
  its own end-of-print cleanup deletes the file, the next power cycle
  has nothing to replay, the loop breaks. A perfect cancel would require
  a `print_stop` MQTT command to the printer mid-ghost — invasive,
  explicitly out of scope.
2026-05-29 11:18:06 +02:00
maziggy 7a7dfed81a fix(archives): resolve raw_data wrapper in fallback-archive filament extraction (#1533 follow-up)
Reporter updated to 0.2.5b1 expecting the #1533 fix to populate filament
  fields on his P2S virtual-printer prints when the .3mf is locked. His
  support bundle showed Bambuddy still creating fallback archives with NULL
  filament fields even though the print-start log line proved AMS-0-T0 had
  PETG loaded at the moment the helper should have read it
  (`AMS 0: T0(type=PETG, color=FFFFFFFF, ...)`).

  Cause: the #1533 helper `_extract_filament_data_from_mqtt(data)` in
  backend/app/main.py only looked at `data["ams"]`, but the dict that
  on_print_start actually receives at runtime is the wrapper shape
  `{"filename", "subtask_name", "remaining_time", "raw_data": <mqtt>,
  "ams_mapping"}` that backend/app/services/bambu_mqtt.py:2971-2980
  constructs. So `data["ams"]` was undefined on every real call and the
  helper silently returned `{}`, leaving the fallback archive's
  filament_type / filament_color NULL — the exact regression the original
  fix was meant to close. The 15 unit tests that shipped with #1533 all
  passed the bare inner shape directly and never exercised the callback
  wiring, so the regression slipped through a green build. Surrounding code
  in the same fallback path (main.py:2064, 2674) already reads
  `data.get("raw_data")` — the original hunk was the outlier that forgot
  the wrapper.

  Fix: helper now resolves `data["raw_data"]["ams"]` first (the callback
  shape) and only falls back to `data["ams"]` when the wrapper isn't
  present, preserving the inner-shape callers from the existing tests.
  Defensive against a non-dict `raw_data` (e.g. partial MQTT decode
  failure) falling through to the inner lookup instead of crashing.

  Tests: 5 new in `TestOnPrintStartCallbackShape` in
  test_fallback_archive_mqtt_filament.py — wrapper payload with ams_mapping
  resolves to the inner AMS state; wrapper without ams_mapping lists all
  loaded slots; the existing inner-shape callers still work after the
  additive wrapper lookup; missing raw_data returns `{}` instead of
  raising; junk raw_data (string) doesn't shadow a present inner `ams`.
  Existing 15 inner-shape tests untouched and green. Full 5378-test
  backend suite green; backend ruff clean.

  What this does NOT fix: per-filament gram usage still needs the actual
  .3mf — the printer locks it during print (P-line firmware behaviour,
  not a Bambuddy bug), and the existing 19 FTP candidate paths +
  directory probes are expected to 550 in that window. Per-print filament
  type and colour are the data point the reporter explicitly called out
  as load-bearing for AMS-expansion planning at his maker space, so this
  is what moves the needle.
2026-05-29 10:55:41 +02:00
maziggy 6d316c5593 fix(dispatch): align SD cleanup with upload path so doubled-extension library rows don't leave ghost prints (#1542)
A library row with archive.filename "Cube (1).gcode.3mf.gcode.3mf" uploaded
  to /Cube_(1).gcode.3mf.3mf (single-iteration strip + append). Post-print
  cleanup looked at /Cube_(1).3mf and /Cube_(1).gcode (subtask_name + ext),
  missed the on-card file, and A1 firmware re-ran it on next power-on.

  - New derive_remote_filename() helper in backend/app/utils/filename.py:
    iterative strip of .gcode.3mf/.3mf suffixes, append single .3mf,
    space->underscore. isinstance() guard raises TypeError on non-str
    input rather than entering the strip loop with a duck-typed
    object that returns truthy sentinels from endswith.
  - Three previously-duplicated upload sites (background_dispatch reprint +
    library, print_scheduler queue) now share the helper.
  - SD cleanup fetches archive.filename and tries the derived path first,
    with the legacy subtask_name + ext paths kept as fallbacks.
  - 10 new unit tests pin the reproducer + edge cases (doubled .gcode.3mf,
    doubled .3mf, raw .gcode preserved, idempotence, Unicode, plus the
    type guard against MagicMock / None / int inputs).
2026-05-28 08:40:43 +02:00
maziggy c42e923e4c fix(archives): MQTT-derived filament type/color on fallback archives (#1533)
When the source .3mf can't be downloaded at print start (P1S/A1/P2S
  firmwares lock the file mid-print), main.py creates a fallback
  PrintArchive with file_path="" and every filament field NULL — even
  though the MQTT payload already has the AMS state and the slicer's
  slot-per-print-filament mapping (data["ams"]["ams"] and
  data["ams_mapping"]).

  New _extract_filament_data_from_mqtt(data, ams_mapping) builds a
  {global_tray_id: (type, color)} map from the AMS units, then narrows
  to slots referenced by ams_mapping (slicer order preserved, -1 VT-tray
  sentinels skipped) or falls back to every loaded slot when no mapping
  is present. Returns comma-separated filament_type and filament_color
  matching the 3MF-extraction shape, so the inventory page, Quick Stats
  rollup, and len(filament_type.split(",")) per-print count behave
  identically for fallback rows.

  The constructor at the fallback site now passes the resulting values
  into the PrintArchive row.

  This does NOT recover per-filament gram usage — that needs the .3mf's
  slice_info.config or a deeper layer-delta integration via usage_tracker.
  The reporter (maker-space lead evaluating Bambuddy partly for AMS
  expansion planning) asked specifically for "the number of filaments
  used", which is what this gives them.

  15 unit tests cover empty/malformed payloads, the no-mapping path,
  mapping filtering and reordering, VT-tray sentinels, dual-AMS global
  ids, column-limit truncation, and defensive garbage handling.
2026-05-26 11:50:41 +02:00
maziggy 554a73070f fix(maintenance): paused prints no longer accumulate runtime hours (#1521)
PAUSE counted toward runtime_seconds equally with RUNNING, inflating
  hours-based maintenance thresholds (rod lube, belt check, nozzle clean)
  by however long overnight or extended pauses lasted. Maintenance items
  track mechanical wear, which is zero while paused, so the predicate
  now excludes PAUSE. Field-comment and docstring trail across main.py /
  models/printer.py / maintenance.py updated to match. Existing
  runtime_seconds values cannot be retroactively split — only future
  accumulation is fixed.

  Adds 3 regression tests pinning PAUSE non-accumulation, RUNNING
  accumulation, and the FINISH state's last_runtime_update clear
  (prevents idle-time back-bill when the printer next goes RUNNING).
2026-05-25 09:04:47 +02:00
maziggy ed08ed3787 fix(timelapse): capture baseline on restart-recovery so post-reboot timelapses attach (follow up issue #1485)
When Bambuddy is restarted mid-print, the first MQTT push from the
  printer carries `_previous_gcode_state = None`. The #1304 guard
  deliberately suppresses on_print_start on that first push to prevent
  duplicate archive creation — but on_print_start is also where
  _capture_timelapse_baseline_at_start runs, so the in-memory
  _timelapse_baselines dict stays empty for the resumed session.

  At PRINT COMPLETE, _scan_for_timelapse_with_retries finds no baseline
  and falls into its "take baseline now" fallback. By that point the
  printer has already uploaded the in-flight MP4, so the snapshot
  includes the new file. Every "Found N files / no new files since
  baseline" retry then fails to detect a diff, and the archive ends up
  with no timelapse attached — pwostran's report (#1485 follow-up): card
  shows the finish snapshot but no video.

  Add a sibling callback on_print_running_observed that bambu_mqtt fires
  in the "Now tracking RUNNING state" branch when on_print_start was
  suppressed. main.py wires it to a thin handler that looks up the
  printer row and calls the existing _capture_timelapse_baseline_at_start.
  Idempotent — skips if a baseline already exists (handles the rare
  same-session race where on_print_start also fires for some reason).

  The printer doesn't upload the timelapse until after PRINT COMPLETE,
  so a baseline captured any time during the print is still pre-upload —
  no narrow window to hit.

  Verified against the in-the-field logs in #1485 (pwostran's 2026-05-23
  support bundle):
    pre-reboot: baseline = 7 files
    reboot
    post-reboot completion: fallback baseline = 8 files (includes new MP4)
    -> all 4 retry attempts report "no new files since baseline"

  10 new tests cover both the MQTT-side fire decision (fires when
  suppressed, doesn't fire when on_print_start handles it, once per
  session, payload shape mirrors on_print_start) and the main.py
  handler (snapshot capture, double-capture guard, missing-printer-row
  guard).
2026-05-23 14:52:04 +02:00
maziggy 4096d8d6bd fix(csp): nonce-based script-src so Cloudflare-injected scripts pass (#1460 follow-up)
Behind Cloudflare, the bot-detection script CF injects into every HTML
  response carries a hash that rotates per request, so it can never be
  allowlisted by hash. Reporters with CF in front had to relax their NPM
  CSP to 'unsafe-inline' as a workaround.

  Per Cloudflare's documented behaviour, when a nonce is present in the
  page's script-src, CF clones it onto its injected <script>. The SPA CSP
  now stamps a fresh per-request nonce via secrets.token_urlsafe(16),
  keeping 'self' for our own scripts (index.html has had no inline scripts
  since the SW registration moved to /sw-register.js in the original
  #1460 PR), so no HTML body rewriting is needed.

  Also folded in: /manifest.json, /sw.js and /sw-register.js now accept
  HEAD as well as GET, so `curl -I` and uptime scanners stop returning
  405 on those routes - a separate red herring during this issue's
  debugging.

  Tests: 3 new in test_security_headers.py - 'nonce-' token stamped into
  SPA script-src while 'self' remains and 'unsafe-inline' does not; nonce
  is fresh per request across 5 sequential calls; HEAD on the three PWA
  routes never returns 405. 22/22 security-header tests green; backend
  ruff clean.
2026-05-23 12:28:58 +02:00
maziggy 16da533c9a fix(static): serve /fonts/*.woff2 — self-hosted Inter font (#1460 follow-up)
The browser console logged "downloadable font: rejected by sanitizer"
  for inter-latin.woff2 on every load. The #1460 PWA fix added @font-face
  rules pointing at /fonts/inter-latin.woff2 and bundled the woff2 files
  into static/fonts/, but main.py only mounts /assets, /img and /icons as
  static directories. With no /fonts mount, /fonts/*.woff2 fell through to
  the SPA catch-all and returned index.html with 200 OK; the browser's
  OpenType sanitizer rejected the HTML-as-a-font.

  Add a /fonts StaticFiles mount alongside /img and /icons. The woff2
  files themselves are valid (verified — Inter variable, latin and
  latin-ext subsets).

  Also bump the service worker STATIC_CACHE version (v26 -> v27). sw.js
  lists the two font URLs in STATIC_ASSETS, and cache.addAll() treats the
  200 OK HTML as a successful fetch — so it had cached index.html under
  the font URLs and served it cache-first. The version bump makes the
  activate handler purge the poisoned cache and re-fetch the real fonts.
2026-05-22 10:43:27 +02:00
maziggy 774eba73c8 feat(diagnostics): event-loop stall watchdog to catch silent backend freezes
Several "container hangs after adding a printer" reports (#1486) share a
  signature with nothing to act on: HTTP goes silent, /health hangs, the
  process may ignore SIGTERM, and the log just stops mid-stream - a frozen
  asyncio loop cannot log a thing.

  loop_watchdog re-arms faulthandler.dump_traceback_later() from an async
  heartbeat. While the loop ticks the timer is cancelled and re-armed before
  it can fire; if the loop stalls for 30s the heartbeat can't re-arm and
  faulthandler's C-level timer thread dumps every thread's stack to stderr,
  so the blocked frame shows up in `docker compose logs`.
2026-05-22 09:28:43 +02:00
maziggy 745ed847e6 fix(archive): stop duplicating the job on a backend restart mid-print (#1485)
A restart during an active print duplicated the running job in the
  archive, and every further restart spawned another. on_print_start
  re-attaches by subtask_id, falling back to a name match plus a 4-hour
  staleness cutoff that cancelled + recreated any name-matched 'printing'
  archive older than 4h - destroying the live archive of every long print.

  Two fixes:
  - start_print records the minted subtask_id (last_dispatch_subtask_id);
    on_print_start falls back to it when the printer hasn't echoed one
    yet, so queue/scheduled archives persist a restart-stable id.
  - Replace the 4h cutoff with a progress-aware check: a name-matched
    'printing' archive resumes whenever the printer reports real (or
    unknown) progress; it is stale only when the printer shows a
    freshly-started print (<1%) on an archive over 2h old.
2026-05-22 09:05:25 +02:00
maziggy 17e39921bb fix(pwa): add in-app install button and self-host the Inter font (#1460)
Bambuddy installed as a PWA on desktop but not on Android. Two causes:

  - Chrome for Android removed the automatic install banner in Chrome 108.
    With no beforeinstallprompt handler, Android had no install path. New
    InstallAppButton captures the event and re-fires it from the sidebar.
  - index.css pulled Inter from fonts.googleapis.com: breaks offline, trips
    CSP, and the service worker answered the failed cross-origin request
    with cached index.html. Inter is now self-hosted; the SW skips all
    cross-origin requests and caches the font; CSP drops the Google hosts.
2026-05-22 07:40:44 +02:00
maziggy bfd3fc755d Fix: capture timelapse baseline on expected-archive on_print_start branch (#1403 follow-up)
The snapshot-diff strategy in _scan_for_timelapse_with_retries needs
  _timelapse_baselines[printer_id] populated at print start so the
  completion-time scan can find the new MP4 by set-difference (mtime is
  unreliable — LAN-only printers don't sync NTP).

  The baseline-capture call was only in on_print_start's new-archive branch.
  Queue / VP-dispatched / reprinted jobs take the expected-archive branch
  which returns earlier, so the dict stayed empty and the completion-time
  scan fell into the "take baseline now" fallback that snapshots after the
  new file has already landed — no diff ever matches.

  Extract the snapshot into _capture_timelapse_baseline_at_start and call
  it from both branches.
2026-05-20 08:12:17 +02:00
MartinNYHC 12a352e5b8 Merge branch 'main' into dev 2026-05-19 14:12:50 +02:00
maziggy 0b33862ae9 fix(archives): assign printer_id when reusing VP-queue archives in print-start (#1403 follow-up)
VP-queue archives are created with printer_id=None at queue-add
  time because the scheduler hasn't picked a printer yet (and even
  for explicit-printer queue items, the archive predates dispatch).
  on_print_start's expected-archive branch updated status,
  started_at, and subtask_id but never assigned printer_id, so
  VP-queue-dispatched archives stayed permanently unassigned.

  That broke every UI/API path gated on archive.printer_id —
  critically the post-print "Scan for timelapse" action: the
  H.264 file is on the printer's SD card and reachable via the
  file browser, but the archive's scan endpoint refused the request
  and the button stayed greyed out forever.

  One-line fix: archive.printer_id = printer_id in the
  expected-archive branch. Guarded against clobbering an
  already-correct value so library-file queue items (which create
  their archive with the printer pre-assigned) are idempotent.
2026-05-19 12:55:16 +02:00
maziggy fc32b388de fix(stats): align Filament Used / By Time / Success Rate with Total Consumed and Total Prints (#1390 follow-up)
Three independent root causes behind the divergences the reporter
  flagged after the archived-spool fix shipped — fixed together.

  (1) Filament Used vs Total Consumed.
  _compute_run_filament_grams returned the slicer estimate for completed
  prints even when inventory had measured the actual AMS weight delta.
  That made Stats and Inventory two different sources of truth: Stats
  showed slicer-estimate grams, Inventory showed AMS-tracked grams, and
  the two never agreed. Reordered the helper so the tracked spool delta
  (same source that drives weight_used behind Total Consumed) takes
  priority for every status. Slicer estimate stays as the fallback when
  no inventory was tracked; partial-progress scale stays as the fallback
  for failed/cancelled with no tracker. The _run_cost block right next
  to it was already tracker-first; only filament_used_grams was
  inconsistent.

  (2) Printer Stats By Time vs Quick Stats Print Time.
  /archives/slim only set actual_time_seconds when status == "completed".
  For failed/cancelled rows the frontend fell back to print_time_seconds
  (the slicer's full-print estimate — wrong number for a print that
  failed at 15%). Quick Stats already summed elapsed duration across
  all statuses, so the two halves of the page disagreed by the
  (estimate - actual-elapsed) gap on every non-completed event. Dropped
  the completed-only gate; failed/cancelled now report measured elapsed.

  (3) Success Rate %.
  Was successful / (successful + failed), excluding cancelled / stopped
  from the denominator. With "Total Prints: N" displayed right above
  the gauge that produced confusing numbers — 4 successful, 0 failed,
  48 cancelled showed 100% out of an apparent 52 prints. Switched to
  successful / total_prints — matches the count the user reads from
  the widget header.
2026-05-19 10:53:36 +02:00
maziggy 6f2cec5eb3 feat(smart-plugs): auto-off after AMS drying completes (#1349)
Reporter Kyobinoyo asked for the equivalent of the existing
  print-finish auto-off but triggered when AMS drying ends.

  Two new SmartPlug columns: auto_off_after_drying (default false),
  off_delay_after_drying_minutes (default 10 — AMS chamber is hot
  post-cycle so longer cooldown than the print-finish default of 5).
  SQLite + Postgres migrations both idempotent.

  Trigger lives in BambuMQTTClient — per-AMS _previous_dry_times
  tracks the dry_time > 0 → 0 falling edge and fires a new
  on_drying_complete(ams_id) callback. Plumbed through
  PrinterManager.set_drying_complete_callback to
  SmartPlugManager.on_drying_complete(printer_id, db), which walks
  linked plugs and respects the per-plug toggle. Catches queue,
  ambient and manual drying identically because it observes firmware
  state, not scheduler intent.

  Frontend: single "Auto Off After Drying" toggle + delay input on
  the smart plug card, next to the existing print-finish auto-off
  section.

  Per-AMS plug routing (separate plug for AMS only, per-AMS targeting
  on dual-AMS printers) deferred — Bambuddy's plug model is
  plug→printer, so the trigger fires whenever any AMS on the linked
  printer finishes a cycle.
2026-05-17 12:26:42 +02:00
maziggy 6569d5d1a7 refactor(timelapse): extract _maybe_start_layer_timelapse + rewrite test
The CI-only failure on test_layer_timelapse_expected_archive came from
  the test driving the entire on_print_start flow through ~12 patches and
  a MagicMock printer, which behaved differently between Python 3.11 (CI)
  and 3.13 (local) — execution stopped silently somewhere in the
  expected-archive path under CI's pytest-xdist parallelism but completed
  locally.

  Fix root-shape instead of fix the symptom:

  1. Extract the three identical start_session call sites in on_print_start
     (expected-archive promotion at 2030, fallback archive at 2554, fresh
     archive at 2644) into one helper _maybe_start_layer_timelapse() with
     the same external_camera_enabled / external_camera_url guard. The
     three inline blocks had already started drifting (#1353 originally
     only fixed one of them on the first pass) — the helper keeps them
     locked together going forward.

  2. Rewrite the test to call the helper directly. Uses SimpleNamespace
     (strict attribute access) instead of MagicMock (default-truthy), no
     DB mocking, no event loop, no parallel-state surface. Four small
     cases instead of two integration-style ones: enabled→starts,
     disabled→skips, URL-missing→skips, camera_type default 'mjpeg'.
2026-05-16 14:35:47 +02:00
maziggy 856b849ffa fix(stats): per-event aggregation so reprints add to Quick Stats instead of overwriting (#1378)
Statistics now aggregate over PrintLogEntry (one row per print event,
  the same table backing the global Print Log) rather than PrintArchive
  (one row per file). A reprint creates a new PrintLogEntry instead of
  overwriting the source archive's runtime fields, so:

  - a 100 g successful print + a 10 g failed reprint correctly sums to
    110 g / 2 prints / 1 successful / 1 failed in Quick Stats and the
    Prometheus /metrics endpoint (previously the failed reprint silently
    replaced the source archive's data; totals dropped from 100 g to 10 g)
  - the archive's card cost/energy_kwh are preserved on reprints (only
    the first run writes them); per-run actuals live on PrintLogEntry
  - failed/cancelled/stopped reprints record partial-aware filament: sum
    of tracked spool deltas when inventory is set up, else estimate
    scaled to progress%, else None — prevents the full slicer estimate
    from inflating totals on a print that stopped at 10 % progress

  PrintLogEntry gains six columns: archive_id (nullable FK, ON DELETE
  SET NULL so log entries survive archive deletion preserving #1343
  soft-delete-vs-stats decoupling), cost, energy_kwh, energy_cost,
  failure_reason, created_by_id. Idempotent SQLite + Postgres migrations.

  New per-archive surface:

  - archive list response carries run_count / last_run_at /
    total_filament_actual_grams / successful_run_count / failed_run_count
    via a single batch JOIN, no N+1
  - new GET /archives/{id}/runs endpoint returns every PrintLogEntry for
    the archive (ARCHIVES_READ permission, newest-first ordering)
  - archive cards render an orange "N prints" badge for archives with
    more than one run; clicking the badge opens a dedicated PrintLogModal
    with date/status/duration/filament/cost columns plus failure_reason
    under failed runs. Also reachable via the context menu's new "Print
    Log" entry (works for single-run archives too), and embedded at the
    top of the Edit Archive modal for context.

  The purge_stats=true delete path now hard-deletes linked PrintLogEntry
  rows up front so the archive's contribution truly leaves the totals;
  without it, ON DELETE SET NULL would orphan the runs and leave them
  counting toward stats.
2026-05-16 11:34:40 +02:00
maziggy f2e3de0a63 fix(camera): start layer timelapse for queue/VP-dispatched prints (#1353)
Reporter @Andlar94 ran the external-camera flow on an A1 dispatched via the
  print queue and got no MP4 output even though the log said "Stitching layer
  timelapse for printer 1" after each print. Support bundle confirmed the
  external camera was working (Obico was polling the snapshot URL fine for
  plate detection).

  Root cause: start_session() only ran in the two new-archive paths in
  on_print_start (fallback_archive at main.py:2510 and regular new-archive at
  2600). The expected-archive branch at main.py:1981-2052 — where every
  reprint and every queue/VP-dispatched print lands — updated the existing
  archive row to status=printing but never started a timelapse session.

  So _background_layer_timelapse ran at print complete, called tl_complete(),
  found nothing in _active_sessions, returned None silently, and the wrapper
  at main.py:3917 produced no log message for the no-session case. Every
  print through the queue silently lost its timelapse — likely the reason
  this hasn't been caught before (direct slice-and-send-to-printer prints
  take the new-archive path and work fine).

  Fix: mirror the same start_session() call in the expected-archive branch,
  guarded by the same external_camera_enabled + external_camera_url check the
  other two paths use.

  Also reworded the snapshot URL help text across all 8 locales to make clear
  that timelapse and plate detection each require their own per-printer
  toggle — the URL is just the image source they pull from when active. The
  previous wording read as if filling in the URL was sufficient.
2026-05-15 12:46:59 +02:00
maziggy f45aaea97c fix(inventory): assign to AMS slot on firmwares that never report state=11 (#1322)
A1 Mini BMCU (01.07.02.00) and P1S Standard AMS (00.00.06.75) always
  report tray.state=3, even for loaded configured slots. The empty-slot
  detection preferred state==11 with tray_type as a fallback only when
  state was absent, so every assign was classified as empty and MQTT
  was skipped — both for "assign to unconfigured slot" and the secondary
  "PETG over a PLA-configured slot won't reconfigure" symptom.

  Empty-slot detection in the assign route and the on_ams_change replay
  now treats the slot as loaded when EITHER state==11 OR tray_type is
  non-empty. Reset-slot case (state=11 + tray_type="") still works
  through the first clause; configured slots on these firmwares now
  work through the second.

  Truly empty unconfigured slots (state!=11 + tray_type="") still hit
  the pending-config path, and the deferred publish now fires when the
  user later configures the slot in Bambu Studio (tray_type goes
  non-empty), since the replay uses the same disjunction.
2026-05-14 14:36:57 +02:00
maziggy b334d7edc9 fix(spoolman): per-print 3MF tracking is the only weight writer (#1119)
Spoolman had two mutually-exclusive weight paths gated on the
  `disable_weight_sync` flag. The default (False) used AMS remain%
  x tray_weight auto-sync, which silently dropped non-BL spools
  because the AMS doesn't report tray_weight without RFID. The
  inventory_remaining fallback would have covered it, but the
  spool_assignment table it reads from is wiped on Spoolman
  activation, so non-BL spools got no weight updates at all.

  Match the internal Filament Inventory: per-print tracking always
  runs, AMS auto-sync no longer writes remaining_weight (it still
  maintains spool metadata and slot assignments). The setting
  becomes a no-op; left in the schema and UI for backwards compat.

  - store_print_data: drop the disable_weight_sync early return
  - sync_ams_tray callsites in main.py + routes/spoolman.py: force
    disable_weight_sync=True so weight is never written by AMS sync
  - new regression test confirming tracking runs with flag=false
2026-05-12 08:15:28 +02:00
MartinNYHC b30a283184 Feature/spoolman inventory UI (#1241)
feat(spoolman-inventory): squashed feature work for rebase onto dev

Squashed all commits from feature/spoolman-inventory-ui onto a single commit
to enable a clean rebase onto dev. Original per-commit history preserved at
backup tag backup/spoolman-inventory-ui-prerebase-20260507-105721.
2026-05-08 11:52:42 +02:00
MartinNYHC dac2a31192 Revert "feat(inventory): unified Spoolman inventory UI + AMS slot assignments…" (#1232)
This reverts commit 55d71498e9.
2026-05-07 11:30:31 +02:00
Sn0rrii 55d71498e9 feat(inventory): unified Spoolman inventory UI + AMS slot assignments + Storage Location + NFC write support + Spoolman Filament Catalog Picker (#1114)
feat(spoolman-inventory): squashed feature work for rebase onto dev

Squashed all commits from feature/spoolman-inventory-ui onto a single commit
to enable a clean rebase onto dev. Original per-commit history preserved at
backup tag backup/spoolman-inventory-ui-prerebase-20260507-105721.
2026-05-07 11:15:24 +02:00
maziggy a3e09891d1 fix(docker): copy gcode_viewer assets into the production image (issue #1218)
The embedded GCode viewer's static assets (gcode_viewer/) were never
  copied into the production Docker image, so /gcode-viewer/ returned a
  bare FastAPI 404 ({"detail":"Not Found"}) and 3D Preview broke for every
  Docker user since the viewer landed in 0.2.4b1. The Vite production
  build doesn't stage the directory either — the dev server serves it via
  a configureServer middleware that's dev-only.

  Dockerfile now copies gcode_viewer/ alongside the React build output.

  Defence in depth: main.py logs an ERROR at startup when
  _gcode_viewer_dir/index.html is missing so future packaging gaps surface
  in docker logs and the support bundle instead of as silent runtime 404s.

  The existing integration test accepted 404 unconditionally
  (assert response.status_code in (200, 404)) so CI never caught the
  missing files. Add test_gcode_viewer_index_served_when_assets_present
  which skips when the directory is intentionally absent (unit-test envs)
  but asserts 200 + non-empty HTML body when the assets do exist on disk —
  so a broken COPY fails CI loudly rather than shipping a broken image.
2026-05-06 14:23:39 +02:00
maziggy 864e5c990e feat(inventory): printable PDF spool labels in 4 sizes (#809)
Closes the longest-standing inventory gap — finding a specific spool
  in a closet of 50 partials. Per-spool icon button on every inventory
  card and table row, plus a "Print labels..." header action that opens
  a multi-select picker pre-loaded with the currently filtered spools.

  Four pre-built templates: AMS holder (30 x 15 mm) for the popular
  Makerworld AMS Filament Label Holder, single box label (62 x 29 mm)
  for Brother PT/QL or Dymo small labels, Avery L7160 (A4, 21 per
  sheet), and Avery 5160 (US Letter, 30 per sheet). Each label carries
  the colour swatch (with multi-colour gradient stripes for spools
  with extra_colors set), brand, material, name, the *spool ID*
  (bsaunder's articulated user-need: telling 8 spools of "PLA White"
  apart, especially partials), and a QR code that deep-links to
  /inventory?spool=<id> for phone-scan round-trips. Box-label adds
  storage location; AMS-holder drops the QR — at 30 x 15 mm there is
  no room for swatch + text + QR without truncating away the spool ID,
  and AMS-bay identification is at arm's length where the swatch and
  ID are enough.

  Server-side rendering via ReportLab + qrcode (already a dep). Pure
  Python, no headless browser, no system libs. Output is byte-identical
  across browsers, Avery sheets align to <0.1 mm, and bulk export is
  one click for one PDF. Two endpoints — POST /inventory/labels (local
  DB) and POST /spoolman/labels (Spoolman-backed) — gated on
  INVENTORY_READ, capped at 500 spools per request, returning
  application/pdf via StreamingResponse. The renderer is decoupled
  from the SQLAlchemy model via a LabelData dataclass so the same code
  path serves both modes.

  Modal picker scales to large libraries: search (substring match
  across name / brand / #ID), material filter chips derived from the
  visible spools, additive Select-all-visible / Deselect-visible /
  Clear-all actions so selections survive filter changes. Restyled
  twice in development — first cut used generic Tailwind which clashed
  with the inventory's bambu-dark palette; second cut switched to
  bambu-dark-secondary / bambu-green / bambu-gray to match.

  Two render bugs found during visual inspection of generated PDFs and
  fixed before commit:

    1. AMS-30x15 template originally produced labels with only swatch
       + QR and no text at all — the side-by-side layout left <5 mm
       for the text column, so the renderer bailed without drawing
       anything. Layout split into tight (h<20mm) and roomy (h>=20mm)
       regimes; tight regime drops the QR and gives the right column
       to brand + material + a 13pt-bold spool ID.

    2. Box-62x29 template aggressively truncated text — swatch + QR
       each at ~14 mm on a 26mm-tall label squeezed the text column
       to ~16 mm, turning "Polymaker Ivory" into "Polymak..." and
       "Polymaker . PLA . Matte" into "Polymaker ...". Swatch capped
       at 16 mm, QR capped at 18 mm and constrained to ~20% of width,
       leaving the text column ~30 mm — full names render without
       truncation.

  Both bugs pinned by regression tests in test_label_renderer.py that
  render with pageCompression=0 so the resulting PDF bytes contain the
  text as ASCII and `assert b"Polymaker" in pdf` works.
2026-05-05 12:29:57 +02:00
maziggy b42aaca521 fix(spool-assign): defer MQTT for empty AMS slot, replay on physical insert
The SpoolBuddy "weigh-then-assign" workflow tried to configure an empty AMS
  slot at assign time, but Bambu firmware silently drops ams_filament_setting
  and extrusion_cali_sel for unloaded slots — the MQTT calls completed and the
  modal closed, yet BambuStudio kept showing the slot as default-PLA forever.

  assign_spool now detects an empty target slot (fingerprint_type empty) and
  persists the SpoolAssignment without publishing MQTT, returning a new
  pending_config flag so the frontend can swap "Assigned!" for "Slot will
  configure when you insert the spool." on_ams_change watches for the slot
  to load (state == 11, which fires for 3rd-party tags too even when
  tray_type stays empty) and replays the deferred ams_filament_setting +
  extrusion_cali_sel — including the printer-kp realignment that converts
  PFUS-prefix cloud user presets to the P-prefix local-preset filament_id
  the slicer actually accepts.

  The full assign-time MQTT block was extracted into
  apply_spool_to_slot_via_mqtt so both the assign endpoint and the
  on_ams_change replay path use the same resolution logic; the helper takes
  ~270 lines of duplication out of assign_spool.
2026-05-04 16:02:18 +02:00
maziggy 713b85387a fix(archives): validate downloaded 3MF plate against gcode_file (#1204)
Two consecutive plates of the same model would create the second print's
archive with the first plate's metadata: subtask_name lags across the
boundary while gcode_file is fresh, so the FTP candidate list (built
from subtask_name first) lands on the previous plate's still-resident
upload. The 3MF parser then locks the wrong _plate_index, name, time
estimate, and per-slot filament data into the archive at creation.

Fix peeks the downloaded 3MF's slice_info plate index, compares against
parse_plate_id(filename) (the plate parsed from /Metadata/plate_N.gcode,
which always reflects what's running), and on mismatch retries FTP with
swap_plate_suffix(subtask_name, expected_plate) — handling both the
spaced "Plate N" and underscored "_plate_N" suffix forms seen in real
subtask_names. If the retry finds a matching 3MF, the wrong file is
dropped and the corrected one feeds the archive; if no match is found
(or no swap is possible) the wrong file is dropped and the existing
no-3MF fallback creates an archive whose name reflects the right plate.

The validation only runs when parse_plate_id() returns a value, so
single-plate / cloud-named / non-Bambu jobs are unaffected.

17 new unit tests in test_archive_plate_validation.py cover both helpers:
plate-index peek across malformed / missing / non-integer / non-zip
inputs, and the suffix swap across both casings, the underscored form,
case-insensitive matching, and rejection of names without a recognised
suffix.
2026-05-04 10:09:25 +02:00
maziggy 7dea33d0d8 feat(vp): mirror live target printer state to slicer in non-proxy modes
In non-proxy VP modes (Immediate / Review / Print Queue), the slicer now
sees real AMS / FTS / nozzle / k-profile state from the target printer
and streams the live camera — full slicer-as-remote functionality without
giving up Bambuddy's queue / archive / dispatch features.

Architecture (cached-as-base, single source of truth). The bridge caches
the latest real push_status and info.get_version response from Bambuddy's
existing per-printer MQTT subscription — no second session on the printer,
firmware in-flight budget unaffected (#1164). _send_status_report serves
a near-byte-identical copy of the cached push with only the upload-state-
machine fields overridden. Command responses (extrusion_cali_get, AMS
write acks, xcam) fan out raw — they carry sequence_ids the slicer is
waiting on. Slicer-issued commands forward to the printer except
project_file / gcode_file, which still terminate locally because the file
lives on Bambuddy. Camera is a raw TCPProxy on bind_ip:322 → printer:322,
same approach proxy mode uses.

Field-shape gotchas pinned in the bridge module's docstring and the
new test file:
  - Real Bambu pushes use json.dumps(indent=4) wire format. Compact JSON
    fails BambuStudio's Send pre-flight silently.
  - net.info[*].ip is the FTP destination IP (little-endian uint32).
    Without rewriting to the VP bind IP, the slicer FTPs straight to
    the real printer.
  - upgrade_state.sn rewritten to VP serial; AMS-hardware sn fields
    (n3f/0.sn etc.) left alone.
  - ipcam.rtsp_url passes through unchanged; BambuStudio overrides the
    URL host with the device IP it bound on, so :322 lands on the VP's
    TCPProxy.
  - extrusion_cali_get must forward; answering it locally hides the
    user's stored per-filament k-profiles.

Setup nuance for camera: the VP's access code must match the target
printer's because the slicer authenticates RTSPS with whatever access
code is in its profile. MQTT and FTP work either way.

Tested e2e with BambuStudio and OrcaSlicer against H2D (dual-nozzle,
AMS 2 Pro + AMS HT) and X1C across all three non-proxy modes — sync,
send, k-profile lookup, AMS configuration from slicer, and live camera
all work. Proxy mode is untouched: SlicerProxyManager owns its own
proxies and never instantiates SimpleMQTTServer or MQTTBridge.

25 new tests in backend/tests/unit/test_vp_mqtt_bridge.py cover lifecycle,
caching, identity / IP rewriting, wire format, slicer→printer routing,
and the LE-uint32 IP encoder against the real H2D capture value.
2026-05-03 13:43:54 +02:00
maziggy a6c53798d4 fix(notifications): print-complete duration uses actual elapsed, not slicer estimate (#1198)
Pre-fix, _background_notifications in main.py:3434 built archive_data
  with print_time_seconds (the slicer's pre-print estimate parsed from
  the 3MF at archive creation), and notification_service.py:909 formatted
  that field straight into the {{duration}} template variable. A print
  cancelled 2 minutes into a 3-hour estimate notified "duration: 3h".

  Compute actual_time_seconds from started_at/completed_at in main.py and
  add it to archive_data. notification_service.py prefers it, falls back
  to print_time_seconds when the actual can't be derived.

  Also add "cancelled" to the list of statuses that get completed_at set
  in update_archive_status — pre-fix only completed/failed/aborted got a
  timestamp, so queue-UI cancellations had no actual elapsed to compute
  from. Audited every completed_at consumer; none depend on NULL to mean
  "cancelled" (status field already carries that signal), and the
  statistics-totals aggregation gets more accurate too as a side effect.

  3 new regression tests in TestNotificationVariableFallbacks pin the
  {{duration}} variable contract (actual wins over estimate; estimate
  falls in when actual is missing; "Unknown" when both absent).
2026-05-03 08:33:00 +02:00
maziggy abc8e97050 feat(camera): optional snapshot URL override for external cameras (#1177)
go2rtc and several IP cameras still emit a warm-up / black frame on every
  fresh MJPEG connection — even with the v0.2.4b2 warm-up-skip fix it
  slipped through intermittently for @nkm8's setup. His own bisect named
  the clean solution: go2rtc exposes /api/frame.jpeg as a dedicated
  single-frame endpoint that never returns the encoder's stale keyframe.

  Adds an optional external_camera_snapshot_url column on printers. When
  set, every single-frame capture path (snapshot endpoint, [SNAPSHOT]
  notification thumbnails, [PHOTO-BG] finish photo, layer timelapse,
  Obico ML, plate-detect / calibrate-plate) routes through _capture_snapshot
  on the override URL via plain HTTP GET, bypassing the warm-up dance.

  Live view stays on the configured stream URL — only single-frame
  captures use the override. Override is camera-type-agnostic. SSRF guard
  applies (existing _sanitize_camera_url allowlist). Empty string treated
  as unset.

  Settings UI: new "Snapshot URL (optional)" input + Test button under
  External Cameras, hidden for camera_type=snapshot since the live URL is
  already a single-frame source. en + de fully translated; 6 other locales
  seeded with English copy.

  5 backend tests pin the routing contract; 3 frontend tests pin the
  input + debounced PATCH. Documented in
  bambuddy-wiki/docs/features/camera.md with the go2rtc example.
2026-05-03 07:59:02 +02:00
maziggy b02350d423 fix(security): allow iframe embedding from trusted origins via env var (#1191)
Bambuddy ships strict anti-clickjacking headers (X-Frame-Options:
  SAMEORIGIN + CSP frame-ancestors 'none') by default. Internet-exposed
  deployments need this; same-LAN HA Webpage-panel users do not, and
  SAMEORIGIN is port-strict so HA on :8123 + Bambuddy on :8000 always
  fails. azurusnova hit exactly that case.

  Add TRUSTED_FRAME_ORIGINS env var (comma-separated scheme://host[:port]).
  When set, drop X-Frame-Options entirely (modern browsers honor
  frame-ancestors and the legacy ALLOW-FROM syntax is deprecated /
  inconsistent across vendors) and emit "frame-ancestors 'self' <list>"
  on every CSP-bearing route. Origin validation is strict: only http(s),
  no paths, no query/fragment, no wildcards. Bad entries get a warning
  and are dropped — startup never fails.

  Default behaviour (no env var) is unchanged: X-Frame-Options:
  SAMEORIGIN + frame-ancestors 'none', so existing Docker / bare-metal
  deployments are not affected.
2026-05-02 12:32:02 +02:00
maziggy 25eab96817 fix(scheduler): raise plate-clear gate for every terminal status (#1171)
The plate-clear gate added in #961 was raised only when a print ended
  with status completed or failed. Aborted prints (printer self-abort
  or a user stopping the print from the printer's own touchscreen) and
  cancelled prints (user stopping via the Bambuddy queue UI) did NOT
  raise the flag, so the queue scheduler dispatched the next pending
  item ~2 seconds later onto a fouled bed.

  The reporter saw two prints (P1P + P1S) auto-start onto fouled beds
  within seconds of touchscreen-aborts, and explicitly flagged the
  risk of damage to the printer. A third printer behaved correctly
  because its previous print had ended "completed" — the asymmetry he
  noticed was the gate working for one terminal status and not the
  other three.

  Touchscreen-aborts are particularly important to gate. Bambuddy's
  existing "user stopped via UI" override (which translates aborted
  to cancelled when _user_stopped_printers is populated) only fires
  for stops through the Bambuddy queue UI; a touchscreen stop reports
  aborted straight through.

  The original code comment claimed user-cancelled prints don't need a
  plate-clear ack because "nothing printed on the bed". That only
  holds if you cancel right at layer 1; a cancel at hour 11 of a
  12-hour print leaves a fully fouled bed.

  The gate is user-clearable on the Printers page, so worst case a
  user who cancels at layer 1 clicks "Clear Plate" once — that's a
  non-issue compared to auto-dispatching onto material.

  Regression coverage in test_print_lifecycle.py::TestPlateClearGate:
  parametrised across all 4 terminal statuses asserting
  set_awaiting_plate_clear(printer_id, True) is called for each, plus
  a defence-in-depth test that an unrecognised future status string
  never silently raises the gate.
2026-05-01 08:13:52 +02:00
maziggy 61c15aac03 feat(slicer): unified Cloud/local/standard presets + harden 3MF profile path
UNIFIED PRESET LISTING (the main feature)

  The initial slicer integration only saw DB-backed local imports — users
  without imported profiles got an empty Slice modal even when their
  Bambu Cloud account or the slicer sidecar carried perfectly usable
  presets. The Slice modal now pulls from three tiers in priority order:

    - cloud:    user's own Bambu Cloud presets, fetched live.
    - local:    DB-backed imports.
    - standard: slicer-bundled stock profiles via the sidecar's new
                GET /profiles/bundled endpoint.

  Listing endpoint: GET /api/v1/slicer/presets

    - Name-based dedup, cloud > local > standard, within-tier order
      preserved exactly. A preset that exists in multiple tiers only
      renders in the highest-priority one.
    - cloud_status (ok / not_authenticated / expired / unreachable)
      drives a precise modal banner instead of an unexplained empty
      list.
    - Cloud branch: per-user cache, 5 min TTL, key
      (user_id, sha256(token)[:16]) so logout/login or token rotation
      auto-invalidates without callback wiring from the cloud-auth
      routes.
    - Bundled branch: global cache, 1 h TTL.
    - Bundled URL respects preferred_slicer (bambu_studio vs orcaslicer)
      so BambuStudio installs see the bambu sidecar's bundled list, not
      OrcaSlicer's.

  Slicing endpoint: POST /library/files/{id}/slice + /archives/{id}/slice

    - Body now accepts source-aware {source, id} triplets per slot:
        printer_preset:  PresetRef
        process_preset:  PresetRef
        filament_preset: PresetRef
    - Legacy *_preset_id integer fields kept for backwards-compat. The
      schema validator normalises bare ints into
      PresetRef(source='local', id=str(int)) so the route handler only
      deals with one shape.

  New preset_resolver service fetches the JSON content per source:

    - cloud:    BambuCloudService.get_setting_detail(id), unwraps the
                `setting` envelope (falls back to top-level for minor
                shape variants).
    - local:    DB read with preset_type slot validation (existing path,
                factored into the new helper).
    - standard: minimal {name, inherits, from: "system"} stub — the
                sidecar's profile-resolver flattens it against
                BUNDLED_PROFILES_PATH/<category>/<name>.json with no
                preset-content round-trip from Bambuddy.

  PERMISSIONS

    - Listing route gate: LIBRARY_UPLOAD (matches the slice action — any
      user who can slice can populate the dropdowns).
    - Cloud branch in BOTH the listing helper and the resolver checks
      CLOUD_AUTH independently — a user with LIBRARY_UPLOAD but not
      CLOUD_AUTH doesn't see the cloud tier (returns 403 if they try
      to slice with a cloud preset) even if a leftover User.cloud_token
      survived a permission revocation. Cloud listing path
      short-circuits the token lookup entirely on the gate-fail branch.

  FRONTEND — SliceModal

    - Calls api.getSlicerPresets() instead of api.getLocalPresets().
    - Dropdowns render <optgroup> per tier with localised section
      labels (Cloud / Imported / Standard).
    - Default selection follows cloud > local > standard priority on
      first load (auto-pick fires once when the data arrives, manual
      choices stick after that).
    - Cloud-status banner renders three variants
      (sign-in / expired / unreachable) only when status != 'ok'.
    - Slice button submits source-aware refs; legacy integer payload
      is preserved server-side for older clients.

  3MF PROFILE-PATH HARDENING (shipped together because they touch the
  same code paths)

  (1) Strip widened. _strip_3mf_embedded_settings only removed
      Metadata/project_settings.config. Real-world Bambu Studio /
      OrcaSlicer 3MFs also carry model_settings.config, slice_info.config,
      and cut_information.xml — any single leftover trips the CLI's
      input validation and the slice falls back to embedded settings,
      making the SliceModal's profile picker theatrical for 3MF inputs.
      Now removes all four configs via a centralised
      _STRIPPABLE_3MF_CONFIGS frozenset with per-file rationale;
      geometry (3D/3dmodel.model), thumbnails, multi-part data
      preserved.

  (2) Sidecar 5xx error capture. slicer_api.py was reading only
      `message` from sidecar 5xx responses and dropping `details`, so
      every CLI failure surfaced as the unhelpful generic
      "Failed to slice the model". New _format_sidecar_error helper
      combines both fields, falls back to plain-text body for
      non-JSON 5xx (nginx 502s, gateway timeouts), replaces the four
      duplicated extraction blocks. Pairs with the orca-slicer-api
      fork's bambuddy/profile-resolver branch which now emits
      `details` on AppError responses (d9c6121) and captures CLI
      stderr in the failure path (fb928c8).

  CARE TAKEN — additive on existing surfaces

    - main.py:               +1 import, +1 router register
    - slicer_api.py:         +list_bundled_profiles, +_format_sidecar_error
                             (dedupes the 4 message-extraction blocks);
                             no existing method behaviour changed
    - library.py:            resolver swap inside _run_slicer_with_fallback,
                             user_id threaded through two callers,
                             strip widened
    - schemas/slicer.py:     PresetRef added, *_preset fields added,
                             legacy *_preset_id kept; validator normalises
    - 4 new files:           schema, route, resolver, tests
    - No existing route URL changed, no existing field removed, no
      behaviour change for clients still sending bare integer ids.

  TESTS

    - 17 unit tests for the listing endpoint helpers
    - 11 unit tests for the source-aware resolver
    - 6 schema tests for SliceRequest legacy + new shapes
    - 3 unit tests for the new sidecar error-detail capture
    - Strip integration test extended to assert all 4 configs go and
      geometry stays
    - 12 frontend tests for SliceModal covering tier-priority
      auto-selection, <optgroup> grouping, fallback paths, source-aware
      payload on submit, manual override across tiers, archive vs
      library routing, error display, all three banner variants

  Verified: 3394 backend + 1531 frontend tests pass, ruff clean,
  frontend production build clean.

  Pairs with three already-pushed commits on the orca-slicer-api fork's
  bambuddy/profile-resolver branch:

    - 5fd6bc6  feat(profiles): add GET /profiles/bundled
    - d9c6121  fix(error): include causeMessage in JSON response as `details`
    - fb928c8  fix(slicing): include CLI stdout/stderr in failure causeMessage
2026-04-27 19:15:31 +02:00
MartinNYHC 8829bc2cc6 Merge branch 'dev' into feature/slicer-api 2026-04-27 17:09:42 +02:00
maziggy 9884018497 fix: cancel-safe get_db + drop sqlalchemy.pool cancellation noise
@Carter3DP's support package showed bambuddy.log filling with two
  distinct cascades on long uploads:

    ERROR sqlalchemy.pool   Exception terminating connection ...
                            CancelledError: Cancelled via cancel scope
                            ... by starlette.middleware.base
                            .BaseHTTPMiddleware.__call__.call_next
    ERROR sqlalchemy.pool   The garbage collector is trying to clean up
                            non-checked-in connection ... will be
                            terminated.
    WARN  backend.app.main  Runtime tracking commit failed:
                            (sqlite3.OperationalError) database is locked

  Single root cause. Starlette's BaseHTTPMiddleware (used under the hood
  by every @app.middleware("http") decorator) cancels the inner task
  scope when a client disconnects mid-request — common on long
  multipart uploads where the client times out before the server's
  response. Pre-fix get_db only caught Exception, but CancelledError
  is BaseException, so cancellation skipped the rollback path entirely.
  The SQLite write lock stayed held until GC reclaimed the connection
  ages later, blocking every other writer in the meantime. On Postgres
  the leak shape is identical; the symptom would be "QueuePool limit
  ... overflow" instead of "database is locked".

  (1) get_db now catches BaseException so CancelledError triggers
      rollback. Both rollback() and close() are wrapped in
      asyncio.shield so the cleanup completes even when the await
      itself is being cancelled by the same cancel scope. SQLite write
      lock is released promptly; connection returns to the pool instead
      of leaking until GC.

  (2) CancelledPoolNoiseFilter (new filter on sqlalchemy.pool) drops
      the residual records that pre-existing pools still emit during
      their own cleanup. Two patterns suppressed:
        - "Exception terminating connection ..." with a CancelledError
          anywhere in the exc_info chain (walks __cause__/__context__
          with a seen-set guard against pathological cycles)
        - "The garbage collector is trying to clean up non-checked-in
          connection ..." (always symptomatic of cancellation; never
          independently actionable)
      Real pool problems — broken connections, OSError on terminate,
      pool exhaustion — keep flowing because they carry a different
      exception chain or a different message prefix.

  13 regression tests across test_get_db_cancel_safety.py (commit on
  clean exit, rollback on regular Exception, rollback on CancelledError,
  close runs even if rollback raises, close failure on clean exit
  doesn't propagate, rollback + close both go through asyncio.shield)
  and test_cancelled_pool_filter.py (drops cancellation-driven
  terminate, drops GC-cleanup, keeps real OSError terminate, keeps
  terminate without exc_info, keeps unrelated pool messages, drops
  chained-cause CancelledError, defensive guard against self-referential
  cause chains).

  Applies to SQLite and PostgreSQL — get_db is dialect-agnostic and
  the filtered messages come from base sqlalchemy.pool not from any
  specific dialect.
2026-04-27 16:32:10 +02:00
maziggy 56800589ff fix(#1113): silence Windows asyncio Proactor cleanup-RST noise
bambuddy.log on Windows fills with

    Exception in callback _ProactorBasePipeTransport._call_connection_lost()
    ConnectionResetError: [WinError 10054] An existing connection was
    forcibly closed by the remote host

  every time a printer / MQTT broker / camera RSTs a TCP socket instead
  of FINing it. The application-layer reconnect (paho-mqtt, httpx)
  handles the actual disconnect fine; the traceback is asyncio
  bookkeeping. Reported by @cadtoolbox who runs 9 printers including 5
  offline X1Es, so the log filled multiple times per minute.

  New backend/app/core/asyncio_handlers.py installs a custom
  loop.set_exception_handler on Windows that pattern-matches three
  signals together (platform == win32, exception is
  ConnectionResetError, asyncio message contains
  _call_connection_lost) and demotes the entry to DEBUG. Genuine
  ConnectionResetErrors raised inside application coroutines have a
  different message string and still surface; BrokenPipeError /
  ConnectionAbortedError on the same cleanup path also still surface.

  Wired from lifespan startup before any task can spawn that might
  trip it. Linux / macOS use the Selector loop, so install is an
  explicit no-op there with a False return.

  9 unit tests in test_asyncio_handlers.py covering signature match,
  rejection of unrelated resets, platform gate, suppress vs.
  pass-through to default handler.
2026-04-27 16:14:31 +02:00
maziggy 6deaa513af ● feat(slicer): server-side slicing via OrcaSlicer / Bambu Studio sidecar
Adds an optional slicer-api/ Compose stack and wires Bambuddy's File
  Manager, Archives, and MakerWorld pages to a new server-side Slice flow.
  Slicing runs as an in-memory background job (POST returns 202 + job_id,
  polled via GET /api/v1/slice-jobs/{id}) so a multi-minute slice no
  longer pins the modal; result lands as a new .gcode.3mf in the same
  folder (or new archive for archive sources) with the embedded
  thumbnail extracted.

  Backend
  - New services: slice_dispatch (in-memory dispatcher, 30min retention
    sweep) and slicer_api (HTTP bridge with 4xx/5xx/connection error
    split that drives the 3MF embedded-settings fallback retry path).
  - New schemas: SliceRequest, SliceResponse, SliceArchiveResponse,
    SliceJobEnqueueResponse.
  - New routes: POST /library/files/{id}/slice,
    POST /archives/{id}/slice, GET /api/v1/slice-jobs/{id} (gated on
    LIBRARY_READ since job IDs are sequential and the body leaks source
    filenames and result IDs).
  - AppSettings + env defaults: use_slicer_api, orcaslicer_api_url,
    bambu_studio_api_url. DB-stored values override env defaults.

  Frontend
  - New SliceModal handles preset gating; enqueues then closes
    immediately.
  - New SliceJobTrackerProvider polls active jobs at app level, surfaces
    a single toast per job (queued -> running -> completed / failed)
    and invalidates library/archives queries on terminal status.
  - Settings -> Workflow -> Slicer card: preferred slicer dropdown,
    Use Slicer API toggle, contextual sidecar URL field.
  - File Manager / Archives / MakerWorld get a Slice button gated on
    the Use Slicer API setting.
  - gcode-viewer adapter learns ?library_file=<id> so sliced library
    files preview inline.

  i18n
  - New slice.* and settings.{useSlicerApi,slicerCard,orcaslicerApiUrl,
    bambuStudioApiUrl,slicerApiUrlDescription,useSlicerApiDescription}
    + fileManager.noPermissionSlice keys across all 8 locales (en, de,
    fr, it, ja, pt-BR, zh-CN, zh-TW). English fully translated, German
    fully translated, the other six seeded with English fallbacks
    pending native translation.

  Tests
  - 10 backend integration tests in test_library_slice_api.py covering
    validation (404/400), happy-path enqueue, sidecar-down, 3MF
    embedded-settings fallback, STL no-fallback, and preset-error ->
    failed job paths.
  - New unit tests in test_slicer_api.py for the HTTP bridge.
  - 5 new SliceModal frontend tests covering preset gating, library +
    archive enqueue paths, error surface, and preset-load failure.
  - Existing SettingsPage tests adjusted: slicer dropdown asserts now
    switch to the Workflow tab first; added a beforeEach URL reset so
    one test's tab click doesn't bleed into sibling tests.

  Sidecar
  - New slicer-api/ folder is self-contained and optional. Two services
    (orca-slicer-api on 3003, bambu-studio-api on 3001 behind --profile
    bambu) build via Docker git-build-context from
    maziggy/orca-slicer-api@bambuddy/profile-resolver. The fork patches
    the OrcaSlicer CLI's profile compatibility quirks (inherits-chain
    resolver, from:User -> system rewrite, '# ' clone-prefix strip,
    sentinel-value strip) empirically required to slice real GUI
    exports without segfaulting the CLI.

  Docs
  - CHANGELOG entry under [0.2.4b1] - Unreleased Added.
  - README File Manager bullet for the new server-side Slice button.
  - bambuddy-website features.html: new card under "Configurable Slicer".
  - bambuddy-wiki: new page features/slicer-api.md + nav entry +
    features index card.

  Notes
  - Opt-in: with Use Slicer API off, the existing "open in desktop
    slicer via URI" flow is the default and unchanged.
  - 3MF inputs that segfault the CLI on --load-settings transparently
    retry with embedded settings; the resulting job carries
    used_embedded_settings: true.
  - Sliced files always export as .gcode.3mf so File Manager picks up
    the embedded thumbnail; file_type is set to "gcode" (blue badge).
2026-04-27 15:28:37 +02:00
maziggy 88b5f56eb2 fix: cancel = layer shift, stuck "1 problem", and dropped child-logger logs
Three bugs that surfaced together while debugging an H2D cancel:

  1. Cancelling a print stamped failure_reason="Layer shift" in archives
     AND left the printer card stuck on "1 problem" forever. Four causes:
     (a) POST /printers/{id}/print/stop never set the user-stopped flag, so
         on_print_complete couldn't override "failed" -> "cancelled".
     (b) HMS-derived failure_reason heuristic mapped any module-0x0C HMS to
         "Layer shift". Module 0x0C is "Motion Controller" broadly (includes
         cameras, markers, AND the cancel-sequence echo 0C00_001B). Real
         layer-shift codes live in module 0x03. Same false-positive class
         existed for "Filament runout" (any 0x07) and "Clogged nozzle" (any
         0x05). Replaced with a 23-code curated short-code map; unknowns
         leave failure_reason=None.
     (c) Cancel-echo HMS codes (0300_400C "The task was canceled.",
         0500_400E "Printing was cancelled.") were polluting state.hms_errors
         via both the hms[] and print_error parse paths. Filter them at
         parse time so the frontend never sees them.
     (d) Frontend bucketed gcode_state="FAILED" as a problem unconditionally.
         Real failures attach an HMS error; user-cancels don't — so FAILED-
         without-HMS now buckets as "finished" and only escalates to "error"
         when there's an active known HMS.

  2. logs/bambuddy.log was silently dropping records from named child
     loggers. TraceIDFilter was attached to root_logger, but Python's
     logging only invokes a Logger's filters on records originating at that
     logger — propagated child-logger records skipped it, formatter raised
     KeyError, handler.handleError dropped the record. Moved the filter
     from root_logger.addFilter() to handler.addFilter() on each handler,
     matching the filter's own docstring guidance.

  derive_failure_reason() extracted as a pure function for testability.
  status="cancelled" now symmetrically yields "User cancelled" alongside
  "aborted".

  20 regression tests across:
  - backend/tests/unit/test_failure_reason_derivation.py (11)
  - backend/tests/unit/services/test_bambu_mqtt.py::TestHMSUserActionFiltering (4)
  - backend/tests/unit/test_trace.py::TestFilterMustBeAttachedToHandlerNotLogger (1)
  - frontend/src/__tests__/pages/PrintersPageBucketing.test.ts (5; includes
    the H2D-cancel-echo "FAILED + only unknown HMS" case)
2026-04-26 14:24:28 +02:00
maziggy e9200449ae fix(deploy): kiosk picks up new builds without operator intervention
Reproduced live during the #1133 rollout: the SpoolBuddy display kept
  serving the pre-fix picker for hours after every cache-clear,
  chromium-restart, and pkill attempt because a chain of stale state
  across HTTP cache + Service Worker + persistent profile prevented
  fresh code from reaching the running tab.

  Three independent changes — any one of them sufficient on a clean
  profile, but all three needed to escape an already-corrupted one:

  (1) backend/app/main.py — index.html now served with
  Cache-Control: no-cache, must-revalidate on both / and the SPA
  catch-all. Vite emits content-hashed JS/CSS bundle filenames so the
  assets themselves are safe to cache forever, but the HTML wrapping
  them is the only file that knows which hash is current. Without
  explicit cache directives Chromium falls back to heuristic caching
  (typically 10% of time since Last-Modified) and on long-running
  kiosks happily serves stale HTML across browser restarts. That stale
  HTML references an old bundle hash which is also still in disk
  cache, so the kiosk runs pre-deploy JS forever without ever knowing
  why.

  (2) frontend/public/sw.js — CACHE_NAME bumped from bambuddy-v25 to
  bambuddy-v26 so any client that fetches the new sw.js drops its old
  CacheStorage. The SW does network-first for HTML/JS/CSS but
  intercepts and falls back to cache, and cache-control on HTTP
  responses doesn't reach into the SW's own cache layer.

  (3) spoolbuddy/install/install.sh — generated kiosk launcher now uses
  --user-data-dir=/tmp/spoolbuddy-kiosk-userdata with a pre-launch
  rm -rf, so every kiosk restart starts from a clean slate (no HTTP
  cache, no SW registration, no IndexedDB). Trade-off is a slightly
  slower first paint and zero offline support; neither matters for a
  single-purpose kiosk facing a backend on the same LAN, and the
  guarantee that next-deploy-just-works is worth far more.

  4 new tests in test_static_html_cache_headers.py: index.html on /
  and SPA catch-all paths emit Cache-Control: no-cache,
  must-revalidate; API routes are unaffected (no leak of HTML cache
  directive onto endpoints we want React Query to cache aggressively).

  For existing kiosks already trapped by an old persistent profile,
  operator runs once: rm -rf ~/.config/chromium && systemctl restart
  getty@tty1.service. The new launcher then picks up automatically.
2026-04-26 11:45:39 +02:00
maziggy 1878d2aab5 feat(observability): trace ID column on every log line + X-Trace-Id header
Builds on the recent uvicorn-access-log-into-bambuddy.log change.
  Until now the access line told us who called an endpoint, but there
  was no way to tie that line to the application records emitted on the
  server side while handling that request. The rogue stop_print mystery
  on 2026-04-26 left exactly that gap: even with access logs piped in,
  correlating "this POST landed" with "this MQTT publish went out 6 ms
  later" required eyeball-matching timestamps across different loggers.

  A new ContextVar + middleware + logging filter wire a trace ID through
  every record:

    * trace_id_middleware mints an 8-char hex ID per request (or honours
      a sane inbound X-Trace-Id for cross-system correlation), stores it
      in trace_id_var (ContextVar), echoes it on the response as
      X-Trace-Id, and resets the var in finally.
    * TraceIDFilter, attached to root + uvicorn.access, copies the
      current trace_id_var value onto every LogRecord so the format
      string [%(trace_id)s] resolves to the right ID per record.
    * Records emitted outside any request scope (startup, MQTT
      callbacks, scheduler) get a stable "-" placeholder so the column
      stays visually aligned and grep stays simple.

  ContextVars are the right plumbing because asyncio copies the current
  context into every asyncio.create_task, so background work spawned
  from inside a request inherits the same ID without explicit threading.
  request.state can't make that hop. The logging filter also has no
  access to the FastAPI request object — it runs synchronously inside
  the stdlib logging machinery — and the ContextVar is the only
  mechanism that bridges async request scope to sync log emission.

  Inbound X-Trace-Id is hard-validated against [A-Za-z0-9_-]+ (max 64
  chars) before being honoured — a hostile/buggy caller cannot smuggle
  log-injection payloads (newlines, control chars, megabyte blobs) into
  bambuddy.log via the trace ID column; values that fail the gate
  silently trigger a freshly minted server-side ID rather than failing
  the request.

  Middleware is decorated AFTER auth_middleware on purpose: Starlette
  stacks @app.middleware decorators LIFO so the last-decorated runs
  first inbound, making trace stamp the OUTERMOST layer — auth log
  lines and every record emitted on the way down to and back from the
  route handler all carry the same ID.

  Output now correlates as:

    2026-04-26 09:51:39,152 INFO [uvicorn.access] [a4f3b1e7] - "POST
      /api/v1/printers/1/print/stop HTTP/1.1" 200
    2026-04-26 09:51:39,158 INFO [bambu_mqtt] [a4f3b1e7] [SERIAL] Sent
      stop print command

  One grep a4f3b1e7 returns the full causality chain.

  30 new tests: 22 unit (ContextVar placeholder, filter copies value,
  asyncio task propagation, concurrent-request isolation, hex generator
  uniqueness, hostile-payload validator, max-length boundary, all four
  write verbs survive, GET/HEAD/OPTIONS dropped, URL-substring false-
  match guards, edge cases) and 8 integration (X-Trace-Id round-trips,
  body matches header, hostile inbound replaced, overlong inbound
  replaced, ContextVar resets after request, generator format stable,
  each request gets unique ID).
2026-04-26 10:01:17 +02:00
maziggy 352e619ad7 fix(inventory): serialise spool auto-assign per printer to fix Postgres race
Bambu MQTT can deliver two ams_data push frames for the same printer
  ~30 ms apart (observed on H2D + dual AMS at K-profile-load / RFID-read
  boundaries). Each frame triggers on_ams_change in main.py, whose
  auto-assign block reads (printer_id, ams_id, tray_id), decides "no
  existing assignment", and INSERTs via auto_assign_spool — and the two
  callbacks raced in their respective sessions, both deciding to insert,
  with the second commit losing on:

      asyncpg.exceptions.UniqueViolationError: duplicate key value
      violates unique constraint
      "spool_assignment_printer_id_ams_id_tray_id_key"
      DETAIL:  Key (printer_id, ams_id, tray_id)=(1, 0, 0) already exists.

  SQLite's WAL serial-write semantics had been silently swallowing the
  race for ~7 weeks since the spool-assignment feature shipped (latent in
  ec82092b "Sync", 2026-02-12). When optional Postgres support landed in
  610431d6 (2026-04-03) and asyncpg started allowing true concurrent
  transactions, it surfaced. Net impact: log noise + one assignment cycle
  skipped, retried on the next on_ams_change.

  Adds a per-printer asyncio.Lock (_ams_assignment_locks keyed by
  printer_id) wrapping the auto-assign critical section. By the time the
  second callback's session runs the SELECT, the first's commit is
  visible and the early-return "existing assignment" branch fires instead
  of a duplicate INSERT.

  The Spoolman sync block further down in on_ams_change intentionally
  stays OUTSIDE the lock — it's network-bound and idempotent, so
  serialising it would block subsequent AMS callbacks for the duration of
  a remote roundtrip. Per-printer scope keeps unrelated printers fully
  parallel. The auto-unlink block above isn't wrapped because its
  DELETE/UPDATE operations don't have the same constraint surface.

  5 new regression tests in test_ams_assignment_lock.py: same-printer-
  same-lock identity, different-printers-different-lock isolation, second
  acquirer waits for first (proves serialisation), different printers run
  truly in parallel under a held lock (proves per-printer scope), and an
  autouse fixture that resets the module-level dict between tests so
  cross-test loop affinity bugs can't surface.
2026-04-26 09:40:02 +02:00
maziggy 30cf384b5a fix: render Swagger UI at /docs with a docs-scoped CSP
The global CSP set script-src 'self', so FastAPI's /docs page rendered
  blank: the inline boot <script> and the cdn.jsdelivr.net swagger-ui
  bundle/CSS were both blocked. /redoc and /docs/oauth2-redirect had the
  same problem.

  Branch the security_headers_middleware to emit a docs-scoped CSP for
  those three paths that allows cdn.jsdelivr.net (scripts + styles), the
  FastAPI/Redoc favicon hosts (images), and 'unsafe-inline' for the
  inline boot script. Every other route keeps the stricter SPA policy
  unchanged.
2026-04-25 11:29:16 +02:00
maziggy 1e3ad697f2 fix(#1089): camera stream fan-out broadcaster
Most Bambu Lab printers only allow one concurrent camera connection, but
  GET /printers/{id}/camera/stream opened a fresh upstream per viewer.
  Two browser tabs → second viewer fails or kicks the first off.

  New MjpegBroadcaster (services/camera_fanout.py) owns one upstream per
  printer and fans MJPEG chunks out to N subscribers. 5 s grace window
  absorbs tab refreshes without reconnecting. Bounded subscriber queues
  drop frames for slow viewers rather than blocking the broadcaster.

  Audit-pass fixes:
  - _stream_start_times set with setdefault() so stream_uptime reflects
    the shared upstream's age, not the most-recent viewer's
  - subscribe() retried once on RuntimeError to close a tiny grace race
  - unsubscribe() returns post-removal count atomically so the detach log
    no longer races with concurrent leavers

  Permission gates unchanged; broadcaster has no FastAPI surface.

  Tests: 13 broadcaster unit tests + 2 integration tests on /camera/stop.
  External-camera path untouched.
2026-04-25 09:56:45 +02:00
maziggy 08601b4772 ● fix(#1111): advance queue item when print fails before reaching RUNNING
When a file sliced for the wrong nozzle size is dispatched, the printer
  goes IDLE -> PREPARE -> FAILED without ever entering RUNNING. Completion
  detection required prev=RUNNING or _was_running=True, so on_print_complete
  never fired and the queue item stayed at "printing" forever -- blocking
  every subsequent pending item for that printer (check_queue seeds
  busy_printers from any row in 'printing').

  Fire completion on FAILED from PREPARE or SLICING too. Restricted to
  those two pre-print states so a stale FAILED on first connection
  (prev=None) still can't accidentally advance an unrelated queue item.

  Also populate PrintQueueItem.error_message from the current HMS error
  list via the existing hms_errors.py lookup, so users see e.g.
  "[0500_4038] The nozzle diameter in sliced file is not consistent
  with the current nozzle setting" instead of a blank failure reason.
2026-04-24 16:02:55 +02:00
maziggy 9e938cbc8c Revert "feat(inventory): unified Spoolman inventory UI + Storage Location + AMS deep-link + SpoolBuddy NFC write support (#1063)"
This reverts commit 89f14c57ad.
2026-04-24 14:33:33 +02:00