Commit Graph
625 Commits
Author SHA1 Message Date
maziggy f243e4e598 fix(asyncio): track strong refs on orphan create_task sites
asyncio holds only a weak reference to tasks returned by
  ``create_task``. Fire-and-forget callers that discard the return value
  let the event loop GC the task before it finishes, logging
  ``Task was destroyed but it is pending!`` with no traceback. The #1648
  support-bundle review surfaced 94 such warnings in 8 days of v0.2.4.5
  -- the silently-vanished exceptions reach support bundles as opaque
  GC notices instead of actionable errors.

  New backend/app/core/tasks.py::spawn_background_task(coro, *, name=None)
  is the one place in the codebase that calls asyncio.create_task. It
  stores the task in a module-level set, attaches a done-callback that
  auto-removes on completion AND surfaces any uncaught exception via the
  logger with the originating traceback, and accepts name= so a leak
  source is traceable through /tracebacks and the log line. Cancelled
  tasks don't log (a shutting-down service is not an error).

  Migrated the 16 truly-orphan create_task call sites to the helper:

    main.py (8):
      reconcile-stale, cooldown-poweroff, energy calc, smart-plug,
      maintenance-check, photo-then-notify, layer-timelapse,
      scan-timelapse, print-scheduler, notify-no-archive (the last one
      was hand-rolling the same pattern with task + no-op done_callback)
    printers.py:3123        apply-pa-after-refresh
    print_queue.py:1034     queue cooldown-poweroff
    firmware_update.py:261  firmware upload
    archive.py:1514         timelapse mp4 convert
    print_scheduler.py:2199 watchdog print-start
    library.py:1614         STL backfill
    smart_plugs.py:259      tasmota scan
    discovery.py:159        subnet scan
    smart_plug_manager.py   x3 plug auto-off-pending
    background_dispatch.py  x2 (lambda-wrapped inside
                            loop.call_soon_threadsafe) upload progress

  Sites that already kept strong refs are unchanged:
    self._tasks.append(asyncio.create_task(...)) -- VP manager,
      tcp_proxy, mqtt_server
    self._x_task = asyncio.create_task(...) on service instances --
      mqtt_bridge, obico_detection, github_backup, archive_purge,
      local_backup, library_trash, discovery service
    Locally assigned + awaited/gathered -- tcp_proxy bidirectional
      pumps, camera_fanout, slice_dispatch, slicer_api progress_task,
      manager._finish_release_task, main.py module-level cleanup loops
2026-06-06 10:41:28 +02:00
maziggy d82f4e032f fix(inventory): handle PFCN cloud preset IDs in assign-via-MQTT (#1648)
Reporter on an H2D + Polymaker PLA Matte spool noticed that assigning
  the spool from the Dashboard left the slicer's filament dropdown
  showing "unknown", but clicking Configure right after made the
  slicer recognize it correctly. "Configure" felt like a mandatory
  follow-up step rather than a refinement.

  Bambu cloud uses three preset-ID shapes:
    GFS…   — Bambu official cloud preset
    PFUS…  — cloud user-created preset
    PFCN…  — cloud shared / partner preset (Polymaker's "(Custom)"
             Bambu Lab H2D variants ship this prefix)

  apply_spool_to_slot_via_mqtt only routed GFS and PFUS through the
  cloud-detail lookup that extracts the underlying filament_id. PFCN
  slipped past the cloud-lookup branch, fell into the local-preset
  int() parse path, raised ValueError, dropped into
  normalize_slicer_filament which returns any P-prefix unchanged, and
  the raw PFCN landed in tray_info_idx. The printer's calibration
  table can't index that, so the slicer rendered "unknown". The
  Configure modal rescued every assign because it does its own
  getCloudSettingDetail and writes the resolved filament_id.

  Extend the cloud-detail-lookup branch (inventory.py:129) and the
  discard safety net (inventory.py:223) to include PFCN alongside
  GFS/PFUS. Three behaviours fall out:

    * Cloud-authenticated: the real filament_id from
      detail["filament_id"] ships as tray_info_idx (Polymaker PLA
      Matte resolves to GFL05).
    * Cloud unavailable: raw PFCN discarded, the slot reuses an
      existing valid P-prefix preset if material matches.

  Source comment now lists all three cloud-ID shapes so the next time
  Bambu invents a new prefix the maintainer doesn't have to re-derive
  the structure from a bug report.
2026-06-06 10:17:09 +02:00
maziggy e895d8350a fix(vp): re-fire FINISH after project_file ack so slicer releases (#1658)
Reporter on Bambu Studio 2.7.1.57 + X1C saw the Send modal stuck at
  "Downloading" after sending to a Queue-mode VP. Delete-from-queue and
  even Auto-Dispatch ON + a successful real print didn't release it.

  Root cause: BS 2.7.x flipped the Send sequence from
    MQTT project_file -> FTP upload -> done
  to
    FTP verify_job -> FTP .3mf -> MQTT project_file.

  The #1280 fix sets gcode_state=FINISH in on_file_received (after the
  FTP upload). Under the new order, the synthetic project_file ack in
  _send_print_response then runs and overwrites _gcode_state back to
  PREPARE. The 1 Hz cached-as-base push stream carries PREPARE forever,
  the slicer never sees the FINISH transition it waits for, and the
  modal sits stuck. Auto-Dispatch ON shares the cause: the real
  printer's PREPARE->RUNNING->FINISH on the bridge gets masked by the
  local _gcode_state override in _send_status_report.

  Re-fire set_gcode_state("FINISH", filename, prepare_percent="100")
  1.5 s after the project_file ack for every non-proxy mode (queue /
  archive / review). The 1.5 s window lets the slicer see at least one
  PREPARE push on the 1 Hz cycle so the transition reads as
  PREPARE -> FINISH, matching what the slicer expects. Proxy mode is
  exempt -- there the real printer drives the bridge state and a
  synthetic FINISH would clobber a real PREPARE/RUNNING transition.

  The scheduler cancels any in-flight timer when a new project_file
  arrives so a retrying slicer doesn't end with two competing FINISH
  timers. The pending timer is also cancelled on stop_server.
2026-06-06 08:51:09 +02:00
maziggy 19073f5c84 fix(vp): auto-derive access code from target printer in non-proxy modes
Non-proxy VPs (Archive / Review / Queue) with a target printer set up
  a live-mirror bridge that forwards the slicer's MQTT and RTSPS auth
  bytes through to the real printer. The slicer holds one code in its
  profile (the one it bound the VP with), and that code has to satisfy
  both the VP listener and the real printer at the far end of the
  bridge. If the codes diverge the bridge silently fails at the second
  hop — slicer reaches .49:8883, FINs before sending a ClientHello,
  retries identically. The wiki framed the code-match requirement as a
  camera-only concern; it isn't, all bridged protocols inherit.

  Fix removes the foot-gun instead of re-documenting it. When a target
  is selected on a non-proxy VP the access-code field switches to a
  read-only display showing the target's code with an Eye-toggle
  reveal; the backend auto-inherits on every create / update (any
  explicit access_code submitted alongside a target is silently
  overridden as belt-and-braces for non-UI clients). The required-when-
  enabling check now treats target-set as satisfying the access-code
  requirement. Standalone (no-target) non-proxy VPs still get the
  editable input + Save button.

  One-shot startup migration corrects any pre-existing mismatched
  rows: SELECTs diverged VPs and logs one INFO line per row for the
  audit trail, then UPDATEs via correlated subquery. Idempotent and
  portable between SQLite and Postgres.
2026-06-05 17:19:16 +02:00
maziggy 12d17bfbe7 fix(photo): source finish photo from forced timelapse + cleanup (#1397)
Bambu's end-gcode lowers the bed at gcode_state=FINISH. Bambuddy's
  live-camera grab captured the bed already dropped, ruining the photo
  framing. Source the photo from a brief Bambu timelapse instead —
  firmware stops timelapse recording AFTER toolhead parks but BEFORE
  bed-drop runs, so the last frame frames the finished print correctly.

  When capture_finish_photo is on AND the user did not opt in to
  timelapse for this print, force timelapse=True at dispatch + mark the
  new PrintArchive.bambuddy_forced_timelapse column. After extraction
  (success or failure), cleanup deletes the locally-attached file,
  clears archive.timelapse_path, and walks the four scanner directories
  (/timelapse, /timelapse/video, /record, /recording) trying FTP DELE
  against the original filename. User-opted-in timelapses pass through
  unchanged.

  Resolver lives at services/background_dispatch.py::resolve_effective_timelapse
  (module-level so the print queue can reuse it). Both dispatch paths
  wired: background_dispatch.py (Print Now / Reprint) AND
  print_scheduler.py:_start_print (the queue). Field testing caught the
  scheduler gap on the first round — AST regression test now asserts
  start_print(timelapse=...) references effective_timelapse, not the raw
  item.timelapse, so a future refactor can't silently drop it.

  Extractor: ffmpeg -i input.mp4 -update 1 -q:v 2 out.jpg. Decoded
  frames overwrite the same output file, so the file left on disk is the
  literal last frame regardless of duration. Bambu records one frame per
  layer-change, so a 16-layer cube produces a 0.6 s timelapse — the
  original -sseof -1.0 approach seeked before the start of the file and
  returned frame 0 (empty bed). Decoding every frame is fine; Bambu
  timelapses are short by construction even on hours-long prints.

  Migration adds bambuddy_forced_timelapse branched on is_sqlite()
  (DEFAULT 0 / DEFAULT FALSE — PG rejects DEFAULT 0 for BOOLEAN).
  Verified live on postgres:16-alpine.

  Photo-task wait_for budget extends 45s -> 75s when timelapse_was_active
  so the notification carries the bed-up photo instead of falling back
  to the live-cam grab on slow links.

  Scope limit, documented in the camera wiki: prints started directly
  on the printer touchscreen / Bambu Handy / Bambu Studio Send bypass
  both dispatch paths, so the override doesn't fire there. Future
  option: mid-print M981 S1 P20000 MQTT toggle in on_print_start.

  Setting description rewritten in all 11 locales to drop the "only
  works when timelapse enabled" caveat (Bambuddy now forces it) and
  explain the kept-or-deleted behaviour.
2026-06-05 13:53:26 +02:00
maziggy 4851f54595 feat(file-manager): split "All Files" into internal-only + External views (#1621)
Reporter linked a NAS and the auto-imported files drowned their own
  Bambuddy uploads in the "All Files" sidebar listing. There was no filter
  to escape it — only per-folder clicks. Restore the pre-external semantics:
  "All Files" now lists managed-storage files only. The combined
  across-every-external view moves to a new sibling sidebar entry,
  "External", that only appears when at least one external folder is linked.

  Backend: GET /api/v1/library/files gains internal_only and external_only
  query flags. Filter is on LibraryFile.is_external. Both flags set is a
  400, not a silent pick-one.

  Frontend: new topLevelView state on FileManagerPage (default internal);
  the query passes the scope only when selectedFolderId is null. Mobile
  selector dropdown uses __top:internal / __top:external sentinels so the
  same state round-trips through option values. Empty-state copy
  distinguishes internal-empty from external-empty.
2026-06-05 10:11:03 +02:00
maziggy aed01f875a fix(vp): resolve hostname/FQDN targets in MQTT bridge IP encoding (#1429 follow-up)
Printers added to Bambuddy by hostname/FQDN (e.g. p1s.fritz.box) hit
  'invalid IPv4' in _ip_to_uint32_le, so the net.info[*].ip rewrite never
  armed and BambuStudio Send went straight to the real printer instead of
  the Bambuddy archive whenever the printer was powered on.

  Add _resolve_target_to_ipv4(target): IPv4 pass-through, else
  socket.getaddrinfo(target, family=AF_INET). AF_INET filter is load-bearing
  because net.info[*].ip is uint32 LE and IPv6 can't round-trip. OSError
  returns None so a transient DNS failure recovers on the next 30s refresh
  tick via the existing not-armed throttle.

  Apply the resolver to both the encode call and the host-interface picker
  (which also assumes dotted-quad). Armed log line now carries
  configured->resolved when they differ, so bad-DNS regressions stay legible
  in 'docker logs'. The unresolvable not-armed reason now names the
  configured value rather than parroting 'invalid IPv4', distinguishing
  'DNS gave a v6 result' from 'user typed garbage'.

  Root-caused by @Mape6; @TrickShotMLG02 confirmed the FQDN workaround
  on the same release. Pre-0.2.4 these setups worked by accident because
  there was no net.info[].ip rewrite at all.
2026-06-05 09:50:59 +02:00
maziggy 54389a54aa fix(library): MakerWorld URL import honours external folder destinations (#1645)
Reported and root-caused by @needo37. Importing a model via the MakerWorld
  URL-download feature into a writable external folder (e.g. SMB/NFS-mounted
  NAS) saved the 3MF into Bambuddy's internal managed library dir, not the
  external mount. The file card showed in the File Manager under the
  external folder, but the bytes never landed on the NAS, and the on-disk
  copy was UUID-renamed so a find by the original basename matched nothing.

  Root cause was save_3mf_bytes_to_library at backend/app/api/routes/library.py:422:
  it accepted folder_id but never loaded the folder, never inspected
  is_external / external_path, hardcoded the destination to
  get_library_files_dir() with a UUID name, and left the LibraryFile row
  with is_external=False. So the row's folder_id pointed at the external
  folder while its bytes and is_external flag both said "managed/internal".
  Same class of bug as #1112, which had been fixed for the multipart-upload
  and move paths but never applied to this byte-import path.

  Fix mirrors the multipart-upload path directly:
  - Load target_folder from folder_id when non-None.
  - Feed it to _resolve_upload_destination(target_folder, filename), which
    already returns (dest, is_external) and enforces the 403-read-only /
    400-unwritable-or-missing / 409-collision rejections.
  - Write bytes to dest (real filename for external, UUID for managed).
  - Persist the row with file_path=_stored_file_path(dest, is_external)
    and is_external=is_external.

  The route-layer read-only guard at makerworld.py:256-260 is preserved -
  it returns the friendlier error before the upstream download burns
  bandwidth - and _resolve_upload_destination's identical check stays as
  defence-in-depth for any future caller that skips the route gate.
  Thumbnails continue to live under the managed get_library_thumbnails_dir()
  regardless of the 3MF's location, matching the upload path.
2026-06-05 09:29:53 +02:00
maziggy 9554ebd05d refactor(inventory): rename /reset-usage to /reset-consumed-counter to match what it actually does (issue #1644)
The old endpoint name implied that calling it would drop weight_used to
  0. In practice it only stamps weight_used_baseline = weight_used so the
  Inventory page's "Total Consumed" widget (weight_used - baseline) reads
  0 going forward, while remaining (label_weight - weight_used) is
  preserved. Calling the endpoint via curl and seeing weight_used
  unchanged in the JSON response is confusing.

  New paths:
  - internal: /api/v1/inventory/spools/{id}/reset-consumed-counter
             /api/v1/inventory/spools/reset-consumed-counter-bulk
  - spoolman: /api/v1/spoolman/inventory/spools/{id}/reset-consumed-counter
             /api/v1/spoolman/inventory/spools/reset-consumed-counter-bulk

  Behaviour is unchanged in both modes; internal stamps the baseline
  directly, Spoolman-mode PATCHes upstream used_weight=0 and the
  _map_spoolman_spool read mapping reconstructs the same "displayed
  consumed = 0, remaining unchanged" Bambuddy-visible shape. Parity
  between modes was already in place and is preserved.

  The Spoolman-client method reset_spool_usage keeps its name because it
  describes what is sent upstream to Spoolman, not what Bambuddy's
  endpoint promises to callers.

  Frontend:
  - api.resetSpoolUsage / bulkResetSpoolUsage (and Spoolman variants)
    renamed to resetSpoolConsumedCounter / bulkResetSpoolConsumedCounter.
  - Button labels: "Reset usage to 0" -> "Reset counter" / "Reset all
    counters" (short, unambiguous); tooltips and confirm-modal bodies
    still spell out the full semantics.
2026-06-05 09:22:05 +02:00
maziggy e802bfc806 fix(ftp): cap TLS to v1.2 for X2D FTPS to dodge WRONG_VERSION_NUMBER on firmware 01.01.00.00 (#1638)
Reporter @vasmarfas saw X2D archive cards land almost empty - only print
  time, no filament weight / layers / MakerWorld link / thumbnail - and
  Spoolman filament-usage tracking went silent on the same printer.

  Support bundle traces the end-to-end: at print start
  backend/app/main.py::on_print_start tries the usual FTP-download dance
  for the 3MF, every implicit-FTPS connect attempt to the X2D fails with
  `[SSL: WRONG_VERSION_NUMBER] wrong version number (_ssl.c:1032)`, ~2
  minutes later "Could not find 3MF file for print" -> "Created fallback
  archive". Fallback path writes file_path="", file_size=0,
  content_hash=NULL, no layers / filament / model-link fields. Spoolman
  tracking degrades from the same root cause - both depend on the 3MF
  metadata parser.

  Proximate cause: Python 3.13's default ssl.create_default_context()
  negotiates TLS 1.3, the X2D's implicit-FTPS server on port 990 rejects
  the ClientHello. Same family as the P2S 01.02.00.00 bug from #1401
  (post-Python-3.13 TLS-1.3 breakage), different wire-level failure mode
  (P2S completes the handshake and truncates with 426; X2D fails the
  handshake outright).

  Same fix shape: add X2D to backend/app/services/ftp_profiles.py with
  cap_tls_v1_2=True, plus N6 -> X2D SSDP alias. Every other model stays
  on negotiated TLS 1.3.

  Honest caveat: hypothesis-driven trial, not a confirmed root-cause fix.
  WRONG_VERSION_NUMBER could equally describe the X2D switching to
  explicit FTPS (AUTH TLS on plaintext greeting) or moving FTPS to a
  different port - either would need a different code path. Reporter has
  been asked to test this build; if the cap doesn't clear it the registry
  slot stays useful and the next diagnostic round goes to openssl
  s_client from a network-adjacent host.
2026-06-05 08:30:19 +02:00
maziggy 18d534c945 feat(orca-cloud): integrate Orca Cloud profile sync across UI, slicer and SpoolBuddy
Reads, lists, and slices with profiles from a user's Orca Cloud account
  (OrcaSlicer 2.4.0-alpha's Supabase-backed sync) alongside the existing
  Bambu Cloud integration. Four sign-in providers (Google / Apple / GitHub /
  email+password); password defaults. Paste-flow PKCE because Orca's
  Supabase project only allowlists localhost redirect_to — open feature
  request at OrcaSlicer/OrcaSlicer#14028.

  Surfaces:
  - Profiles tab: new "Orca Cloud" tab next to "Bambu Cloud" with the same
    rich layout (search + 5 filter dropdowns + 3-column grouped grid +
    read-only detail modal)
  - SliceModal: 4-tier preset picker (orca_cloud > local > bambu cloud >
    standard); separate status banner per cloud; metadata-aware pre-pick
    scores Orca filaments above local (Orca's sync_pull returns full
    content inline so filament_type / filament_colour come for free, no
    per-setting fetch rate-limit dance)
  - ConfigureAmsSlotModal: orca_cloud as a new preset source (prefixed
    orca_<UUID> to match local_/builtin_); generic Bambu filament-ID
    derivation from parsed material (printer firmware can't grok Orca
    UUIDs); slot mapping persists preset_source='orca_cloud'
  - SpoolForm / SpoolBuddyWriteTagPage: Orca filaments merge into the
    cloud preset list via Promise.allSettled (OrcaProfileMeta is
    structurally identical to SlicerSetting)

  Backend:
  - services/orca_cloud.py: OrcaCloudService with PKCE / token exchange /
    single-use refresh rotation / get_user_info / list_profiles via the
    bare /sync/pull bootstrap path
  - routes/orca_cloud.py: 7 endpoints (auth/start, auth/finish,
    auth/password, status, logout, profiles, profiles/{id}); router-level
    _cloud_api_key_gate + per-route cloud_caller() so API-keyed callers
    (SpoolBuddy kiosk) properly resolve their owner User; just-in-time
    refresh with atomic persist-before-API-call
  - routes/slicer_presets.py: _fetch_orca_cloud_presets mirrors the Bambu
    Cloud fetcher (status vocabulary, 5min cache, permission shortcut);
    _dedupe_by_name extended to 4 tiers; UnifiedPresetsResponse gains
    orca_cloud + orca_cloud_status
  - services/preset_resolver.py: PresetRef.source extended with
    "orca_cloud"; _resolve_orca_cloud walks list + filters
  - 8 columns on users table for tokens (5 persistent) + transient PKCE
    handshake state with 10-min TTL (3); dialect-branched DATETIME /
    TIMESTAMP; auth-disabled mode falls back to Settings table
  - orca_cloud:auth permission folded into can_access_cloud API-key scope
    (same trust dimension)
2026-06-04 15:50:44 +02:00
maziggy 51730a7bf1 feat(ams): gate humidity/temperature alarms on AMS-has-filament (#1619)
The hourly AMS sensor recorder dispatched humidity and temperature alarms
  for every unit above threshold without checking whether the unit was
  actually loaded. Empty AMS units still report ambient readings, so users
  with one loaded + one empty AMS got useful alarms for the loaded one and
  hourly noise for the empty one. Disabling the whole alarm category killed
  both — not a real choice.

  New _ams_has_filament helper inspects tray_exist_bits (hex bitmap, "0" =
  empty) with fallback to the tray array's tray_type strings for shapes
  where the bitmap is missing. The recorder gates the alarm dispatch on
  this check per-AMS-unit, so a multi-AMS printer with one loaded + one
  empty still alarms on the loaded one.

  Sensor history still records regardless of the gate so the System page
  humidity charts stay continuous — only the outbound notification is
  suppressed. 9 unit tests cover the bitmap-zero case, bitmap-missing
  fallback, garbage/blank/int bitmap edges, and defensive malformed tray.
2026-06-04 11:30:26 +02:00
maziggy db1c664fce feat(vp): surface why MQTT bridge IP encoding didn't arm (#1429 defensive)
_refresh_ip_encoding had 4 silent early-returns. When the rewrite
  silently no-op'd on a user's setup, the only signal was the absence of
  the "armed" INFO line, and diagnosing which path was firing meant
  grepping the source.

  Each path now emits one INFO line naming the specific reason. A
  _not_armed_reason dedup field throttles to one line per state change,
  so an idle unarmed bridge doesn't spam every 30s refresh tick. Cleared
  on successful arm so regressions re-emit.

  Not a fix for #1429 itself — the bridge logic is unchanged; this just
  turns the silent failure into visible signal so the next "fix didn't
  work for me" report can be triaged in one round-trip.
2026-06-04 11:24:20 +02:00
maziggy 51abc4b7a8 feat(drying): enable AMS drying for H2C at firmware 01.02.00.00+ (issue #1624)
H2C was in _DRYING_UNSUPPORTED_MODELS. Move to _DRYING_MIN_FIRMWARE
  with the same 01.02.00.00 floor as H2S / P2S. Both SSDP model codes
  the H2C advertises (O1C single-nozzle, O1C2 dual-nozzle) get the
  same gate so supports_drying() fires correctly regardless of which
  form is stored on the printer record.
2026-06-04 10:57:46 +02:00
maziggy c571ad86dd feat(diagnostic): printer_publishing check + countdown UI (#1622)
The existing connection diagnostic proved TCP + TLS + auth + SUBSCRIBE but
  not that the printer was actually publishing reports. A wrong-cased serial
  passes mqtt_auth because the broker accepts the subscription regardless;
  the user-visible symptom is empty AMS / no K-profiles / no custom filaments
  in the slicer Device tab because the VP cached state is empty. Bambuddy
  already logged the actionable hint at bambu_mqtt.py:498 but only to
  container logs.

  New printer_publishing check turns that warning into a structured
  diagnostic result. Pass = bridge has seen at least one report since the
  last (re)connect; fail = zero reports across the wait window with fix-text
  pointing at the case-sensitive serial. Bounded 10s poll on the on-demand
  UI route, no wait on the support-package gathering path so bundling stays
  fast. Exits the moment a message arrives — typical wall-clock is 1-2s.

  Frontend renders an elapsed-seconds counter plus a "Listening for status
  report — up to 10s" hint during the pending state so the wait doesn't look
  hung. PUBLISH_WAIT_DEFAULT_SECONDS pinned on both sides.

  report_messages_since_connect exposed as a public property on
  BambuMQTTClient so the diagnostic doesn't reach into private state.
2026-06-04 10:54:04 +02:00
maziggy a1cb5d5b4d fix(backup): interpret scheduled-backup HH:MM as local time, not UTC (#1602 follow-up)
The Scheduled Local Backups time-of-day picker was interpreted as UTC by
  _calculate_next_run, so a UTC+3 user had to enter 18:00 to get a 21:00
  local backup. The UI labeled the field "UTC" but it was still surprising.

  Picker is now interpreted in the container's local timezone, resolved
  from the TZ env var via zoneinfo.ZoneInfo (same source the Support page's
  environment.timezone shows). UTC fallback when TZ is unset or
  unrecognised. The /local-backup/status endpoint exposes the resolved
  zone, and the UI renders it next to the field via a new
  backup.localTimeHint i18n key with real translations in all 10
  non-English locales.

  One-time behaviour change for users who entered a UTC time as a
  workaround: the first scheduled cycle after upgrade will run at their
  local TZ offset earlier than expected. Re-enter the time as local once
  and it is correct from then on. No migration is shipped; migrating
  around a DST boundary would be ambiguous.
2026-06-04 10:17:10 +02:00
maziggy cc25cbe774 fix(archives): #1608 suppress card time-accuracy badge for multi-run archives
compute_time_accuracy in routes/archives.py compares the archive row's
  own started_at / completed_at (which reflect the latest run only)
  against archive.print_time_seconds (which the #1593 parser fix
  correctly stores as the sum across plates). For a 3-plate file printed
  plate-by-plate the ratio is ~300%, producing a "+188%" card badge that
  means nothing — apples to oranges. The 5-500% sanity band catches
  truly broken values but lets this deterministic N×100% shape through.
  Reporter's archive #65 was 3 plates over 9 runs.

  compute_time_accuracy gains an optional run_aggregate argument and
  returns both actual_time_seconds and time_accuracy as null when the
  aggregate reports more than one logged run. The frontend already falls
  through to print_time_seconds for the time display
  (actual_time_seconds || print_time_seconds) and gates the badge on
  time_accuracy being truthy, so multi-run archives now show the slicer
  estimate with no badge. Single-run archives keep the original
  behaviour verbatim.

  The fix is applied at every call site that renders an archive card:
  archive_to_response now threads run_aggregate through, and the three
  endpoints that previously didn't load the aggregate (archives.py
  search fast-path and FTS path, single-archive PATCH, and
  projects.list_project_archives) now batch-load it via the existing
  _load_run_aggregates helper.

  The stats endpoint's per-run accuracy aggregation at archives.py:940
  already uses PrintLogEntry.duration_seconds with its own 50-200% band
  filter and is untouched.
2026-06-03 10:14:57 +02:00
maziggy 9cc4b6aa60 fix(virtual-printer): #1610 add Bambu cipher pin to every slicer-facing TLS context
The #620 patch fixed the OpenSSL-3.x-strips-plain-RSA-AES-GCM cipher
  mismatch on the printer-facing TLSProxy client context. The same fix
  was never applied to the four other slicer-facing TLS contexts. On
  hardened distros (Fedora / RHEL with update-crypto-policies, hardened
  Alpine builds) where the system narrows DEFAULT to forward-secrecy
  only, the slicer's ClientHello finds no overlap with what Bambuddy
  offers and the handshake aborts with the slicer reporting code=-1
  before any application data flows. The reporter pinpointed the missing
  set_ciphers call in bind_server.py against the #620 lineage; the
  audit-wide sweep here extends the same fix to mqtt_server.py,
  tcp_proxy._create_server_ssl_context (the missing other half of #620),
  and ftp_server.py.

  For the three new contexts (bind / mqtt / proxy-server) the cipher
  string is DEFAULT:AES256-GCM-SHA384:AES128-GCM-SHA256 — verbatim match
  with the #620 client-side fix. For FTPS the original HIGH baseline is
  kept (HIGH:AES256-GCM-SHA384:AES128-GCM-SHA256:!aNULL:!MD5:!RC4) so the
  cipher set stays a strict superset of what shipped before — HIGH
  offers ~58 suites DEFAULT doesn't (CCM / ARIA / CAMELLIA / DSS) that
  no Bambu slicer is known to pick, but narrowing a compat surface
  without proof would violate the existing don't-remove-compat-pinning
  rule. TLS version pins (TLSv1_2 minimum across all four, TLSv1_2 max
  on FTPS for the BambuStudio PSK-reuse compat) and verify-mode settings
  are unchanged — only the cipher list is widened.
2026-06-03 10:04:37 +02:00
maziggy fee7c722f5 fix(usage-tracker): #1607 filter empty AMS slots from position-based fallback
When no explicit slot-to-tray mapping is captured (path 5 of 6 in
  _track_from_3mf — fires before the request-topic subscription that catches
  ams_mapping is accepted), the tracker builds available_trays from
  build_ams_tray_lookup and uses position to map the slicer's Nth filament
  to the Nth available tray. The helper enumerated every AMS tray by id
  regardless of whether a spool was loaded, so AMS slots 0-2 loaded + slot 3
  empty + external yielded available_trays = [0, 1, 2, 3, 254]. The slicer
  compacts its filament UI to hide empty AMS slots, so its 4th filament is
  the external — but position mapping routed it to AMS0-T3 (the empty slot)
  instead of 254 (external). No spool assigned there → usage silently
  skipped → external never decremented.

  Filter the fallback to slots with a non-empty tray_type. build_ams_tray_lookup
  stays unchanged for its other callers (spoolman_tracking.store_print_data,
  routes/printers, spool_assignment_notifications); the filter is applied at
  the usage-tracker call site only. Mirrors the existing vt_tray filter in
  build_ams_tray_lookup line 174.
2026-06-03 09:45:10 +02:00
maziggy cdc27eb517 fix(virtual-printer): #1429 follow-up — auto-resolve VP IP when bind_address is 0.0.0.0
The original #1429 fix's _refresh_ip_encoding early-returned when
  mqtt_server.bind_address was "0.0.0.0" or empty (the default for VPs created
  without a dedicated bind IP). On a flat-LAN install that's the typical case,
  so the encoding never armed, _rewrite_net_info_ips was a no-op on every push,
  and the slicer kept following the real printer IP to its SD card. @Mape6
  reported this on the 2026-06-02 daily that supposedly fixed the bug.

  New helper _resolve_host_interface_for_target() consults the existing
  network_utils.find_interface_for_ip() to pick the host interface in the
  printer's subnet. _refresh_ip_encoding falls back to it when bind_address
  is unspecified; an explicit bind IP still wins. INFO log line distinguishes
  the two paths ("armed: ... (bind_address)" vs "(auto-resolved)") so future
  bundles directly answer which IP the rewrite picked.

  Tests: 4 new under TestBindAddressAutoResolve — rewrite arms via auto-resolved
  IP at bind_address=0.0.0.0; stays disabled if no interface matches (no crash);
  explicit bind_ip still takes precedence; helper returns None defensively when
  find_interface_for_ip does.
2026-06-03 09:07:57 +02:00
maziggy 673001e3cd fix(maintenance): #1596 persist wiki_url on Custom Type create + correct
inline mutation type

  POST /api/v1/maintenance/types hard-coded the MaintenanceType
  constructor and silently dropped `wiki_url`, so the Documentation URL
  field disappeared after save. PATCH worked because it uses
  `data.model_dump(exclude_unset=True) + setattr`, which is why editing
  a freshly-created type DID save the URL — masking the bug under any
  "save then immediately fix it" retest. Reporter @BurntOutHylian
  pre-triaged the issue to the exact constructor call at
  routes/maintenance.py:206-213; fix is the missing `wiki_url=data.wiki_url`
  argument.

  Frontend nit from the same report: MaintenancePage.tsx:1131's
  `updateTypeMutation` declared `data: Partial<{ name; default_interval_hours;
  interval_type; icon }>` — omitting `wiki_url`. The value reached the
  API correctly at runtime because `api.updateMaintenanceType` accepts
  `Partial<MaintenanceTypeCreate>` (which has wiki_url), but the inline
  type lied about the payload shape. Extended the inline `Partial<{...}>`
  to include `wiki_url?: string | null`. Pure type fix — no runtime change.
2026-06-02 14:58:10 +02:00
maziggy a837a3acbd ● fix(library): #1600 thumbnail extraction for external-folder .gcode.3mf
files + unify file_type classification across ingest paths

  #1600: external-folder sliced outputs landed
  with no thumbnail. Cause: four backend ingest paths classified
  LibraryFile.file_type differently for the same .gcode.3mf family.
  upload / ZIP-extract / in-process used os.path.splitext()[1] which
  returns .3mf for foo.gcode.3mf and stored file_type="3mf", matching
  the thumbnail-extraction gate at library.py:1467 (file_type == "3mf").
  External-folder scan explicitly detected the compound and stored
  file_type="gcode.3mf" — preserving "sliced output" identity — but
  then skipped both the "3mf" gate and the "gcode" gate, so the file
  landed with thumbnail_path = None. Same compound-extension drift that
  bit #1543 in the 3D preview, in a surface that audit didn't trace
  back to.

  Unified fix:

  - New classify_file_type(filename) helper in routes/library.py is the
    single source of truth. Returns "gcode.3mf" for sliced outputs and
    ext[1:] otherwise.
  - Applied to every ingest path: upload (line 1704), ZIP-extract
    (1998), external-folder scan (the bug site — the manual compound
    check is replaced), and in-process save_3mf_from_bytes (471, used
    by MakerWorld import).
  - External-scan thumbnail gate widened to
    `if file_type in ("3mf", "gcode.3mf"):` — a .gcode.3mf IS a 3MF zip
    with Metadata/plate_1.png; ThreeMFParser doesn't care about the
    trailing extension.
  - gcode-download endpoint at GET /library/files/{id}/gcode had the
    same drift in reverse: gate was `elif file.file_type == "3mf":` so
    a row stored with file_type="gcode.3mf" (the external-scan path's
    pre-unification behaviour, and the canonical going forward) got
    rejected with HTTP 400. Widened to the same compound-aware tuple.

  One-shot DB migration in core/database.py::run_migrations backfills
  existing legacy rows:

    UPDATE library_files
       SET file_type = 'gcode.3mf'
     WHERE file_type = '3mf'
       AND LOWER(filename) LIKE '%.gcode.3mf'

  Idempotent (post-update rows no longer match the file_type='3mf'
  predicate, so re-runs at every boot are no-ops) and dialect-neutral
  (LOWER + LIKE are identical under SQLite and Postgres). Without the
  backfill, users would have a permanent split state: old uploads at
  '3mf', new uploads at 'gcode.3mf' — which would double-bucket sliced
  outputs in the dashboard stats query at line 4615 and show two
  entries in the file-manager filter dropdown for the same conceptual
  type.

  Frontend untouched. FileManagerPage.tsx and ProjectDetailPage.tsx
  already accept both '3mf' and 'gcode.3mf' per the #1543 fix. After
  the migration the DB only contains canonical values, so the legacy
  '3mf' branches in the frontend become dead code for sliced files —
  they stay as defence-in-depth in case any future ingest path I
  missed reverts to the legacy classifier.
2026-06-02 14:51:08 +02:00
maziggy 5d6d928b3f fix(virtual-printer): #1429 net.info[*].ip cache leak + mode wire-value rename
#1429 (reported by @TrickShotMLG02, confirmed by @Mape6 on a flat single-LAN
  that rules out subnet / mDNS-reflector theories): with the physical printer
  off the slicer's "Send" landed in Bambuddy's archive; once the printer
  powered on every subsequent "Send" went straight to the printer's SD card
  and bypassed Bambuddy. Bundle analysis: mape6-before showed clean FTP
  receive + archive lines, mape6-after had zero FTP attempts to Bambuddy
  once the printer was online.

  Cause: mqtt_bridge.py::_resolve_client encoded _target_ip_uint32_le /
  _vp_ip_uint32_le ONLY on client-identity change and early-returned on
  every refresh tick when the same client object was still bound. If
  target_client.ip_address was empty at first bind (DB row stale, or client
  constructed before SSDP refresh filled it in), the encoding stayed None,
  the net.info[*].ip rewrite block was skipped, the cache filled with the
  real printer IP, sticky-key preservation kept the poisoned net value
  alive across every subsequent incremental push, and the slicer followed
  the leaked IP. Only Bambuddy-restart-with-printer-off cleared it — the
  workaround both reporters independently arrived at. Same shape on
  multi-NIC printers (X1C, H2D Pro): the rewrite only matched entries
  whose ip equalled _target_ip_uint32_le, so a secondary interface IP
  Bambuddy never saw would leak through unchanged.

  Bridge fix:
  - _resolve_client calls a new _refresh_ip_encoding() on every refresh
    tick, even when client identity is unchanged; self-heals once
    ip_address becomes valid.
  - _refresh_ip_encoding() sweeps the existing _latest_print_state when
    encoding becomes valid for the first time. Without the sweep,
    sticky-key preservation keeps the pre-arm poisoned cache alive
    forever — incremental pushes that don't include net carry the bad
    value forward.
  - _rewrite_net_info_ips() rewrites EVERY non-zero net.info[].ip entry
    that doesn't already equal the VP IP, not just entries matching
    _target_ip_uint32_le. Multi-NIC printers stop leaking secondary
    interfaces. Zero-IP placeholders are left alone so "active interface"
    detection still works.
  - INFO logging on encoding arm/update and on cache sweep so future
    bundles directly answer "did the rewrite fire?".

  Mode wire-value rename (#1429 follow-up, separate confusion source):
  - Both reporters' support bundles showed mode: immediate while the UI
    said "Archive"; @TrickShotMLG02 quoted: "I have no idea why it says
    immediate in the support-info.json file. In the webui the printer is
    set to archive". UI button "Archive" had always saved immediate, and
    "Queue" had always saved print_queue. Canonical wire values are now
    archive / review / queue / proxy, matching the button labels 1:1.
  - New normalize_vp_mode() + VP_MODE_* constants in
    models/virtual_printer.py; manager.py normalises on construction so
    a legacy row read pre-migration still dispatches correctly.
  - core/database.py::run_migrations rewrites existing virtual_printers
    and settings rows; idempotent (re-runs are no-ops); identical SQL
    under SQLite and Postgres.
  - API routes accept both legacy and canonical on input, normalise
    before storage. GET /settings/virtual-printer normalises on read so
    the frontend's mode-button highlight works for stale legacy values.
  - Three frontend VP components (VirtualPrinterSettings,
    VirtualPrinterCard, VirtualPrinterAddDialog) switched click handlers
    and type aliases to canonical; each got its own normalizeMode()
    helper so a stale-cached settings payload still highlights the right
    button. Two pre-existing `printer.mode === 'queue' ? 'review'`
    legacy mappings in VirtualPrinterCard were the source of a test
    failure caught mid-implementation where the new canonical 'queue'
    was being mis-aliased back to 'review' and hiding the auto-dispatch
    + force-color-match toggles.

  mode handler is NOT the dispatch bug: manager.py::_archive_file (the
  handler for archive mode) doesn't dispatch to the physical printer.
  The "files end up on the printer's SD card" symptom was the IP-leak
  from the bridge cache. The mode rename is purely clarity / support-
  bundle accuracy.
2026-06-02 14:33:20 +02:00
maziggy 1e08c25a9f fix(stats): #1593 multi-plate parser + per-run project rollup + carry-over system totals + accuracy band
Two stacked causes under-reported multi-plate prints in the project
  rollup and the archive card.

  Root cause 1 - parser only read plate 1.

  ThreeMFParser._parse_slice_info used root.find(".//plate") and pulled
  prediction / weight from that one element. Any multi-plate file's
  archive-level print_time_seconds / filament_used_grams reflected
  plate 1 alone. The /plates endpoint already looped findall and was
  correct, which is why the plate carousel showed the right numbers
  while the archive card was wrong.

  Fix: loop every <plate> and sum prediction + weight. Per-plate
  concepts (plate_number, _plate_index, printable_objects) only set
  when there's exactly one plate - for multi-plate exports the
  archive represents all plates and a single index doesn't apply at
  the file level. bed_type keeps the first plate's value as a
  best-effort default. Malformed prediction / weight on individual
  plates skip cleanly rather than poison the sum.

  Root cause 2 - project rollup aggregated PrintArchive, not the
  per-run log.

  compute_project_stats and list_projects quick-stats summed
  PrintArchive.print_time_seconds / filament_used_grams / cost /
  energy_* WHERE project_id. A reprint reuses the source archive row
  and writes a new PrintLogEntry, so 3 sequential runs collapsed to 1
  archive - and that archive's numbers were already plate-1-only from
  cause 1. The Archive Print Log path was already correct because it
  drove off print_log_entries (archives.py:420 comment).

  Fix: both compute_project_stats and the list_projects quick-stats
  block inner-join print_log_entries -> print_archives WHERE
  archives.project_id. total_archives becomes COUNT(PrintLogEntry.id),
  failed_prints counts runs in failed/aborted/cancelled/stopped,
  completed_items is SUM(PrintArchive.quantity) for runs where
  status='completed', time/filament/cost/energy from PrintLogEntry.
  Orphan log rows (archive_id IS NULL post archive deletion) are
  excluded by the inner join.

  Same-shape fixes carried forward (no follow-ups per project rule):

  system.py system-info totals: total_print_time / total_filament
  had the same bug shape - summed PrintArchive directly so reprints
  collapsed to one row. Now sums PrintLogEntry.duration_seconds /
  filament_used_grams. The semantic shift is also a correctness
  improvement: the field now reflects time the printer actually spent
  printing, not slicer-estimated time.

  archives.py time-accuracy metric: estimate / actual per run where
  estimate = PrintArchive.print_time_seconds. Post-parser-fix
  multi-plate archives have file-level estimate but per-run actual =
  one plate, so ratio = N x 100% for an N-plate file. The calc now
  clamps each row to the [50%, 200%] plausibility band before
  contributing to the printer-level average; single-plate accuracy
  (the case the metric is designed for) stays fully included.

  Backfill: users with AMS spool tracking - the reporter's case - have
  per-run filament_used_grams from the tracked spool delta, so stats
  become correct immediately. Users without tracking fall back to the
  archive estimate and undercount until they reprint. Archive card
  still reads PrintArchive.filament_used_grams directly so old
  multi-plate archives keep plate-1-only numbers until reslice -
  forward-only as the reporter accepted.
2026-06-02 13:35:06 +02:00
maziggy c9bc5eb4e9 fix(webhook): #1584 PrinterState dataclass attribute access in status / stop / cancel
webhook.py treated printer_manager.get_status() return as a dict and
  called .get(...) on it. The return is a PrinterState dataclass
  (backend/app/services/bambu_mqtt.py), so the call raised AttributeError
  and Starlette surfaced it as a generic 500 for every printer with a
  status row. Non-existent printers correctly returned 404 because the
  early "Printer not found" branch fired before the crash.

  Reporter's repro matched exactly: id 1 (existing printer) returned 500,
  id 2 and id 3 (no row) returned 404. Verified end-to-end against a live
  PG-backed instance with the reporter's key shape — same 500 before the
  patch, 200 with the correct payload after.

  8 crash sites across 3 routes:
    - webhook_get_printer_status   GET  /printer/{id}/status      5 sites
    - webhook_stop_print           POST /printer/{id}/stop        2 sites
    - webhook_cancel_print         POST /printer/{id}/cancel      2 sites

  Every status.get("X", default) replaced with status.X if status else
  default. Pydantic response schema unchanged; PrinterState's dataclass
  defaults cleanly cover the "registered but never connected" branch so
  the status route now returns 200 with connected=false, state=null
  rather than crashing.
2026-06-02 12:57:51 +02:00
maziggy 171848a0fc fix(slicer-presets): #1581 SliceModal refresh + invalidate on local-profile delete/import
Two-part fix for the reporter's "removed profiles still show on the slice
  menu" symptom.

  Local half (real bug). LocalProfilesView's import and delete mutations
  invalidated ['localPresets'] (the management view's own query) but not
  ['slicerPresets'] (the SliceModal's unified preset query, staleTime 60s).
  A freshly-deleted preset kept rendering in the slice dropdown until that
  staleTime elapsed plus a refocus/remount. Both mutations now also call
  queryClient.invalidateQueries({queryKey: ['slicerPresets']}).

  Cloud half (opt-in cache bypass). _fetch_cloud_presets keeps a 5-minute
  per-(user, token) in-process cache (slicer_presets.py:69, balances
  "users see freshly-saved presets quickly" against "busy install doesn't
  hit Bambu Cloud once per modal open"). Users delete cloud presets in
  Bambu Studio / Bambu Handy, not in Bambuddy, so there's no event hook
  to invalidate on. Rather than shorten the TTL globally, the listing
  endpoint gains an opt-in ?refresh=true query param that bypasses both
  the cloud cache AND the 1-hour bundled-preset cache for that one call;
  the fresh result is still written back so subsequent normal callers
  keep hitting the cache.

  New SliceModal "Refresh" button. Lives in the preset section header
  next to the cloud-status banner. Calls getSlicerPresets({refresh: true})
  and writes the fresh slots into the ['slicerPresets'] cache via
  queryClient.setQueryData so the spinner stops immediately rather than
  triggering a second refetch. RefreshCw icon spins while in-flight;
  disabled during slice enqueue to prevent double-fire.
2026-06-02 12:42:16 +02:00
maziggy 396e9aa09e security: harden path-traversal class across routes + services; fifth CI backstop
Two attacker-controlled strings were being joined to library_dir with no
  resolve + containment check in the project ZIP import endpoint:

    - linked_folders[*].name from the request's project.json
    - per-entry zf.namelist() paths from the ZIP itself

  An absolute path in either field collapsed the join (Path("/lib") / "/etc"
  becomes Path("/etc") because pathlib discards the left side when the right
  is absolute) and the next write_bytes landed wherever the attacker chose.

  Adjacent finding from the routes audit: GET /archives/{id}/photos/{filename}
  had NO validation on filename and FileResponse-served arbitrary paths -
  the DELETE counterpart at least gated on the photos membership check.

  Adjacent finding from the services audit: ArchiveService.attach_timelapse
  wrote archive_dir / filename where filename ultimately came from a printer's
  FTP listing (compromised-printer threat model) or the /timelapse/select
  query param. A malicious printer that exposes a directory entry with ..
  segments could write the timelapse outside the archive directory.

  New backend/app/utils/safe_path.py::safe_join_under(parent, *parts) is the
  single source of truth: rejects empty / null-byte / absolute parts up-front,
  joins under parent, resolves both sides, asserts is_relative_to. Returns the
  resolved canonical path on success, raises HTTPException(400) on escape, or
  PathTraversalError when http=False (for service-layer callers that need to
  match a non-HTTP return contract).

  Wired into the import vectors, both archive photo handlers, and the
  attach_timelapse service. The full audit sweep inspected every Path/Name
  join in backend/app/api/routes/ AND backend/app/services/ - 25 route-layer
  sites + 8 service-layer sites confirmed safe and tagged with
  # SEC-PATH-OK: <reason> so future audits trust the inline guard at a glance.

  Fifth CI backstop test_route_path_arithmetic_is_safe_joined_or_marked
  AST-walks both layers and fails the build on any <dir-like>/<bare variable>
  join that doesn't either route through safe_join_under or carry the marker.
  The services layer is in scope because it receives values verbatim from the
  routes AND from external sources Bambuddy has no control over (the printer
  FTP-listing case above).

  SECURITY.md gets a fifth rule + a fifth row in the CI test mapping table;
  the rule now names the printer FTP-listing case explicitly so future
  services-layer audits set the right expectation.

--------------

  fix(library): suppress warning storm when bulk-uploading ZIPs of empty/stub STL files

  Uploading a ZIP of stub or empty STL files (e.g. the 24-byte
  "solid test\nendsolid test" shape) produced one WARNING per file in
  stl_thumbnail.py::generate_stl_thumbnail. The warnings were technically
  correct - trimesh returns a valid Mesh with zero vertices, the safeguard
  matches, and the function returns None so the library entry is still
  created without a thumbnail - but the volume turned a successful upload
  into thousands of WARNING lines in the journal.

  Two changes:

  1. The per-file "Failed to load STL or empty mesh" message in
     stl_thumbnail.py is now logger.debug instead of logger.warning. It's
     a per-file content observation, not an actionable error; the caller
     already handles None correctly. The branch now catches the rare
     "large enough but trimesh still can't parse it" case, visible in
     debug logs without spamming production.

  2. New module constant MIN_USABLE_STL_BYTES = 200 (smallest binary STL
     with one triangle is 134B, smallest ASCII ~150B; 200 is a safe floor
     below any real STL). The three thumbnail call sites in library.py
     (extract_zip_file, single-file upload, _backfill_external_stl_thumbnails)
     pre-skip files below this size before calling generate_stl_thumbnail.
     Stubs never enter the trimesh pipeline at all.

  Behavior is unchanged for real STLs: any file >=200 bytes runs through
  the existing pipeline, MAX_VERTICES still triggers simplification at
  100k vertices for the 256x256 thumbnail render, large files still get
  thumbnails.

------------

  fix(stl-thumbnail): silence matplotlib first-import noise (writable cache + font_manager log level)

  On first STL upload, three matplotlib-internal log lines surfaced:

    WARNING [matplotlib] /opt/claude/.config/matplotlib is not a writable directory
    INFO    [matplotlib.font_manager] Failed to extract font properties from NotoColorEmoji.ttf
    INFO    [matplotlib.font_manager] generated new fontManager

  The writable-dir warning fired because Bambuddy's $HOME isn't writable for
  matplotlib's default config path; matplotlib fell back to /tmp/matplotlib-XXX
  which lost the font cache on every host reboot, so font_manager rebuilt it
  each cold start - producing another batch of INFO lines.

  Fix is two small additions in stl_thumbnail.py before the matplotlib import:

  1. New _configure_matplotlib_cache() sets MPLCONFIGDIR to
     settings.base_dir/.cache/matplotlib (mkdir if missing) so the cache
     persists across container restarts and the writable-dir warning never
     fires. Respects an externally-set MPLCONFIGDIR so operators who chose
     their own path aren't overridden. Best-effort with a debug fallback if
     settings can't be imported or the mkdir fails.

  2. logging.getLogger("matplotlib.font_manager").setLevel(WARNING) at module
     import demotes the per-font INFO scan that fires when font_manager
     builds its cache cold. Real font warnings (>= WARNING) still surface.

  3 new tests: font_manager logger at WARNING after module import;
  _configure_matplotlib_cache creates the directory under base_dir and sets
  MPLCONFIGDIR; an externally-set MPLCONFIGDIR is preserved verbatim.
  5516 backend tests green, frontend gates clean.
2026-06-02 12:12:54 +02:00
maziggy db9b20631c fix(cloud): #1575 surface actionable error when Bambu Cloud Cloudflare challenge swallows the JSON response
When Cloudflare in front of bambulab.com returns a "Just a moment..." interstitial
  instead of the JSON the API normally produces, the parse error in
  verify_totp / verify_code / login_request used to surface as the opaque "Invalid
  response from Bambu Cloud" or a generic 401 from BambuCloudAuthError. Reporter
  hit this with three back-to-back TOTP attempts; a curl from a different network
  with the same honest Bambuddy UA returns clean JSON, so the trigger is CF-side
  (per-IP / TLS-fingerprint / rate / transient mitigation), not our code.

  Add a small _detect_cloudflare_challenge() helper that inspects the response for
  four CF markers (body "Just a moment...", body "challenges.cloudflare.com", 403
  with cf-mitigated header, 503 with cf-ray header) and returns a message that
  attributes the block to Bambu Lab's Cloudflare protection, suggests waiting a few
  minutes, and points the user at a same-network browser sign-in as the standard
  workaround. Wired into all three JSON-parse sites; verify_totp previously had a
  defensive catch, login_request and verify_code now do too.

  No header changes, no impersonation, no retry loop - pure diagnostics. Stays
  clearly on the right side of Bambu Lab's "no falsified client identity" line.
2026-06-02 10:55:25 +02:00
maziggy be15a375a6 fix(oidc): #1569 populate User.email from standard 'email' claim when email_claim is preferred_username
When an operator configures `Email Claim = preferred_username` (e.g. Authentik) the
  primary `_resolve_provider_email` correctly rejects the identity value as non-email
  shaped and returns None, leaving auto-provisioned users with `email=None` even though
  the same token carries a valid standard `email` claim.

  Add a narrow fallback in the auto-create-users branch only: when
  `provider.email_claim != "email"` and the primary returned None, resolve the standard
  `email` claim with the same Fall A/B shape + email_verified enforcement and use it for
  `User.email` and `UserOIDCLink.provider_email`.

  The auto-link-existing-accounts gate is left on the primary `provider_email`, so the
  GHSA Fall-B / Fall-C guards remain intact - the fallback never feeds account matching.
2026-06-02 10:23:21 +02:00
maziggy b7d7c82501 fix(security): WebSocket auth gate + audit-driven hardening sweep 2026-06-02 10:02:17 +02:00
maziggy ec51394196 fix(security): GHSA-r2qv-8222-hqg3 — allowlist API-key permissions (CVSS 9.9)
API-key permission gates went from a 17-entry admin denylist with the three
  documented scope flags (can_read_status / can_queue / can_control_printer)
  enforced only inside /api/v1/webhook/* to an explicit per-Permission
  allowlist consulted by every dependency:

    - core/auth.py: _APIKEY_SCOPE_BY_PERMISSION maps every non-admin
      Permission to one scope flag on APIKey; unmapped = 403.
      _check_apikey_permissions now takes the api_key and checks the flag.
    - require_any_permission_if_auth_enabled + require_ownership_permission
      were returning None for any valid key with zero scope check; both now
      invoke _check_apikey_permissions and fail closed.
    - Two new scope flags on api_keys: can_manage_library (LIBRARY_UPLOAD /
      UPDATE_OWN / DELETE_OWN / MAKERWORLD_IMPORT) and can_manage_inventory
      (INVENTORY_CREATE / UPDATE / DELETE / FORECAST_WRITE — required by
      SpoolBuddy kiosks). Default TRUE, backfilled from can_queue so existing
      "queue-only" keys keep working and hardened "read-only" keys do not
      silently gain writes.
    - CLOUD_AUTH now routed through can_access_cloud for defence-in-depth
      alongside the existing _cloud_api_key_gate.
    - Migration column-existence check (_api_keys_column_exists) gates the
      backfill so user-edited values are never overwritten on restart.

  Structural drift backstop: test_every_permission_has_a_classification fails
  CI on any new Permission added without an explicit scope mapping —
  prevents the denylist-shape regression that grew the prior surface.

  Backend 5469 tests green; ruff clean. Frontend build green; i18n parity
  green across 9 locales (5005 leaves each, +6 new keys). Wiki permissions
  table + allowlist callout + upgrade notes updated.
2026-06-02 08:28:24 +02:00
maziggy f6f6f92e38 test(pytest): silence upstream starlette httpx2 deprecation noise 2026-05-30 14:19:19 +02:00
maziggy e10462678f fix(security): GHSA-6mf4-q26m-47pv — fail-closed on auth-probe DB errors (CVSS 9.8)
is_auth_enabled() and auth_middleware both caught every exception during
  the auth-state probe and returned the "allow" answer instead of denying
  the request. Reporter's PoC floods /api/v1/auth/login to exhaust file
  descriptors, forcing the next SQLite connect to raise, then hits a
  protected endpoint during the fail-open window with no token — granting
  unauthenticated access to admin-account creation, API-key creation, DB
  backup download, and printer control. CWE-636 / CWE-755. Affects >= 0.1.6.

  Fix:
  - is_auth_enabled (backend/app/core/auth.py): only returns False for the
    legitimate "settings row absent" case; any actual exception propagates
    so the caller can deny the request.
  - auth_middleware (backend/app/main.py): returns 503 on any probe failure
    instead of await call_next(request).

  4 new regression tests in test_auth_fail_closed.py pin the contract
  (propagates DB exceptions, returns False for no-row, True for "true",
  False for "false"). 1 existing security test renamed and updated to
  accept either 500 or 503 (both fail-closed) and to verify the
  SQLAlchemy detail does not leak in the body.

  Codebase grep confirmed no other auth-decision predicate has the same
  fail-open shape: _validate_api_key returns None on catch (→ 401 fail-
  closed downstream), is_advanced_auth_enabled propagates correctly,
  permissions.py has no catch-alls.

  Reported by @wondercrash via private advisory.
2026-05-30 13:49:09 +02:00
maziggy 597762685c fix(virtual-printer): #1558 Send pre-flight + slicer-surface audit bundle
#1558: cached-as-base push_status only forced gcode_state=IDLE while letting
  the real printer's live-progress fields (mc_percent, stg_cur, layer_num, ...)
  leak through. Bambu Studio's Send pre-flight read them as busy and refused.
  The cached branch now overrides the activity-field set the same way it
  already overrode storage indicators (#1228) and protocol fields.

  Same bundle ships a multi-round VP audit that found adjacent bugs in the
  same family:

  - #1558: cached branch zeroes mc_print_stage / mc_percent / mc_remaining_time / stg / stg_cur / layer_num / total_layer_num / print_error
  - MQTT auth: per-IP rate-limit (5/60s lockout), hmac.compare_digest, access_code redacted in DEBUG log
  - FTP cmd_STOR streams chunks to disk + 4 GiB cap (was buffering whole upload)
  - Sticky-keys allowlist extended with upgrade_state / xcam / hw_switch_state / nozzle_diameter / nozzle_type / online / ams_status
  - _pending_files cleanup in finally for archive / queue / dispatch handlers
  - _add_to_print_queue position uses MAX+1 (was hardcoded 1)
  - DELETE VP removes orphan PendingUpload rows + upload_dir from disk
  - Per-VP cert regenerates on shared-CA rotation (real signature verification, not DN match)
  - DHCP target-IP refresh + queue_force_color_match toggle now restart proxy VPs
  - Per-slicer bridge-response routing (multi-slicer cross-leak fix via sequence_id map)
  - Child-service readiness barrier (FTP / MQTT / Bind / SSDP) — no false is_running before sockets bind
  - H2D Pro O1E / O2D model codes added (experimental, needs field confirmation)
  - FTP passive port range widened 50000-51000; docker-compose + wiki updated
  - VP refresh_loop crash now unbinds raw_message_handler; tailscale catches asyncio.TimeoutError; SlicerProxyManager lifecycle hardening
2026-05-30 13:34:10 +02:00
maziggy ed232718f5 fix(prints): connected-edge reconciliation closes the missed-PRINT-COMPLETE loop behind smart-plug ghost prints (#1542 follow-up)
Reporter ran a fresh trace after the doubled-extension fix landed and
  found a distinct second cause behind his ghost prints, hitting 4-of-4
  of his A1s. Timeline:

    22:50 PRINT START
    ...print runs all night...
    23:13 / 00:47 / 09:35 MQTT disconnects (A1 keepalives are unstable)
    print finishes during one of those disconnect windows → PRINT COMPLETE
      is never observed
    smart plug detects idle → cuts power
    power resumes for the next scheduled print → firmware auto-replays the
      leftover .3mf from the SD card
    09:46 Bambuddy reconnects to a fresh PRINT START for the ghost

  The existing IDLE-after-RUNNING completion check at
  backend/app/services/bambu_mqtt.py:3022 was meant to catch the simple
  disconnect-then-finish case via `_previous_gcode_state` preserved across
  reconnects, but with multiple disconnect/reconnect cycles + a smart-plug
  power-off that Bambuddy can't distinguish from any other transient drop,
  the IDLE window that branch needs simply never reaches it. The SD .3mf
  lingers, the firmware ghost-replays every power cycle, and the loop
  repeats.

  Fix: new connected-edge reconciliation pass.

  * `_is_active_archive_stale(archive, state)` — pure decision function
    with three triggers:
      (1) printer state is terminal (IDLE / FINISH / FAILED)
      (2) printer running with a different `subtask_id` than the archive —
          Bambu firmware mints a fresh subtask_id for each print including
          the ghost replay, so a mismatch is unambiguous
      (3) printer running but `subtask_name` is empty — printer doesn't
          know what it's running, archive reference is broken
    Conservative on PAUSE / PREPARE / SLICING and on RUNNING with matching
    subtask. False-positive cost = one misreported "aborted" status that
    the next real PRINT COMPLETE would have overwritten anyway. False-
    negative cost = the ghost-print loop.

  * `reconcile_stale_active_prints(printer_id)` — queries archives in
    `status="printing"` for the printer, runs the decision function, and
    synthesises `on_print_complete(status="aborted", _reconciled=True)`
    for each stale match. Reuses the existing PRINT COMPLETE chain (SD
    cleanup, status update, usage tracker, notifications) — no reimplem-
    entation. Per-archive try/except so one failure doesn't block the
    rest. Returns 0 when status is None / disconnected — the connected
    edge is the only legitimate trigger.

  * `on_printer_status_change` now runs a connected-edge check at the
    start. New `_printer_reconciled_since_connect: dict[int, bool]`
    tracker flips False → True on the first connected status update for
    this connection and back to False on disconnect, so reconciliation
    fires exactly once per (re)connection. The flag is set BEFORE the
    task is spawned so concurrent status updates within the same
    connection don't re-trigger it (no await between check and set,
    asyncio guarantees atomicity). The reconciliation runs as
    `asyncio.create_task` so the hot WebSocket dedup / broadcast path
    isn't blocked.

  * Single handler covers both startup and reconnect — when the first
    MQTT connection completes after startup the printer pushes status,
    the connected edge fires, reconciliation runs. No separate startup
    hook needed.

  Idempotency: when the existing #3022 branch DOES fire on a clean
  disconnect-then-IDLE-on-reconnect, it lands `archive.status` to terminal
  synchronously. The async-scheduled reconciliation then queries
  `WHERE status="printing"` and finds 0 rows → no-op. The narrow race
  where reconcile's query lands between a real on_print_complete's
  archive update and its archive lookup produces at most a duplicate
  notification (no double SD-cleanup since FTP delete on a missing file
  is a no-op 550).

  Ghost-print collateral worth being explicit about: if the ghost is
  already running when reconciliation fires, the synthesised SD-cleanup
  hits 550-file-locked (same root cause as the #1542 first case). The
  cleanup retries 3× then logs "lingering". The ghost runs to completion,
  its own end-of-print cleanup deletes the file, the next power cycle
  has nothing to replay, the loop breaks. A perfect cancel would require
  a `print_stop` MQTT command to the printer mid-ghost — invasive,
  explicitly out of scope.
2026-05-29 11:18:06 +02:00
maziggy 7a7dfed81a fix(archives): resolve raw_data wrapper in fallback-archive filament extraction (#1533 follow-up)
Reporter updated to 0.2.5b1 expecting the #1533 fix to populate filament
  fields on his P2S virtual-printer prints when the .3mf is locked. His
  support bundle showed Bambuddy still creating fallback archives with NULL
  filament fields even though the print-start log line proved AMS-0-T0 had
  PETG loaded at the moment the helper should have read it
  (`AMS 0: T0(type=PETG, color=FFFFFFFF, ...)`).

  Cause: the #1533 helper `_extract_filament_data_from_mqtt(data)` in
  backend/app/main.py only looked at `data["ams"]`, but the dict that
  on_print_start actually receives at runtime is the wrapper shape
  `{"filename", "subtask_name", "remaining_time", "raw_data": <mqtt>,
  "ams_mapping"}` that backend/app/services/bambu_mqtt.py:2971-2980
  constructs. So `data["ams"]` was undefined on every real call and the
  helper silently returned `{}`, leaving the fallback archive's
  filament_type / filament_color NULL — the exact regression the original
  fix was meant to close. The 15 unit tests that shipped with #1533 all
  passed the bare inner shape directly and never exercised the callback
  wiring, so the regression slipped through a green build. Surrounding code
  in the same fallback path (main.py:2064, 2674) already reads
  `data.get("raw_data")` — the original hunk was the outlier that forgot
  the wrapper.

  Fix: helper now resolves `data["raw_data"]["ams"]` first (the callback
  shape) and only falls back to `data["ams"]` when the wrapper isn't
  present, preserving the inner-shape callers from the existing tests.
  Defensive against a non-dict `raw_data` (e.g. partial MQTT decode
  failure) falling through to the inner lookup instead of crashing.

  Tests: 5 new in `TestOnPrintStartCallbackShape` in
  test_fallback_archive_mqtt_filament.py — wrapper payload with ams_mapping
  resolves to the inner AMS state; wrapper without ams_mapping lists all
  loaded slots; the existing inner-shape callers still work after the
  additive wrapper lookup; missing raw_data returns `{}` instead of
  raising; junk raw_data (string) doesn't shadow a present inner `ams`.
  Existing 15 inner-shape tests untouched and green. Full 5378-test
  backend suite green; backend ruff clean.

  What this does NOT fix: per-filament gram usage still needs the actual
  .3mf — the printer locks it during print (P-line firmware behaviour,
  not a Bambuddy bug), and the existing 19 FTP candidate paths +
  directory probes are expected to 550 in that window. Per-print filament
  type and colour are the data point the reporter explicitly called out
  as load-bearing for AMS-expansion planning at his maker space, so this
  is what moves the needle.
2026-05-29 10:55:41 +02:00
maziggy b663605318 fix(virtual-printer): honour client-negotiated MQTT keepalive instead of hardcoded 60s (#1548)
OrcaSlicer connects, exchanges pushall + get_version, then sits idle waiting
  for status pushes from the (virtual) printer. The VP MQTT server's read
  loop used `asyncio.wait_for(reader.read(1), timeout=60)` regardless of what
  the client negotiated, and `_handle_connect` explicitly skipped the
  keepalive field in the CONNECT payload, so every idle slicer connection was
  torn down at exactly 60s.

  - Parse the 2-byte big-endian keepalive from CONNECT; return it from
    _handle_connect alongside the auth bool.
  - Use 1.5x the negotiated keepalive as the per-packet read timeout per
    MQTT spec sec 4.4. Treat keep_alive == 0 as no timeout (spec sec 3.1.2.10).
  - Retain the 60s default for the initial read before CONNECT arrives, so
    a TCP-connect-without-CONNECT still gets reaped.
  - 7 new tests: 4 unit-level for the parser (success, opt-out=0, auth-fail
    tuple shape, malformed CONNECT) + 3 integration-style for the read loop
    (long keepalive survives the old 60s mark, short keepalive closes idle
    in ~3s, PINGREQ resets the window so DISCONNECT decides the exit).
2026-05-28 09:13:44 +02:00
maziggy 6d316c5593 fix(dispatch): align SD cleanup with upload path so doubled-extension library rows don't leave ghost prints (#1542)
A library row with archive.filename "Cube (1).gcode.3mf.gcode.3mf" uploaded
  to /Cube_(1).gcode.3mf.3mf (single-iteration strip + append). Post-print
  cleanup looked at /Cube_(1).3mf and /Cube_(1).gcode (subtask_name + ext),
  missed the on-card file, and A1 firmware re-ran it on next power-on.

  - New derive_remote_filename() helper in backend/app/utils/filename.py:
    iterative strip of .gcode.3mf/.3mf suffixes, append single .3mf,
    space->underscore. isinstance() guard raises TypeError on non-str
    input rather than entering the strip loop with a duck-typed
    object that returns truthy sentinels from endswith.
  - Three previously-duplicated upload sites (background_dispatch reprint +
    library, print_scheduler queue) now share the helper.
  - SD cleanup fetches archive.filename and tries the derived path first,
    with the legacy subtask_name + ext paths kept as fallbacks.
  - 10 new unit tests pin the reproducer + edge cases (doubled .gcode.3mf,
    doubled .3mf, raw .gcode preserved, idempotence, Unicode, plus the
    type guard against MagicMock / None / int inputs).
2026-05-28 08:40:43 +02:00
maziggy 2241924312 fix(library): reject FAT32-illegal filename chars at rename/upload/queue time (#1540)
Bambu printer SD cards are FAT32/exFAT, which forbids < > : " / \ | ? *
  plus control chars and trailing dots/spaces. Library rename only blocked
  path separators, so a name like L|R.3mf was accepted and only failed
  later at FTP upload with 553 Could not create file - far from the rename
  action that caused it. Bambu Studio refuses these names in its save
  dialog; Bambuddy now does the same.

  New backend/app/utils/filename.py centralises validation. Wired into
  update_file, upload_file, print_library_file, and queue add. Existing
  rows with bad names are left alone (no silent rewrite of user data);
  users get an actionable 400 pointing at rename.

  Frontend rename modal mirrors the same set client-side with inline error.
  New fileManager.invalidFilenameChar i18n key translated across all 9
  locales. 26 new tests in test_filename_validation.py.
2026-05-27 09:54:07 +02:00
maziggy c42e923e4c fix(archives): MQTT-derived filament type/color on fallback archives (#1533)
When the source .3mf can't be downloaded at print start (P1S/A1/P2S
  firmwares lock the file mid-print), main.py creates a fallback
  PrintArchive with file_path="" and every filament field NULL — even
  though the MQTT payload already has the AMS state and the slicer's
  slot-per-print-filament mapping (data["ams"]["ams"] and
  data["ams_mapping"]).

  New _extract_filament_data_from_mqtt(data, ams_mapping) builds a
  {global_tray_id: (type, color)} map from the AMS units, then narrows
  to slots referenced by ams_mapping (slicer order preserved, -1 VT-tray
  sentinels skipped) or falls back to every loaded slot when no mapping
  is present. Returns comma-separated filament_type and filament_color
  matching the 3MF-extraction shape, so the inventory page, Quick Stats
  rollup, and len(filament_type.split(",")) per-print count behave
  identically for fallback rows.

  The constructor at the fallback site now passes the resulting values
  into the PrintArchive row.

  This does NOT recover per-filament gram usage — that needs the .3mf's
  slice_info.config or a deeper layer-delta integration via usage_tracker.
  The reporter (maker-space lead evaluating Bambuddy partly for AMS
  expansion planning) asked specifically for "the number of filaments
  used", which is what this gives them.

  15 unit tests cover empty/malformed payloads, the no-mapping path,
  mapping filtering and reordering, VT-tray sentinels, dual-AMS global
  ids, column-limit truncation, and defensive garbage handling.
2026-05-26 11:50:41 +02:00
maziggy 4343bd60b1 fix(notifications): honest UA + Cloudflare-challenge detection on ntfy (#1534)
The notification service's httpx client was the only outbound client in
  the codebase still leaking python-httpx/<version> as User-Agent; all
  other clients identify as Bambuddy/1.0 since the May 2026 compliance
  pass. Bring it in line.

  The reporter's ntfy server was behind a Cloudflare Tunnel and CF returned
  its JS challenge page (Just a moment...) to every API request — confirmed
  by reproducing the same 403 with curl. Cloudflare can't be solved from a
  backend, so add detection for the challenge shape (Server: cloudflare or
  cf-mitigated header, or <!DOCTYPE html>...Just a moment... body) and
  return an actionable error message that points at the real fix on the
  user's CF side instead of dumping the raw HTML.

  Normal 403s (auth failures with plain text bodies) still surface the
  original body so genuine errors stay debuggable.
2026-05-26 11:21:03 +02:00
maziggy e9beb1e8fc fix(archives): handle fallback archives in source-3MF upload (#1531)
Archives created from prints Bambuddy didn't archive (cloud / Handy /
  SD-card prints) carry file_path="". The two source-upload routes
  computed the destination as (base_dir / archive.file_path).parent /
  "source", which collapsed to base_dir.parent / "source" for fallback
  rows — sending the file to /app/source/ (outside the data volume,
  orphaned on container restart) and raising 500 on the final
  relative_to.

  Centralise the destination math in _resolve_source_3mf_path. Normal
  archives keep the <archive>/source/<filename> layout. Fallback
  archives land at <base_dir>/archive/no_source/<id>/<filename>, which
  stays inside the data volume and is addressable by every existing
  read site. The helper also asserts the resolved directory is under
  base_dir.resolve() so a corrupted row fails with a clear message
  instead of writing outside the volume.

  Both upload routes (upload_source_3mf and upload_source_3mf_by_name)
  now route through the helper. Two regression tests in
  TestUploadSourceThreeMF pin both branches.
2026-05-26 11:04:46 +02:00
maziggy 4387a09162 fix(spoolbuddy): route weight sync by inventory mode exclusively (#1530)
POST /spoolbuddy/scale/update-spool-weight tried the local DB first
  and only fell back to Spoolman on a local miss. Combined with
  nfc/tag-scanned's post-#1119 always-Spoolman routing, a stale local
  Spool row sharing a numeric id with a Spoolman spool would absorb
  the sync silently while the Spoolman row stayed unchanged.

  Mirror the routing already used by nfc/tag-scanned: pick the branch
  via _get_spoolman_client_or_none() and never cross. Local mode now
  returns 404 on a local miss instead of falling through.

  New TestUpdateSpoolWeightSpoolman.test_stale_local_row_does_not_shadow_spoolman
  asserts both directions: Spoolman gets the update, the colliding
  local row's weight_used and last_scale_weight are untouched.
2026-05-26 10:36:37 +02:00
maziggy 554a73070f fix(maintenance): paused prints no longer accumulate runtime hours (#1521)
PAUSE counted toward runtime_seconds equally with RUNNING, inflating
  hours-based maintenance thresholds (rod lube, belt check, nozzle clean)
  by however long overnight or extended pauses lasted. Maintenance items
  track mechanical wear, which is zero while paused, so the predicate
  now excludes PAUSE. Field-comment and docstring trail across main.py /
  models/printer.py / maintenance.py updated to match. Existing
  runtime_seconds values cannot be retroactively split — only future
  accumulation is fixed.

  Adds 3 regression tests pinning PAUSE non-accumulation, RUNNING
  accumulation, and the FINISH state's last_runtime_update clear
  (prevents idle-time back-bill when the printer next goes RUNNING).
2026-05-25 09:04:47 +02:00
maziggy ca08f1f340 fix(test): stop sys.modules-deleting backend.app.main in test_code_quality
+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup

  Root cause of the 4 CI failures on PR #1514 (all in
  test_print_start_assigns_printer_id_to_vp_archive.py +
  test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
  in test_code_quality.py was deleting backend.app.main from sys.modules
  and re-importing it via importlib.import_module. That created NEW
  module-level dicts (_timelapse_baselines, _expected_prints,
  _active_prints, …) and re-ran root_logger.addHandler — hence the
  duplicate log lines at the same microsecond in captured stderr.

  Any sibling test that bound those names via "from backend.app.main
  import _timelapse_baselines" before the reimport now held a reference
  to the OLD dict; production code (reached via "from backend.app.main
  import on_print_start") resolved the symbol through the NEW module
  instance. Production mutated the new dict, the test read the old one,
  the assertion saw None / un-mutated mock_archive.

  Locally with -n 30, xdist load-balanced test_code_quality.py to a
  different worker process so the collision never happened (which is
  why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
  made the collision deterministic.

  Fix: drop the "del sys.modules[name]" step. importlib.import_module
  already returns the cached module if cached, or runs the import
  machinery if not — either way, any import-time error surfaces. The
  "fresh import" framing was theatre; in practice every module in the
  list is already imported by other tests/fixtures before this test
  runs, so we were never actually getting a fresh import anyway — just
  destruction.

  CI workflow tightening (separate concern, same PR since both touch
  the test infrastructure):

  - Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
    lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
    runners. --tb=short is sufficient for failure context.
  - Sharded backend-tests into a 4-way matrix via pytest-split (new
    dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
    all 4 run in parallel so wall-clock drops from 362s -> ~100s.
  - fail-fast: false on the matrix so a single failing shard doesn't
    hide failures in the other three — PRs see the complete failure
    picture in one push.
2026-05-24 12:51:24 +02:00
maziggy 22f222e4ac fix(test): snapshot _timelapse_baselines inside the patch context to dodge CI race
test_running_observed_captures_baseline_on_restart_recovery was reading
  _timelapse_baselines.get(1) after the patch() with-block exited.
  Locally and under low parallelism this works fine — the dict still
  holds what _capture_timelapse_baseline_at_start wrote. CI under
  xdist's default load-balancing scheduling intermittently saw the
  dict empty by the time the top-level assert ran, even though the
  production code logged "Baseline at print start: 3 video files for
  printer 1" right before returning. The duplicate log line at the
  same microsecond in the captured stderr is the tell — module state
  is being re-touched between the handler completing and the test
  asserting, almost certainly via the session-scoped event_loop
  fixture in conftest.py interacting badly with the per-file
  autouse _clear_baselines teardown of a sibling test on the same
  worker.

  The test is verifying the handler captured the baseline at the
  moment it returned, so capture the relevant value at exactly
  that point — inside the with-block, immediately after the await.
  That's immune to whatever happens to the module-level dict
  afterward.
2026-05-24 12:30:12 +02:00
maziggy eb98521e93 fix(test): use /nonexistent/ instead of /tmp/ to satisfy Bandit B108
The test_returns_empty_when_3mf_missing test sets a deliberately
  non-existent file_path on a PrintArchive to verify
  compute_deficit_for_queue_item handles the missing-3MF branch
  gracefully. The path just needs to fail an existence check — the
  /tmp/ prefix was incidental.

  Bandit B108 ("insecure temp file usage") regex-matches /tmp/,
  /var/tmp/, and /dev/shm/. Dropping /tmp/ in favour of /nonexistent/
  keeps the test behaviour identical (still a guaranteed-missing
  path, still triggers the missing-file branch) while clearing the
  GitHub Advanced Security finding on PR #1514 without adding a
  # nosec annotation.
2026-05-24 12:03:11 +02:00
maziggy 3b9633a178 ● feat(support): include sanitized connection / VP / log-health diagnostics in support bundle and bug report (#1506 follow-up)
The three diagnostic surfaces shipped earlier this month
  (6bc6a1d6 VP setup diagnostic, e222a0ef log-health scanner,
  ed31b8f4 connection diagnostic in the bug-report bubble) were
  only ever shown to the *user*. A bug report arriving in the
  maintainer's inbox carried raw logs but no diagnostic results —
  the user-visible "your X1C can't reach MQTT" finding never made
  it into the issue body, so the maintainer had to ask the user
  to re-run and paste.

  New `services/diagnostic_snapshot.collect_diagnostic_snapshot`
  runs all three concurrently with a per-probe 15 s wall-clock cap
  (so total ≈ max(per-cap), not sum — fleet size doesn't matter)
  and is fail-soft per probe: a crash inside one printer's check
  emits `{"printer_id": N, "error": "..."}` for that entry rather
  than nuking the whole snapshot. The snapshot is then added as a
  `diagnostics` top-level key by `_collect_support_info()`, so both
  flows (POST /support/bundle and POST /bug-report/submit via
  `support_info=...`) pick it up without their own changes.

  Private-data sanitization
  -------------------------
  The diagnostic schemas embed raw IPv4 in five field shapes that
  must not land in a submitted GitHub issue or a shared support ZIP:

    - PrinterDiagnosticResult.ip_address (top-level)
    - DiagnosticCheck.params.printer_ip (network-mode check)
    - DiagnosticCheck.params.host_ip (network-mode check)
    - VPDiagnosticResult per-check params.bind_ip (VP setup)
    - IPs embedded in log-health sample lines

  The first two carry the printer's own IP (already in the
  existing `collect_sensitive_strings` table via the Printer rows);
  host_ip and bind_ip are NOT in the DB so a sensitive_strings-only
  pass missed them.

  Fix: `_sanitize_recursive` walks the full snapshot tree, masks
  DB-known values with the same `[PRINTER]/[IP]/[SERIAL]/[ACCESS_CODE]`
  labels the log sanitizer applies (via the shared
  `collect_sensitive_strings`), then an IPv4-regex pass catches any
  IP the DB didn't cover — most importantly the Bambuddy host IP
  returned by `_get_host_ip()` and the VP `bind_ip` the user picked
  at setup. Recursive walk so arbitrary nested dicts/lists don't
  slip through future schema additions.

  Live-DB smoke test against the dev fleet: zero raw IPv4 instances
  in the serialized snapshot output; all five field shapes plus the
  embedded log samples render as `[IP]`.

  Progress indicators
  -------------------
  The bubble's "submitting" view and the System page's Download
  button now render a static four-line checklist showing what's
  running (printer connectivity → VP setup → log scan →
  submit / build ZIP). Static, not faked phase progress — we can't
  actually track server-side phases without SSE and the honest
  "here's what's happening" list communicates the longer wait
  without lying about percentage complete.

  9 new i18n keys, real translations in all 9 locales (no English
  fallback). parity script clean at 4993 leaves per locale.

  Tests: 6 new in test_diagnostic_snapshot.py
    - empty-input shape stable (the three top-level keys always present)
    - per-printer / per-VP result coverage (lists match input lengths)
    - fail-soft on a single-probe crash (other entries + log-health
      still complete)
    - timed_out marker when a probe exceeds the per-probe cap
      (test patches the cap to 0.05 s)
    - end-to-end IP sanitization across all five field shapes plus
      log-sample IPs, with a final JSON-serialize-and-regex sweep
      asserting zero raw IPv4 escapes anywhere in the result
    - concurrent execution proof (4 × 0.2 s probes complete in
      < 0.5 s; sequential would be 0.8 s)
2026-05-24 10:16:30 +02:00
maziggy 03896d1af0 fix(scheduler): use inventory weight for "Prefer Lowest Filament" sort (#1508)
Reporter has a P1S with an inventory spool cloned to slot 1 and the
  original (much further used) in slot 4, the preference enabled, and
  the dispatch picked slot 1 every time. The sort's been blind to
  Bambuddy inventory weights — it reads MQTT `tray.remain`, the
  printer firmware's RFID-decremented value, which has two limitations:

  - Bambu RFID only. Non-RFID spools report -1 and get clamped to a
    sentinel; multiple non-RFID trays then tie in the sort and Python's
    stable sort collapses to AMS-slot insertion order, so slot 1 wins.
  - Even when set, it's the printer's counter, not Bambuddy's
    `label_weight - weight_used` (internal) or Spoolman's
    `remaining_weight`. The two diverge whenever the user re-spools,
    swaps cardboard, or runs a print outside Bambuddy.

  The reporter is on internal-inventory mode with non-RFID spools — both
  failure modes apply, hence slot 1 every time.

  Fix: when a slot is bound to an inventory spool the inventory record's
  remaining weight becomes the sort signal. New async helper
  `_build_inventory_remain_overrides(db, printer_id, loaded)` returns
  `{global_tray_id: remaining_grams}` for bound slots — internal mode
  joins SpoolAssignment → Spool once per dispatch; Spoolman mode joins
  SpoolmanSlotAssignment then reuses `_spoolman_remaining_grams` from
  filament_deficit.py for parity.

  New `_prefer_lowest_sort_key` does a two-tier comparison: inventory-
  tracked spools sort BEFORE MQTT-only spools, then ascending by
  remaining within each tier, then ascending by ams_id*4+tray_id as the
  deterministic slot tie-breaker. The tier flag dominates so grams
  (inventory) and percent (MQTT) never get cross-compared — no unit
  conversion needed.

  MQTT-only behaviour is preserved exactly: remain=-1 still maps to the
  101 sentinel and slot order still decides on ties. Users who haven't
  bound any inventory spool see no change. The DB lookup runs only when
  prefer_lowest_filament is enabled.

  External / VT slots are skipped (tracked separately from AMS bindings).
2026-05-24 09:18:38 +02:00
maziggy eae96da56e fix(camera): probe ffmpeg for the right RTSP socket-timeout flag (#1504)
A previous attempt swapped `-timeout` → `-stimeout` unconditionally to
  fix EADDRINUSE on the reporter's transitional ffmpeg. That broke every
  install on a modern ffmpeg (5+/6+/7+) — current Debian/Ubuntu/Homebrew
  — where `-stimeout` was removed and `-timeout` is back to meaning
  socket I/O. Verified locally: `ffmpeg -stimeout ...` errors
  "Unrecognized option 'stimeout'" on ffmpeg 7.1.
  install on a modern ffmpeg (5+/6+/7+) — current Debian/Ubuntu/Homebrew
  — where `-stimeout` was removed and `-timeout` is back to meaning
  socket I/O. Verified locally: `ffmpeg -stimeout ...` errors
  "Unrecognized option 'stimeout'" on ffmpeg 7.1.

  ffmpeg has shipped THREE arrangements of this option over time and
  Bambuddy supports the full range:

  - Pre-deprecation (early 4.x and earlier): `-timeout` is socket I/O.
  - Transitional (~late-4.x, Jammy-era): `-timeout` is deprecated and
    repurposed to RTSP listen-mode timeout; any non-zero value implies
    `-listen`, which makes ffmpeg bind the TLS-proxy port and fail with
    EADDRINUSE. `-stimeout` is the replacement socket I/O option.
  - Modern (5.x / 6.x / 7.x): `-stimeout` REMOVED. `-timeout` is back to
    socket I/O — the original meaning.

  So no single literal is correct on all installs.

  Fix: `rtsp_socket_timeout_flag()` in services/camera.py probes
  `ffmpeg -h demuxer=rtsp` once and picks `-stimeout` when ffmpeg
  advertises it (covers transitional + older builds that kept it as an
  alias), else `-timeout` (modern + pre-deprecation). Cached at module
  level for the process lifetime — ffmpeg doesn't swap mid-run.

  The function returns the option name without a leading dash; callers
  prepend it themselves so a formatting bug can't pass an empty flag.

  Wired into both RTSP ffmpeg call sites in lockstep: routes/camera.py
  (printer camera) and services/external_camera.py (external RTSP),
  which use the same TLS-proxy + ffmpeg pattern and would hit the same
  regression on either ffmpeg cohort.

  Tests: 8 in test_ffmpeg_rtsp_timeout_flag.py — 6 probe unit tests
  (prefers stimeout when advertised, falls back to timeout on modern,
  defaults to timeout when ffmpeg missing or probe raises, caches across
  calls, trailing-space substring guard against `-listen_timeout`
  false-positives), 2 parametrised guards against either RTSP ffmpeg
  argv re-hard-coding a literal instead of consuming the probe. 37
  probe + existing external-camera tests green.
2026-05-24 08:49:21 +02:00