Commit Graph
246 Commits
Author SHA1 Message Date
maziggy 950286ad40 fix(queue): prevent duplicate dispatch and stale progress on batch prints
Two related queue issues surfaced when scheduling an ASAP print with
  quantity > 1 on an H2D:

  1. Double-dispatch — both items in the batch ended up in 'printing'
     status on the same printer, logged as "BUG: Multiple queue items in
     'printing' status for printer N". The scheduler seeded its busy
     set empty each tick and relied on _is_printer_idle() reading live
     MQTT state, but H2D / P1 series lag several seconds between the
     print command and IDLE → RUNNING, so the next check_queue() tick
     saw IDLE and dispatched the second batch item onto the already-
     running printer. check_queue() now seeds busy_printers with every
     printer_id that has a row in 'printing' status before iterating,
     so any printer with an outstanding dispatched job is excluded
     regardless of what MQTT currently reports.

  2. Progress bar flashed 100% — immediately after dispatch the queue
     item's per-row progress bar showed the prior print's final mc_percent
     for a few seconds, then snapped back to 0% when the new print
     started ticking. QueuePage.tsx now gates progress / remaining_time /
     layer fields on status.state being RUNNING or PAUSE; in any other
     state (FINISH from the prior print, IDLE, PREPARE while heating)
     the bar renders at 0% with no stale ETA or layer count.

  Regression coverage added in test_phantom_print_hardening.py
  (TestBusyPrinterSeedingFromPrintingItems, 3 tests): seeding query
  returns only printers with 'printing' rows, empty when none exist,
  and end-to-end check_queue() does not call _start_print for a pending
  item whose printer already has a 'printing' row even when
  _is_printer_idle() is forced True.
2026-04-22 17:51:33 +02:00
maziggy 0918907dab fix(scheduler): watchdog falsely reverts slow H2D dispatches, causing reprints (#1078)
_watchdog_print_start reverted queue items to "pending" at 45 s if
  gcode_state hadn't changed, assuming the MQTT project_file was swallowed
  by a half-broken session (#887/#967). H2D Pro firmware (01.01.00.00)
  routinely keeps state=FINISH for 48-55 s after actually accepting the
  command before transitioning to PREPARE. The watchdog reverted items
  the printer had already started physically printing; the archive updated
  normally via _active_prints, but the queue item was now "pending" again,
  and the next scheduler tick after plate-clear re-dispatched the same
  item as if it had never run. With one item left in the queue that looked
  like a reprint of the just-finished job; with multiple items the
  symptom was masked by item N+1 getting dispatched during the race.

  Add a second "command landed" signal: subtask_id advancing past the
  pre-dispatch value. Bambuddy already mints a unique submission_id per
  project_file publish (#1042) and the printer echoes it back on the next
  push_status as soon as it starts processing the command - well before
  gcode_state transitions on slow-transition models. _start_print now
  captures pre_subtask_id alongside pre_state and passes both to the
  watchdog, which exits early on either a state change or a subtask_id
  advance.

  Raise default timeout 45 s → 90 s as belt-and-braces for printers that
  neither flip state nor echo subtask_id inside the polling window.
  Genuinely half-broken sessions (both signals unchanged across the full
  90 s) still revert + force-reconnect exactly as before.

  Transient subtask_id=None during reconnect is not mis-detected as a
  change. pre_subtask_id=None falls back to state-only checking so the
  fix is safe for printers that haven't reported a subtask_id yet.

  New test_scheduler_watchdog.py pins the eight behaviours that matter:
  pickup via state change; pickup via subtask_id change with state still
  FINISH (the exact #1078 case); revert when neither signal changes;
  default timeout is 90 s; pre_subtask_id=None state-only fallback;
  current subtask_id=None not treated as change; printer disconnect
  mid-watchdog leaves DB untouched; item that already moved on is not
  clobbered.
2026-04-22 08:38:50 +02:00
maziggy 4e86e8cb16 fix(printers): Clear-Plate button delayed 30s–5min after print completes (#939 follow-up)
PR #939 added the awaiting_plate_clear gate but stored it on
  PrinterManager, not on PrinterState. printer_state_to_dict() — which
  builds every WebSocket printer_status payload — never emitted the flag,
  so the frontend's WS merge preserved the stale false value. The only
  path that surfaced true was the 30s HTTP fallback poll, and incoming WS
  ticks kept bumping React Query's dataUpdatedAt, pushing the refetch out
  further on chatty printers.

  Emit awaiting_plate_clear from printer_state_to_dict by reading
  printer_manager.is_awaiting_plate_clear(printer_id) directly; returns
  False when no id is passed. No frontend change needed — the existing WS
  merge carries the flag end-to-end and the button now appears the instant
  the printer transitions to FINISH.

  Regression tests assert the WS dict always contains the key and surfaces
  True when the manager has the flag set for that printer_id.

  Affects every printer (A1/H2D/X1C) equally — transport-agnostic path.
2026-04-21 18:10:02 +02:00
maziggy bf5135cb12 fix(inventory): malformed rgba no longer bricks the Filaments page (#1055)
A single legacy spool with a 7-char rgba ('FFFFFFF', missing one F)
  caused GET /api/v1/inventory/spools to 500 with a pydantic
  ResponseValidationError, leaving the reporter with a blank Filaments
  page and "Add Spool" silently failing. Root cause spans three layers:

  1. Write path: SpoolUpdate.rgba had no pattern constraint (only
     SpoolCreate did), so PATCH could plant malformed values in the DB.
  2. Frontend: ColorSection hex input's `val.length <= 6 ? 'FF' : ''`
     emitted 7-char rgba for 5-char input (XXXXX + FF = 7) and for
     7-char typed input (no alpha appended).
  3. Read path: SpoolResponse inherited the write-side pattern, so a
     single bad row 500'd the entire list endpoint instead of being
     tolerated through serialize.

  SpoolUpdate.rgba now carries the same ^[0-9A-Fa-f]{8}$ pattern as
  SpoolCreate. The hex input emits a fully-formed 8-char RRGGBBAA on
  every keystroke — 8-char paste passes through, 7-char drops the
  stray, shorter input pads RGB with '0' and appends FF alpha.
  SpoolResponse.rgba is now Optional[str] with no pattern — write-side
  validation is the right place for format rules; responses must
  tolerate historical rows.

  Tests: 16 schema tests (SpoolCreate/Update reject, SpoolResponse
  tolerate), 7 frontend tests covering every input length 0–8 plus
  non-hex strip. A user who already has a bad row in their DB now sees
  it render with a default color instead of having to hand-edit SQLite.
2026-04-21 14:40:03 +02:00
maziggy 1682b6956f fix(dispatch): clean up transient library upload from Direct-Print flow (#730)
The "Print" button on a printer card (and drag-drop-onto-card) used
  FileUploadModal to persist the file as a LibraryFile, then dispatched
  through POST /library/files/{id}/print. The LibraryFile row + disk file
  were left behind after every one-off print, polluting File Manager with
  entries the user never asked to save.

  FilePrintRequest.cleanup_library_after_dispatch (default False) opts
  into post-dispatch cleanup. When set, _run_print_library_file stages
  db.delete(lib_file) in the same transaction as archive_print so a
  mid-flight FTP / start_print failure rolls both back cleanly, commits
  together, then unlinks the library disk file + thumbnail after commit
  succeeds. External library files (is_external=True) are never touched.

  Only the Printers-page Direct-Print PrintModal sets the flag. Every
  other api.printLibraryFile caller (File Manager Print, Project Detail
  Print) leaves it unset — their entries are there by user intent.

  Also moves formatPrintName out of PrintersPage.tsx into a new
  utils/printName.ts module — fa1c46d9 (#881) exported it inline so its
  test could import it, tripping react-refresh/only-export-components.
2026-04-21 14:10:04 +02:00
maziggy fa1c46d9a5 feat(printers): show plate name on card for multi-plate active prints (#881)
When two printers were running different plates of the same multi-plate
  3MF, the Printers page cards displayed the same file name on both and
  there was no way to tell them apart. The Queue view already had this
  information by cross-referencing the archive's plate list; the card
  didn't have the linkage.

  Expose `current_archive_id` (resolved by matching the MQTT `subtask_id`
  against `PrintArchive.subtask_id` — the bridge introduced in #972 for
  restart-resume) and `current_plate_id` (parsed from `gcode_file` by a
  new shared `parse_plate_id` helper) on the status endpoint. The helper
  is also called from the WebSocket push path so plate transitions
  reflect within 100 ms instead of waiting 30 s for the next REST poll;
  the archive id itself stays REST-only since it's stable for the life
  of a print and shouldn't make the push path touch the DB.

  The card fetches plate metadata via the same `api.getArchivePlates()`
  call QueuePage uses — shared React Query cache keeps it cheap across
  polls — and renders the actual plate name (or a "Plate N" fallback)
  only when `is_multi_plate` is true. Single-plate prints stay clean.
  Falls back to the previous `plate_N.gcode` regex path when there's no
  archive linkage (e.g. prints started directly from the printer LCD).

  Tests cover the plate-id extraction across Bambu Studio path shapes
  (backend parse_plate_id, printer_state_to_dict wiring) and the label
  override precedence in formatPrintName (frontend).
2026-04-21 09:46:42 +02:00
maziggy 87a5aa36e9 fix(ams): HT slot shows "Generic" after configuring custom preset (#1053)
After configuring an AMS-HT slot with a custom cloud preset, the slot
  card and Configure modal kept showing "Generic PLA" even though the
  printer and slicer had the correct preset. The `/slot-presets` response
  keyed HT entries at `ams_id * 4 + tray_id = 512`, but frontend lookups
  used `ams_id` directly (128 on PrintersPage via getGlobalTrayId, 64 on
  SpoolBuddy via a one-off formula). All three agreed for regular AMS, so
  the mismatch only surfaced on HT — the saved preset never reached the
  UI and the render fell through to `tray.tray_type`.

  Backend now keys via a helper that mirrors frontend `getGlobalTrayId`.
  SpoolBuddy's AMS page switches to the shared helper. Regression test
  covers regular, HT, and external slot keys.
2026-04-20 17:32:34 +02:00
maziggy 7026a6de77 fix(printer): bed-jog "Home Z" could crash bed into toolhead on H2C/H2D/H2S/X1 (#1052)
Critical safety fix. The bed-jog dialog's "Home Z" button sent a bare
  `G28 Z` over gcode_line. On Bambu printers where the Z endstop is at
  the top (bed moves UP into it — H2C, H2D, H2S, X1 family), `G28 Z`
  skips the toolhead-park step that a full `G28` runs first, so the bed
  rises at full speed with nothing getting out of the way. The reporter
  only escaped damage because the toolhead happened to be parked on the
  purge chute.

  The /printers/{id}/home-axes endpoint and BambuClient.home_axes() now
  always send bare `G28` regardless of the axes argument, triggering the
  firmware's safe multi-step routine (park toolhead → home XY → home Z).
  The axes argument is kept for API compat but ignored; invalid values
  still return 400.

  Frontend retitles the button "Auto Home" and updates the dialog copy
  in all 7 locales so users aren't surprised when X/Y motion happens
  before Z. Parameterized regression test asserts z/xy/all all produce
  bare G28.
2026-04-20 12:56:31 +02:00
maziggy d0f35e5d60 fix(mqtt): cap task_id at int32 max to prevent P1S dispatch stalls (#1042) 2026-04-20 08:46:57 +02:00
maziggy d3425c7f44 fix(ftp): wait for zombie thread to complete before giving up on download (#1014) 2026-04-20 08:39:09 +02:00
maziggy ea78fe720c fix(obico): clear Status banner on next successful detection cycle (#172) 2026-04-20 08:19:55 +02:00
maziggy 74527d4124 fix(smart-plug): restore MQTT subscriptions for per-type topic configs on startup (#1010)
Users integrating a Shelly plug through an external MQTT broker
  (ioBroker, Zigbee2MQTT, HA's MQTT broker, etc.) lost the plug's
  power/state/energy readings after every Bambuddy restart. The only
  fix was opening Settings → Smart Plugs, renaming the topic to a dummy
  value, saving, renaming back, and saving again.

  Root cause: three code paths configure an MQTT smart plug's
  subscriptions — the startup restore in main.py, the create route,
  and the update route — and they had drifted. The create/update
  routes used the newer per-type model (mqtt_power_topic /
  mqtt_energy_topic / mqtt_state_topic with per-type paths,
  multipliers and mqtt_state_on_value) while the startup restore was
  still on the legacy single-topic model. Worse, the restore loop
  short-circuited on `if plug.mqtt_topic:`, skipping any plug whose
  topics were only set in the new per-type fields — exactly the shape
  of a Shelly-via-ioBroker config, which publishes power and state on
  separate topics. The "rename, save, rename back" workaround routed
  through the update endpoint and re-established the subscription the
  correct way.

  Extracted the topic-resolution + service.subscribe() call into
  subscribe_plug_to_mqtt() in mqtt_smart_plug.py and routed all three
  paths through it so the schema can't drift again. The helper keeps
  the legacy `mqtt_topic` field working as a fallback for all three
  data types — matching the behaviour the startup restore used to
  have via subscribe()'s internal `effective_*_topic or topic`
  collapsing, and matching the change-detection dict already used
  during updates.

  Regression tests cover: per-type topics restored without a legacy
  topic, legacy single-topic backward compat, per-type multipliers
  overriding legacy, per-type winning when both are set, the
  empty-config skip case, and topic-list de-duplication.
2026-04-19 13:51:58 +02:00
maziggy 936b748127 fix(archive): truncation of large 3MF uploads on sendfile short-return (#1032)
On bare-metal Raspberry Pi OS bookworm / armv7l / Python 3.11, 3MF
  files larger than a few megabytes arrived complete via the
  virtual-printer FTP server but the copy into data/archives/ was
  silently truncated. The archive row was still written, the printer
  card looked fine, and the problem only surfaced later when opening
  the archive — the subsequent zipfile.ZipFile() in
  GET /archives/{id}/plates raised BadZipFile and the UI came up blank
  with no thumbnail, plate list, or filament data.

  Two things conspired:

  1. archive_print() used shutil.copy2, which takes Python's sendfile()
     fast path on Linux. On the reporter's kernel/fs combination
     sendfile returned a short count on the first call for the upload
     sizes hit in practice and the destination ended up truncated.
     Small files completed in one syscall and were fine.
  2. ThreeMFParser.parse() caught the resulting BadZipFile in a bare
     `except Exception: pass`, so the archive pipeline kept going with
     empty metadata and left the bad file on disk — nothing in the
     logs hinted anything had gone wrong until a support bundle came
     in and the "Failed to parse plates" warning fired much later.

  The archive copy is now an explicit chunked read/write with fsync —
  sendfile is not in the path. After the copy, if the source was a
  valid ZIP but the destination isn't, we refuse to create the archive
  row, remove only the truncated file (and the archive directory if
  empty — archive_dir is created with exist_ok=True so rmtree would be
  unsafe if a same-second same-filename collision happened), and log
  both sizes at ERROR so the condition is obvious in future support
  bundles. The parser's silent catch now logs at WARNING for the same
  reason.

  All nine archive_print() call sites already check `if archive:` or
  `if not archive:`, so returning None for corrupted ZIPs propagates
  cleanly without behaviour changes elsewhere.

  Regression tests cover single-chunk and multi-chunk copies, mtime
  preservation via copystat, overwrite of an existing destination, a
  ZIP roundtrip through a multi-megabyte 3MF, the new parser WARNING,
  and a truncation sentinel verifying that zipfile.is_zipfile() flips
  to False on a half-written ZIP — the exact post-condition
  archive_print now trusts.
2026-04-19 13:40:43 +02:00
maziggy c7ad449e4e fix(firmware): parse P2S/X2D wiki anchors without dash and full-width parens (#1030)
The wiki scraper silently returned no versions for P2S and X2D, causing
  Bambuddy to fall back to the Bambu Lab download page, which still listed
  01.01.01.00 as "latest" even though 01.02.00.00 shipped on 2026-04-09.

  Two regex mismatches in _fetch_all_versions_from_wiki():

  1. Heading anchor ids require an optional dash between version bytes and
     date. H2D/X1/H2C/H2S use "h-01020000-20260409"; P2S and X2D publish
     "h-0102000020260409" (no dash).
  2. The text fallback only matched ASCII parens around release dates, but
     P2S, X2D, A1 and A1-mini render dates in full-width parens (YYYYMMDD)
     (U+FF08/U+FF09).

  Anchor regex now accepts an optional dash; fallback accepts both paren
  styles. Added regression tests for both shapes.
2026-04-19 12:27:06 +02:00
maziggy 10c261dcf2 chore(tests): suppress B108 on dummy /tmp test fixtures 2026-04-19 09:48:10 +02:00
maziggy 68920f8c62 Fix virtual printer dropping null-terminated MQTT payloads from OrcaSlicer Linux (#927)
OrcaSlicer's Linux BBLNetworkPlugin publishes MQTT payloads with the
  C-string null terminator included in the length, so decoded messages
  arrived as `{…}\x00`. The strict json.loads() raised JSONDecodeError
  and the publish handler silently returned — pushall, get_version, and
  project_file were never answered, and the slicer hit its 60 s sync
  timeout. Print_queue mode only (proxy mode tunnels MQTT). The b069b521
  serial-adaptation fix was correct but ran past this earlier silent
  failure.

  _handle_publish now strips trailing \x00/whitespace before parsing and
  logs the raw payload on any remaining decode failure so future silent
  variants are visible in support bundles.
2026-04-19 08:04:05 +02:00
Minidoracat baf0716a9a feat(cloud): support China region for token-based login (#1013)
feat(cloud): support China region for token-based login

The /cloud/token endpoint always used the global Bambu API endpoint,
so users with China-region access tokens could not validate their
token. The password login flow already exposes a region selector; this
brings the token flow to parity.
2026-04-18 12:30:01 +02:00
maziggy 115d6fe627 fix(mqtt): unique per-submission IDs for archive reprints (#1011)
Archive reprints and library-file prints built the MQTT project_file
  command with hardcoded project_id="0", subtask_id="0", task_id="0".
  Printers key per-job state (including gcode_start_time) on those IDs,
  so reprints looked like continuations of the same job and third-party
  MQTT observers (OctoEverywhere) reported compounding durations across
  repeat replays — a 40 min job reprinted from archive showed ~1h40m,
  and a second reprint of the same file showed ~4h. BambuStudio mints
  fresh IDs per submission; bambu_mqtt.start_print() now does the same
  using an epoch-millisecond timestamp for all three fields. md5 is
  deliberately left empty to avoid activating firmware md5-validation
  against a digest we can't compute without re-reading the upload.

  Added 6 regression tests in TestStartPrintUniqueIdentityFields
  covering non-zero IDs, md5 stays empty, uniqueness across successive
  submissions, numeric-string format, and blast-radius guard on
  unrelated payload fields. Updated CHANGELOG.
2026-04-18 09:09:21 +02:00
maziggy a2c7fd4542 fix(obico): revert POST-bytes approach — Obico /p/ is GET-only
The 0.2.3b4 #1003 "fix" POSTed JPEG bytes as multipart form data,
  but Obico's /p/ endpoint is declared methods=['GET'] upstream and
  reads ?img=URL from the query string. Every POST was 405'd by
  Flask's router before any handler ran, which is why the Obico
  container logs were silent while Bambuddy kept reporting
  "ML API call failed for printer N:" with a blank suffix —
  raise_for_status() on the 405 produced an exception whose str()
  rendered empty.

  Restored the pre-#1003 nonce-URL approach (commit 3e434458):
  capture locally with a 20s timeout we control, stash the JPEG
  under a single-use 32-byte nonce, hand Obico a
  GET /api/v1/obico/cached-frame/{nonce} URL that resolves in
  <50ms so its hardcoded 5s read timeout never races RTSP.

  Also guards against future silent exceptions: the error format
  now falls back to type(exc).__name__ when str(exc) is empty.
  Detection also early-returns with an explicit error if
  external_url is unset instead of handing Obico a URL it can't
  resolve.

  The #1003 reverse-proxy scenario (Authelia/Authentik/CF Access
  in front of Bambuddy) is addressed by documenting that the
  /api/v1/obico/cached-frame/ path must be whitelisted from
  external auth at the proxy layer — it is already public on
  Bambuddy's side.

  Backend: services/obico_detection.py, api/routes/obico.py,
  main.py (PUBLIC_API_PATTERNS).
  Frontend: FailureDetectionSettings banner + client.ts type +
  all 7 locales restored.
  Tests: 15 unit + 5 integration tests pass.
2026-04-18 08:50:46 +02:00
maziggy 464d56ea0d fix(install): make SpoolBuddy kiosk usable on first boot in full-mode install
Full-mode install booted into an unusable kiosk:
  - Chromium opened before uvicorn → "can't connect to localhost"
  - After reload, requires_setup=true hijacked /spoolbuddy → /setup
  - Touch-only Pi has no keyboard to complete the setup wizard
  - Declining auth left the user at / instead of the kiosk

  Fixes, bundled:

  1. backend/app/cli.py kiosk-bootstrap now, in one DB transaction:
     - creates a scoped API key (can_read_status=True, rest false)
     - upserts setup_completed=true
     so AuthContext never redirects and the kiosk URL loads directly. Users
     who want auth can still enable it from the admin UI; the provisioned
     key keeps working.

  2. install.sh full-mode runs the CLI as the bambuddy service user after
     create_bambuddy_service and sed-replaces the CHANGE_ME_AFTER_SETUP
     placeholder in spoolbuddy/.env.

  3. The generated spoolbuddy-kiosk-launch polls ${backend_url}/health for
     up to 60s before exec'ing chromium, so cold boots wait for uvicorn
     instead of flashing ERR_CONNECTION_REFUSED.

  Standalone mode was unaffected — users supply a real key from their
  existing Bambuddy before install.
2026-04-18 08:14:53 +02:00
maziggy 3502ab33c5 fix(install): auto-provision SpoolBuddy kiosk API key in full-mode install
Full-mode install wrote CHANGE_ME_AFTER_SETUP as SPOOLBUDDY_API_KEY because
  no admin exists yet to create a real one. On reboot the kiosk launched with
  that placeholder, AuthContext rejected it, and the user hit the Bambuddy
  login page instead of the kiosk. Standalone mode was unaffected — users
  paste a real key from their existing Bambuddy before install.

  Adds backend/app/cli.py with a kiosk-bootstrap subcommand that creates a
  scoped APIKey row directly in the DB (can_read_status=True, everything else
  false) and prints the full key to stdout. install.sh full-mode runs it as
  the bambuddy service user after create_bambuddy_service, captures the key,
  and sed-replaces the placeholder in spoolbuddy/.env. Idempotent with
  --force for re-installs.

  Drops the outdated "create an API key and edit .env" next-step block since
  the kiosk is now provisioned automatically.
2026-04-18 07:40:38 +02:00
maziggy a95a3c52ee fix(mqtt): detect zombie sessions via ams_filament_setting response tracking (#887)
After hours idle the MQTT connection can degrade so telemetry still
  flows but published commands never reach the printer.  The existing
  dev-mode probe only ran on first connect; this adds tracking for
  user-initiated ams_filament_setting commands — two consecutive
  unanswered commands (10 s timeout each) trigger force_reconnect.
2026-04-17 09:29:46 +02:00
maziggy 475e34ebda fix(obico): POST image bytes directly to ML API instead of callback URL (#1003)
The ML API previously called back into Bambuddy to fetch snapshots,
  which failed behind reverse proxies with external auth (Authelia, etc.).
  Now the detection loop captures the JPEG locally and POSTs it directly
  as multipart form data — no callback URL, no nonce cache, no
  external_url dependency.
2026-04-17 09:06:31 +02:00
maziggy ef37ffa7c7 fix(obico): exclude snapshot capture PIDs from stream cleanup (#172)
The periodic camera cleanup task scans /proc for ffmpeg processes and
  kills any not in the active-streams registry. The Obico detection
  service's capture_camera_frame_bytes() spawns short-lived ffmpeg for
  snapshots but never registered the PID — so cleanup killed it as
  "orphaned" mid-capture (SIGKILL, exit -9), producing false errors and
  missed detection frames.

  Track capture PIDs in _active_capture_pids and exclude them from the
  cleanup kill list.
2026-04-17 08:38:22 +02:00
maziggy 3e434458a4 fix(obico): capture snapshots locally and serve via nonce URL (#172)
Obico's ML API has a hardcoded 5s read timeout on the URL it fetches, which
  our /camera/snapshot regularly exceeds on cold calls (TLS proxy + ffmpeg +
  RTSP keyframe wait). The detection loop now captures the JPEG locally with
  a 20s timeout we control, stashes the bytes under a single-use 32-byte
  nonce, and hands Obico a new /api/v1/obico/cached-frame/{nonce} URL that
  returns the cached bytes instantly. The 5s ceiling is no longer a factor.

  The nonce is the credential (URL-safe, 256 bits of entropy, single-use,
  30s TTL) so the endpoint can be unauthenticated without widening the
  camera access surface. Replaces the previous camera-stream-token snapshot
  URL approach, which remained vulnerable to the upstream 5s timeout even
  when auth was disabled.

  Thanks to @fblix for the detailed reproducer with timeout numbers.
2026-04-16 11:09:54 +02:00
maziggy 6fb814c5ea feat(printer): add X2D support — camera, dual-nozzle, K-profile, maintenance (#988)
The Bambu Lab X2D (launched April 2026, dual-nozzle, enclosed, hardened
  steel rod gantry, AMS 2 Pro compatible) identifies itself as internal
  model code N6 via SSDP/MQTT, and real serials begin with 20P9. None of
  these identifiers existed in Bambuddy's registries, so the camera
  service fell back to the chamber-image protocol on port 6000 (X2D
  doesn't speak it), firmware-check logged "Unknown printer model: N6",
  and the dual-nozzle K-profile paths — gated on the H2D serial prefix
  "094" — would have treated X2D as single-nozzle.

  Backend:
  - Register N6 → X2D across every registry (PRINTER_MODEL_ID_MAP,
    PRINTER_MODEL_MAP, STEEL_ROD_MODELS, ETHERNET_MODELS,
    CHAMBER_TEMP_SUPPORTED_MODELS, firmware-check API keys + wiki path,
    virtual-printer SSDP/product/serial tables, DB vp_model_fixes).
  - supports_rtsp(): match the X2 display-name prefix and the N6 internal
    code; camera now routes to RTSP on port 322.
  - Dual-nozzle serial prefix check in bambu_mqtt.delete_kprofile and
    kprofiles.set_kprofile broadened to ("094", "20P9") — X2D now takes
    the H2D-style cali_idx in-place edit path.
  - is_h2d model gate in bambu_mqtt.start_print extended with "X2D" so
    timelapse / bed_leveling / flow_cali / vibration_cali / layer_inspect
    are sent as integers and external-spool ams_id 254/255 routing is
    preserved (H2D-style deputy-nozzle addressing).

  X2D uses hardened steel rods like P2S — it is intentionally placed in
  STEEL_ROD_MODELS, not CARBON_ROD_MODELS. A regression-guard test pins
  the classification.

  Frontend:
  - mapModelCode in PrintersPage and SpoolBuddyAmsPage handle N6 and X2D.
  - Enclosure-door badge and airduct-mode whitelists include X2D.
  - MaintenancePage.getMaintenanceWikiUrl routes X2D to P2S wiki URLs for
    steel-rod lubrication, belt tension, cold-pull, and PTFE tube
    (exported to enable direct unit testing).

  Tests:
  - test_printer_models.py: TestX2DModel (10 assertions).
  - test_bambu_mqtt.py: X2D in start_print ams_mapping and is_h2d gate;
    TestDeleteKProfileDualNozzleDetection across H2D, X2D, P2S, X1C.
  - MaintenancePageWikiUrls.test.tsx: 15 assertions covering X2D, P2S
    regression, X1C/H2D/A1Mini regression, and model-name normalisation.

  Docs:
  - README: added X2 series to the supported printers table.
  - CHANGELOG: new entry under 0.2.3b4 Fixed.

  Credit to @krautech for the report and debug bundle, and to @legend813
  for PR #989 which seeded most of the registry changes — rod-type
  classification was corrected (steel, not carbon) and the dual-nozzle /
  K-profile / is_h2d gaps were added on top.
2026-04-16 10:40:32 +02:00
maziggy 46c246c504 fix(archive): resume on subtask_id, short-circuit 550, cache 3mf (#972)
Second wave of #972 — reproducer on a 37.5 MB BambuStudio print to an A1
  showed three stacking root causes when Bambuddy restarts mid-print.

  1. Archive start_time lost on container restart. The name-based dedup
     cancelled any "printing" archive older than 4h and recreated it with
     started_at=now(), so a 13h print that saw a restart 10h in ended up
     showing ~1.5h duration. Persist MQTT subtask_id on every archive and
     match on that first, regardless of age — same id means same print,
     resume in place. Also revives Stale-cancelled rows for users
     upgrading mid-print.

  2. 3MF FTP search tried non-existent paths for ~48 min. Order was
     /cache → /model → /data → /data/Metadata → / with 11×30s retries
     each; BambuStudio actually pushes to / on A1, so the real path was
     tested last. Reorder to / first, and raise a new FileNotOnPrinterError
     sentinel from download_to_file on 550 so with_ftp_retry short-circuits
     via non_retry_exceptions. 425 / SSL EOF / connection resets still
     retry as before.

  3. Cover endpoint and archive flow downloaded the same 36 MB twice and
     competed for the printer's single FTP socket, producing 425 errors
     that fed cause-2's retry storm. Add an in-memory _threemf_path_cache
     keyed on (printer_id, normalized filename); whichever flow fetches
     first populates it, the other reuses the file read-only. Eviction
     runs on on_print_complete and deletes the temp file.

  Backend: 14 new tests across test_bambu_ftp.py and a new
  test_subtask_archive_resume.py. Existing suite: 2737 pass. ruff clean,
  frontend build clean.
2026-04-16 09:36:44 +02:00
maziggy 1b43488016 fix(printers): recover large-3mf metadata after FTP timeout (#972)
Two-part root cause for missing photos/filament/cost on large prints
  (#972). The configured ftp_timeout was only plumbed through as the FTP
  socket timeout; the asyncio.wait_for wrapping run_in_executor stayed on
  its 60s hardcoded default, so the user's 300s setting never applied.
  Worse, asyncio.wait_for cannot cancel run_in_executor threads — after
  the 60s outer timeout fired, the executor thread kept running
  ftplib.retrbinary and frequently completed the download ~30–60s later,
  but by then the async wrapper had returned False. with_ftp_retry kept
  re-attempting the same path, each retry truncating the file the zombie
  thread had just written, and the archive was ultimately persisted as a
  fallback with no 3MF.

  download_file_async now accepts timeout at each call site (plumbed from
  ftp_timeout) and salvages post-timeout success via an explicit
  completion flag the executor thread sets only after download_to_file
  returns True. Per-attempt completion dict so a prot_p zombie can't
  flip the flag for a later prot_c attempt. A cosmetic // prefix in the
  directory-search download path is also fixed by replacing string
  concatenation with posixpath.join.
2026-04-15 07:52:28 +02:00
maziggy 899c2c6480 revert(printers): remove SD card badge entirely
Four attempts at making the printer-card SD badge stable on H2D all failed:
  the final straw was powering on an A1 causing every connected H2D to flip to
  red simultaneously. Bambu firmware SD signaling is not reliably derivable
  from MQTT — the legacy `sdcard` field is sporadic and inconsistently typed,
  and home_flag bits 8-9 are cleared on heartbeat pushes regardless of card
  state with no clean way to distinguish heartbeats from full status reports.

  Remove the badge from the Printers page card and the Printer Info modal,
  drop `sdcard` from the frontend PrinterStatus type, and strip all home_flag
  derivation and heartbeat-handling code from the MQTT parser.

  `state.sdcard` is retained on the backend and populated only from a plain
  truthy read of the `sdcard` field, because firmware_update.py uses it as a
  precondition before starting firmware installs.
2026-04-14 18:41:46 +02:00
maziggy 0d7c0d4054 fix(printers): stop H2D SD badge from flipping red on heartbeat bursts
Third follow-up on the H2D SD card badge. The prior 3-strike downgrade still
  lost the race: on idle printers, a nearby printer coming online (e.g. an A1
  reconnecting) triggered an MQTT activity burst that let idle H2Ds accumulate
  ≥3 heartbeat home_flag pushes before the next full push_status, flipping every
  H2D badge to red at once.

  Reworked the derivation:
    - the top-level `sdcard` field is authoritative when present (truthy check
      handles bool / int / "HAS_SDCARD_NORMAL" string variants)
    - home_flag bits 8-9 are only consulted on full push_status payloads
      (detected via ≥2 of gcode_state, mc_percent, nozzle_temper, print_type,
      stg_cur, ams)
    - bare heartbeat pushes carrying home_flag alone no longer affect SD state

  Removed the now-dead `_home_flag_seen` latch and 3-strike counter. Tests in
  TestSdCardParsing rewritten to cover the new semantics.
2026-04-14 18:25:36 +02:00
maziggy dd349954a6 fix(printers): stop H2D SD-card badge flipping red on heartbeat pushes
H2D sends heartbeat-style home_flag pushes where bits 8-9 are clear
  even when a card is inserted, so a single heartbeat flipped the badge
  to red until the next full push. Downgrades true->false now require
  three consecutive clear reads; upgrades apply immediately.
2026-04-14 12:32:37 +02:00
maziggy b5ccc38e4a feat(support): include all settings (redacted) + SpoolBuddy devices in support bundle
Settings dump now retains every key from the Settings table and replaces
  sensitive values with [REDACTED] instead of dropping the row. New config
  flags automatically surface in future bundles without a code change.

  Adds integrations.spoolbuddy with per-device firmware, NFC/scale hardware,
  calibration, online state and uptime — anonymized (no hostnames, IPs or
  device IDs). Both /support/bundle and the bug-report bubble benefit, since
  they share _collect_support_info().
2026-04-14 12:22:07 +02:00
maziggy 44bb179364 feat: build-plate Z-jog control from printer card (#791)
Adds a compact "Bed" badge in the printer-card controls row
  between print speed and Stop/Pause. Opens a popover with up/down
  arrows and a 1 / 10 / 50 mm step selector.

  When the Z axis has not been homed since the last print, the
  first jog per session opens a Bambu Studio-style modal with
  Home Z / Move anyway / Cancel. "Move anyway" bypasses soft
  endstops (M211 S0 ... M211 S1) for a single move and is
  remembered for the browser session.

  Backend:
  - POST /printers/{id}/bed-jog?distance=N[&force=bool]
    Emits G91 / G1 ZN F600 / G90 (with optional M211 wrap).
    Distance validated server-side (non-zero, |N| <= 200 mm).
  - POST /printers/{id}/home-axes?axes=z|xy|all
    Emits G28 variants.
  Both gated behind Permission.PRINTERS_CONTROL.

  Frontend:
  - New indigo-themed badge + popover in PrintersPage.
  - Not-homed confirmation modal with sessionStorage "warned" flag.
  - i18n keys under printers.bedJog.* in all 7 locales.

  Tests:
  - backend/tests/unit/test_bed_jog.py — 13 tests covering
    404 / 400 / 500 / success paths for both endpoints, plus
    gcode-payload assertions for force on/off.

  Docs:
  - README feature list, CHANGELOG (0.2.3b4 Unreleased),
    printer-control wiki page, website features.html.
2026-04-14 12:05:06 +02:00
maziggy 0198226db1 Root cause: the Obico detection service handed the ML API a bare /camera/snapshot URL, and Bambuddy's auth returned 401 on that unauthenticated GET. The ML API then
surfaced it as "Failed to get image", which Bambuddy reported back as a 400.

  Fix: the snapshot endpoint already accepts a reusable camera-stream token (the same mechanism used by <img>-based camera consumers, since browsers can't send auth headers
  on image loads). The detection service now appends that token to the URL it gives the ML API. The token is cached on the service, refreshed 5 min before its 60-min expiry,
   and is simply ignored when Bambuddy auth is disabled — so no behavior change for users without auth.

  Refs #172
2026-04-14 11:36:46 +02:00
maziggy d74ab06072 feat(firmware): list all announced versions with usable/unavailable status, support rollback
Firmware update modal now shows every version from Bambu's wiki release
  history, each badged Usable/Unavailable/Installed. Selecting a usable row
  — newer or older than current — swaps the release notes and enables
  install for that version, so rollback no longer requires hand-flashing.

  Wiki scraper tightened to only read heading-anchor ids (h-XXXXXXXX-YYYYMMDD)
  instead of any XX.XX.XX.XX substring, eliminating false positives like an
  AMS firmware version mentioned in an H2D changelog being listed as H2D
  firmware.

  Refs #568
2026-04-14 11:10:24 +02:00
maziggy b71b721658 fix(security): stop leaking webhook tokens via httpx debug logging
Settings → Support → Debug Logging elevated httpx/httpcore to DEBUG,
  which makes httpx log every outbound request URL. For Discord and
  generic webhook notifications the bearer token is embedded in the URL
  path, so users who turned on debug logging to capture a support bundle
  were writing their webhook tokens straight into bambuddy.log.

  Pin httpx/httpcore to WARNING regardless of the debug toggle. paho.mqtt
  still honours debug. Users who enabled debug logging while notifications
  were sending must rotate any exposed Discord/webhook URLs — the token
  is the path, so the whole URL has to be regenerated in the provider UI.
2026-04-14 09:13:55 +02:00
maziggy 9cc7efdf5f fix(printers): stop SD card badge flapping on H2D
Parse `sdcard` from `home_flag` bits 8-9 (HAS_SDCARD_NORMAL /
  HAS_SDCARD_ABNORMAL) when available and fall back to a type-tolerant
  truthy check on the top-level `sdcard` field. Firmware ships that
  field inconsistently (bool, int `1`, or string `"HAS_SDCARD_NORMAL"`),
  so the previous `is True` identity check flipped the badge to red on
  every report that carried a non-bool value.
2026-04-14 08:17:34 +02:00
Sn0rrii ba1c97c808 feat: Two-Factor Authentication (TOTP, Email OTP) and OIDC/SSO – full implementation with admin UI (#933)
feat: Two-Factor Authentication (TOTP, Email OTP) and OIDC/SSO – full implementation with admin UI (#933)
2026-04-13 13:24:28 +02:00
maziggy 8af0966e68 feat(printers): airduct mode + status badges + force refresh on printer card
Surface four Home Assistant-style controls on the Printers page card:

  - SD Card badge in the top status row (green / red, icon-only).
  - Enclosure Door badge in the top status row (green / yellow, icon-only).
    Detection per printer family — X1/X1C/X1E read home_flag bit 23, all
    others read top-level `stat` (hex string) bit 23 — so X1 firmware that
    does not flip stat bit 23 stops false-triggering "open". WebSocket
    status-change dedup key now includes door_open so toggling the door
    alone publishes a push, no 30s REST-poll wait.
  - Airduct Mode badge beside the speed control (cooling / heating)
    for P2S/H2D/H2C/H2S; one-click dropdown calls the existing
    set_airduct MQTT command via a new POST /printers/{id}/airduct-mode
    route.
  - Force Refresh entry in the kebab menu — calls the existing
    /printers/{id}/refresh-status endpoint to request a pushall snapshot
    without forcing a reconnect.

  Tests: door-open parsing (X1 home_flag, non-X1 stat, ignore mismatched
  source, invalid hex) and airduct route (validation, not-connected,
  success, failure).
2026-04-13 12:50:17 +02:00
maziggy eec7793955 feat(obico): AI print-failure detection via self-hosted Obico ML API (#172)
Adds a Failure Detection tab under Settings that wires Bambuddy to a
  self-hosted Obico ml_api container — no cloud, no account, no WebSocket.
  While a print is running, the detection service periodically hands the
  printer's camera snapshot URL to the ML API and smooths scores over
  time (30-frame warmup + EWM, alpha=2/13, short/long rolling means) so
  one noisy frame can't trigger an action. When the smoothed score
  crosses HIGH, the configured action fires exactly once per print:
  notify, pause, or pause-and-cut-power (via linked smart plugs).

  - Backend: new obico_detection + obico_smoothing + obico_actions
    services, /obico/status and /obico/test-connection routes
    (SETTINGS_READ / SETTINGS_UPDATE), six obico_* AppSettings fields
    with validators for sensitivity/action/enabled_printers.
  - Frontend: FailureDetectionSettings component (enable, ML URL + test,
    sensitivity, action, poll interval, per-printer monitor list, live
    status + detection history), new sidebar tab with service-active
    bullet, toast on save.
  - Tests: 17 detection unit tests + 15 smoothing unit tests + 4
    frontend component tests.
  - Docs: README bullet, CHANGELOG entry, wiki page under Analytics,
    website features.html entry.
2026-04-13 09:54:26 +02:00
maziggy de7fff0be4 fix: persist plate-clear gate so Auto Off power cycles can't bypass the queue confirmation (#961)
With Auto Off enabled and another job queued, the smart plug cut power when a
  print finished and immediately re-powered the printer because the scheduler
  saw pending items. The printer booted fresh into IDLE and the next job
  auto-dispatched, bypassing the "Clear Plate & Start Next" confirmation.

  Root cause: the plate-clear gate lived only in PrinterManager._plate_cleared
  (in-memory set) and _is_printer_idle treated IDLE as unconditionally idle. On
  power cycle the in-memory flag was lost and the IDLE-on-boot state skipped
  the gate entirely.

  Fix:
  - Replace the in-memory flag with an awaiting_plate_clear column on the
    printers table, rehydrated into the PrinterManager at startup.
  - Set the flag in on_print_complete for completed/failed prints (not user
    cancellations); clear it on ack and on scheduler dispatch.
  - _is_printer_idle now short-circuits to not-idle whenever require_plate_clear
    is on and the flag is set, regardless of the currently reported state —
    so the gate holds through power cycles, Bambuddy restarts, and the printer
    booting back into IDLE.
  - /printers/{id}/clear-plate no longer requires the printer to report
    FINISH/FAILED; it accepts the ack whenever the flag is raised.
  - Frontend widgets (PrinterQueueWidget, Layout, BulkPrinterToolbar) gate on
    the flag rather than reported state.

  Tests: added regression tests for IDLE+awaiting=True (the #961 case) and
  full DB round-trip tests for the persistence layer.
2026-04-13 09:12:12 +02:00
maziggy 774a639e9a . 2026-04-12 14:16:17 +02:00
maziggy 99c193b535 refactor(colors): color_catalog is the single source of truth (#857)
The Printer tab AMS popup and spool auto-provisioner resolved color
  names from hardcoded tray_id_name tables with a suffix-code fallback —
  and suffix codes like "R1" are not globally unique across material
  families. A17-R1 (PLA Translucent Cherry Pink) fell through the
  fallback and resolved to "Scarlet Red" (A01-R1, PLA Matte), baking
  the wrong name into auto-created inventory spools.

  The fix removes the hardcoded tables entirely. Backend resolves color
  names via the existing color_catalog table by hex; frontend fetches a
  compact {hex: name} map once per session via a new
  GET /inventory/colors/map endpoint (auth-gated but not on
  inventory:read — read-only views need it too) and stores it in a
  ColorCatalogProvider context. A useSyncExternalStore hook cascades a
  re-render into pages mounted before the fetch completes so they
  refresh from HSL-fallback names once the catalog loads.

  Existing auto-provisioned spools keep their stored names; only new
  provisioning and live display benefit. Co-Authored-By is intentionally
  omitted here per project convention — set it via git config if needed.
2026-04-11 12:10:45 +02:00
maziggy 3d893b22f3 fix(auth): make password_hash nullable on upgraded SQLite installs (#794)
LDAP auto-provisioning hit a NOT NULL constraint error on upgraded SQLite
  installs because the existing migration only ran on PostgreSQL. The SQLite
  branch now patches sqlite_master via writable_schema and bumps schema_version
  so the change takes effect without a restart. Fresh installs were unaffected.
2026-04-11 11:33:49 +02:00
maziggy 8266d225d2 fix(energy): date-range energy in total mode + restart-resilient per-print tracking (#941)
The Statistics page reported "Gesamt" (All Time) kWh correctly but showed
  zero for Today/Week/Month in total-consumption mode. Two bugs drove it:

  1. The starting plug counter was kept in an in-memory dict
     `_print_energy_start` that was lost on any backend restart mid-print, so
     the per-print `energy_kwh` delta silently never got computed. The stats
     endpoint's fallback path `SUM(PrintArchive.energy_kwh)` therefore summed
     to zero for users running in total mode.
  2. Total-consumption mode has no per-print delta by design — it includes
     idle/preheat/standby — so the fallback to archive rows was the wrong
     strategy even when the data existed.

  Fix, in two parts:

  - Persist `energy_start_kwh` on the archive row and read it back from a
    fresh session at print end. Deletes `_print_energy_start` and its 5
    call sites, replacing them with a single `_record_energy_start()` helper.
    Per-print tracking is now restart-resilient regardless of tracking mode.
  - Add hourly `smart_plug_energy_snapshots` table + `_snapshot_loop()` in
    SmartPlugManager. Rewrote the `/archives/stats` energy branch as
    `_sum_snapshot_deltas()` which computes per-plug
    `max(0, last-in-range - baseline)` where baseline is the latest snapshot
    at or before the range start, falling back to the earliest-ever snapshot
    and signalling `energy_data_warming_up` when no pre-range baseline
    exists (fresh upgrade). MQTT plugs are skipped from snapshots since they
    only report "today" and have no lifetime counter.

  Frontend: QuickStatsWidget renders an AlertTriangle next to Energy Used /
  Energy Cost with a tooltip when `energy_data_warming_up` is true, so the
  "low values right after upgrading" situation is explained in-product.
  Fully localised across 7 UI languages.

  Tests: new backend unit tests cover the snapshot delta arithmetic
  (baseline/endpoint, counter reset clamp, multi-plug, warming-up fallback,
  endpoint windowing), per-print restart resilience via expunge_all, and the
  snapshot task lifecycle (start idempotent, stop cancels). Frontend tests
  assert the warning icon appears only when the flag is set and only on the
  energy tiles.

  Docs: updated `CHANGELOG.md`, `README.md`, wiki `features/energy.md`,
  wiki `features/statistics.md`, and website `features.html` with the new
  behaviour and warming-up explanation.
2026-04-11 11:14:54 +02:00
maziggy b069b5217c y Fix virtual printer "Synchronizing device information" timeout in Orca (#927)
OrcaSlicer's "Send job" flow sat on "Synchronizing device information…"
  until it gave up, even though FTP upload worked when the user clicked
  "Send job anyway". The virtual printer's MQTT server gated all incoming
  command handling on `f"device/{self.serial}/request" in topic` — if the
  slicer's cached serial for the VP didn't exactly equal the VP's computed
  self.serial (model prefix + per-VP serial_suffix), every get_version,
  pushall, and project_file publish was silently dropped. Nothing was
  logged past the initial "MQTT publish to …" line, so the slicer never
  received a push_status or get_version response on its subscribed
  device/{serial}/report topic and hit its sync timeout. Responses were
  also unconditionally published on device/{self.serial}/report, so even
  when the inbound check happened to pass, replies targeted a topic the
  slicer wasn't listening on if its serial had drifted.

  Both directions are now serial-adaptive:

  - `_handle_publish` accepts any authenticated publish on a
    `device/*/request` topic and extracts the serial from the topic itself
    rather than comparing against self.serial.
  - A per-connection `_client_serials` dict tracks the serial the slicer
    actually uses, populated from the first SUBSCRIBE or PUBLISH seen on
    each connection and cleared on disconnect/stop.
  - `_send_status_report`, `_send_version_response`, `_send_print_response`
    now take an optional `serial` parameter (defaulting to self.serial)
    so every outgoing publish — including the periodic 1-second status
    push — targets the topic the slicer subscribed to.
  - The version response's embedded `module[].sn` fields now also carry
    the client's serial so the payload is internally consistent with the
    topic.
  - When the client's serial differs from self.serial an INFO log records
    the adaptation so it's visible in future support bundles.

  The working case (slicer's cached serial equals self.serial, as in my
  own H2D-1 Proxy setup) is bit-for-bit identical to the old behavior —
  the new check is strictly more permissive and only affects cases the
  old code silently dropped.

  Regression tests cover:
  - `_extract_serial_from_topic` valid/invalid topic shapes
  - mismatched-serial publish → handler runs, response topic and sn field
    both use the client's serial
  - non-`/request` topics → still rejected
  - pushall → status_report routed to the client's subscribed topic
  - `_client_serials` cleared on stop()
2026-04-10 12:13:54 +02:00
maziggy 7e7372c469 Fix SpoolBuddy update always pulling main instead of current branch
detect_current_branch() was reading .git/HEAD from settings.base_dir,
  which points at the data volume (DATA_DIR=/app/data in Docker) and
  never contains .git. The repo is at /app, so the lookup always failed
  and the code fell through to the GIT_BRANCH env-var → "main" fallback.
  The SpoolBuddy device was therefore checking out `main` regardless of
  which branch Bambuddy itself was running.

  The old subprocess-based implementation had the same bug but it was
  masked: the stock Docker image has no `git` binary, so `git rev-parse`
  raised FileNotFoundError, the except clause swallowed it, and the
  fallback kicked in. Swapping to filesystem reads exposed the wrong
  lookup path.

  Add a module-level _APP_DIR constant (parents[3] of the module file,
  same depth as config.py uses for its own _app_dir) and read `.git/HEAD`
  from there. A regression test plants a decoy .git in the data dir and
  asserts we still pick up the real one from the app root.

  Per your NO GIT WRITES rule, nothing is staged or committed.
2026-04-10 10:56:48 +02:00
maziggy 44cb26c7c3 Fix SpoolBuddy update Docker failure — set LOGNAME/USER/HOME in image
Follow-up to the asyncssh migration. asyncssh.connect() internally
  calls getpass.getuser() for ~/.ssh/config host matching, regardless
  of the explicit `username=` passed for the remote login. Under an
  arbitrary Docker PUID with no /etc/passwd entry, getpass.getuser()
  raises "No username set in the environment" (OSError in Python 3.13+,
  previously a bare KeyError).

  Fix: set LOGNAME=bambuddy, USER=bambuddy, HOME=/app in the Dockerfile.
  getpass.getuser() tries env vars before pwd.getpwuid(), so the lookup
  never touches the passwd database and works for any PUID the operator
  picks — no helper code, no image rebuild for different UIDs.

  Also pass config=[] to asyncssh.connect() so it does not try to load
  ~/.ssh/config (whose default path needs a resolvable home directory).

  An earlier draft of this fix added a Python helper that caught the
  KeyError and injected LOGNAME at module import. That was both more
  code than needed and broken on Python 3.13, which wraps the KeyError
  in an OSError the helper didn't catch — so the module import itself
  crashed, producing a 500 on /spoolbuddy/devices/{id}/update. Reverted
  in favour of the one-line ENV fix.
2026-04-10 10:46:06 +02:00
maziggy 60d034d40c Fix SpoolBuddy update Docker failure — asyncssh local-username lookup
Follow-up to the previous commit that swapped the `ssh`/`ssh-keygen`
  subprocesses for asyncssh. asyncssh.connect() internally calls
  getpass.getuser() to resolve the *local* username for ~/.ssh/config
  host matching, regardless of the explicit `username=` we pass for the
  remote login. Under an arbitrary Docker PUID with no /etc/passwd
  entry, getpass.getuser() tries LOGNAME/USER/LNAME/USERNAME (all unset
  in python:3.13-slim) and falls back to pwd.getpwuid(), which raises
  KeyError. asyncssh rewraps that as "Unknown local username: set one
  of LOGNAME, USER, LNAME, or USERNAME in the environment" — which
  surfaced in the UI as "ssh connection failed: no username set in the
  environment".

  Fix is two-part:

  - _ensure_local_username_env() runs at module import. If getpass
    .getuser() already works, or any of LOGNAME/USER/LNAME/USERNAME is
    set, it is a no-op. Otherwise it sets LOGNAME=bambuddy so asyncssh
    can proceed. Native installs are untouched.

  - asyncssh.connect() is now called with config=[] to skip the
    default ~/.ssh/config load, which relies on a resolvable home
    directory that may not exist under arbitrary Docker PUIDs.

  Three new unit tests cover the env-var fallback, including the case
  where the operator has set USER but the passwd lookup still fails.
2026-04-10 10:35:57 +02:00
maziggy a78a4bff2e Fix SpoolBuddy update still failing in Docker after keypair fix
Commit 67749565 eliminated ssh-keygen from the SpoolBuddy remote-update
  flow, but the update path still shelled out to the OpenSSH `ssh` client
  for every command. Like ssh-keygen, the `ssh` binary calls
  getpwuid(getuid()) during startup and aborts with "No user exists for
  uid <N>" when the container runs under an arbitrary PUID that isn't in
  /etc/passwd (python:3.13-slim only ships a root entry, so any
  `user: "1000:1000"` compose setup trips the same error).

  detect_current_branch() had a related problem: when the git repo is
  bind-mounted into the container, .git exists inside Docker, so the code
  tried to run `git rev-parse`. Git isn't in the image, so the subprocess
  silently fell back to the GIT_BRANCH env var — and if git ever were
  added, it could hit the same getpwuid trap.

  The entire update path is now subprocess-free:

  - _run_ssh_command uses asyncssh (pure-Python, built on the already
    installed cryptography library). Connection errors map to rc=255 to
    match `ssh`'s convention; asyncio.timeout handles the timeout path.
  - detect_current_branch reads .git/HEAD directly (handling git-worktree
    `gitdir:` pointer files too), keeping the same GIT_BRANCH → "main"
    fallback chain.
  - shutil and the inline `import subprocess` are gone from the module.

  Regression tests assert that neither keypair creation, branch
  detection, nor command execution spawns any subprocess. Native installs
  are unaffected.
2026-04-10 10:25:08 +02:00