Commit Graph
774 Commits
Author SHA1 Message Date
maziggy 75b0175e3d fix(camera): bound the post-kill wait on ffmpeg cleanup (#2580)
After an RTSP read timeout the stream cleanup killed the stalled ffmpeg
and then awaited process.wait() unbounded. A SIGKILLed ffmpeg stuck in
uninterruptible I/O on a dead RTSP socket can take arbitrarily long to
be reaped, so the fan-out stream coroutine sat parked in that wait (12
hours in the reported case) while every new viewer attached to the
stalled broadcaster and received no frames.

Bound the post-kill wait to 2s in all three places it existed: the
stream generator's _terminate_ffmpeg (the reported hang), the camera
stop endpoint (which would hang the recovery request itself; now uses
the shared helper instead of an inline copy), and the orphan-cleanup
janitor (whose hang would disable the safety net). On timeout the
zombie is abandoned; the janitor's /proc scan reaps it next pass and
the stream proceeds to its normal reconnect.
2026-07-17 07:03:10 +02:00
maziggy e97413edc7 fix(queue): enforce sliced-model compatibility on cross-model dispatch (#2578)
A queue item's "Any <model>" button labeled itself from the file's slice
metadata while the scheduler used the row's target_model, so an X1C-sliced
item targeting H2D showed "Any X1C" above "assign to first idle H2D". The
mismatch itself was created silently: sliced-for metadata loads async, and
switching to model mode before it arrived pre-selected the alphabetically
first model (H2D on a mixed farm), after which the model dropdown hid
itself. Nothing validated compatibility, so the scheduler would hand X1C
G-code to an H2D.

Frontend: never default the target silently, keep the dropdown visible in
model mode (incompatible models disabled), label from the actual target,
warn on mismatch, block submit when incompatible.

Backend: new GCODE_COMPAT_FAMILIES table (X1/X1C/X1E/P1P/P1S interchange;
everything else exact-match; missing metadata never blocks). Queue create
and update reject incompatible targets with 400; the scheduler holds back
pre-existing mismatched rows with an actionable waiting_reason instead of
dispatching them.
2026-07-17 06:51:51 +02:00
maziggy a6e7d671f2 fix(jog): stop disabling firmware endstops; warn that limits aren't enforced (#2579)
Manual jog could drive an axis past its travel limit into a collision.
Instrumenting the exact G-code to an H2D showed Bambuddy sending a clean
move at the limit (G91 / G1 Z-1.00 F600 / G90, no M211) that the printer
ran straight past, while its own touchscreen refuses the identical move.
This is a Bambu firmware bug: soft endstops are not enforced on G-code
received over MQTT, and no axis position is reported, so the move cannot
be clamped firmware- or client-side from position.

Two changes: (1) jogs no longer wrap moves in M211 S0/S1 — that disabled
the firmware's soft endstops globally, breaking even the touchscreen's
limits until a power cycle; a bare move keeps the touchscreen protected.
(2) The jog panel shows a prominent warning that travel limits are not
enforced during manual moves due to the firmware bug. Client-side
dead-reckoning enforcement is tracked separately.
2026-07-16 15:21:47 +02:00
maziggy 34818927c2 Revert "fix(db): configurable connection pool + auth_enabled cache for large farms (#2572)"
This reverts commit 4ad43c96de.
2026-07-16 09:30:36 +02:00
maziggy b3c0429373 fix(camera): release DB connection before streaming, not after (#2572)
/camera/stream took its printer row via Depends(get_db). get_db is a
yield dependency, so its session stayed open until the response body
finished streaming — for a live MJPEG stream, as long as the browser
tab is open (hours). Every open camera tile pinned one pooled DB
connection idle-in-transaction, draining the pool on large farms.

Fetch the printer in a short-lived async_session() and release the
connection before returning the StreamingResponse. expire_on_commit=
False keeps the already-loaded columns readable during the stream.
2026-07-16 08:38:42 +02:00
maziggy 4ad43c96de fix(db): configurable connection pool + auth_enabled cache for large farms (#2572)
Large PostgreSQL farms exhausted the fixed pool (pool_size=10 +
max_overflow=20): with ~93 printers every connection sat idle in
transaction and unrelated requests waited out the 30s pool timeout or
failed in the auth middleware.

- Make pool sizing env-configurable (DB_POOL_SIZE / DB_MAX_OVERFLOW /
  DB_POOL_TIMEOUT / DB_POOL_RECYCLE); raise the Postgres default to
  20 + 80 with pool_pre_ping + pool_recycle=1800.
- Cache the auth_enabled probe (30s) to drop a per-request DB round-trip.
  Only enabled=True is cached, so staleness fails closed; set_auth_enabled
  invalidates immediately.
- Add GET /api/v1/system/db-pool exposing resolved config + live
  checked_out/checked_in/overflow gauges without consuming a connection.

Session-hygiene (connections held across MQTT/FTP/camera/3MF I/O) is a
separate follow-up.
2026-07-16 08:16:48 +02:00
maziggy c2b23e5e61 Security hardening (maziggy/bambuddy-security #5) 2026-07-16 07:50:51 +02:00
maziggy 09b739b95d fix(cloud): stop reporting an expired Bambu Cloud sign-in as connected (issue #2562)
An expired token was indistinguishable from a working one. set_token()
stamped token_expiry = now + 30 days every time a stored token was loaded,
so the expiry reset on every request and is_authenticated could never
return False. /cloud/status answered "connected" for as long as any token
existed, while every cloud call 401'd — and the user was shown Bambu's own
{"error": "Please login."} verbatim.

Bambu is now the authority: /cloud/status validates the token upstream
(cached 5m), and any 401 from any authenticated call durably records the
credential as dead via users.cloud_token_invalid_at, so MakerWorld, cloud
profiles, slicer presets and firmware checks all agree at once. An
unreachable Bambu is treated as unknown, never as expired, so an outage
cannot sign a working session out.

The user-facing message now names the Profiles page, where the Bambu Cloud
sign-in actually lives; the old text pointed at a Settings page that does
not exist. Same stale path corrected in the wiki.
2026-07-14 11:29:56 +02:00
maziggy ce807fb1cc fix(queue): upload to printers in parallel, cap wedge retries, make debug logs survive a farm
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.

Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.

FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.

Two problems the same bundle exposed:

A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.

The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.

Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).

Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
2026-07-14 10:35:58 +02:00
maziggy a0d4b3d837 fix(queue): scope force-colour overrides to the plate the item prints (#2551)
Queueing several plates of one 3MF built a single filament-override list from
every selected plate and posted that same list with each plate's item. A
force_color_match entry blocks dispatch until the printer has that exact colour
loaded, so a single-colour plate waited on the whole batch's palette. The same
shared list also widened required_filament_types, making a PLA plate refuse
every printer that lacked a sibling plate's PETG.

Narrow the overrides to the slots the plate actually consumes, on create and on
update -- in the backend, where the 3MF is, so it holds for every writer of the
queue. Dispatch already re-parsed requirements per plate and keyed overrides by
slot, so the dropped entries were inert there. When the plate's slots cannot be
read the overrides are kept whole: an item waiting on a colour it does not need
is visible, one that silently lost a forced colour prints in the wrong filament.

Items queued before this would stay stuck with a waiting reason that explains
nothing, so a startup migration re-scopes the pending ones. Printing and
finished items keep their overrides -- that is a record of what they dispatched
with, not an instruction.
2026-07-13 09:18:17 +02:00
maziggy c640ddc1f7 fix(projects): carry tags, due date and priority in the list payload (#2536)
The edit dialog is shared between the projects list and the project detail
page and seeds itself from whichever project object it is handed. The list
payload never carried tags, due_date or priority, so editing from the list
showed a blank tags field -- and, unreported, submitted the dialog's default
priority over a stored high/urgent one. The component read those fields
through a cast, so the compiler never flagged that they were always absent.

Put them on ProjectListResponse and ProjectListItem, drop the casts, and let
an explicit null clear tags and due date the way it already clears budget and
url -- an emptied field was previously sent as undefined and silently reverted.
The template list was missing target_parts_count, which the same dialog edits.
2026-07-13 08:59:36 +02:00
maziggy 5bbfeefa65 fix(backup): diagnose an unwritable backup path instead of quoting errno 30 (#2544)
Nightly backups to a mounted NAS share ran from May and then stopped, failing
with [Errno 30] Read-only file system. The reporter checked folder permissions
-- correctly: the mount is gid=backup,dir_mode=0775, the service user is in that
group, and his own shell writes to the share fine.

Errno 30 is EROFS. A permission problem is errno 13. EROFS means the filesystem
refused the write, and it refused because we told it to: our systemd unit ships
ProtectSystem=strict, which mounts everything read-only inside the service's
mount namespace and carves back out only ReadWritePaths=<install> <data> <logs>.
A NAS share is not one of those three. Reads are unaffected -- which is why the
UI happily listed his existing backups from the share while being unable to
write a new one -- and his shell is outside the namespace entirely, so every
check he could think to run said the directory was fine.

Both installers write the unit file wholesale, so a ReadWritePaths line added by
hand disappeared on the next install, taking the backups with it. They now back
the old unit up (.bak-<timestamp>) and carry the operator's extra writable paths
forward, reporting which ones they kept. The unit template documents the
carve-out.

The output directory is probed with a real write when it is saved and when the
backup card loads, so an unwritable path is caught there rather than at 03:00
for a week. On failure the card names the cause and hands over the fix with the
operator's path already in it (systemctl edit bambuddy -> ReadWritePaths=...),
and a failed run reports the same diagnosis rather than the raw OSError. EROFS
outside systemd, permission-denied, out-of-space, not-a-directory and missing are
told apart, in all 11 locales.

Docker: a backup path that is not bind-mounted is writable -- the write lands in
the container's ephemeral layer and is lost on the next compose up. The probe
compares the directory's device against the container root and warns, with the
compose snippet that mounts it properly.
2026-07-12 08:44:53 +02:00
maziggy aba00598bb fix(smart-plugs): read a REST plug's lifetime counter, and derive Today/Yesterday from it (issue #2539)
A Shelly reports one energy figure — aenergy.total, a lifetime counter in Wh
that never resets. Bambuddy had a single REST energy field and filed whatever
it found under "today", so the value never reset at midnight, and Yesterday
and Total stayed at zero: get_energy() simply never set those keys.

With `total` unpopulated, the hourly snapshot recorder skipped the plug, so
the Statistics page's energy figure was zero as well, not just the Settings
card.

Split the REST energy config in two: rest_energy_path still means "used
today", rest_energy_total_path means "lifetime counter". A Shelly has only
the latter; a Tasmota behind a REST bridge has both; sharing a URL costs one
fetch, not two.

Then derive Today and Yesterday from that counter using the snapshots we were
already taking: today = counter now - counter at the last local midnight;
yesterday = the gap between the two previous midnights. Local midnight, not
UTC — a UTC boundary rolls Today over at 02:00 in Berlin. The snapshot loop
now ticks on the local hour so a reading lands on the boundary instead of up
to an hour early. A counter that goes backwards (factory reset) reports
nothing rather than a negative.

Collateral, found while verifying on both engines: the smart-plug DateTime
columns are naive UTC but the code wrote aware datetimes into them. SQLite
drops the offset; asyncpg raises DataError. So on Postgres every snapshot
capture raised inside the loop's except, and every status poll raised on
last_checked — the whole subsystem was dead on the database we recommend for
multi-printer installs. All plug timestamps are naive UTC now.

Existing REST users with a cumulative path in the today field must move it to
the new lifetime field; the form and wiki now name which counter each wants.
2026-07-11 14:22:31 +02:00
maziggy d09db436c3 feat(camwall): serve the Cam Wall at /camwall, and on a token-authenticated kiosk
Cam Wall had no URL — the only way in was the toggle on the Printers page,
so it could not be bookmarked, linked, or shown on a wall-mounted screen.

Add a standalone /camwall route. Signed in, it is the wall as it was. For a
TV or Pi with no login, it authenticates with a long-lived token in the URL.

A kiosk needs the printer list and per-printer status, both of which sit
behind PRINTERS_READ. Rather than widen camera_stream to cover GET /printers
— whose response carries serial_number and ip_address, which have no business
on a screen in a shared room — add a read-only feed at
GET /api/v1/camwall/printers that serves only what a tile draws, and gate it
on a new camwall token scope. The print filename is not served at all: a token
wall renders the compact overlay, so the part on the bed is never named.

The scope is separate rather than a widening: camera_stream tokens are already
in the wild, minted to hand out video, and must not gain the ability to
enumerate a fleet by name. camera_stream is refused by the feed; camwall
passes the stream gate so its own tiles fill.

Kiosk walls drop the settings popover and click-through entirely (not merely
hidden — a passive screen must carry no focusable control it cannot act on),
cap the overlay at compact, and poll rather than open a WebSocket. maxLive,
interval and status can be set from the URL, clamped to the popover's ranges.
2026-07-11 13:38:15 +02:00
maziggy ca3f6e5ee0 fix(drying): P1 AMS drying is screen-only — stop offering it (#2533)
The reporter found what his P1S was doing, and it is in Bambu's P1 manual:
"P1S connected AMS drying functions may only be controlled from the P1S screen."
The firmware acks ams_filament_drying with result: success and then discards it,
which is why three commands on an idle printer left the AMS 2 Pro at dry_status 0.
No command can start a cycle on a P1, on any firmware, so don't offer one.

supports_drying() now excludes the P1 series outright, replacing the 01.08+ gate
carried since #292 — that version is when P1 firmware gained AMS 2 Pro support,
not remote drying, and it was never checked against a live P1. Both drying routes
refuse with a specific 400 instead of publishing a message the printer will drop;
queue and ambient auto-drying skip P1s via the same helper.

A new drying_screen_only flag keeps the control on the card, disabled, saying why
— a P1 owner needs to learn where to dry, not watch the button disappear. A cycle
started at the printer still shows with its countdown; only Stop goes away, since
a P1 ignores stop exactly as it ignores start.

Also corrects the wiki firmware matrix, which listed P1P/P1S as supported and
(separately) P2S/H2S/H2C as unsupported. 8 tests.
2026-07-11 09:33:52 +02:00
maziggy ce31de65c2 fix(skip-objects): scope the object list to the plate being printed (#2522)
extract_printable_objects_from_3mf() has accepted a plate_number since it was
written and no caller ever passed one, so it took root.find(".//plate") — the
first plate in the file. Passing one would not have helped either: the lookup
was .//plate[@plate_idx='N'], a predicate on an attribute neither Bambu Studio
nor OrcaSlicer writes. The index lives in a <metadata key="index"> child, as
threemf_tools and filament_requirements already read it, so the selector never
matched and fell back to plate 1 regardless.

On an all-plates .gcode.3mf that meant Skip Objects offered the wrong plate's
objects, with that plate's marker positions drawn over the correct plate's
thumbnail (/cover resolves the plate properly via resolve_plate_id, the object
list did not). The reporter printed a one-object plate and was shown the four
copies from another plate of the same file.

Select the plate on its index metadata, and pass resolve_plate_id(state) at all
three call sites so the list and the thumbnail share one resolver. Also stop
peek_plate_index_in_3mf() reporting plate 1 for a multi-plate file: it backs the
running one, so an all-plates upload printing plate 2+ lost its archive entirely.
2026-07-11 09:18:35 +02:00
maziggy 50c3e94d33 fix(cloud): log expected preset misses at DEBUG, keep real faults at WARNING
Failed to get cloud preset ... 400 {"message":"missing"} is the expected
answer, not a fault: many official presets are only addressable with a
printer-variant suffix (GFSL05 exists solely as GFSL05_07 @BBL A1), and
personal P-prefixed presets belong to the account that sliced the file.
Phase 3 already resolves both from local presets, so the lookup miss is
routine -- and one WARNING per AMS tray per tooltip refresh teaches
operators to ignore the log.

BambuCloudError now carries the upstream status_code. The preset lookup
logs HTTP 400 at DEBUG; expired tokens, 5xx and transport failures stay
at WARNING.

Not fixed here: resolving the variant suffix. It selects a printer profile
and the response carries that profile's pressure_advance, so guessing a
suffix would report another printer's K value.
2026-07-10 08:24:20 +02:00
maziggy e6136b660b fix(cloud): carry Bambu Cloud credentials across the auth on/off boundary
get_stored_token() reads the global Settings rows when auth is disabled and
User.cloud_token when it is enabled, so completing /auth/setup switched which
store the /cloud/* routes consult without moving the token. An account linked
before enabling auth was stranded: build_authenticated_cloud() returned None,
get_filament_info() skipped its cloud phase and answered 200 from local
fallbacks, and /cloud/devices began returning 401 -- all silently.

setup_auth() now migrates the global token onto the owning admin and deletes
the global rows; disable_auth() mirrors the hand-off back. Neither guesses:
setup migrates only when it creates the admin or exactly one exists, disable
declines to overwrite an existing global token. Region survives both hops.

Instances that already crossed the transition must re-link once.
2026-07-10 08:15:35 +02:00
maziggy e1e2c12d25 fix(library-tags): declare response_model=None on the 204 DELETE route
Under `from __future__ import annotations` the `-> None` return annotation
reaches FastAPI as the string "None", which resolves to NoneType -- truthy,
so APIRoute asserts a 204 may carry no response body and the app fails to
import. fastapi >= 0.116 guards against this; the 0.109-0.115 releases
requirements.txt still allows do not.
2026-07-09 16:29:13 +02:00
maziggy f6c6cfbad3 fix(ams): show "?" not "Empty" for non-RFID spools using tray_exist_bits (#2527)
A spool with no readable RFID was reported by the standard AMS with an empty
tray_type and state=9 — structurally identical to a truly-empty slot at the
tray level — so the AMS card rendered it "Empty" while Bambu Studio correctly
showed "?". The authoritative "a spool is physically here" signal is firmware's
AMS-level tray_exist_bits bitmask (what Studio uses), but Bambuddy inferred
emptiness from the per-tray state/tray_type. Confirmed from the reporter's
bundle: tray_exist_bits=f (all four slots present) with tray_is_bbl_bits=5
(only slots 0,2 Bambu) — the present-but-non-Bambu slots were the ones shown
Empty. Supersedes closed #1838.

apply_tray_exist_bits() already parses the bitmask to clear stale fields on
absent slots; it now also annotates each slot with an authoritative `exists`
bool, gated behind a new annotate_exists flag so only the printer-card path
sets it. The VP bridge leaves it off, so the `exists` key never reaches the
slicer wire format. `exists` flows through the AMSTray schema/serialization to
the frontend, where getEmptySlotKind() uses it: exists===true + no tray_type
-> "?" (present, unconfigured), exists===false -> "Empty", exists absent ->
the previous state=9/10 heuristic (AMS-HT and missing-bitmask paths unchanged).
H2D/X1C already reported present-unknown slots with a non-9 state and took the
"?" path; with the fix they reach it via `exists` and are unaffected.
2026-07-09 09:27:13 +02:00
maziggy 9e7f6cafd9 fix(backup): preserve NOT NULL/DEFAULT/FK/UNIQUE in Postgres→SQLite backup (#2526)
On a PostgreSQL install, create_backup_zip() exports a portable SQLite copy
so backups move between engines. It rebuilt each table with only column name
+ type + PK, dropping NOT NULL, server_default/DEFAULT, foreign keys, and
unique constraints. Restore onto SQLite page-copies that schema straight onto
the live database, and post-restore init_db() can't repair it (create_all is
CREATE TABLE IF NOT EXISTS). So server_default columns like
spoolbuddy_devices.created_at (server_default=func.now()) ended up with no
DEFAULT: SQLAlchemy omits them on INSERT, the DB wrote NULL, and the next read
500'd on Pydantic validation. Every server_default column was exposed the same
way; the FK/unique loss followed from the same simplified CREATE TABLE.

Build the portable schema with Base.metadata.create_all() against a SQLite
engine instead of the hand-rolled loop, so it emits the exact DDL a native
SQLite install gets (NOT NULL, DEFAULT func.now() -> CURRENT_TIMESTAMP, FKs,
unique constraints, indexes). The data-export insert path is unchanged, and
the #1333 OIDC-icon guard is preserved automatically (LargeBinary -> BLOB),
which lets the now-redundant _sqlalchemy_type_to_sqlite_type() helper be
removed. Fixes newly-created backups; a backup from an older build still
carries the degraded schema, so re-take backups after upgrading.

Replace the #1333 type-mapping unit tests with three that inspect the real
backup schema via metadata.create_all + PRAGMA table_info: icon_data is BLOB,
created_at keeps its CURRENT_TIMESTAMP DEFAULT, a NOT NULL non-PK column stays
NOT NULL.
2026-07-09 09:02:34 +02:00
maziggy d03b108965 Fix external-folder scan deleting README.md records; index markdown (#2520)
.md was missing from _SCANNABLE_EXTENSIONS, so scanning an external
folder skipped markdown during the walk and the cleanup pass deleted
its LibraryFile row (assuming it was gone from disk), 404ing the Folder
Readme panel. Add .md to the scannable set so pre-existing markdown is
indexed, and gate cleanup deletion on actual disk presence rather than
absence from the extension-filtered found_paths, so any non-scannable
upload still on disk survives a scan.
2026-07-09 07:20:55 +02:00
maziggy 917bfd7666 feat(labels): scannable QR on 203 dpi thermal printers + monochrome mode (#1870)
The 40x30 mm box label rendered its QR too densely for low-res thermal
printers — the modules bled together and wouldn't scan. Two causes: the QR
was 20% of inner width (~7.5 mm on the narrowest template, half of the
others) and used ERROR_CORRECT_M. Fix adaptively so all templates benefit:
give the roomy-layout QR a 12 mm minimum size (box_40x30 -> 12 mm, ~3.5
dots/module at 203 dpi) and switch label QRs to ERROR_CORRECT_L (same
payload, chunkier modules; a label needs no M-level recovery). Keep the
quiet-zone border at 2 — the size+L gains suffice without risking scans.

Also add a Monochrome (black & white printer) option to the label dialog:
drops the colour swatch (a useless grey block on B&W) and widens the text;
the hex-code line still carries the colour. Threaded through the renderer,
route, API client, and modal, with translations in all 11 locales.
2026-07-07 10:31:19 +02:00
maziggy 168d9d8f8e fix(auth): let API keys manage projects via new can_manage_projects scope (#1893)
PROJECTS_CREATE/UPDATE/DELETE were in _APIKEY_DENIED_PERMISSIONS with no
entry in _APIKEY_SCOPE_BY_PERMISSION, so every project mutation returned a
generic 403 for any API key regardless of granted permissions -- the same
regression class as archives (#1888) and library (#1832).

Add a per-key can_manage_projects scope. Project routes gate on plain
PROJECTS_* (no OWN/ALL split), so all three CRUD permissions map to the one
scope; membership edits (add-archives) gate on PROJECTS_UPDATE and are
covered. PROJECTS_READ is unchanged (already under can_read_status).

Column defaults TRUE for new keys; existing rows backfill to FALSE so the
upgrade never silently widens scope. Migration is BOOLEAN (SQLite + Postgres
safe), verified on fresh SQLite and Postgres 17. Bundled SpoolBuddy kiosk key
set to False. Settings API-key UI gets a Manage Projects toggle + Projects
badge; 11-locale i18n. RBAC scope matrix + drift guards extended.
2026-07-05 09:58:16 +02:00
maziggy d568307eac fix(smart-plug): don't cut power when a print restarts, honor per-plug cooldown setting (#1890)
The print-queue "auto off after this job" trigger used a second, inline
auto-off implementation (main.py, print_scheduler.py, print_queue.py)
that hardcoded wait_for_cooldown(50C, 600s) — ignoring each plug's
configured off_delay_mode / off_delay_minutes / off_temp_threshold — and
ignored the return value, powering off on the 600s timeout regardless of
print state. A print that failed and was reprinted from the touchscreen
got its power cut mid-print. The inline tasks were also uncancellable, so
a reprint couldn't abort a pending off.

Consolidate all three into SmartPlugManager.schedule_off_after_queue_job,
which schedules via the plug's configured strategy (shared with
on_print_complete through _schedule_off_per_mode) and is cancellable via
_pending_off. Add printer_manager.is_print_active() and guard the actual
power-off in _delayed_off and _temp_based_off so no path cuts power on a
loaded print. Move the on_print_start cancellation ahead of the auto_on
gate so a reprint always aborts a pending off.
2026-07-03 08:32:51 +02:00
maziggy 6358e9544e fix(auth): allow API keys to delete/edit archives via new can_manage_archives scope (#1888)
DELETE /api/v1/archives/{id} rejected every API key with 403
"API keys cannot be used for administrative operations", regardless of
the print's owner or the key's scopes. ARCHIVES_DELETE_ALL/_OWN (and the
create/update variants) were on the denylist and absent from the scope
allowlist, so require_ownership_permission fell through to the generic
admin-denied 403 — the whole archive-management surface was unreachable
for API keys. Same regression class as the #1832 library/maintenance
carve-outs.

Add a can_manage_archives per-key scope: ARCHIVES_CREATE, ARCHIVES_
UPDATE_OWN/_ALL and ARCHIVES_DELETE_OWN/_ALL move from the denylist to
the allowlist under it (OWN and ALL fold into the same scope, matching
can_manage_library). ARCHIVES_PURGE stays admin-only — it drops the
print's Quick Stats contribution, mirroring LIBRARY_PURGE. Column
defaults TRUE for UI-created keys; existing rows backfill to FALSE so the
upgrade never silently widens scope. Bundled SpoolBuddy kiosk key stays
minimally scoped (False). Migration is dialect-agnostic and verified on
fresh SQLite and Postgres 17.

Adds the Settings API-key toggle + badge (11-locale i18n) and extends the
RBAC scope matrix to cover all five archive-management permissions.
2026-07-03 08:01:54 +02:00
maziggy bdac27ebee fix(slicer): preserve PVA-for-support intent across re-slice of source 3MF (#1881)
Three bugs on the same PLA-model + PVA-support flow, discovered in
sequence:

(A) substitute_unused_plate_filaments inspected only object geometry
    (per-object extruder metadata + paint_color triangles) so a support-
    only slot was silently treated as "unused" and the user's PVA profile
    got overwritten with slot 1's PLA.

(B) _extract_filament_info stripped filament_is_support==1 entries,
    hiding PVA from unsliced source archive cards even when the project
    explicitly configured it.

(C) --load-settings is authoritative over the source's project_settings.
    config, and Bambu's shipped process presets ship enable_support=0
    (supports are a per-print decision, not per-quality). So even with
    (A) fixed, the sliced output had supports disabled and the PVA slot
    loaded but never consumed. Inverts BambuStudio GUI's semantics where
    the project overrides the preset.

Fixes:
- New extract_support_filament_slots_from_3mf reads enable_support +
  support_filament + support_interface_filament from project_settings.
  config; substitute_unused_plate_filaments unions it into the geometry-
  derived set.
- _extract_filament_info returns all configured filament types + colours.
- New _patch_process_support_settings overlays four fields (enable_
  support, support_filament, support_interface_filament, support_type)
  from the source 3MF onto the picked process preset JSON before
  --load-settings sees it. Deliberately targeted to what fixes #1881
  without widening to a full project-over-preset merge.
2026-07-02 11:32:22 +02:00
maziggy 006c3113a0 feat(api-keys): can_manage_maintenance scope for HA-style automations (#1832 follow-up)
Carve MAINTENANCE_CREATE/UPDATE/DELETE out of the admin denylist so
HA automations can log "cleaned nozzle" / reset a counter via API key
without granting broader printer control. Follows the same shape as
can_manage_library and can_manage_inventory: new column, allowlist
entry, UI checkbox, wiki row, RBAC test coverage.

Distinct backfill: these perms were EXPLICITLY denied for every API
key before this change (no existing integration relies on them), so
existing rows migrate to FALSE — no silent scope widening on upgrade.
New keys default to TRUE, matching the safe-on-by-default pattern.
Bundled SpoolBuddy kiosk key gets False explicitly (kiosk doesn't need it).
2026-07-01 09:21:09 +02:00
maziggy 61a7f2e4ac feat(scheduler): preheat & heat-soak before queued prints with per-filament chamber targets + airduct flap control (#1468)
New scheduler stage that heats the bed (and the chamber, on supported
printers) and holds at temperature before each queued print starts —
the heat-soak engineering filaments need for adhesion and warp
control. Bambuddy waits between FTP upload and start_print, so the
soak runs while the printer is otherwise idle. M191 is silently
ignored by Bambu firmware, so doing this at the orchestration layer
is the only place it works.

Resolution order at dispatch:

1. PrintQueueItem.preheat_override ∈ {inherit, on, off}.
   'off' skips entirely; 'inherit' falls back to the global
   preheat_enabled toggle; 'on' forces the stage even when the
   global is off.

2. chamber_target = item.preheat_chamber_target_override
                 ?? max(filament_map[normalize(t.tray_type)] for loaded slots)
                 ?? 0.
   Mixed PA+PLA picks PA's 50 (max-across-slots — PA's chamber
   requirement is binding, PLA doesn't suffer being warm). PLA-only
   derives 0 and skips the chamber phase automatically.

3. Three hardware tiers for chamber heat:
   - Active chamber heater (H2C/H2D/H2D Pro/H2S/X2D/X1E) → M141 +
     chamber-sensor wait
   - Chamber sensor only (X1C/P2S) → no M141, passive bed-radiation
     wait with hard max-wait cap
   - No chamber sensor (P1S/P1P/A1/A1 Mini) → bed + soak timer only

4. Airduct flap (H2C/H2D/H2D Pro/H2S/X2D/P2S) auto-switches to
   match the chamber target — heating mode for engineering
   filaments, cooling mode for PLA. Bambu firmware does NOT
   auto-switch the flap with M141, so without this an ABS print
   on a previously-cooling flap fights the open exhaust, and a
   PLA print on a previously-hot flap recirculates ABS heat.
   Idempotent: only fires set_airduct_mode when current ≠ desired.

Settings → Workflow → Queue & Dispatch → Preheat & Heat Soak card:
master enable toggle (default off — disabled installs see no change),
per-filament chamber-target editor (replaces a single global int that
shipped in the first cut and couldn't serve PA + PLA in the same
config), preheat_max_wait_seconds, preheat_soak_seconds. The Print
Options panel in PrintModal gets a Preheat sub-section with the
tri-state Inherit/On/Off control and an optional chamber-target
override input.

DB migration: PrintQueueItem gains preheat_override VARCHAR(10)
DEFAULT 'inherit' and preheat_chamber_target_override INTEGER NULL.
Idempotent via _safe_execute. Existing rows behave exactly as before
the migration.

Best-effort throughout: printer drops, refused M141 or set_airduct,
missing bed temp, lost MQTT state mid-wait all log and return cleanly.
Normal upload + start path runs after this returns regardless.
2026-06-29 12:35:43 +02:00
maziggy a45d32efd0 fix(hms): wrong-plate Ignore actually ignores + buttons read as buttons + ack-detection survives transient re-pause (#1869)
The HMS error modal had three compounding bugs that surfaced when a
user forced a wrong-plate HMS (0500_8051) and tried to dispatch the
per-fault actions.

(1) IGNORE_RESUME did not ignore. Bambuddy redirected the action on
state=PAUSE to a plain `resume` command, citing a #1830 verdict that
BambuStudio's "err-bearing shape" was firmware-silently-rejected.
BambuStudio source disagrees: DeviceErrorDialog.cpp:600 dispatches
IGNORE_RESUME via command_hms_ignore, whose wire shape is
{command:"ignore", err:"<decimal>", param:"reserve", job_id:...}.
That's a distinct command from `resume` — the firmware suppresses
the next re-check AND auto-resumes in one operation. Plain resume
means "re-check normally", which is exactly why the wrong-plate
detection re-fired 1-2 s after the user clicked Ignore. The #1830
"err-bearing shape rejected" test almost certainly sent the err as
a hex shortcode; BambuStudio passes std::to_string(int m_error_code)
i.e. the DECIMAL form, which is what the firmware matches against.

(2) Action buttons read as inert badges. The button className used
`hover:${buttonHoverColor}` — a template-literal interpolation
Tailwind's JIT scanner can't see as a literal string, so the
per-severity hover utility never reached the compiled CSS. Same
bg/text color as the severity badge above and no border made it
read as another label. No disabled state and no spinner during the
2.5 s ack wait left clicks sitting silently inert.

(3) Ack-detection 502'd on legitimate ack. The route compared
(gcode_state, hms_errors-len) before vs after publish; wrong-plate
re-pause round-tripped both fields to their pre-publish values
inside the 2.5 s window → false 502 even though the firmware fully
ack'd. PROBLEM_SOLVED_RESUME working but IGNORE_RESUME 502'ing on
the same fault was the same race resolving differently.

Fixes:

bambu_mqtt.py — new hms_ignore_command() publishes the BambuStudio
shape; existing hms_ignore(persistent) renamed to hms_idle_ignore
(unchanged shape, used by NO_REMINDER_NEXT_TIME per
DeviceErrorDialog.cpp:588). Dispatch routes IGNORE_RESUME,
IGNORE_NO_REMINDER_NEXT_TIME, and DONT_REMIND_NEXT_TIME to
hms_ignore_command (BambuStudio routes all three to the same
command_hms_ignore — the "don't remind" half is the firmware's
job). NO_REMINDER_NEXT_TIME stays on hms_idle_ignore type=0. Hex →
decimal err conversion at the helper layer with a defensive
fallback. job_id=None → empty string (matches BambuStudio's
std::string default).

HMSErrorModal.tsx — getSeverityInfo loses the dead buttonHoverColor
field. Action button uses static
`bg-white/10 hover:bg-white/20 active:bg-white/30 text-white
border border-white/20`, wires
`disabled={!hasPermission||mutation.isPending}`, and renders
`<Loader2/>` only on the button whose (action,print_error) matches
mutation.variables.

printers.py — ack-detection probes `client._last_message_time`
(bumped on every MQTT push regardless of payload) rather than
diffing state fields. The pushall that follows every command
guarantees a fresh push lands inside the 2.5 s window on any
healthy printer; only firmware-silent-drop leaves the timestamp
untouched, which is the 502 path #1830 wanted.
2026-06-29 10:59:29 +02:00
maziggy 425a3ac404 fix(slicer): surface real CLI rejections + hard-skip mismatched filaments in auto-pick (#1851)
Two compounding bugs let an H2C-bound filament land in slot 1 of an A1
slice silently. (1) `_slicer_rejection_message` discarded the actual CLI
diagnostic - `filament preset Generic PLA @BBL H2C (slot 1) is not
compatible with printer Bambu Lab A1 0.4 nozzle.` - when the sidecar's
headline error_string was Bambu Studio's catch-all
`The input preset file is invalid and can not be parsed.` placeholder.
The real reason was in the stdout `[error] run NNNN:` line, trimmed off
before reaching the SliceJob's error_detail. (2) `pickFilamentForSlot`
used a soft `-100` mismatch penalty rather than a hard skip, leaving
the "never auto-fill an incompatible preset while a compatible one
exists" contract implicit. The unused-slot substitution in
`substitute_unused_plate_filaments` then propagated whatever slot 1
held across every unused slot - one bad pick poisoned the array.

(1) Mine `[error] <msg>` (with or without `run NNNN:`) from the full
pre-trim response; substitute the placeholder, keep meaningful
headlines. (2) Partition candidates into compatible/unknown vs
mismatch; prefer compatible whenever the bucket is non-empty, fall
back to mismatch only on graceful-degrade. Picker helpers moved out
of `SliceModal.tsx` into `utils/slicePresetPicker.ts` so the modal
file stays component-only (react-refresh lint).
2026-06-29 08:23:29 +02:00
maziggy b23cb69a66 fix(permissions): self-heal Administrators to ALL_PERMISSIONS on upgrade + Pipelines runs dashboard polish
Administrators system group sync
- Fresh installs already bootstrap with ALL_PERMISSIONS, so they always have
  every permission. Upgrades previously only got what one-off backfill blocks
  in seed_default_groups() explicitly listed (library:purge, archives:purge,
  the OWN/ALL read-flag block, orca_cloud:auth, pipelines:*). Any Permission
  enum member added without a matching block silently stayed missing on
  existing admin rows. The most recent gap was printer_sensor_history:read
  (Sensor History charts returned 403 for upgraded admins).
- seed_default_groups() now syncs Administrators to ALL_PERMISSIONS on every
  startup: append every Permission value that isn't already on the row.
  Additive only -- hand-added custom permissions are preserved.
- The pure-admin one-off backfills (library:purge / archives:purge block,
  the OWN/ALL + orca_cloud:auth + legacy-read-flag block, the Administrators
  branch of the pipeline backfill) are retired since the sync subsumes
  them. Non-admin backfills (Operators / Viewers OWN-tier reads, Operators
  orca_cloud:auth, pipelines for non-admin groups, makerworld:*, clear_plate
  cross-group adders) are untouched.
- Tests: test_administrators_printer_sensor_history_read_backfilled
  (regression for the reported gap),
  test_administrators_sync_covers_every_current_permission (generic
  invariant -- any future new permission lands on admin without needing
  a one-off test), test_administrators_sync_is_additive_only (custom
  permissions preserved). 12/12 backfill-migration + 102/102 broader
  permission tests green; ruff clean.

Pipelines runs dashboard
- PipelineRunsPage.tsx: the Pipeline / Status / Target filter row's three
  native <select> elements are replaced with a bambu-themed FilterDropdown
  (button trigger, floating menu, optgroup-style headers for the Target
  picker, hover + selected states with a check mark, closes on outside
  click and Escape). Same value/onChange contract -- visual only.
- SlicerPipelinesPanel.tsx: wrap list?.pipelines ?? [] in useMemo so the
  reference is stable when the data is stable. Fixes the
  react-hooks/exhaustive-deps warning where the inline fallback returned
  a fresh empty array every render, invalidating both downstream useMemo
  caches (target-options + filtered-pipelines list).
2026-06-28 11:18:18 +02:00
maziggy 3ef197e4e0 feat(slicer): Pipelines — multi-copy + class targeting + fanout + runs dashboard + retry-failed + WS updates (#1425 PR C — completes the v3 design)
PR A/B turned the slice modal's preset bundle into a one-click dispatch
with a pinned target printer. PR C closes the original issue: operators
type in a number of copies, Bambuddy slices once and distributes prints
across a fleet per the pipeline's chosen fanout strategy. A new dashboard
surfaces every run with filters, expandable per-copy status, cancel,
and retry-failed-copies. WS pushes keep everything live.

Backend
- copies field on POST /run, capped by new pipeline_max_copies setting
  (default 50, hard cap 1000). PipelineRun.parent_run_id chains retries.
- SlicerPipelineUpdate accepts target_kind (specific_printer /
  printer_class), target_model_class, fanout_strategy.
- Eligibility matcher branches: class-targeting enumerates matching
  Printer rows, runs per-printer checks via a status_lookup closure,
  returns printer_reports[]. New issue kinds: no_class_matches,
  class_not_set.
- _pick_assignments distributes copies per strategy:
  - max_parallel: target_model set, printer_id None — scheduler picks
  - round_robin: copy i → eligible[i % N], fixed printer_id
  - fill_one_first: all copies pinned to eligible[0]
  All three reuse the slice-once path through slice_dispatch.enqueue.
- New routes:
  - GET /pipeline-runs (paginated, filterable by pipeline + status)
  - POST /pipeline-runs/{id}/retry-failed (creates child run with
    copies = failed+cancelled count, parent_run_id set)
  - Cancel cascades to all N queue entries (only pending/queued)
- _roll_up_run_status computes run-level status from per-job statuses;
  introduces partial_failure for "some completed, some failed".
- ws_manager.broadcast_to_user emits pipeline_run_updated on every
  state transition with the full materialised response.

Frontend
- Pipeline editor: target_kind radio + class picker (filtered to
  installed models) + fanout-strategy radio. Read-only row shows
  "X1C · Round robin" for class pipelines.
- RunWithPipelineModal: copies number input bounded by
  settings.pipeline_max_copies. Accepts class-targeted pipelines.
- Settings → Workflow → Queue & Dispatch: new "Slicer Pipeline limits"
  card with the max-copies input.
- New /pipelines/runs dashboard page (sidebar entry, gated on
  pipelines:read). Two-filter dropdown, 25-per-page pagination, per-row
  expandable to job list, Cancel + Retry-failed buttons.
- useWebSocket case for pipeline_run_updated invalidates both
  pipeline-runs-all and pipeline-runs/{id} query keys.
2026-06-27 16:52:05 +02:00
maziggy 4bbf0f031e feat(slicer): Pipelines — archive entry point + slicer progress toast (#1425 PR B follow-up)
Two real gaps from the PR B drop:

1. Run-with-pipeline only existed in the file manager. Operators who keep
   working files in archives had to copy them to the library to use a
   pipeline.

2. Triggering a slice via a pipeline produced a silent multi-second-to-
   minute wait. The manual SliceModal flow shows the sticky
   "Slicing X - Generating G-code 75%" persistent toast; the pipeline
   path went through asyncio.create_task directly and never registered
   with SliceJobTracker.

Archive entry point
- POST /slicer-pipelines/{id}/check-eligibility and /run accept
  source_archive_id as an alternative to source_library_file_id (XOR,
  enforced by Pydantic validator).
- PipelineRun.source_archive_id is a new nullable FK column with the
  ALTER TABLE migration in run_migrations (idempotent via _safe_execute,
  works on SQLite + Postgres).
- _resolve_source branches: archive path reads source_3mf_path with
  fallback to file_path, mirroring routes/archives.py.
- ArchiveCard's context menu picks up a "Run with pipeline" item next to
  Slice (only on source archives), gated on useSlicerApi + pipelines:run.
  Slice (only on source archives), gated on useSlicerApi + pipelines:run.
- Path-safety: SEC-PATH-OK markers added at both LibraryFile.file_path
  and archive.source_3mf_path join sites, citing the upload-time
  validators.

Progress toast
- Pipeline orchestration is now the `run` callable of a
  slice_dispatch.enqueue call — the same dispatcher SliceModal uses —
  instead of a bare asyncio.create_task. The SliceJob lifecycle drives
  the existing progress toast end to end with no separate notification
  surface for pipeline runs.
- PipelineRun.slice_job_id is set before the 202 returns.
- RunWithPipelineModal calls useSliceJobTracker().trackJob() from
  runMutation.onSuccess.
- RunWithPipelineModal source prop is now {kind, id, filename}
  mirroring SliceModal.SliceSource; api.checkPipelineEligibility +
  api.runPipeline take a discriminated-union source argument.
2026-06-27 15:01:14 +02:00
maziggy d6bdb7e200 feat(slicer): Slicer Pipelines — save & reuse a preset bundle in one click (#1425 PR A)
The SliceModal forces the user to pick four slots every time (printer /
process / filament(s) / bed type). For fleet production that's tedious
and error-prone. Pipelines let an operator save a named bundle and apply
it with one click on the next file.

PR A is bundle-and-management only. PR B adds single-target dispatch,
PR C adds multi-copy batch with capability-matched fanout. Future-PR
columns (target_kind / target_printer_id / target_model_class /
fanout_strategy) ship in this migration so PR B+ is code-only, not a
schema bump.

Backend
- New model SlicerPipeline + slicer_pipelines table; soft-delete via
  is_deleted so PR B+ run history can still resolve metadata.
- Pydantic schemas reuse the existing PresetRef shape from
  schemas/slicer.py.
- CRUD routes at /api/v1/slicer-pipelines/ — list (newest first by id
  DESC), create (201), get-by-id, partial PUT, soft-delete (204).
- Three new permissions: PIPELINES_READ / PIPELINES_WRITE / PIPELINES_RUN.
  Administrators + Operators get all three; Viewers get READ.
  Backfill in seed_default_groups() so existing installs upgrade
  cleanly. All three denied to API keys for now.

Frontend
- Settings → Workflow splits into two horizontal sub-tabs mirroring
  the Authentication tab pattern: "Queue & Dispatch" (existing
  Workflow content) and "Pipelines" (new). URL deep-link via
  ?tab=queue&sub=pipelines.
- SlicerPipelinesPanel — list, inline rename, delete, stale-preset
  warning when a referenced preset no longer resolves.
- SliceModal gets "Apply pipeline ▾" + "Save as pipeline". Apply
  fills all four slot states; the filament list right-pads from
  current state so a pipeline with fewer entries than the current
  source's slot count keeps the existing tail.
2026-06-27 13:56:18 +02:00
maziggy 9033b0f81e feat(toast): restore upload-progress toast for scheduler dispatches (#1625 follow-up)
FTP push to the printer into the server-side scheduler tick. That
removed the browser-side upload the old XHR-progress modal listened
to — users only saw the queue item flip to "active" with no visibility
into the FTP push + the H2D/H2D Pro 80-210 s project_file digestion
window before the printer actually started.

Port the legacy bg-dispatch toast rendering from
0b43ac0d:frontend/src/contexts/ToastContext.tsx lines 510-650 back in
place verbatim — same DOM tree, same Tailwind classes, same
formatFileSize bytes line, same uppercase status chip, same collapse
chevron, same awaitingPrinter derivation, same auto-dismiss. The only
adapt is the event ingestion: a useEffect maps the four scheduler-side
WS events to the legacy DispatchToastJob shape.

The toast materializes when the FTP push to the printer ACTUALLY
STARTS (queue_item_uploading) — NOT on POST /queue. A draft that
fired at queue-add made the toast jump to "Dispatched" before any
upload had happened.

Four backend WS events drive it: uploading (carries printer_name +
total_bytes), upload_progress (throttled at 200 ms / 256 KB to match
legacy background_dispatch.py:614-615 1:1, first call always emits,
completion always emits; an _UploadProgressBridge bridges from the
FTP executor thread to the asyncio loop), acked (printer transitioned
out of pre_state), failed (with a reason key the toast looks up as
dispatchToast.failed.{reason}). No queue_item_dispatched event: the
legacy path kept status=processing from upload start until printer
ack, "Awaiting printer..." derives from upload_progress_pct >= 99.9
(legacy uploadDoneAwaitingPrinter trick).

Per-user routing: WS connect resolves the principal username to
User.id once and stashes it on websocket.state, so
ws_manager.broadcast_to_user filters O(connections). Auth-disabled
installs route user_id=None to all connections — matches the legacy
single-user behaviour. The watchdog receives created_by_id through a
new kwarg so the static method can still emit acked without
re-fetching the queue item.
2026-06-27 11:28:00 +02:00
maziggy d4ad41d850 fix(hms): action buttons actually reach the printer (#1830)
Three distinct bugs combined into one user-facing failure: clicking
Stop / Problem-solved-and-resume / Ignore-and-resume returned 200 OK
but the printer didn't act, modal stayed up, print stayed paused.
Verified by injecting candidate command shapes on device/<sn>/request
against a live H2D paused on a wrong-plate HMS (print_error=0x05008051).

(1) hms_resume / hms_stop dispatched the "err"-bearing shape that
BambuStudio doesn't actually send; Bambu firmware silently rejects it.
Both now send the plain shape ({"print":{"command":"<x>","param":"",
"sequence_id":"0"}}). PAUSE -> FAILED in 1.7s for stop, PAUSE -> RUNNING
in <2s for resume.

(2) IGNORE_RESUME mapped to idle_ignore, which is BambuStudio's
"dismiss a warning" command and only works for non-pause warnings.
hms_ignore now branches on state.state == "PAUSE": paused -> plain
resume; not-paused -> idle_ignore with the full-length err.

(3) 64-bit hms[]-array faults were truncated to a non-matching err.
short_code in _parse_status discarded 32 of the 64 identifier bits, so
the firmware didn't match it to the active fault. HMSError.full_code
now carries the canonical hex identifier (16 chars for hms[] faults,
8 chars for print_error faults). Catalog lookup tries 16-char first,
falls back to 8-char. HmsActionBody.print_error pattern relaxed to
^[0-9A-Fa-f]{8}([0-9A-Fa-f]{8})?$.

(4) execute_hms_action returned publish-success as success, masking
every silent-rejection bug above as 200 OK. Route now snapshots
(state.state, len(state.hms_errors)) before dispatch, awaits
HMS_ACTION_ACK_WAIT_SECONDS (default 2.5s, module-level so tests
override), and returns 502 with "Printer did not acknowledge HMS
action within 2.5s" if state didn't move.
2026-06-27 09:18:57 +02:00
maziggy 510005f043 fix(printers): cam wall — offline tile chip + don't kill shared
streams when one viewer closes

1) Offline tiles now show OFF (not LIVE)
   CameraWall.modeByPrinter assigned 'live' to any visible printer
   without considering status.connected, so a disconnected X1C wasted
   a live-budget slot AND rendered the red LIVE chip on top of the
   WifiOff placeholder. Disconnected printers now map to 'paused' and
   don't decrement liveBudget — the existing WifiOff + Off chip
   rendering takes over.

2) /camera/stop no longer kills other viewers' streams
   The cam-wall tile, EmbeddedCameraViewer, and the /camera/:id popup
   all subscribe to the same fan-out broadcaster for a printer.
   /camera/stop used to unconditionally shutdown_broadcaster() + kill
   every ffmpeg process for the printer, so closing the embedded viewer
   while the cam-wall tile of the same printer was live force-killed
   the source the tile was pulling from — the tile's <img> errored.

   New get_subscriber_count(key) accessor in camera_fanout.py exposes
   the broadcaster's subscriber list length. /camera/stop now reads
   that first; when >= 1 subscriber is still attached, return
   {stopped: 0, skipped: true} and leave the broadcaster + ffmpeg
   processes alone. The leaving viewer's HTTP teardown still runs the
   natural iter_subscriber.finally -> unsubscribe path, so its slot is
   released; the broadcaster keeps serving the other viewers. Single-
   viewer close still hits the immediate force-teardown (count is 0).
2026-06-26 16:01:28 +02:00
Zelda 3ddf8d847e [Feature]: HMS Actions (#1743) 2026-06-26 14:40:25 +02:00
maziggy 1c683f063c fix(queue): ownership gates + TOCTOU lock + /reorder validator (#1625-followup)
Three issues from the post-merge audit of the unified-dispatch PR, all
pre-existed on dev but became more impactful once every print routes
through the queue:

1. Start/Stop ownership gates. /queue/{id}/stop required QUEUE_UPDATE_ALL
   (admin-only) -- operators saw the Stop button in the queue UI but got
   403 on click. /queue/{id}/start required QUEUE_UPDATE_OWN with no
   ownership check -- _OWN holders could start anyone's queue items via
   direct API. Both routes now use require_ownership_permission, mirroring
   /cancel. Stop is strict (rejects unowned items for _OWN); start preserves
   #1670's VP-import flow where _OWN can start NULL-owner items and claim
   ownership at click-time. Frontend QueuePage Start/Stop buttons flip
   from printers:control to canModify('queue', 'update', created_by_id).

2. TOCTOU race on insert_position. Concurrent ASAP inserts to the same
   scope both computed MAX(position) from before the other committed; in
   an empty scope, both inserted at position=1 (duplicate). Wraps the
   read+update in a transaction-scoped Postgres pg_advisory_xact_lock
   keyed on the printer_id. Different printers don't contend. SQLite
   serializes writes implicitly so the path is no-op there. Dialect is
   checked against the live session binding, not the is_sqlite() helper,
   because the test fixture overrides get_db to SQLite while
   settings.database_url still points at Postgres.

3. /reorder duplicate-position validator. POST /queue/reorder set position
   from the payload in a loop with no uniqueness validation -- a buggy
   drag-drop client could leave the queue with ambiguous ordering (the
   scheduler's ORDER BY (printer_id, position) ties break by row order).
   New model_validator on PrintQueueReorder rejects duplicates at the
   schema layer with 422 + "Duplicate positions in reorder request: [N, ...]".
2026-06-26 13:06:40 +02:00
Ed 4c67d8a4e1 feat: Unify print dispatch through the scheduler (#1625) 2026-06-26 12:31:48 +02:00
maziggy 70857af393 feat(auth): SSO autologin + disable local username/password login (#1589)
Adds a global local_login_enabled setting plus a per-provider
  is_autologin flag on OIDCProvider so operators who run their own SSO
  enabled, or if the calling admin has no UserOIDCLink — either would
  lock everyone out. App-layer invariant: at most one provider can carry
  is_autologin; setting it on one clears it on every other.

  /auth/advanced-auth/status surfaces both new fields so the LoginPage
  decides UI in one query. The env-var bypass flips the reported
  local_login_enabled back to true so the SPA matches what the route
  will accept.
2026-06-25 14:54:27 +02:00
maziggy fd61812d01 feat(drying): show active-cycle filament + target temperature on the AMS drying badge
Bambu's per-tick AMS push carries only the dry_time countdown — the
  filament name and target temperature the user chose are never echoed on
  the wire. The AMS card had no source of truth for them and rendered the
  bare "Drying · 11h 35m left". The badge now shows
  "Drying · PETG @ 65°C · 11h 35m left", matching the cycle the user
  actually started.

  BambuMQTTClient caches {ams_id: {filament, temp}} on send_drying_command
  (mode=1), clears on mode=0 and on the dry_time falling edge to 0 — the
  same per-AMS edge detector that drives the smart-plug-after-drying
  callback. PrinterManager.get_drying_targets exposes it, the four
  printer_state_to_dict call sites thread it through, AMS schema gains
  dry_target_temp + dry_filament, and routes/printers.py builds the same
  fields into the manually-constructed AMSUnit response.

  When no cached target exists (drying started in a previous backend
  lifetime, or initiated outside Bambuddy), the badge falls back to the
  first loaded tray's tray_type + RFID-recommended drying_temp — the
  heuristic the popover already uses to seed defaults.

  i18n: printers.drying.targetSummary = "{{filament}} @ {{temp}}°C" in
  all 11 locales. Parity check 5356 leaves per locale.

  Note: a user reported the H2D's own physical display still labels the
  cycle by the loaded tray's filament (e.g. "PLA" instead of the
  Bambuddy-requested "PETG"). The wire payload is correct end-to-end —
  journalctl shows filament: "PETG" sent and result: success ACKed — and
  the badge in Bambuddy's own UI now reflects what we actually sent,
  independent of the firmware's display choice.
2026-06-25 13:27:32 +02:00
maziggy 8d6f701f1d feat(drying): continue drying while printing + gate rotate-spool when tray loaded (issue #1816)
Continue Auto-Drying while a print is running on capable hardware.
  New Settings > Print Queue > "Continue drying while printing" toggle
  (default OFF). Extends _check_auto_drying in print_scheduler.py to
  evaluate running printers when supports_drying_while_printing(model,
  firmware) returns true. Strict allowlist verified per Bambu wiki
  release notes for "Print While Drying" / "printing while filament is
  drying": H2D 01.03.00.00+, H2C/H2S/P2S/H2D Pro 01.02.00.00+, X2D/A2L
  01.01.00.00+, X1C 01.11.02.00+. P1*, A1, A1 Mini, X1 (non-C), X1E
  intentionally excluded. Mid-print drying temperature is capped at
  max(40, preset_temp - 5) to protect spools from heat damage inside the
  hot enclosure during a print, matching Bambu's own "lower drying
  temperature during printing" guidance.

  Rotate-spool toggle in the drying popover is now disabled when any tray
  in the targeted AMS has filament threaded into the feed tube
  (tray.state === 11). The whole AMS rotates as one mechanism, so a
  single loaded slot locks the entire unit. Previously the toggle was
  always clickable and the firmware rejected with dry_sf_reason=[3]
  (ConsumableAtAmsOutlet) after the click. The first cut keyed on the
  printer-level tray_now but missed the H2D's typical post-print state
  where tray_now resets to 255 while filament stays in the tube — the
  per-tray state field reports it correctly. Submission also clamps
  rotateTray off so a stale-true state from a previous AMS can't leak
  through.

  Backend: supports_drying_while_printing in printer_manager.py covers
  display names and internal SSDP/MQTT codes (O1D, O1E/O2D, O1C/O1C2,
  O1S, N6, BL-P001, N7, N9). New print_drying_enabled boolean in
  settings schema. Frontend: toggle on SettingsPage, gate + clamp on
  PrintersPage drying popover using existing amsData cache. i18n: 3 new
  keys x 11 locales, no English fallback. Tests: 7 cases on the gate
  matrix (TestSupportsDryingWhilePrinting), 4 cases on the scheduler
  mid-print path (TestMidPrintDrying), 9 cases on the rotate gate state
  transitions. Full backend pytest -n 30 green (4251/4251), ruff clean,
  frontend npm run build clean, i18n parity 5355 leaves per locale.
2026-06-25 12:47:26 +02:00
Keybored 6c5b40dd57 [Fix] Forecasting: Group spools by color and rework UI (#1814) 2026-06-25 11:35:57 +02:00
maziggy 8a26e7d753 fix(inventory): stop popping the unknown-tag modal for slots with no RFID
The 7cb905a follow-up mounted the global unknown-tag modal listener, which
  turned an existing always-on broadcast for no-tag slots from a silent no-op
  into a perpetual popup loop — every push for a slot with a generic
  non-RFID spool (or zero-filled tag) re-prompted, and confirming each one
  created a fresh ghost spool with an empty tag.

  - main.py on_ams_change: drop the no-tag else-branch broadcast. No identity,
    no prompt; the slot stays unassigned until a real tag is read.
  - inventory.py + spoolman.py /spools/from-slot: 400 when the slot has no
    usable tag_uid / tray_uuid so stale frontends can't recreate the ghost
    spool by re-confirming a queued prompt.
  - test_inventory_from_slot_no_tag: lock the guard in (zero-filled + empty
    string).
2026-06-25 09:57:15 +02:00
maziggy 2fe9896917 fix(queue): close #1818 — Resume after failure clears the gate
Single failure on a printer with require_previous_success queue items
  permanently skipped every downstream + every new item — the
  _check_previous_success lookback always walked back to the original
  failed row (skipped is excluded from the lookback), and no code path
  could dismiss that failure.

  Three pieces:

  1. PrintQueueItem.gate_acknowledged Boolean column (default False).
     SQLite/Postgres-safe ALTER, dialect-branched DEFAULT.

  2. _check_previous_success skips rows where gate_acknowledged=True so
     acknowledged failures walk past the lookback. Fresh post-resume
     failures still gate independently.

  3. POST /api/v1/queue/printer/{printer_id}/resume — gated on
     QUEUE_UPDATE_ALL — acknowledges failed/aborted items for that
     printer AND restores items where
     status='skipped' AND error_message='Previous print failed or was
     aborted' back to pending in one transaction. Returns
     {acknowledged, restored}.

  Frontend banner above the active Queue tab surfaces blocked printers,
  fires a warning-variant ConfirmModal, and shows a precise toast on
  success.
2026-06-25 09:00:57 +02:00
maziggy fb3821630f feat(inventory): batch / mass edit on the Filament tab (#1795)
Bulk operations on the Inventory page in both built-in and Spoolman modes.
  Reporter wanted ten-of-the-same-spool edits without ten round-trips through
  the per-spool editor.

  Frontend
  - New checkbox column on the inventory table (header / row / group). Sticky
    toolbar appears when at least one row is selected with Edit / Print labels /
    Reset usage / Archive (or Restore in the Archived tab) / Delete / Clear.
    Selection clears on any filter / tab / search change so the count can't
    drift from what is on screen.
  - BulkEditSpoolsModal is a three-state-per-field form. The user opts in per
    field by ticking its checkbox or just typing into it; only ticked + non-
    empty fields are sent. Clearing fields in bulk is intentionally NOT
    supported per the issue discussion.
  - A new SearchableSelect renders all categorical fields (material, sub-type,
    brand, category, slicer preset name, slicer filament, storage location)
    with the same dropdown pattern the per-spool editor uses - text input +
    chevron + filtered button list, click-outside / Escape closes. No native
    select anywhere in the modal. Options merge the canonical constants from
    spool-form/constants.ts with whatever already exists in the user's
    inventory. Slicer-preset dropdowns fetch the same sources as the per-spool
    form (Bambu Cloud + Orca Cloud + local + built-in) through buildFilament
    Options() and three useQuery calls gated on isOpen.
  - onSuccess handlers surface three outcomes: all-succeeded (green toast),
    partial-success (yellow toast with ok / failed counts), all-failed (red
    toast that keeps the selection and modal open so the user can retry).
    The first cut silently dropped errors / not_found arrays - audited and
    fixed before merge.
  - Invalid rgba hex is flagged inline with a red border + helper text and
    the Apply button is gated on a hasDroppedTickedField guard, so silently
    dropping a ticked field is no longer possible.
  - bulkResetConsumedCounterMutation.onSuccess now closes the confirm modal +
    clears selection, matching the other three bulk mutations.

  Backend
  - Four new endpoints per inventory mode (eight total):
      POST /api/v1/inventory/spools/bulk-update         INVENTORY_UPDATE
      POST /api/v1/inventory/spools/bulk-delete         INVENTORY_UPDATE
      POST /api/v1/inventory/spools/bulk-archive        INVENTORY_UPDATE
      POST /api/v1/inventory/spools/bulk-restore        INVENTORY_UPDATE
      POST /api/v1/spoolman/inventory/spools/bulk-*     FILAMENTS_UPDATE
  - Built-in update runs the same prepare_internal_spool_payload(...) +
    weight_used / weight_locked auto-stamp as the per-spool PATCH.
  - Spoolman update loops the per-spool update_spool route function so the
    filament re-linking / extra-dict / extra-lock / shared-filament rules
    stay byte-identical to single-spool edits.
  - Per-spool failures inside the batch are collected. Spoolman bulk-delete /
    archive / restore now catch non-HTTPException too (matches bulk-update) -
    a mid-batch httpx.ConnectError or TimeoutError no longer aborts the route
    with a 500 and skips the WS broadcast.
  - Both modes broadcast a single inventory_changed WS event at the end of
    the batch.
2026-06-23 11:34:15 +02:00
maziggy 7cb905ad0c feat(inventory): toggle to disable auto-add of unknown RFID spools + global confirmation modal (issue #1764)
New setting "Auto-add unknown RFID spools" under Settings -> Filament -> Filament Tracking,
  default ON for back-compat. When turned off, the backend stops auto-creating an inventory
  record for an unknown RFID tag and instead broadcasts an unknown_tag WS event that pops
  a global confirmation modal in the Bambuddy UI showing the printer / AMS-X label / slot /
  material / colour. Add or Cancel; no nag on every MQTT push.

  Backend
  - Module-level _unknown_tag_last_broadcast dict dedupes per (printer, slot, tag). Set is
    committed AFTER ws_manager.broadcast() returns so a crashed broadcast doesn't poison
    the dedup and permanently silence the slot.
  - Empty-slot MQTT push clears that slot's entry, so remove+reinsert reliably re-prompts.
  - Successful matches via get_spool_by_tag / find_matching_untagged_spool / create_spool
    also clear the entry so a future tag swap re-prompts.
  - Tray data (tray_type, tray_color, tray_sub_brands, tray_count) shipped in the WS payload
    directly so the modal renders the real material / colour instead of relying on the
    React Query cache that lags the WS event by several seconds.
  - Two new endpoints back the modal's confirm action:
      POST /api/v1/inventory/spools/from-slot     (INVENTORY_UPDATE)
      POST /api/v1/spoolman/spools/from-slot      (FILAMENTS_UPDATE)
    Both look up the slot's tray data server-side and create + auto-assign atomically.
  - Spoolman /from-slot now raises HTTP 500 when the slot-assignment INSERT fails instead
    of returning success while the DB rolled back the binding.
  - sync_ams_tray gained an optional auto_add_unknown_rfid kwarg (default True so existing
    callers are unaffected); auto-sync and both manual sync routes thread the setting.

  Frontend
  - useUnknownTagPrompt hook listens for the unknown-tag CustomEvent, reads the tray fields
    out of the event detail, and feeds a single-modal queue. No long-lived dismissed set;
    the backend dedup handles spam suppression.
  - UnknownSpoolModal wraps the existing ConfirmModal with a material + colour-swatch
    preview block.
  - Mounted in Layout.tsx alongside useSponsorPrompt so SpoolBuddy kiosk / login / setup
    routes are excluded.
  - getAmsLabel moved to utils/amsHelpers.ts; ConfigureAmsSlotModal.tsx and PrintersPage.tsx
    both import the shared version (canonical AMS-A / HT-A / External labels).
  - AppSettings TS interface gained spoolman_enabled, auto_add_unknown_rfid, spoolman_url
    so the runtime cast in the hook is no longer needed.
  - SpoolmanSettings.tsx gets a new toggle row in the Filament Tracking card, visible in
    both built-in and Spoolman branches; auto-save + toast already wired.
2026-06-23 09:59:05 +02:00
maziggy 4dcd37bc87 feat(system): NTP-gate state on /api/v1/system/appliance
Extends the appliance endpoint that landed in the previous commit with a
  time_synced field, sourced from /run/bambuddy/time-synced (the appliance's
  ntp-gate.sh writes this once chronyd reports sync, or with a "warning"
  marker after the 3-minute timeout). The RPi 5 has no battery-backed RTC,
  so on a fresh boot the system clock is wrong until NTP catches up -- JWT
  expiries and TLS certificate validity windows depend on this being right.
  Exposing the gate lets the SPA render a "time not synced" indicator while
  that's still true and clear it once "ok" comes through.

  backend/app/core/local_config.py

  New read_ntp_gate(path) function alongside read_local_toml. Three states:

    "ok"       chrony reported sync within the 3-minute window
    "warning"  3-minute timeout elapsed without sync; user already waited
               and the wizard proceeded with a degraded clock
    None       file absent (non-appliance install), OSError, empty content,
               unknown marker, or binary garbage -- "unknown / don't gate"

  Defensive read mode (errors="replace") survives non-utf8 content without
  crashing. Module docstring broadened from "local.toml reader" to "small
  readers for appliance-set state files".

  backend/app/api/routes/system.py

  /system/appliance now returns:

    {hostname, timezone, locale, time_synced}

  with the same no-auth posture: bootstrap surfaces (i18n init, time-sync
  banner) read this before auth might be set up, and the contents are
  non-secret (user-set defaults + a public sync flag). The endpoint
  docstring expands to explain the RTC motivation -- otherwise the
  time_synced field reads like a leftover.
2026-06-22 14:23:30 +02:00