Commit Graph
1543 Commits
Author SHA1 Message Date
maziggy 3f21e0b8ff follow-up(queue): cache per-plate 3MF metadata so queue polling stops re-parsing every row (#2573) 2026-07-17 08:26:23 +02:00
maziggy f8e49fed91 fix(queue): cache per-plate 3MF metadata so queue polling stops re-parsing every row (#2573)
The Queue listing serialized each item by opening its 3MF and re-parsing
slice_info.config three times (print time, filament usage, bed type) on
every poll, per connected client, even for unchanged files. Add a single
combined extract_plate_metadata_from_3mf() cached by (path, plate_id,
mtime_ns, size); the three legacy helpers delegate to it. An unchanged
queue now does no repeat 3MF parsing.
2026-07-17 08:26:17 +02:00
maziggy 00251fe808 feat(orca-cloud): pair via RFC 8628 device flow, replacing the paste-based sign-in
OrcaSlicer shipped a first-class external-app pairing API (OAuth 2.0 Device
Authorization Grant), so the Supabase-PKCE copy-paste flow is replaced end to
end. Connecting is now: click Connect, approve a short code on the Orca Cloud
settings page, done — no redirect, no callback paste, no client secret, works
from a LAN IP / localhost / behind a proxy.

Backend: services/orca_cloud.py rewritten to device-code request + poll (the
four RFC outcomes) + refresh_token grant + introspection + external sync pull;
routes expose /device/start and /device/poll (device_code kept server-side in
the reused orca_cloud_pending_* columns, no migration). Requests sync:read
(read-only feature). Prod endpoint by default, ORCA_CLOUD_API_BASE overrides
to staging. Wired the shared httpx client (fixes a per-request socket leak).

Frontend: device-code connect UI + api client methods; all 11 locales updated.
2026-07-17 08:16:35 +02:00
maziggy da128e50c7 fix(inventory): broadcast assignment change when auto-unlink clears a stale slot (#2575)
The #2575 reconciliation correctly deletes a stale external-spool
assignment in on_ams_change, but did so silently: spool_assignment_changed
was only broadcast by the manual REST assign/unassign endpoints, and the
frontend's spool-assignments cache is invalidated only by that event. So
after an external-spool type swap the DB was correct but every open browser
kept rendering the unlinked spool on the slot until an unrelated refetch —
which the reporter read as "the fix didn't work" (a browser refresh showed
the right state all along).

Broadcast spool_assignment_changed for each auto-unlinked slot after the
commit. No frontend change — the handler already invalidates the cache.
2026-07-17 07:26:24 +02:00
maziggy 75b0175e3d fix(camera): bound the post-kill wait on ffmpeg cleanup (#2580)
After an RTSP read timeout the stream cleanup killed the stalled ffmpeg
and then awaited process.wait() unbounded. A SIGKILLed ffmpeg stuck in
uninterruptible I/O on a dead RTSP socket can take arbitrarily long to
be reaped, so the fan-out stream coroutine sat parked in that wait (12
hours in the reported case) while every new viewer attached to the
stalled broadcaster and received no frames.

Bound the post-kill wait to 2s in all three places it existed: the
stream generator's _terminate_ffmpeg (the reported hang), the camera
stop endpoint (which would hang the recovery request itself; now uses
the shared helper instead of an inline copy), and the orphan-cleanup
janitor (whose hang would disable the safety net). On timeout the
zombie is abandoned; the janitor's /proc scan reaps it next pass and
the stream proceeds to its normal reconnect.
2026-07-17 07:03:10 +02:00
maziggy e97413edc7 fix(queue): enforce sliced-model compatibility on cross-model dispatch (#2578)
A queue item's "Any <model>" button labeled itself from the file's slice
metadata while the scheduler used the row's target_model, so an X1C-sliced
item targeting H2D showed "Any X1C" above "assign to first idle H2D". The
mismatch itself was created silently: sliced-for metadata loads async, and
switching to model mode before it arrived pre-selected the alphabetically
first model (H2D on a mixed farm), after which the model dropdown hid
itself. Nothing validated compatibility, so the scheduler would hand X1C
G-code to an H2D.

Frontend: never default the target silently, keep the dropdown visible in
model mode (incompatible models disabled), label from the actual target,
warn on mismatch, block submit when incompatible.

Backend: new GCODE_COMPAT_FAMILIES table (X1/X1C/X1E/P1P/P1S interchange;
everything else exact-match; missing metadata never blocks). Queue create
and update reject incompatible targets with 400; the scheduler holds back
pre-existing mismatched rows with an actionable waiting_reason instead of
dispatching them.
2026-07-17 06:51:51 +02:00
maziggy a6e7d671f2 fix(jog): stop disabling firmware endstops; warn that limits aren't enforced (#2579)
Manual jog could drive an axis past its travel limit into a collision.
Instrumenting the exact G-code to an H2D showed Bambuddy sending a clean
move at the limit (G91 / G1 Z-1.00 F600 / G90, no M211) that the printer
ran straight past, while its own touchscreen refuses the identical move.
This is a Bambu firmware bug: soft endstops are not enforced on G-code
received over MQTT, and no axis position is reported, so the move cannot
be clamped firmware- or client-side from position.

Two changes: (1) jogs no longer wrap moves in M211 S0/S1 — that disabled
the firmware's soft endstops globally, breaking even the touchscreen's
limits until a power cycle; a bare move keeps the touchscreen protected.
(2) The jog panel shows a prominent warning that travel limits are not
enforced during manual moves due to the firmware bug. Client-side
dead-reckoning enforcement is tracked separately.
2026-07-16 15:21:47 +02:00
maziggy e336eb53e7 fix(inventory): reconcile external-spool assignment when its filament type changes (#2575)
Assigning a new filament to the external spool (e.g. generic ABS over
generic TPU) left the previous inventory spool assigned. The reconciliation
that unlinks a stale external-spool assignment lives in on_ams_change, but
that callback only fired on regular AMS-unit changes: its change-hash never
included the external spool (vt_tray/vir_slot), and the external-spool data
is stored after the AMS handler runs.

Detect external-spool identity changes (type, colour, tag, or reset to
empty) and re-fire on_ams_change so the stale assignment is unlinked. The
fill percentage (remain) is excluded from the fingerprint so a running
print doesn't trigger it on every push.
2026-07-16 09:47:21 +02:00
maziggy 34818927c2 Revert "fix(db): configurable connection pool + auth_enabled cache for large farms (#2572)"
This reverts commit 4ad43c96de.
2026-07-16 09:30:36 +02:00
maziggy a5176349dd fix(notifications): capture the snapshot outside the DB session (#2572)
The progress-milestone and HMS-error notification paths in
on_printer_status_change held a session across the ~15s camera snapshot
taken for the notification image, pinning a pooled connection per
milestone/error per printer.

The snapshot needs no DB: read the printer in a short session, release
it, grab the snapshot with none held, then open a fresh session for the
notification send (and lift the db-free MQTT publish out too). Pinned by
a test that fails if the snapshot runs while a session is open.
2026-07-16 09:20:12 +02:00
maziggy c1d25af7f1 fix(finish-photo): release DB connection during the camera capture (#2572)
The background finish-photo task held one session open across the whole
capture pipeline (timelapse extraction, up to 20s stage-22 wait,
external-camera/RTSP grab) — tens of seconds of a pooled connection
idle-in-transaction per finishing print.

Read the setting/printer/archive in a short session, release it, run the
capture with no session held, then re-open a fresh short session only to
append the photo. Logic unchanged.
2026-07-16 09:07:07 +02:00
maziggy afcea50eb7 fix(timelapse): don't hold a DB connection across the FTP scan (#2572)
_scan_for_timelapse_with_retries opened one session and held it across
the FTP directory listing and the multi-MB video download — once per
retry attempt, per completed print — pinning a pooled connection
idle-in-transaction for the whole transfer.

Read the archive + printer in a short session, release it, do the FTP
list/download with no session held, then re-open a fresh short session
only to attach the file. Existing scan tests already cover the
read/download/attach path.
2026-07-16 08:58:54 +02:00
maziggy b3c0429373 fix(camera): release DB connection before streaming, not after (#2572)
/camera/stream took its printer row via Depends(get_db). get_db is a
yield dependency, so its session stayed open until the response body
finished streaming — for a live MJPEG stream, as long as the browser
tab is open (hours). Every open camera tile pinned one pooled DB
connection idle-in-transaction, draining the pool on large farms.

Fetch the printer in a short-lived async_session() and release the
connection before returning the StreamingResponse. expire_on_commit=
False keeps the already-loaded columns readable during the stream.
2026-07-16 08:38:42 +02:00
maziggy 4ad43c96de fix(db): configurable connection pool + auth_enabled cache for large farms (#2572)
Large PostgreSQL farms exhausted the fixed pool (pool_size=10 +
max_overflow=20): with ~93 printers every connection sat idle in
transaction and unrelated requests waited out the 30s pool timeout or
failed in the auth middleware.

- Make pool sizing env-configurable (DB_POOL_SIZE / DB_MAX_OVERFLOW /
  DB_POOL_TIMEOUT / DB_POOL_RECYCLE); raise the Postgres default to
  20 + 80 with pool_pre_ping + pool_recycle=1800.
- Cache the auth_enabled probe (30s) to drop a per-request DB round-trip.
  Only enabled=True is cached, so staleness fails closed; set_auth_enabled
  invalidates immediately.
- Add GET /api/v1/system/db-pool exposing resolved config + live
  checked_out/checked_in/overflow gauges without consuming a connection.

Session-hygiene (connections held across MQTT/FTP/camera/3MF I/O) is a
separate follow-up.
2026-07-16 08:16:48 +02:00
maziggy c2b23e5e61 Security hardening (maziggy/bambuddy-security #5) 2026-07-16 07:50:51 +02:00
maziggy a31a3a522c Bumped version 2026-07-15 13:15:46 +02:00
maziggy 001f2f120c fix(inventory): record the mapped spool's material, not the sliced type (#2563)
A slot mapped to a different filament than it was sliced for (PLA slice
routed to the only loaded PETG slot) was logged under the sliced material
in the archive, Print Log and material stats, even though the correct spool
was debited. Once usage tracking resolves every used slot to a spool, adopt
the spool's material as the archive filament_type, exactly as the spool
colour is already adopted (#1494). All-or-nothing; both inventory backends;
flows through to the Print Log and stats. No schema/UI/i18n change.
2026-07-15 07:29:16 +02:00
maziggy c0a50edbe8 fix(queue): re-check the queue quickly after a dispatch instead of always waiting 30s (issue #2555)
The scheduler slept a fixed 30s after every pass, so each printer that
freed up during a batch waited up to a full interval before its next job
was dispatched — on a farm, that idle gap stacked into the "several long
minutes" reporters saw between requesting prints and them starting (#2555).

check_queue() now reports whether it dispatched anything; run() loops again
after 3s on a productive pass and falls back to 30s otherwise. Fast ticks
only continue while the queue is actively draining, so this can't tight-loop:
a pass that dispatches nothing (all pending items behind busy printers, or a
wedged head-of-line job holding its printer) reverts to the normal interval.
2026-07-15 07:13:22 +02:00
maziggy 62a64006b8 fix(camera): stop /camera/stop from letting a second socket open (#2521)
The single-connection barrier from the last round was correct and was being
bypassed. shutdown_broadcaster() popped the broadcaster out of the registry
and only then awaited its teardown, so while the socket was still closing the
slot sat empty: a /camera/stream request landing in that window minted a
broadcaster with no predecessor and dialled port 6000 immediately. A page
reload fires /camera/stop and the new stream request concurrently, so a P1S
ended up holding two connections, kept feeding the orphan, and starved the
live viewer until its TCP keepalive reaped the dead one ~20 min later. The
stopped broadcaster now stays in the registry so the successor chains behind
its socket close.

The camera page also rendered the <img> src before the stream token arrived
whenever auth was disabled, then swapped it once the token landed — aborting
the in-flight request and issuing a second one. With auth off both reached the
backend, so every load attached two viewers to a one-socket printer. The src
now waits for the token query to settle.

Subscribers only checked for client disconnect after yielding a frame or on a
30s idle timeout, so a viewer that left during a black stream stayed counted —
and /camera/stop trusts that count to decide whether to tear the upstream down.
2026-07-14 11:52:56 +02:00
maziggy 09b739b95d fix(cloud): stop reporting an expired Bambu Cloud sign-in as connected (issue #2562)
An expired token was indistinguishable from a working one. set_token()
stamped token_expiry = now + 30 days every time a stored token was loaded,
so the expiry reset on every request and is_authenticated could never
return False. /cloud/status answered "connected" for as long as any token
existed, while every cloud call 401'd — and the user was shown Bambu's own
{"error": "Please login."} verbatim.

Bambu is now the authority: /cloud/status validates the token upstream
(cached 5m), and any 401 from any authenticated call durably records the
credential as dead via users.cloud_token_invalid_at, so MakerWorld, cloud
profiles, slicer presets and firmware checks all agree at once. An
unreachable Bambu is treated as unknown, never as expired, so an outage
cannot sign a working session out.

The user-facing message now names the Profiles page, where the Bambu Cloud
sign-in actually lives; the old text pointed at a Settings page that does
not exist. Same stale path corrected in the wiki.
2026-07-14 11:29:56 +02:00
maziggy ce807fb1cc fix(queue): upload to printers in parallel, cap wedge retries, make debug logs survive a farm
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.

Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.

FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.

Two problems the same bundle exposed:

A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.

The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.

Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).

Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
2026-07-14 10:35:58 +02:00
maziggy a0d4b3d837 fix(queue): scope force-colour overrides to the plate the item prints (#2551)
Queueing several plates of one 3MF built a single filament-override list from
every selected plate and posted that same list with each plate's item. A
force_color_match entry blocks dispatch until the printer has that exact colour
loaded, so a single-colour plate waited on the whole batch's palette. The same
shared list also widened required_filament_types, making a PLA plate refuse
every printer that lacked a sibling plate's PETG.

Narrow the overrides to the slots the plate actually consumes, on create and on
update -- in the backend, where the 3MF is, so it holds for every writer of the
queue. Dispatch already re-parsed requirements per plate and keyed overrides by
slot, so the dropped entries were inert there. When the plate's slots cannot be
read the overrides are kept whole: an item waiting on a colour it does not need
is visible, one that silently lost a forced colour prints in the wrong filament.

Items queued before this would stay stuck with a waiting reason that explains
nothing, so a startup migration re-scopes the pending ones. Printing and
finished items keep their overrides -- that is a record of what they dispatched
with, not an instruction.
2026-07-13 09:18:17 +02:00
maziggy c640ddc1f7 fix(projects): carry tags, due date and priority in the list payload (#2536)
The edit dialog is shared between the projects list and the project detail
page and seeds itself from whichever project object it is handed. The list
payload never carried tags, due_date or priority, so editing from the list
showed a blank tags field -- and, unreported, submitted the dialog's default
priority over a stored high/urgent one. The component read those fields
through a cast, so the compiler never flagged that they were always absent.

Put them on ProjectListResponse and ProjectListItem, drop the casts, and let
an explicit null clear tags and due date the way it already clears budget and
url -- an emptied field was previously sent as undefined and silently reverted.
The template list was missing target_parts_count, which the same dialog edits.
2026-07-13 08:59:36 +02:00
maziggy 5bbfeefa65 fix(backup): diagnose an unwritable backup path instead of quoting errno 30 (#2544)
Nightly backups to a mounted NAS share ran from May and then stopped, failing
with [Errno 30] Read-only file system. The reporter checked folder permissions
-- correctly: the mount is gid=backup,dir_mode=0775, the service user is in that
group, and his own shell writes to the share fine.

Errno 30 is EROFS. A permission problem is errno 13. EROFS means the filesystem
refused the write, and it refused because we told it to: our systemd unit ships
ProtectSystem=strict, which mounts everything read-only inside the service's
mount namespace and carves back out only ReadWritePaths=<install> <data> <logs>.
A NAS share is not one of those three. Reads are unaffected -- which is why the
UI happily listed his existing backups from the share while being unable to
write a new one -- and his shell is outside the namespace entirely, so every
check he could think to run said the directory was fine.

Both installers write the unit file wholesale, so a ReadWritePaths line added by
hand disappeared on the next install, taking the backups with it. They now back
the old unit up (.bak-<timestamp>) and carry the operator's extra writable paths
forward, reporting which ones they kept. The unit template documents the
carve-out.

The output directory is probed with a real write when it is saved and when the
backup card loads, so an unwritable path is caught there rather than at 03:00
for a week. On failure the card names the cause and hands over the fix with the
operator's path already in it (systemctl edit bambuddy -> ReadWritePaths=...),
and a failed run reports the same diagnosis rather than the raw OSError. EROFS
outside systemd, permission-denied, out-of-space, not-a-directory and missing are
told apart, in all 11 locales.

Docker: a backup path that is not bind-mounted is writable -- the write lands in
the container's ephemeral layer and is lost on the next compose up. The probe
compares the directory's device against the container root and warns, with the
compose snippet that mounts it properly.
2026-07-12 08:44:53 +02:00
maziggy ba1394db3e fix(shutdown): exec uvicorn as PID 1 in Docker, and bound the graceful-shutdown wait
Two defects, both invisible until you ask the app to stop.

Docker never shut down gracefully at all. CMD ["sh","-c","uvicorn ..."] left
the shell as PID 1 with uvicorn as its child, and dash does not forward
signals, so docker stop SIGTERMed the shell and uvicorn never heard about it.
Measured on the shipped image: the full 10s grace period, exit 137, and no
"Shutting down" line in the log. Every stop, restart and image update was a
hard kill -- no WAL checkpoint, no MQTT disconnect, no virtual-printer
teardown. `exec` makes uvicorn PID 1; the rebuilt image now stops in 1s with
exit 0 and checkpoints the WAL.

Separately, uvicorn's timeout_graceful_shutdown defaults to None -- wait
forever for in-flight requests. An MJPEG camera stream is a response that
never completes (httptools' connection shutdown() only flips keep_alive on an
in-flight cycle, it never closes the transport), so one open camera tile
pinned the process until systemd SIGKILLed at 90s. The ordering makes it
unfixable from inside the app: uvicorn fires the lifespan shutdown -- the code
that tears the streams down -- only after connections drain.

All six launchers now pass --timeout-graceful-shutdown 5: Dockerfile,
deploy/bambuddy.service, the systemd unit and launchd plist from
install/install.sh, the SpoolBuddy installer's unit, and the Windows NSSM
registration. On timeout uvicorn cancels the request tasks; the camera
generators already unwind cleanly on CancelledError.

TimeoutStopSec raised to 30s on the units and stop_grace_period: 30s added to
compose, as backstops rather than the mechanism. On Windows NSSM's default
1500ms AppStopMethodConsole was force-killing uvicorn mid-teardown; raised to
15s, with the WM_CLOSE and thread-message stages skipped (uvicorn is a console
app with neither a window nor a message loop).
2026-07-11 14:44:45 +02:00
maziggy aba00598bb fix(smart-plugs): read a REST plug's lifetime counter, and derive Today/Yesterday from it (issue #2539)
A Shelly reports one energy figure — aenergy.total, a lifetime counter in Wh
that never resets. Bambuddy had a single REST energy field and filed whatever
it found under "today", so the value never reset at midnight, and Yesterday
and Total stayed at zero: get_energy() simply never set those keys.

With `total` unpopulated, the hourly snapshot recorder skipped the plug, so
the Statistics page's energy figure was zero as well, not just the Settings
card.

Split the REST energy config in two: rest_energy_path still means "used
today", rest_energy_total_path means "lifetime counter". A Shelly has only
the latter; a Tasmota behind a REST bridge has both; sharing a URL costs one
fetch, not two.

Then derive Today and Yesterday from that counter using the snapshots we were
already taking: today = counter now - counter at the last local midnight;
yesterday = the gap between the two previous midnights. Local midnight, not
UTC — a UTC boundary rolls Today over at 02:00 in Berlin. The snapshot loop
now ticks on the local hour so a reading lands on the boundary instead of up
to an hour early. A counter that goes backwards (factory reset) reports
nothing rather than a negative.

Collateral, found while verifying on both engines: the smart-plug DateTime
columns are naive UTC but the code wrote aware datetimes into them. SQLite
drops the offset; asyncpg raises DataError. So on Postgres every snapshot
capture raised inside the loop's except, and every status poll raised on
last_checked — the whole subsystem was dead on the database we recommend for
multi-printer installs. All plug timestamps are naive UTC now.

Existing REST users with a cumulative path in the today field must move it to
the new lifetime field; the form and wiki now name which counter each wants.
2026-07-11 14:22:31 +02:00
maziggy d09db436c3 feat(camwall): serve the Cam Wall at /camwall, and on a token-authenticated kiosk
Cam Wall had no URL — the only way in was the toggle on the Printers page,
so it could not be bookmarked, linked, or shown on a wall-mounted screen.

Add a standalone /camwall route. Signed in, it is the wall as it was. For a
TV or Pi with no login, it authenticates with a long-lived token in the URL.

A kiosk needs the printer list and per-printer status, both of which sit
behind PRINTERS_READ. Rather than widen camera_stream to cover GET /printers
— whose response carries serial_number and ip_address, which have no business
on a screen in a shared room — add a read-only feed at
GET /api/v1/camwall/printers that serves only what a tile draws, and gate it
on a new camwall token scope. The print filename is not served at all: a token
wall renders the compact overlay, so the part on the bed is never named.

The scope is separate rather than a widening: camera_stream tokens are already
in the wild, minted to hand out video, and must not gain the ability to
enumerate a fleet by name. camera_stream is refused by the feed; camwall
passes the stream gate so its own tiles fill.

Kiosk walls drop the settings popover and click-through entirely (not merely
hidden — a passive screen must carry no focusable control it cannot act on),
cap the overlay at compact, and poll rather than open a WebSocket. maxLive,
interval and status can be set from the URL, clamped to the popover's ranges.
2026-07-11 13:38:15 +02:00
maziggy f2e113ce20 feat(vp): mirror live print progress to the slicer (#1887)
A server-mode VP with a target printer bound showed the print as a bare
filename in Bambu Studio / OrcaSlicer -- no stage, percentage, layer count or
time remaining. The data was already in the bridge cache; we were overwriting
it with zeros, because passing it through made the slicer read the VP as busy
and hide the Send button (#1558).

Both slicers gate the progress panel and the Send button on one predicate,
MachineObject::is_in_printing() -- gcode_state in RUNNING/PAUSE/SLICING/PREPARE
-- so there is no field-level way to have both. FINISH is the one state in the
gap: StatusPanel::update_subtask() renders the panel for it, and
SelectMachineDialog::update_show_status() does not disable Send. The VP already
parks at FINISH after each upload (#1280 / #1658), so it only needed the real
numbers underneath it.

While the target prints and no upload is in flight, the report now holds
gcode_state=FINISH and passes mc_print_stage, mc_percent, mc_remaining_time,
stg, stg_cur, layer_num and total_layer_num through from the cache. Mirroring
is suppressed during PREPARE and for 5s after the last upload transition, so
the slicer still receives the FINISH carrying its own subtask_name and releases
its send modal. print_error is never mirrored -- it would raise a modal error
dialog for a fault the VP did not throw.
2026-07-11 10:06:03 +02:00
maziggy ca3f6e5ee0 fix(drying): P1 AMS drying is screen-only — stop offering it (#2533)
The reporter found what his P1S was doing, and it is in Bambu's P1 manual:
"P1S connected AMS drying functions may only be controlled from the P1S screen."
The firmware acks ams_filament_drying with result: success and then discards it,
which is why three commands on an idle printer left the AMS 2 Pro at dry_status 0.
No command can start a cycle on a P1, on any firmware, so don't offer one.

supports_drying() now excludes the P1 series outright, replacing the 01.08+ gate
carried since #292 — that version is when P1 firmware gained AMS 2 Pro support,
not remote drying, and it was never checked against a live P1. Both drying routes
refuse with a specific 400 instead of publishing a message the printer will drop;
queue and ambient auto-drying skip P1s via the same helper.

A new drying_screen_only flag keeps the control on the card, disabled, saying why
— a P1 owner needs to learn where to dry, not watch the button disappear. A cycle
started at the printer still shows with its countdown; only Stop goes away, since
a P1 ignores stop exactly as it ignores start.

Also corrects the wiki firmware matrix, which listed P1P/P1S as supported and
(separately) P2S/H2S/H2C as unsupported. 8 tests.
2026-07-11 09:33:52 +02:00
maziggy ce31de65c2 fix(skip-objects): scope the object list to the plate being printed (#2522)
extract_printable_objects_from_3mf() has accepted a plate_number since it was
written and no caller ever passed one, so it took root.find(".//plate") — the
first plate in the file. Passing one would not have helped either: the lookup
was .//plate[@plate_idx='N'], a predicate on an attribute neither Bambu Studio
nor OrcaSlicer writes. The index lives in a <metadata key="index"> child, as
threemf_tools and filament_requirements already read it, so the selector never
matched and fell back to plate 1 regardless.

On an all-plates .gcode.3mf that meant Skip Objects offered the wrong plate's
objects, with that plate's marker positions drawn over the correct plate's
thumbnail (/cover resolves the plate properly via resolve_plate_id, the object
list did not). The reporter printed a one-object plate and was shown the four
copies from another plate of the same file.

Select the plate on its index metadata, and pass resolve_plate_id(state) at all
three call sites so the list and the thumbnail share one resolver. Also stop
peek_plate_index_in_3mf() reporting plate 1 for a multi-plate file: it backs the
running one, so an all-plates upload printing plate 2+ lost its archive entirely.
2026-07-11 09:18:35 +02:00
maziggy 3fd3ec06b9 fix(ftp): stop a slow upload from being retried on top of itself (#2529)
upload_file_async carried a flat 600s wall-clock deadline and ran the
transfer via asyncio.wait_for(run_in_executor(...)). wait_for cancels the
future, not the executor thread. A 96 MB 3MF to an A1 over WiFi sustains
~75 KB/s and needs ~20 minutes, so the await gave up at ~70 MB, returned
False, and with_ftp_retry started a second STOR of the same file onto the
same printer while the first was still streaming. The reporter filmed two
transfers of one job climbing in parallel at 2% and 72%; the print never
landed and the printer read as having a flaky network.

The deadline is now derived from the file size against a 25 KB/s floor, so a
slow-but-healthy transfer can finish — a link that has actually died is
caught within socket_timeout by the blocking sendall, which is what should be
detecting failure. A deadline expiry now stops the transfer for real: the
worker is signalled, raises UploadCancelled from its progress callback, and
upload_file's existing cancel path breaks the send loop and deletes the
partial file. with_ftp_retry never retries that, and a per-printer lock makes
overlapping uploads impossible however they were triggered.
2026-07-11 09:05:49 +02:00
maziggy 50c3e94d33 fix(cloud): log expected preset misses at DEBUG, keep real faults at WARNING
Failed to get cloud preset ... 400 {"message":"missing"} is the expected
answer, not a fault: many official presets are only addressable with a
printer-variant suffix (GFSL05 exists solely as GFSL05_07 @BBL A1), and
personal P-prefixed presets belong to the account that sliced the file.
Phase 3 already resolves both from local presets, so the lookup miss is
routine -- and one WARNING per AMS tray per tooltip refresh teaches
operators to ignore the log.

BambuCloudError now carries the upstream status_code. The preset lookup
logs HTTP 400 at DEBUG; expired tokens, 5xx and transport failures stay
at WARNING.

Not fixed here: resolving the variant suffix. It selects a printer profile
and the response carries that profile's pressure_advance, so guessing a
suffix would report another printer's K value.
2026-07-10 08:24:20 +02:00
maziggy e6136b660b fix(cloud): carry Bambu Cloud credentials across the auth on/off boundary
get_stored_token() reads the global Settings rows when auth is disabled and
User.cloud_token when it is enabled, so completing /auth/setup switched which
store the /cloud/* routes consult without moving the token. An account linked
before enabling auth was stranded: build_authenticated_cloud() returned None,
get_filament_info() skipped its cloud phase and answered 200 from local
fallbacks, and /cloud/devices began returning 401 -- all silently.

setup_auth() now migrates the global token onto the owning admin and deletes
the global rows; disable_auth() mirrors the hand-off back. Neither guesses:
setup migrates only when it creates the admin or exactly one exists, disable
declines to overwrite an existing global token. Region survives both hops.

Instances that already crossed the transition must re-link once.
2026-07-10 08:15:35 +02:00
maziggy 77f8a3a3ff fix(tls): declare TLS 1.2 as the minimum for printer FTPS and MQTT
ssl.create_default_context() leaves minimum_version at MINIMUM_SUPPORTED,
so the floor came from the OpenSSL build rather than from Bambuddy. On
identical OpenSSL 3.5.6, python:3.13-slim-trixie reports TLSv1_2 while a
bare-metal venv reports MINIMUM_SUPPORTED -- Docker installs were floored
at 1.2, bare-metal and appliance installs were not.

Set minimum_version explicitly in ImplicitFTP_TLS and the MQTT client. On
the P2S/X2D profiles that also cap maximum_version this becomes an exact
TLS 1.2 pin. Probed against an X1C and an H2D on :990 and :8883: both
complete only on TLS 1.2 and reject 1.0, 1.1 and 1.3; live FTPS login
through the new path succeeds on both.

Also correct a stale comment in ftp_profiles.py -- X1C and H2D refuse
TLS 1.3, so cap_tls_v1_2 is a no-op there, contrary to what it claimed.
2026-07-10 07:55:14 +02:00
maziggy e1e2c12d25 fix(library-tags): declare response_model=None on the 204 DELETE route
Under `from __future__ import annotations` the `-> None` return annotation
reaches FastAPI as the string "None", which resolves to NoneType -- truthy,
so APIRoute asserts a 204 may carry no response body and the app fails to
import. fastapi >= 0.116 guards against this; the 0.109-0.115 releases
requirements.txt still allows do not.
2026-07-09 16:29:13 +02:00
maziggy f6c6cfbad3 fix(ams): show "?" not "Empty" for non-RFID spools using tray_exist_bits (#2527)
A spool with no readable RFID was reported by the standard AMS with an empty
tray_type and state=9 — structurally identical to a truly-empty slot at the
tray level — so the AMS card rendered it "Empty" while Bambu Studio correctly
showed "?". The authoritative "a spool is physically here" signal is firmware's
AMS-level tray_exist_bits bitmask (what Studio uses), but Bambuddy inferred
emptiness from the per-tray state/tray_type. Confirmed from the reporter's
bundle: tray_exist_bits=f (all four slots present) with tray_is_bbl_bits=5
(only slots 0,2 Bambu) — the present-but-non-Bambu slots were the ones shown
Empty. Supersedes closed #1838.

apply_tray_exist_bits() already parses the bitmask to clear stale fields on
absent slots; it now also annotates each slot with an authoritative `exists`
bool, gated behind a new annotate_exists flag so only the printer-card path
sets it. The VP bridge leaves it off, so the `exists` key never reaches the
slicer wire format. `exists` flows through the AMSTray schema/serialization to
the frontend, where getEmptySlotKind() uses it: exists===true + no tray_type
-> "?" (present, unconfigured), exists===false -> "Empty", exists absent ->
the previous state=9/10 heuristic (AMS-HT and missing-bitmask paths unchanged).
H2D/X1C already reported present-unknown slots with a non-9 state and took the
"?" path; with the fix they reach it via `exists` and are unaffected.
2026-07-09 09:27:13 +02:00
maziggy 9e7f6cafd9 fix(backup): preserve NOT NULL/DEFAULT/FK/UNIQUE in Postgres→SQLite backup (#2526)
On a PostgreSQL install, create_backup_zip() exports a portable SQLite copy
so backups move between engines. It rebuilt each table with only column name
+ type + PK, dropping NOT NULL, server_default/DEFAULT, foreign keys, and
unique constraints. Restore onto SQLite page-copies that schema straight onto
the live database, and post-restore init_db() can't repair it (create_all is
CREATE TABLE IF NOT EXISTS). So server_default columns like
spoolbuddy_devices.created_at (server_default=func.now()) ended up with no
DEFAULT: SQLAlchemy omits them on INSERT, the DB wrote NULL, and the next read
500'd on Pydantic validation. Every server_default column was exposed the same
way; the FK/unique loss followed from the same simplified CREATE TABLE.

Build the portable schema with Base.metadata.create_all() against a SQLite
engine instead of the hand-rolled loop, so it emits the exact DDL a native
SQLite install gets (NOT NULL, DEFAULT func.now() -> CURRENT_TIMESTAMP, FKs,
unique constraints, indexes). The data-export insert path is unchanged, and
the #1333 OIDC-icon guard is preserved automatically (LargeBinary -> BLOB),
which lets the now-redundant _sqlalchemy_type_to_sqlite_type() helper be
removed. Fixes newly-created backups; a backup from an older build still
carries the degraded schema, so re-take backups after upgrading.

Replace the #1333 type-mapping unit tests with three that inspect the real
backup schema via metadata.create_all + PRAGMA table_info: icon_data is BLOB,
created_at keeps its CURRENT_TIMESTAMP DEFAULT, a NOT NULL non-PK column stays
NOT NULL.
2026-07-09 09:02:34 +02:00
maziggy 6127e30abf fix(diagnostic): skip external-storage check on P1S/P1P instead of fail (#2524)
P1-series printers have a MicroSD slot but no reachable control to enable
"Store sent files on external storage": current P1 firmware (through
01.10.00.00) never publishes support_save_remote_print_file_to_storage, so
the Bambu Studio toggle never renders, and the P1S has no screen — leaving
store_to_sdcard stuck False with no way for the user to change it. The
external_storage check reported a permanently-unresolvable fail.

Add NO_REMOTE_STORAGE_TOGGLE_MODELS (P1S, P1P) + has_remote_storage_toggle(),
kept distinct from the no-slot NO_EXTERNAL_STORAGE_MODELS. When a model has a
slot but no reachable toggle and the option is off, the check now emits skip
with params reason=unsupported_model rather than fail, and overall no longer
escalates. A P1S reporting the option on still passes. Model-scoped and
default-open, so X1/P2S/H2 (where the fail is actionable) are unaffected; if
a future firmware surfaces the capability, drop the model and it reactivates.
The frontend DiagnosticChecklist renders a reason-specific message variant
(external_storage.skip_unsupported_model) so P1 users see an accurate
explanation instead of the generic "needs a live MQTT connection" skip text.
The fix propagates to the support-bundle diagnostic snapshot automatically.
2026-07-09 08:43:59 +02:00
maziggy 5dd0370397 fix(finish-photo): bank in-print frame for FINISH-state fallback (#1867)
A1 Mini firmware never emits stg_cur=22, so every completion hits the
FINISH-state fallback — which fires after the End G-code (SwapMod plate
swap) runs, capturing the swapped plate. The last-layer edge trigger
depends on catching one transient MQTT packet and is dropped
intermittently, reverting to the post-swap grab.

Bank a rolling in-print camera frame per printer, refreshed on layer
change. Because it's layer-driven it freezes when printing ends (no more
layer increases during the swap), so the last banked frame is the
finished print. The FINISH-state finish-photo path now prefers the
banked frame over a live grab; stage_22/last_layer still live-grab.
2026-07-09 08:17:59 +02:00
maziggy 616cebdf3f Fix P1/A1 camera black screen from fan-out churn on single-connection cams (#2521)
Chamber-image printers came up black on load and only recovered ~20 min
later. Two causes: (1) late fan-out subscribers got an empty queue and
waited for the next frame, so the browser never fired onLoad and the
stall-detector reconnect-looped, churning short-lived viewers; (2) the
churn reopened the port-6000 socket before the old one closed, so the
printer fed an orphaned socket until its TCP keepalive reaped it.

Prime late subscribers with the last pumped frame; make a replacement
broadcaster's pump wait for the predecessor's socket to close before
dialing (bounded 10s); require two consecutive stalled reads before the
frontend reconnects.
2026-07-09 07:59:41 +02:00
maziggy d03b108965 Fix external-folder scan deleting README.md records; index markdown (#2520)
.md was missing from _SCANNABLE_EXTENSIONS, so scanning an external
folder skipped markdown during the walk and the cleanup pass deleted
its LibraryFile row (assuming it was gone from disk), 404ing the Folder
Readme panel. Add .md to the scannable set so pre-existing markdown is
indexed, and gate cleanup deletion on actual disk presence rather than
absence from the extension-filtered found_paths, so any non-scannable
upload still on disk survives a scan.
2026-07-09 07:20:55 +02:00
maziggy bb3e2a710e Support non-0.4mm nozzles in AMS Slot config + guard dispatch (#1899)
The Configure AMS Slot picker was hardwired to 0.4mm (nozzleDiameter
prop never passed from PrintersPage / SpoolBuddyAmsPage), so a 0.6
machine could only set 0.4 profiles on its trays. Resolve the real
installed nozzle per-AMS (ams_extruder_map on dual-nozzle) and pass it
in. Separately, nothing validated the sliced nozzle against the
installed one, so a mismatch reached the printer as a cryptic HMS
_8012 "Failed to get AMS mapping table". Add a fail-safe pre-dispatch
guard in _start_print that fails the item with an actionable message
before upload; no slice diameter or no reported nozzles = no-op.
2026-07-08 08:57:56 +02:00
maziggy 39ce37885e Bumped version 2026-07-07 13:22:13 +02:00
maziggy 70312b0a11 fix(scheduler): skip per-nozzle filter under FTS so dispatch feeds the right spool (#2186)
On a dual-nozzle H2C with a Filament Track Switch, a queued print targeting
one nozzle fed a same-type wrong-colour spool: the backend mapping
(_match_filaments_to_slots) hard-filtered candidate trays to the requested
extruder, excluding the correct spool loaded in the other nozzle's AMS — which
the FTS can route across. Confirmed from the reporter's captures: same model
mapped to AMS-A slot 2 (black) on the left nozzle but AMS-B slot 3 (red) on the
right. The #1162 FTS-skip existed only in the frontend mapping, never the
queue-dispatch path.

Read fila_switch.installed in _compute_ams_mapping_for_printer and skip the
per-nozzle filter when an FTS is present. Single-nozzle printers are unaffected
(no nozzle_id in the 3MF, no FTS). Regression tests in TestFtsNozzleBypass.
2026-07-07 12:17:46 +02:00
maziggy 5cf429f696 feat(labels): scannable QR on 203 dpi thermal printers + monochrome mode (#1870)
The 40x30 mm box label rendered its QR too densely for low-res thermal
    printers — the modules bled together and wouldn't scan. Two causes: the QR
    was 20% of inner width (~7.5 mm on the narrowest template, half of the
    others) and used ERROR_CORRECT_M. Fix adaptively so all templates benefit:
    give the roomy-layout QR a 12 mm minimum size (box_40x30 -> 12 mm, ~3.5
    dots/module at 203 dpi) and switch label QRs to ERROR_CORRECT_L (same
    payload, chunkier modules; a label needs no M-level recovery). Keep the
    quiet-zone border at 2 — the size+L gains suffice without risking scans.

    Also add a Monochrome (black & white printer) option to the label dialog:
    drops the colour swatch (a useless grey block on B&W) and widens the text;
    the hex-code line still carries the colour. Threaded through the renderer,
    route, API client, and modal, with translations in all 11 locales.
2026-07-07 11:07:08 +02:00
maziggy d8d3cde830 fix(scheduler): default require_plate_clear to False to match schema/UI (#1865)
check_queue() read the plate-clear setting with _get_bool_setting(default=True),
    but SettingsSchema.require_plate_clear defaults False and the whole frontend
    treats a missing value as off. Since _get_bool_setting returns its default when
    no DB row exists, installs that never saved the setting enforced the plate-clear
    gate the UI showed as disabled — FINISH-state printers never dispatched and no UI
    control existed to clear awaiting_plate_clear. Read the setting with default=False
    so the enforced behavior matches the schema and the toggle. Both defaults shipped
    together in #752; this aligns them.
2026-07-07 11:06:48 +02:00
maziggy 2119ddd4f9 fix(vp): populate bind-interface list on macOS (route non-Linux to psutil)
get_network_interfaces() only sent Windows to the psutil path; macOS fell into
    the Linux ioctl branch, whose SIOCGIFADDR/SIOCGIFNETMASK ioctls are Linux-only.
    macOS/BSD have fcntl but different ioctl numbers, so every call raised OSError
    and the function returned an empty list — the VP bind-interface dropdown showed
    nothing. Route all non-Linux platforms through the cross-platform psutil path.
2026-07-07 11:05:35 +02:00
maziggy e3fe2971db fix(camera): transcode non-JPEG external snapshots to JPEG (#1902)
External cameras in HTTP-snapshot mode failed to load with a repeating
    "connection lost" when the endpoint served PNG/WebP/BMP stills instead of
    JPEG (common on IP cameras and reverse-proxied snapshot URLs). The URL
    rendered fine directly in a browser, but Bambuddy's MJPEG stream wraps
    every part in a hard-coded Content-Type: image/jpeg boundary, so a
    non-JPEG payload labelled as JPEG made the browser reject the frame and
    tear down the whole multipart/x-mixed-replace stream.

    _capture_snapshot now transcodes non-JPEG stills to JPEG via OpenCV
    (already a dependency). Genuine JPEG snapshots keep a byte-for-byte fast
    path; truly undecodable responses (HTML error pages, auth redirects) fall
    back to the previous raw-return behaviour with a single clear warning
    instead of a per-frame log flood.
2026-07-07 11:04:58 +02:00
maziggy 6e03ecdb8d fix(vp): stop uvloop from silently truncating VP FTP uploads (#1896)
Native (non-Docker) installs launched uvicorn without --loop asyncio, so
    uvicorn[standard] auto-selected uvloop. uvloop's SSL layer drops
    already-received but still-buffered data when the client closes the data
    connection without a TLS close_notify while the reader is flow-control
    paused on slow storage. cmd_STOR writes each chunk to disk inside the read
    loop, so a slow consumer falls behind, the tail is lost, read() returns a
    clean EOF, and the loop exits with no exception -- the server acked 226 for
    a file it truncated itself, then archived, queued, and forwarded the corrupt
    3MF to the real printer.

    Fix in two independent layers:

    1. Remove the trigger: add --loop asyncio to every native launch path,
       matching the Dockerfile -- deploy/bambuddy.service, install/install.sh
       (systemd + launchd), spoolbuddy/install/install.sh, the Windows NSSM
       service, README, CONTRIBUTING dev command.

    2. Defense in depth (loop-independent): cmd_STOR now validates that a
       received .3mf opens as a ZIP (reads the central directory, no
       decompression) before replying 226. A truncated/corrupt file is dropped
       and answered with 426, and on_file_received never runs -- so a broken
       upload surfaces as an immediate slicer-side send error instead of being
       archived and pushed to the printer. Scoped to .3mf; other filetypes pass
       through unchanged.
2026-07-07 11:04:04 +02:00
maziggy c5b02d9473 fix(auth): let API keys manage projects via new can_manage_projects scope (#1893)
PROJECTS_CREATE/UPDATE/DELETE were in _APIKEY_DENIED_PERMISSIONS with no
    entry in _APIKEY_SCOPE_BY_PERMISSION, so every project mutation returned a
    generic 403 for any API key regardless of granted permissions -- the same
    regression class as archives (#1888) and library (#1832).

    Add a per-key can_manage_projects scope. Project routes gate on plain
    PROJECTS_* (no OWN/ALL split), so all three CRUD permissions map to the one
    scope; membership edits (add-archives) gate on PROJECTS_UPDATE and are
    covered. PROJECTS_READ is unchanged (already under can_read_status).

    Column defaults TRUE for new keys; existing rows backfill to FALSE so the
    upgrade never silently widens scope. Migration is BOOLEAN (SQLite + Postgres
    safe), verified on fresh SQLite and Postgres 17. Bundled SpoolBuddy kiosk key
    set to False. Settings API-key UI gets a Manage Projects toggle + Projects
    badge; 11-locale i18n. RBAC scope matrix + drift guards extended.
2026-07-07 11:03:46 +02:00