Commit Graph
100 Commits
Author SHA1 Message Date
maziggy 82d6d98266 Restore the pointer cursor on interactive controls (#2791)
Hovering most of Bambuddy gave an arrow rather than a hand. Not everywhere,
which is what made it read as sloppiness rather than a bug: the update pill was
inert while the buttons beside it were fine, a bed or nozzle tile responded but
the history-graph button in its corner did not, and dropdowns went either way
with no pattern behind it.

The pattern was there. Tailwind v3's Preflight set `button { cursor: pointer }`.
v4 dropped it to match the browser default, which for a button is `default`.
Bambuddy has been on v4 since the frontend was built, and `src/index.css` never
had a base layer restoring it, so a button only looked clickable where someone
had written `cursor-pointer` by hand. 15 of 934 had. 0 of 149 selects, and 19
of 130 checkbox/radio inputs. The 233 ad-hoc `cursor-pointer` usages are why it
looked arbitrary instead of uniformly broken.

One `@layer base` rule now covers button, select, checkbox, radio, summary and
[role=button]. base sits below utilities, so `cursor-not-allowed` and the
`disabled:cursor-*` variants still win; the `:not(:disabled)` guard catches the
disabled controls that carry no such utility. Verified against the built bundle
rather than the source -- the rule lands inside @layer base, and
`.cursor-not-allowed` is emitted after it.

Click-outside backdrops are deliberately excluded. 90 of the 96 remaining
onClick divs are `fixed inset-0` overlays; a full-screen sheet advertising
itself as a button is worse than one that says nothing. Of the rest, 50 are
stopPropagation wrappers and 3 are the temperature tiles, which already set the
cursor through `statusControlClass` -- which is exactly why those tiles worked
while the button nested inside them did not. That left two real ones: Card, now
conditional on an onClick actually being passed, and the queue card, whose
existing `sm:cursor-default` kept the desktop intent.

Separately, from the same report. FilamentHoverCard draws the slot menu twice,
and the two paths had drifted into opposite orders: Configure above Assign Spool
on an empty slot, the reverse on a filled one, so the menu reshuffled itself
depending on whether the slot held filament. Both now lead with the spool
action. Tests assert the order on each path, so one can no longer move without
the other -- checked by reinstating the old order and confirming the empty-slot
test fails.

Those buttons also used justify-center, which centred each label independently
and left the icons in a ragged column; they are justify-start now. Their hover
was a 10% opacity step that was very hard to see, now 20%. And the favourites
star previews yellow on hover, suppressed when the user lacks archives:update.
2026-08-08 10:06:45 +02:00
maziggy 14d0d14365 Power on a printer for jobs queued to a printer class (#2786)
Queue a print against a printer class -- "Any X1C", or a Slicer Pipeline whose
target type is Printer class -- with every printer of that class switched off,
and nothing happened. The job sat pending and no smart plug was touched, while
the same file pinned to a specific printer powered that printer on within one
queue check. The reporter's log holds both halves: thirteen minutes of the item
being polled as (133, None, ...) and passed over, then a PATCH onto printer 2,
then "Printer 2 offline, attempting to power on via smart plug(s)" on the very
next tick. Same item, same plug, same Auto On setting.

Powering a printer on had only ever been written inside `if item.printer_id:`.
The model-based branch below it walks the same queue but its matcher classes an
offline printer as a reason to keep waiting -- printers_offline collects the
*name*, for the waiting reason -- and nothing on that path ever looks at plugs.

_wake_printer_for_model adds it. The model query moves into _printers_for_model
so the matcher and the wake step answer "which printers can this job run on"
from one place: a job can only be woken onto a printer the matcher would also
have considered. Candidates that failed the cross-model gate are excluded --
switching a printer on for a file that can never legally run on it leaves the
job just as stuck, with the printer now drawing power.

Two things it does that the fixed-printer branch does not:

A printer awaiting plate-clear acknowledgment is skipped. Waking it buys
nothing; it boots into IDLE and is held by the gate. That is what the reporter's
log shows for the eighty minutes after their manual edit -- "printer 2 not
available -- connected=True, state=IDLE, awaiting_plate_clear=True" every thirty
seconds to the end of the capture. The flag is Bambuddy-side and persisted, so
it is readable while the printer is still off.

At most one printer per pass, because each wake blocks the queue loop for the
boot wait. Several queued jobs bring several printers up over the following
minutes rather than a whole shelf at once.

A failed power-on opens a 600s per-printer cool-off. Without it the walk is by
id, the pass spends its single attempt on the same broken printer every time,
and a healthy sibling two slots down is never reached -- one unreachable plug
starves its whole model, and costs a 180s boot timeout out of every 30s pass.
Entries expire on read: a printer inside its cool-off is skipped before the
power-on is reached, so a live entry can never be overwritten by a success.

The failed printer is deliberately NOT added to busy_printers. It is off, not
busy; labelling it busy would misdescribe it in every later item's waiting
reason and, because an all-busy reason is treated as needing no user action,
suppress the notification too.

Assignment is left to the next pass. AMS trays arrive with the first status push
after connect, so matching filament against a printer that booted five seconds
ago can reject the printer we just woke.

Finally, the waiting reason separates "Offline: X1C-1" from "Offline, no Auto On
smart plug: X1C-2". Those are different problems and only the second is one the
user has to go and fix -- it was also the first question the reporter had to be
asked, and the queue could not answer it.

Tests cover the wake, the plate-clear skip in both gate states, all-candidates-
awaiting-plate-clear waking nothing, one wake per pass, the starvation case over
two passes, cool-off expiry, no-Auto-On-plug being left alone and named, an
incompatible sliced model waking nothing, connected printers being left alone,
scheduled-for-later and manual-start jobs switching nothing on, and a regression
pin on the fixed-printer branch.
2026-08-08 09:35:43 +02:00
maziggy 91acac2b35 Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing
but a filename. Bambuddy opened port 990, the printer accepted the connection
and answered with something that was not TLS, and connect() logged a warning
and returned False -- indistinguishable, to every caller, from "the file is
not at this path". So the 3MF lookup walked all six filename variants across
five directories with four retries each, the cover endpoint ran its own
sixteen-path sweep, and the timelapse scan added four more, all against a
sixteen-path sweep, and the timelapse scan added four more, all against a
printer that could not have answered any of them. One reporter's log carried
1813 identical handshake failures, another's 3511.

The evidence says this is the printer's own file service getting stuck, not a
model, firmware or TLS-configuration problem. In #2780's bundle the same two
printers ran clean from 22 July to 4 August and failed again from the 5th; a
second bundle shows an X2D serving files for five days, flipping on 19 July,
then failing every connection for eight days with zero successes. The same
models and firmware appear in roughly twenty other bundles with no occurrences
at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D
and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly
this symptom, and the reporter's own debug line proves they do not.

An ssl.SSLError from connect() now opens a five-minute cool-off for that
printer. Subsequent connects return False without touching the network, so a
wedged printer is contacted twice an hour instead of hundreds of times a
minute, and the single warning that is logged names the remedy. The cool-off
is dropped on expiry rather than kept, so the map holds one key per currently
wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup
abandons the remaining paths and skips the directory-walk fallback, the cover
endpoint returns 503 naming the file service instead of a 404 that reads as
"this print has no thumbnail", and the timelapse scan separates 503 (cannot
reach the printer) from 404 (no timelapse directory) -- one 500 used to cover
both, which is what the reporter hit when reproducing.

The Connection Diagnostic completed a bare TCP connect to 990, which is why it
reported the port green throughout: the port is open, it is what is behind it
that is broken. It now completes a real implicit-TLS handshake using the
model's own ftp_profiles cap, so a pass means the FTP client would also get
through. An open port that cannot negotiate reports warn with reason no_tls,
selecting a new message in all 13 locales that points at a printer restart
rather than at the firewall. No login is attempted, so this stays valid in the
pre-save Add Printer flow.

The cool-off tests run against a real socket that accepts on 990 and replies
with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than
mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off
map too -- every test here talks to 127.0.0.1, so one left behind would make
the next test's connect() a no-op.
2026-08-08 09:00:03 +02:00
maziggy 9c86a05657 Check filament deficit for Library-backed queue items (#2779)
A job needing 20.5 g was dispatched onto a spool holding 9 g and the printer
started. _resolve_source_3mf returned LibraryFile.file_path verbatim, but that
column stores a path relative to base_dir -- so it resolved against the process
working directory, found nothing, and compute_deficit_for_queue_item treated a
missing source as "nothing to verify" and returned no deficit.

Every library-backed queue item was affected: Slicer Pipeline jobs, which are
always library-backed, and everything added through the Library's bulk Add to
queue. Both callers share the resolver, so the Play button on the queue was as
blind as the auto-dispatcher. Archive-backed items (print history, VP intake)
resolved correctly and were never affected, and neither was PrintModal, which
resolves the file on its own path.

The library branch now uses the same idiom as the eleven other readers of
file_path -- absolute stays, relative joins base_dir. The join carries a
SEC-PATH-OK marker: the value is DB-stored and generated by the Library ingest,
and it is already what resolves the file for upload, so the check has to
resolve it identically or it is not checking what gets printed.

A source that is configured but absent now logs a warning naming the item and
the resolved path. It still dispatches, because the upload needs the same file
seconds later and fails there, where blocking would strand a queue on a moved
file -- but a safety check that skips itself must not do so in silence, which
is what hid this for every library-backed item.

Tests cover the relative path (the reporter's 20.5 g against 9 g), the absolute
path against a base_dir the file is not under, and the missing-source warning.
The existing cases all used archives with absolute paths, which is the gap the
bug lived in.
2026-08-06 12:21:50 +02:00
maziggy 306b9ba7fd Accept Forgejo tokens scoped to a single repository (#2775)
ForgejoBackend.test_connection asked GET /user who the token belonged to
before asking whether the token could reach the repository, and treated a 403
there as fatal. A Forgejo v15 repository-scoped token may only carry
read/write on issues and repositories, so it 403s on /user -- and was rejected
despite reaching its own repository fine, which is all a backup needs: the push
path uses the Contents API and restore reads commits, trees and blobs, all
under /repos/{owner}/{repo}. That /user call was the only one in the whole
provider layer.

The probe stays, because a 401 from it is genuinely conclusive and names a bad
token before the repo call has to guess -- Forgejo v15+ hides a private repo
behind 404 rather than 403, so the repo call cannot always tell those apart.
Every other status now falls through to the repo check.

Two additions keep the messages as sharp as before: the repo call's own 401 is
mapped to "Invalid access token" instead of a generic API error, and the 404
names write:repository and the scoped-to-another-repository case, mentioning a
possibly-invalid token only when /user did not confirm the identity.

The token hint under the field was one shared string reading "fine-grained
token with Contents read/write" -- GitHub's advice, shown to Gitea, Forgejo and
GitLab users too. It is now per provider via PROVIDER_TOKEN_HINT_I18N_KEY,
following the existing repo-URL placeholder map, translated in all 13 locales.

Tests pin the repository-scoped token connecting, a transient /user status not
blocking the repo call, both 404 wordings, and the repo-call 401; a frontend
test switches providers and asserts the hint follows.
2026-08-06 12:08:36 +02:00
maziggy 9beb001a17 Record who queued a file from the Library and the webhook API
PrintQueueItem.created_by_id is what the queue:read_own / queue:update_own /
queue:delete_own permissions filter on, but only three of the paths that create
queue items were setting it.

The Library's bulk "Add to queue" required Permission.QUEUE_CREATE and then
bound the dependency to `_`, discarding the user, so every item it created was
ownerless -- and invisible to the person who added it if their permissions are
scoped to their own work. That is the one path built for adding many files at
once, which is where it was hardest to notice.

The webhook queue endpoint has no request user, but APIKey.user_id records the
key's owner, which is the acting identity everywhere else the key is used, so
its items are credited to that owner. Keys minted before per-user ownership
have no user_id and their items stay ownerless.

The virtual-printer path is left as-is on purpose. VirtualPrinter carries no
owner, and the obvious substitute is wrong rather than incomplete: one admin
typically configures the VP while everyone sends prints through it, so
crediting those to the admin would make the "added by" column lie and put other
people's jobs in the admin's own queue. Existing NULL rows are not backfilled
-- there is no record of who created them, and the ownerless case is already
handled throughout.

Tests pin both fixed paths and the two cases that must stay ownerless (auth
disabled, legacy key).
2026-08-06 11:06:15 +02:00
maziggy afa0ba0dc0 Nest projects under a master project and roll their figures up (#1264)
Projects were flat. The parent_id column and the sub-project list existed
but nothing could set a parent outside the API, and a master project's
stats only ever covered its own prints.

The project dialog gets a parent picker, and a project with sub-projects
gets a second card covering the whole tree -- jobs, parts, time, filament,
cost, and progress against every target in the tree added together. That
card is separate from the project's own stats, which keep their existing
meaning; widening them would have restated the figures of anyone who had
already nested projects over the API. Each listed sub-project carries its
own branch's roll-up, so the rows add up to the card above them.

On the Projects page a sub-project is drawn inside its parent's group
rather than as another card in the grid -- two cards columns apart cannot
show that they belong together, whatever the caption says. A sub-project
whose parent the status filter has hidden stays put and names its parent
instead.

compute_project_stats now goes through the same grouped aggregation as the
roll-up rather than its own copy of the SQL, since the two must agree.

Three things the interface made reachable:

- PATCH refused only a project as its own direct parent, so A -> B -> A
  was two calls away. A cycle has no root to roll up to, and the walk
  keeps its seen-set for databases that already contain one.
- A sub-project's percentage was completed quantities against the plate
  target, disagreeing with the page it linked to.
- Deleting a mid-tree project orphaned its children at top level; they
  now move up to its own parent.
2026-08-06 10:00:45 +02:00
maziggy 312f09a4ac Updated CHANGELOG 2026-08-06 08:50:01 +02:00
maziggy b5163b94f8 fix(backup): report the categories a failed restore already committed (#2656)
The service reports what landed on a part-way failure -- categories commit as
they finish, so results names the ones on disk -- and the modal gated the
whole result panel on success, so it showed the failure message and dropped
them.

The cache invalidation was inside that same branch, which is the half that
mattered: a run that committed the settings category and then failed left the
app rendering pre-restore settings, with no reload and no re-read, which is
the failure the modal's own reload-on-close exists to prevent.

Gate on what was written instead. A refusal that never reached a category
still carries an empty results and still keeps the form, so the mutex and
backup-in-flight cases are unchanged. A partial does not read as a success:
the tick becomes a warning and a line says the listed categories are the ones
on disk.

---

fix(backup): keep the local owner when the backup names one we cannot resolve (#2656)

An owner the backup names but this instance has no user for was written as
NULL, and overwrite is a blanket setattr -- so restoring over a local archive
that had a perfectly good owner took it away, which is the 404-for-its-own-
owner failure this column is carried across to fix. Resolving by username
widened the trigger from a stale id to any user renamed since the backup.

It is the same state as an absent key: the backup has not told us who owns
this. So it takes the same action -- the column is not written at all.
Overwrite keeps the local owner, insert lands ownerless with the note, and an
explicit null still writes, so overwrite still means "match the backup".

The notes move to the insert path with it. On overwrite nothing was taken
away, so there is nothing to warn about, which is the rule the absent-key
case already follows.
2026-08-06 08:46:02 +02:00
maziggy 1eea194953 Resolve a spool's material to a known drying preset before starting a cycle (#2774)
The drying popover prefilled its material from the loaded spool without
checking the preset table had that material. An AMS-HT holding Support for
PLA/PETG (tray_type PLA-S) fell back to PLA's temperature but kept PLA-S as
the material, and the dropdown displays its first option when handed a value
outside its list -- so it read PLA while PLA-S was sent. Same gap for every
composite: PETG-CF prefilled at PLA's 45C.

Resolve the tray_type to a key the table has before setting either value.
Support materials and composites resolve to their base, nylon is aliased
under its several spellings, and anything unrecognised falls back to PLA --
the coolest row, so an unknown material under-dries rather than deforming a
PLA spool.

Also record request-topic messages in the MQTT debug log. That topic carries
every command a printer is given, including Bambu Studio's, and returned
before the logging block -- so a capture could show only what the printer
said, never what it was told.
2026-08-06 08:19:01 +02:00
maziggy 16c8c6f2ea Security hardening (maziggy/bambuddy-security #9) 2026-08-05 15:21:42 +02:00
maziggy 684a328d6f Broadcast AMS slot changes that keep the same material
Configuring a slot from the printer card left the card showing the old
filament until a reload or the 30s fallback poll. The command reached the
printer and the printer applied it; the update just never got broadcast.

on_printer_status_change deduplicates WebSocket pushes against a status_key
whose AMS part carried id, tray_type and state. Configure Slot writes none of
those -- it writes tray_info_idx, tray_color, tray_sub_brands and cali_idx. So
PLA to another brand or colour of PLA produced an identical key and was
dropped, while PLA to PETG came through. Reset always worked because it clears
tray_type.

Those four fields only move when someone configures a slot or swaps a spool,
so this costs no broadcasts mid-print. remain stays out of the key for the
opposite reason.
2026-08-05 14:53:05 +02:00
maziggy 3bbe00784f Write printer status straight through while the tab is hidden (#2754)
Removing the requestAnimationFrame wrapper fixed the total stall but left the
100ms coalescing timer in the path, and a hidden page's timers are clamped to
once a second at best -- once a minute past five minutes hidden. The reporter
still saw a tab title at 2% beside a page at 40%.

The coalescing guards against a render cascade, which a hidden tab cannot
have, so it is skipped there and kept while visible.

The existing hidden-tab tests advanced fake timers, which simulates the timer
the browser was throttling; the new one never advances the clock.
2026-08-05 14:40:05 +02:00
maziggy cd004df817 Show Home Assistant sensors on the printer card (#1148, #448)
Binds binary_sensor and reading-carrying sensor entities to a printer and
renders their state on its card, worded by Home Assistant's device_class.
Optional per-sensor alert condition drives a notification on the transition
into the alert state and an opt-in interlock that holds queued prints while
alerting -- a hold with a readable waiting_reason, never a failure, and only
ever on a sensor that was read successfully.

Sensors get their own table rather than a wider entity pattern on SmartPlug:
get_smart_plug_by_printer would otherwise hand the card's power button a door
contact to switch.

The hold is passed to the model matcher directly rather than merged into
busy_printers: _check_auto_drying reads that set as "is currently printing"
and would put an idle-but-held printer down the mid-print drying path.

The notification_providers migration spells its default FALSE, not 0 --
Postgres rejects an integer default for a boolean and _safe_execute swallows
the error.
2026-08-05 14:26:38 +02:00
maziggy 945d4ca6eb Route external spools to a nozzle when the printer has no AMS (#2771)
Five X2Ds with no AMS, each printing from its external spool holder,
took a job sent to a named printer and refused the same job sent to
"Any X2D": the file uploaded, the firmware answered 0700_8012 "Failed
to get AMS mapping table", and the item failed after three attempts.

A named-printer job carries a mapping the frontend resolved at queue
time, so the scheduler's matcher never runs. A model-based job has no
printer until dispatch, so the matcher does run -- and could not see an
external spool on a dual-nozzle printer. _build_loaded_filaments derived
dual-nozzle status from ams_extruder_map, which is built from AMS info
bits, so a printer with zero AMS units reported an empty map; every
external spool got extruder_id=None, and the nozzle-aware hard filter in
_match_filaments_to_slots discarded it because None equals neither 0 nor
1. The mapping came back all -1, was cleared to None, and the print
command went out as use_ams:true with no ams_mapping and no
ams_mapping2 at all.

This is the backend half of #1257, which fixed the same logic in
useFilamentMapping.ts and left this copy behind. Mirror its inference:
a populated nozzles[1].nozzle_diameter, a non-empty ams_extruder_map, or
more than one vt_tray entry. Replaying the reporter's own push-status
now yields extruder 1 for Ext-L and 0 for Ext-R, and a nozzle-1
requirement resolves to [254] -- what their working named-printer
dispatch sent. Single-nozzle printers keep extruder_id=None; nozzles
always has two entries, so its length alone must not be the signal.

Also stop dispatching a job the firmware is certain to reject. When the
matcher ran, matched nothing, and the printer has no AMS, fail the item
with the filament and nozzle it wants instead of spending an upload and
two retries on it -- that path already ended in a failed item, just an
opaque one. With an AMS attached the firmware error still stands, since
there the user can load a spool and press Resume. Fail-safe like the
nozzle-diameter guard (#1899): every branch short of a positive finding
returns None and dispatches as before.

_apply_filament_overrides is extracted from _compute_ams_mapping_for_printer
so the message names the filament the matcher looked for rather than the one
the 3MF was sliced with.
2026-08-05 13:15:04 +02:00
maziggy 0596ff424e Say why a drying cycle ended when the firmware cuts it short (#2770)
An H2D started a twelve-hour PETG dry at 65 degC and the AMS gave up on it
twenty minutes in, with 700 of the 720 minutes still on the clock. It cooled,
humidity climbed back over the threshold, auto-drying started another
twelve-hour cycle, and that one went the same way; the reporter's AMS
temperature history shows the loop running all morning.

The log had one line for it: "AMS 0 drying complete", which is exactly what it
says for a dry that ran its full twelve hours. Nothing in a support bundle told
the two apart, and the one number that does -- the time still remaining -- was
written into that line as the previous value, where it reads like a duration
rather than a shortfall. The reporter took 700 for seconds and concluded the
cycle had lasted twelve minutes.

Bambuddy did not stop that cycle; every stop it sends is logged with the full
outgoing command and there was none. So ending it was the printer's decision,
and the account of why lives in three things already received and parsed and
never written down: the drying phase and sub-phase from the AMS info hex, the
per-unit dry_sf_reason constraint codes, and the live HMS errors.

A cycle that ends with most of its countdown left now logs all three alongside
how much of the requested duration ran. One that reaches its duration keeps the
single line it has always had. A stop Bambuddy sent is named as ours -- it is
short of its duration too, and on the telemetry alone is indistinguishable from
the firmware abandoning the cycle, so without tracking it the print-takes-priority
stop and the Stop button would both have been blamed on the printer.

Diagnostics only. Nothing about when drying starts or stops has changed, and the
restart loop is not addressed: what the firmware objects to has to be established
before Bambuddy can sensibly decide how long to wait before trying again.
2026-08-05 12:38:45 +02:00
maziggy 0dc911fce1 Updated CHANGELOG 2026-08-05 12:21:21 +02:00
maziggy fb4c130bb2 fix(slicer): drop the legacy library:read from the slice gate (#2725)
The desktop handoff accepted library:read alongside read_all/read_own, on
the reasoning that default groups do not carry it and requiring it would
lock out Operators and Viewers. The permission grants nothing in that
position: the slicer-token endpoint gates on
require_ownership_permission(LIBRARY_READ_ALL, LIBRARY_READ_OWN), and
neither that dependency nor User.has_permission expands the legacy name,
so a group holding only library:read gets a 403 there. It cannot reach
the File Manager to try, either - GET /library/folders gates on the same
pair - and the library:read -> library:read_own migration in
core/database.py runs only over the groups named in DEFAULT_GROUPS, so a
custom role that still carries it stays stuck rather than being upgraded.
custom role that still carries it stays stuck rather than being upgraded.

Accepting it only enabled a menu item the server refuses, and the failure
is indistinguishable from "no slicer installed" once the catch hands the
unauthenticated URL over. Removed, with the comment recording the reason
so the next reader does not re-add it, and a test that pins it.

---

refactor(slicer): share one sliceable-file-type rule (#2725)

The File Manager and the 3D preview decide the same thing about the same
file and each held its own list of extensions - which is how they came to
disagree, offering a desktop handoff for an STL whose own preview showed
"Open in Slicer" greyed out. Making the two lists identical fixed the
symptom and left the drift, so SLICEABLE_FILE_TYPES now lives in
utils/slicer.ts with isSliceableFileType for a stored file_type and
isSliceableFilename for a name.

The filename form still rules out the compound extensions explicitly,
since .gcode.3mf ends with .3mf; the type form does not need to, because
classify_file_type stores that one whole.

Both test files mocked the whole slicer module, which would have replaced
the new predicates with undefined - switched to importOriginal so only
openInSlicer is stubbed. That is the better shape regardless: the tests
now exercise the rule the component runs instead of a copy declared
beside them.

Carries the rebuilt bundle. The CSS hash moves with it - the split button
introduces Tailwind classes the previous build had no reason to emit.
2026-08-05 12:11:55 +02:00
maziggy a74dc7932f Cover nested data structures in the HA notify pass-through (#1441)
The three tests around it use flat scalars, which is also all the field's
placeholder and the wiki showed, so nothing recorded that the value is
forwarded verbatim rather than treated as a key/value list. A reporter asked
whether action buttons work; they always have, and now that is pinned.

The changelog entry said "nested options work" and left it there. It now names
actions and the two things that decide whether the buttons do anything - the
mobile_app_notification_action automation, and iOS needing a registered
category - since neither is set from Bambuddy and both are what a reader would
otherwise have to discover the way the reporter did.
2026-08-05 11:51:30 +02:00
maziggy 5efbd353ed Survive a directory that defines no posixGroup class (#2769)
Every LDAP user on an lldap directory was rejected with "Incorrect
username or password", on an install where Test Connection passed and
where the same bind DN, filter and group membership all checked out
under ldapsearch. The directory never saw the request.

_extract_user_info searches for POSIX groups alongside the memberOf
ones, and both of those filters name the posixGroup object class. ldap3
fetches the schema at connect time (get_info=ALL) and validates class
names in a filter against it while building the request, raising
LDAPObjectClassError before anything is sent. lldap marks every account
it creates as posixAccount -- which is what makes us look for POSIX
groups at all -- but defines no group class beyond groupOfNames. The
exception escaped authenticate_ldap_user, and the login route reports
any LDAP failure as bad credentials.

A directory with no posixGroup class has no posixGroup entries, which is
exactly the answer those searches would have returned. Catch it, log it
once, and carry on with the memberOf groups collected above. Both
searches sit inside the one try: they name the same class, so once one
is rejected the other cannot succeed, and attempting it would only
produce a second identical exception to swallow.

Not a regression from 848f55810. The memberUid filter has named the
class since b6599dd41 and runs for every user whether or not they have a
gidNumber, so a directory of this shape has never been able to log in;
the primary-group lookup only added a second trigger. Test Connection
was unaffected throughout because (objectClass=*) is a presence filter
and never reaches the value validator.

The mock connection gained a hook that raises on a filter substring,
standing in for that client-side validation. Reset in the fixture and
default None, so existing tests are unchanged.
2026-08-05 11:14:57 +02:00
maziggy 6fb6b845e7 Read the printer's own slot mapping when Spoolman has none (#2768)
A sliced file numbers its filaments 1..4; which AMS tray each came from is
a separate decision made when the job is sent. store_print_data learns it
from one of two sources, both of which require the print command to pass
through us: the mapping Bambuddy chose itself, or the one it intercepted
on the printer's local request topic. A job dispatched from Bambu Studio
while the printer is cloud-bound satisfies neither -- the command travels
through Bambu's broker and never reaches the topic we subscribe to.

slot_to_tray is then NULL and _resolve_global_tray_id guesses by position:
filament 1 from the first loaded tray, filament 2 from the second. The
reporter's X1C was loaded in the order 2, 4, 1, AMS-HT, so all four slots
were charged to the wrong spool. Their log carries the printer's own
answer, mapping=[1, 3, 0, 32768], sitting unread.

usage_tracker has consulted that field since it started resolving mappings
at completion, along with a colour match against the loaded trays for the
models that never publish it (A1, A1 Mini, P1S, P2S). Only the Spoolman
writer, which resolves at print start, never learned to -- and main.py
gates usage_tracker behind Spoolman being off, so enabling Spoolman is
what costs you the better resolver.

_resolve_slot_to_tray_fallback gives it both, at completion rather than at
print start: a printer keeps publishing the last job's mapping while it
sits idle, so reading it early would risk stamping the previous print's
mapping onto this one. A mapping we or the slicer actually recorded is
never second-guessed.

Applied in _report_partial_usage too. Cancelled and failed prints feed the
same slot_to_tray to the same resolver and mis-charged just as readily.

The resolved mapping and its source are now logged at print start and at
completion. "source: none" at start is the signal that completion will
have to fall back, and it was the one line that would have turned this
report into a five-minute triage.

Not addressed: editing the mapping after the fact, which the reporter also
asked for. ArchiveUpdate exposes neither filament field and there is no way
to re-run an attribution, so that is a feature rather than a fix.
2026-08-05 11:04:59 +02:00
maziggy fce7ea0200 Stop the drying badge inventing a temperature on a uniform AMS (#2759)
The follow-up to the same report: a second AMS 2 Pro, no aux power,
loaded entirely with PLA and drying at the 45C the reporter picked,
showed 45C and then switched to 55C.

Bambu never echoes back a cycle's filament or temperature, so both come
from the target cached when the command went out, and the fallback for
a missing cache reads the loaded trays. The first pass narrowed that
fallback to units whose spools agree on a filament, which fixed the
mixed-unit case in the original report but left the uniform case
answering with the spools' RFID-recommended drying_temp -- 55C here.
Agreement across slots is evidence of what is being dried, because the
dryer heats all of them. It is no evidence of the temperature, which is
picked freely in the popover, so the recommendation was never more than
a guess wearing the same confident "PLA @ 55C" as a known target.

uniform_tray_drying_hint therefore becomes uniform_tray_filament_hint
and returns the filament alone. The badge names a temperature only when
we sent it, and otherwise shows the filament and the countdown.

Both status builders also stopped filling the two fields independently.
Entering the fallback when either was missing let a cached filament pair
with a guessed temperature and render as though both were known; the
temperature now simply has no fallback to reach.

The badge required both fields before rendering anything, so dropping
the temperature would have blanked it rather than shortening it -- the
frontend now renders each on its own terms. No new translation key: the
filament type is a passthrough.

This changes what is shown when the cached target is missing, not why
it goes missing. If the reporter was on the fixed build, the falling-
edge gate is still letting a zero through on an unpowered unit, which
needs a log covering the start of the cycle.
2026-08-05 08:12:42 +02:00
maziggy cb508c5d16 brace-expansion override ^5.0.8 -> ^5.0.9 (GHSA-rgw5-rvv9-x895, DoS).
5.0.8's maxLength cap was applied in combine(), where output is merged, but
not to the two arrays built before it runs: comma alternatives each got their
own full allowance and were concatenated with no running total, and padded
sequences never consulted maxLength at all. So a ~25 KB pattern still OOMs the
process -- fatally, past the reach of try/catch -- and a ~400 KB one blocks the
event loop for over two minutes. 5.0.9 bounds both as they are built.

Dev-only and transitive here: it reaches us as eslint -> minimatch@5 ->
brace-expansion, the only input it sees is our own lint globs, and it is not in
the shipped bundle. The ci.yml audit gate runs --omit=dev, so this never would
have failed CI; it surfaced through Dependabot.

The overrides floor is bumped alongside the lockfile so a clean install can't
resolve back to the vulnerable 5.0.8.
2026-08-05 07:48:23 +02:00
maziggy 78cbd82259 Scale the printer card's body text and icons with its size (#1848)
Switching a card from M to XL made it wider, enlarged the printer name
and the thumbnail, and left everything else where it was. The AMS slot
labels, temperatures, filament names, status text and every small button
stayed pinned between 8 and 11 pixels -- under the smallest size used
anywhere else in the app -- so a full-width card carried the same tiny
text as the compact one. Browser zoom does not answer this: it enlarges
the whole page and so preserves the very disparity being reported.

The card root now carries ten custom properties derived from cardSize,
and the 200 fixed sizes in its subtree reference them: text-[10px]
becomes text-[length:var(--pc-t10,10px)], w-3 h-3 becomes
w-[var(--pc-i3,0.75rem)]. L draws the body 20% larger and XL 40%,
icons included, so the controls grow with the text instead of staying
fiddly to hit.

Custom properties rather than an em-based root font-size. Setting
font-size on the card would silently reshape any text that declares no
size of its own, and would break for portalled content. Each converted
class names its old fixed value as the fallback, so anything rendering
outside a card root is untouched -- which is what leaves the portalled
temperature popover exactly as it is. Its four sites stay fixed on
purpose, as does the page chrome; the conversion was scoped from the
function declarations rather than line numbers, and afterwards only
those four intended sites still hold a literal px value.

S and M stay at 1.0. S is the dense fleet view where density is the
point and M is the default, so an existing install looks identical until
the user reaches for a size that is already asking for more room -- the
same control the request asked this to follow.

The AMS-HT card needed separate work, because its temperature and
humidity readings sit beside the slot rather than under it. That single
slot was the only growable item on its row, so it took every spare pixel
and pushed the readings hard against the card's edge; it is now capped
at roughly two ordinary slots, which keeps them clear at any card width.
The card itself is capped at one full AMS card's width, so a unit that
wraps onto a line of its own no longer stretches that slot across the
whole card.

The AMS slot minimums are deliberately NOT scaled. Raising them was
tried and reverted: those cards already grow to fill their row, so
3.5rem is a floor they sit well above, and raising it only cost a unit
its place on the row -- which is what pushed the AMS-HT onto a line by
itself and exposed the stretching above. A test pins them at 3.5rem at
XL so this reads as a decision rather than a missed spot.
2026-08-04 14:09:10 +02:00
maziggy 45b678692c Scale the printer card's body text and icons with its size (#1848)
Switching a card from M to XL made it wider, enlarged the printer name
and the thumbnail, and left everything else where it was. The AMS slot
labels, temperatures, filament names, status text and every small button
stayed pinned between 8 and 11 pixels -- under the smallest size used
anywhere else in the app -- so a full-width card carried the same tiny
text as the compact one. Browser zoom does not answer this: it enlarges
the whole page and so preserves the very disparity being reported.

The card root now carries ten custom properties derived from cardSize,
and the 200 fixed sizes in its subtree reference them: text-[10px]
becomes text-[length:var(--pc-t10,10px)], w-3 h-3 becomes
w-[var(--pc-i3,0.75rem)]. L draws the body 20% larger and XL 40%,
icons included, so the controls grow with the text instead of staying
fiddly to hit.

Custom properties rather than an em-based root font-size. Setting
font-size on the card would silently reshape any text that declares no
size of its own, and would break for portalled content. Each converted
class names its old fixed value as the fallback, so anything rendering
outside a card root is untouched -- which is what leaves the portalled
temperature popover exactly as it is. Its four sites stay fixed on
purpose, as does the page chrome; the conversion was scoped from the
function declarations rather than line numbers, and afterwards only
those four intended sites still hold a literal px value.

S and M stay at 1.0. S is the dense fleet view where density is the
point and M is the default, so an existing install looks identical until
the user reaches for a size that is already asking for more room -- the
same control the request asked this to follow.

Wiki notes the scaling in the card-size table and why it differs from
browser zoom. Tests pin the variable values at every size, including
that S and M still emit the pre-change sizes.
2026-08-04 13:39:54 +02:00
maziggy 3db8ac9da7 Let the external spool be hidden from the printer card (#1782)
An external spool holder that never gets used still takes a full card's
width in the Filaments row, next to the AMS units that are actually in
use. An eye icon at the right-hand end of that row's header now hides
it, and clicking it again brings it back -- the affordance stays in
place rather than moving to a settings page, so the choice is
discoverable and reversible where it applies.

Per printer rather than global. A global flag would suit a toolbar
button, but an icon on the card that silently rearranged every other
card would surprise; it is keyed by printer id in one localStorage
entry, the same shape as printerCollapsedSections, and sits alongside
the other browser-local printer-page view preferences.

The toggle is offered only when the printer has at least one AMS. On a
machine with no AMS the external spool is the entire filament section,
so hiding it would leave an empty row with no control to undo it. The
icon and the hide condition read the same canHideExternalSpool, so a
preference stored before an AMS was unplugged cannot blank the row
either -- the spool reappears instead.

The store lives in a new utils/printerCardPrefs.ts rather than in the
9,157-line page. It re-reads before writing so two cards toggled in one
session cannot clobber each other's entry, deletes the key instead of
storing false, and treats a malformed or unavailable localStorage as
"nothing hidden" so a private-mode browser cannot throw out of a render.
2026-08-04 13:08:30 +02:00
maziggy 9bc96aeb83 Hold error and warning toasts for twice as long
Every pop-up notification auto-dismissed after three seconds regardless
of what it said. That suits "Settings saved" -- a confirmation of
something the user just did, skimmed rather than read -- but errors and
warnings are a different kind of message. They carry a reason, often one
relayed from the printer or the backend, and they run to a couple of
lines. Three seconds was not enough to finish reading one, and there is
no notification history to go back to once it slides away.

Errors and warnings now hold for six seconds; success and info keep the
three-second default. The duration was a bare literal in showToast and
is now derived from the toast type, with the long window expressed as
twice the base so the two cannot drift apart if the base is retuned.

showPersistentToast never had an auto-dismiss timer and is untouched, as
is the background dispatch toast -- its timer measures "the summary has
stopped changing" rather than reading time. Manual dismissal is
unchanged for every type.
2026-08-04 12:54:24 +02:00
maziggy 33ab5f1ead Add temperatures to the streaming overlay and a URL builder (#1422)
The overlay at /overlay/{printer} draws live print data over a
full-screen camera view for OBS, a wall display or any browser source.
It has been tunable since it shipped -- which fields, what size, what
frame rate -- but only through query parameters documented in the wiki,
and temperatures were not among the fields on offer. The request asked
for temperatures first and for the field set to be selectable in the web
UI; this addresses both.

Nozzle, bed and chamber readings join the list. The target is drawn only
while the heater is still climbing, so a settled hotend reads "220°C"
for the rest of the print instead of the noisier "220 / 220°C" -- 219.6
against a target of 220 rounds to the same number, and repeating it says
nothing. Both nozzles appear on a dual-nozzle machine. They are drawn
whether or not a print is running, because a preheating printer is
exactly when they are worth watching, and each reading appears only when
the printer genuinely reports one: chamber temperature stays absent on
P1 and A1 models, which publish a chamber_temper with no sensor behind
it, so the overlay never puts a measurement on screen that does not
exist. Labels reuse the heater chart's strings rather than inventing a
second vocabulary for the same three things.

The feed sends an allow-list rather than the temperatures dict. That
dict doubles as the MQTT client's working memory -- derived heater flags
and private target-set timestamps live alongside the readings -- and an
overlay token is a narrower grant than a login, so it gets exactly what
the overlay draws and does not pick up fields as the dict grows. The
same chamber-sensor gate the full status payload already applies is
applied here. The integration test that asserts the payload's exact key
set, which exists to catch that surface widening silently, is updated
deliberately.

Temperatures are not in the default field set, so an overlay URL already
pasted into a scene renders identically after upgrading.

Settings -> API Keys -> Streaming Overlay now builds the URL: printer,
field checkboxes, size, frame rate, camera toggle, an optional token,
and a copy button. It persists nothing and calls nothing new -- the URL
is the configuration, which keeps a scene reproducible by copy-paste and
lets two displays show different fields off one token. Fields are
emitted in the overlay's own top-to-bottom order rather than click
order, and parameters left at their default are omitted, so the same
selection always produces the same URL. The preview alongside it stays
off until asked for: an always-live iframe would hold a subscriber on
the printer's single camera connection for as long as the settings tab
stayed open.

The preview needed one narrow security-header change. Every SPA route
sent frame-ancestors 'none', which is stricter than the SAMEORIGIN in
X-Frame-Options beside it and refuses even a same-origin frame, so the
preview showed Firefox's "another site has embedded it" page instead of
the overlay. The overlay path now sends 'self', mirroring /gcode-viewer,
which admits a framer only on this origin -- Bambuddy's own UI. Every
other path keeps 'none', and embedding the overlay from another host
still requires TRUSTED_FRAME_ORIGINS.
2026-08-04 12:38:36 +02:00
maziggy 48c231d8ce Stop a drying cycle reporting itself finished a minute in (#2759)
Starting the dryer on an AMS 2 Pro holding two PETG and two PLA spools
and picking PLA showed "PLA @ 45°C" for about a minute and then switched
to "PETG @ 65°C" for the remaining twelve hours.

Bambu never echoes back which filament or temperature a cycle is
running, so the badge reads the target cached when the command went out,
and that cache had been dropped. Between accepting the command and
settling its countdown the firmware publishes one update with the
remaining time at zero while the unit is still in its Checking phase --
the reporter's log has 720, then 0, then 719, and four seconds later the
same unit's info hex decodes to dry_status 2, Drying. The falling-edge
detector read that zero as the cycle ending. Losing the cached target
left the badge to guess from the first loaded slot, which was PETG, and
its RFID-recommended 65°C. The same false ending fired
on_drying_complete, so anyone with smart-plug auto-off-after-drying
switched on had power scheduled to cut one minute into a twelve-hour
dry; the reporter had it off, which is the only reason this reads as a
cosmetic bug.

A remaining time of zero now ends a cycle only when the unit is not also
reporting an active phase. dry_status comes from the same info hex
already parsed a few lines above, so this costs nothing to check.
Stopping and Error are deliberately not treated as active -- those
should end it -- and a unit that reports no phase at all still ends its
cycles, so the gate can only ever suppress on positive evidence that the
cycle is live. A suppressed edge leaves the remembered dry_time alone,
exactly as the #1462 absent-value skip does, so the push that really
ends the cycle still sees a non-zero previous.

The fallback guess is tightened to match. On a mixed unit the first tray
is evidence of nothing, and naming a temperature the cycle is not using
is worse than naming none, so it now answers only when every loaded
spool is the same filament and otherwise leaves the badge showing the
countdown alone. Both the websocket and REST status builders carried
their own copy of that loop; they now share one helper, which also takes
the temperature from the first slot that carries an RFID one rather than
giving up when slot 1 holds a third-party spool.
2026-08-04 11:47:17 +02:00
maziggy 71a06f3638 Add batch orders with a quantity per plate (#342)
Printing a multi-plate file in different quantities per plate meant
queueing each plate separately and tracking the counts by hand: one
shared Quantity field cannot say "plate 1 once, plate 2 twice, plate 3
three times". Each selected plate now carries its own quantity, and the
submission becomes an order on a new Batches tab.

The point is the distinction the old flat batch could not express.
print_batch_plates stores how many runs of each plate were wanted,
separately from what was queued, so a run that fails, is cancelled or is
skipped does not satisfy a target -- the order goes on saying it owes a
print instead of quietly under-delivering. Queue remaining re-queues
exactly what is missing, for the whole order or one plate, by cloning
the most recent item for that plate: that inherits the printer or model
target, AMS mapping, filament overrides and print options along with the
validation they already passed, rather than re-serialising twenty fields
through a template that would drift from the model the first time
someone adds a column. Clones append to the end of the relevant
printer's queue and take the same advisory lock the add-to-queue route
does; positions are per-printer sequences, not global.

Cost is measured, not estimated. print_log_entries gains queue_item_id,
set where the queue item is already in scope, so each run's material and
energy are attributed through the item that produced them -- an
unrelated reprint of the same archive never lands in an order's total,
and a multi-plate order gets each plate's own cost rather than the whole
file's via the plate-scoped estimate from #2614. Before any run has
completed there is no honest figure, so cost reads as unknown instead of
a fabricated 0.00.

The Batches tab wires up GET /queue/batches, which has been unreferenced
since the batch MVP shipped, along with six locale keys that were
translated and never used. It is a separate tab because an order
outlives the queue that produced it: once its runs finish they leave the
active queue, so Queue and History each hold half the picture.

completed was not a reachable status before now, so every batch created
since April is still marked active however long ago its last print
finished -- 73 of them on the development install. A startup pass closes
out the finished ones: those whose runs all completed become completed,
and groupings whose items were all cancelled become cancelled, which is
what they are. Not applied to orders, which state their intent
independently of their runs and still owe the work. Only batches with
nothing queued or printing are considered, and repeating the pass also
catches an order whose last run landed while the process was down.
Batches with neither items nor targets are no longer listed at all --
empty shells left when a grouping's items went with their source
archive.

Dispatch applies the same source-file gates as POST /queue/. It creates
queue items, so without them it would be a weaker door to the same
outcome; the archive and library-file checks move into shared helpers
so a third route cannot drift from them.
2026-08-04 11:11:36 +02:00
maziggy ebc73e46ba Say when AMS drying was running and a print never started (#2758)
Dispatching to an X2D with two AMS units mid-drying failed silently: the
file uploaded, the printer accepted it and stayed idle. The watchdog waits
for an active state or HMS_MQTT_VERIFY_FAILED, and a drying refusal is
neither, so it timed out, re-uploaded the whole 3MF twice more, and closed
with advice about the printer screen and the SD card. Studio, asked
directly, said it could not start the job because of the drying.

Latch the AMS dry_time telemetry across both watchdog phases and name the
units in the give-up message, plus an INFO log on every failed window so
the correlation reaches a support bundle from the first attempt.

Detection only, no gate. These models support drying CONTINUING through a
print (supports_drying_while_printing covers X2D from 01.01.00.00), so
drying is not incompatible with printing and stopping it before every
dispatch would tear down cycles the hardware is happy to run. One of the
two units was also drying without its external PSU, which would make this a
power budget problem at start-of-print calibration rather than a drying one
-- dry_sf_reason 1/8 exist for exactly that. The message names both
possibilities rather than asserting one.

Also correct _sync_drying_state's docstring, which claimed to adopt drying
it did not start; it only prunes. Behaviour unchanged -- populating it would
let the scheduler stop a cycle the user started by hand.
2026-08-04 08:49:49 +02:00
maziggy 28a6ca6f4d Add CAP_NET_BIND_SERVICE everywhere the service is defined (#2549)
The Virtual Printer binds 990 and 322, below 1024, which a service running
as a normal user may not do without CAP_NET_BIND_SERVICE. Without it the
rest of Bambuddy works and only the VP is dead -- sockets never open, the
slicer never finds the printer, and the sole trace is one journal line.

332a7c6ac added the line to install/install.sh in March under the heading
"Fix install.sh missing AmbientCapabilities". Three other places define the
same unit and none of them got it: the manual template, the combined
Bambuddy + SpoolBuddy installer, and the unit the wiki tells you to paste.
The wiki additionally claimed the capability was always included.

Also diagnose it. The VP diagnostic reported only that nothing was listening
on 990, which reads identically to a port conflict. It now checks CapEff for
the capability and names it as the cause -- but stays quiet when the port is
answering (an iptables REDIRECT is the documented alternative and that host
works) and when the capability is held (the port is down for another reason
and blaming this would misdirect). Skips where there is no procfs rather
than putting a systemd instruction in front of a macOS user.
2026-08-04 08:33:30 +02:00
maziggy 4af782cc2f Report a refused AMS filament setting instead of discarding it (#2756)
Configuring a slot publishes ams_filament_setting and the printer answers
with a verdict. The answer was received and dropped at DEBUG, so a refusal
left no trace at the level support bundles are collected at: the reporter
saw six Configure Slot attempts on an X1C all return success, all read back
by the #2582 verification as holding the previous profile, and no record of
what the printer said about any of them.

Promote a non-success response to INFO with result, reason, ams_id and
tray_id. Refusals only -- unlike extrusion_cali_set (#2718) and
ams_filament_drying (#1447) this command is not rare, since every spool
assignment and K-profile re-apply sends one, so promoting each ack would
bury the interesting line.

The developer-mode probe is excluded: it sends this same command to the
external slot expecting a refusal on P1 firmware, so promoting it would
put an alarming line in every P1 bundle on every reconnect. Matched by
sequence id, which user commands cannot collide with -- they publish a
hardcoded "0".

Diagnostics only; no change to which commands are sent or how they are built.
2026-08-04 08:08:55 +02:00
maziggy c28e053126 Keep live updates flowing while the tab is in the background (#2754)
Printer status, query invalidations and the message queue all ran their
work inside requestAnimationFrame. A hidden tab gets no rendering
opportunities, so the browser holds those callbacks instead of merely
throttling them: the socket stayed open, messages kept arriving, and
every cache write parked in a pending frame until the tab was shown
again — at which point they all ran at once. The tab-title progress
reads ['printerStatus', id] and nothing else, so it simply froze.

The frames came in with the print-completion freeze fix, where the
load-bearing part was the coalescing (100ms throttle, 3s debounce,
500ms stagger). That is untouched; the frames only deferred each write
by ~16ms and are gone. Not made visibility-aware on purpose — a frame
scheduled just before hiding would fire after the writes that took the
hidden path and clobber newer status with older.

The six rAF stubs in the tests ran frames synchronously, which is why
nothing caught this. Replaced with coverage that stubs rAF to never
fire, as a hidden tab does.
2026-08-04 07:56:27 +02:00
maziggy 8dde48587e Move the bug-report trigger out of the contended corner (#2750)
The floating disc is pinned bottom-right, which is where most controls
live — it covered ~83% of the Profiles scroll-to-top button at the same
z-index, and being viewport-fixed it also sits on card action buttons
that scroll under it. Below the sidebar-compact breakpoint the trigger
moves into the top bar; at 1144px and up nothing changes.

Not a hide switch: the bubble is the only entry to the report form, and
that form runs the printer diagnostic, the log scan and the debug
capture. Hiding it yields reports with nothing attached.

The panel stays at the Layout root — the header is a fixed z-40
stacking context and would bury a nested z-50 panel under every modal.
Also fixes the panel hanging 16px off-screen on phones: w-full resolves
against the viewport, so right-4 pushed its left edge negative.
2026-08-03 15:44:10 +02:00
maziggy dbf674561c Sort the inventory by colour, not colour name (#2729)
The Color column was missing from the page's sort-extractor map, so its
header ignored clicks. Sorts by family first — rainbow, then browns,
then neutrals light to dark — with the hue sort running inside each.

The issue asked for a straight hue/saturation/lightness sort. Measured
against a real 30-spool inventory that puts Titan Gray (hue 210, sat
0.04) among the blues and a warm grey next to the reds, and splits the
oranges around brown. Neutrals order by lightness because their hue is
noise.

Families come from the classifier that already names colours missing
from the catalog, so the Color and Color Name columns cannot disagree.
2026-08-03 15:24:37 +02:00
maziggy 689f5276e4 Show the compose directory in the Docker update command (#2664)
The printed command only works from the directory holding the compose
file, which is the thing the user came to the page not knowing. Adds a
copy button, a saved Compose directory setting, BAMBUDDY_COMPOSE_DIR,
and best-effort detection from a bind mount's host path.

Compose records the directory on every container it creates, but reading
that label needs the Docker socket mounted in — root-equivalent access
for a convenience string. The mountinfo guess is a prefill only: its root
field is relative to the mounted device, so a compose dir on its own
mount loses that prefix, and nothing in the container can detect it.

The field is restricted to path characters. It is the one setting whose
purpose is to be pasted into a root shell, so "/opt/bambuddy; rm -rf /"
would otherwise render as a plausible update command.
2026-08-03 15:12:02 +02:00
maziggy a08d3e62f3 Show the Print Log's per-run cost and energy, and let users pick columns (#2636)
The list and update endpoints serialised field by field and never named
cost / energy_kwh / energy_cost, so values Bambuddy had been recording
all along went out as nulls. Both now validate from the ORM row, which
removes the chance to omit a field rather than patching the three that
were missing.

Adds a Filament Used column plus a Columns picker for Cost, Energy,
Energy Cost and Finished, persisted per browser.

Also fixes the log view being unreachable with zero archives: the empty
state ran before the view check, hiding a log that outlives the archives
it refers to.

---

Sort the Print Log by any column (#2636)

Adds sort_by / sort_dir to the print-log endpoint, driven by clickable
column headers. Server-side because paging is: ordering the rows the
client holds would sort one page rather than the log.

Empty values are held last in both directions — Postgres sorts NULLs
high and SQLite low, so the same click would otherwise open on blanks
on one backend and values on the other. id DESC breaks ties so paging
through a low-cardinality sort can't repeat or skip a row.
2026-08-03 14:53:21 +02:00
maziggy e95c42c021 Add auto-orient and auto-arrange to server-side slicing (#2548)
Both are per-slice checkboxes, off by default, forwarded as the sidecar's
orient / arrange form fields. An unticked box is sent by omission: the
sidecar treats any present value as truthy, so a literal "false" would
have arranged every slice.

Arrange unions with the #1493 cross-class decision rather than replacing
it, and the per-plate slice-all loop is now keyed on the arrange flag
itself — the project-wide collapse belongs to --arrange, not to the
cross-class case. The loop also covers the embedded-settings path, whose
crash-retry is suppressed there since a single --slice 0 retry would
return one consolidated plate.
2026-08-03 14:10:02 +02:00
maziggy 78fc5aab75 Updated CHANGELOG 2026-08-03 12:41:38 +02:00
maziggy ea63355fde Fix cross-model queue items being misrepresented and editable into a broken state (#671)
The edit dialog offered a printer picker and target-model dropdown for
an item with alternatives. Saving left a row with variants AND a
printer_id, and the scheduler's fixed-printer branch wins that race, so
it dispatched a row whose library_file_id is still null and failed in
the upload. PATCH now refuses printer/model changes on such an item —
comparing against the current value, since the dialog re-sends
target_model unchanged — and the route eager-loads variants, without
which the guard could not see them and every PATCH response dropped the
alternatives from its payload.

Names come from a shared helper now. A cross-model item holds neither
archive_id nor library_file_id until dispatch, so five separate inlined
fallbacks all rendered "File #null"; they now read "x1c.gcode.3mf +1
more".

The queue also grouped these under "Any H2D" — the first candidate
mirrored onto the row — filing a job under a printer it might never run
on. It groups as "Any H2D / X1C", matching the row beneath it.
2026-08-03 12:33:19 +02:00
maziggy ef7c1b21f1 Add cross-model print alternatives to the File Manager and print modal (#671, #2570)
Selecting several sliced files and pressing Print now creates one queue
item carrying all of them, instead of hiding the Print button the moment
a second file is selected. The printer picker is replaced by the ordered
candidate list, since choosing these files is already the answer to
"which printer" and the only question left is which is preferred.

Per-candidate configuration is the plate only. Model-based assignment
sends no AMS mapping — the printer is unknown until dispatch, where the
scheduler derives it — so a per-candidate mapping editor would collect
choices it then discards. Filament overrides stay shared: "this job
needs PETG" holds for every slice of the same job.

Adds Group as versions for durable grouping, a versions badge counting
the whole group rather than the rows on screen, and a queue card label
naming every model a pending item is waiting on.
2026-08-03 11:33:39 +02:00
maziggy a9b57ccd3c Add variant-group endpoints and cross-model queue creation (#671, #2570)
Adds /library/variant-groups for declaring that several sliced files are
the same job for different printers, and a variants payload on queue
creation that turns such a set into one queue item with a candidate per
file.

The candidate set is validated as a set: one file per printer model, each
file sliced for the model it is offered as, and at least one model with
an active printer. A cross-model item deliberately holds no file of its
own, because print_queue.library_file_id is ON DELETE CASCADE and would
destroy the whole job when a single alternative is deleted.

Fixes internal printer-model codes never being resolved on queue create
and update: normalize_printer_model returns unknown input unchanged, so
the or-chain never reached the code map and a "C13" target matched no
printer and waited forever.

Skips candidates whose file is trashed or missing. Library deletes are
soft, and SQLite runs with PRAGMA foreign_keys off, so neither case is
covered by the schema; the hard-delete paths now also drop the rows.

Adds library_files.variant_target_model so a user can say which printer
a file without slicer metadata is for, kept out of file_metadata so the
assertion is never mistaken for parsed data.
2026-08-03 11:10:11 +02:00
maziggy 752e345d1a Add cross-model variant resolution to the queue scheduler (#671)
Adds print_queue_variants: the candidate files a queue item may run, each
with its own model, plate, AMS mapping and nozzle mapping. The scheduler
walks them in priority order and takes the first whose model has an idle
printer, then folds that candidate onto the queue row before the
selection commit — so upload, archive creation, print history and reprint
keep seeing an ordinary single-file item.

Candidates are ordered least-attempted first, so a printer that accepts
the file and never starts hands the job to the alternative on the next
lap instead of spending the item's whole retry budget on the machine
that is wedged. The item-level DISPATCH_MAX_ATTEMPTS bound is unchanged.

An item whose candidate files have all been deleted is held pending with
an actionable reason rather than failing deep in the upload, and waiting
notifications name the job and every model it is waiting on.
2026-08-03 10:54:45 +02:00
maziggy da07c5884b Add variant-group data model for cross-model queue alternatives (#671)
Adds file_variant_groups plus variant_group_id / variant_position on
library_files, so a set of files that are the same job sliced for
different printers can be resolved to whichever printer frees up first.

Backfills groups from the sliced_from_library_file_id provenance that
slice_and_persist and the pipeline runner have been writing into
file_metadata since they shipped, and which nothing has ever read.
Only sources with two or more children carrying distinct
sliced_for_model values are grouped: a single candidate is not a
choice, and two slices for the same printer give the resolver no basis
to prefer one.
2026-08-03 10:17:28 +02:00
maziggy b2f84cd37f Updated CHANGELOG 2026-08-03 09:07:03 +02:00
maziggy dd541b20a8 Document prefer_filename_for_name in the OpenAPI schema
The param was a bare bool, so /docs showed an undocumented boolean on
both upload routes. #2609 is about external integrations, and the
interactive docs are where those callers look — a docstring only reaches
someone reading the source. Wraps both in Query(False, description=...),
matching how this file documents its other query params.

Also records why these two routes take the flag per-request while the FTP
review flow and virtual-printer dispatch derive it from the VP-scoped
virtual_printer_archive_name_source setting, and drops the db_session
fixture the four new tests requested but never used.
2026-08-03 09:04:29 +02:00
maziggy b04664c64a Raise the chamber-temperature ceiling from 60 to 65 C
Every field that takes a chamber target stopped at 60: the per-filament
chamber map and per-print override in Preheat & Heat Soak, the chamber
quick-select presets, and the printer-card chamber control. 60 is the
X1E's ceiling and the X1E was the only heated-chamber model when that
limit was written; the H2 series and X2D heat to 65, so the top of their
range was unreachable.

The ceiling now lives in one constant per side (MAX_CHAMBER_TEMP_C in
backend/app/utils/printer_models.py and frontend/src/utils/printer.ts)
rather than as a literal at each call site. X1E firmware clamps a higher
request to its own maximum, so a shared ceiling is safe.

Also fixes a live bug at PrintersPage.tsx:7985: parsePresetTriple was
bounded to 60 there, and it rejects the whole triple on any out-of-range
entry, so a saved 65 preset would have silently reverted the printer
card to the defaults while Settings still showed 65.
2026-08-03 08:52:23 +02:00
maziggy ad375f6ca7 Housekeeping 2026-08-02 12:29:32 +02:00
maziggy e52b73e21f Stop the Bambu Cloud TOTP tests reaching the network
verify_totp fetches a CSRF token from the bambulab.com web origin before
posting the code (#2696) and returns early when it cannot get one. These
tests patch only post, so the pre-flight GET went out for real: it succeeded
wherever bambulab.com was reachable and returned a tokenless 403 on a CI
runner, where six tests then asserted on a post that never happened.

Stub the handshake for the module. It is covered end to end, no-token path
included, in tests/unit/test_cloud_totp_csrf.py.
2026-08-02 11:45:53 +02:00
maziggy bbbb9d35c7 Bound the scheme repetition in the log credential-redaction pattern. As an
unbounded repetition the match was quadratic in the subject length: on a run
of scheme-legal characters the engine restarted at every offset and consumed
to the end before failing to find "://". ffmpeg echoes the configured camera
URL into its stderr and the whole blob reaches the pattern before any
truncation, so the subject length is attacker-influenced.
2026-08-02 11:17:49 +02:00
maziggy bf424493ba Suppress Bandit B104 false positive in the SSRF guard tests
The parametrize list feeds "0.0.0.0" to TasmotaService._validate_ip and
asserts it is refused. B104 matches the literal wherever it occurs and
cannot distinguish a rejection fixture from a bind address.

Split the list across lines so the token carries its own nosec with the
reason; the single-line form was 117 chars against a 120 limit.
2026-08-02 11:09:32 +02:00
maziggy 9e18ff7fa0 Updated CHANGELOG 2026-08-02 10:56:51 +02:00
maziggy f5580914af Updated CHANGELOG 2026-08-02 10:52:21 +02:00
maziggy f52cebd319 yyUpdated CHANGELOG 2026-08-02 10:50:01 +02:00
maziggy 77135aaf8f Fix unawaited coroutine warning in printer offline notification tests
on_printer_status_change builds reconcile_stale_active_prints(...) as a
call argument, so the coroutine is constructed even when the spawn helper
is mocked out. A bare MagicMock retained it in call_args and it finalised
unawaited during a later test's GC, surfacing as a
PytestUnraisableExceptionWarning attributed to test_printer_sensor_history.

Patch spawn_background_task with a side_effect that closes the coroutine,
and correct the _state() docstring, which claimed state="IDLE" kept the
reconcile-edge branch quiescent when it does the opposite.
2026-08-02 10:23:09 +02:00
maziggy dbd2fd19d4 Post work PR #2740 2026-08-02 08:32:52 +02:00
maziggy 30daed2756 Fix per-job queue ETA showing for jobs that cannot start now
The scheduler only writes waiting_reason on the model-based assignment
path, so a job pinned to a specific printer sits behind a running print
with no marker at all. Every such job rendered an identical "starts now"
ETA that was wrong by the length of everything ahead of it.

Decide eligibility on the page instead: an item gets an ETA only when its
printer is idle and it is the item the scheduler would dispatch next,
following the same ordering the scheduler uses. Staged and future-
scheduled items do not block the item behind them, matching the
scheduler, and items conditional on a previous print are excluded.

The value also froze at first render, since react-query's structural
sharing keeps the queue reference stable and nothing re-rendered the row.
formatETA now accepts a base instant and the page drives it from a 30s
clock shared by every visible row.

Retire the borrowed printers.estimatedCompletion tooltip for a queue key
that says what the number means, translated into all 13 locales.
2026-08-02 08:28:26 +02:00
maziggy d36632db0f Updated README 2026-08-01 12:53:58 +02:00
maziggy 9259763179 Updated README 2026-08-01 12:53:37 +02:00
maziggy aef4f3a3e9 fix(oidc): strip the required BAMBUDDY_OIDC_* values and register the local-login bypass
A Kubernetes Secret written as a block scalar carries a trailing newline, and
the schema bounds the four required variables by max_length only, so an
unstripped issuer_url was stored and enabled and then raised httpx.InvalidURL
on the first click of the SSO button -- the authorize-time failure the
all-or-nothing rule exists to prevent. Whitespace-only values got through the
same way, contradicting the reader's own "an empty required var counts as
unset". The optional variables have always treated blank as unset; the
required ones now do too.

Also registers BAMBUDDY_LOCAL_LOGIN (#1589) in the typo guard, which logged
"possible typo" for it on every boot while listing every BAMBUDDY_OIDC_*
variable as legitimate.
2026-08-01 11:37:01 +02:00
maziggy 43cb216ae9 fix(settings): stop the Settings page reverting changes made elsewhere (issue #2716)
While the Settings page was mounted it held its own copy of every
setting and synced it from the server exactly once, on first load
(:887-900). A debounced effect then diffed the live ['settings']
cache against that copy and PUT all 77 keys it manages on any
difference, with no way to tell a user edit from a value that had
changed on the server. Anything written server-side while the page
sat open was silently reverted ~500ms later (#2716, reporter
@jmoore-skild).

No interaction was needed to trigger it. The query inherits a 60s
staleTime and react-query's default refetchOnWindowFocus, and ~30
other observers share the key, so a window refocus or a refetch from
any of them moved the cache and the page wrote its page-load snapshot
back over all 77 keys -- showing "Settings saved" while doing it.

The page now tracks the last server snapshot it reconciled with. A
field still equal to that baseline has not been touched since, so a
newer server value is adopted; a field the user has edited keeps
their value and is saved over the top, so the newer of the two writes
wins either way. Typing into a text field while a refetch lands stays
safe, which is what the previous behaviour was protecting -- an
in-progress edit is by definition different from the baseline.

The baseline is seeded from the raw server row rather than from the
copy the page patches a browser-detected external_url into, so that
detection still reads as a local change and is still persisted.

The payload builder and the comparison key lists are unchanged. The
diff simply measures against the baseline instead of the live cache,
so no field can silently stop saving.

Removing the adoption step was verified to reintroduce the revert,
and removing the post-save baseline advance to reintroduce a resend
loop; both are covered by frontend tests asserting on the request
bodies rather than on rendered values.
2026-08-01 10:55:08 +02:00
maziggy 18938a10ee fix(kprofiles): stop reporting rejected K-profile writes as saved
Saving a K-profile was fire-and-forget. set_kprofiles_batch published
and returned True, and the printer's extrusion_cali_set answer was
logged at DEBUG and dropped, so a write the printer refused was
reported to the user as saved (#2718, reporter @jmoore-skild).

The reason it could not simply be gated on: the answer itself was
wrong. Single-nozzle firmware returned result:"fail" with
reason:"invalid tray_id" on writes that demonstrably applied.
Measured against an X1C and an H2D over MQTT, the cause is the
tray_id:-1 Bambuddy itself put in the payload. Sending three
otherwise identical writes isolated it: tray_id:-1 fails, tray_id:0
succeeds, and cali_idx:-1 is accepted either way, so only that one
field is at fault. The H2D ignores the value entirely; the X1C
validates it, complains, and applies the write anyway. BambuStudio
always sends a real tray_id and defaults it to 0 for a manually
entered profile.

With tray_id:0 the acknowledgement is honest, and the printer echoes
back the sequence_id we sent -- confirmed for extrusion_cali_get,
_set and _del on both printer classes -- so it can be matched to the
write that caused it. Writes now return their sequence_id and the
routes await the verdict, turning a real failure into an error that
carries the printer's own reason. A printer that stays silent is
still treated as success: no answer is not evidence of refusal, and
firmware that never answers must not turn every save into an error.

Raises the ack to INFO. It sat at DEBUG, so the one line that
explains a failed save was absent from every support bundle -- the
same reasoning that put ams_filament_drying at INFO for #1447.

Also fixes extrusion_cali_set building its payload from
str(self._sequence_id) without incrementing first, reusing the
previous command's id. Harmless while nothing correlated on it,
fatal now that the write path does.

Adds supports_nozzle_flow_type() for the Standard / High Flow choice,
which the K-Profiles UI previously showed as "Not reported by
printer" -- not a value anyone can save. Most printers omit the
nozzle identity from their calibration table entirely, and the slicer
treats that as Standard rather than unknown; Bambuddy now does the
same and keeps the choice editable. The field is hidden only where
the model ships a single nozzle variant, using the slicer's own rule
(len(nozzle_volume) // len(nozzle_diameter) > 1 over the machine
preset) evaluated across every bundled Bambu profile. That puts only
A1, A1 Mini and A2L on the hidden side -- it is not the single-
versus-dual-nozzle split, since P1P, P1S, P2S, X1, X1C, X1E and H2S
are all single-nozzle and all carry two variants. Editing a profile
also no longer writes back an empty nozzle_id.

Wiki records that on printers which omit the field the chosen flow
type is discarded by the firmware and reads back as Standard, in
Bambu Studio as well, so it does not get filed as a bug again.
2026-08-01 10:35:22 +02:00
maziggy af282b3527 fix(kprofiles): populate the filament picker from all preset tiers (issue #2719)
Add K-Profile built its Filament dropdown from the profiles already on
the printer, so on a printer with none the field was empty, required
and unsatisfiable (#2719, reporter @jmoore-skild). The modal's own
hint described the dead end: create the profile in Bambu Studio first.

The dropdown now uses the app-wide lookup order -- local imported,
Orca Cloud, Bambu Cloud, hardcoded built-in table -- same as the AMS
slot picker and the SliceModal tier groups. The built-in table is
compiled into the backend, so the list can never be empty: a new
printer with no cloud account and nothing imported still gets a first
profile.

Not fixed the way the report suggested. Seeding from
/printers/available-filaments would have offered only what happens to
be in an AMS right now, which on the reported printer is nothing; its
tray_info_idx is empty or a cloud user preset rather than a filament
id; it aggregates across every printer of the same model; and it is
gated on QUEUE_CREATE, which the K-Profiles page does not hold.

The printer indexes its calibration table by filament_id, so the
picked preset is reduced to one before anything is sent. Built-in
entries and Bambu official cloud presets carry one; a cloud user
preset needs its detail fetched (never base_id -- that collapses a
custom preset onto its inherited generic, #1053); imported and Orca
presets have no Bambu id at all and take the closest generic for
their material, via the same table the AMS slot configure flow uses
so the two agree. A filament that resolves to nothing is refused with
a named error rather than written under a wrong id.

Collapses duplicates from two separate causes. A cloud account
carries one copy of each filament per printer model, and with the
"@BBL <model>" suffix stripped for display those rows are
indistinguishable -- deduped within each tier by resolved filament id,
by display name for user presets that have none. Cloud setting_ids
also carry a "_NN" variant suffix, so the built-in tier's
already-covered check never matched and listed the same filament
again; the bare id is now recorded alongside.

Groups the options by source with an optgroup per tier, styled in
index.css: browsers render optgroup labels small, grey and italic,
which buries the one thing distinguishing a "Bambu PLA Basic" you
imported from the one the built-in table ships.

Drops the second getKProfiles(printer, "0.4") query that existed only
to seed the old dropdown. It ran concurrently with the main fetch
whenever a non-0.4mm nozzle was selected -- the two-requests-in-flight
case that made K-profile fetches time out.

---

fix(ui): cancel a dialog's deferred close when it unmounts

The AMS slot configure and K-Profile dialogs hold a success state
briefly and then close themselves -- 1.5s to 4s after the command
goes out, so the printer has time to process it before the list
refetches. Each did that with a bare setTimeout closing over setState
and the parent's onClose, and nothing cancelled it.

The timer therefore ran whether or not the dialog was still there.
Dismissing it inside that window, or the printer card re-rendering
underneath it, left a pending close that fired later and dismissed
whatever dialog was open by then. It also threw outright when the
surrounding environment was gone first: a test tearing down its DOM
before the 1.5s elapsed produced "ReferenceError: window is not
defined" out of react-dom's resolveUpdatePriority, reported as an
unhandled error against a suite that otherwise passed.

Routes all five through a useCancellableTimeout hook -- two in
ConfigureAmsSlotModal, three in KProfileModal, the latter with the
longest windows and so the widest exposure. Scheduling replaces any
pending timer and unmounting clears it.
2026-08-01 09:36:42 +02:00
maziggy a35ba8fa5f fix(kprofiles): read the nozzle diameter the printer actually sent (issue #1748)
Every K-profile came back as 0.4mm on printers running any other
nozzle (#1748, reporters @Liquidmasl and @jmoore-skild). The printer
puts nozzle_diameter on the extrusion_cali_get envelope only; the
per-filament entries carry setting_id, filament_id, name, k_value,
n_coef and cali_idx, and nothing else. The parser read the field per
entry with a hardcoded "0.4" fallback, so the fallback fired on every
profile of every response. The envelope value was already in scope,
read into response_nozzle and used only to match the request.

This never reproduced on H2D because that firmware does include the
field per entry. Both construction sites are in the same handler, so
the code path is shared; what differs is the payload, and every
single-nozzle model omits it.

The display was the least of it. Editing is delete-and-re-add on
single-nozzle printers, and the dialog rebuilt nozzle_id and
nozzle_diameter from its own greyed-out selects, so saving an
untouched 0.6mm profile rewrote it on the printer as HH00-0.4.
Deleting aimed extrusion_cali_del at the wrong nozzle the same way.
Both now pass through what the printer reported. The cali_idx cascade
in inventory.py, spoolman_inventory.py and spoolman.py matches on
nozzle_diameter, so on a 0.6 or 0.8 nozzle it never found the
printer-side entry and the assignment silently failed to stick --
that is the "cannot auto-map a K-profile" half of the report, fixed
at the source without touching those three call sites.

nozzle_id has no source in the payload at all, and state.nozzles
carries material (hardened_steel), not flow, so it cannot honestly
produce HH/HS. Rather than keep inventing one, the UI now says the
printer did not report it: the card shows the diameter alone, the
dialog shows "Not reported by printer", and the High Flow / Standard
filter is hidden instead of being offered as a control that can only
ever empty the list. Import stops stamping HH00 on profiles whose
source reported none.

Also correlates K-profile requests by sequence_id. Responses were
matched by nozzle diameter through a single shared expectation slot,
so a second request overwrote the first's and the first's valid
answer was discarded as a mismatch -- the "Failed to get K-profiles
after 3 attempts" in the same logs, with the printer having answered
correctly both times. Pending state is now one entry per request,
keyed by the id we already send, with the nozzle match kept as a
fallback for firmware that does not echo it back.

Fixes the flow-type select naming a new profile with the opposite
label, which contradicted the identical expression 44 lines above it.
2026-08-01 08:49:44 +02:00
maziggy aa07415270 Updated CONTRIBUTING.md 2026-08-01 08:28:48 +02:00
maziggy b8225d9e9f Housekeeping 2026-08-01 08:28:20 +02:00
maziggy a94f1ef4ff Updated CONTRIBUTING.md 2026-08-01 08:28:06 +02:00
maziggy 455a9e4ba7 fix(backup): collect cloud profiles from every connected account (#2717)
Enabling Cloud Profiles for a Git backup produced nothing, and said it had
worked. Two independent faults, either one sufficient.

The collector looked for a "setting" list. The Bambu Cloud listing endpoint
is keyed by preset type instead, each key holding private and public arrays,
so the loop body never executed once — and the entries carry no type of
their own either, which routes/cloud.py already knew: it takes the type from
the outer key and maps Bambu's "print" to process. Two bugs on one line.

It also asked build_authenticated_cloud for the credential store used when
authentication is disabled. With auth on, tokens live on User rows, so the
collector returned at "Cloud not authenticated" before ever reaching the bad
key. Every multi-user install was collecting from zero accounts.

Neither failure surfaced. backup_metadata.json recorded the configured flag
rather than the outcome, so it claimed cloud_profiles: true on runs that
wrote nothing, and the log read "Collected cloud profiles: 0 filament, 0
printer, 0 process" at INFO — which is exactly what a successful backup of
an empty account looks like.

Cloud profiles now come from every connected account across both clouds. The
toggle predates Orca Cloud entirely, and Orca has the same three preset
types, so both are collected and grouped the same way:

    cloud_profiles/bambu/user-3/{filament,printer,process}.json
    cloud_profiles/orca/user-3/{filament,printer,process}.json

Accounts are keyed by Bambuddy user id, "global" when auth is off. Never by
email: a backup repository can be public, and the Bambu listing's user_id is
dropped for the same reason. Both credential stores are read on every run,
because a Settings row survives someone enabling auth later and dropping it
would silently stop backing that account up.

Bambu costs one get_setting_detail per private preset. The listing is
metadata only, and without base_id and setting the backup is a list of names
that create_setting cannot rebuild from. Public presets are skipped — Bambu's
bundled catalogue is the same hundreds of entries for everyone, always
re-downloadable, not recreatable under your account, and would rewrite the
repository on every run. Orca needs no second call; its sync-pull carries
each profile's content inline. Where the Orca route drops a profile whose
content.type it cannot map, the backup writes it to other.json instead:
silently omitting a profile because Orca added a type is the same class of
bug as this one.

Failures are contained per account and per preset, and counted rather than
swallowed. A partial backup that looks complete is how this stayed invisible.

The metadata now reports what was collected, per cloud and per account, and a
run that collects nothing while the category is enabled warns with the reason
instead of an INFO line that reads like success.

The checkbox gated on the viewer's own Bambu sign-in, which is not the same
question as whether there is anything to back up — with auth enabled the
accounts belong to individual users, and an administrator who never signed
in personally saw the category disabled with plenty in scope. It now gates
on the total across both clouds and shows the counts. That comes from its
own endpoint rather than a field on /config, since /config answers null
until the first save and would disable the toggle during the very setup it
belongs to. Counts only, never identities.

One deliberate restraint. _build_authenticated_service clears stored
credentials when a refresh is rejected, which is right for a route — the
user is on the page and can pair again — and wrong for a scheduled job.
Orca reports every rejection with one composite reason ("unknown, expired,
revoked, or already used"), so a genuine revocation cannot be told apart
from a lost token-rotation race, and acting destructively on a signal that
cannot be disambiguated is the #2562 mistake in a different cloud. It also
gains nothing: the Profiles route hits the same failure and clears it then,
with the user present. Background callers now pass clear_on_auth_failure=
False and skip the account. A successful refresh is still persisted either
way — by that point the old token is consumed, so dropping the new pair
would break a working pairing for real.

Restore is not part of this. Nothing reads cloud_profiles/* yet; the format
carries base_id/setting for Bambu and content for Orca so that it can.
2026-07-31 16:59:33 +02:00
maziggy 82656c8760 Updated BACKERS 2026-07-31 16:22:55 +02:00
maziggy 3c49990387 Updated BACKERS 2026-07-31 16:22:22 +02:00
maziggy 4f2c073a34 fix(vp): gate the slicer's AMS pick behind the toggle and scope its badges (#2700)
Round-3 review of the "Save AMS mapping" PR.

The queue item's ams_mapping was set unconditionally, on the reasoning that
honouring the slicer's own pick is a correctness fix rather than a feature.
It is both. Storing a resolved mapping makes _ensure_ams_mapping return
early, so _compute_ams_mapping_for_printer never runs — and that function is
where prefer_lowest_filament lives, along with the AMS-filament-backup gate
that qualifies it (#1766), the inventory-remain overrides, and the per-slot
force-colour overrides. Every existing queue-mode VP pointed at a printer
would have quietly lost all of it on upgrade, without a setting to turn it
back on.

So save_ams_mapping now gates the queue item too, not just the archive
persistence. Off is exactly the old behaviour. The correctness case the PR
was written for — two spools of the same red PLA, and the slot the user
picked in the slicer thrown away — is still fixed, for anyone who asks for
it.

Force color match wins over it when both are on. Its only effect on a
fixed-printer item is the filament_overrides written onto the queue item,
and those are read inside the function a stored mapping skips, so the two
toggles sitting next to each other on the same card silently cancelled. The
dispatch now matches strictly, as asked, while the slicer's pick is still
saved onto the archive — that is what the toggle's name promises, and a
later reprint is a separate decision from this print. The queue-add fallback
applies the same rule to a request that carries force-colour overrides.

A mapping shorter than a plate's highest slot id cannot address that plate's
own slots, and _ensure_ams_mapping would have kept it anyway, since it only
rejects an all-unresolved one. Each plate now checks the length it needs and
falls back to a computed mapping if the array does not reach. Bambu Studio
sends a file-global array, so this normally never fires; it also means a
multi-plate Send All degrades safely if that ever stops being true.

The badges claimed more than they delivered. Both rendered whenever a saved
mapping existed, ignoring which printer it belonged to, while the tooltips
promised the reprint would reuse those exact spools — true only on the
printer the trays were resolved against. The queue row's flag is now
computed against that row's own printer, which is precisely when dispatch
reuses the mapping, and the archive card names the printer instead of
implying any of them will do. It hides itself when that printer no longer
exists. Retranslated in all 13 locales.

Frontend tests, which the PR had none of. The printer-scoping rule is now a
pure function rather than an inline expression, covered for the mismatched
printer, the no-printer-selected case that would otherwise compare undefined
against undefined, and malformed extra_data. The toggle's undo bookkeeping
is covered for unresolved slots, short mappings, and hand-made picks —
preserved when the toggle never wrote that slot, replaced when it did, which
is behaviour worth pinning either way.

Also reverts all three queue-mode switches when a save fails, not just the
new one; without it the card shows a setting the server rejected.
2026-07-31 16:16:39 +02:00
maziggy c457cf54bf feat(support): record process memory, threads and children in bundles (#2734)
A bundle described everything except the process it runs in. So a report of
memory climbing over days until the OOM killer fires arrives with no way to
act on it: the numbers that name the mechanism only exist while it is
happening, and by the time anyone asks, the container has been restarted.

The new `process` section carries what actually separates the candidates.
Resident against virtual memory: 650MB RSS with 12.9GB VMS is address
space — thread stacks or allocator arenas — not a heap full of live data,
and that reading is the opposite of the one the reporter drew from the same
figures. Thread count and child-process count then split those two apart,
and a census of live objects by type names what a growing heap is filling
up with. Open files, sockets and uptime round it out.

Three constraints worth keeping:

The heap census is skipped above 2GB. gc.get_objects() materialises every
tracked object, so it costs most on exactly the process that can least
afford it — a bundle generated to diagnose runaway memory must not be the
allocation that tips the host over. Everything else is still collected, and
the skip is recorded with its reason rather than silently omitted.

Children are recorded by executable name only. An ffmpeg command line
carries the camera URL, and with it the camera's password.

Collection runs off the event loop and every metric is independently
best-effort. psutil raises on hardened kernels and in restricted
containers, and the bundle is how someone reports a problem in the first
place — it has to be produced even when half the numbers are unavailable.

This does not fix #2734, and nothing here should be read as having found
its cause. The bundle's own evidence contradicts both proposed causes: the
orphan janitor ran 7 times in 26 days over 725 stream-ends and killed no
orphaned ffmpeg, which is not the #776 signature; and the 5 "database is
locked" errors all fall between two OOM kills, making them a symptom of the
memory pressure rather than a source of it.
2026-07-31 15:32:08 +02:00
maziggy ce3e59884a fix(slicer): bound slices by silence, not by total slicing time (#2730)
A heavy MakerWorld model — one Bambu Studio also takes a long time over —
failed after five minutes with "Slicer sidecar unreachable". The sidecar
was reachable the whole time and still slicing when we hung up on it.

SlicerApiService carried a hardcoded 300s timeout, passed to httpx as a
bare float so it covered connect, read, write and pool alike. On a single
long request that is not a health check, it is a cap on how long a model is
allowed to take. And because httpx.ReadTimeout subclasses RequestError,
expiry landed in the same handler as a refused connection and was reported
as an unreachable sidecar — so the reporter went and updated their sidecar
container, which was never the problem.

The information to do better was already being collected. _poll_progress
polls /slice/progress/{id} once a second alongside the blocking POST to
drive the live progress toast, so at minute five Bambuddy had fresh
evidence the slicer was working. It killed the request anyway.

So the read timeout comes off the HTTP call and the poller supervises
instead: the deadline moves forward on every progress update, and only
genuine silence ends the wait. A model that keeps reporting runs to
completion however long it takes. Connect and pool keep short timeouts —
a sidecar that will not accept a connection is unreachable and should
still say so quickly.

Only a *changed* progress payload counts as alive. The sidecar re-serves
its last snapshot on every poll, so counting repeats would leave the
watchdog unable to detect a stall at all.

The window is floored at three poll intervals: liveness can only be
observed as fast as the poller ticks, so anything shorter would expire in
the gap between two polls and fail every slice instantly.

New setting slicer_stall_timeout_minutes (Settings > Workflow > Slicer),
default 15, range 1-240, alongside the sidecar URL and gated on
use_slicer_api like its neighbours. Sidecars too old to report progress
have no liveness signal, so for those the same number bounds total elapsed
time — the old behaviour, configurable and no longer 300s flat. The
message says which case applies and where to change it.

SlicerTimeoutError is its own type and maps to 504, not 502: the sidecar
answered throughout, we stopped waiting. Connection failures keep
SlicerApiUnavailableError. The preview slice path gets the same treatment.
2026-07-31 15:14:01 +02:00
maziggy 284709f850 fix(projects): drop deleted prints from their project, and refresh the view (#2731)
Deleting a print that belonged to a project left it on the project page as
a card with a missing thumbnail, and there was no way to remove it.

Deleting a print is a soft delete by default (#1343): the files go from
disk, the row stays so global Quick Stats keeps counting its filament,
time and cost. Every other consumer filters those rows out. The projects
module filtered none of them — the only deleted_at check in the whole file
was for LibraryFile — so a deleted print kept its project_id and kept
being listed, pointing at a thumbnail that no longer existed. The same
broken previews appeared on the overview cards, and in the timeline, where
the entry links to an archive that no longer opens. Unassigning was
impossible because the only UI that can change a print's project lives on
the Archives page, which correctly hides deleted prints: visible on the
project, unreachable from anywhere.

All eight project-scoped archive queries now filter, counts included. That
last part is a deliberate divergence from #1343, where the whole point of
the soft delete is that the contribution survives: a project is a piece of
work with a definite membership, not a lifetime total, so a project that
lists eleven prints must not claim twelve. The reasoning is recorded at
the constant so nobody later "fixes" it back.

remove_archives_from_project keeps working on hidden rows on purpose — it
is the repair path for links written before this. The BOM print_name
lookups are left alone; naming a since-deleted print is still correct.

Two more consumers had the same gap. The CSV/Excel export handed back rows
the interface says are gone — filtered at the base query, since the export
is the list you are looking at saved to a file. Per-project failure
analysis measured a failure rate against prints deleted from the project,
and disagreed with the project's own numbers; only the project-scoped
branch filters, global analysis still counts every run including orphans
as #1390 established.

Finally, the project page needed a manual reload to catch up. staleTime is
60s and the delete mutations invalidated only ['archives'], so a project
visited within the minute served its cached copy, print still there. The
project-assign mutations had the mirror-image bug: ['projects'] refreshed
the overview cards but never ['project', id]. Both now go through one
shared helper covering every project-derived key, as bare prefixes so all
cached project ids are matched.
2026-07-31 14:50:25 +02:00
maziggy 3abab1fd45 fix(printers): recover MQTT sessions that stopped reconnecting (#2732)
The reporter's printer lost its session to a keep-alive timeout at 02:19
and did not come back until 11:24 — nine hours offline, with the web UI
open throughout.

check_staleness() was never going to catch it. Its first line is
`if self.state.connected and self.is_stale()`, so it only ever handles the
half-broken session that is still connected but has gone quiet. This
client had connected=False from 02:19:42 (the offline notification fired a
minute later), so every call returned immediately, and paho's own retry was
the only thing left watching. When that stopped making progress nothing
noticed.

Adds a sweep every 60s that rebuilds a client when all four hold: it is
disconnected, it had a working session before, it has been silent for five
minutes, and its MQTT port still answers. The port check is what keeps this
from becoming a nuisance — a switched-off printer is left to paho, so a
farm powering down overnight causes no client churn and no log spam. The
five-minute grace sits well past the 60s stale timeout and the 30s max
reconnect backoff, so a session recovering on its own is never interrupted.

The rebuild goes through force_reconnect_stale_session from async context,
which takes the hard-reset path: fresh client_id and paho's QoS 1 queue
dropped, so a project_file left unacked on the dead session cannot replay
into the new one and trip 0500_4003 (#1136). Rate-limited per printer,
cooldown cleared when the printer returns, and the sweep continues past a
client that throws rather than abandoning the rest of the farm. The log
line names how long the printer was gone and the last connect error, so a
session that dies repeatedly leaves a trail.

check_port gains a public alias in printer_diagnostic rather than having
the watchdog reach for the private name.

Also corrects the Developer Mode path added in the previous commit: the
wiki documents it under Settings > Network, not Settings > General. The
menu path is dropped from the translated string entirely, since it varies
by model and firmware and the wiki carries the detail.
2026-07-31 14:28:17 +02:00
maziggy 5e2b7b53e6 fix(printers): surface the printer's own "command verification failed"
A P1S on firmware 01.10.00.00 rejected every control command and said so:
HMS 0500-0500-0001-0007, "MQTT command verification failed". Bambuddy
received that, dropped it, and reported a healthy printer instead.

The frontend filtered it out. This code's meaning lives in attr's low half
(0500) and code's high half (0001), both of which the MMMM_EEEE short form
discards, so it collapsed to "0500_0007" — no catalog entry, no firmware
actions, and filterKnownHMSErrors drops uncatalogued action-less errors.
Catalog lookups now try full_code first, in both the description and the
filter, and errors matched that way display the four-group code the
printer's own screen shows. The remedy line is ours, not Bambu's: their
wiki says to update Studio or Handy, which does not apply to a print sent
from Bambuddy.

The developer-mode probe made it worse. It read anything that was not an
explicit refusal as confirmation, and this firmware answers the probe with
an empty result while refusing everything else — so an inference drawn
from a non-answer became "developer_mode: pass" in the support bundle of a
printer that had not accepted a command all day. The probe now has three
outcomes: explicit success enables, explicit verify-failure disables,
anything else stays unknown and the diagnostic reports skip.

The HMS is authoritative over that inference in both directions. It forces
developer_mode False when present, and clears back to unknown when the
printer stops reporting it, so enabling Developer Mode and restarting the
printer is picked up without restarting Bambuddy.

Dispatch no longer treats a refusal as a wedge. The watchdog latches the
HMS across both phases and fails the item on the first attempt naming the
code and the fix, rather than spending three uploads and 270s a lap to
arrive at a message about SD cards. The check runs after the active-state
exit in both phases, so a lingering HMS can never abort a print that is
visibly running.

Also: the "wrong or mis-cased serial number" hint no longer fires in the
moment after a reconnect. _report_messages_since_connect is reset by
_on_connect, so a reconnect landing microseconds before the staleness
check leaves it at 0 for reasons that have nothing to do with the serial —
this reporter's healthy printer was told to go check its serial 1 ms after
reconnecting.
2026-07-31 14:14:58 +02:00
maziggy 11dc612bc4 feat(obico): authenticate to a token-protected ML API (#2733)
Obico's ml_api container takes an optional ML_API_TOKEN environment variable.
With it set, ml_api/auth.py answers a bare 401 to any request whose
Authorization header isn't "Bearer <token>"; with it unset it ignores the
header entirely. Bambuddy never sent one, so pointing it at a protected server
meant deleting the token there — which the reporter had set for their Home
Assistant integration and did not want to undo.

Settings -> Failure Detection gains an ML API Token field. When it is empty no
header is sent, so an unconfigured install's request stays byte-identical to
what shipped before the setting existed.

This failed in the worst possible way, and that is the more important half of
the change. Obico decorates /p/ with token_required but leaves /hc/ open. Test
Connection pinged /hc/, so it reported success against a server that was
rejecting every real detection call, the settings looked right, and detection
silently never ran. The only symptom was a generic "ML API call failed" buried
in the status card.

So the test now proves what it claims. After health passes it probes GET /p/
with no img parameter: the auth decorator runs before the handler, so 401 means
the token was rejected and 422 ("Invalid request params") means it was
accepted. No inference work is done either way. A probe that itself errors
reports the token as unknown rather than as working — the UI says it could not
be checked instead of claiming success.

The detection loop checks for 401 before raise_for_status, so a rejected token
is reported as a rejected token, naming the setting and the environment
variable, instead of surfacing "401 Unauthorized" with no hint of what to do.
The message never contains the token; a test pins that.

The setting name carries "token", so the support bundle's keyword redactor
masks it with no new rule. Resolving "field omitted" to the saved token is the
route's job, keeping test_connection a pure outbound call with no database
access.

Second fix, same issue: support bundles misreported which printers Obico
watches. The bundle split obico_enabled_printers on commas and read an empty
value as "no printers". The settings UI writes a JSON array, and empty means
*all* printers — the default — so a working Obico setup showed obico_enabled
false against every printer in its own bundle. That is the reporter's bundle
exactly, and it points anyone reading it at the wrong subsystem. The bundle now
parses the setting the way ObicoDetectionService does, keeps a comma fallback
for any install that stored the legacy shape, and factors in the global switch.
2026-07-31 13:39:17 +02:00
maziggy 6844aa292f fix(ams): offer every K profile the printer holds for a generic filament preset (#2710)
The reporter's A1 mini has nine Flow Dynamics calibrations, all of them saved
under Generic PLA and named after the spool's colour — "Dark Brown", "Glow",
"Marble". Bambu Studio lists all nine for that slot. Configure AMS Slot offered
one: the profile already bound to the slot. After a slot reset it offered none,
leaving the slicer as the only way to assign a K value.

Two independent faults, both tripped by picking a built-in generic preset.

The filament-id match discarded Bambu's generic GFx99 ids as too broad. But the
comparison already requires both sides to carry the same id, so that exclusion
could only ever fire when the selected preset was itself the generic one —
precisely the case where the match is right. The printer keeps one calibration
table per filament id, so a slot on Generic PLA should offer everything
calibrated under Generic PLA. Equal ids now match, generic or not.

The name fallback was dead for the same presets: parsePresetName reads the
leading "Generic" in "Generic PLA" as a manufacturer, which put the matcher into
brand-gated mode and demanded the word GENERIC appear in the profile name. No
real profile has it. "Generic" is no longer treated as a brand, so profiles still
match on material when a printer reports no filament_id with its calibrations.

The one profile that did appear came from the #1689 safety net that always
surfaces the slot's active cali_idx — which is also why a reset slot, having no
active profile, showed an empty list.

Neither fix can be complete on its own, because profile names are free text and
nothing ties "Marble" to a material. The picker now also lists every remaining
profile on the printer under "Other K profiles on this printer", so a profile
that exists can always be selected. Applying one from that group needs no new
backend work: configure_ams_slot already realigns the slot's filament context to
the chosen profile's, which is what makes the cali_idx stick.

Options are keyed by name+k_value rather than the bare name, so two profiles
sharing a name are no longer indistinguishable in the select. Both render blocks
carry the change — the modal duplicates the picker for its full-screen variant.

isMatchingCalibration gets the same generic-id rule for the spool form's PA
suggester, with two guards. A new generic-id-to-material table means a PETG spool
can never claim GFL99 profiles just because both sides stored a generic id
(Nylon and PA compare as one material). And a spool that names its own brand
keeps the stricter name path, so its suggestions stay brand-specific rather than
becoming the printer's whole generic table.
2026-07-31 13:15:35 +02:00
maziggy db6cdb0745 fix(camera): take the finish photo when the print ends, not when its last layer starts (#2547)
The photo fired the moment layer_num reached total_layer_num. That edge is
where the printer *starts* its final layer, not where it finishes it: the
reporter's H2C capture shows it arriving at 92% with mc_remaining_time=2,
three minutes and seventeen seconds and one filament change before the print
actually ended, so the frame caught the toolhead mid-print over the model.

The trigger also latched _finish_photo_captured, which locked out both the
stage-22 and FINISH triggers for the rest of the print — so on firmware that
never reports an end-of-print filament unload (H2C and A1 Mini confirmed)
nothing could replace the bad frame.

Remove the last-layer trigger. The photo is now taken at the FINISH-state
trigger, which every model sends and which lands after the toolhead parks.

Since Bambu's end G-code drops the plate ~100mm just before that, restore the
framing before capturing: absolute G90/G1 Z to max_z_height + 10mm clearance,
settle, capture, then drop it back so the print is as reachable as the printer
left it. Absolute is the safety argument — that Z is a height the toolhead
occupied seconds earlier, so it is inside the travel limits by construction and
leaves the nozzle above the part, and it is unambiguous across model families
because Z is the nozzle-to-bed gap whether the bed moves or the toolhead does.
M211 is never touched (#2579). This is what #1145, #1397 and #1565 asked for.

The height is only trusted when two independent sources agree: the archive is
matched by the finished print's subtask_name by equality (not LIKE, so "Cube"
cannot resolve to "Cube v2"), and its layer count from the 3MF must match the
layer count the printer reported over MQTT. Matching on "most recent archive
for this printer" was not safe — on_print_complete pops the _active_prints
binding concurrently, and a print Bambuddy failed to archive would have
resolved to its predecessor. A wrong height is the one failure that could drive
the nozzle into the model.

The move is additionally skipped when the print height is unknown, when a queue
item is pending for the printer, when the printer has left FINISH, and when the
new finish_photo_restore_plate setting is off.

for every FINISH-state capture — which is what shipped the mid-print photo —
the bank is used only when the dispatcher recorded that it injected End G-code
into this print, since a SwapMod snippet may have ejected the plate. The flag is
handed over in two steps (mark_pending at dispatch, adopt at print start) so it
can never outlive its print: a job started from the slicer or SD card adopts
False rather than inheriting its predecessor's answer. Those prints also skip
the plate move outright, bank or no bank.

The bank now refreshes on mc_percent advances as well as layer changes, via a
new on_print_progress callback. Layer changes stop the instant the final layer
begins, which left the #1867 fallback frame stale by the whole length of that
layer; progress keeps ticking there and freezes before the End G-code runs, so
a swapped plate still cannot reach the bank. The last-layer throttle exemption
is dropped, since it would now fire a grab on every percent tick.

On the timelapse path the moment producer returns early, so the consumer does
the restore itself before its live-grab fallback — the documented usual outcome
on P1-series, where the video has not transferred by the time the notification
goes out and the shipped photo was of an already-dropped plate. The two waits
are now derived from the settle window and the video poll timeout rather than
hardcoded; at the old flat 75s that fallback was guaranteed to be cut off
mid-settle.

extract_max_z_height_from_3mf reads only a bounded prefix of the plate G-code,
since a sliced plate is routinely tens of megabytes and the header is ~40 lines.
It returns None for missing, unparseable, zero and negative values so callers
must treat "don't know" as such rather than defaulting.
2026-07-31 12:55:05 +02:00
maziggy 432e956eff fix(camera): rotate every still exactly once, and cover the sources that let ffmpeg write the file
Review follow-ups on applying camera_rotation to finish photos and
layer-timelapse frames.

Rotating the frame popped from _stage22_finish_frames rotated one of its
sources twice. The cache has two kinds of feeder: live grabs, which are raw,
and the #1867 in-print bank, whose bytes come from
_capture_snapshot_for_notification and have already been rotated on the way
in. The consumer cannot tell them apart, so on the finish_state trigger - the
path the bank exists to serve, on firmware that never emits stg_cur=22 - a 180
degree rotation cancelled itself out and the photo was upside-down again,
which is the reported symptom exactly; 90 and 270 landed 180 out. Rotation now
happens where each frame is captured, so every entry in the cache carries one
rotation whatever produced it, and the invariant is stated both where the
cache is declared and where it is consumed.

Two finish-photo sources were still writing unrotated files: the built-in
camera's own capture_finish_photo, and the still extracted from a
printer-recorded timelapse - which is the *preferred* source for a built-in
camera print, so a user with a rotation set got a correctly oriented photo or
not depending on which source happened to win. Neither ever holds the frame as
bytes; ffmpeg writes the file and they return a filename. apply_camera_rotation_to_file
handles that case and is best-effort - a failed rotate leaves the unrotated
file rather than losing a delivered photo. The archived video itself is the
printer's own file and is not re-encoded, so it still plays at the camera's
native orientation; the CHANGELOG says so rather than leaving it to be
discovered.

apply_camera_rotation logs at debug, not info. It was on a path that runs once
per layer, where a tall print would have put hundreds of lines in the log for
something the surrounding capture already reports at debug.

The moved rotation logic had no test of its own - every existing test patches
it out and asserts the call, so a flipped sign or a dropped expand=True would
have shipped green. test_camera_rotation.py drives the real round trip: a
corner marker pins which way it turns, the dimensions pin that the frame is
not cropped, and an undecodable frame comes back by identity because a capture
path must not lose a frame to a failed rotate.

Tests for the fix itself sit on both sides of the cache. The producer half is
driven directly; the consumer half is a closure nested inside on_print_complete
with nothing able to reach it, so it is pinned by an AST guard - checked
against the source because the alternative is no check at all. Reverting
main.py to the pre-fix shape fails three of the five, the guard among them.

The three new tests used Path("/tmp/test") for a patched base_dir, which Bandit
flagged (B108); they take tmp_path now.
2026-07-31 09:30:06 +02:00
maziggy 08c9ec6749 fix(camera): protect an in-progress stitch from the orphan sweep, and only sweep this feature's own files
Review follow-ups on the orphaned timelapse session cleanup.

The sweep's own docstring said min_age_seconds made it safe to call mid-run.
It did not. on_print_complete drops the session from _active_sessions before
handing frames_dir to ffmpeg, so for the length of a stitch the directory
matches no active session, and its mtime is the last layer's frame write -
which on a tall print's final layer is easily older than the margin. The
default margin is 300s and the stitch timeout is also 300s, so the two were
tied with no headroom at all: a sweep landing in that window deleted ffmpeg's
input from under it. _finalizing_sessions now covers the stitch, set as the
session leaves _active_sessions and cleared in a finally so a failed stitch
cannot leak the marker and make that printer's leftovers permanently
un-sweepable. The docstring names all three guards and which gap each covers,
including that the margin does have real headroom for the two cases it suits -
a session mid-creation, and the freshly written .mp4 awaiting attach.

The file branch now requires the timelapse_<session_id>.mp4 shape its own
comment describes. It previously deleted any file under
timelapse_frames/<printer_id>/ past the margin; nothing else writes there
today, but age alone is not a reason to delete a file this feature did not
create.

Dropped ignore_errors=True from the rmtree. It made the surrounding
except OSError unreachable, so a read-only mount or a permissions problem was
counted and logged as a successful removal - and that log is the only evidence
an operator has of what was deleted.

Tests 5 -> 9: sparing a session mid-stitch, the finalizing marker cleared even
when the stitch raises, unrelated files left alone, and a failed removal not
counted. The failure test's rmtree stub honours the real contract and returns
silently when ignore_errors=True, because that silent no-op is exactly what the
old call could never observe; a stub that raised unconditionally would have
passed against both versions and proved nothing.

main.py is unchanged: it has no module-level logger, and the inline
logging.getLogger(__name__) the sweep uses is the idiom throughout lifespan.
2026-07-31 09:02:58 +02:00
maziggy 4888485a54 fix(camera): redact credentials, contain failures, and stop the external-camera test claiming a connection it never opened
Review follow-ups on the external-camera capture coalescing.

The coalescing was transplanted from camera.py, which is keyed by printer IP
and so has nothing to hide in a log line. These keys carry the camera URL, and
an RTSP camera URL routinely embeds user:pass@ - so the five new log lines
printed the password, one of them at warning level, where it reaches support
bundles. All five now go through _log_key(), which redacts before truncating:
slicing first can cut the URL short of the @ the pattern anchors on and leave
the password intact, which is why every other URL log in the module already
does it in that order.

_capture_frame_uncoalesced gained the blanket catch its camera.py counterpart
has. That is load-bearing once captures are shared: the wrapper hands one
task's outcome to every caller waiting on it and can only give a follower its
own turn for an outcome it recognises, so an escaping exception reached all of
them at once and none retried - one caller's failure becoming N. The per-type
helpers catch narrowly (aiohttp.ClientError / OSError / timeouts), so the
guarantee belongs here rather than resting on their coverage. CancelledError
is re-raised ahead of it, since the wrapper distinguishes a cancelled leader
from a failed one.

test_connection reports whether it shared a capture. It reaches capture_frame
like any other consumer, so a test landing while Obico is polling got that
frame back and answered "connected" for a connection it never made - the one
answer a connection test must not give silently. It still shares rather than
forcing its own capture, because forcing one would open the second handle to a
single-reader device that this whole mechanism exists to prevent. The response
carries `coalesced`, which also gives capture_in_flight() the consumer its
camera.py counterpart has in the Diagnose tool, and the Test button says
"shared with a capture already running" instead of a bare success.

Tests 12 -> 20: an unexpected error reported as a failed capture, a raising
leader whose follower still gets a frame, the three coalesced states, and
redaction on each log line that can carry a URL. The raising-leader test
patches _capture_rtsp_frame rather than _capture_frame_uncoalesced, since a
stand-in installed in the latter's place sits above the catch and would test
the wrapper against a shape it can no longer be handed.
2026-07-31 08:47:57 +02:00
maziggy 3daae22f3d Security hardening (maziggy/bambuddy-security #8) 2026-07-31 08:18:05 +02:00
maziggy c0ed97f287 Post work PR #2691 2026-07-30 18:09:15 +02:00
maziggy ac3e3cc60f fix(printers): don't retract a fan kit on a partial airduct frame
device.airduct is pushed field by field - the modeCur handler reads it with
an "in" check for that reason - so a frame can carry parts without carrying
every fan. Absence in that list is what tells us a kit is not fitted, and
taken from a truncated frame it made both accessory badges vanish mid-print
and started rejecting fan=aux2 on a printer that has the fan.

A parts list now counts as a full inventory only when it carries ids 1 (part
cooling) and 2 (aux). Neither is optional on a machine that reports an
airduct at all, and both appear in every layout in the support-package
archive - P2S base 1,2 / P2S+kit 1,2,3 / X2D 1,2,3,10 / H2C,H2D,H2S 1,2,3,6.
Anything narrower is a diff frame: its speeds are applied, presence is left
alone. Presence can still be added from a partial frame; only retraction
needs the full list, so a kit that really is removed still disappears.

Also compose showChamberFan from both model lists rather than branching
between them, so the P2S/X2D entries in MODELS_WITH_CHAMBER_FAN stay
reachable instead of reading as dead, and note in the fan-speed docstring
that the aux2 gate also rejects between connect and the first airduct push.
2026-07-30 18:04:38 +02:00
maziggy db538e43f1 fix(slice): give the slice modal one filament row per project slot (#2712)
The filament list is positional from the modal down to the CLI's
filament_N.json parts, but for a source that already carries slice_info the
requirements endpoint returns only the slots the plate consumes. A
MakerWorld model declaring four filaments and painting with slot 4 alone
therefore showed one dropdown, whose PETG pick the CLI bound to slot 1 —
slot 4 sliced with the profile baked into the source, and the print came out
PLA.

The endpoint now takes full_slots, which widens that answer to every
project slot with used_in_plate flags, and only the slice modal passes it.
Print-time AMS matching shares the endpoint and keeps the used-only list, so
it still asks for exactly the spools the job needs.
2026-07-30 17:34:26 +02:00
maziggy 83142c726c fix(slice): report a finished slice once, not once per queued poll
setInterval does not await an async callback. Slicing a large project
blocks the backend for seconds, so poll ticks piled up behind one stalled
request, each holding a snapshot taken while the job was still active.
They resolved together, and every one of them ran the completion path —
one toast and two query invalidations each. A 20s stall against the 1.5s
interval produced 13 "Sliced X" toasts from a single slice.

Only one poll round is now in flight at a time, which also stops queueing
requests against a backend that is already saturated. Completion is
recorded once per job id, and a round still awaiting a response when the
effect tears down now returns instead of acting.
2026-07-30 17:14:07 +02:00
maziggy ad785a95cb fix(queue): withdraw an expected print when the command never goes out
feat(db): warn when the connection pool can outgrow the PostgreSQL server

fix(mqtt): an unusable layer_num must not drop the printer connection

test: patch settings.base_dir via monkeypatch so it unwinds on error

test: restore the config module after reloading it
2026-07-30 16:31:09 +02:00
maziggy beca3a8d73 fix(mqtt): keep the layer total that arrives with the print-start frame (#2702)
fix(support): redact push_status values, not the serialised JSON (#2702)
2026-07-30 15:11:08 +02:00
maziggy 88dc56d6e1 Security hardening (maziggy/bambuddy-security #7)
fix(settings): accept JSON booleans on the Spoolman settings endpoint
2026-07-30 13:52:56 +02:00
maziggy d9da60dd8d fix(camera): reuse the live view's frame for external-camera captures (#2707)
On a printer with an external camera, watching the live view while a print
ran meant the layer timelapse recorded almost nothing and the finish photo
went out with no image. The reporter measured 0 of 87 layer captures on one
print and 0 of 105 on another, both watched throughout. A USB camera allows
one V4L2 handle, so a capture during a live view fails outright.

The built-in camera has had this rule since #1348 and #1271: reuse the
viewer's buffered frame rather than opening a second connection. It was
never extended to the external paths, and it could not have been -- the
buffer it depends on was only ever populated by the built-in paths.
generate_mjpeg_stream yields multipart-wrapped chunks, so the route layer
could not recover the JPEG, and a guarded caller would have found an empty
buffer and skipped every time.

So the stream now publishes each raw frame through a new on_frame callback
(parallel to on_process from #2675), and the six one-shot consumers reuse
it: layer timelapse, the finish-photo moment and its background fallback,
the notification snapshot, Obico polling, and the plate check. A viewer
attached with nothing buffered yet skips that one attempt rather than
competing -- kicking the viewer off is worse than missing a frame.

on_frame exceptions are logged and swallowed, like iter_subscriber's
on_unsubscribe: buffering is a side effect and must never be able to take
the live stream down with it. The external stream's teardown now releases
the buffered frame too, ownership-checked so a concurrent viewer of the
same printer keeps its own.

Two side effects on paths not touched here, both improvements: the snapshot
endpoint and the finish-photo fallback chain consult get_buffered_frame and
can now serve an external camera's live frame. plate_detection's docstring
already claimed this behaviour while implementing it only for the built-in
fallback; that drift is resolved.
2026-07-30 13:05:15 +02:00
maziggy 40e7b60e8c fix(camera): drain a streaming ffmpeg's stderr continuously (#2707)
ffmpeg is spawned with stderr=PIPE and it was only read on the error
paths, so for the life of a working stream nobody read that pipe. ffmpeg
writes its banner, the input analysis, then a progress line at a steady
rate; a 64 KiB pipe fills eventually, ffmpeg blocks writing to it, frames
stop, and the stream's own 30s timeout fires -- logged as "RTSP read
timeout" with no hint that we starved it ourselves.

How long that takes is unmeasured and evidently long: one H2D upstream
ran 21m36s without stalling, and an earlier 512 B/s extrapolation of mine
was mostly the one-off startup banner. So this is a bounded resource being
treated as unbounded, not a fault anyone has reported.

_FfmpegStderrTail drains the pipe continuously and keeps a 16 KiB rolling
tail. That tail is what the error paths now report, which is better
material than before: it holds what ffmpeg said as things went wrong,
where the on-demand read returned whatever was printed first -- usually
the banner, which the summariser strips anyway.

Three readers wanted this one pipe, and asyncio rejects concurrent reads
on a StreamReader, so the collector is authoritative: it registers by pid,
_read_ffmpeg_stderr returns its tail when present and otherwise reads the
pipe unchanged, and _terminate_ffmpeg skips its own stderr drain when the
collector owns it (the collector keeps draining through teardown, which is
all wait() needs). The generator starts it after the immediate-failure
check, which reads the pipe directly because the process is already dead,
and releases it after _terminate_ffmpeg.

text() goes through _summarize_ffmpeg_stderr like every other stderr log
here, so the access code ffmpeg echoes in its input URL stays masked.
aclose() awaits the cancelled pump rather than firing and forgetting, so
no pending task survives into loop teardown.
2026-07-30 12:52:11 +02:00
maziggy f26bcbbcce fix(camera): one registry key per stream, not per printer (issue #2707)
Closing a camera view and reopening it immediately could leave the new
stream unregistered while it was running and delivering frames. The
damage was all indirect: is_stream_active() reported no viewer, so Obico
polling and snapshots opened a second camera connection against the live
view (the thing #1348 and #1271 exist to prevent); the janitor's /proc
scan found an ffmpeg missing from _active_streams and killed the live
stream as an orphan; and /camera/stop reported "Stopped 0" with a stream
running.

The fan-out stream id was f"{printer_id}-fanout" -- constant per printer,
so every successive stream shared one registry key, and the departing
generator's finally popped whatever was under it, including its
successor's entry. The same finally also cleared the per-printer frame
buffer unconditionally, discarding the new stream's frame. It needed the
two streams to overlap, which the 4s teardown made easy.

Each stream now gets its own key via _new_fanout_stream_id(), so a
generator can only clean up after itself -- the external-camera path
already does this (#2675) and this brings the fan-out path in line. The
per-printer dicts are released through _release_printer_frame_state(),
which checks that no other stream for the printer is still running; both
the RTSP and chamber-image cleanups had the same unconditional pop.

Also hoisted time and uuid to module level and dropped four
function-local `import time` statements. A local import shadows the name
for the whole function, so any use on a branch that doesn't reach the
import raises UnboundLocalError -- a real hazard in camera_stream, whose
external-camera branch imported both while the RTSP path needs them too.
A test pins camera_stream as free of function-local imports.
2026-07-30 12:35:27 +02:00
maziggy 18cc906fad fix(camera): drain ffmpeg's pipes during teardown (#NNNN)
Closing a camera view logged "ffmpeg didn't terminate gracefully,
killing" followed by "ffmpeg did not exit within 2.0s of SIGKILL;
abandoning wait", on every single close. Both waits expired every time,
so teardown took a fixed 4.00s -- and since the firmware allows one
camera connection, that was 4s in which nothing else could use it.

ffmpeg is spawned with stdout and stderr as pipes and the teardown paths
have stopped reading them, so it sits blocked in write() on a full 64 KiB
pipe. SIGTERM cannot be acted on there: the handler only sets a flag that
the main loop polls, and the loop never gets back to the check. SIGKILL
does kill it, but asyncio resolves Process.wait()'s waiter through
_try_finish(), which requires every pipe transport to report
disconnected; paused, unread pipes never reach EOF, so wait() blocks with
returncode already set. A negative-control test shows returncode=-9 at
the instant the abandon fires.

Draining both pipes while stopping the process fixes both halves: 4.00s
becomes ~0.15s. The signal ladder and its bounds stay as backstops, so a
genuinely wedged process still cannot hang a stream, a Stop request or
the janitor.

This corrects _FFMPEG_KILL_TIMEOUT's premise and #2580's conclusion. That
12-hour hang was the unbounded form of this same self-inflicted stall, not
an ffmpeg stuck in uninterruptible I/O -- the process observed doing it
was in state S, which cannot survive a delivered SIGKILL. Bounding the
wait capped the symptom without removing the cause.
2026-07-30 10:31:22 +02:00
maziggy 73afa95047 fix(camera): share one connection between concurrent one-shot captures (#2705)
Bambu firmware allows exactly one camera connection. The existing guards
(is_stream_active / try_get_active_buffered_frame, #1271 and #1348) only
stop a one-shot capturer from competing with the fan-out broadcaster.
Nothing coordinated the capturers with each other, so with no viewer
attached every consumer correctly concluded it was not competing with a
viewer and then collided with the others. On the reporter's P2S an Obico
poll and a snapshot opened two RTSP sockets 207 ms apart, which knocked
over the fan-out stream feeding the camera wall; it was then reaped for
having received no frames for 58s.

capture_camera_frame_bytes() now coalesces: the first caller opens the
connection, callers arriving while it is in flight await the same result.
Eight paths reach that function independently - Obico polling, the
snapshot route, the finish-photo moment and its disk-writing sibling,
plate detection, the camera test and the diagnose tool - so the
single-flight sits at the bottom of the stack and no call site changes.

Keyed by IP, since that is what the firmware's limit applies to and the
function never sees a printer_id. The key excludes the timeout on
purpose: the call sites disagree about it, from 10s to 30s, so keying on
it would mean the Obico-vs-snapshot pair from the report never coalesced
at all.

It coalesces, it does not cache. A call arriving after the previous
capture finished still captures fresh, because plate detection and the
finish-photo path judge a running print from these frames and a stale one
there is worse than a slow one - #1397 was a finish photo taken seconds
late showing the bed already lowered.

Each caller waits on its own deadline rather than inheriting whichever
one happened to open the connection, and shield() means giving up leaves
the capture running for whoever else is still waiting. A follower whose
leader fails takes a turn of its own instead of inheriting a failure it
never had a chance to avoid; the leader has finished by then, so there is
nothing left to compete with. Bounded at two rounds. That also covers the
follower whose timeout is longer than the leader's, which coalescing
alone cannot. Cancellation is disambiguated via leader.cancelled(), so a
follower's own cancellation propagates while a cancelled leader is
treated as a failed one.

The leader is deliberately not wrapped in a second wait_for: the
implementation already enforces the timeout internally, where it can also
kill the ffmpeg process, and an outer deadline would abandon the
subprocess instead of killing it.

The diagnose tool now marks a stage whose frame came from a capture
already in flight as coalesced_capture. The pass is real evidence the
camera works, but duration_ms is then mostly time spent queueing, and a
diagnostic must not report a connection it never opened - the same reason
that file declares its live_stream_active shortcut instead of quietly
passing. Failures are not annotated, since a follower whose leader fails
goes on to capture on its own.
2026-07-30 09:49:56 +02:00
maziggy 9ef06449ef fix(filament): a unique preset match no longer counts as a colour match (#2687)
The Filament Mapping panel reported "(Ready)" with a green tick for a slot
where the slice wanted dark red and the auto-matched tray held dark green.
Manually picking that same tray reported the mismatch correctly, which is
what made it obvious something was inconsistent.

Auto-match ranks candidates by tray_info_idx first, and a uniquely-matching
preset was accepted as definitive on the premise "same preset = same spool =
same colour". The preset names the variant, not the spool: GFA00 is PLA
Basic, GFA01 PLA Matte, GFA17 PLA Translucent, in every colour Bambu sells.
The reporter's own bundle has eight GFA00 trays in eight colours. With one
Matte spool loaded, every Matte requirement idx-matched it and the colour
comparison was never reached - which is why this surfaced on PLA Matte and
not on Basic, where several spools are usually loaded and the match falls
through to the branch that does compare colours.

The verdict now comes from the tray that was selected rather than from which
rule selected it, and both branches share one comparison so they cannot
drift apart again. Selection is unchanged - the right variant still wins per
mismatch and the slot stays selected.

A requirement with no colour at all is treated as satisfied rather than
mismatched; 3MFs that omit it parse to "" and there is nothing to disagree
with. That also affects the manual branch, which used to flag it.

No dispatch change: _get_missing_force_color_slots already required an exact
colour, so force colour match was gated correctly throughout.
2026-07-30 09:26:46 +02:00
maziggy 9ef03067f7 feat(file-manager): show last activity on folder rows via the existing date toggle (issue #2680)
Follow-up to #2680: the calendar toggle only put dates on the file pane, so
the folder tree had no way to show the timestamp it was already sorting on.
FolderTreeItem now takes showModified and renders latest_activity_at under
the folder name, threaded through the recursive call so nested folders get it
too. No backend change - the field was already on the wire from the sort fix.

Folders are labelled "last activity", not "last modified", and get their own
i18n key. The value is the newest timestamp among the folder, its files and
everything below it, so a folder can read as newer than its own directory
mtime - calling that "modified" would look like a fresh instance of the
ls -lt mismatch the issue was originally about. Folders with no activity
render nothing rather than an Invalid Date placeholder.

The name span moved into a flex column so the second line does not disturb
the row's link badge, file count or kebab menu. That broke a folder-delete
test that reached the row via parentElement, now fixed to use closest().

Separately, the #996 collapse describe left an implementation on the
module-global localStorage.getItem mock, which silently collapsed the folder
tree for every describe after it. It resets in afterEach now; without that,
any later test asserting on nested folders fails for reasons unrelated to
what it is testing.
2026-07-30 09:09:47 +02:00
maziggy ef9357849b Post work PR #2693 2026-07-29 14:42:34 +02:00
maziggy 24322c71cb fix(tab-progress): drop the redundant status poll and quieten the test suite
The hook is mounted globally in WebSocketProvider, so refetchInterval on its
per-printer status queries added one request per printer every 30s on every
page. The Printers page already runs that fallback on the same query key, and
useWebSocket writes ['printerStatus', id] straight into the cache, so the poll
bought nothing outside the Printers page and cost a request per printer per
tab everywhere else.

Also captures document.title at mount instead of restoring to a hardcoded
'Bambuddy', so the default no longer has to be kept in sync with index.html.

jsdom has no canvas backend, so getContext('2d') logged a "Not implemented"
jsdomError with a full React stack on every run of the hook's tests, and the
favicon branch bailed on the null context and went untested. Stubbing
getContext/toDataURL removes the noise and lets the ring code run, so the
favicon swap and the restore-on-toggle-off path are now asserted.
2026-07-29 14:38:18 +02:00