2187 Commits
Author SHA1 Message Date
maziggy 14da007c7d fix(tests): assemble the upload batch instead of racing a sleep
The concurrent-dispatch tests went red on the Docker shard with
"expected all 6 printers to be uploaded to concurrently, but the
high-water mark was 4", on a scheduler that was dispatching all six
correctly.

peak == 6 is the claim that the sixth dispatch reaches its upload
before the first one finishes, and the dispatches do not arrive
together: each runs a preamble of database work first. So the
assertion was a race between that spread and a fixed 0.15 s sleep.
Measured here the spread is ~14 ms; on the runner it passed 150 ms.
Padding the sleep would only move the threshold and slow every test
that uses it.

_UploadRecorder(assemble=True) now holds each call until every
upload the pass launched has arrived, reading len(_inflight), which
is filled synchronously at launch and is therefore the batch size.
That states the property the peak assertions are about, directly and
with no time in it, and it ends sooner than the sleep it replaces -
the file runs 7.0s to 5.9s. _BATCH_DEADLINE_SECONDS bounds it so a
dispatch that really has gone serial fails on its peak assertion
instead of hanging to the pytest timeout.

test_check_queue_returns_without_awaiting_the_uploads keeps the
sleep, with the reason recorded next to it: it reads the rows while
the uploads are open, and an assembled batch releases as soon as it
is complete, which would let the dispatches flip those rows
mid-assertion.

The test engine also takes the application's own _set_sqlite_pragmas
listener rather than a copy, so the file is opened the way the
running system opens one - WAL, synchronous = NORMAL, a 15 s busy
timeout. SQLite's defaults spend an fsync per commit and lock the
whole file, which made a dispatch's preamble cost more than the
upload it precedes. That is what turned the previous commit's move
to a file-backed database into a visible failure.

Verified in both directions rather than by a green run: with two
printers' preambles delayed by a second, the old recorder reports
exactly the "high-water mark was 4" the runner saw and the shared
library test reports 3 == 4, while the assembling recorder passes
the same scenario. Without the induced skew, ruff is clean, the full
backend suite is green at 12224 passed, and ten consecutive runs of
the file pinned to one core show no failures.
2026-09-25 14:21:28 +02:00
maziggy 35c93a5362 fix(tests): give each concurrent-dispatch session its own connection
The scheduler's concurrent-dispatch tests failed on CI and passed
locally, on a commit that touched nothing but a text file. Six tests
in tests/unit/test_scheduler_concurrent_dispatch.py went red with
every queue item logged as "Status set to 'printing'" and four of
them read back as "pending", alongside "cannot commit transaction -
SQL statements in progress" and a PrintArchive that could not be
refreshed.

The fixtures built their farm on sqlite+aiosqlite:///:memory:, and
SQLAlchemy backs an in-memory SQLite with a StaticPool: one DBAPI
connection handed to every session, with nothing keeping them apart.
That was harmless while check_queue awaited its uploads inline,
because only one session was ever live at a time. Under the
refillable upload pool the uploads run as concurrent background
tasks with a session each, so their transactions interleave on that
single connection - a sibling session's close() rolls back another's
flushed-but-uncommitted UPDATE, and a commit() landing while another
session still holds a cursor raises the commit error above. Whether
the interleaving lands badly comes down to core count and
interpreter version, which is why a 30-core box on 3.13 stayed green
and a 4-vCPU runner on 3.11 did not.

The four fixtures now put the database in the test's own tmp_path,
which gets an AsyncAdaptedQueuePool and a connection per session -
what the application itself runs with (_resolve_pool_kwargs in
backend/app/core/database.py: pool_size 20, max_overflow 200). So
the harness is being brought in line with production rather than
having its assertions relaxed; nothing in the scheduler changes,
and the concurrency under test was correct throughout.

Checked against the failure mode rather than against a green run: an
isolated repro of the same shape - six writer sessions and one
reader session closing mid-transaction - yields all-pending on the
in-memory engine and all-printing on a file. The new risk is real
SQLite write contention on one file, so the file was run twelve
times pinned to two cores with no failures and no lock errors.

tests/unit/test_scheduler_busy_reasons_3018.py carries the same
harness on an in-memory engine, but dispatches to a single printer,
so it has no second session to race; left as is.
2026-09-25 14:00:07 +02:00
maziggy 724ce6f6ef fix(ams): stop showing the humidity drop index as a percentage (issue #3140)
Bambu sends two humidity fields that are not the same quantity.
    humidity_raw is relative humidity in percent; humidity is a 1-5 drop
    index, and it runs the other way -- OpenBambuAPI's push_info sample
    pairs "humidity:30%" with "humidity_idx:4", so a high index means dry
    where a high percentage means wet.

    Four call sites used the index whenever no percentage arrived. A unit
    sending only the index therefore rendered as "2%" in the green band
    while being the second-wettest of the five steps, charted an average of
    index values as a percentage, and sat under every humidity threshold
    forever, since no index can reach one -- the alarm and auto-drying could
    not fire for such a unit at all.

    - utils/ams_humidity: one leaf helper, a percentage or None. The index
      is never converted; None is what every caller already handles.
    - routes/printers, printer_manager, print_scheduler, main, bambu_mqtt:
      all five readings go through it, so the card, the websocket, the
      chart, the alarm and auto-drying cannot answer differently.
    - main: a unit that reports the index and no usable percentage says so
      once per unit in the log, with its firmware versions requested. No
      supported printer is known to do this, and "known" is doing work
      there -- the alternative is a card that goes blank with no trace.

    Three faults found while checking what else those paths touched:

    - main: humidity_raw=float(x) if x else None stored NULL for a numeric
      0% while writing 0.0 to humidity on the same row.
    - main: that same expression was unguarded, unlike the parse above it,
      so a non-numeric humidity_raw raised inside record_ams_history and
      aborted the pass for every printer, not just the one that sent it.
    - routes/ams_history: the averages were tested for truthiness, so a
      window averaging exactly 0 reported no average while the min and max
      beside it reported 0.0.

    An affected unit now reports no humidity rather than a number that means
    the opposite: the indicator is hidden, the chart leaves a gap, the alarm
    and auto-drying skip the unit. Temperature is untouched. Auto-drying's
    outcome is unchanged either way -- an index could never cross the
    threshold -- so only the intent moves.

    No supported printer is known to be affected; the report came from an
    install running X1Plus, which Bambuddy does not support. Verified
    against 7927 recorded samples from seven AMS units including an AMS-HT:
    not one used the fallback. Two percentages that did fall through to the
    index no longer do -- a reading with a decimal point, and "38.0", which
    int() rejected.
2026-09-24 14:53:17 +02:00
maziggy e4ce5c56e5 fix(virtual-printer): give each install's CA a name of its own (issue #3014)
A slicer holding the CA of two Bambuddy installs could only connect to
    one of them. Each CA worked on its own; together, one stopped, with the
    generic "Connect ... failed! [SN:..., code=-1]" that an install whose
    CA was never imported gives.

    Every install signed as exactly CN=Virtual Printer CA. A slicer's trust
    store is a flat list of certificates and OpenSSL resolves an issuer by
    Subject DN: it takes the first authority whose name matches and fails
    the chain when that one turns out not to have signed the certificate,
    rather than trying the next match. Whichever CA landed second in the
    file lost -- decided by nothing but the order they were appended in.
    Reproduced with openssl verify against a bundle holding two CAs: the
    first leaf verifies, the second fails with "certificate signature
    failure".

    - certificate.py: a newly generated CA takes a suffix from its own key
      identifier (CN=Virtual Printer CA D55808BE) and publishes that
      identifier, which the printer certificate points back at.
    - Existing CAs are untouched, so nothing has to be re-imported. A
      printer certificate signed by one keeps exactly the shape it has
      today: the authority key identifier is added only when the CA has an
      identifier to name.
    - tests: unique names per install, the identifier reaching the leaf,
      an existing CA being reused unchanged, and both chains verifying
      through openssl from a single trust store.

    The collision goes away as soon as one of the two CAs is newer than
    this change. Two installs that both predate it still collide until one
    has its bbl_ca.crt/.key deleted and regenerated, which is a re-import
    for that one -- documented in the wiki.
2026-09-24 14:52:37 +02:00
maziggy 7d75b313ed fix(notifications): send the ntfy priority the dialog was collecting (issue #3139)
The per-event Priority header from #990 never reached ntfy. The dialog
    builds its rows from the provider's event toggles and stores the map
    under those names -- on_print_complete, on_print_failed -- while every
    sender is called with the bare event name, print_complete. The lookup
    missed for all 18 events the dialog offers, so every notification went
    out at the ntfy server's default with the configured priority sitting
    untouched in the database.

    Both ends looked healthy, which is why it shipped. The stored config
    held exactly what was set, and the tests were green because they called
    _send_ntfy directly with the prefixed name -- the one spelling the
    running system never produces.

    - notification_service.py: accept either spelling, bare first, so
      existing configs keep working and nothing needs migrating.
    - schemas/notification.py: document both key forms, and which one the
      UI writes.
    - tests: use the bare names, and add one that runs from a finished
      print through to the outgoing request. Without the fix it fails on a
      header dict holding only Title, which is the assertion that was
      missing.

    The daily digest is unchanged: send_digest sends with no event_type at
    all, and the dialog offers no priority for it -- it is one message for
    several events.
2026-09-24 14:52:17 +02:00
maziggy 7d8fb15a84 refactor(models): break the schema cycle that backup and restore sort through
print_archives.library_file_id -> library_files.folder_id ->
    library_folders.archive_id -> print_archives. Three nullable SET NULL
    links, each reasonable alone, that together made a loop
    metadata.sorted_tables could not sort: it dropped those edges, warned on
    every backup and every restore, and could return an order placing a
    child before its parent -- which once imported library_files ahead of
    library_folders and killed a restore on a ForeignKeyViolation.

    The restore no longer depends on that order (it strips every foreign key
    before importing and adds them back after), but the backup export sorts
    the same way, and the warning ends with "may raise an error in a future
    release" -- which would break backup and restore on one upgrade.

    Marking one edge use_alter removes it from the sort graph, not from the
    database: PostgreSQL emits it as ALTER TABLE ADD CONSTRAINT, as it
    already did for every constraint on these three tables, and SQLite
    inlines it into CREATE TABLE, so ON DELETE SET NULL holds on both.
    Verified against PostgreSQL 16 and SQLite.
2026-09-24 14:51:59 +02:00
maziggy 8e839d58db fix(slicer): read the plate model once per part for the post-slice thumbnail (issue #3135)
The post-slice thumbnail loaded the sliced 3MF with trimesh, whose reader
    mishandles the Bambu Studio / OrcaSlicer layout: for every component that
    references a file in 3D/Objects/ it re-parses that file and appends all
    of its meshes again. N copies of a part came back as N^2 copies of its
    triangles, and each component carried every other part of the same file.
    25 bins of 10k faces loaded as 6.4M faces; the render took 54 s and
    8.4 GB on the event loop, and the server was OOM-killed.

    Parse the 3MF directly (lxml, entities/DTD/network off, streamed and
    freed element by element), walk objects, components and build items
    (p:path on either) with their transforms, and keep each mesh once.
    Decimate per unique mesh to its share of a face budget before placing
    instances, flip winding on mirrored placements, and skip the thumbnail
    above a face ceiling checked both on what decimation can reach and on
    what it delivered.

    Both slice routes run the render in a thread. The renderer uses
    matplotlib's Figure/Agg API instead of pyplot, whose process-global
    figure state let a threaded plate render and stl_thumbnail's
    event-loop render lay out and close each other's figures.
2026-09-24 14:51:33 +02:00
maziggy fb12f8d4f3 fix(queue): keep the filament override when a model job moves to one printer (issue #3133)
Switching an "Any P2S" job to a specific P2S cleared its filament
    override: "Specific Printer" empties the target model and the reset
    effect counted that as a model change. Printer mode also matched trays
    against the 3MF's colours, never sent the override, and left the old one
    on the row.

    The reset now compares against the last model actually targeted, so the
    switch keeps the override while a real model or plate change still clears
    it. Printer-mode tray matching (single, per-plate, multi-printer and the
    selector's per-printer editor) runs against the requirements with the
    overrides applied, mirroring the scheduler's _apply_filament_overrides; an
    entry naming the slot's own filament is not a swap and keeps its
    tray_info_idx. Printer-mode submits carry the user's overrides, and the
    create endpoint stores them for a printer-targeted job, so a dispatch-time
    recompute of an unresolved mapping looks for the same filament.

    Saving re-attaches the tray_info_idx an unchanged entry already had, so a
    virtual printer's force-colour PLA-variant pin (#2650) survives an edit in
    either assignment mode. The printer card's compatibility filter skips
    printer-targeted jobs: it mirrors the model scheduler, and hiding a job on
    filament would hide it from the printer it is going to run on.
2026-09-24 14:51:12 +02:00
maziggy 6e8d543d2a fix(ams): stop showing the humidity drop index as a percentage (issue #3140)
Bambu sends two humidity fields that are not the same quantity.
humidity_raw is relative humidity in percent; humidity is a 1-5 drop
index, and it runs the other way -- OpenBambuAPI's push_info sample
pairs "humidity:30%" with "humidity_idx:4", so a high index means dry
where a high percentage means wet.

Four call sites used the index whenever no percentage arrived. A unit
sending only the index therefore rendered as "2%" in the green band
while being the second-wettest of the five steps, charted an average of
index values as a percentage, and sat under every humidity threshold
forever, since no index can reach one -- the alarm and auto-drying could
not fire for such a unit at all.

- utils/ams_humidity: one leaf helper, a percentage or None. The index
  is never converted; None is what every caller already handles.
- routes/printers, printer_manager, print_scheduler, main, bambu_mqtt:
  all five readings go through it, so the card, the websocket, the
  chart, the alarm and auto-drying cannot answer differently.
- main: a unit that reports the index and no usable percentage says so
  once per unit in the log, with its firmware versions requested. No
  supported printer is known to do this, and "known" is doing work
  there -- the alternative is a card that goes blank with no trace.

Three faults found while checking what else those paths touched:

- main: humidity_raw=float(x) if x else None stored NULL for a numeric
  0% while writing 0.0 to humidity on the same row.
- main: that same expression was unguarded, unlike the parse above it,
  so a non-numeric humidity_raw raised inside record_ams_history and
  aborted the pass for every printer, not just the one that sent it.
- routes/ams_history: the averages were tested for truthiness, so a
  window averaging exactly 0 reported no average while the min and max
  beside it reported 0.0.

An affected unit now reports no humidity rather than a number that means
the opposite: the indicator is hidden, the chart leaves a gap, the alarm
and auto-drying skip the unit. Temperature is untouched. Auto-drying's
outcome is unchanged either way -- an index could never cross the
threshold -- so only the intent moves.

No supported printer is known to be affected; the report came from an
install running X1Plus, which Bambuddy does not support. Verified
against 7927 recorded samples from seven AMS units including an AMS-HT:
not one used the fallback. Two percentages that did fall through to the
index no longer do -- a reading with a decimal point, and "38.0", which
int() rejected.
2026-09-24 11:59:30 +02:00
maziggy fc6b953816 fix(virtual-printer): give each install's CA a name of its own (issue #3014)
A slicer holding the CA of two Bambuddy installs could only connect to
one of them. Each CA worked on its own; together, one stopped, with the
generic "Connect ... failed! [SN:..., code=-1]" that an install whose
CA was never imported gives.

Every install signed as exactly CN=Virtual Printer CA. A slicer's trust
store is a flat list of certificates and OpenSSL resolves an issuer by
Subject DN: it takes the first authority whose name matches and fails
the chain when that one turns out not to have signed the certificate,
rather than trying the next match. Whichever CA landed second in the
file lost -- decided by nothing but the order they were appended in.
Reproduced with openssl verify against a bundle holding two CAs: the
first leaf verifies, the second fails with "certificate signature
failure".

- certificate.py: a newly generated CA takes a suffix from its own key
  identifier (CN=Virtual Printer CA D55808BE) and publishes that
  identifier, which the printer certificate points back at.
- Existing CAs are untouched, so nothing has to be re-imported. A
  printer certificate signed by one keeps exactly the shape it has
  today: the authority key identifier is added only when the CA has an
  identifier to name.
- tests: unique names per install, the identifier reaching the leaf,
  an existing CA being reused unchanged, and both chains verifying
  through openssl from a single trust store.

The collision goes away as soon as one of the two CAs is newer than
this change. Two installs that both predate it still collide until one
has its bbl_ca.crt/.key deleted and regenerated, which is a re-import
for that one -- documented in the wiki.

Reported by @Steven-Pierce.
2026-09-24 11:22:40 +02:00
maziggy 1b20d1a968 fix(notifications): send the ntfy priority the dialog was collecting (issue #3139)
The per-event Priority header from #990 never reached ntfy. The dialog
builds its rows from the provider's event toggles and stores the map
under those names -- on_print_complete, on_print_failed -- while every
sender is called with the bare event name, print_complete. The lookup
missed for all 18 events the dialog offers, so every notification went
out at the ntfy server's default with the configured priority sitting
untouched in the database.

Both ends looked healthy, which is why it shipped. The stored config
held exactly what was set, and the tests were green because they called
_send_ntfy directly with the prefixed name -- the one spelling the
running system never produces.

- notification_service.py: accept either spelling, bare first, so
  existing configs keep working and nothing needs migrating.
- schemas/notification.py: document both key forms, and which one the
  UI writes.
- tests: use the bare names, and add one that runs from a finished
  print through to the outgoing request. Without the fix it fails on a
  header dict holding only Title, which is the assertion that was
  missing.

The daily digest is unchanged: send_digest sends with no event_type at
all, and the dialog offers no priority for it -- it is one message for
several events.
2026-09-24 10:00:33 +02:00
maziggy 58ea7a360d refactor(models): break the schema cycle that backup and restore sort through
print_archives.library_file_id -> library_files.folder_id ->
library_folders.archive_id -> print_archives. Three nullable SET NULL
links, each reasonable alone, that together made a loop
metadata.sorted_tables could not sort: it dropped those edges, warned on
every backup and every restore, and could return an order placing a
child before its parent -- which once imported library_files ahead of
library_folders and killed a restore on a ForeignKeyViolation.

The restore no longer depends on that order (it strips every foreign key
before importing and adds them back after), but the backup export sorts
the same way, and the warning ends with "may raise an error in a future
release" -- which would break backup and restore on one upgrade.

Marking one edge use_alter removes it from the sort graph, not from the
database: PostgreSQL emits it as ALTER TABLE ADD CONSTRAINT, as it
already did for every constraint on these three tables, and SQLite
inlines it into CREATE TABLE, so ON DELETE SET NULL holds on both.
Verified against PostgreSQL 16 and SQLite.
2026-09-23 16:58:14 +02:00
maziggy 0eb083b32e fix(slicer): read the plate model once per part for the post-slice thumbnail (issue #3135)
The post-slice thumbnail loaded the sliced 3MF with trimesh, whose reader
mishandles the Bambu Studio / OrcaSlicer layout: for every component that
references a file in 3D/Objects/ it re-parses that file and appends all
of its meshes again. N copies of a part came back as N^2 copies of its
triangles, and each component carried every other part of the same file.
25 bins of 10k faces loaded as 6.4M faces; the render took 54 s and
8.4 GB on the event loop, and the server was OOM-killed.

Parse the 3MF directly (lxml, entities/DTD/network off, streamed and
freed element by element), walk objects, components and build items
(p:path on either) with their transforms, and keep each mesh once.
Decimate per unique mesh to its share of a face budget before placing
instances, flip winding on mirrored placements, and skip the thumbnail
above a face ceiling checked both on what decimation can reach and on
what it delivered.

Both slice routes run the render in a thread. The renderer uses
matplotlib's Figure/Agg API instead of pyplot, whose process-global
figure state let a threaded plate render and stl_thumbnail's
event-loop render lay out and close each other's figures.
2026-09-22 10:47:00 +02:00
maziggy 89d94796ee fix(queue): keep the filament override when a model job moves to one printer (issue #3133)
Switching an "Any P2S" job to a specific P2S cleared its filament
override: "Specific Printer" empties the target model and the reset
effect counted that as a model change. Printer mode also matched trays
against the 3MF's colours, never sent the override, and left the old one
on the row.

The reset now compares against the last model actually targeted, so the
switch keeps the override while a real model or plate change still clears
it. Printer-mode tray matching (single, per-plate, multi-printer and the
selector's per-printer editor) runs against the requirements with the
overrides applied, mirroring the scheduler's _apply_filament_overrides; an
entry naming the slot's own filament is not a swap and keeps its
tray_info_idx. Printer-mode submits carry the user's overrides, and the
create endpoint stores them for a printer-targeted job, so a dispatch-time
recompute of an unresolved mapping looks for the same filament.

Saving re-attaches the tray_info_idx an unchanged entry already had, so a
virtual printer's force-colour PLA-variant pin (#2650) survives an edit in
either assignment mode. The printer card's compatibility filter skips
printer-targeted jobs: it mirrors the model scheduler, and hiding a job on
filament would hide it from the printer it is going to run on.
2026-09-22 09:45:56 +02:00
maziggy bec17de946 fix(queue): send the copy count for a cross-model print (issue #3101)
Selecting sliced files for two printer models and asking for 25 copies
    queued one item. The queue emptied as soon as it dispatched and the
    Batches tab stayed empty, because no batch is created at quantity 1.

    A multi-plate file moves the run count off the modal's Quantity field
    onto a stepper beside each plate (#342), hiding the field. The
    cross-model submit (#671) posts that field, which in this combination
    nothing can set, so it stayed at its initial 1. The modal read "19 runs
    in total" above a button that queued one.

    Per-plate steppers do not fit a cross-model job: its plate is chosen per
    candidate, in the alternatives list, so there is one number to give.
    Exclude cross-model from the per-plate mode and the global field comes
    back.

    Drop the plate selector in that mode too. Its choice never reached the
    request; it only keyed the filament-requirements query, so picking plate
    3 for a candidate while plate 1 stayed ticked above produced overrides
    computed from a plate the job would not print. That query now follows
    the primary file's own dropdown.

    Dispatch needed nothing -- it already gives each copy its own candidate
    rows -- but naming did. A cross-model job carries neither archive_id nor
    library_file_id, because the candidates are the files, so both branches
    that name a batch missed and every such order would have read "Batch" in
    the tab the reporter went looking in. Name it after the first candidate.

    The existing cross-model tests all mock a single-plate file, which is
    why the pair was never covered; the multi-plate case is added.
2026-09-21 15:28:49 +02:00
maziggy f98381f3d1 fix(queue): send the copy count for a cross-model print (issue #3101)
Selecting sliced files for two printer models and asking for 25 copies
queued one item. The queue emptied as soon as it dispatched and the
Batches tab stayed empty, because no batch is created at quantity 1.

A multi-plate file moves the run count off the modal's Quantity field
onto a stepper beside each plate (#342), hiding the field. The
cross-model submit (#671) posts that field, which in this combination
nothing can set, so it stayed at its initial 1. The modal read "19 runs
in total" above a button that queued one.

Per-plate steppers do not fit a cross-model job: its plate is chosen per
candidate, in the alternatives list, so there is one number to give.
Exclude cross-model from the per-plate mode and the global field comes
back.

Drop the plate selector in that mode too. Its choice never reached the
request; it only keyed the filament-requirements query, so picking plate
3 for a candidate while plate 1 stayed ticked above produced overrides
computed from a plate the job would not print. That query now follows
the primary file's own dropdown.

Dispatch needed nothing -- it already gives each copy its own candidate
rows -- but naming did. A cross-model job carries neither archive_id nor
library_file_id, because the candidates are the files, so both branches
that name a batch missed and every such order would have read "Batch" in
the tab the reporter went looking in. Name it after the first candidate.

The existing cross-model tests all mock a single-plate file, which is
why the pair was never covered; the multi-plate case is added.
2026-09-21 15:26:09 +02:00
maziggy fd3efe4a92 fix(archives): keep the project name when the wrong-plate guard rejects a 3MF (issue #3126)
Bambu Studio files a sliced print on the X2D's internal eMMC, which FTPS
    does not serve. The bounded probe found a same-named file on the card --
    an earlier slice of the same project, plate 4, against a running plate 1
    -- and #1204's guard correctly refused it rather than archive another
    plate's thumbnail, filament and cost.

    It then blanked subtask_name because swap_plate_suffix returned None. But
    None also means "this name carries no plate suffix", and such a name holds
    no stale plate number to be wrong about. The project name was dropped, the
    row fell through to the gcode_file path, and the archive was titled
    plate_1.

    Keep the name for the title only. subtask_name itself stays disowned,
    because it is what every lookup here is built from and it keys
    _active_prints, where the cover endpoint's own download of that same name
    would find this archive and hand the contradicted file to
    _recover_fallback_archive -- which checks a candidate is a readable 3MF
    and never which plate it holds. A corrected name is still registered:
    that one points at the plate actually running.

    Also name the X2D alongside H2-series and P2S in the Archives banner, the
    connection diagnostic and the storage-verdict docs -- it stores slicer
    sends the same way, and an X2D owner was told the explanation did not
    apply.
2026-09-21 14:54:06 +02:00
maziggy d23ba7645f Merge commit '8f9d79ccf0e6258bfd2562fb8c898ee2aa811685' into 1.2.5.6 2026-09-21 14:53:48 +02:00
maziggy 857647596a Merge pull request #2845 from pascalheidmann/refactor/modular-import
(Refactor): modularize import ("Makerworld tab")
2026-09-21 14:53:11 +02:00
maziggy ebc72e1d41 fix(archives): keep the project name when the wrong-plate guard rejects a 3MF (issue #3126)
Bambu Studio files a sliced print on the X2D's internal eMMC, which FTPS
does not serve. The bounded probe found a same-named file on the card --
an earlier slice of the same project, plate 4, against a running plate 1
-- and #1204's guard correctly refused it rather than archive another
plate's thumbnail, filament and cost.

It then blanked subtask_name because swap_plate_suffix returned None. But
None also means "this name carries no plate suffix", and such a name holds
no stale plate number to be wrong about. The project name was dropped, the
row fell through to the gcode_file path, and the archive was titled
plate_1.

Keep the name for the title only. subtask_name itself stays disowned,
because it is what every lookup here is built from and it keys
_active_prints, where the cover endpoint's own download of that same name
would find this archive and hand the contradicted file to
_recover_fallback_archive -- which checks a candidate is a readable 3MF
and never which plate it holds. A corrected name is still registered:
that one points at the plate actually running.

Also name the X2D alongside H2-series and P2S in the Archives banner, the
connection diagnostic and the storage-verdict docs -- it stores slicer
sends the same way, and an X2D owner was told the explanation did not
apply.
2026-09-21 11:18:35 +02:00
maziggy d44873b36e fix(inventory): one structured 409 for a tag another spool holds (issue #3110)
The two tag-link routes answered the same conflict differently. The
    built-in one said "Tag UID already linked to another active spool" and
    named nobody -- while holding the conflicting spool row it had just
    loaded -- and Spoolman mode named the spool inside a different English
    sentence. Neither was machine-readable, so a client had to parse prose
    to learn which spool to look at, and could only do it in one mode.

    Both now raise one shared constructor: code tag_already_linked, the
    holder's id, and which identifier collided. That is the detail shape
    insufficient_filament and printer_connection_failed already use, so
    ApiError parses it with no frontend change.

    Two active spools can carry one tag -- no unique index on either
    column, no conflict check on PATCH /spools/{id}, and /spools/bulk
    copies one payload including the tag into every row it creates -- and
    the lookup read that with scalar_one_or_none(), which raises on two
    rows. The exception escaped into the auth middleware's fail-closed
    handler, so the caller was told the authentication service was
    unavailable. Both lookups are now ordered and take the first row, as
    get_spool_by_tag earlier in the same file always has.

    Naming the lowest id means the Spoolman scan reads every row where it
    used to stop at its first match, so it now reads extra.tag defensively:
    that field is edited outside Bambuddy, and a single null further down
    the list would otherwise take the request down in place of the 409.

    The kiosk reads the new code: a refused link showed a flat "Failed to
    assign spool" and now names the spool holding the tag, reusing the
    inventory.tagAlreadyLinked key that no code referenced.
2026-09-20 13:41:08 +02:00
maziggy 999f5e0afe fix(ams): read the firmware presence bit, not the tray state (issue #3084)
Swapping a Bambu spool for one the AMS cannot read left Assign Spool
    publishing no ams_filament_setting at all. The printer kept showing "?"
    on its screen and in the slicer, and only Configure, which publishes
    unconditionally, put anything there.

    Four places asked the tray's `state` field whether a spool was in the
    slot. It cannot answer that. An AMS-HT reports its LOADED tray as 9
    rather than 11, because it does not feed into a shared buffer the way a
    4-slot AMS does -- the merge has skipped its own state heuristic for HT
    units since #2594 for exactly this reason. And the field is partly our
    own writing: apply_tray_exist_bits stamps state=9 on every slot whose
    tray_exist_bits bit is 0, and when the bit comes back it refreshes only
    the `exists` annotation beside it. Either way the slot sits at
    exists=True, state=9 until something configures it.

    That 9 also kept the deferred-configuration replay from firing -- its
    own "has a spool appeared" test was the same heuristic -- which is the
    deadlock #1322 removed from the assign path, still in place one step
    further along. And it is what deleted the assignments in #3100: with the
    replay never firing, the row kept the empty fingerprint it was stored
    with, and the first tray report naming a filament was read as a swap.

    All four now read tray_exist_bits first, which is the mask firmware
    answers this question with and the one the printer card has drawn its
    "?" from since #2527. The bit is allowed to overrule an "empty" state
    and nothing else: a bit reading empty deliberately does not start
    suppressing pushes that go out today, because the cost of computing a
    bit position wrong is a slot that silently stops configuring, against a
    saving of one message firmware would have dropped.

    A blank tray report from a slot the bit calls occupied no longer unlinks
    anything, off a print as well as during one, in both inventory modes --
    Spoolman's parse_ams_tray calls a tray with no type empty, so a tag-less
    spool assigned through the UI had its row deleted by the first idle push
    after it went in. A filament the AMS cannot identify is not a filament
    that was removed.
2026-09-20 13:40:30 +02:00
maziggy 7a20e731b5 fix(finance): show the currency the install is configured for (issue #3123)
The Finance page was the only surface in Bambuddy that read its currency
    from a data row rather than the `currency` setting, and it fell back to EUR
    where every other page falls back to USD. One variable drives every amount
    on that page, so the personal balance, the cost-center budgets and the whole
    transaction list were wrong together on any install not set to euros. It now
    takes the configured currency from /settings/ui-flags, which is readable by
    anyone who can see Finance -- /settings needs SETTINGS_READ, which a
    cost_centers:read_own user does not have.

    The backend was the other half. Of the four places that settle on a
    currency, three wrote a hardcoded "EUR": the wallet the API mints on demand,
    the wallet a print charge mints when none exists, and the balance returned
    for a user with no wallet row at all. All four now go through one resolver,
    which lives beside the rest of the balance logic.

    The wallet's currency column is removed outright rather than merely ignored.
    An install has one currency and nothing here converts between them, so a
    per-wallet copy could only ever drift from the setting -- and a column
    nothing reads is a trap for whoever finds it next. A startup migration drops
    it on both SQLite and PostgreSQL, after the raw CREATE TABLE that would
    otherwise re-add it on an install whose finance tables predate the ORM.
    SQLite builds older than 3.35 have no DROP COLUMN and keep it, harmlessly,
    since it has a default and no reader.

    Saving settings now invalidates the ui-flags query too. Nothing did, so a
    changed currency sat behind that query's staleTime before showing up. The
    sponsor prompt's own EUR fallback is now USD, matching AppSettings.
2026-09-20 13:40:10 +02:00
maziggy eb42104e01 fix(vp): offer every host IP as a bind target, not one per adapter (issue #3121)
Each enabled virtual printer needs its own IP address, and the documented
    way to get several is to add secondary addresses to the adapter already in
    use. Linux reads those back through `ip -j addr show`, which reports every
    address. Windows and macOS have no `ip` command and fell through to a psutil
    enumeration that stopped at the first IPv4 of each adapter, so a host with
    three addresses on one NIC offered exactly one bind target and the second
    virtual printer could only fail with "Bind IP ... is already in use".

    The psutil path now collects every address, marking the ones after an
    adapter's first as aliases the way the iproute2 path does. Callers that want
    interfaces rather than addresses -- the discovery scan and the support bundle
    -- project the primaries back out, so their view is unchanged.

    The ioctl fallback that a Linux host without iproute2 used to get is now
    reached only when psutil itself is missing, which gains that host aliases
    too. The interface-name exclusions stay Linux-only: they are Linux device
    names, and a Windows adapter called "Local Area Connection" matches the "lo"
    prefix.
2026-09-20 13:39:50 +02:00
maziggy 2d385cf978 fix(install): sign the Python that macOS grants local network access to (issue #3114)
macOS attributes Local Network permission to a code signature and judges a
    launchd-spawned process on its own, rather than letting it inherit the grant
    of the Terminal that started it. Homebrew ships Python unsigned on Intel, so
    there is no identity for the grant to attach to: every connection to a LAN
    address is dropped with no error the application can log and no permission
    prompt. The printer reads as unreachable and nothing says why, and the entry
    in Privacy & Security cannot be made to work because it refers to an identity
    that no longer resolves.

    install.sh signs during a macOS install; update_macos.sh re-checks on every
    update, because `brew upgrade python` installs a fresh unsigned binary under
    a new versioned path.

    Both sign only what is currently unsigned. That gate is load-bearing: on
    arm64 the linker ad-hoc signs every binary and the identity is a hash of the
    file, so re-signing would rotate it and revoke a working grant on each update.
    A python.org build carries a real Developer ID and must not be downgraded for
    the same reason.

    The interpreter and the framework's Python.app are both signed. The first is
    what sys._base_executable resolves to and what the reporter's TCC log names;
    the second is what his fix actually targeted. Which one macOS attributes
    could not be established from either, and signing both costs nothing.

    -----

    fix(diagnostics): name the macOS permission that silently blocks the printer (issue #3114)

    The port checks reported all three ports unreachable while the subnet check
    passed, and port_mqtt's fix text sent the reporter after firewalls and IP
    addresses. On a macOS native install that pattern has a cause neither of
    those covers: no Local Network grant, denied with no error and no prompt.

    A new macos_local_network check, appended on macOS only so no permanently
    dimmed row appears for anyone else. It passes when the control port answered,
    which is proof the permission is in place and means the signature probe never
    runs on a healthy diagnostic. Otherwise it probes the interpreter: an
    unsigned one gets the repair that fixes it, a signed one gets System Settings
    — the arm64 case, where the identity is a hash of the binary, so a Python
    upgrade presents macOS with a new application and strands the old grant.

    Always warn, never fail, and only once port_mqtt has already failed, so this
    can never be why a green diagnostic turns red. A printer that is simply
    switched off produces the same all-ports-dead pattern, which is why the
    signature, not the pattern, is what earns the specific advice. An
    undeterminable signature is reported as the generic case rather than as
    unsigned: that advice rewrites a file in the user's Python installation and
    must not be offered on a guess.
2026-09-20 13:39:22 +02:00
maziggy 25c36ba5de fix(library): stop reporting success for a bulk add that queued nothing (issue #3112)
POST /library/files/add-to-queue reported every per-file rejection in an
    errors array and returned 200 regardless. A caller that checks the status
    code saw a successful request, no visible failure, and no queue item.
    That is a 400 now when nothing at all was added, with the same reasons in
    the body. A call that created some items still succeeds, because it did.

    The items it created were aimed at nothing. The route always wrote
    printer_id=None with no target_model, and the scheduler dispatches on one
    or the other -- so those rows matched neither branch and could never be
    picked up by anything. They sat in Unassigned until someone opened each
    one by hand.

    The request takes an optional printer_id or target_model for the batch,
    and with neither it aims each file at the model its own G-code declares.
    Only when a printer of that model is active: owning no H2D is the user's
    situation rather than their mistake, so the file still queues as the
    unassigned row it has always been, rather than gaining a target nothing
    can answer.

    Three gates POST /queue/ has applied for a while now apply here too,
    because an item reaching the scheduler through this route has to be as
    printable as one reaching it through that one: the cross-model check that
    stops a file sliced for one printer being dispatched to another (#2578),
    the filename check that would otherwise surface as a failed upload hours
    later (#1540), and the filament requirements the scheduler matches before
    handing a model-based item to hardware.

    Nothing inside Bambuddy calls this endpoint -- the Library's own Print
    action goes through the queue API with a printer already chosen -- which
    is how it came to drift this far from it.

    -----

    fix(library): scope add-to-queue file reads to the caller

    The bulk add resolved its files by raw id. Every other read in this
    module goes through the ownership gate, and so does the single-item
    queue path; this one did not.

    Invisible rows are dropped before the loop, so they report as the plain
    "File not found" an unknown id already gets.
2026-09-20 13:38:54 +02:00
maziggy 0b9ba0e1ef fix(diagnostics): read the subnet the host is actually on (issue #3092)
The Network subnet check told the reporter that 192.168.98.170 and
    192.168.96.9 were on different networks and to go configure routing
    between them. They are four hundred addresses apart inside one
    192.168.96.0/22 LAN.

    An IPv4 address does not carry its prefix, and the check supplied /24
    for both sides. That is the most common LAN and not the only one, and
    the guess is wrong in both directions: it splits a /22 and it merges a
    /25. Read the prefix off the interface that owns the address instead.

    find_local_ipv4_network() enumerates every interface, including the ones
    EXCLUDED_INTERFACE_PREFIXES hides. That list keeps docker0 and friends
    out of the Virtual Printer's bind dropdown; here the caller is asking
    about an address the kernel has already picked as a route source, and
    answering "unknown" because it sits on a bridge would be a worse answer
    than the truth. When nothing claims the address the check skips, which
    is what it always did with no host IP at all -- it must not assert a
    split it cannot see.

    The same check chose which of Bambuddy's own addresses to compare by
    probing a route toward 10.255.255.255, which on a multi-homed host is
    not the interface the printer is on. It asks for the route toward the
    printer now. On a two-NIC dev box that alone was warning about a printer
    sitting on the second card's own subnet.

    The probe takes IPv4 literals only. connect() on a name would resolve
    it on the event loop, and _same_subnet rejects names anyway, so nothing
    is lost. Resolving the prefix shells out to `ip -j addr show`, so it
    moves off the loop too.

    -----

    fix(diagnostics): name the container engine instead of asking about Docker (issue #3092)

    "Not running in Docker - not applicable", said to a Bambuddy inside a
    Podman container. It reads as "you are on bare metal", and it sent the
    reporter looking for his problem somewhere else.

    Podman runs Bambuddy in exactly the two shapes Docker does, and the
    shape is the thing that breaks printer discovery and the Virtual
    Printer. detect_container_runtime() names the engine -- Docker, Podman,
    Kubernetes, containerd, LXC, or a container it cannot place -- and the
    check became Container network mode.

    is_running_in_docker() is deliberately left alone rather than rewritten
    on top of it. Three callers key real behaviour off that flag, and one of
    them switches the Add Printer flow from SSDP to subnet scanning. SSDP
    works for a host-networked Podman container, so answering True there
    would take a working feature away to fix a sentence. Widening it is a
    separate decision from naming the engine, so it is made separately.

    Mode detection keeps the original signal first, which also makes the
    Docker path incapable of regressing: a Docker host always has a docker0,
    so a container that sees one shares its namespace, and the new rules can
    only turn a warning into a pass. That signal says nothing about Podman,
    which creates no such interface on a host running no bridge containers --
    which is how host networking came to be reported as bridge. The general
    form of the same idea answers for Podman: an interface whose iflink
    equals its ifindex was created in this namespace, and a NAT-networked
    container only ever receives one end of a veth pair. tun/tap is skipped,
    because a container may run its own WireGuard and that tun is native to
    a namespace it is not evidence of. The interface also has to be the one
    the kernel just named -- sysfs is namespace-tagged but a bind-mounted
    host /sys is not, and reading a colliding name's numbers would be
    reading another namespace's answer.

    What is still unreadable now says so and suggests host networking if
    discovery is failing, rather than guessing bridge and telling a healthy
    install to recreate itself. An LXC or LXD system container is named and
    told the question does not apply: it is on the LAN like a small virtual
    machine, so there is no network mode to recommend -- and its subnet
    check still runs.

    An engine we cannot name is a sentinel the frontend localizes, not a
    word interpolated into thirteen other languages.

    The support bundle carries the engine name beside the Docker flag, so
    the next report of this shape is answerable from the bundle.
2026-09-20 13:38:08 +02:00
maziggy 31ea824db8 fix(spoolbuddy): show the colour name the rest of Bambuddy shows (issue #3090)
SpoolBuddy said "Unknown color" under a correctly-coloured swatch for
    spools the inventory page names without trouble.

    The name was never in the spool record. Bambu's RFID tags frequently
    carry no readable colour name -- some carry an internal code instead --
    so Bambuddy has always resolved the swatch's own hex against the colour
    catalog, and the kiosk was rendering the empty column. Exactly one
    SpoolBuddy file already did it right, which is what marks this as an
    inconsistency rather than a kiosk simplification.

    Route every SpoolBuddy colour-name display through resolveSpoolColorName,
    which also stops the spools that do carry a code from showing "A06-D0" at
    the user. The write-tag edit form keeps the raw stored value on purpose:
    offering a derived name for editing invites the user to save it as though
    they had typed it.

    Spoolman has no colour-name field at all, so _map_spoolman_spool puts the
    spool's subtype there and sets color_name_is_synthesized. That flag now
    travels on the tag-matched broadcast, and resolveSpoolColorName takes a
    third argument to honour it -- a synthesised name loses to the catalog
    and survives only as a last resort. Spoolman installs were reading
    "Silk+" as a colour on the Inventory page and the AMS hover card too, so
    those call sites pass the flag as well.

    Searching by a colour you can read on screen now finds it, in the kiosk
    and in Bambuddy: the shared inventory filter matches the resolved name as
    well as the stored one. That makes the filter depend on the catalog,
    which loads asynchronously, so the three memoised call sites take its
    version as a dependency -- without that, a query typed before the catalog
    arrives keeps its empty result and reproduces the very symptom being
    fixed.

    The fallback label was hardcoded English in components that already
    import useTranslation; it is now spoolbuddy.spool.unknownColor in all 14
2026-09-20 13:37:39 +02:00
maziggy e0377d25db fix(camera): stream an external RTSP camera that describes itself late (issue #3082)
An external camera could pass the connection test, play in VLC, and show a
    black live view that gave up after a few seconds.

    The two RTSP paths were not asking ffmpeg for the same thing. The one-shot
    _capture_rtsp_frame passed no probe settings and got ffmpeg's defaults;
    _stream_rtsp hard-coded -probesize 32 -analyzeduration 0. Thirty-two bytes
    is enough for a camera that puts its H.264 parameters in the SDP, and not
    enough for one that sends them in-band a moment later -- a WebRTC source
    republished through go2rtc, in the reporter's case. ffmpeg then starts no
    decoder and yields nothing at all, which is why the test button kept
    passing while the live view stayed black.

    Those settings were never chosen for external cameras: they arrived with
    the P2S TLS proxy (#661) as fast-start tuning for the printer camera path,
    where the source is a known Bambu model, and were copied here in the same
    commit. This path has no model to tune against and belongs on the
    defaults, which are a ceiling rather than a wait -- a camera that
    announces itself in the first packet still starts as fast as it did.

    Drop the probe cap from _stream_rtsp. Keep -fflags nobuffer and -flags
    low_delay, which bear on how long ffmpeg sits on frames it already has
    rather than how long it may look before it has any. Leave
    _capture_rtsp_frame and the per-model printer profiles alone.

    Tests pin the absence of both flags, the presence of the low-latency ones,
    and the property underneath: both RTSP paths must probe alike, or passing
    the test button again means nothing about the live view. The ffmpeg
    subprocess fakes move to backend/tests/_fixtures/external_camera.py so the
    SSRF suite and this one share one definition.
2026-09-20 13:37:15 +02:00
maziggy c204c79363 fix(jog): send the nozzle-bed gap the API promises on every model (issue #1334)
POST /printers/{id}/bed-jog takes a signed nozzle-bed gap, documented since it
    was written: positive asks for more room between the nozzle and the plate. On
    an A1 it did the opposite. The reporter sent distance=5 for clearance and
    watched the toolhead come down.

    The sign had been flipped on A1 models since the original report on this issue,
    where an A1 Mini owner clicked an arrow labelled "move the plate up" and watched
    the nozzle dive. That is a labelling problem -- a bed-slinger's plate does not
    move in Z at all, so closing the gap shows up as the toolhead descending -- and
    it was solved in the transport layer, which turned a parameter documented as
    model-independent into one that meant the opposite thing on part of the fleet.

    Z is the nozzle-to-bed distance on every Bambu model, by definition of the
    coordinate system rather than by convention: G1 Z+ opens the gap whether the bed
    drops away from a fixed nozzle (X1/P1/H2, whose end G-code parks with
    G1 Z{max_layer_z + 100}) or the nozzle rises off a fixed bed (A1/A2L). The
    finish-photo plate restore already relies on exactly that and carries no model
    branch. So distance goes onto the wire unchanged and one call means one physical
    outcome everywhere: positive is the safe direction on every printer.

    Which way an arrow points is a different question, about the machine in front of
    the user rather than about G-code, so the printer card answers it and asks for
    the gap it wants. The buttons move what you would expect them to move, exactly
    as before; on a bed-slinger they now say toolhead rather than plate.

    The A2L never had the old fix. It slings its bed the same way the A1 does, but
    the inversion listed the A1 names and the A2L was not among them, so its up
    arrow has been sending the toolhead at the plate for as long as the machine has
    been supported. The new classifier also covers the alternate internal codes
    A04 / A11 / A12, which LINEAR_RAIL_MODELS and SINGLE_NOZZLE_FLOW_MODELS both
    carry and the old gate did not.

    is_bed_slinger is gone from the backend rather than widened: with the route
    model-independent it had no caller, and a kinematics helper sitting unused in
    the service layer invites the next person to assume the backend handles
    direction. It does not, deliberately.

    Separately, the soft-endstop comments on both jog routes claimed the firmware
    clamps a bare move at the travel limit. It does not, and #2579 measured that:
    an H2D at its Z limit ran straight past a clean G91/G1 Z-1.00/G90, while its own
    touchscreen refuses the identical move. What #2579 removed was M211 S0, which
    disabled the limits globally and took the touchscreen's protection with them.
    The jog popover has warned about this correctly the whole time; only the code
    comments disagreed with it.
2026-09-20 13:36:46 +02:00
maziggy 1f88b9846f fix(queue): tell a pinned queue item why it is waiting (issue #3074)
A job queued as "Any X1C" explains itself when it cannot start: the
    model-based branch builds a reason for every candidate printer and puts it
    on the row, so the queue shows "Busy: X1C-01" or "Waiting for filament:
    X1C-02 (needs PETG)". The same job pinned to one printer showed nothing.
    It sat at Pending with waiting_reason NULL for as long as that printer was
    busy, which from the outside is indistinguishable from a queue that has
    stopped working -- the reporter watched fourteen minutes of it while his
    X1C ran a print he had started from its own screen.

    The fixed-printer branch had six ways out and none of them wrote the field.
    The sensor interlock (#1148) was its only writer, and it cleared the field
    up front on every pass where no sensor was holding the printer, so NULL was
    not an oversight on those paths but a guarantee.

    Every exit now writes, through one helper. The reasons reuse the
    model-based branch's vocabulary so _is_busy_only() keeps deciding what is
    worth a notification: a printer that is printing, drying, or working
    through the item ahead of this one reads as "Busy: <printer>" and stays
    silent, because it resolves itself. A printer that is off with no Auto On
    plug, and one whose plug could not switch it on, are worth saying.

    A finished plate nobody has acknowledged is split out from plain busy and
    named as itself. _is_printer_idle() returns the same plain False for that
    and for a running print, but they are not the same thing to the person
    looking at the queue: one clears itself and the other needs somebody to
    walk over to the printer.

    That notification fires on the transition into asking, where a busy-only
    reason counts as not asking. Testing whether the item was waiting at all --
    which is what the model-based branch does -- would never fire it here:
    nobody's queue goes straight from idle to an unconfirmed plate, it waits
    behind the print first. The cost is that a printer dropping offline,
    returning busy and dropping again asks twice rather than once.

    The interlock stays silent. It has never sent this notification, and a
    change about what the queue displays is not the place to start.

    Clearing the field up front is gone with it. It existed so a shut door
    could not leave "Waiting on Enclosure Door" standing while the printer
    stayed busy with something else, and the new rule carries that guarantee
    instead -- whichever exit runs next overwrites it, and the dispatch path
    clears it.

    Two paths clear it that the report did not mention. A staged item and a
    future-scheduled one skip before this branch and never reach it again, so
    anything written on an earlier pass would outlive its condition for the
    life of the row. That includes the filament-deficit check, which stages the
    item itself.

    The notification is wrapped: a queue that cannot say why it is waiting is
    the bug being fixed, and a queue that stops dispatching because a provider
    timed out would be a worse one.

    On the frontend, the queue timeline drops any pending item carrying a
    reason, on the grounds that such an item will not auto-dispatch. That held
    while only the model-based branch wrote the field; "Busy: <printer>" is
    now the commonest reason there is, and it describes the very chain the
    timeline forecasts, so the rule would have emptied the view for anyone
    whose queue is pinned. It now asks whether the reason needs the user, via
    a small shared reader of the same shape the scheduler encodes.

    Which job goes out, and when, is unchanged: running the previous scheduler
    and this one over the same 768 states dispatches the same items in the same
    order with the same statuses, across 1452 rows that now carry a reason.
2026-09-20 13:36:03 +02:00
maziggy 0b8cc823e7 fix(queue): print a plate whose filaments are all on the external spool (issue #3087)
The reporter's P1S heated up, sat at Heatbed preheating for ten and a half
    minutes, then paused with 07FF_8012, "Failed to get AMS mapping table".
    Resuming only reheated it. Prints that fed from the AMS were fine.

    The plate was one filament of a seven-filament MakerWorld project, mapped by
    hand to the external spool. slice_info.config numbers filaments across the
    whole project, so the mapping for that plate is [-1,-1,-1,-1,-1,-1,254]: six
    placeholders and the spool holder. The command builder decides whether a print
    needs the AMS by asking whether the mapping is entirely external, and six -1s
    answer no. So the print went out as use_ams=true carrying a flat mapping of
    nothing but -1 -- 254 is deliberately never sent raw, the firmware reads it as
    AMS tray 0 -- which is exactly the mapping table the firmware then could not
    find.

    The builder cannot fix this itself. Down there a -1 is either padding for a
    filament this plate does not print, which is BambuStudio's own convention, or a
    slot that never resolved to a tray, and sending the second one to the spool
    holder is what #2589 exists to prevent. They are the same byte.

    The scheduler knows. extract_filament_requirements drops every filament with
    used_g <= 0, so it names precisely the slots the plate prints. When all of those
    are an explicit 254/255, dispatch now sends use_ams=false and the print runs.
    When one of them resolved to nothing, the flag is left alone and the firmware
    rejects the print as it does today -- deliberately, because that is the case
    where guessing would print a filament in the wrong material without saying so.

    Single-nozzle only, mirroring the reconcile in the command builder: on a
    two-extruder printer use_ams selects which nozzle to feed rather than whether to
    use the AMS, so an H2D with a spool on each side must keep the flag it was
    given. Judged generously from the model name and from live telemetry -- a second
    nozzle reporting a diameter, an extruder map, or more than one external feed --
    because a wrong yes only preserves existing behaviour while a wrong no would
    reroute the print. H2S stays single-nozzle (#1386).

    Nothing in the command builder changed. Its own reconcile keeps the exact
    semantics #2589, #2595 and #797 gave it, and now usually agrees with a decision
    that was already made one layer up. Where the file has no parseable filament
    list the mapping is left exactly as before, the same evidence-only convention
    as #2771, and the parse itself sits behind a check for anything external at all
    so an AMS-only print never opens the file.

    Covered end to end at the dispatcher, including the reporter's seven-filament
    shape, a plate mixing the spool holder with an AMS tray, a consumed slot that
    never resolved, both external feeds on a dual-nozzle machine, and a 3MF with no
    filament list at all.
2026-09-20 13:35:25 +02:00
maziggy 14b322d86a fix(mqtt): never wait for a wedged paho network thread (issue #3068)
The reporter's A1 had been offline 38 hours and still answered on 8883, so
    the connection watchdog did exactly what it exists for: rebuild the session
    with a fresh client, since anything left in the old one's QoS 1 queue would
    otherwise replay onto the next print (#1136). The rebuild ended in paho's
    loop_stop(), which sets a terminate flag and then joins the network thread
    with no timeout.

    That thread only reads the flag between iterations of loop_forever, so it
    cannot read it while parked inside reconnect() -> _ssl_wrap_socket() ->
    do_handshake(). paho gives that handshake the keepalive as its socket
    timeout -- 30s here -- and a socket timeout is per operation, renewed by
    every byte the peer sends. A printer that answers TCP and then trickles
    holds the join open for as long as it likes.

    The join ran on the asyncio thread. Bambuddy stopped answering anything --
    UI, API, /health -- while the process stayed up, which is why a
    restart: unless-stopped container never restarted.

    Retiring a client no longer waits for it. The replacement is built at once
    and the old one is shut down on a thread of its own that nobody joins. Its
    callbacks are detached first, inline: blocking until the network thread was
    gone is what used to guarantee a client we had let go of could no longer
    touch our state, and with the teardown detached a zombie that finishes its
    handshake would otherwise auto-reconnect and report itself connected behind
    its replacement's back. disconnect() still goes out, still promptly, because
    that is what stops paho's auto-reconnect and the replay with it.

    The reported watchdog is one of six callers. The queue's dispatch recovery
    and check_staleness -- reached from an ordinary status poll -- share
    _hard_reset_client; editing, deleting and hand-disconnecting a printer share
    disconnect(); the relay and smart-plug services had the same join on their
    shutdown path, where a wedged broker stopped the process from exiting at
    all. #1445 was this join too, from the add-printer probe, and its off-loop
    teardown stays as it is.

    disconnect() stays quiet on the way out, as it always effectively did.
    paho's callback used to land during the join, but it suppresses itself for a
    clean disconnect of a printer that reported in the last ten seconds, so a
    healthy printer disconnected by hand never announced itself offline.
    Announcing it now would tell the user their printer had gone offline a
    minute after they disconnected it on purpose (#1752).

    A retirement that takes more than five seconds logs which printer it was.
    The whole point is that the next one of these should not have to be
    diagnosed from a thread dump.
2026-09-20 13:34:55 +02:00
maziggy 743d657b46 fix(drying): dry a composite spool as its base material (issue #3067)
The reporter's AMS-HT would not auto-dry PA6-CF, and drying the same spool by
    hand worked. The scheduler reduced a tray to a preset key by splitting on spaces
    only, so "PA6-CF" stayed "PA6-CF", matched none of the eight rows the preset
    table has, and the tray was read as holding nothing worth drying. The AMS was
    then passed over on every scheduler sweep, silently, because every caller reads
    "no row" as "nothing to do for this tray".

    It was never only nylon. Of the 41 types a printer can report, 33 had no row
    under that rule, and 20 of those have a base material sitting right there: every
    -CF, -GF and -AERO variant of PLA, PETG, ABS, ASA, PC and PA.

    Doing it by hand worked because the drying popover has resolved composites since
    exact key first, so a row the user added for the exact type still wins, then the
    suffix, then an alias map reading PA6, PA11, PA12, PAHT, PPA and Nylon as PA.

    PPA is the one alias that is a judgement rather than a spelling. Polyphthalamide
    is a distinct polymer, not a grade of nylon -- but it is an aromatic polyamide,
    it takes up moisture the same way, and PA's row is the hottest the table has.

    A material with no row and no alias is still skipped rather than dried at a
    number nothing here can source. That is where this parts company with the
    popover, which falls back to PLA because a dropdown has to show something.

    The table is user-editable JSON, so a preset row can be present and empty. That
    has always meant "skip this material" and still does: the temp and hours reads
    fall back per field to 55C/12h, which would dry a PLA spool at 55 degrees.

    Two more callers had the same line and move with it: per-filament humidity
    thresholds, where the override set for a material never applied to that
    material's composites, and the chamber preheat target, which had the suffix half
    of this from #2902 but not the aliases.

    The preheat test for that half reimplemented the lookup inline rather than
    calling it, so it would have passed whatever the function did. It calls it now.
2026-09-20 13:34:32 +02:00
maziggy 2b2617ecb7 fix(archives): come back for a 3MF whose transfer ran out of time (issue #3063)
The reporter's P1S had the sliced file on its card and was serving it. The 19MB
    transfer just did not finish inside the budget while the printer was also running
    its camera, its status messages and the upload of the job itself. Bambuddy wrote
    an empty fallback archive and never looked again -- then downloaded that same file
    successfully three times over the next two minutes and discarded every copy,
    because the only code that would have attached one had already run.

    The recovery machinery was there. It was armed for exactly one give-up, the FTPS
    cool-off, on the grounds that the three storage verdicts are settled: a job on
    internal eMMC never appears at any FTPS path, and sweeping for it again is what
    where the file is demonstrably still on the card.

    The sweep already had the signal and never used it. A file that is genuinely not
    there is answered with 550, which surfaces as FileNotOnPrinterError and is caught
    by name; a timeout returns falsy instead. So "the printer says no such file" and
    "we never got a straight answer" are distinguishable without guessing, and only
    the second schedules anything.

    Not scheduled either for a 3MF that downloaded fine and turned out to be another
    plate's. Recovery checks that a candidate is a readable 3MF but not which plate it
    holds, and the names a retry would use are the same stale ones that fetched the
    contradicted file -- so it would put back exactly what #2957 discards.

    The ladder follows the cause: a cool-off has to expire, so its first attempt sits
    past the 300s; nothing has to expire here, and this reporter's file completed 48
    seconds after the budget was spent.

    The archives banner gets its own wording for this, because the old text sends an
    owner whose card is working to switch on a setting that is already on. It names
    the Connection Timeout setting instead.
2026-09-20 13:34:10 +02:00
maziggy ad3b295312 fix(archives): let Items Printed go to 0 for a ruined plate (issue #3051)
A jam can destroy everything on the plate while the printer still reports the
    job as a success, so the honest count of usable parts is zero. The edit dialog
    floored the field at one, and a project's completed-items count sums that
    column, so there was no way to record that a job produced nothing.

    The floor was in the dialog only; the API stored whatever it was given, which
    also meant a negative count was accepted and would have subtracted from the
    project totals. The column is now bounded at zero instead.

    Filament Trends counted prints as `quantity || 1`, which would have read a
    deliberate 0 as "unset" and charged the ruined plate as one print while the
    project page counted none.
2026-09-20 13:33:46 +02:00
maziggy 2f7ec240fd fix(ams): stop reading the printer's command acks as status (issue #3040)
Every project_file carried "cfg": "0" — the device-config bitmask, which
    Bambu Studio has never sent and the firmware ignores. The printer echoes a
    command's fields back in its ack, and the ack was ingested as telemetry, so
    bit 18 read as "AMS Filament Backup off" 25 ms after every dispatch.

    Families that repeat cfg in their periodic status (P2S, H2C, X2D) corrected
    themselves a second later; the P1S, A1, A1 Mini and A2L send it only in a
    full status dump, so the wrong value stuck and silently disabled the
    prefer-lowest-remaining gate. The A1 family, which reports no cfg at all and
    is meant to stay "unknown", was pinned to a definite "off".

    Acks are no longer read as status, for the backup bit or the per-job
    timelapse flag they also echo, and cfg is gone from the print command.
2026-09-20 13:32:50 +02:00
maziggy 5ec99a7e06 fix(ams): resolve a slot's K profile by index when the printer does not file per hotend (issue #3044)
An X2D with two AMS 2 Pro, one per hotend, showed a K value on every slot
    of the first and nothing on any slot of the second. Configure Slot was
    worse than blank there: the picker offered no matching profile, the slot
    read as though nothing were bound, and choosing one changed nothing the
    user could see. Both symptoms are one rule.

    A calibration index can mean two different profiles on a dual-nozzle
    machine -- on the maintainer's H2C, index 16 is the left hotend's black
    PLA at K=0.018 and 15 is the right's at K=0.020 -- so the index is
    resolved against the slot's own hotend, and a miss shows nothing rather
    than the other nozzle's number. That is right whenever the printer files
    its calibrations per hotend. This one files them per filament: the second
    AMS's slots point at the same entries as the first, every entry tagged
    with one extruder, and requiring a match found nothing at all.

    The hotend now has to appear in the table the printer actually sent
    before it is used to narrow anything. Where it does not, the index stands
    on its own, which is what BambuStudio does for this same card --
    AMSItem.cpp resolves it through get_pa_k_n_value_by_cali_idx, matching
    cali_idx and nothing else. Where it does, nothing changes: the H2C case
    still blanks rather than borrowing, and the other hotend's profiles stay
    reachable under Other K profiles. The relaxed path still refuses an
    answer when the candidates disagree on a value.

    The premise that the table is always numbered per nozzle had been written
    into three comments and two layers of code; it is corrected where it
    appears.

    Alongside it, in the same picker: the K-profile options rendered the
    hotend suffix twice in the matching group and three times under Other, so
    every option on a dual-nozzle printer read "... . Left . Left".
2026-09-20 13:31:43 +02:00
maziggy 4440c95976 fix(queue): skip preheat entirely when no loaded filament wants a chamber (issue #3041)
Preheat & Heat Soak delayed every PLA print by five to seven minutes and
    gave nothing back. The filament map correctly derived a chamber target of
    0, and the chamber phase correctly skipped -- but the stage then heated
    the bed, waited for it, and held the full soak anyway, because the soak
    had no idea it was holding for a chamber nobody asked for. The print's
    own G-code sets the bed the moment it starts, so the bed phase only moved
    the warm-up ahead of the FTP upload instead of overlapping with it.

    A 0 that comes out of the filament map now skips the stage before any
    command goes out. The one thing the skip still does is put the airduct
    flap back to cooling on the models that have one -- an H2D left in
    heating mode by the ABS job before it would otherwise cook the PLA that
    follows, and that costs one MQTT command and no waiting.

    Explicit instructions are untouched. A chamber target of 0 typed into a
    print's own override still heats the bed and runs the soak, which is what
    the queue documentation has always promised it does, as does forcing a
    print's Preheat override to On. Prints that want chamber heat are
    unaffected, including the P1S/P1P/A1 tier where the bed and the soak
    timer are the whole mechanism.

    The existing unit tests all ran with soak_seconds=0, which is why the
    production default was never exercised; the PLA test now runs at the real
    default and asserts nothing is dispatched and nothing is slept.

    Surfaced in the UI on the way through: the Settings hint claimed the
    derived 0 skipped "the chamber phase", and the per-print chamber override
    field said nothing about a typed 0 meaning bed-only -- a user reaching
    for 0 to turn preheat off got the delay instead.
2026-09-20 13:30:36 +02:00
maziggy ad09406672 fix(slicer): strip zero-valued filament-index sentinels, and sanitise the preview slice too (issue #3030)
Bambu Studio writes 0 into wall_filament, sparse_infill_filament and
    solid_infill_filament to mean "use whichever filament the object is set
    to". Bambu Studio and OrcaSlicer 2.4 define these min 0 and accept it;
    OrcaSlicer 2.3 and earlier used the 1-based scheme (min 1, default 1)
    and reject it with "0 not in range [1.000000,...]". Sidecar images are
    version-tagged, so an install can be pinned to one of those builds.

    Same shape as the -1 inherit markers from #1201 with a different marker,
    so the allowlist becomes a key-to-marker map rather than one global
    constant. The buckets must not bleed: a -1 on a filament index is a real
    value, and a 0 on a raft field is a setting the user chose.

    The key is removed rather than rewritten, which is what makes it safe on
    every build. The CLI then uses its own default: 0 where 0 was legal
    (unchanged), 1 on the older builds, which is what "the active filament"
    means under that scheme.

    The preview slice never ran the sanitiser at all, so a file that sliced
    fine could still fail its automatic plate preview and fall back to the
    painted-face heuristic. It matters more there than in a real slice: the
    preview runs on the file's own embedded settings, so there is no
    --load-settings pass that could supply a replacement for a field the
    range validator has already rejected. That also explains the reported
    "same error on a later attempt of an unchanged file" without any second
    copy of the keys -- the validator that emits it reads the merged global
    config, which per-object model_settings.config overrides never reach.

    The sanitiser moves to utils/threemf_tools so the service can use it
    without importing a route module, and both preview callers pick it up
    from one place. Drops _strip_3mf_embedded_settings and its constant,
    which have had no callers since the strip-everything experiment was
    reverted.
2026-09-20 13:30:11 +02:00
maziggy 417d03d174 fix(slicer): keep protocol-handler download tokens valid for their whole TTL (issue #3029)
The Slice and Open in Slicer actions mint a short-lived token and put it in
    the URL, because a protocol handler cannot carry an Authorization header.
    That token was spent by the first request to reach the endpoint, which made
    the handoff depend on the slicer fetching the URL exactly once. Nothing
    guarantees that: Bambu Studio's downloader retries three times after a
    failed attempt, transfers get resumed, on-access scanners fetch. The first
    request won and the slicer was handed a 403.

    verify_slicer_download_token takes a keyword-only single_use flag. The
    default still consumes via DELETE...RETURNING; single_use=False verifies
    with a SELECT and leaves the row for the rest of its five-minute TTL. The
    stored row is the same either way, so the endpoint decides, not the mint.

    The three protocol-handler downloads pass single_use=False: a library file,
    an archive's sliced 3MF, an archive's source 3MF. Resource binding and
    expiry are untouched. The two browser downloads keep consuming, because
    what they hand over is itself consumed -- the prepared printer bundle is
    deleted the moment it has been streamed.

    Also: add "/source-dl/" to PUBLIC_API_PATTERNS. Those patterns match by
    substring and the source 3MF route's segment is source-dl, which does not
    contain "/dl/", so with auth enabled the middleware rejected the slicer's
    header-less request before the route's token check ran. Open source 3MF in
    slicer could never work on an install with authentication on.
2026-09-20 13:29:25 +02:00
maziggy e1fad9d68f fix(auth): decouple media routes from the camera stream token (issue #3025)
Thirteen routes with nothing to do with a camera took the camera stream
    token as their credential -- library and archive thumbnails, plate
    previews and plate thumbnails, timelapses, print photos, archive QR
    codes, project covers, print-log thumbnails, printer covers and
    external-link icons. A browser cannot put an Authorization header on an
    <img src>, so these need a credential that fits in the URL, and the
    camera token was the only one that existed. Minting one costs
    camera:view, so a user granted library access to their own files got a
    grid of broken images until they were also handed the live camera.

    Adds a media token: minted by POST /auth/media-token behind plain
    authentication, and identified -- it records the principal the way the
    websocket token does rather than being anonymous the way the camera
    token is. Each route now gates on the permission and ownership rules of
    the resource it serves, through the same _ensure_*_visible helpers its
    header-authenticated siblings already use. The three camera routes keep
    the camera token, and require_camera_stream_token_if_auth_enabled now
    documents that it is for those only.

    The media dependencies accept ordinary Authorization / X-API-Key headers
    as well as ?token=, delegating that path to the existing checkers, so
    API-key scope rules and the per-printer allowlist are unchanged.

    Long-lived camera_stream, camwall and overlay tokens are deliberately
    not accepted on the media routes -- those are handed to kiosks, walls
    and Home Assistant to display video. The cam wall, streaming overlay and
    kiosk views use only the three camera routes and are unaffected.

    Frontend: withMediaToken alongside withStreamToken, and
    useStreamTokenSync fetches a media token for every signed-in user while
    asking for a camera token only when the user can mint one, which also
    stops the 403 that fired on every page load for everyone else.

    Also fixed, same class:
    - /printers/{id}/files/plate-thumbnail/{i} is rendered in an <img> but
      had a header-only guard, so the file manager's plate thumbnails 401'd
      whenever auth was enabled. It now takes a media token too.
    - getProjectCoverImageUrl returned a URL ending in ?token=, and the
      project edit dialog appended its own ?v= cache-buster after it, so the
      second ? landed inside the token value. The version is now a parameter
      applied before the token.

    Tests: 15 integration tests for the token boundary, permission
    enforcement and per-row scoping; 10 frontend tests for the URL split and
    the two-query hook. test_cover_image_get_uses_stream_token_gate is
    renamed and repointed at the media gate -- what it pins, that the
    credential has to fit in a URL, is unchanged.
2026-09-20 13:28:50 +02:00
maziggy ef6446d30a fix(auth): let the sidebar read install flags without settings:read (issue #3023)
cost_centers:read_own exists so a non-admin can see their own wallet, balance
    and cost-centre spend, and the Finance page honoured it -- typing the URL
    worked and rendered their balance. The sidebar never offered the entry.

    It decides whether to show Finance by reading billing_enabled from
    GET /settings, which requires SETTINGS_READ. A non-admin gets 403 there, so
    the value arrived undefined, `undefined !== true` held, and the entry was
    hidden from precisely the users the permission was written for. The permission
    map and the route guard were both already right; only discovery was broken.

    Three more fields came from that same 403, and one of them failed the other way
    up. The Notifications gate tests `=== false`, which undefined never satisfies,
    so an administrator who switched user notifications off still left the entry
    showing to the non-admins it governs. Nobody reported that one, and no
    administrator could have reproduced either: administrators can read /settings.
    The remaining two were quieter -- the sponsor prompt fell back to EUR whatever
    the install uses, and the update check ran where it had been turned off.

    SETTINGS_READ cannot be the price of knowing whether billing is on. It also
    grants sight of the SMTP, LDAP and MQTT credentials, which is the reason
    /settings/ui-preferences exists at all.

    So: a second endpoint, GET /settings/ui-flags, carrying those four fields and
    asking only that the caller be signed in, via the existing
    require_auth_if_enabled. Layout drops its /settings query altogether, which
    closes the class rather than the two instances that happened to be visible.

    Deliberately not four more fields on /ui-preferences. That endpoint is served
    to anyone at all on the recorded grounds that its contents are "public defaults
    that ship with the app" (test_route_auth_coverage.py), and its field set is
    pinned by a test written to make anyone adding to it stop and think. These
    fields are not defaults -- they say how this deployment is configured -- so
    they get their own endpoint at their own trust level instead of stretching that
    charter to fit them. require_auth_if_enabled also keeps the auth-disabled case
    that /ui-preferences was ungated for: "works when there is no auth" and
    "readable by anyone" are different statements, and conflating them is what put
    a settings read in front of a permission that never needed one.

    Twelve tests. Backend pins that the operator can read the flags, that the same
    operator still gets 403 from /settings, that an anonymous caller is refused
    when auth is on, that it answers when auth is off, the exact field set, that no
    credential ever appears, and that the public endpoint did not quietly gain
    these fields. Frontend pins Finance visible for cost_centers:read_own with
    /settings returning 403, and Notifications hidden when the flag is off -- each
    waiting on a positive signal before asserting an absence, so the negative cases
    cannot pass before the query resolves.

    Reported by @lonix, who traced it to the queryKey and the route gate.
2026-09-20 13:28:17 +02:00
maziggy 23a6633f42 fix(queue): say when an unscheduled item runs instead of calling it ASAP (issue #3018)
The print dialog offers ASAP, Queue and Schedule. ASAP and Queue differ only
    in where the item is inserted, and neither is stored on the item -- scheduleType
    is a frontend-only concept, and grep finds no "asap" anywhere in the backend. So
    the queue's time column had nothing to read but scheduled_time, and labelled
    every unscheduled item "ASAP": the name of the one mode the user may well have
    chosen against.

    Someone who picked Queue then watched their row appear as ASAP and start
    immediately, and concluded Bambuddy had overridden them. Two reporters wrote
    that same sentence thirteen months apart, and #2557 was closed as A2L-specific
    after the first of them -- kilrah replied there with an X1C before filing this.

    The column answers when an item runs, so it now says that. The key is renamed
    whenFree rather than just retranslated: left called asap, the next translator
    puts ASAP back.

    The dispatch is unchanged, because it was right. A print scheduled for later
    does not reserve the printer until then; an unscheduled item behind it uses the
    idle printer rather than leaving an X1C dark until 6 AM. Two of the new tests
    pin that, so it does not get "fixed" later on the strength of a report like this
    one.

    What genuinely could not answer the question was the queue's own log. Its
    per-printer line called every entry in busy_printers "not available" -- but that
    set holds both printers that cannot take work and printers the pass has just
    claimed for some, which are opposite facts. It also read printer state at
    logging time rather than at the decision, so #3018's bundle carries

        Queue: printer 1 not available — connected=True, state=IDLE, ...
        Launching 1 upload(s) (pool 0/4 in flight)
        Starting queue item 18

    a printer reported unavailable, evidence that it was available, and a dispatch
    to it, in three consecutive lines. It is the first line anyone greps for "why
    did my item not go out".

    Each of the nine sites that removes a printer from a pass now records why, and
    the summary reports a claim as a reservation and everything else as an
    obstruction with its reason. The live fields stay, since a bundle reader wants
    them next, but are labelled as read now rather than offered as the cause.
    print_scheduler.py:1210 already documented that these two meanings differ -- the
    dispatching_printers snapshot exists for it. This carries that distinction into
    the log.
2026-09-20 13:27:43 +02:00
maziggy eab55cef75 fix(archives): report a refused FTPS handshake as the printer, not the slicer (issue #2780)
The Archives banner picks its wording from a priority list of the causes it
    knows. REASON_FTPS_COOLOFF was added by #2957 and never put in that list, so an
    install whose empty archives all came from a printer refusing the TLS handshake
    matched nothing, got reason: null, and fell to the original wording: the slicer
    did not leave the .gcode.3mf on the card, switch on "Store sent files on
    external storage", here is installation step 4.

    Every clause of that is wrong for this cause. The slicer did write the file --
    reason he read the whole thing as Bambuddy being broken. The setting was
    already on. And there is nothing on his side to change: the printer's file
    service answered port 990 with something that is not TLS, so no lookup ever
    ran and where the file went was never tested. It is #2899's mistake -- an
    error message describing a cause that was ruled out before it was printed --
    in a surface that did not get that pass.

    The slug now leads the list rather than joining the end of it. The other three
    describe an install working as configured and each ends in something the
    operator can change; this one reports a fault nobody can yet explain, which is
    both the more urgent thing to say and the thing that produces a useful report.
    The banner also dismisses one-shot into localStorage, so a reason ranked below
    another is not deferred to next time -- it is never shown to that user again.
    Ranking it first cannot bury a permanent cause in exchange: a successful
    recovery clears the row's markers (#2957), so a row still carrying this slug is
    one whose retry failed too, days after the print.

    New wording in all fourteen languages says the printer refused the connection,
    that this is not a slicer setting and not something the operator did, that
    Bambuddy comes back for the file when the five-minute pause clears so a brief
    episode fills itself in, and that a card still empty means the refusal outlasted
    the retry. It links to the handshake entry in the troubleshooting guide instead
    of to the installation guide.

    The client's getNo3MFWarning type still declared the old three-slug union, which
    made all three new comparisons provably dead -- caught by tsc, not by any test.

    Four tests. One pins the slug reaching the banner, one pins it outranking the
    three settled causes, one pins those three keeping their order behind it, and
    one asserts the rendered wording carries no slicer advice at all.

    Also corrects the wiki page these reports are pointed at. It said to power-cycle
    the printer; the reporter who prompted that advice power-cycled both of his and
    the failure continued unchanged, and bambu_ftp.py has carried the retraction in
    a comment since. The page now states what was actually measured -- that a
    version mismatch reports itself differently, that every printer probed refuses
    TLS 1.3 and completes on 1.2 so there is no version to fall back from, and that
    three P2S units failed while three more on the same switch never did -- says
    plainly that the trigger is unknown, and names the one cleartext-probe line
    worth collecting.
2026-09-20 13:27:10 +02:00
maziggy 58df1cb866 feat(ftp): log how every FTP session closes (issue #3009)
disconnect() and _abandon_connection() logged nothing, at any level. A
    session closed cleanly and a socket genuinely abandoned therefore produced
    identical output -- none -- and the only way to tell them apart was to read
    the source.

    That is how #3009 was filed. Its trace shows a print completion opening two
    FTP connections, deleting one file, and then nothing until the printer was
    powered off 21 minutes later, read as connections left open and offered as a
    mechanism for the 0500-C010 SD-card error that #645 has been chasing since
    April. The two connections are the post-print SD cleanup in main.py walking
    its candidate filenames, each through delete_file_async, which closes in a
    finally; running that against the mock FTPS server shows the server logging
    "FTP session closed (disconnect)" for both the 250 and the 550, holding zero
    sessions afterwards. Nothing in a support bundle could have shown that.

    Both close paths now log one DEBUG line: the printer, whether QUIT was
    acknowledged or the socket had to be dropped without it, why, and how long
    the session was held. Every connect in a debug log now has a matching close.

    The duration comes from a stamp taken when the control socket opens rather
    than after login, so a session that dies during login is accounted for too;
    where no socket was ever established the line says "held unknown" rather
    than claiming a number. The four connect() failure paths pass their own
    reason, so a close line stands on its own next to the warning above it.

    Nine tests, seven of which fail against the unlogged version. The other two
    assert silence -- a bare disconnect(), and a connect skipped by the handshake
    cool-off -- where no socket was opened and a close line would pair with no
    connect.
2026-09-20 13:26:42 +02:00
maziggy 8186ef817e fix(orca-cloud): close the HTTP client when an authenticated build fails
OrcaCloudService owns an httpx client from construction, and every path in
    _build_authenticated_service after that point can raise: no stored refresh
    token, a rejected refresh, an unreachable Orca, and the token-rotation write.
    On success the caller closes the client. On failure nobody is ever handed it,
    so all four paths leaked one into the connection pool.

    That went unnoticed while the only callers were routes, where the trigger is a
    person retrying a broken sign-in a handful of times. It stopped being harmless
    in 9434875f, which added a caller in spool assignment -- one build per
    Orca-referenced spool, failing on every assignment for as long as the stored
    credentials cannot be refreshed.

    The unwind guard catches BaseException rather than Exception: a cancelled
    request leaks the client just as surely as a failed refresh, and cancellation
    during shutdown is exactly when dangling sockets are least welcome. The close
    inside it is guarded in turn, so a failing cleanup cannot replace the error the
    caller needs to see -- least of all a CancelledError, which has to keep
    propagating for cancellation to work at all.

    Six tests. Four fail against the unguarded builder, verified by reverting the
    guard and re-running; the other two pin the surrounding contract (a failing
    close must not mask the real error, and a successful build must leave the
    client open for its caller) and pass either way. The shared _expired_service
    helper now gives the mock an awaitable close(), so the four pre-existing
    refresh tests exercise the same path.

    Also corrects two comments and the changelog entry from 9434875f, which
    overstated what the captures support. They claimed Bambu Cloud returns a
    preset's filament_id in either of two places and only one was read. The
    responses recorded in #1053 show something narrower: a Studio-created preset
    carries it on the envelope, and an Orca-created one has none at all -- the
    envelope says null and `setting` is a delta from the base. The `setting` lookup
    stays as belt-and-braces for a shape no captured response has needed yet, but
    it is not why a custom profile reached the slicer as its base. That is the
    OrcaSlicer preset format having no filament_id field, filed upstream as
    OrcaSlicer PR #13315.

    The eight-character truncation is now evidenced across three models rather than
    one -- an A1 storing PFUS9DDC of PFUS9DDC938FE3AB8F, a P1S storing PFUS7A65 of
    PFUS7A65290D3DADC4, and an H2D storing 8219C45D of an Orca profile UUID.
2026-09-20 13:26:02 +02:00
maziggy a3e6fae5cb fix(ams): resolve a custom filament's own id from every preset source (issue #3003)
A custom filament profile reaches an AMS slot as itself through exactly one
    field, tray_info_idx, and every source we can read that id from was reading it
    from the wrong place or not reading it at all.

    Bambu Cloud returns a preset's own filament_id either on the response envelope
    or inside the preset JSON under `setting`, and only the envelope was read.
    Presets of the second shape fell through to the base_id branch and reached the
    slicer as the Bambu filament they inherit from. filament_type next door already
    handled both spreads; filament_id now does too.

    Orca Cloud was absent from the resolver entirely. A spool stores the bare
    profile UUID, which matched no branch and fell through normalize_slicer_filament
    -- a function that passes anything it does not recognise straight through -- so
    a 36-character UUID went into the field. Orca profiles carry their own
    filament_id in the slicer JSON that OrcaProfileDetail already exposes under
    `setting`, so the lookup is the same one the Bambu branch does. It is
    best-effort: no pairing, a dead token or a missing orca_cloud:auth permission
    degrades to the fallback rather than failing the assignment, and it passes
    clear_on_auth_failure=False because a background caller cannot tell a real
    revocation from a lost refresh-rotation race.

    configure_ams_slot sent the cloud setting_id as tray_info_idx when it found no
    real filament id. That field is 8 characters on the printer -- exactly the width
    of a local preset id, less than half a cloud one. Measured on the reporter's A1:
    sent PFUS9ddc938fe3ab8f, the tray read back PFUS9DDC, acknowledged as a success.
    The slot then resolved to nothing, so the slicer showed Generic anyway and the
    calibration table, keyed by the same field, lost the slot. It now falls back to
    the slot's existing filament id or the generic for the material, and the route's
    guard was aligned with the resolver's so both refuse the same four shapes from
    one shared definition.

    This reverses the contract #1053 pinned. Six tests asserted that the PFUS
    belonged in tray_info_idx; the A1 capture shows it never worked, so they were
    rewritten with the measurement in their docstrings.

    Verified against 874 AMS trays across twelve models in the support archive: 92
    already carry a custom "P" + 7 hex filament id, which is what confirms the
    mechanism works and this is a lookup failure rather than a platform limit. No
    tray on any model carries a setting_id, so a profile with no filament_id of its
    own still cannot be told apart from its base.
2026-09-20 13:25:27 +02:00
maziggy 7490c93081 ci: balance the backend test shards by measured time, not test count
Backend Tests (shard 1/4) timed out after 10 minutes on e8b901f54
    ("Updated BACKERS"), a docs-only commit. The step was not hung: it
    reached 98%, every test passing, and was killed about 12 seconds short
    of finishing.

    pytest-split balances by duration only when it has a durations file,
    and there was none. Without one it splits by test COUNT -- 2926 / 2926
    / 2926 / 2925, exactly 25% each, which is what the comment here claimed
    was good enough. Count is not time. Measured over the full suite
    (11703 tests, 1045.6s):

        shard 1   2926 tests   658.2s
        shard 2   2926 tests   173.1s
        shard 3   2926 tests   113.3s
        shard 4   2925 tests   101.1s

    Shard 1 was carrying 63% of the suite's runtime -- a 6.5x spread -- and
    had been walking toward the cap for weeks: 475s, 404s, 445s, then 587s
    on 08-30, thirteen seconds under it, and 613s here. CI printed the
    reason in its own log on every run: "[pytest-split] No test durations
    found."

    Commit tests/.test_durations and pass --durations-path explicitly:

        shard 1   1273 tests   261.6s
        shard 2   1070 tests   261.5s
        shard 3   1522 tests   262.0s
        shard 4   7838 tests   260.6s

    A 1.01x spread. The lopsided test counts are the point: shard 4 takes
    thousands of fast unit tests, shard 1 keeps the slow integration ones.
    The four groups still partition the suite exactly -- union is 11703
    with nothing dropped or duplicated, and all four pass.

    --durations-path has to be explicit because pytest-split defaults it to
    $CWD/.test_durations, and the two matrices run from different
    directories: the native job from backend/, the Docker job from the
    image root. Left implicit, the Docker one finds nothing and silently
    falls back to the count split. Node IDs match across both because
    backend/tests/pytest.ini pins rootdir to backend/tests either way, so a
    single file serves them both; verified the file clears .dockerignore
    and lands at /app/backend/tests/.test_durations at full size.

    timeout-minutes 10 -> 15 for headroom, since the file goes stale as
    tests are added. Staleness degrades slowly rather than breaking --
    unknown tests are treated as average.
2026-09-20 13:22:17 +02:00
maziggy 5584dca898 Bumped version 2026-09-20 13:19:02 +02:00