The concurrent-dispatch tests went red on the Docker shard with
"expected all 6 printers to be uploaded to concurrently, but the
high-water mark was 4", on a scheduler that was dispatching all six
correctly.
peak == 6 is the claim that the sixth dispatch reaches its upload
before the first one finishes, and the dispatches do not arrive
together: each runs a preamble of database work first. So the
assertion was a race between that spread and a fixed 0.15 s sleep.
Measured here the spread is ~14 ms; on the runner it passed 150 ms.
Padding the sleep would only move the threshold and slow every test
that uses it.
_UploadRecorder(assemble=True) now holds each call until every
upload the pass launched has arrived, reading len(_inflight), which
is filled synchronously at launch and is therefore the batch size.
That states the property the peak assertions are about, directly and
with no time in it, and it ends sooner than the sleep it replaces -
the file runs 7.0s to 5.9s. _BATCH_DEADLINE_SECONDS bounds it so a
dispatch that really has gone serial fails on its peak assertion
instead of hanging to the pytest timeout.
test_check_queue_returns_without_awaiting_the_uploads keeps the
sleep, with the reason recorded next to it: it reads the rows while
the uploads are open, and an assembled batch releases as soon as it
is complete, which would let the dispatches flip those rows
mid-assertion.
The test engine also takes the application's own _set_sqlite_pragmas
listener rather than a copy, so the file is opened the way the
running system opens one - WAL, synchronous = NORMAL, a 15 s busy
timeout. SQLite's defaults spend an fsync per commit and lock the
whole file, which made a dispatch's preamble cost more than the
upload it precedes. That is what turned the previous commit's move
to a file-backed database into a visible failure.
Verified in both directions rather than by a green run: with two
printers' preambles delayed by a second, the old recorder reports
exactly the "high-water mark was 4" the runner saw and the shared
library test reports 3 == 4, while the assembling recorder passes
the same scenario. Without the induced skew, ruff is clean, the full
backend suite is green at 12224 passed, and ten consecutive runs of
the file pinned to one core show no failures.
The scheduler's concurrent-dispatch tests failed on CI and passed
locally, on a commit that touched nothing but a text file. Six tests
in tests/unit/test_scheduler_concurrent_dispatch.py went red with
every queue item logged as "Status set to 'printing'" and four of
them read back as "pending", alongside "cannot commit transaction -
SQL statements in progress" and a PrintArchive that could not be
refreshed.
The fixtures built their farm on sqlite+aiosqlite:///:memory:, and
SQLAlchemy backs an in-memory SQLite with a StaticPool: one DBAPI
connection handed to every session, with nothing keeping them apart.
That was harmless while check_queue awaited its uploads inline,
because only one session was ever live at a time. Under the
refillable upload pool the uploads run as concurrent background
tasks with a session each, so their transactions interleave on that
single connection - a sibling session's close() rolls back another's
flushed-but-uncommitted UPDATE, and a commit() landing while another
session still holds a cursor raises the commit error above. Whether
the interleaving lands badly comes down to core count and
interpreter version, which is why a 30-core box on 3.13 stayed green
and a 4-vCPU runner on 3.11 did not.
The four fixtures now put the database in the test's own tmp_path,
which gets an AsyncAdaptedQueuePool and a connection per session -
what the application itself runs with (_resolve_pool_kwargs in
backend/app/core/database.py: pool_size 20, max_overflow 200). So
the harness is being brought in line with production rather than
having its assertions relaxed; nothing in the scheduler changes,
and the concurrency under test was correct throughout.
Checked against the failure mode rather than against a green run: an
isolated repro of the same shape - six writer sessions and one
reader session closing mid-transaction - yields all-pending on the
in-memory engine and all-printing on a file. The new risk is real
SQLite write contention on one file, so the file was run twelve
times pinned to two cores with no failures and no lock errors.
tests/unit/test_scheduler_busy_reasons_3018.py carries the same
harness on an in-memory engine, but dispatches to a single printer,
so it has no second session to race; left as is.
Bambu sends two humidity fields that are not the same quantity.
humidity_raw is relative humidity in percent; humidity is a 1-5 drop
index, and it runs the other way -- OpenBambuAPI's push_info sample
pairs "humidity:30%" with "humidity_idx:4", so a high index means dry
where a high percentage means wet.
Four call sites used the index whenever no percentage arrived. A unit
sending only the index therefore rendered as "2%" in the green band
while being the second-wettest of the five steps, charted an average of
index values as a percentage, and sat under every humidity threshold
forever, since no index can reach one -- the alarm and auto-drying could
not fire for such a unit at all.
- utils/ams_humidity: one leaf helper, a percentage or None. The index
is never converted; None is what every caller already handles.
- routes/printers, printer_manager, print_scheduler, main, bambu_mqtt:
all five readings go through it, so the card, the websocket, the
chart, the alarm and auto-drying cannot answer differently.
- main: a unit that reports the index and no usable percentage says so
once per unit in the log, with its firmware versions requested. No
supported printer is known to do this, and "known" is doing work
there -- the alternative is a card that goes blank with no trace.
Three faults found while checking what else those paths touched:
- main: humidity_raw=float(x) if x else None stored NULL for a numeric
0% while writing 0.0 to humidity on the same row.
- main: that same expression was unguarded, unlike the parse above it,
so a non-numeric humidity_raw raised inside record_ams_history and
aborted the pass for every printer, not just the one that sent it.
- routes/ams_history: the averages were tested for truthiness, so a
window averaging exactly 0 reported no average while the min and max
beside it reported 0.0.
An affected unit now reports no humidity rather than a number that means
the opposite: the indicator is hidden, the chart leaves a gap, the alarm
and auto-drying skip the unit. Temperature is untouched. Auto-drying's
outcome is unchanged either way -- an index could never cross the
threshold -- so only the intent moves.
No supported printer is known to be affected; the report came from an
install running X1Plus, which Bambuddy does not support. Verified
against 7927 recorded samples from seven AMS units including an AMS-HT:
not one used the fallback. Two percentages that did fall through to the
index no longer do -- a reading with a decimal point, and "38.0", which
int() rejected.
A slicer holding the CA of two Bambuddy installs could only connect to
one of them. Each CA worked on its own; together, one stopped, with the
generic "Connect ... failed! [SN:..., code=-1]" that an install whose
CA was never imported gives.
Every install signed as exactly CN=Virtual Printer CA. A slicer's trust
store is a flat list of certificates and OpenSSL resolves an issuer by
Subject DN: it takes the first authority whose name matches and fails
the chain when that one turns out not to have signed the certificate,
rather than trying the next match. Whichever CA landed second in the
file lost -- decided by nothing but the order they were appended in.
Reproduced with openssl verify against a bundle holding two CAs: the
first leaf verifies, the second fails with "certificate signature
failure".
- certificate.py: a newly generated CA takes a suffix from its own key
identifier (CN=Virtual Printer CA D55808BE) and publishes that
identifier, which the printer certificate points back at.
- Existing CAs are untouched, so nothing has to be re-imported. A
printer certificate signed by one keeps exactly the shape it has
today: the authority key identifier is added only when the CA has an
identifier to name.
- tests: unique names per install, the identifier reaching the leaf,
an existing CA being reused unchanged, and both chains verifying
through openssl from a single trust store.
The collision goes away as soon as one of the two CAs is newer than
this change. Two installs that both predate it still collide until one
has its bbl_ca.crt/.key deleted and regenerated, which is a re-import
for that one -- documented in the wiki.
The per-event Priority header from #990 never reached ntfy. The dialog
builds its rows from the provider's event toggles and stores the map
under those names -- on_print_complete, on_print_failed -- while every
sender is called with the bare event name, print_complete. The lookup
missed for all 18 events the dialog offers, so every notification went
out at the ntfy server's default with the configured priority sitting
untouched in the database.
Both ends looked healthy, which is why it shipped. The stored config
held exactly what was set, and the tests were green because they called
_send_ntfy directly with the prefixed name -- the one spelling the
running system never produces.
- notification_service.py: accept either spelling, bare first, so
existing configs keep working and nothing needs migrating.
- schemas/notification.py: document both key forms, and which one the
UI writes.
- tests: use the bare names, and add one that runs from a finished
print through to the outgoing request. Without the fix it fails on a
header dict holding only Title, which is the assertion that was
missing.
The daily digest is unchanged: send_digest sends with no event_type at
all, and the dialog offers no priority for it -- it is one message for
several events.
print_archives.library_file_id -> library_files.folder_id ->
library_folders.archive_id -> print_archives. Three nullable SET NULL
links, each reasonable alone, that together made a loop
metadata.sorted_tables could not sort: it dropped those edges, warned on
every backup and every restore, and could return an order placing a
child before its parent -- which once imported library_files ahead of
library_folders and killed a restore on a ForeignKeyViolation.
The restore no longer depends on that order (it strips every foreign key
before importing and adds them back after), but the backup export sorts
the same way, and the warning ends with "may raise an error in a future
release" -- which would break backup and restore on one upgrade.
Marking one edge use_alter removes it from the sort graph, not from the
database: PostgreSQL emits it as ALTER TABLE ADD CONSTRAINT, as it
already did for every constraint on these three tables, and SQLite
inlines it into CREATE TABLE, so ON DELETE SET NULL holds on both.
Verified against PostgreSQL 16 and SQLite.
The post-slice thumbnail loaded the sliced 3MF with trimesh, whose reader
mishandles the Bambu Studio / OrcaSlicer layout: for every component that
references a file in 3D/Objects/ it re-parses that file and appends all
of its meshes again. N copies of a part came back as N^2 copies of its
triangles, and each component carried every other part of the same file.
25 bins of 10k faces loaded as 6.4M faces; the render took 54 s and
8.4 GB on the event loop, and the server was OOM-killed.
Parse the 3MF directly (lxml, entities/DTD/network off, streamed and
freed element by element), walk objects, components and build items
(p:path on either) with their transforms, and keep each mesh once.
Decimate per unique mesh to its share of a face budget before placing
instances, flip winding on mirrored placements, and skip the thumbnail
above a face ceiling checked both on what decimation can reach and on
what it delivered.
Both slice routes run the render in a thread. The renderer uses
matplotlib's Figure/Agg API instead of pyplot, whose process-global
figure state let a threaded plate render and stl_thumbnail's
event-loop render lay out and close each other's figures.
Switching an "Any P2S" job to a specific P2S cleared its filament
override: "Specific Printer" empties the target model and the reset
effect counted that as a model change. Printer mode also matched trays
against the 3MF's colours, never sent the override, and left the old one
on the row.
The reset now compares against the last model actually targeted, so the
switch keeps the override while a real model or plate change still clears
it. Printer-mode tray matching (single, per-plate, multi-printer and the
selector's per-printer editor) runs against the requirements with the
overrides applied, mirroring the scheduler's _apply_filament_overrides; an
entry naming the slot's own filament is not a swap and keeps its
tray_info_idx. Printer-mode submits carry the user's overrides, and the
create endpoint stores them for a printer-targeted job, so a dispatch-time
recompute of an unresolved mapping looks for the same filament.
Saving re-attaches the tray_info_idx an unchanged entry already had, so a
virtual printer's force-colour PLA-variant pin (#2650) survives an edit in
either assignment mode. The printer card's compatibility filter skips
printer-targeted jobs: it mirrors the model scheduler, and hiding a job on
filament would hide it from the printer it is going to run on.
Bambu sends two humidity fields that are not the same quantity.
humidity_raw is relative humidity in percent; humidity is a 1-5 drop
index, and it runs the other way -- OpenBambuAPI's push_info sample
pairs "humidity:30%" with "humidity_idx:4", so a high index means dry
where a high percentage means wet.
Four call sites used the index whenever no percentage arrived. A unit
sending only the index therefore rendered as "2%" in the green band
while being the second-wettest of the five steps, charted an average of
index values as a percentage, and sat under every humidity threshold
forever, since no index can reach one -- the alarm and auto-drying could
not fire for such a unit at all.
- utils/ams_humidity: one leaf helper, a percentage or None. The index
is never converted; None is what every caller already handles.
- routes/printers, printer_manager, print_scheduler, main, bambu_mqtt:
all five readings go through it, so the card, the websocket, the
chart, the alarm and auto-drying cannot answer differently.
- main: a unit that reports the index and no usable percentage says so
once per unit in the log, with its firmware versions requested. No
supported printer is known to do this, and "known" is doing work
there -- the alternative is a card that goes blank with no trace.
Three faults found while checking what else those paths touched:
- main: humidity_raw=float(x) if x else None stored NULL for a numeric
0% while writing 0.0 to humidity on the same row.
- main: that same expression was unguarded, unlike the parse above it,
so a non-numeric humidity_raw raised inside record_ams_history and
aborted the pass for every printer, not just the one that sent it.
- routes/ams_history: the averages were tested for truthiness, so a
window averaging exactly 0 reported no average while the min and max
beside it reported 0.0.
An affected unit now reports no humidity rather than a number that means
the opposite: the indicator is hidden, the chart leaves a gap, the alarm
and auto-drying skip the unit. Temperature is untouched. Auto-drying's
outcome is unchanged either way -- an index could never cross the
threshold -- so only the intent moves.
No supported printer is known to be affected; the report came from an
install running X1Plus, which Bambuddy does not support. Verified
against 7927 recorded samples from seven AMS units including an AMS-HT:
not one used the fallback. Two percentages that did fall through to the
index no longer do -- a reading with a decimal point, and "38.0", which
int() rejected.
A slicer holding the CA of two Bambuddy installs could only connect to
one of them. Each CA worked on its own; together, one stopped, with the
generic "Connect ... failed! [SN:..., code=-1]" that an install whose
CA was never imported gives.
Every install signed as exactly CN=Virtual Printer CA. A slicer's trust
store is a flat list of certificates and OpenSSL resolves an issuer by
Subject DN: it takes the first authority whose name matches and fails
the chain when that one turns out not to have signed the certificate,
rather than trying the next match. Whichever CA landed second in the
file lost -- decided by nothing but the order they were appended in.
Reproduced with openssl verify against a bundle holding two CAs: the
first leaf verifies, the second fails with "certificate signature
failure".
- certificate.py: a newly generated CA takes a suffix from its own key
identifier (CN=Virtual Printer CA D55808BE) and publishes that
identifier, which the printer certificate points back at.
- Existing CAs are untouched, so nothing has to be re-imported. A
printer certificate signed by one keeps exactly the shape it has
today: the authority key identifier is added only when the CA has an
identifier to name.
- tests: unique names per install, the identifier reaching the leaf,
an existing CA being reused unchanged, and both chains verifying
through openssl from a single trust store.
The collision goes away as soon as one of the two CAs is newer than
this change. Two installs that both predate it still collide until one
has its bbl_ca.crt/.key deleted and regenerated, which is a re-import
for that one -- documented in the wiki.
Reported by @Steven-Pierce.
The per-event Priority header from #990 never reached ntfy. The dialog
builds its rows from the provider's event toggles and stores the map
under those names -- on_print_complete, on_print_failed -- while every
sender is called with the bare event name, print_complete. The lookup
missed for all 18 events the dialog offers, so every notification went
out at the ntfy server's default with the configured priority sitting
untouched in the database.
Both ends looked healthy, which is why it shipped. The stored config
held exactly what was set, and the tests were green because they called
_send_ntfy directly with the prefixed name -- the one spelling the
running system never produces.
- notification_service.py: accept either spelling, bare first, so
existing configs keep working and nothing needs migrating.
- schemas/notification.py: document both key forms, and which one the
UI writes.
- tests: use the bare names, and add one that runs from a finished
print through to the outgoing request. Without the fix it fails on a
header dict holding only Title, which is the assertion that was
missing.
The daily digest is unchanged: send_digest sends with no event_type at
all, and the dialog offers no priority for it -- it is one message for
several events.
print_archives.library_file_id -> library_files.folder_id ->
library_folders.archive_id -> print_archives. Three nullable SET NULL
links, each reasonable alone, that together made a loop
metadata.sorted_tables could not sort: it dropped those edges, warned on
every backup and every restore, and could return an order placing a
child before its parent -- which once imported library_files ahead of
library_folders and killed a restore on a ForeignKeyViolation.
The restore no longer depends on that order (it strips every foreign key
before importing and adds them back after), but the backup export sorts
the same way, and the warning ends with "may raise an error in a future
release" -- which would break backup and restore on one upgrade.
Marking one edge use_alter removes it from the sort graph, not from the
database: PostgreSQL emits it as ALTER TABLE ADD CONSTRAINT, as it
already did for every constraint on these three tables, and SQLite
inlines it into CREATE TABLE, so ON DELETE SET NULL holds on both.
Verified against PostgreSQL 16 and SQLite.
The post-slice thumbnail loaded the sliced 3MF with trimesh, whose reader
mishandles the Bambu Studio / OrcaSlicer layout: for every component that
references a file in 3D/Objects/ it re-parses that file and appends all
of its meshes again. N copies of a part came back as N^2 copies of its
triangles, and each component carried every other part of the same file.
25 bins of 10k faces loaded as 6.4M faces; the render took 54 s and
8.4 GB on the event loop, and the server was OOM-killed.
Parse the 3MF directly (lxml, entities/DTD/network off, streamed and
freed element by element), walk objects, components and build items
(p:path on either) with their transforms, and keep each mesh once.
Decimate per unique mesh to its share of a face budget before placing
instances, flip winding on mirrored placements, and skip the thumbnail
above a face ceiling checked both on what decimation can reach and on
what it delivered.
Both slice routes run the render in a thread. The renderer uses
matplotlib's Figure/Agg API instead of pyplot, whose process-global
figure state let a threaded plate render and stl_thumbnail's
event-loop render lay out and close each other's figures.
Switching an "Any P2S" job to a specific P2S cleared its filament
override: "Specific Printer" empties the target model and the reset
effect counted that as a model change. Printer mode also matched trays
against the 3MF's colours, never sent the override, and left the old one
on the row.
The reset now compares against the last model actually targeted, so the
switch keeps the override while a real model or plate change still clears
it. Printer-mode tray matching (single, per-plate, multi-printer and the
selector's per-printer editor) runs against the requirements with the
overrides applied, mirroring the scheduler's _apply_filament_overrides; an
entry naming the slot's own filament is not a swap and keeps its
tray_info_idx. Printer-mode submits carry the user's overrides, and the
create endpoint stores them for a printer-targeted job, so a dispatch-time
recompute of an unresolved mapping looks for the same filament.
Saving re-attaches the tray_info_idx an unchanged entry already had, so a
virtual printer's force-colour PLA-variant pin (#2650) survives an edit in
either assignment mode. The printer card's compatibility filter skips
printer-targeted jobs: it mirrors the model scheduler, and hiding a job on
filament would hide it from the printer it is going to run on.
Selecting sliced files for two printer models and asking for 25 copies
queued one item. The queue emptied as soon as it dispatched and the
Batches tab stayed empty, because no batch is created at quantity 1.
A multi-plate file moves the run count off the modal's Quantity field
onto a stepper beside each plate (#342), hiding the field. The
cross-model submit (#671) posts that field, which in this combination
nothing can set, so it stayed at its initial 1. The modal read "19 runs
in total" above a button that queued one.
Per-plate steppers do not fit a cross-model job: its plate is chosen per
candidate, in the alternatives list, so there is one number to give.
Exclude cross-model from the per-plate mode and the global field comes
back.
Drop the plate selector in that mode too. Its choice never reached the
request; it only keyed the filament-requirements query, so picking plate
3 for a candidate while plate 1 stayed ticked above produced overrides
computed from a plate the job would not print. That query now follows
the primary file's own dropdown.
Dispatch needed nothing -- it already gives each copy its own candidate
rows -- but naming did. A cross-model job carries neither archive_id nor
library_file_id, because the candidates are the files, so both branches
that name a batch missed and every such order would have read "Batch" in
the tab the reporter went looking in. Name it after the first candidate.
The existing cross-model tests all mock a single-plate file, which is
why the pair was never covered; the multi-plate case is added.
Selecting sliced files for two printer models and asking for 25 copies
queued one item. The queue emptied as soon as it dispatched and the
Batches tab stayed empty, because no batch is created at quantity 1.
A multi-plate file moves the run count off the modal's Quantity field
onto a stepper beside each plate (#342), hiding the field. The
cross-model submit (#671) posts that field, which in this combination
nothing can set, so it stayed at its initial 1. The modal read "19 runs
in total" above a button that queued one.
Per-plate steppers do not fit a cross-model job: its plate is chosen per
candidate, in the alternatives list, so there is one number to give.
Exclude cross-model from the per-plate mode and the global field comes
back.
Drop the plate selector in that mode too. Its choice never reached the
request; it only keyed the filament-requirements query, so picking plate
3 for a candidate while plate 1 stayed ticked above produced overrides
computed from a plate the job would not print. That query now follows
the primary file's own dropdown.
Dispatch needed nothing -- it already gives each copy its own candidate
rows -- but naming did. A cross-model job carries neither archive_id nor
library_file_id, because the candidates are the files, so both branches
that name a batch missed and every such order would have read "Batch" in
the tab the reporter went looking in. Name it after the first candidate.
The existing cross-model tests all mock a single-plate file, which is
why the pair was never covered; the multi-plate case is added.
Bambu Studio files a sliced print on the X2D's internal eMMC, which FTPS
does not serve. The bounded probe found a same-named file on the card --
an earlier slice of the same project, plate 4, against a running plate 1
-- and #1204's guard correctly refused it rather than archive another
plate's thumbnail, filament and cost.
It then blanked subtask_name because swap_plate_suffix returned None. But
None also means "this name carries no plate suffix", and such a name holds
no stale plate number to be wrong about. The project name was dropped, the
row fell through to the gcode_file path, and the archive was titled
plate_1.
Keep the name for the title only. subtask_name itself stays disowned,
because it is what every lookup here is built from and it keys
_active_prints, where the cover endpoint's own download of that same name
would find this archive and hand the contradicted file to
_recover_fallback_archive -- which checks a candidate is a readable 3MF
and never which plate it holds. A corrected name is still registered:
that one points at the plate actually running.
Also name the X2D alongside H2-series and P2S in the Archives banner, the
connection diagnostic and the storage-verdict docs -- it stores slicer
sends the same way, and an X2D owner was told the explanation did not
apply.
Bambu Studio files a sliced print on the X2D's internal eMMC, which FTPS
does not serve. The bounded probe found a same-named file on the card --
an earlier slice of the same project, plate 4, against a running plate 1
-- and #1204's guard correctly refused it rather than archive another
plate's thumbnail, filament and cost.
It then blanked subtask_name because swap_plate_suffix returned None. But
None also means "this name carries no plate suffix", and such a name holds
no stale plate number to be wrong about. The project name was dropped, the
row fell through to the gcode_file path, and the archive was titled
plate_1.
Keep the name for the title only. subtask_name itself stays disowned,
because it is what every lookup here is built from and it keys
_active_prints, where the cover endpoint's own download of that same name
would find this archive and hand the contradicted file to
_recover_fallback_archive -- which checks a candidate is a readable 3MF
and never which plate it holds. A corrected name is still registered:
that one points at the plate actually running.
Also name the X2D alongside H2-series and P2S in the Archives banner, the
connection diagnostic and the storage-verdict docs -- it stores slicer
sends the same way, and an X2D owner was told the explanation did not
apply.
The two tag-link routes answered the same conflict differently. The
built-in one said "Tag UID already linked to another active spool" and
named nobody -- while holding the conflicting spool row it had just
loaded -- and Spoolman mode named the spool inside a different English
sentence. Neither was machine-readable, so a client had to parse prose
to learn which spool to look at, and could only do it in one mode.
Both now raise one shared constructor: code tag_already_linked, the
holder's id, and which identifier collided. That is the detail shape
insufficient_filament and printer_connection_failed already use, so
ApiError parses it with no frontend change.
Two active spools can carry one tag -- no unique index on either
column, no conflict check on PATCH /spools/{id}, and /spools/bulk
copies one payload including the tag into every row it creates -- and
the lookup read that with scalar_one_or_none(), which raises on two
rows. The exception escaped into the auth middleware's fail-closed
handler, so the caller was told the authentication service was
unavailable. Both lookups are now ordered and take the first row, as
get_spool_by_tag earlier in the same file always has.
Naming the lowest id means the Spoolman scan reads every row where it
used to stop at its first match, so it now reads extra.tag defensively:
that field is edited outside Bambuddy, and a single null further down
the list would otherwise take the request down in place of the 409.
The kiosk reads the new code: a refused link showed a flat "Failed to
assign spool" and now names the spool holding the tag, reusing the
inventory.tagAlreadyLinked key that no code referenced.
The button gated on tag_uid alone. A spool linked by its 32-character
Bambu tray UUID carries none -- Bambuddy splits a stored tag by length,
so a 32-char value becomes tray_uuid and tag_uid stays empty. In
Spoolman mode that is every Bambu Lab spool synced from the AMS; on the
reporter's instance, 35 of 39 tagged spools, none of which could have
its tag cleared from the dialog. The documented workaround was to edit
extra.tag in Spoolman's own interface.
Everything around the button already treated those spools as tagged.
The Tag ID column renders whichever identifier is present, and the
payload the button sends nulls both fields -- which both inventory
modes honour: the built-in PATCH applies them through exclude_unset,
and the Spoolman route keys its tag-removal branch off either field
being explicitly null.
Either identifier now enables it, and clearing still removes both.
Swapping a Bambu spool for one the AMS cannot read left Assign Spool
publishing no ams_filament_setting at all. The printer kept showing "?"
on its screen and in the slicer, and only Configure, which publishes
unconditionally, put anything there.
Four places asked the tray's `state` field whether a spool was in the
slot. It cannot answer that. An AMS-HT reports its LOADED tray as 9
rather than 11, because it does not feed into a shared buffer the way a
4-slot AMS does -- the merge has skipped its own state heuristic for HT
units since #2594 for exactly this reason. And the field is partly our
own writing: apply_tray_exist_bits stamps state=9 on every slot whose
tray_exist_bits bit is 0, and when the bit comes back it refreshes only
the `exists` annotation beside it. Either way the slot sits at
exists=True, state=9 until something configures it.
That 9 also kept the deferred-configuration replay from firing -- its
own "has a spool appeared" test was the same heuristic -- which is the
deadlock #1322 removed from the assign path, still in place one step
further along. And it is what deleted the assignments in #3100: with the
replay never firing, the row kept the empty fingerprint it was stored
with, and the first tray report naming a filament was read as a swap.
All four now read tray_exist_bits first, which is the mask firmware
answers this question with and the one the printer card has drawn its
"?" from since #2527. The bit is allowed to overrule an "empty" state
and nothing else: a bit reading empty deliberately does not start
suppressing pushes that go out today, because the cost of computing a
bit position wrong is a slot that silently stops configuring, against a
saving of one message firmware would have dropped.
A blank tray report from a slot the bit calls occupied no longer unlinks
anything, off a print as well as during one, in both inventory modes --
Spoolman's parse_ams_tray calls a tray with no type empty, so a tag-less
spool assigned through the UI had its row deleted by the first idle push
after it went in. A filament the AMS cannot identify is not a filament
that was removed.
The Finance page was the only surface in Bambuddy that read its currency
from a data row rather than the `currency` setting, and it fell back to EUR
where every other page falls back to USD. One variable drives every amount
on that page, so the personal balance, the cost-center budgets and the whole
transaction list were wrong together on any install not set to euros. It now
takes the configured currency from /settings/ui-flags, which is readable by
anyone who can see Finance -- /settings needs SETTINGS_READ, which a
cost_centers:read_own user does not have.
The backend was the other half. Of the four places that settle on a
currency, three wrote a hardcoded "EUR": the wallet the API mints on demand,
the wallet a print charge mints when none exists, and the balance returned
for a user with no wallet row at all. All four now go through one resolver,
which lives beside the rest of the balance logic.
The wallet's currency column is removed outright rather than merely ignored.
An install has one currency and nothing here converts between them, so a
per-wallet copy could only ever drift from the setting -- and a column
nothing reads is a trap for whoever finds it next. A startup migration drops
it on both SQLite and PostgreSQL, after the raw CREATE TABLE that would
otherwise re-add it on an install whose finance tables predate the ORM.
SQLite builds older than 3.35 have no DROP COLUMN and keep it, harmlessly,
since it has a default and no reader.
Saving settings now invalidates the ui-flags query too. Nothing did, so a
changed currency sat behind that query's staleTime before showing up. The
sponsor prompt's own EUR fallback is now USD, matching AppSettings.
Each enabled virtual printer needs its own IP address, and the documented
way to get several is to add secondary addresses to the adapter already in
use. Linux reads those back through `ip -j addr show`, which reports every
address. Windows and macOS have no `ip` command and fell through to a psutil
enumeration that stopped at the first IPv4 of each adapter, so a host with
three addresses on one NIC offered exactly one bind target and the second
virtual printer could only fail with "Bind IP ... is already in use".
The psutil path now collects every address, marking the ones after an
adapter's first as aliases the way the iproute2 path does. Callers that want
interfaces rather than addresses -- the discovery scan and the support bundle
-- project the primaries back out, so their view is unchanged.
The ioctl fallback that a Linux host without iproute2 used to get is now
reached only when psutil itself is missing, which gains that host aliases
too. The interface-name exclusions stay Linux-only: they are Linux device
names, and a Windows adapter called "Local Area Connection" matches the "lo"
prefix.
macOS attributes Local Network permission to a code signature and judges a
launchd-spawned process on its own, rather than letting it inherit the grant
of the Terminal that started it. Homebrew ships Python unsigned on Intel, so
there is no identity for the grant to attach to: every connection to a LAN
address is dropped with no error the application can log and no permission
prompt. The printer reads as unreachable and nothing says why, and the entry
in Privacy & Security cannot be made to work because it refers to an identity
that no longer resolves.
install.sh signs during a macOS install; update_macos.sh re-checks on every
update, because `brew upgrade python` installs a fresh unsigned binary under
a new versioned path.
Both sign only what is currently unsigned. That gate is load-bearing: on
arm64 the linker ad-hoc signs every binary and the identity is a hash of the
file, so re-signing would rotate it and revoke a working grant on each update.
A python.org build carries a real Developer ID and must not be downgraded for
the same reason.
The interpreter and the framework's Python.app are both signed. The first is
what sys._base_executable resolves to and what the reporter's TCC log names;
the second is what his fix actually targeted. Which one macOS attributes
could not be established from either, and signing both costs nothing.
-----
fix(diagnostics): name the macOS permission that silently blocks the printer (issue #3114)
The port checks reported all three ports unreachable while the subnet check
passed, and port_mqtt's fix text sent the reporter after firewalls and IP
addresses. On a macOS native install that pattern has a cause neither of
those covers: no Local Network grant, denied with no error and no prompt.
A new macos_local_network check, appended on macOS only so no permanently
dimmed row appears for anyone else. It passes when the control port answered,
which is proof the permission is in place and means the signature probe never
runs on a healthy diagnostic. Otherwise it probes the interpreter: an
unsigned one gets the repair that fixes it, a signed one gets System Settings
— the arm64 case, where the identity is a hash of the binary, so a Python
upgrade presents macOS with a new application and strands the old grant.
Always warn, never fail, and only once port_mqtt has already failed, so this
can never be why a green diagnostic turns red. A printer that is simply
switched off produces the same all-ports-dead pattern, which is why the
signature, not the pattern, is what earns the specific advice. An
undeterminable signature is reported as the generic case rather than as
unsigned: that advice rewrites a file in the user's Python installation and
must not be offered on a guess.
POST /library/files/add-to-queue reported every per-file rejection in an
errors array and returned 200 regardless. A caller that checks the status
code saw a successful request, no visible failure, and no queue item.
That is a 400 now when nothing at all was added, with the same reasons in
the body. A call that created some items still succeeds, because it did.
The items it created were aimed at nothing. The route always wrote
printer_id=None with no target_model, and the scheduler dispatches on one
or the other -- so those rows matched neither branch and could never be
picked up by anything. They sat in Unassigned until someone opened each
one by hand.
The request takes an optional printer_id or target_model for the batch,
and with neither it aims each file at the model its own G-code declares.
Only when a printer of that model is active: owning no H2D is the user's
situation rather than their mistake, so the file still queues as the
unassigned row it has always been, rather than gaining a target nothing
can answer.
Three gates POST /queue/ has applied for a while now apply here too,
because an item reaching the scheduler through this route has to be as
printable as one reaching it through that one: the cross-model check that
stops a file sliced for one printer being dispatched to another (#2578),
the filename check that would otherwise surface as a failed upload hours
later (#1540), and the filament requirements the scheduler matches before
handing a model-based item to hardware.
Nothing inside Bambuddy calls this endpoint -- the Library's own Print
action goes through the queue API with a printer already chosen -- which
is how it came to drift this far from it.
-----
fix(library): scope add-to-queue file reads to the caller
The bulk add resolved its files by raw id. Every other read in this
module goes through the ownership gate, and so does the single-item
queue path; this one did not.
Invisible rows are dropped before the loop, so they report as the plain
"File not found" an unknown id already gets.
The Network subnet check told the reporter that 192.168.98.170 and
192.168.96.9 were on different networks and to go configure routing
between them. They are four hundred addresses apart inside one
192.168.96.0/22 LAN.
An IPv4 address does not carry its prefix, and the check supplied /24
for both sides. That is the most common LAN and not the only one, and
the guess is wrong in both directions: it splits a /22 and it merges a
/25. Read the prefix off the interface that owns the address instead.
find_local_ipv4_network() enumerates every interface, including the ones
EXCLUDED_INTERFACE_PREFIXES hides. That list keeps docker0 and friends
out of the Virtual Printer's bind dropdown; here the caller is asking
about an address the kernel has already picked as a route source, and
answering "unknown" because it sits on a bridge would be a worse answer
than the truth. When nothing claims the address the check skips, which
is what it always did with no host IP at all -- it must not assert a
split it cannot see.
The same check chose which of Bambuddy's own addresses to compare by
probing a route toward 10.255.255.255, which on a multi-homed host is
not the interface the printer is on. It asks for the route toward the
printer now. On a two-NIC dev box that alone was warning about a printer
sitting on the second card's own subnet.
The probe takes IPv4 literals only. connect() on a name would resolve
it on the event loop, and _same_subnet rejects names anyway, so nothing
is lost. Resolving the prefix shells out to `ip -j addr show`, so it
moves off the loop too.
-----
fix(diagnostics): name the container engine instead of asking about Docker (issue #3092)
"Not running in Docker - not applicable", said to a Bambuddy inside a
Podman container. It reads as "you are on bare metal", and it sent the
reporter looking for his problem somewhere else.
Podman runs Bambuddy in exactly the two shapes Docker does, and the
shape is the thing that breaks printer discovery and the Virtual
Printer. detect_container_runtime() names the engine -- Docker, Podman,
Kubernetes, containerd, LXC, or a container it cannot place -- and the
check became Container network mode.
is_running_in_docker() is deliberately left alone rather than rewritten
on top of it. Three callers key real behaviour off that flag, and one of
them switches the Add Printer flow from SSDP to subnet scanning. SSDP
works for a host-networked Podman container, so answering True there
would take a working feature away to fix a sentence. Widening it is a
separate decision from naming the engine, so it is made separately.
Mode detection keeps the original signal first, which also makes the
Docker path incapable of regressing: a Docker host always has a docker0,
so a container that sees one shares its namespace, and the new rules can
only turn a warning into a pass. That signal says nothing about Podman,
which creates no such interface on a host running no bridge containers --
which is how host networking came to be reported as bridge. The general
form of the same idea answers for Podman: an interface whose iflink
equals its ifindex was created in this namespace, and a NAT-networked
container only ever receives one end of a veth pair. tun/tap is skipped,
because a container may run its own WireGuard and that tun is native to
a namespace it is not evidence of. The interface also has to be the one
the kernel just named -- sysfs is namespace-tagged but a bind-mounted
host /sys is not, and reading a colliding name's numbers would be
reading another namespace's answer.
What is still unreadable now says so and suggests host networking if
discovery is failing, rather than guessing bridge and telling a healthy
install to recreate itself. An LXC or LXD system container is named and
told the question does not apply: it is on the LAN like a small virtual
machine, so there is no network mode to recommend -- and its subnet
check still runs.
An engine we cannot name is a sentinel the frontend localizes, not a
word interpolated into thirteen other languages.
The support bundle carries the engine name beside the Docker flag, so
the next report of this shape is answerable from the bundle.
SpoolBuddy said "Unknown color" under a correctly-coloured swatch for
spools the inventory page names without trouble.
The name was never in the spool record. Bambu's RFID tags frequently
carry no readable colour name -- some carry an internal code instead --
so Bambuddy has always resolved the swatch's own hex against the colour
catalog, and the kiosk was rendering the empty column. Exactly one
SpoolBuddy file already did it right, which is what marks this as an
inconsistency rather than a kiosk simplification.
Route every SpoolBuddy colour-name display through resolveSpoolColorName,
which also stops the spools that do carry a code from showing "A06-D0" at
the user. The write-tag edit form keeps the raw stored value on purpose:
offering a derived name for editing invites the user to save it as though
they had typed it.
Spoolman has no colour-name field at all, so _map_spoolman_spool puts the
spool's subtype there and sets color_name_is_synthesized. That flag now
travels on the tag-matched broadcast, and resolveSpoolColorName takes a
third argument to honour it -- a synthesised name loses to the catalog
and survives only as a last resort. Spoolman installs were reading
"Silk+" as a colour on the Inventory page and the AMS hover card too, so
those call sites pass the flag as well.
Searching by a colour you can read on screen now finds it, in the kiosk
and in Bambuddy: the shared inventory filter matches the resolved name as
well as the stored one. That makes the filter depend on the catalog,
which loads asynchronously, so the three memoised call sites take its
version as a dependency -- without that, a query typed before the catalog
arrives keeps its empty result and reproduces the very symptom being
fixed.
The fallback label was hardcoded English in components that already
import useTranslation; it is now spoolbuddy.spool.unknownColor in all 14
An external camera could pass the connection test, play in VLC, and show a
black live view that gave up after a few seconds.
The two RTSP paths were not asking ffmpeg for the same thing. The one-shot
_capture_rtsp_frame passed no probe settings and got ffmpeg's defaults;
_stream_rtsp hard-coded -probesize 32 -analyzeduration 0. Thirty-two bytes
is enough for a camera that puts its H.264 parameters in the SDP, and not
enough for one that sends them in-band a moment later -- a WebRTC source
republished through go2rtc, in the reporter's case. ffmpeg then starts no
decoder and yields nothing at all, which is why the test button kept
passing while the live view stayed black.
Those settings were never chosen for external cameras: they arrived with
the P2S TLS proxy (#661) as fast-start tuning for the printer camera path,
where the source is a known Bambu model, and were copied here in the same
commit. This path has no model to tune against and belongs on the
defaults, which are a ceiling rather than a wait -- a camera that
announces itself in the first packet still starts as fast as it did.
Drop the probe cap from _stream_rtsp. Keep -fflags nobuffer and -flags
low_delay, which bear on how long ffmpeg sits on frames it already has
rather than how long it may look before it has any. Leave
_capture_rtsp_frame and the per-model printer profiles alone.
Tests pin the absence of both flags, the presence of the low-latency ones,
and the property underneath: both RTSP paths must probe alike, or passing
the test button again means nothing about the live view. The ffmpeg
subprocess fakes move to backend/tests/_fixtures/external_camera.py so the
SSRF suite and this one share one definition.
POST /printers/{id}/bed-jog takes a signed nozzle-bed gap, documented since it
was written: positive asks for more room between the nozzle and the plate. On
an A1 it did the opposite. The reporter sent distance=5 for clearance and
watched the toolhead come down.
The sign had been flipped on A1 models since the original report on this issue,
where an A1 Mini owner clicked an arrow labelled "move the plate up" and watched
the nozzle dive. That is a labelling problem -- a bed-slinger's plate does not
move in Z at all, so closing the gap shows up as the toolhead descending -- and
it was solved in the transport layer, which turned a parameter documented as
model-independent into one that meant the opposite thing on part of the fleet.
Z is the nozzle-to-bed distance on every Bambu model, by definition of the
coordinate system rather than by convention: G1 Z+ opens the gap whether the bed
drops away from a fixed nozzle (X1/P1/H2, whose end G-code parks with
G1 Z{max_layer_z + 100}) or the nozzle rises off a fixed bed (A1/A2L). The
finish-photo plate restore already relies on exactly that and carries no model
branch. So distance goes onto the wire unchanged and one call means one physical
outcome everywhere: positive is the safe direction on every printer.
Which way an arrow points is a different question, about the machine in front of
the user rather than about G-code, so the printer card answers it and asks for
the gap it wants. The buttons move what you would expect them to move, exactly
as before; on a bed-slinger they now say toolhead rather than plate.
The A2L never had the old fix. It slings its bed the same way the A1 does, but
the inversion listed the A1 names and the A2L was not among them, so its up
arrow has been sending the toolhead at the plate for as long as the machine has
been supported. The new classifier also covers the alternate internal codes
A04 / A11 / A12, which LINEAR_RAIL_MODELS and SINGLE_NOZZLE_FLOW_MODELS both
carry and the old gate did not.
is_bed_slinger is gone from the backend rather than widened: with the route
model-independent it had no caller, and a kinematics helper sitting unused in
the service layer invites the next person to assume the backend handles
direction. It does not, deliberately.
Separately, the soft-endstop comments on both jog routes claimed the firmware
clamps a bare move at the travel limit. It does not, and #2579 measured that:
an H2D at its Z limit ran straight past a clean G91/G1 Z-1.00/G90, while its own
touchscreen refuses the identical move. What #2579 removed was M211 S0, which
disabled the limits globally and took the touchscreen's protection with them.
The jog popover has warned about this correctly the whole time; only the code
comments disagreed with it.
A job queued as "Any X1C" explains itself when it cannot start: the
model-based branch builds a reason for every candidate printer and puts it
on the row, so the queue shows "Busy: X1C-01" or "Waiting for filament:
X1C-02 (needs PETG)". The same job pinned to one printer showed nothing.
It sat at Pending with waiting_reason NULL for as long as that printer was
busy, which from the outside is indistinguishable from a queue that has
stopped working -- the reporter watched fourteen minutes of it while his
X1C ran a print he had started from its own screen.
The fixed-printer branch had six ways out and none of them wrote the field.
The sensor interlock (#1148) was its only writer, and it cleared the field
up front on every pass where no sensor was holding the printer, so NULL was
not an oversight on those paths but a guarantee.
Every exit now writes, through one helper. The reasons reuse the
model-based branch's vocabulary so _is_busy_only() keeps deciding what is
worth a notification: a printer that is printing, drying, or working
through the item ahead of this one reads as "Busy: <printer>" and stays
silent, because it resolves itself. A printer that is off with no Auto On
plug, and one whose plug could not switch it on, are worth saying.
A finished plate nobody has acknowledged is split out from plain busy and
named as itself. _is_printer_idle() returns the same plain False for that
and for a running print, but they are not the same thing to the person
looking at the queue: one clears itself and the other needs somebody to
walk over to the printer.
That notification fires on the transition into asking, where a busy-only
reason counts as not asking. Testing whether the item was waiting at all --
which is what the model-based branch does -- would never fire it here:
nobody's queue goes straight from idle to an unconfirmed plate, it waits
behind the print first. The cost is that a printer dropping offline,
returning busy and dropping again asks twice rather than once.
The interlock stays silent. It has never sent this notification, and a
change about what the queue displays is not the place to start.
Clearing the field up front is gone with it. It existed so a shut door
could not leave "Waiting on Enclosure Door" standing while the printer
stayed busy with something else, and the new rule carries that guarantee
instead -- whichever exit runs next overwrites it, and the dispatch path
clears it.
Two paths clear it that the report did not mention. A staged item and a
future-scheduled one skip before this branch and never reach it again, so
anything written on an earlier pass would outlive its condition for the
life of the row. That includes the filament-deficit check, which stages the
item itself.
The notification is wrapped: a queue that cannot say why it is waiting is
the bug being fixed, and a queue that stops dispatching because a provider
timed out would be a worse one.
On the frontend, the queue timeline drops any pending item carrying a
reason, on the grounds that such an item will not auto-dispatch. That held
while only the model-based branch wrote the field; "Busy: <printer>" is
now the commonest reason there is, and it describes the very chain the
timeline forecasts, so the rule would have emptied the view for anyone
whose queue is pinned. It now asks whether the reason needs the user, via
a small shared reader of the same shape the scheduler encodes.
Which job goes out, and when, is unchanged: running the previous scheduler
and this one over the same 768 states dispatches the same items in the same
order with the same statuses, across 1452 rows that now carry a reason.