Every slot of the A2L's AMS Lite rendered as "?" in Bambu Studio through the
Virtual Printer while Bambuddy's own AMS card was correct, and a filament set
by hand in Studio reverted about a second later.
The A2L reports its AMS Lite as physical unit id 16 but packs the slot presence
bits at base 24, so bambu_mqtt normalises the id to 6 at the ingest boundary and
every internal reader gets the right bits. The VP bridge is not downstream of
that: BambuMQTTClient._on_message fans raw payload bytes out to raw-message
handlers before parsing, so mqtt_bridge._on_printer_raw does its own json.loads
and still holds id 16. It then called the shared apply_tray_exist_bits, which
computed 16*4 = bits 64-67 -- never set -- concluded all four slots were empty,
and wiped tray_type / tray_color / tray_info_idx / tag_uid / tray_uuid / remain
from the copy sent to the slicer. That runs on every push, which is why a manual
pick could not survive the next 1 Hz cached-as-base report.
apply_tray_exist_bits now folds the unit id through normalize_am_unit_id, so 16
and 6 land on the same bit base whichever id the caller holds. The bridge's
cached ids stay physical on purpose -- Studio addresses the Lite as 16, sending
ams_get_rfid {ams_id: 16} through the VP -- so normalising the cache instead
would have broken the slicer's own command path.
Confirmed from the reporter's debug log, which shows the cleanup clearing slots
at bits 64-67 under the VP's log label. Before #2670 added the
0 <= ams_id <= 15 range guard this wiped the slots; after it, unit 16 fell out
of the guard and the A2L got no empty-slot cleanup at all -- two different wrong
answers, both fixed here.
Signing in to Bambu Cloud with an authenticator-app account failed every
time with "Invalid code", whatever the code was. Bambu Lab added double-
submit CSRF protection to the bambulab.com web origin - which is where,
and only where, this service posts the two-factor code. The endpoint
refused the request with 403 "CSRF error: missing_cookie" before it ever
evaluated the code, and Bambuddy reported that refusal as a bad code.
Verified against the live endpoint with a deliberately invalid key: a
bare POST returns missing_cookie; GET /api/csrf mints a bbl_csrf_token
cookie; a POST carrying only the cookie returns missing_header; a POST
carrying the cookie plus an x-bbl-csrf-token header reaches application
logic. Landing on the sign-in page first - the intuitive fix - does not
help, as that page sets only Cloudflare's __cf_bm. Of five header
spellings tried, only x-bbl-csrf-token is accepted, so the tests pin it.
verify_totp now performs that handshake against the same origin it will
post to (bambulab.cn for the China region - a token minted by the global
site is a cookie the .cn endpoint never issued), and declines to submit
the code at all when no token can be obtained rather than burning the
user's 30-second TOTP window on a request that is certain to be refused.
A CSRF refusal now also says the code was never checked instead of
masquerading as a wrong code, which is what sent the reporter chasing
clock drift and leading-zero parsing.
Only TOTP sign-ins were affected. Every other cloud call, the email-code
two-factor path included, goes to api.bambulab.com, which is not gated,
and existing stored tokens were unaffected throughout.
The region-routing test's MockTransport needed teaching about the
handshake: it returns one canned response for every request and set no
cookie, so the fix correctly refused to POST and the test lost the URL it
asserts on. It now mints a token for /api/csrf and additionally checks
the handshake stays on the .cn origin.
Gitea/Forgejo served under a ROOT_URL path prefix (e.g. https://host/gitea)
place repos at /<prefix>/owner/repo. The Gitea backend assumed a host-root
layout: parse_repo_url required exactly two path segments (so subpath URLs
raised "Cannot parse repository URL") and get_api_base dropped the prefix,
yielding https://host/api/v1 instead of https://host/gitea/api/v1.
Treat the final two path segments as owner/repo and keep leading segments as
a base-path prefix; derive the API base as {scheme}://{host}{prefix}/api/v1.
Root-hosted instances are unaffected. Forgejo inherits the fix; GitHub/GitLab
are untouched.
With LDAP auth in use, the debug log carried the full user DN on successful
auth -- e.g. "(DN: CN=Joe Schmoe,CN=Users,DC=ad,DC=example,DC=com, ...)". A DN's
leaf CN is the user's real name, PII on par with the email address already
redacted, and it passed straight into an uploaded support bundle. The log
sanitizer (shared by the support bundle and the in-app bug report) had no DN
pattern; DNs also leak via ldap3 exception strings and group-mapping logs.
- sanitize_log_content: redact LDAP DNs to [DN] -- a run of >=2 attr=value RDN
components (CN/OU/DC/UID/...). The value class excludes <>;+ (RFC 4514 requires
them escaped in a value) so the final comma-unbounded component doesn't swallow
trailing log text such as "-> GroupName". Ordinary key=value lines are untouched.
- ldap_service: stop logging the raw DN on successful auth (username + group
count suffices), keeping the PII off disk even before bundle sanitization.
Closing an external USB (V4L2) camera view abruptly could leave its ffmpeg
running and holding /dev/videoN open -- LED stuck on, and reopening the view
failed or took 10-30s fighting for exclusive device access. Same class of leak
as #776 (built-in RTSP path), but the external path was never wired into that
fix: external streams registered into none of the _active_streams /
_disconnect_events / spawned-PID registries, so /camera/stop returned
{"stopped": 0} for a live USB stream and the orphan janitor's /proc net matched
only rtsp(s)://bblp: cmdlines. Cleanup ran only via the stream generator's own
finally, which an abrupt disconnect can skip.
- Thread an on_process callback + stop_event through generate_mjpeg_stream into
_stream_usb / _stream_rtsp; register the process before the startup probe so a
process that hangs on a locked device (not just one that exits) is reapable.
- Register external streams into the shared registries under a unique
{printer_id}-ext-{token} id so /camera/stop and cleanup_orphaned_streams find
and kill them; stop_event prevents the reconnect loops from respawning.
- Extend the /proc safety-net scan to also match USB (-f v4l2) ffmpeg, excluding
still-active streams and unrelated ffmpeg.
The slice client only checked the sidecar's HTTP status, not its body. When the
sidecar -- or a reverse proxy in front of it -- returned 200 OK with a body that
wasn't a real 3MF (a stock/misconfigured sidecar, a proxy error page, a truncated
response, or an OrcaSlicer/Bambu Studio CLI crash emitting no output), Bambuddy
wrote that tiny blob to a .gcode.3mf, stored it as a valid sliced file (the
3MF-parse failure was swallowed as "no thumbnail"), and let it be queued and FTP'd
to the printer -- producing the ~28-byte files that "did nothing" and then failed
at print time. Separately, a genuine 413 comes from the proxy in front of the
sidecar rejecting the multi-MB upload (model + profiles), so raising the body
limit on the wrong proxy layer had no effect.
- Factor the duplicated status handling in slice_with_profiles /
slice_without_profiles into one _handle_slice_response.
- When a 3MF export was requested, validate the body is a real ZIP; otherwise
raise SlicerApiServerError with an actionable message instead of persisting
a corrupt file.
- Special-case 413 with a message naming client_max_body_size on the proxy
directly in front of the sidecar (Cloudflare cap noted).
The folder tree's "sort by recent activity" and the file pane's date sort
put external (mapped/NAS) files in a near-random order instead of ls -t's
newest-first. Nothing captured the files' on-disk mtime: the sort keyed off
the DB updated_at/created_at, which for a bulk external scan is the same
scan instant for every row, so a whole block tied and sorted arbitrarily;
only rows Bambuddy had later touched individually looked "partially right."
The tree also bubbled up only immediate child-file activity, so a file added
deep in a subtree never lifted its parent folders.
- Add nullable fs_modified_at to LibraryFile and LibraryFolder (dialect-
branched migration, mirroring the #2615 dispatching_at pattern).
- External scan records each file's and directory's real os.stat().st_mtime
and refreshes it on every re-scan, so a file edited over the mount
re-sorts and existing installs backfill on the next scan.
- list_folders computes each folder's activity as a recursive newest-
descendant roll-up (post-order), so a fresh deep file lifts every ancestor.
- Folder tree sort and the file pane's date sort now use the real mtime,
falling back to created_at for managed uploads with none.
- New toolbar toggle shows/hides each item's last-modified date in the right
pane (grid + list), with strings in all locales.
Store the mtime as naive UTC to match the other timestamp columns so activity
comparisons never mix naive and aware values on either dialect. Covered by
integration tests (mtime capture, re-scan refresh, deep-file recursive bubble,
folder mtime) and a frontend test proving fs_modified_at is preferred over
created_at.
After #2594 every empty-slot clearing path skipped AMS-HT units, so a removed
HT spool never cleared on the printer card while Bambu Studio correctly showed
Empty. Root cause: the HT presence bit is packed as a single consecutive bit in
tray_exist_bits at 16 + (ams_id - 128) (HT-A=16, HT-B=17, ...), not the regular
ams_id*4 position -- so the bitmask cleanup, which skipped id>=128, never
touched it. The HT's state field is firmware-variant (loaded reports 11 on H2D,
9 on the #2594 firmware) and it keeps echoing stale tray_type after removal, so
the bitmask is the only reliable, firmware-independent signal. Confirmed against
a live H2D capture (loaded=0x10f7f, empty=0xf7f) and the OrcaSlicer
DevFilaSystem.cpp reference (is_exists = tray_exist_bits >> (16 + (ams_id-128))).
- apply_tray_exist_bits: handle AMS-HT (128-135) at 16+(ams_id-128) instead of
skipping it; clears an empty HT and, because a loaded HT keeps its bit set,
never wrongly clears one (keeps the #2594 fix intact). Ids outside the known
regular (0-15) and HT (128-135) ranges are left untouched rather than guessed.
- Build the AMS change-hash from the merged state, not the raw payload, so a
removal signalled only by the bitmask (firmware still echoing tray_type) still
flips the hash and fires on_ams_change to unbind the spool_assignment row.
- Emit the exists presence bit in the websocket status serializer (REST already
did) so the card renders "Empty" instead of "?" where state is ambiguous.
Fetching K-profiles probes every nozzle size in turn (extrusion_cali_get
for 0.2/0.4/0.6/0.8mm), and each response echoes the requested diameter at
the top level. _process_message passed every "print" message -- including
these responses -- to _update_state, which treats a top-level
nozzle_diameter as the installed hardware. So the last size probed (0.8)
overwrote the real nozzle in memory; a genuine status push corrected it and
the next K-profile fetch broke it again, which is why it flickered between
0.8, empty and the correct 0.4. Since 1.2.5 the #1899 mismatch guard then
refused to dispatch, failing prints with a bogus "printer has 0.8mm".
Handle extrusion_cali_get responses only via the K-profile parser and skip
_update_state for them, mirroring the existing get_accessories guard. The
nozzle size now comes solely from the real status push. In-memory only --
affected printers self-correct on the next push after updating.
Follow-up to 0f203ce: force color match now picks the right AMS slot, not
just the right printer. The slot mapper cleared tray_info_idx when applying
an override, so on a printer holding two same-colour PLA spools of different
variants (Basic GFA00 / Matte GFA01 / Silk GFA06) it could map to the wrong
one. It now keeps the variant for force_color_match overrides (both the 3MF
and no-3MF fallback paths) so the matcher pins the matching tray, and falls
back to type+colour when that variant isn't loaded. A manual filament swap
(a preference override) still clears the idx so it matches the swapped-in
spool rather than the old one.
The printer-card queue-compatibility hint applies the same variant rule.
Force color match dispatched onto the wrong PLA sub-variant: a job sliced
for White PLA Matte was treated as an exact match by printers loaded with
White PLA Basic or Silk+, because Bambu reports every variant as
tray_type "PLA" and the distinction lives only in tray_info_idx
(GFA00=Basic, GFA01=Matte, GFA06=Silk).
Two places dropped the field: the VP queue built each force override
without the parsed tray_info_idx, and _get_missing_force_color_slots
compared loaded trays on (type, colour) only.
Carry tray_info_idx into the override and require it to match when both
the override and a candidate tray have one; a blank idx on either side
(custom/third-party spools, older 3MFs) falls back to the historical
type+colour behaviour, so those setups are unaffected.
The bed_levelling/flow_cali/nozzle_offset_cali boolean->tristate migration
builds two UPDATE statements with an f-string interpolating the column
name. Bandit flags these as B608 (SQL injection) at medium severity, which
failed the release-gate scan in test_security.sh.
The interpolated _col only ever iterates the hardcoded _tristate_cols tuple,
never user input, and SQL identifiers can't be passed as bound parameters.
Suppress with `# nosec B608` (matching the existing settings.py convention)
plus an inline rationale. No behavior change.
test_launcher_shutdown_config.py and test_systemd_backup_paths.py read
repo-root launcher/config files (Dockerfile, docker-compose.yml,
deploy/bambuddy.service, install/install.sh, installers/windows/...,
spoolbuddy/install/install.sh). The Docker test image built from
Dockerfile.test copies only backend/, pyproject.toml, gcode_viewer/ and
requirements, so all 15 tests failed in test_docker.sh with "launcher
moved or was removed" — the files simply aren't in the image.
Guard both modules with skipif on frontend/package.json, which is present
in every source checkout but never in the test image. Native runs
(test_backend.sh, every commit) still execute the tests in full and catch
a genuinely moved/deleted launcher; the release-gate Docker run skips them
instead of failing on files it deliberately doesn't ship.
Pairs the top-down plate preview with the slicer's per-object pick mask
(Metadata/pick_N.png), whose pixel colours encode the same identify_id the
firmware's skip command takes, so a click resolves to a real object rather
than an inferred bounding box. Several objects can be selected before one
confirmation; selected and already-skipped items are highlighted on the
plate; the checklist stays available when no mask exists.
view=pick serves only the active plate's mask and 404s otherwise, unlike
every other view. A render returned in a mask's place would be decoded as
object IDs — dark pixels yield small integers that collide with real ones —
and a click would then skip an arbitrary object, mid-print, irreversibly.
The 404 is what tells the UI to fall back to the checklist.
Click mapping goes through the contained rect, since the canvas paints at
mask resolution under object-contain; clicks on a letterbox bar are rejected
rather than clamped onto whichever object touches the border. Confirming
names the object when one is selected and counts them when several are,
which is what plates of identically-named clones need.
No printer-control command path was added or changed; the layer, permission
and existing skip-command guards are untouched.
Slicing for a P2S failed with "filament preset Bambu PLA Basic @BBL X1C 0.2
nozzle (slot 1) is not compatible with printer Bambu Lab P2S 0.4 nozzle" —
naming a profile shown nowhere in the dialog. The picked profile was
"Overture PLA Matte @0.2", whose inheritance chain roots in that X1C profile.
The dialog classifies a profile by its compatible_printers list and falls back
to reading the printer out of its name. That name carries no model, and the
list — present on the imported copy — is not shipped by every source: Bambu
Cloud omits it deliberately (rate limits), and Orca Cloud shipped it but
Bambuddy only mined filament type and colour from the same content.
Orca Cloud entries now carry their own compatible_printers, and the existing
same-name enrichment bridge carries the list onto entries that lack one, in
both directions between the cloud tiers. A bare "@<size>" name tag is read as
a nozzle size as a last resort: it can rule a printer out but never rules one
in, and implausible values are ignored rather than guessed at.
An end-of-print auto-off on a plug that powers a filter fan marked the linked
printer offline and forced its state to "unknown". The mark was unrecoverable:
connected heals on the next MQTT message but state does not (only frames
carrying gcode_state rewrite it, and steady-state push_status frames are
partial), so the printer stayed "unknown" until a manual Force Refresh and the
queue never dispatched to it again.
The offline mark is now an explicit presumption: mark_power_off records the
state it overwrites and _on_message undoes it as soon as the printer sends
another report on its own topic, since inbound traffic proves the power was
never cut. A reconnect discards the saved state, so a genuine power cut is
unaffected. Each plug also gains a controls_printer_power flag (default true,
backfilled) that gates all five power-off paths, and the queue's power-on step
now picks the flagged plug instead of whichever linked plug came first.
The A2L reports its 4-slot AMS Lite as unit id 16, but its slot-presence
bitmasks sit at bit base 24 (id 6) and it reports tray_now as a local 0-3
slot. Fed the raw id 16, the ams_id*4+slot convention probed bits 64-67
(always zero) and marked loaded slots empty; the local tray_now was read as
global, so usage deducted from the wrong spool (or not at all); and the
ams_id<=7 DB constraint rejected id-16 Spoolman links.
Normalise the Lite 16->6 at the MQTT ingest boundary so global tray ids land
at 24-27 - matching the firmware's own bit base, working with every existing
ams_id*4+slot consumer, colliding with nothing, and passing the DB
constraint. Globalise tray_now to 24+slot, widen the valid-tray guards, label
the unit "AMS Lite", and build the confirmed ams_mapping2 {ams_id:16,
slot_id:0-3} / flat 0-3 for dispatch. Outbound slot commands translate 6->16
on the wire via a single helper. Self-scoping: only unit id 16 is touched, so
all other printers/AMS types are unaffected. One uncaptured wire field (the
physical global tray on load/cali) is extrapolated and isolated to the helper.
Assigning a spool to an AMS tray pushed ams_filament_setting +
extrusion_cali_sel and reported success immediately, whether or not the
tray accepted it. A silently-dropped assignment never surfaced, and since
a print only deducts from the spool on the exact tray it pulls from, it
also recorded no filament usage - which made the whole thing feel random.
Read the AMS telemetry back after every assign (inventory assign_spool and
the Configure Slot modal) and toast the outcome: loaded when the tray
echoes the pushed tray_info_idx, a warning when the filament loaded but the
K-profile (cali_idx) did not, or not-confirmed after ~30s. Verification
uses the periodic per-tray push (the command ack hardcodes sequence_id 0
and can't be correlated); an on-demand pushall is nudged so it lands
quickly. Covers regular AMS, AMS-HT and external slots; stays silent rather
than inventing a failure if the printer goes quiet. The read-back check
runs on every AMS push because the change-hash excludes tray_info_idx.
Since #2562, a Bambu Cloud sign-in flipped to "expired" and forced constant
re-logins even while cloud features worked. #2562 made a 401 durably record
the stored token as dead, but treated *any* 401 from any cloud/MakerWorld call
as expiry. Bambu 401s for benign reasons (endpoint/region/scope refusals,
Cloudflare edge, transient blips), so one stray 401 -- including from a
background poll -- signed the whole cloud integration out until manual re-login.
The flag lives in the DB, so a setup with more than one instance against the
same database signed the user out across all of them.
Invalidate only on Bambu's documented expiry body {"code":4,"error":"Please
login."}. A plain/unparseable 401 is treated as transient: the request fails
but the session stays signed in. validate_token maps a signature-less 401 to
None (unknown), never expired. A shared is_expiry_401() gates both the Bambu
Cloud and MakerWorld services (same token). Genuine expiry is still detected
and surfaced exactly as before.
Bed levelling, flow calibration, and nozzle-offset calibration were on/off
only, so the sole way to run bed levelling was to force a full level before
every print. Bambu Studio has always offered a third "Auto" state that lets
the printer skip the calibration when it was done recently -- the state most
users actually want. Make these three options tri-state (off/on/auto),
defaulting to auto, and leave vibration/layer-inspect/timelapse as on/off
(Bambu Studio exposes no auto for those).
Wire encoding follows Bambu Studio's source exactly: each option sends a JSON
bool (true only for "on") plus a companion int -- off=0, on=1, auto=2. The
bool fields stay booleans (the #1478 H2S regression); only the companion int
widened from {0,1} to {0,1,2}. #1721's observation that stage 8/39 stays
queued when sending 2 is the auto contract (queued, skipped at runtime if
recent), not a broken "off".
- schemas: TriState = Literal[off/on/auto] with a BeforeValidator coercing
legacy bool / 0-1 / true-false so old clients and un-migrated rows validate
- model + migration: boolean columns -> String; SQLite via column affinity +
data backfill, PostgreSQL via ALTER COLUMN TYPE guarded on information_schema
(verified on both dialects); settings rows normalised true/false -> on/off
- MQTT: start_print takes the tri-state strings and emits the paired bool+int
- Virtual Printer: reconstructs the slicer's auto/on/off from the int companion
(auto_bed_leveling / extrude_cali_flag) in both capture paths
- frontend: CalibrationMode type; off/auto/on segmented controls in the print
dialog, queue bulk-edit, and Settings -> Workflow; calibrationMode_* strings
in all 11 locales
The /overlay/{id} route renders without a login, but everything it draws is
auth-gated: printer status and name (PRINTERS_READ), one setting (SETTINGS_READ),
and the camera stream (a camera-stream token). A signed-in browser rides its JWT
from local storage; OBS is a fresh browser with no session, so the overlay stayed
blank whenever authentication was enabled. Cloudflare/remote access was never the
cause -- an incognito window fails identically.
Give the overlay a self-contained kiosk-token mode, mirroring the Cam Wall:
- New `overlay` long-lived-token scope, kept separate from `camwall`: the overlay
names the printed file on screen, which a Cam Wall token is trusted never to
expose, so folding it in would silently widen every existing wall token.
- New token-authed GET /printers/{id}/overlay-status returning exactly the fields
the overlay draws and nothing else; added to the auth-middleware allowlist so it
reaches its own RequireOverlayTokenIfAuthEnabled gate.
- StreamOverlayPage reads ?token= and, in that mode, authenticates its status and
camera calls with the token and skips the WebSocket (the 2s poll is the feed).
The logged-in path is unchanged.
- Token-mint UI (Settings > API Keys) offers the scope with a ready-made
/overlay/{id}?token= URL copied once on creation.
A queue row stays status='pending' for the whole FTP upload; status only flips
to 'printing' at the end. The edit routes only blocked non-pending rows, so a
PATCH during the upload window was accepted while the in-flight dispatch kept
using its snapshotted printer -- splitting the queue row from the archive /
expected-print / physical command across two printers, and enabling a duplicate
dispatch on restart. The #1853 CAS guards cancellation, not reassignment.
Add a dispatching_at claim, stamped atomically (WHERE status='pending' AND
dispatching_at IS NULL) before any slow I/O and cleared on every exit. While
held, the single-item PATCH returns 409 (re-checked just before the write),
bulk edits skip the row, and the scheduler won't re-select it. Startup
reconciliation clears claims orphaned by a crash mid-dispatch. The row stays
pending throughout, so no status/UI/completion/reconciliation path changes.
New column print_queue.dispatching_at (nullable, dialect-safe DDL). Covered by
scheduler tests (claim exclusivity, non-pending rejection, release-on-exit,
skip-already-claimed, startup stale-clear) and API tests (reassign 409,
printer_id unchanged, bulk skip, unclaimed row still edits).
A single plate dispatched from a multi-plate 3MF could log the entire file's
filament against that one plate. When the AMS tracker measured nothing, a
completed run's PrintLogEntry.filament_used_grams fell back to
PrintArchive.filament_used_grams -- the sum over every plate (correct for the
archive card / project rollup, #1593) -- ignoring the archive's plate_id. So
each printed plate of a 22-plate file logged the full ~12 kg; cost inherited
the same whole-file value.
Forward: when the archive has a plate_id and its 3MF is on disk, the completed-
run fallback uses that plate's own slicer estimate (extract_plate_metadata_from_3mf)
and scales cost by the plate's share of the whole. Tracker-measured runs and
single-plate archives are unchanged.
Backfill: a startup migration repairs rows already written -- completed entries
whose stored grams exactly equal the archive's whole-file value, with a plate_id
and an on-disk 3MF, get recomputed to plate-scoped grams + cost. The exact-match
guard never touches tracker-measured or partial rows; idempotent, data-only,
identical on SQLite and Postgres, and logs the correction.
Server-side slicing always applied the picked printer/process/filament
triplet via --load-settings, which overrides the designer's embedded
project_settings.config — so a MakerWorld model set up for 5 walls came
out at the picked profile's default 2. That override is correct for
re-slicing a design onto your own printer/AMS, but there was no way to
slice a file the way its author configured it.
SliceModal now offers a "Use the file's built-in settings" checkbox when
the source 3MF carries embedded settings AND the picked printer matches
the design's target model. It routes to the existing embedded-settings
slice path (previously only a crash fallback), so walls/infill/filament
come from the file. Ticking it locks all four preset dropdowns — printer
included, since it's unused on this path and changing it would drop the
match and hide the toggle. The printer-match gate stops embedded settings
being honoured across models (wrong bed); there is no cross-printer
re-targeting on this path.
- schema: use_embedded_settings on SliceRequest
- route: embedded_mode branch; crash-fallback guarded against re-running
- frontend: gated checkbox locking all four dropdowns, resets on mismatch
- 2 i18n keys across all 11 locales
- tests: backend (flag skips triplet / ignored for STL) + frontend
(toggle offered on match, locks dropdowns + sends flag / hidden on mismatch)
A print queued from a specific plate of a multi-plate 3MF showed as Plate 1
in Print History after cancellation: the archive derives its plate from the
filename, but a whole multi-plate 3MF uploads under one name with no plate
suffix, so the parser defaulted to plate 1 and nothing copied the queue
item's plate_id onto the archive (which had no plate field).
Add a nullable print_archives.plate_id, copy it from the queue item at
dispatch (archive- and library-file paths), expose it in the archive API,
and render it in Print History. A startup backfill copies the plate onto
existing archives from their linked queue rows. Column add + backfill are
identical on SQLite and Postgres.
Also fix a related lifecycle bug: stopping a printing item while the printer
was offline left the linked archive stuck at "printing" (queue row
cancelled, but no MQTT completion ever arrives to reconcile the archive).
The offline-stop path now closes the archive out directly; the online path
still defers to the MQTT completion event.
check_queue awaited asyncio.gather() over the whole selected batch before
returning, so the scheduler run loop was blocked until the slowest FTP
upload in the batch finished. On a large farm a 513s upload left 15 of 16
configured upload slots idle for 8.5 minutes while other printers came
free — the setting behaved as a per-batch cap, not a worker pool.
Launch uploads as independent background tasks tracked in a _inflight pool.
Each tick excludes in-flight item rows and their printers from selection,
launches at most limit - len(_inflight) new uploads, and returns
immediately, so a freed slot refills on the next fast tick. The no-double-
dispatch invariant the batch-await provided (rows stay pending until upload
completes) is now carried by the in-flight exclusion; the pending->printing
CAS, busy-printer guard (#2598), per-printer hold, auto-drying exclusion,
and per-item failure isolation are all preserved per task.
Rewrites the concurrent-dispatch tests around pool/reservation/refill
semantics and adds coverage for slot refill, in-flight exclusion, and the
non-blocking return.
The Configure AMS Slot modal sends built-in / local / Orca-generic presets
with a GF* tray_info_idx but an empty setting_id, and configure_ams_slot
forwarded that empty value to ams_filament_setting. The firmware treats a
filament-id-without-setting-id slot as half configured: it shows the new
material briefly, then reverts to its previously stored profile.
Back-fill setting_id from the resolved tray_info_idx via
filament_id_to_setting_id when the client sent none (e.g. GFB99 -> GFSB99),
mirroring the derivation the inventory/assignment path already does. Doing
it server-side also protects API callers and future frontends. P* user
presets and already-GFS* values are left unchanged, and an explicit
setting_id still passes through untouched.
The AMS merge clears a tray on a partial {id, state} update when state != 11
(the 4-slot AMS "emptied slot" signal, #784). An AMS-HT (single-tray high-temp
dry box, id >= 128) reports its loaded tray as state=9, so the partial the
printer sends on power-on was misread as "emptied" and wiped the HT-A spool's
tray_type/RFID/assignment seconds after power-on.
Skip the state-heuristic for HT units (id >= 128). Genuine HT removal still
clears via the explicit tray_type="" update and tray_exist_bits cleanup;
regular AMS (id < 128) is unchanged.
start_print() published project_file guarding only on connection state, so a
re-dispatch onto a printer that had already started — e.g. a watchdog revert
(#2555) after the printer sat in FINISH past accepting the job — collided with
the live print. The firmware answers 0500_4004 ("Device is busy and cannot
start a new task"), which on an A1 mini cancels the running job.
Defense-in-depth at the paths that can reach a busy printer:
- bambu_mqtt: refuse to publish project_file when gcode_state is
PREPARE/SLICING/RUNNING/PAUSE and return without sending. This is the one
publish choke point every dispatch path funnels through (queue scheduler,
manual start, webhook, Virtual-Printer forward). IDLE/FINISH/FAILED still
start.
- print_scheduler: re-check the live printer state right before the FTP upload
and defer a busy printer (leave the item pending for a later tick) instead of
uploading and dispatching. If the printer goes busy in the upload window and
the start is refused, revert the item to pending rather than marking it
failed — a busy printer is a deferral, not a failure.
A transport-level MQTT QoS-1 replay on reconnect would bypass the client guard,
but the dispatch/watchdog reconnect path already hard-resets the client with a
fresh session, so it has no inflight project_file to replay.
Three more idle-in-transaction / thundering-herd paths from farm testing:
- print_scheduler: _start_print commits before the FTP delete/upload and
_preheat_and_soak commits before the heat-soak wait, so the per-item
session no longer sits idle-in-transaction across preheat + upload.
- cloud/filament-info: rollback the request transaction after the token
read and before the sequential Bambu Cloud calls; single-flight
concurrent misses for the same setting_id through one shared call.
- printers/cover: coalesce identical in-flight cover requests so followers
serve from the cache the leader fills instead of duplicating the
multi-path FTP + 3MF extraction.
Also adds pool_use_lifo (PostgreSQL default on, DB_POOL_USE_LIFO override,
shown in /system/db-pool) so a bursty farm keeps a small hot connection set.
Two or three concurrent UI logins exhausted the PostgreSQL pool on the
reporter's 93-printer farm: QueuePool limit of size 10 overflow 20 reached,
with all 30 sessions idle in transaction on the auth_enabled SELECT. Three
regressions had landed on dev after an earlier configurable-pool change was
reverted and never re-applied (only the route-by-route session fixes were).
- Pool sizing is env-configurable again (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); the PostgreSQL default returns to
20 + 80 with pool_pre_ping and pool_recycle=1800, and GET
/api/v1/system/db-pool reports resolved config + live gauges without
checking out a connection. SQLite unchanged (20 + 200).
- is_auth_enabled caches for 30s again. Only enabled=True is ever cached, so
a stale read can only fail closed (require auth), never open; set_auth_enabled
invalidates immediately. An autouse test fixture resets the module cache
between tests to keep ordering deterministic.
- Every authenticated request checked out two pooled connections: the
permission dependency held one and the revoked-jti check opened another.
is_jti_revoked now reuses the caller's session; the token dependencies and
the auth-middleware gateway were restructured to open one session and pass
it in, so each request makes a single checkout.
A print sliced against a Virtual Printer carries use_ams=false — a VP
advertises no AMS, so the slicer sends it and VP intake stamps it on the
queue item. But an "Any [model]" item is colour-matched to a real printer
at dispatch, resolving a real AMS slot in ams_mapping. The command builder
only ever forced use_ams off (all-external) and never back on, so the stale
false shipped with a real-tray mapping and the printer aborted at layer 0 on
the empty external spool.
For single-nozzle printers the mapping is now authoritative: a real tray
(0-253) forces use_ams=true, explicit external (254/255) forces it false,
and an unresolved -1 does neither (preserving the #2589 contract).
Dual-nozzle is untouched — use_ams is nozzle routing there. The correction
sits at the single command-builder choke point, covering the VP, queue, and
manual paths.
The firmware's runout HMS text says "insert into the same AMS slot", which is
wrong under AMS Filament Backup: the firmware won't re-accept the depleted slot
and advances to the next compatible one. Bambuddy parsed print.ams.tray_now only
and dropped tray_tar/tray_pre, so the expected slot never reached the UI.
Capture tray_tar/tray_pre on PrinterState and, while paused, resolve them to
global tray IDs (expected_tray/previous_tray) on both the REST and WebSocket
status payloads via a shared resolver: single-AMS passthrough, multi-AMS
snow-mapping resolution, AMS-HT/external passthrough, and an honest null when the
slot can't be placed. The AMS graphic highlights the expected slot (amber) and
the ran-out slot (red); the HMS modal re-describes runout codes to name both,
falling back to "check the printer" when unresolved. Runout copy translated in
all 11 locales.
Reporter @Jostxxl confirmed tray_pre=1/tray_tar=2 during the pause (ran out in
Slot 2, printer expected Slot 3).
reconcile_stale_active_prints closes out stale status="printing" archives by
synthesising an aborted on_print_complete, which logged a PrintLogEntry whose
duration was completed_at - started_at — the whole multi-day disconnect gap,
since a reconciled archive's real end time is unknown. Across a farm of stale
rows this inflated Total Print Time by hundreds of hours, and the Stats total
recomputed the same value from the timestamps even when duration was NULL/0.
Reconciled completions now log duration_seconds=0, the two Stats time paths
trust a stored 0 instead of recomputing, and reconciled aborts get an honest
"Stale - reconciled ..." failure_reason instead of "User cancelled". Genuine
long prints are untouched (no cap; still-running >24h prints aren't stale).
H2C (firmware 01.02.00.00) had no per-model FTP profile and ran on the
Python-default TLS 1.3, hitting the same vsFTPd session-reuse fault the P2S
(#1401) and X2D (#1638) were already capped for. The intermittent FTPS
failure dropped prints to the no-3MF fallback archive, so slice data was
missing — hence no filament in the Print Log and no inventory deduction.
Add an H2C cap_tls_v1_2 profile plus its O1C/O1C2 SSDP aliases. H2D is left
on the default profile (negotiates TLS 1.3 without the fault).
The remaining routes of the idle-in-transaction class: the file-manager,
storage, camera-snapshot and timelapse routes each took their printer row
via Depends(get_db) and then talked FTP/camera on the same held session, so
a farm dashboard polling cover/snapshot tiles (offline printers included)
crept the pool to exhaustion over ~23h. They now read in a short session and
release before the I/O; timelapse re-opens a fresh session only for the write.
Also caps the four bare-executor FTP helpers with asyncio.wait_for so a
saturated 48-worker pool can't pin a caller (and its DB connection)
indefinitely, and runs the synchronous smtplib send off the event loop with
an explicit timeout so a wedged relay can't freeze the loop.
A P1S queue row with use_ams=true but ams_mapping=[-1] was silently
printed with no AMS, starting against the empty external feed and pausing
with a runout. Two faults combined:
- start_print treated -1 (unresolved) the same as >=254 (explicit
external) when deciding to force use_ams=False. Only genuine external
now downgrades; -1 never does.
- The scheduler trusted a stored [-1] as "already resolved" and passed it
through. It now recomputes from live AMS trays whenever the stored
mapping is entirely unresolved, and clears it if nothing matches rather
than sending a doomed command.
Frontend: the Print dialog no longer serializes an all-[-1] mapping while
the printer status is still loading (the hook returns no mapping), and
submit waits for AMS status with a "Waiting for AMS status" notice.
Tests: new backend + frontend regression coverage; corrected one existing
test that pinned the old [-1] -> use_ams=False behavior.
Pushover rejects priority-2 (Emergency) messages unless they carry retry
and expire. _send_pushover never sent them, so setting priority 2 always
failed with Pushover's "retry and expire are required" error. Now at
priority 2 we send retry/expire (default 60s/3600s, clamped to Pushover's
30-10800s range), surfaced as two provider fields shown only when priority
is 2. Added PushoverConfig schema fields, i18n labels across all locales,
and unit tests.
GET /printers/{id}/cover took its printer row via Depends(get_db), whose
yield-dependency session stays open for the whole request — including the
3MF cover download (up to 8 remote paths x retries with backoff, minutes
under FTP contention). One pooled connection sat idle-in-transaction the
entire time; on a large farm a wall of dashboards drained the pool. The
route now fetches the printer in a short-lived async_session() and releases
the connection before the download (expire_on_commit=False keeps printer.*
readable). Pinned by a signature-inspection guard that fails if get_db is
ever re-added.
fix(print-start): release the DB connection across plate detection and 3MF download (#2572)
on_print_start held one session from top to bottom of the handler, across
two slow I/O blocks that need no database: the plate-detection camera grab
and, on the new-archive path, the multi-path 3MF FTP download (its own
comments cite worst cases of tens of minutes). The connection sat idle-in-
transaction for both, once per starting print. It now commits at each
boundary — only read SELECTs have run on those paths (every write branch
returns earlier), so the commit persists nothing and simply returns the
connection to the pool for the I/O; the next query re-acquires, and
expire_on_commit=False keeps printer.* readable.
fix(startup): connect to printers concurrently so the API serves within seconds (#2572)
init_printer_connections awaited each printer's connection serially, and
connect_printer ends in a fixed 1s settle wait. The MQTT connect is non-
blocking (connect_async + loop_start), so that 1s x fleet size was pure
serial dead air the FastAPI lifespan blocked on before uvicorn began
serving — ~100s before port 8000 responded on a 93-printer farm. The
connections are now started with asyncio.gather, so the step takes ~1s
regardless of fleet size. return_exceptions=True isolates each result: one
unreachable printer no longer aborts the rest, or startup itself.
The Queue listing serialized each item by opening its 3MF and re-parsing
slice_info.config three times (print time, filament usage, bed type) on
every poll, per connected client, even for unchanged files. Add a single
combined extract_plate_metadata_from_3mf() cached by (path, plate_id,
mtime_ns, size); the three legacy helpers delegate to it. An unchanged
queue now does no repeat 3MF parsing.
OrcaSlicer shipped a first-class external-app pairing API (OAuth 2.0 Device
Authorization Grant), so the Supabase-PKCE copy-paste flow is replaced end to
end. Connecting is now: click Connect, approve a short code on the Orca Cloud
settings page, done — no redirect, no callback paste, no client secret, works
from a LAN IP / localhost / behind a proxy.
Backend: services/orca_cloud.py rewritten to device-code request + poll (the
four RFC outcomes) + refresh_token grant + introspection + external sync pull;
routes expose /device/start and /device/poll (device_code kept server-side in
the reused orca_cloud_pending_* columns, no migration). Requests sync:read
(read-only feature). Prod endpoint by default, ORCA_CLOUD_API_BASE overrides
to staging. Wired the shared httpx client (fixes a per-request socket leak).
Frontend: device-code connect UI + api client methods; all 11 locales updated.
The #2575 reconciliation correctly deletes a stale external-spool
assignment in on_ams_change, but did so silently: spool_assignment_changed
was only broadcast by the manual REST assign/unassign endpoints, and the
frontend's spool-assignments cache is invalidated only by that event. So
after an external-spool type swap the DB was correct but every open browser
kept rendering the unlinked spool on the slot until an unrelated refetch —
which the reporter read as "the fix didn't work" (a browser refresh showed
the right state all along).
Broadcast spool_assignment_changed for each auto-unlinked slot after the
commit. No frontend change — the handler already invalidates the cache.
After an RTSP read timeout the stream cleanup killed the stalled ffmpeg
and then awaited process.wait() unbounded. A SIGKILLed ffmpeg stuck in
uninterruptible I/O on a dead RTSP socket can take arbitrarily long to
be reaped, so the fan-out stream coroutine sat parked in that wait (12
hours in the reported case) while every new viewer attached to the
stalled broadcaster and received no frames.
Bound the post-kill wait to 2s in all three places it existed: the
stream generator's _terminate_ffmpeg (the reported hang), the camera
stop endpoint (which would hang the recovery request itself; now uses
the shared helper instead of an inline copy), and the orphan-cleanup
janitor (whose hang would disable the safety net). On timeout the
zombie is abandoned; the janitor's /proc scan reaps it next pass and
the stream proceeds to its normal reconnect.