A print queued from a specific plate of a multi-plate 3MF showed as Plate 1
in Print History after cancellation: the archive derives its plate from the
filename, but a whole multi-plate 3MF uploads under one name with no plate
suffix, so the parser defaulted to plate 1 and nothing copied the queue
item's plate_id onto the archive (which had no plate field).
Add a nullable print_archives.plate_id, copy it from the queue item at
dispatch (archive- and library-file paths), expose it in the archive API,
and render it in Print History. A startup backfill copies the plate onto
existing archives from their linked queue rows. Column add + backfill are
identical on SQLite and Postgres.
Also fix a related lifecycle bug: stopping a printing item while the printer
was offline left the linked archive stuck at "printing" (queue row
cancelled, but no MQTT completion ever arrives to reconcile the archive).
The offline-stop path now closes the archive out directly; the online path
still defers to the MQTT completion event.
check_queue awaited asyncio.gather() over the whole selected batch before
returning, so the scheduler run loop was blocked until the slowest FTP
upload in the batch finished. On a large farm a 513s upload left 15 of 16
configured upload slots idle for 8.5 minutes while other printers came
free — the setting behaved as a per-batch cap, not a worker pool.
Launch uploads as independent background tasks tracked in a _inflight pool.
Each tick excludes in-flight item rows and their printers from selection,
launches at most limit - len(_inflight) new uploads, and returns
immediately, so a freed slot refills on the next fast tick. The no-double-
dispatch invariant the batch-await provided (rows stay pending until upload
completes) is now carried by the in-flight exclusion; the pending->printing
CAS, busy-printer guard (#2598), per-printer hold, auto-drying exclusion,
and per-item failure isolation are all preserved per task.
Rewrites the concurrent-dispatch tests around pool/reservation/refill
semantics and adds coverage for slot refill, in-flight exclusion, and the
non-blocking return.
The Configure AMS Slot modal sends built-in / local / Orca-generic presets
with a GF* tray_info_idx but an empty setting_id, and configure_ams_slot
forwarded that empty value to ams_filament_setting. The firmware treats a
filament-id-without-setting-id slot as half configured: it shows the new
material briefly, then reverts to its previously stored profile.
Back-fill setting_id from the resolved tray_info_idx via
filament_id_to_setting_id when the client sent none (e.g. GFB99 -> GFSB99),
mirroring the derivation the inventory/assignment path already does. Doing
it server-side also protects API callers and future frontends. P* user
presets and already-GFS* values are left unchanged, and an explicit
setting_id still passes through untouched.
The AMS merge clears a tray on a partial {id, state} update when state != 11
(the 4-slot AMS "emptied slot" signal, #784). An AMS-HT (single-tray high-temp
dry box, id >= 128) reports its loaded tray as state=9, so the partial the
printer sends on power-on was misread as "emptied" and wiped the HT-A spool's
tray_type/RFID/assignment seconds after power-on.
Skip the state-heuristic for HT units (id >= 128). Genuine HT removal still
clears via the explicit tray_type="" update and tray_exist_bits cleanup;
regular AMS (id < 128) is unchanged.
start_print() published project_file guarding only on connection state, so a
re-dispatch onto a printer that had already started — e.g. a watchdog revert
(#2555) after the printer sat in FINISH past accepting the job — collided with
the live print. The firmware answers 0500_4004 ("Device is busy and cannot
start a new task"), which on an A1 mini cancels the running job.
Defense-in-depth at the paths that can reach a busy printer:
- bambu_mqtt: refuse to publish project_file when gcode_state is
PREPARE/SLICING/RUNNING/PAUSE and return without sending. This is the one
publish choke point every dispatch path funnels through (queue scheduler,
manual start, webhook, Virtual-Printer forward). IDLE/FINISH/FAILED still
start.
- print_scheduler: re-check the live printer state right before the FTP upload
and defer a busy printer (leave the item pending for a later tick) instead of
uploading and dispatching. If the printer goes busy in the upload window and
the start is refused, revert the item to pending rather than marking it
failed — a busy printer is a deferral, not a failure.
A transport-level MQTT QoS-1 replay on reconnect would bypass the client guard,
but the dispatch/watchdog reconnect path already hard-resets the client with a
fresh session, so it has no inflight project_file to replay.
Three more idle-in-transaction / thundering-herd paths from farm testing:
- print_scheduler: _start_print commits before the FTP delete/upload and
_preheat_and_soak commits before the heat-soak wait, so the per-item
session no longer sits idle-in-transaction across preheat + upload.
- cloud/filament-info: rollback the request transaction after the token
read and before the sequential Bambu Cloud calls; single-flight
concurrent misses for the same setting_id through one shared call.
- printers/cover: coalesce identical in-flight cover requests so followers
serve from the cache the leader fills instead of duplicating the
multi-path FTP + 3MF extraction.
Also adds pool_use_lifo (PostgreSQL default on, DB_POOL_USE_LIFO override,
shown in /system/db-pool) so a bursty farm keeps a small hot connection set.
Two or three concurrent UI logins exhausted the PostgreSQL pool on the
reporter's 93-printer farm: QueuePool limit of size 10 overflow 20 reached,
with all 30 sessions idle in transaction on the auth_enabled SELECT. Three
regressions had landed on dev after an earlier configurable-pool change was
reverted and never re-applied (only the route-by-route session fixes were).
- Pool sizing is env-configurable again (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); the PostgreSQL default returns to
20 + 80 with pool_pre_ping and pool_recycle=1800, and GET
/api/v1/system/db-pool reports resolved config + live gauges without
checking out a connection. SQLite unchanged (20 + 200).
- is_auth_enabled caches for 30s again. Only enabled=True is ever cached, so
a stale read can only fail closed (require auth), never open; set_auth_enabled
invalidates immediately. An autouse test fixture resets the module cache
between tests to keep ordering deterministic.
- Every authenticated request checked out two pooled connections: the
permission dependency held one and the revoked-jti check opened another.
is_jti_revoked now reuses the caller's session; the token dependencies and
the auth-middleware gateway were restructured to open one session and pass
it in, so each request makes a single checkout.
A print sliced against a Virtual Printer carries use_ams=false — a VP
advertises no AMS, so the slicer sends it and VP intake stamps it on the
queue item. But an "Any [model]" item is colour-matched to a real printer
at dispatch, resolving a real AMS slot in ams_mapping. The command builder
only ever forced use_ams off (all-external) and never back on, so the stale
false shipped with a real-tray mapping and the printer aborted at layer 0 on
the empty external spool.
For single-nozzle printers the mapping is now authoritative: a real tray
(0-253) forces use_ams=true, explicit external (254/255) forces it false,
and an unresolved -1 does neither (preserving the #2589 contract).
Dual-nozzle is untouched — use_ams is nozzle routing there. The correction
sits at the single command-builder choke point, covering the VP, queue, and
manual paths.
The firmware's runout HMS text says "insert into the same AMS slot", which is
wrong under AMS Filament Backup: the firmware won't re-accept the depleted slot
and advances to the next compatible one. Bambuddy parsed print.ams.tray_now only
and dropped tray_tar/tray_pre, so the expected slot never reached the UI.
Capture tray_tar/tray_pre on PrinterState and, while paused, resolve them to
global tray IDs (expected_tray/previous_tray) on both the REST and WebSocket
status payloads via a shared resolver: single-AMS passthrough, multi-AMS
snow-mapping resolution, AMS-HT/external passthrough, and an honest null when the
slot can't be placed. The AMS graphic highlights the expected slot (amber) and
the ran-out slot (red); the HMS modal re-describes runout codes to name both,
falling back to "check the printer" when unresolved. Runout copy translated in
all 11 locales.
Reporter @Jostxxl confirmed tray_pre=1/tray_tar=2 during the pause (ran out in
Slot 2, printer expected Slot 3).
reconcile_stale_active_prints closes out stale status="printing" archives by
synthesising an aborted on_print_complete, which logged a PrintLogEntry whose
duration was completed_at - started_at — the whole multi-day disconnect gap,
since a reconciled archive's real end time is unknown. Across a farm of stale
rows this inflated Total Print Time by hundreds of hours, and the Stats total
recomputed the same value from the timestamps even when duration was NULL/0.
Reconciled completions now log duration_seconds=0, the two Stats time paths
trust a stored 0 instead of recomputing, and reconciled aborts get an honest
"Stale - reconciled ..." failure_reason instead of "User cancelled". Genuine
long prints are untouched (no cap; still-running >24h prints aren't stale).
H2C (firmware 01.02.00.00) had no per-model FTP profile and ran on the
Python-default TLS 1.3, hitting the same vsFTPd session-reuse fault the P2S
(#1401) and X2D (#1638) were already capped for. The intermittent FTPS
failure dropped prints to the no-3MF fallback archive, so slice data was
missing — hence no filament in the Print Log and no inventory deduction.
Add an H2C cap_tls_v1_2 profile plus its O1C/O1C2 SSDP aliases. H2D is left
on the default profile (negotiates TLS 1.3 without the fault).
The remaining routes of the idle-in-transaction class: the file-manager,
storage, camera-snapshot and timelapse routes each took their printer row
via Depends(get_db) and then talked FTP/camera on the same held session, so
a farm dashboard polling cover/snapshot tiles (offline printers included)
crept the pool to exhaustion over ~23h. They now read in a short session and
release before the I/O; timelapse re-opens a fresh session only for the write.
Also caps the four bare-executor FTP helpers with asyncio.wait_for so a
saturated 48-worker pool can't pin a caller (and its DB connection)
indefinitely, and runs the synchronous smtplib send off the event loop with
an explicit timeout so a wedged relay can't freeze the loop.
A P1S queue row with use_ams=true but ams_mapping=[-1] was silently
printed with no AMS, starting against the empty external feed and pausing
with a runout. Two faults combined:
- start_print treated -1 (unresolved) the same as >=254 (explicit
external) when deciding to force use_ams=False. Only genuine external
now downgrades; -1 never does.
- The scheduler trusted a stored [-1] as "already resolved" and passed it
through. It now recomputes from live AMS trays whenever the stored
mapping is entirely unresolved, and clears it if nothing matches rather
than sending a doomed command.
Frontend: the Print dialog no longer serializes an all-[-1] mapping while
the printer status is still loading (the hook returns no mapping), and
submit waits for AMS status with a "Waiting for AMS status" notice.
Tests: new backend + frontend regression coverage; corrected one existing
test that pinned the old [-1] -> use_ams=False behavior.
Pushover rejects priority-2 (Emergency) messages unless they carry retry
and expire. _send_pushover never sent them, so setting priority 2 always
failed with Pushover's "retry and expire are required" error. Now at
priority 2 we send retry/expire (default 60s/3600s, clamped to Pushover's
30-10800s range), surfaced as two provider fields shown only when priority
is 2. Added PushoverConfig schema fields, i18n labels across all locales,
and unit tests.
The GitHub runner's Python toolcache ships setuptools 79.0.1, which
pip-audit flags for PYSEC-2026-3447 (fixed in 83.0.0), failing the
blocking Backend Security job. A fix version exists, so upgrade
setuptools in the install step rather than --ignore-vuln. Applied to
both ci.yml (blocking) and security.yml (scheduled scan).
GET /printers/{id}/cover took its printer row via Depends(get_db), whose
yield-dependency session stays open for the whole request — including the
3MF cover download (up to 8 remote paths x retries with backoff, minutes
under FTP contention). One pooled connection sat idle-in-transaction the
entire time; on a large farm a wall of dashboards drained the pool. The
route now fetches the printer in a short-lived async_session() and releases
the connection before the download (expire_on_commit=False keeps printer.*
readable). Pinned by a signature-inspection guard that fails if get_db is
ever re-added.
fix(print-start): release the DB connection across plate detection and 3MF download (#2572)
on_print_start held one session from top to bottom of the handler, across
two slow I/O blocks that need no database: the plate-detection camera grab
and, on the new-archive path, the multi-path 3MF FTP download (its own
comments cite worst cases of tens of minutes). The connection sat idle-in-
transaction for both, once per starting print. It now commits at each
boundary — only read SELECTs have run on those paths (every write branch
returns earlier), so the commit persists nothing and simply returns the
connection to the pool for the I/O; the next query re-acquires, and
expire_on_commit=False keeps printer.* readable.
fix(startup): connect to printers concurrently so the API serves within seconds (#2572)
init_printer_connections awaited each printer's connection serially, and
connect_printer ends in a fixed 1s settle wait. The MQTT connect is non-
blocking (connect_async + loop_start), so that 1s x fleet size was pure
serial dead air the FastAPI lifespan blocked on before uvicorn began
serving — ~100s before port 8000 responded on a 93-printer farm. The
connections are now started with asyncio.gather, so the step takes ~1s
regardless of fleet size. return_exceptions=True isolates each result: one
unreachable printer no longer aborts the rest, or startup itself.
pkgs.tailscale.com intermittently returns 504, which aborted the whole
image build even though the Tailscale CLI is optional (the code falls
back to self-signed without it). Retry the fetch, and on sustained
failure continue building without the CLI instead of failing.
The Queue listing serialized each item by opening its 3MF and re-parsing
slice_info.config three times (print time, filament usage, bed type) on
every poll, per connected client, even for unchanged files. Add a single
combined extract_plate_metadata_from_3mf() cached by (path, plate_id,
mtime_ns, size); the three legacy helpers delegate to it. An unchanged
queue now does no repeat 3MF parsing.
OrcaSlicer shipped a first-class external-app pairing API (OAuth 2.0 Device
Authorization Grant), so the Supabase-PKCE copy-paste flow is replaced end to
end. Connecting is now: click Connect, approve a short code on the Orca Cloud
settings page, done — no redirect, no callback paste, no client secret, works
from a LAN IP / localhost / behind a proxy.
Backend: services/orca_cloud.py rewritten to device-code request + poll (the
four RFC outcomes) + refresh_token grant + introspection + external sync pull;
routes expose /device/start and /device/poll (device_code kept server-side in
the reused orca_cloud_pending_* columns, no migration). Requests sync:read
(read-only feature). Prod endpoint by default, ORCA_CLOUD_API_BASE overrides
to staging. Wired the shared httpx client (fixes a per-request socket leak).
Frontend: device-code connect UI + api client methods; all 11 locales updated.
The #2575 reconciliation correctly deletes a stale external-spool
assignment in on_ams_change, but did so silently: spool_assignment_changed
was only broadcast by the manual REST assign/unassign endpoints, and the
frontend's spool-assignments cache is invalidated only by that event. So
after an external-spool type swap the DB was correct but every open browser
kept rendering the unlinked spool on the slot until an unrelated refetch —
which the reporter read as "the fix didn't work" (a browser refresh showed
the right state all along).
Broadcast spool_assignment_changed for each auto-unlinked slot after the
commit. No frontend change — the handler already invalidates the cache.
After an RTSP read timeout the stream cleanup killed the stalled ffmpeg
and then awaited process.wait() unbounded. A SIGKILLed ffmpeg stuck in
uninterruptible I/O on a dead RTSP socket can take arbitrarily long to
be reaped, so the fan-out stream coroutine sat parked in that wait (12
hours in the reported case) while every new viewer attached to the
stalled broadcaster and received no frames.
Bound the post-kill wait to 2s in all three places it existed: the
stream generator's _terminate_ffmpeg (the reported hang), the camera
stop endpoint (which would hang the recovery request itself; now uses
the shared helper instead of an inline copy), and the orphan-cleanup
janitor (whose hang would disable the safety net). On timeout the
zombie is abandoned; the janitor's /proc scan reaps it next pass and
the stream proceeds to its normal reconnect.
A queue item's "Any <model>" button labeled itself from the file's slice
metadata while the scheduler used the row's target_model, so an X1C-sliced
item targeting H2D showed "Any X1C" above "assign to first idle H2D". The
mismatch itself was created silently: sliced-for metadata loads async, and
switching to model mode before it arrived pre-selected the alphabetically
first model (H2D on a mixed farm), after which the model dropdown hid
itself. Nothing validated compatibility, so the scheduler would hand X1C
G-code to an H2D.
Frontend: never default the target silently, keep the dropdown visible in
model mode (incompatible models disabled), label from the actual target,
warn on mismatch, block submit when incompatible.
Backend: new GCODE_COMPAT_FAMILIES table (X1/X1C/X1E/P1P/P1S interchange;
everything else exact-match; missing metadata never blocks). Queue create
and update reject incompatible targets with 400; the scheduler holds back
pre-existing mismatched rows with an actionable waiting_reason instead of
dispatching them.
Manual jog could drive an axis past its travel limit into a collision.
Instrumenting the exact G-code to an H2D showed Bambuddy sending a clean
move at the limit (G91 / G1 Z-1.00 F600 / G90, no M211) that the printer
ran straight past, while its own touchscreen refuses the identical move.
This is a Bambu firmware bug: soft endstops are not enforced on G-code
received over MQTT, and no axis position is reported, so the move cannot
be clamped firmware- or client-side from position.
Two changes: (1) jogs no longer wrap moves in M211 S0/S1 — that disabled
the firmware's soft endstops globally, breaking even the touchscreen's
limits until a power cycle; a bare move keeps the touchscreen protected.
(2) The jog panel shows a prominent warning that travel limits are not
enforced during manual moves due to the firmware bug. Client-side
dead-reckoning enforcement is tracked separately.
Assigning a new filament to the external spool (e.g. generic ABS over
generic TPU) left the previous inventory spool assigned. The reconciliation
that unlinks a stale external-spool assignment lives in on_ams_change, but
that callback only fired on regular AMS-unit changes: its change-hash never
included the external spool (vt_tray/vir_slot), and the external-spool data
is stored after the AMS handler runs.
Detect external-spool identity changes (type, colour, tag, or reset to
empty) and re-fire on_ams_change so the stale assignment is unlinked. The
fill percentage (remain) is excluded from the fingerprint so a running
print doesn't trigger it on every push.
The progress-milestone and HMS-error notification paths in
on_printer_status_change held a session across the ~15s camera snapshot
taken for the notification image, pinning a pooled connection per
milestone/error per printer.
The snapshot needs no DB: read the printer in a short session, release
it, grab the snapshot with none held, then open a fresh session for the
notification send (and lift the db-free MQTT publish out too). Pinned by
a test that fails if the snapshot runs while a session is open.
The background finish-photo task held one session open across the whole
capture pipeline (timelapse extraction, up to 20s stage-22 wait,
external-camera/RTSP grab) — tens of seconds of a pooled connection
idle-in-transaction per finishing print.
Read the setting/printer/archive in a short session, release it, run the
capture with no session held, then re-open a fresh short session only to
append the photo. Logic unchanged.
_scan_for_timelapse_with_retries opened one session and held it across
the FTP directory listing and the multi-MB video download — once per
retry attempt, per completed print — pinning a pooled connection
idle-in-transaction for the whole transfer.
Read the archive + printer in a short session, release it, do the FTP
list/download with no session held, then re-open a fresh short session
only to attach the file. Existing scan tests already cover the
read/download/attach path.
/camera/stream took its printer row via Depends(get_db). get_db is a
yield dependency, so its session stayed open until the response body
finished streaming — for a live MJPEG stream, as long as the browser
tab is open (hours). Every open camera tile pinned one pooled DB
connection idle-in-transaction, draining the pool on large farms.
Fetch the printer in a short-lived async_session() and release the
connection before returning the StreamingResponse. expire_on_commit=
False keeps the already-loaded columns readable during the stream.
Large PostgreSQL farms exhausted the fixed pool (pool_size=10 +
max_overflow=20): with ~93 printers every connection sat idle in
transaction and unrelated requests waited out the 30s pool timeout or
failed in the auth middleware.
- Make pool sizing env-configurable (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); raise the Postgres default to
20 + 80 with pool_pre_ping + pool_recycle=1800.
- Cache the auth_enabled probe (30s) to drop a per-request DB round-trip.
Only enabled=True is cached, so staleness fails closed; set_auth_enabled
invalidates immediately.
- Add GET /api/v1/system/db-pool exposing resolved config + live
checked_out/checked_in/overflow gauges without consuming a connection.
Session-hygiene (connections held across MQTT/FTP/camera/3MF I/O) is a
separate follow-up.
A slot mapped to a different filament than it was sliced for (PLA slice
routed to the only loaded PETG slot) was logged under the sliced material
in the archive, Print Log and material stats, even though the correct spool
was debited. Once usage tracking resolves every used slot to a spool, adopt
the spool's material as the archive filament_type, exactly as the spool
colour is already adopted (#1494). All-or-nothing; both inventory backends;
flows through to the Print Log and stats. No schema/UI/i18n change.
The scheduler slept a fixed 30s after every pass, so each printer that
freed up during a batch waited up to a full interval before its next job
was dispatched — on a farm, that idle gap stacked into the "several long
minutes" reporters saw between requesting prints and them starting (#2555).
check_queue() now reports whether it dispatched anything; run() loops again
after 3s on a productive pass and falls back to 30s otherwise. Fast ticks
only continue while the queue is actively draining, so this can't tight-loop:
a pass that dispatches nothing (all pending items behind busy printers, or a
wedged head-of-line job holding its printer) reverts to the normal interval.
The single-connection barrier from the last round was correct and was being
bypassed. shutdown_broadcaster() popped the broadcaster out of the registry
and only then awaited its teardown, so while the socket was still closing the
slot sat empty: a /camera/stream request landing in that window minted a
broadcaster with no predecessor and dialled port 6000 immediately. A page
reload fires /camera/stop and the new stream request concurrently, so a P1S
ended up holding two connections, kept feeding the orphan, and starved the
live viewer until its TCP keepalive reaped the dead one ~20 min later. The
stopped broadcaster now stays in the registry so the successor chains behind
its socket close.
The camera page also rendered the <img> src before the stream token arrived
whenever auth was disabled, then swapped it once the token landed — aborting
the in-flight request and issuing a second one. With auth off both reached the
backend, so every load attached two viewers to a one-socket printer. The src
now waits for the token query to settle.
Subscribers only checked for client disconnect after yielding a frame or on a
30s idle timeout, so a viewer that left during a black stream stayed counted —
and /camera/stop trusts that count to decide whether to tear the upstream down.
An expired token was indistinguishable from a working one. set_token()
stamped token_expiry = now + 30 days every time a stored token was loaded,
so the expiry reset on every request and is_authenticated could never
return False. /cloud/status answered "connected" for as long as any token
existed, while every cloud call 401'd — and the user was shown Bambu's own
{"error": "Please login."} verbatim.
Bambu is now the authority: /cloud/status validates the token upstream
(cached 5m), and any 401 from any authenticated call durably records the
credential as dead via users.cloud_token_invalid_at, so MakerWorld, cloud
profiles, slicer presets and firmware checks all agree at once. An
unreachable Bambu is treated as unknown, never as expired, so an outage
cannot sign a working session out.
The user-facing message now names the Profiles page, where the Bambu Cloud
sign-in actually lives; the old text pointed at a Settings page that does
not exist. Same stale path corrected in the wiki.
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.
Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.
FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.
Two problems the same bundle exposed:
A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.
The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.
Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).
Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
The ghcr pulls chart used a symlog y axis with a hardcoded tick list, both
copied from github-repo-stats, which builds the rest of the report. Those
settings suit views and clones - small, spiky, frequently zero - but not
container pulls, which sit in a tight band far above zero.
Consequences on the published report: the series (7,946 to 16,160) lives
entirely inside the top decade of the log scale, so a 2x swing rendered as a
14px wobble on a 200px chart and read as a flat line. Of the nine fixed ticks,
six were squashed against the baseline and 50000 fell outside the domain and
never drew, leaving a single usable gridline.
Switch to a linear scale, drop the fixed ticks so Vega derives them from the
actual domain, and format labels with SI prefixes (5k / 10k / 15k). The zero
baseline and the 10% headroom are unchanged, so the axis stays honest; the same
swing now spans 92px and the growth from ~9k to ~15k pulls/day is legible.
queueing plates we cannot map (#2552)
The override panel disappeared for a multi-plate selection in Any [model]
mode, but only once the dialog had been opened before -- which the reporter
saw as "after the file was queued or printed". The filament requirements are
keyed on the selected plate, which is null as soon as two plates are ticked.
On a cold cache the modal cannot yet tell the file is multi-plate and fetches
the whole file's requirements for one render; the panel rendered from that
union. On a warm cache it knows from the first render, the whole-file fetch
never runs, and the panel had nothing to render. Visibility was decided by a
cache race, and the "working" case listed filaments from plates the user had
not selected.
Model mode now renders one panel per selected plate from that plate's own
requirements, and each queued plate carries only the overrides for the slots
it prints, so a colour forced on one plate no longer blocks another.
Reviewing the per-plate machinery turned up four more holes, all closed here:
a manual tray pick survived a change of printer, and a global tray id names a
different spool on a different machine; a plate whose filaments could not be
read was indistinguishable from one needing none and was queued with neither
mapping nor forced colours, so Print now waits for every selected plate to
answer and names the one it cannot read; the insufficient-filament check still
weighed the whole file against a mapping the plates no longer use, and now
follows what each plate dispatches, summing demand per tray; and the
per-printer tray editor no longer appears for a multi-plate fan-out, where its
choices were collected and then discarded.
Selecting several plates hid the filament mapping panel but did not stop the
modal sending a mapping. With no single plate selected it fell back to the
whole file's filament list -- the union of every plate -- and matched against
that. Tray assignment is stateful, so where plate 1 prints red on slot 1 and
plate 2 prints red on slot 2, slot 1 claimed the only red spool and slot 2 fell
through to a type-only match on black. That one mapping went out with every
plate, and the scheduler uses a stored mapping verbatim, so plate 2 printed in
the wrong colour -- decided by a panel the user never saw.
Fetch each selected plate's requirements and map them separately: one panel per
plate, named after it, with its own tray overrides, and each queue item carries
its own plate's mapping. A fan-out across several printers would be a panel per
plate per printer, so those items carry no mapping and the scheduler maps each
plate against the printer it picks. Model mode is unchanged -- no printer means
no trays to map onto.
The tray matcher existed twice and this needed a third caller, so extract it
once and have both existing paths delegate; its 62 tests pass unchanged.
The bug only reproduces with a realistic query cache -- the shared test harness
sets gcTime: 0, which evicts the union and makes the modal look innocent -- so
the new modal tests bring their own client.
Queueing several plates of one 3MF built a single filament-override list from
every selected plate and posted that same list with each plate's item. A
force_color_match entry blocks dispatch until the printer has that exact colour
loaded, so a single-colour plate waited on the whole batch's palette. The same
shared list also widened required_filament_types, making a PLA plate refuse
every printer that lacked a sibling plate's PETG.
Narrow the overrides to the slots the plate actually consumes, on create and on
update -- in the backend, where the 3MF is, so it holds for every writer of the
queue. Dispatch already re-parsed requirements per plate and keyed overrides by
slot, so the dropped entries were inert there. When the plate's slots cannot be
read the overrides are kept whole: an item waiting on a colour it does not need
is visible, one that silently lost a forced colour prints in the wrong filament.
Items queued before this would stay stuck with a waiting reason that explains
nothing, so a startup migration re-scopes the pending ones. Printing and
finished items keep their overrides -- that is a record of what they dispatched
with, not an instruction.
The edit dialog is shared between the projects list and the project detail
page and seeds itself from whichever project object it is handed. The list
payload never carried tags, due_date or priority, so editing from the list
showed a blank tags field -- and, unreported, submitted the dialog's default
priority over a stored high/urgent one. The component read those fields
through a cast, so the compiler never flagged that they were always absent.
Put them on ProjectListResponse and ProjectListItem, drop the casts, and let
an explicit null clear tags and due date the way it already clears budget and
url -- an emptied field was previously sent as undefined and silently reverted.
The template list was missing target_parts_count, which the same dialog edits.