5 Commits
Author SHA1 Message Date
maziggy 222f28c7f8 fix(tests): assemble the upload batch instead of racing a sleep
The concurrent-dispatch tests went red on the Docker shard with
"expected all 6 printers to be uploaded to concurrently, but the
high-water mark was 4", on a scheduler that was dispatching all six
correctly.

peak == 6 is the claim that the sixth dispatch reaches its upload
before the first one finishes, and the dispatches do not arrive
together: each runs a preamble of database work first. So the
assertion was a race between that spread and a fixed 0.15 s sleep.
Measured here the spread is ~14 ms; on the runner it passed 150 ms.
Padding the sleep would only move the threshold and slow every test
that uses it.

_UploadRecorder(assemble=True) now holds each call until every
upload the pass launched has arrived, reading len(_inflight), which
is filled synchronously at launch and is therefore the batch size.
That states the property the peak assertions are about, directly and
with no time in it, and it ends sooner than the sleep it replaces -
the file runs 7.0s to 5.9s. _BATCH_DEADLINE_SECONDS bounds it so a
dispatch that really has gone serial fails on its peak assertion
instead of hanging to the pytest timeout.

test_check_queue_returns_without_awaiting_the_uploads keeps the
sleep, with the reason recorded next to it: it reads the rows while
the uploads are open, and an assembled batch releases as soon as it
is complete, which would let the dispatches flip those rows
mid-assertion.

The test engine also takes the application's own _set_sqlite_pragmas
listener rather than a copy, so the file is opened the way the
running system opens one - WAL, synchronous = NORMAL, a 15 s busy
timeout. SQLite's defaults spend an fsync per commit and lock the
whole file, which made a dispatch's preamble cost more than the
upload it precedes. That is what turned the previous commit's move
to a file-backed database into a visible failure.

Verified in both directions rather than by a green run: with two
printers' preambles delayed by a second, the old recorder reports
exactly the "high-water mark was 4" the runner saw and the shared
library test reports 3 == 4, while the assembling recorder passes
the same scenario. Without the induced skew, ruff is clean, the full
backend suite is green at 12224 passed, and ten consecutive runs of
the file pinned to one core show no failures.
2026-09-25 14:21:03 +02:00
maziggy 330d3e7e6f fix(tests): give each concurrent-dispatch session its own connection
The scheduler's concurrent-dispatch tests failed on CI and passed
locally, on a commit that touched nothing but a text file. Six tests
in tests/unit/test_scheduler_concurrent_dispatch.py went red with
every queue item logged as "Status set to 'printing'" and four of
them read back as "pending", alongside "cannot commit transaction -
SQL statements in progress" and a PrintArchive that could not be
refreshed.

The fixtures built their farm on sqlite+aiosqlite:///:memory:, and
SQLAlchemy backs an in-memory SQLite with a StaticPool: one DBAPI
connection handed to every session, with nothing keeping them apart.
That was harmless while check_queue awaited its uploads inline,
because only one session was ever live at a time. Under the
refillable upload pool the uploads run as concurrent background
tasks with a session each, so their transactions interleave on that
single connection - a sibling session's close() rolls back another's
flushed-but-uncommitted UPDATE, and a commit() landing while another
session still holds a cursor raises the commit error above. Whether
the interleaving lands badly comes down to core count and
interpreter version, which is why a 30-core box on 3.13 stayed green
and a 4-vCPU runner on 3.11 did not.

The four fixtures now put the database in the test's own tmp_path,
which gets an AsyncAdaptedQueuePool and a connection per session -
what the application itself runs with (_resolve_pool_kwargs in
backend/app/core/database.py: pool_size 20, max_overflow 200). So
the harness is being brought in line with production rather than
having its assertions relaxed; nothing in the scheduler changes,
and the concurrency under test was correct throughout.

Checked against the failure mode rather than against a green run: an
isolated repro of the same shape - six writer sessions and one
reader session closing mid-transaction - yields all-pending on the
in-memory engine and all-printing on a file. The new risk is real
SQLite write contention on one file, so the file was run twelve
times pinned to two cores with no failures and no lock errors.

tests/unit/test_scheduler_busy_reasons_3018.py carries the same
harness on an in-memory engine, but dispatches to a single printer,
so it has no second session to race; left as is.
2026-09-25 13:59:12 +02:00
maziggy 4a0b14ed0e fix(queue): make upload concurrency a refillable pool, not a per-batch cap (#2602)
check_queue awaited asyncio.gather() over the whole selected batch before
returning, so the scheduler run loop was blocked until the slowest FTP
upload in the batch finished. On a large farm a 513s upload left 15 of 16
configured upload slots idle for 8.5 minutes while other printers came
free — the setting behaved as a per-batch cap, not a worker pool.

Launch uploads as independent background tasks tracked in a _inflight pool.
Each tick excludes in-flight item rows and their printers from selection,
launches at most limit - len(_inflight) new uploads, and returns
immediately, so a freed slot refills on the next fast tick. The no-double-
dispatch invariant the batch-await provided (rows stay pending until upload
completes) is now carried by the in-flight exclusion; the pending->printing
CAS, busy-printer guard (#2598), per-printer hold, auto-drying exclusion,
and per-item failure isolation are all preserved per task.

Rewrites the concurrent-dispatch tests around pool/reservation/refill
semantics and adds coverage for slot refill, in-flight exclusion, and the
non-blocking return.
2026-07-19 08:34:38 +02:00
maziggy c0a50edbe8 fix(queue): re-check the queue quickly after a dispatch instead of always waiting 30s (issue #2555)
The scheduler slept a fixed 30s after every pass, so each printer that
freed up during a batch waited up to a full interval before its next job
was dispatched — on a farm, that idle gap stacked into the "several long
minutes" reporters saw between requesting prints and them starting (#2555).

check_queue() now reports whether it dispatched anything; run() loops again
after 3s on a productive pass and falls back to 30s otherwise. Fast ticks
only continue while the queue is actively draining, so this can't tight-loop:
a pass that dispatches nothing (all pending items behind busy printers, or a
wedged head-of-line job holding its printer) reverts to the normal interval.
2026-07-15 07:13:22 +02:00
maziggy ce807fb1cc fix(queue): upload to printers in parallel, cap wedge retries, make debug logs survive a farm
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.

Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.

FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.

Two problems the same bundle exposed:

A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.

The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.

Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).

Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
2026-07-14 10:35:58 +02:00