mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-09-30 11:12:35 +02:00
The reporter's 19-printer farm started prints "one by one", up to an hour apart. check_queue awaited each dispatch inline, and a dispatch includes the FTP upload, so every printer queued behind every other printer's transfer despite being an independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s, 157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen of those in series is ~80 minutes, and the next upload started 131 ms after the previous one finished. The delay is linear in fleet size, which is why it got worse the more printers he selected. Dispatch is now collected during the (still sequential) selection loop and run concurrently afterwards, capped by queue_max_concurrent_uploads - Settings -> Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate is untouched; only the transfers overlap. The pass still awaits its uploads before returning: _start_print flips the row pending -> printing only after the upload, so an early return would let the next tick re-dispatch the same rows. FTP work moves to its own thread pool. It was on asyncio's default executor - min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which was survivable only while uploads were serial. Two problems the same bundle exposed: A printer that accepts project_file but never starts (#1678) was retried forever: 270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his "printer who, since the morning, still not launch" - and on a farm each lap also eats an upload slot the other printers are waiting on. Attempts are now counted on the queue item; after three it fails with a message pointing at the printer instead of queueing a fourth re-upload. The debug bundle we asked him for held 4m49s of history. The push_status dumps fired on every frame rather than on change - several while their own comment claimed otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five minutes on 19 printers. They now log transitions only. The bundle also read just the live log while three rotated backups sat next to it, under a byte budget four times larger than the file it was reading. Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs (dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap). Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies with no settings row, a failed printer does not cancel its siblings, no early return), 4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating. Each verified to fail against the unfixed code - the first end-to-end log assertion I wrote passed without the fix and had to be tightened.