Files
bambuddy/backend
maziggy 0918907dab fix(scheduler): watchdog falsely reverts slow H2D dispatches, causing reprints (#1078)
_watchdog_print_start reverted queue items to "pending" at 45 s if
  gcode_state hadn't changed, assuming the MQTT project_file was swallowed
  by a half-broken session (#887/#967). H2D Pro firmware (01.01.00.00)
  routinely keeps state=FINISH for 48-55 s after actually accepting the
  command before transitioning to PREPARE. The watchdog reverted items
  the printer had already started physically printing; the archive updated
  normally via _active_prints, but the queue item was now "pending" again,
  and the next scheduler tick after plate-clear re-dispatched the same
  item as if it had never run. With one item left in the queue that looked
  like a reprint of the just-finished job; with multiple items the
  symptom was masked by item N+1 getting dispatched during the race.

  Add a second "command landed" signal: subtask_id advancing past the
  pre-dispatch value. Bambuddy already mints a unique submission_id per
  project_file publish (#1042) and the printer echoes it back on the next
  push_status as soon as it starts processing the command - well before
  gcode_state transitions on slow-transition models. _start_print now
  captures pre_subtask_id alongside pre_state and passes both to the
  watchdog, which exits early on either a state change or a subtask_id
  advance.

  Raise default timeout 45 s → 90 s as belt-and-braces for printers that
  neither flip state nor echo subtask_id inside the polling window.
  Genuinely half-broken sessions (both signals unchanged across the full
  90 s) still revert + force-reconnect exactly as before.

  Transient subtask_id=None during reconnect is not mis-detected as a
  change. pre_subtask_id=None falls back to state-only checking so the
  fix is safe for printers that haven't reported a subtask_id yet.

  New test_scheduler_watchdog.py pins the eight behaviours that matter:
  pickup via state change; pickup via subtask_id change with state still
  FINISH (the exact #1078 case); revert when neither signal changes;
  default timeout is 90 s; pre_subtask_id=None state-only fallback;
  current subtask_id=None not treated as change; printer disconnect
  mid-watchdog leaves DB untouched; item that already moved on is not
  clobbered.
2026-04-22 08:38:50 +02:00
..