Files
bambuddy/backend/app
maziggy 9d0418688c fix(#1134): propagate background-dispatch watchdog timeout as job failure
Follow-up to #1042. The post-dispatch watchdog _verify_print_response was
  fire-and-forget — it correctly detected when the printer never transitioned
  (HMS error pending, half-broken MQTT session, plate-clear gate, SD card
  fault) and force-reconnected the MQTT session, but the dispatch job had
  already been marked successful on the optimistic MQTT-publish-acknowledged
  path. The UI carried on showing "Print started successfully" while the
  printer sat idle.

  The watchdog now returns bool and is awaited inline by both call sites in
  _run_reprint_archive and _run_print_library_file. On False the call sites
  raise a RuntimeError carrying a user-actionable message ("Printer did not
  acknowledge print command — state still {pre_state}. Check the printer for
  a pending error...") which routes through the existing _run_active_job →
  _mark_job_finished(failed=True) → background_dispatch WS broadcast path.
  Library-file flow rolls back the freshly-created archive on timeout so no
  phantom row is left behind for a print that never started.

  The watchdog now also accepts subtask_id advancing past pre_subtask_id as a
  definitive "command landed" signal — same as the queue-side watchdog at
  print_scheduler.py:1992 — so slow H2D FINISH→PREPARE transitions (~50 s
  observed) don't false-fail when the printer has clearly accepted the
  project_file but is still in FINISH. Default timeout raised from 15 s to
  90 s to match the queue-side watchdog and give the same headroom on both
  dispatch paths. Brief mid-window MQTT disconnects keep polling instead of
  immediately failing — matches what the queue watchdog already does and
  avoids false-failing on transient telemetry gaps.

  11 new tests in test_background_dispatch_watchdog.py: state-change pickup,
  subtask_id-change pickup with state still FINISH, neither-changed timeout
  plus force_reconnect_stale_session call, pre_subtask_id=None backwards-
  compat, post-dispatch subtask_id=None not counting as a change, brief
  disconnect not short-circuiting the window, persistent disconnect for the
  full window returning False, default-timeout=90s contract, _run_reprint_archive
  raises RuntimeError with the captured pre-state args on watchdog False,
  _run_reprint_archive happy path doesn't rollback, _run_active_job marks
  the job failed with the message when _process_job raises RuntimeError.
2026-04-26 08:58:08 +02:00
..
…