Commit Graph
2 Commits
Author SHA1 Message Date
maziggy 69b6b5a334 fix(#1150): skip MQTT reconnect on watchdog timeout when project_file landed
Background: P1P firmware can take ~135 s after a project_file MQTT publish
  to actually start parsing the uploaded .3mf — gcode_state stays IDLE and
  subtask_id doesn't advance until parse completes. The dispatch watchdogs
  treated the missed transition as a #887/#936 half-broken session and called
  force_reconnect_stale_session, which interrupts the printer's in-progress
  parse and triggers 0500_4003 ("can't parse print file") on the printer side.

  Both #1150 (slow parse) and #887/#936 (zombie session) look identical from
  state and subtask_id alone — both have stale state and stale subtask_id with
  fresh telemetry. The distinguishing signal is the printer's gcode_file
  field: it updates in push_status when the project_file command actually
  lands on the printer, but stays unchanged when the publish was silently
  swallowed.

  Both watchdogs (_verify_print_response in background_dispatch and
  _watchdog_print_start in print_scheduler) now capture pre_gcode_file from
  printer_manager.get_status() before sending the publish, then on timeout
  compare it against the last good status seen during the poll loop. If the
  file changed, the command landed → log a #1150 warning, skip the forced
  reconnect to avoid 0500_4003 mid-parse. If unchanged, fall through to the
  original force_reconnect_stale_session call so the half-broken-session
  recovery is preserved exactly.

  Caveat documented in code: in a retry-same-file slow-parse scenario the
  gcode_file looks identical pre/post-publish, so the watchdog falls through
  to the reconnect path and the user still hits 0500_4003 on that retry.
  Accepted to avoid breaking the half-broken-session recovery, which is the
  more impactful regression of the two.

  The new pre_gcode_file kwarg has a default of None on both watchdog
  functions, so any caller that doesn't pass it keeps the original
  reconnect-on-timeout behavior verbatim.

  4 new unit tests cover both watchdogs: skip on gcode_file change (#1150
  fix), reconnect when unchanged (#936 protection preserved), skip when
  pre=None and current is non-None (printer just connected), reconnect when
  pre_gcode_file arg is omitted (backward-compat). All 439 existing
  dispatch / scheduler / mqtt tests pass unchanged.
2026-04-28 09:26:25 +02:00
maziggy 9d0418688c fix(#1134): propagate background-dispatch watchdog timeout as job failure
Follow-up to #1042. The post-dispatch watchdog _verify_print_response was
  fire-and-forget — it correctly detected when the printer never transitioned
  (HMS error pending, half-broken MQTT session, plate-clear gate, SD card
  fault) and force-reconnected the MQTT session, but the dispatch job had
  already been marked successful on the optimistic MQTT-publish-acknowledged
  path. The UI carried on showing "Print started successfully" while the
  printer sat idle.

  The watchdog now returns bool and is awaited inline by both call sites in
  _run_reprint_archive and _run_print_library_file. On False the call sites
  raise a RuntimeError carrying a user-actionable message ("Printer did not
  acknowledge print command — state still {pre_state}. Check the printer for
  a pending error...") which routes through the existing _run_active_job →
  _mark_job_finished(failed=True) → background_dispatch WS broadcast path.
  Library-file flow rolls back the freshly-created archive on timeout so no
  phantom row is left behind for a print that never started.

  The watchdog now also accepts subtask_id advancing past pre_subtask_id as a
  definitive "command landed" signal — same as the queue-side watchdog at
  print_scheduler.py:1992 — so slow H2D FINISH→PREPARE transitions (~50 s
  observed) don't false-fail when the printer has clearly accepted the
  project_file but is still in FINISH. Default timeout raised from 15 s to
  90 s to match the queue-side watchdog and give the same headroom on both
  dispatch paths. Brief mid-window MQTT disconnects keep polling instead of
  immediately failing — matches what the queue watchdog already does and
  avoids false-failing on transient telemetry gaps.

  11 new tests in test_background_dispatch_watchdog.py: state-change pickup,
  subtask_id-change pickup with state still FINISH, neither-changed timeout
  plus force_reconnect_stale_session call, pre_subtask_id=None backwards-
  compat, post-dispatch subtask_id=None not counting as a change, brief
  disconnect not short-circuiting the window, persistent disconnect for the
  full window returning False, default-timeout=90s contract, _run_reprint_archive
  raises RuntimeError with the captured pre-state args on watchdog False,
  _run_reprint_archive happy path doesn't rollback, _run_active_job marks
  the job failed with the message when _process_job raises RuntimeError.
2026-04-26 08:58:08 +02:00