Background: P1P firmware can take ~135 s after a project_file MQTT publish
to actually start parsing the uploaded .3mf — gcode_state stays IDLE and
subtask_id doesn't advance until parse completes. The dispatch watchdogs
treated the missed transition as a #887/#936 half-broken session and called
force_reconnect_stale_session, which interrupts the printer's in-progress
parse and triggers 0500_4003 ("can't parse print file") on the printer side.
Both #1150 (slow parse) and #887/#936 (zombie session) look identical from
state and subtask_id alone — both have stale state and stale subtask_id with
fresh telemetry. The distinguishing signal is the printer's gcode_file
field: it updates in push_status when the project_file command actually
lands on the printer, but stays unchanged when the publish was silently
swallowed.
Both watchdogs (_verify_print_response in background_dispatch and
_watchdog_print_start in print_scheduler) now capture pre_gcode_file from
printer_manager.get_status() before sending the publish, then on timeout
compare it against the last good status seen during the poll loop. If the
file changed, the command landed → log a #1150 warning, skip the forced
reconnect to avoid 0500_4003 mid-parse. If unchanged, fall through to the
original force_reconnect_stale_session call so the half-broken-session
recovery is preserved exactly.
Caveat documented in code: in a retry-same-file slow-parse scenario the
gcode_file looks identical pre/post-publish, so the watchdog falls through
to the reconnect path and the user still hits 0500_4003 on that retry.
Accepted to avoid breaking the half-broken-session recovery, which is the
more impactful regression of the two.
The new pre_gcode_file kwarg has a default of None on both watchdog
functions, so any caller that doesn't pass it keeps the original
reconnect-on-timeout behavior verbatim.
4 new unit tests cover both watchdogs: skip on gcode_file change (#1150
fix), reconnect when unchanged (#936 protection preserved), skip when
pre=None and current is non-None (printer just connected), reconnect when
pre_gcode_file arg is omitted (backward-compat). All 439 existing
dispatch / scheduler / mqtt tests pass unchanged.
Follow-up to #1042. The post-dispatch watchdog _verify_print_response was
fire-and-forget — it correctly detected when the printer never transitioned
(HMS error pending, half-broken MQTT session, plate-clear gate, SD card
fault) and force-reconnected the MQTT session, but the dispatch job had
already been marked successful on the optimistic MQTT-publish-acknowledged
path. The UI carried on showing "Print started successfully" while the
printer sat idle.
The watchdog now returns bool and is awaited inline by both call sites in
_run_reprint_archive and _run_print_library_file. On False the call sites
raise a RuntimeError carrying a user-actionable message ("Printer did not
acknowledge print command — state still {pre_state}. Check the printer for
a pending error...") which routes through the existing _run_active_job →
_mark_job_finished(failed=True) → background_dispatch WS broadcast path.
Library-file flow rolls back the freshly-created archive on timeout so no
phantom row is left behind for a print that never started.
The watchdog now also accepts subtask_id advancing past pre_subtask_id as a
definitive "command landed" signal — same as the queue-side watchdog at
print_scheduler.py:1992 — so slow H2D FINISH→PREPARE transitions (~50 s
observed) don't false-fail when the printer has clearly accepted the
project_file but is still in FINISH. Default timeout raised from 15 s to
90 s to match the queue-side watchdog and give the same headroom on both
dispatch paths. Brief mid-window MQTT disconnects keep polling instead of
immediately failing — matches what the queue watchdog already does and
avoids false-failing on transient telemetry gaps.
11 new tests in test_background_dispatch_watchdog.py: state-change pickup,
subtask_id-change pickup with state still FINISH, neither-changed timeout
plus force_reconnect_stale_session call, pre_subtask_id=None backwards-
compat, post-dispatch subtask_id=None not counting as a change, brief
disconnect not short-circuiting the window, persistent disconnect for the
full window returning False, default-timeout=90s contract, _run_reprint_archive
raises RuntimeError with the captured pre-state args on watchdog False,
_run_reprint_archive happy path doesn't rollback, _run_active_job marks
the job failed with the message when _process_job raises RuntimeError.