fix(queue): two-phase dispatch watchdog so a printer that accepts project_file but never starts doesn't wedge the queue (#1678)

_watchdog_print_start now treats subtask_id-advance as Phase A
  "command landed", not as final success. Phase B (180s) keeps watching
  for the active-state transition; if it never arrives — printer
  accepted the file but stalled (cloud+LAN re-auth after a power cycle
  on old firmware was the reported trigger) — revert the queue item to
  'pending' instead of leaving it stuck in 'printing' until container
  restart. Phase B skips force_reconnect because subtask_id-advance
  proves the project_file landed; a reconnect mid-parse would trigger
  0500_4003 (#1150). Phase A's H2D 50s FINISH→PREPARE tolerance (#1078)
  is preserved by Phase B's 180s headroom.
This commit is contained in:
maziggy
2026-06-08 07:54:24 +02:00
parent a104ccc95c
commit fdaff37975
3 changed files with 165 additions and 47 deletions
+2
View File
@@ -5,6 +5,8 @@ All notable changes to Bambuddy will be documented in this file.
## [0.2.5b1] - Unreleased
### Fixed
- **Print queue no longer wedges in "Currently Printing" when a printer accepts `project_file` but never starts (#1678, reported by @kleinwareio)** — Reporter on two P1S, one was power-cycled mid-print and came back online; from then on Bambuddy showed the next queue item as "Currently Printing" at 0% while the printer card showed "Idle / Ready to print". The same file also re-appeared in the Queued list as Pending after the user resubmitted. Only restarting the Bambuddy container ever recovered it. Support log + screenshots confirm: at dispatch time MQTT `project_file` was ACK'd, printer pushed `gcode_state=IDLE, gcode_file=<our-file>, subtask_id=<our-submission-id>` — i.e. the file landed on the printer but the printer never transitioned IDLE → PREPARE → RUNNING. **Root cause: `_watchdog_print_start` returned SUCCESS as soon as `subtask_id` advanced.** The subtask_id-as-pickup-signal was added for H2D, which can sit at `FINISH` for ~50 s after accepting `project_file` before flipping to PREPARE (#1078) — but it's strictly a "command landed" signal, not "actually printing". When the printer accepts the file but then wedges (cloud+LAN re-auth dance after a power cycle, old firmware, partial network outage), the watchdog returned success, the queue row stayed at `status='printing'`, the in-memory `_expected_prints` entry stayed registered (TTL is 2 hours and only clears the dict, not the DB row), and every subsequent queue item was blocked because the printer was still "in flight". This reporter's firmware (01.07.00.00, current is 01.08.x+) and `bambu_cloud_token`-enabled cloud+LAN mode make the post-power-cycle wedge measurably more likely on their box, but the queue-wedge bug applies to any printer that accepts a file but stalls before starting. **Fix: split the watchdog into two phases.** Phase A (up to `timeout`, default 90 s, unchanged behaviour) waits for either an active-state transition OR a `subtask_id` advance — if neither happens the publish was lost on a half-broken MQTT session (#887/#936) and we revert + force-reconnect (the original #967 recovery path). Phase B (new, up to `phase_b_timeout`, default 180 s) only runs when Phase A exited via subtask_id-alone: keep watching for the active-state transition. 180 s is ~3.5× the worst observed H2D FINISH → PREPARE delay (#1078), so the H2D path stays green. If Phase B times out the queue item is reverted to `pending` so the user can retry without restarting Bambuddy — and Phase B explicitly does NOT force a MQTT reconnect because subtask_id-advance proves the project_file landed and a forced reconnect mid-parse triggers 0500_4003 (#1150). Phase A's existing `gcode_file`-changed discriminator (#1150) stays put for the no-subtask-id-advance case. **Tests:** `test_scheduler_watchdog.py` 14 cases (up from 13) — the #1078 H2D regression test rewritten to step the status through Phase A (subtask_id advance with state=FINISH) then Phase B (state flips to RUNNING) and pin success; new `test_reverts_when_subtask_advanced_but_state_never_active` pins the #1678 wedge case (subtask_id advances, state stays IDLE for the full Phase B window → revert + NO force_reconnect call); new `test_default_phase_b_timeout_is_180_seconds` pins the new default so a future refactor doesn't silently shrink the H2D headroom. Existing #967 / #1150 / #1370 / disconnect / fallback / discriminator regression coverage all stays green. Wider scheduler + queue + dispatch test surface (305 tests) stays green; ruff clean.
- **Service-worker activate handler no longer hangs first-install browsers (demo site stuck spinner + Firefox Corrupted-Content)** — Reproduced live on the demo platform: a visitor lands on `{session}.demo.bambuddy.cool/`, the Printers page renders, but clicking any sidebar entry sticks the next page on a spinner; only a manual reload recovers. In Firefox the same race surfaces as a "Corrupted Content Error" with `sw.js` stuck in `activating` for the entire session. **Root cause:** the `client.navigate(client.url)` call added to the `activate` handler in `sw.js` (commit `18d534c9`, shipped 2026-06-04 alongside the Orca Cloud landing) was intended to force kiosks running an old SW to reload after a deploy, but its only guard was `client.url && typeof client.navigate === 'function'` — neither distinguishes a first install from an upgrade. On every fresh origin (every demo session is a new subdomain, but also any browser visiting Bambuddy for the first time, or after clearing site data) the activate handler still fired the forced navigation: Chromium raced it against React Router's in-flight SPA mount and wedged the page; Firefox's `event.waitUntil` deadlocked on `await client.navigate(...)` because the SW intercepts its own document fetch while still `activating`, the document load aborts, and the SW never reaches `activated`. The "first install on a never-controlled client" guard the commit's comment claimed simply didn't exist in code. **Fix: split the lifecycle correctly.** `sw.js` activate handler is reduced to cache cleanup + `clients.claim()` (matches the standard PWA lifecycle and lets activation complete in low single-digit ms regardless of in-flight document state). The deploy-pickup reload moves to `sw-register.js`: capture `hadController = !!navigator.serviceWorker.controller` at script load (true ⇔ a previous SW was controlling the document), listen for `controllerchange`, and only `location.reload()` when `hadController` was true. A returning kiosk hits a new deploy → had a controller → reloads as before. A first-install visitor (no prior SW, or hard-refresh, or first demo session) → no controller → no forced navigation → React mount completes cleanly. `CACHE_NAME` bumped `bambuddy-v29 → bambuddy-v30` and `STATIC_CACHE` `bambuddy-static-v28 → bambuddy-static-v29` so existing browsers fetching the new `sw.js` drop the old CacheStorage in the same pass — without the bump the SW file byte content might equal the cached one and the upgrade installs nothing. The SpoolBuddy-kiosk unregister branch at the top of `sw-register.js` is unchanged (still wipes registrations on `/spoolbuddy` paths). The `notificationclick` handler in `sw.js` (open-tab-on-push) still uses `client.navigate(url)` — different code path, unrelated, unchanged.
- **VP archive/queue names with `&` no longer render as `&amp;amp;` + tooltip corrected for BambuStudio 2.7.x reality (#1658 follow-up, reported by @IndividualGhost1905)** — Two bugs surfaced on the same screenshot set: (A) Metadata-mode archive and queue names showed `PCB Vise &amp;amp; Solder Station` where the 3MF's Title metadata is `PCB Vise & Solder Station`. **Root cause:** `ThreeMFParser._parse_3dmodel` (`backend/app/services/archive.py:495-538`) parsed the XML `<metadata name="Title">…</metadata>` payload via regex and stripped whitespace but never called `html.unescape()`. The raw `&amp;` landed in the DB; React then auto-escaped the `&` again on render, producing `&amp;amp;`. The sibling parser `ProjectPageParser` (line 754) already had a loop-until-stable unescape and a comment explaining why ("content is often triple-encoded" — observed BambuStudio behavior), the makerworld-fields path just didn't share it. **Fix:** module-level `import html` and the same loop-until-stable unescape pattern in `_parse_3dmodel`, applied uniformly to all `<metadata>` values so `Title`, `Designer`, and any future fields all get peeled the same way. The loop terminates as soon as `html.unescape()` stops changing the string, so single-, double-, and triple-encoded payloads all converge to the correct value; plain ASCII passes through untouched. (B) Filename-mode showed the slugified project title (`PCB_Vise_&_Solder_Station`) instead of the user-typed Send-dialog text ("Main Parts"). **This is NOT a Bambuddy bug** — BambuStudio source confirms it. `PrintJob.cpp:314-325` (`src/slic3r/GUI/Jobs/PrintJob.cpp`) reads `BBL_DESIGNER_MODEL_TITLE_TAG` (defined as `"Title"` in `bbs_3mf.hpp`) from the 3MF, slugifies it (space → `_`, unusable chars `<>[]:/\|?*"` → `_`, collapse runs of `_`, truncate to 100 chars), and **unconditionally overwrites** the user-typed `m_project_name` with it before sending. `params.project_name` becomes both the FTP filename and the MQTT `subtask_name`. The user-typed string never leaves BambuStudio when a Title metadata exists — there is no MQTT field carrying it, so Bambuddy has no recovery path. The previous tooltip ("handy if you renamed the job in the 'send to printer' dialog") promised something BambuStudio strips, and the previous reply to the reporter dismissed this as "OrcaSlicer-style upload, working as designed" which was wrong on BambuStudio 2.7.1.57. **Fix:** tooltip rewritten in all 11 locales (de / en / es / fr / it / ja / ko / pt-BR / tr / zh-CN / zh-TW) to spell out the BambuStudio behavior — both modes often produce the same string because BS overwrites the Send-dialog name with the 3MF Title field when present. **Tests**: 3 new in `test_archive_service.py::TestThreeMFMetadataHTMLUnescape` — `Title` with `&amp;` unescapes to `&` (the reporter's exact case), `Title` with triple-encoded `&amp;amp;amp;` peels all three layers (the BambuStudio worst-case ProjectPageParser already documents), plain `Title=Benchy` passes through unchanged (regression guard against accidentally munging non-encoded payloads). Full 104-test archive suite green; ruff clean; i18n parity holds (5065 leaves × 11 locales); frontend build clean.
+81 -39
View File
@@ -2315,31 +2315,39 @@ class PrintScheduler:
pre_subtask_id: str | None = None,
pre_gcode_file: str | None = None,
timeout: float = 90.0,
phase_b_timeout: float = 180.0,
poll_interval: float = 3.0,
) -> None:
"""Revert a queue item if the printer never acknowledges the start command.
Bambuddy optimistically marks the queue item as "printing" right after the
MQTT project_file publish succeeds locally. If the printer drops/ignores the
command (half-broken MQTT session — #887/#936), the state never transitions
and the item would otherwise stay stuck in "printing" forever (#967).
MQTT project_file publish succeeds locally. The watchdog runs in two phases:
Exit paths (printer picked up the job — no revert):
- gcode_state changed from pre_state, OR
- subtask_id advanced past pre_subtask_id — the printer echoes our
per-dispatch identity back on push_status, so a subtask_id change is
a definitive "command landed" signal even while state is still FINISH.
H2D can sit at FINISH for ~50 s after accepting project_file before
transitioning to PREPARE, which used to trip the state-only watchdog
and caused the scheduler to revert + re-dispatch the item; the next
successful dispatch then looked like a reprint of the just-finished
job (#1078).
Phase A (up to ``timeout``): wait for either an active-state transition
or a ``subtask_id`` advance past ``pre_subtask_id``. State alone is the
primary signal; subtask_id advance handles the H2D case where state can
sit at FINISH for ~50 s after the printer accepted ``project_file``
before flipping to PREPARE (#1078). If neither happens, the MQTT publish
was lost on a half-broken session (#887/#936) — revert and force
reconnect (the #967 recovery path).
Timeout raised from 45 s → 90 s as belt-and-braces for slow transitions
that also don't emit an early subtask_id tick.
Phase B (up to ``phase_b_timeout``, only if Phase A exited on subtask_id
alone): keep watching for the active-state transition. subtask_id alone
proves the file landed but not that the printer started — and a printer
that accepts the command but stays at IDLE/FINISH indefinitely (e.g.
cloud+LAN re-auth dance after a power cycle on old firmware, #1678)
used to leave the queue item stuck in 'printing' forever because the
old watchdog returned success as soon as subtask_id advanced. If Phase
B times out, revert the queue item so the user can retry without
restarting Bambuddy. Skip ``force_reconnect`` here: the file landed and
a forced reconnect mid-parse triggers 0500_4003 (#1150).
Phase A timeout raised from 45 s → 90 s as belt-and-braces for slow
transitions that also don't emit an early subtask_id tick.
"""
deadline = time.monotonic() + timeout
last_status = None
landed_on_subtask = False
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
await asyncio.sleep(poll_interval)
status = printer_manager.get_status(printer_id)
@@ -2362,14 +2370,28 @@ class PrintScheduler:
scheduler._release_dispatch_hold(printer_id)
return
if pre_subtask_id is not None and status.subtask_id is not None and status.subtask_id != pre_subtask_id:
# Printer picked up the job (subtask_id advanced). H2D can
# sit at FINISH for ~50 s after accepting project_file
# before transitioning to PREPARE, but the subtask_id flips
# to our submission_id almost immediately (#1078).
scheduler._release_dispatch_hold(printer_id)
return
# Phase A exit — printer accepted the file (subtask_id flipped
# to our submission id). Don't return yet: the printer may
# have accepted the command but never actually start (e.g.
# cloud+LAN re-auth dance after a power cycle, #1678). Phase
# B watches for the active-state transition.
landed_on_subtask = True
break
# No transition. Revert the item so the scheduler can retry.
if landed_on_subtask:
phase_b_deadline = time.monotonic() + phase_b_timeout
while time.monotonic() < phase_b_deadline:
await asyncio.sleep(poll_interval)
status = printer_manager.get_status(printer_id)
if not status:
scheduler._release_dispatch_hold(printer_id)
return
last_status = status
if status.state in _ACTIVE_PRINT_STATES:
scheduler._release_dispatch_hold(printer_id)
return
# No active-state transition. Revert the item so the scheduler can retry.
# Drop the in-memory hold so the retry isn't blocked by it.
scheduler._release_dispatch_hold(printer_id)
@@ -2410,24 +2432,44 @@ class PrintScheduler:
# session breaks ongoing prints on the same printer.
return
total_timeout = timeout + (phase_b_timeout if landed_on_subtask else 0.0)
if revert_outcome == "reverted":
logger.warning(
"Queue item %s: printer %d did not respond to print command within "
"%.0fs (state still %s, subtask_id still %s) — reverted to 'pending' "
"for retry (#967)",
queue_item_id,
printer_id,
timeout,
pre_state,
pre_subtask_id,
)
if landed_on_subtask:
logger.warning(
"Queue item %s: printer %d accepted project_file (subtask_id "
"advanced) but never transitioned to an active state within "
"%.0fs — printer wedged post-acceptance; reverted to 'pending' "
"for retry (#1678)",
queue_item_id,
printer_id,
total_timeout,
)
else:
logger.warning(
"Queue item %s: printer %d did not respond to print command within "
"%.0fs (state still %s, subtask_id still %s) — reverted to 'pending' "
"for retry (#967)",
queue_item_id,
printer_id,
timeout,
pre_state,
pre_subtask_id,
)
# Same #1150 / #887/#936 discriminator as background_dispatch: if the
# printer's gcode_file changed since pre-dispatch, the project_file
# command landed and the printer is parsing — a forced reconnect
# mid-parse triggers 0500_4003. If gcode_file is unchanged, the
# publish was silently swallowed (#887/#936) and the original
# force_reconnect recovery is what we want.
# Phase B was entered iff subtask_id advanced, which means the
# project_file landed on the printer. A forced reconnect at this point
# would interrupt the printer's parse and trigger 0500_4003 (#1150) —
# skip the recovery entirely.
if landed_on_subtask:
return
# Phase A timeout path — same #1150 / #887/#936 discriminator as
# background_dispatch: if the printer's gcode_file changed since
# pre-dispatch, the project_file command landed and the printer is
# parsing — a forced reconnect mid-parse triggers 0500_4003. If
# gcode_file is unchanged, the publish was silently swallowed
# (#887/#936) and the original force_reconnect recovery is what we
# want.
client = printer_manager.get_client(printer_id)
current_gcode_file = getattr(last_status, "gcode_file", None) if last_status else None
publish_landed = current_gcode_file is not None and current_gcode_file != pre_gcode_file
+82 -8
View File
@@ -76,11 +76,25 @@ class TestWatchdogExitsEarlyOnPickup:
assert item.status == "printing"
@pytest.mark.asyncio
async def test_exits_on_subtask_id_change_even_if_state_still_finish(self, db_session):
"""Regression for #1078: H2D keeps state=FINISH for ~50 s after accepting
project_file, but subtask_id flips to our new submission_id almost
immediately. That must short-circuit the revert."""
get_status = MagicMock(return_value=_status("FINISH", "NEW_SUBTASK_12345"))
async def test_h2d_finish_to_running_via_subtask_id_then_active_state(self, db_session):
"""Regression for #1078 (preserved through the two-phase rewrite for #1678):
H2D keeps state=FINISH for ~50 s after accepting project_file, but
subtask_id flips to our new submission_id almost immediately. The
watchdog must NOT revert on the basis of state staying at FINISH —
Phase A exits on the subtask_id advance, Phase B then keeps watching
and exits SUCCESS as soon as the printer transitions to PREPARE /
RUNNING within the longer Phase B window.
"""
# First poll: state still FINISH, subtask_id advanced (Phase A → B).
# Second poll: state has flipped to RUNNING (Phase B success).
get_status = MagicMock(
side_effect=[
_status("FINISH", "NEW_SUBTASK_12345"),
_status("RUNNING", "NEW_SUBTASK_12345"),
]
+ [_status("RUNNING", "NEW_SUBTASK_12345")] * 10,
)
with (
patch("backend.app.services.print_scheduler.printer_manager.get_status", get_status),
patch("backend.app.services.print_scheduler.async_session", db_session),
@@ -91,15 +105,16 @@ class TestWatchdogExitsEarlyOnPickup:
pre_state="FINISH",
pre_subtask_id="OLD_SUBTASK_99999",
timeout=0.3,
phase_b_timeout=0.3,
poll_interval=0.05,
)
async with db_session() as db:
item = await db.get(PrintQueueItem, 1)
assert item.status == "printing", (
"subtask_id advanced past pre_subtask_id — the printer accepted our "
"project_file and the watchdog must not revert the queue item even "
"though state is still FINISH (#1078)"
"Phase A exit on subtask_id advance + Phase B observing the "
"active-state transition is the H2D success path — watchdog "
"must keep the item 'printing' (#1078)"
)
@@ -228,6 +243,65 @@ class TestWatchdogRevertsWhenStuck:
sig = inspect.signature(PrintScheduler._watchdog_print_start)
assert sig.parameters["timeout"].default == 90.0
@pytest.mark.asyncio
async def test_default_phase_b_timeout_is_180_seconds(self):
"""Phase B (subtask_id advanced, waiting for active state) must
comfortably exceed the H2D FINISH→PREPARE delay (~50 s observed)
before declaring a printer-side wedge. 180 s gives ~3.5× headroom
and reverts the queue item in well under the previous 2-hour
expected_print TTL (#1678)."""
import inspect
sig = inspect.signature(PrintScheduler._watchdog_print_start)
assert sig.parameters["phase_b_timeout"].default == 180.0
@pytest.mark.asyncio
async def test_reverts_when_subtask_advanced_but_state_never_active(self, db_session):
"""Regression for #1678: P1S on old firmware, power-cycled mid-print,
cloud+LAN re-auth dance in flight. Printer accepts project_file
(gcode_file updates, subtask_id advances to our submission id) but
never transitions from IDLE/FINISH to PREPARE/RUNNING. The pre-fix
watchdog returned SUCCESS as soon as subtask_id advanced and the
queue item stayed in 'printing' until container restart. Phase B now
keeps watching; if the active-state transition never arrives, the
item reverts to 'pending' so the user can retry without restarting.
"""
get_status = MagicMock(
return_value=_status("IDLE", "NEW_SUBTASK_12345", gcode_file="/new.3mf"),
)
client = MagicMock() # NOT None — must verify reconnect isn't called
get_client = MagicMock(return_value=client)
with (
patch("backend.app.services.print_scheduler.printer_manager.get_status", get_status),
patch("backend.app.services.print_scheduler.printer_manager.get_client", get_client),
patch("backend.app.services.print_scheduler.async_session", db_session),
patch("backend.app.core.database.async_session", db_session),
):
await PrintScheduler._watchdog_print_start(
queue_item_id=1,
printer_id=42,
pre_state="IDLE",
pre_subtask_id="OLD_SUBTASK_99999",
pre_gcode_file="/old.3mf",
timeout=0.2,
phase_b_timeout=0.2,
poll_interval=0.05,
)
async with db_session() as db:
item = await db.get(PrintQueueItem, 1)
assert item.status == "pending", (
"subtask_id advanced (Phase A → B) but state never reached an "
"active value — printer-side wedge; the queue item must be "
"reverted to 'pending' (#1678)"
)
assert item.started_at is None
# File landed (subtask_id advance proves this), so a forced reconnect
# would trigger 0500_4003 mid-parse (#1150) — skip.
client.force_reconnect_stale_session.assert_not_called()
class TestWatchdogFallbackBehaviour:
"""Backwards-compat and defensive behaviour around missing data."""