mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-10-08 23:21:58 +02:00
fix(mqtt): lift paho inflight ceiling to prevent QoS=1 session wedge (#1164)
paho's default max_inflight_messages=20 silently fills on Bambu's broker after ~16-20 cumulative commands per session, leaving publish() returning success while packets sit in paho's internal queue. force_reconnect heals it because the inflight queue is per-session, but it costs one wasted user action to trigger. Lifting the ceiling to 1000 keeps QoS=1 untouched (deliberately chosen for cross-model reliability — A1, P1S, X1C, H2D, P2S, X2D all need it) and removes the inflight queue as the bottleneck without changing wire-protocol behaviour. The 0.2.4b2 watchdog reconnect stays as defence-in-depth. Diagnosis credit: RosdasHH's QoS=1/0/2 bisect on #1164.
This commit is contained in:
+1
-1
@@ -39,7 +39,7 @@ All notable changes to Bambuddy will be documented in this file.
|
||||
|
||||
- **Printer card always shows the first plate's thumbnail when printing a multi-plate 3MF** ([#1166](https://github.com/maziggy/bambuddy/issues/1166), reported by @smandon) — On printers running firmware that drops the plate path from `print.gcode_file` (the reporter's case: P1S 01.10.00.00, but the same shape appears on other firmware revisions), the printer reports `gcode_file: MyModel.3mf` instead of `gcode_file: /Metadata/plate_4.gcode`. The `/printers/{id}/cover` route's regex (`plate_(\d+)\.gcode`) found nothing in the bare `.3mf` filename, defaulted to plate 1, and the printer card showed `Metadata/plate_1.png` from the 3MF — even though the user dispatched plate 4. Same problem hit `current_plate_id` on the status response (printer card detail row showed plate 1). Two-pronged fix on a precedence ladder: **(1) Bambuddy now records the plate it dispatched** — `start_print()` writes `(dispatched_plate_id, dispatched_subtask)` onto `PrinterState` at publish time, and a new `resolve_plate_id(state)` helper prefers that record over the gcode_file regex when `dispatched_subtask == state.subtask_name` (the subtask check rejects stale entries from a prior Bambuddy-dispatched print bleeding into a Studio-direct dispatch). **(2) After the 3MF lands on disk, the cover route scans the zip for a unique `Metadata/plate_*.gcode` entry**: per-plate archives sliced separately in Bambu Studio bundle thumbnails for *every* plate but only the *active* plate's gcode, so a single match unambiguously identifies the plate even when no Bambuddy dispatch exists (Studio-direct flow). Final fallback is plate 1, unchanged. The cover-byte cache key was also simplified — `plate_num` was removed from the key now that resolution is late-bound; `clear_cover_cache()` already runs on every print start, so different plates of the same project always re-fetch a fresh thumbnail. Coverage: 5 unit tests in `test_printer_manager.py::TestResolvePlateId` (dispatch precedence, stale-subtask guard, gcode regex fallback, default-1 path, missing-subtask guard), 4 unit tests in `test_bambu_mqtt.py::TestStartPrintRecordsDispatchedPlate` (dispatch record set/cleared/overwritten/skipped on disconnect), 2 integration tests in `test_printers_api.py` (dispatch wins over plate-1 default; 3MF-scan fallback for per-plate archive without dispatch). Studio-direct multi-plate prints (no dispatch record AND multiple plate gcodes in the 3MF) still default to plate 1 — matches the firmware's own ambiguity, not regressed by this change.
|
||||
|
||||
- **AMS slot configuration intermittently fails to reach the printer after several configs in a row** ([#1164](https://github.com/maziggy/bambuddy/issues/1164), reported by @RosdasHH) — Configuring AMS slots a handful of times (the reporter saw it almost every 6th change) would silently stop reaching the printer; ~1 minute later the filament colours on the printer would briefly jump between slots, then settle. Root cause was the zombie-session watchdog at `bambu_mqtt.py:861` introduced for [#887](https://github.com/maziggy/bambuddy/issues/887). When an `ams_filament_setting` response took >10 s (normal under load — concurrent K-profile fetches, busy printer, network jitter) the watchdog incremented an `_ams_cmd_unanswered` counter and zeroed `_last_ams_cmd_time` so it wouldn't re-trigger on the next status push. The bug: the response handler that reset the counter was guarded by `and self._last_ams_cmd_time > 0` — so when the late response *did* arrive (after the watchdog had already zeroed the timer), the counter stayed armed at 1. The next slow response on any `ams_filament_setting` command — possibly minutes or hours later, on an entirely unrelated config attempt — would take the counter to 2 and trigger `force_reconnect_stale_session()`. The user-visible symptoms match exactly: configs stop landing (because MQTT reconnects mid-publish, dropping the in-flight command and surfacing as `Cannot set AMS filament setting: not connected` if the user retries during the ~1 min reconnect window), then the queued state finally lands when the reconnect completes (the "filament colours jumping around" the reporter described). Fix is to drop the `_last_ams_cmd_time > 0` guard: any `ams_filament_setting` response — late or not — proves the channel is alive, so the counter must reset. Watchdog still trips on a real zombie session (no responses at all for two consecutive >10 s windows). Regression test in `test_bambu_mqtt.py::TestZombieSessionDetection::test_late_response_after_watchdog_clears_counter_issue_1164` simulates the exact sequence (watchdog fires → late response arrives → second slow response on a fresh command) and asserts the counter resets to 0 on the late response and the second command doesn't tip the threshold to 2. Other 10 zombie-detection tests still pass unchanged.
|
||||
- **AMS slot configuration intermittently fails to reach the printer after several configs in a row** ([#1164](https://github.com/maziggy/bambuddy/issues/1164), reported by @RosdasHH) — Configuring AMS slots a handful of times (the reporter saw it almost every 6th change) would silently stop reaching the printer; ~1 minute later the filament colours on the printer would briefly jump between slots, then settle. Root cause was the zombie-session watchdog at `bambu_mqtt.py:861` introduced for [#887](https://github.com/maziggy/bambuddy/issues/887). When an `ams_filament_setting` response took >10 s (normal under load — concurrent K-profile fetches, busy printer, network jitter) the watchdog incremented an `_ams_cmd_unanswered` counter and zeroed `_last_ams_cmd_time` so it wouldn't re-trigger on the next status push. The bug: the response handler that reset the counter was guarded by `and self._last_ams_cmd_time > 0` — so when the late response *did* arrive (after the watchdog had already zeroed the timer), the counter stayed armed at 1. The next slow response on any `ams_filament_setting` command — possibly minutes or hours later, on an entirely unrelated config attempt — would take the counter to 2 and trigger `force_reconnect_stale_session()`. The user-visible symptoms match exactly: configs stop landing (because MQTT reconnects mid-publish, dropping the in-flight command and surfacing as `Cannot set AMS filament setting: not connected` if the user retries during the ~1 min reconnect window), then the queued state finally lands when the reconnect completes (the "filament colours jumping around" the reporter described). Fix is to drop the `_last_ams_cmd_time > 0` guard: any `ams_filament_setting` response — late or not — proves the channel is alive, so the counter must reset. Watchdog still trips on a real zombie session (no responses at all for two consecutive >10 s windows). Regression test in `test_bambu_mqtt.py::TestZombieSessionDetection::test_late_response_after_watchdog_clears_counter_issue_1164` simulates the exact sequence (watchdog fires → late response arrives → second slow response on a fresh command) and asserts the counter resets to 0 on the late response and the second command doesn't tip the threshold to 2. Other 10 zombie-detection tests still pass unchanged. **Follow-up: cumulative session wedge after ~16-20 commands** — the watchdog fix above heals real zombie sessions, but @RosdasHH continued to see the wedge fire on healthy sessions after enough cumulative commands (configs + spool assignments share the same threshold: "8 + 3", "12 + 1", "16 + 0" all tripped it). His QoS=1 vs QoS=0 vs QoS=2 bisect was the breakthrough — the wedge only happens at QoS=1. paho-mqtt's default `max_inflight_messages` is 20, and Bambu's broker has racy PUBACK matching that leaves some inflight slots unreleased per session, so after ~16-20 cumulative commands the queue silently fills and `publish()` returns success while packets sit in paho's internal queue (force_reconnect heals it because the inflight queue is per-session — the printer had already processed every command, it just couldn't receive any new ones until the session reset). Lifted the ceiling to 1000 via `client.max_inflight_messages_set(1000)` immediately after `mqtt.Client()` construction (`bambu_mqtt.py:3074-3079`). Keeps QoS=1 untouched (the cross-model reliability we deliberately chose for AMS configuration — A1, P1S, X1C, H2D, P2S, X2D all need it) and removes the ceiling as the bottleneck without changing wire-protocol behaviour. The watchdog reconnect from the original fix above stays as defence-in-depth for sessions that go truly zombie. Diagnosis credit: @RosdasHH's careful bisect.
|
||||
|
||||
## [0.2.4b1] - 2026-04-29
|
||||
|
||||
|
||||
@@ -3072,6 +3072,13 @@ class BambuMQTTClient:
|
||||
protocol=mqtt.MQTTv311,
|
||||
)
|
||||
|
||||
# Bambu's broker has racy PUBACK matching with paho's QoS=1 inflight
|
||||
# tracking (#1164). The default ceiling of 20 wedges sessions after
|
||||
# ~16-20 cumulative commands; lifting it well above any realistic
|
||||
# session count keeps QoS=1 working without changing wire-protocol
|
||||
# behaviour across printer models.
|
||||
self._client.max_inflight_messages_set(1000)
|
||||
|
||||
self._client.username_pw_set("bblp", self.access_code)
|
||||
self._client.on_connect = self._on_connect
|
||||
self._client.on_disconnect = self._on_disconnect
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user