mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-09-30 11:12:35 +02:00
Reprinting from archives sometimes failed immediately with a MicroSD R/W
exception, with the printer's MQTT push referencing a 3MF from a
different unrelated archive. Once it started, every subsequent reprint
hit the same error until the container was restarted.
Root cause from @smandon's support package: paho-mqtt's client-side QoS
1 queue. When the printer's command channel goes half-broken (telemetry
flowing, publishes silently dropped — same #887/#936 pattern),
background_dispatch.py:993 hits its 15s deadline and calls
force_reconnect_stale_session(). That function was force-closing the
underlying socket so paho's auto-reconnect would kick in, but the same
mqtt.Client instance, same client_id, and same in-process QoS 1 queue
stayed alive across the reconnect. Any unacked publish from the broken
session — typically the just-sent project_file for the new archive —
got replayed verbatim on the new connection. The queue accumulates
across multiple stuck dispatches in one Python process, so by the
second or third stuck reprint there were several stale
project_file/resume/stop/clean_print_error commands queued together;
the printer latched onto whichever stale path it processed last,
couldn't find the file on its SD card, and emitted 0500_4003. Container
restart was the only thing that wiped paho's in-process queue.
Replaced socket-close with a context-aware reconnect via a new
_reset_client_for_reconnect() router:
Async-context callers (dispatch deadline, FastAPI handlers via
check_staleness) → hard-reset: client.disconnect() (broker drops
session, clean_session=True), client.loop_stop() (kills paho's
network thread and its queue), null _client, fresh connect() with
incremented client_id. New connection is genuinely empty, no replay.
Paho-network-thread callers (dev-mode probe + ams_filament_setting
zombie detection inside _update_state) → socket-close fallback.
loop_stop() from inside the network thread would self-join and
deadlock, so the safe pattern there is "close the socket and let
paho's loop detect it and auto-reconnect on the same client".
Routing decision uses asyncio.get_running_loop() — paho's callback
thread has no loop, every legitimate hard-reset caller does.
7 regression tests:
- TestForceReconnectRouting (3): sync-context → socket-close fallback,
async-context → hard-reset with disconnect()+loop_stop()+null,
state-disconnected broadcast fires once on either path
- TestHardResetClientDirect (3): helper directly — old client gets
disconnect()+loop_stop(), _client cleared, failing disconnect()
doesn't propagate so background_dispatch's await chain can't break
- TestZombieSessionDetection / TestDeveloperModeProbeTimeout (updated):
paho-thread context still goes through socket-close, preserving the
legacy contract for those paths