Files
bambuddy/backend
maziggy 527f8ea471 fix(mqtt): #1136 reprint fails with 0500_4003 SD R/W after stuck dispatch
Reprinting from archives sometimes failed immediately with a MicroSD R/W
  exception, with the printer's MQTT push referencing a 3MF from a
  different unrelated archive. Once it started, every subsequent reprint
  hit the same error until the container was restarted.

  Root cause from @smandon's support package: paho-mqtt's client-side QoS
  1 queue. When the printer's command channel goes half-broken (telemetry
  flowing, publishes silently dropped — same #887/#936 pattern),
  background_dispatch.py:993 hits its 15s deadline and calls
  force_reconnect_stale_session(). That function was force-closing the
  underlying socket so paho's auto-reconnect would kick in, but the same
  mqtt.Client instance, same client_id, and same in-process QoS 1 queue
  stayed alive across the reconnect. Any unacked publish from the broken
  session — typically the just-sent project_file for the new archive —
  got replayed verbatim on the new connection. The queue accumulates
  across multiple stuck dispatches in one Python process, so by the
  second or third stuck reprint there were several stale
  project_file/resume/stop/clean_print_error commands queued together;
  the printer latched onto whichever stale path it processed last,
  couldn't find the file on its SD card, and emitted 0500_4003. Container
  restart was the only thing that wiped paho's in-process queue.

  Replaced socket-close with a context-aware reconnect via a new
  _reset_client_for_reconnect() router:

    Async-context callers (dispatch deadline, FastAPI handlers via
    check_staleness) → hard-reset: client.disconnect() (broker drops
    session, clean_session=True), client.loop_stop() (kills paho's
    network thread and its queue), null _client, fresh connect() with
    incremented client_id. New connection is genuinely empty, no replay.

    Paho-network-thread callers (dev-mode probe + ams_filament_setting
    zombie detection inside _update_state) → socket-close fallback.
    loop_stop() from inside the network thread would self-join and
    deadlock, so the safe pattern there is "close the socket and let
    paho's loop detect it and auto-reconnect on the same client".

  Routing decision uses asyncio.get_running_loop() — paho's callback
  thread has no loop, every legitimate hard-reset caller does.

  7 regression tests:
  - TestForceReconnectRouting (3): sync-context → socket-close fallback,
    async-context → hard-reset with disconnect()+loop_stop()+null,
    state-disconnected broadcast fires once on either path
  - TestHardResetClientDirect (3): helper directly — old client gets
    disconnect()+loop_stop(), _client cleared, failing disconnect()
    doesn't propagate so background_dispatch's await chain can't break
  - TestZombieSessionDetection / TestDeveloperModeProbeTimeout (updated):
    paho-thread context still goes through socket-close, preserving the
    legacy contract for those paths
2026-04-26 15:00:20 +02:00
..
2025-11-28 10:23:59 +01:00