mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-09-30 11:12:35 +02:00
The reporter's A1 had been offline 38 hours and still answered on 8883, so the connection watchdog did exactly what it exists for: rebuild the session with a fresh client, since anything left in the old one's QoS 1 queue would otherwise replay onto the next print (#1136). The rebuild ended in paho's loop_stop(), which sets a terminate flag and then joins the network thread with no timeout. That thread only reads the flag between iterations of loop_forever, so it cannot read it while parked inside reconnect() -> _ssl_wrap_socket() -> do_handshake(). paho gives that handshake the keepalive as its socket timeout -- 30s here -- and a socket timeout is per operation, renewed by every byte the peer sends. A printer that answers TCP and then trickles holds the join open for as long as it likes. The join ran on the asyncio thread. Bambuddy stopped answering anything -- UI, API, /health -- while the process stayed up, which is why a restart: unless-stopped container never restarted. Retiring a client no longer waits for it. The replacement is built at once and the old one is shut down on a thread of its own that nobody joins. Its callbacks are detached first, inline: blocking until the network thread was gone is what used to guarantee a client we had let go of could no longer touch our state, and with the teardown detached a zombie that finishes its handshake would otherwise auto-reconnect and report itself connected behind its replacement's back. disconnect() still goes out, still promptly, because that is what stops paho's auto-reconnect and the replay with it. The reported watchdog is one of six callers. The queue's dispatch recovery and check_staleness -- reached from an ordinary status poll -- share _hard_reset_client; editing, deleting and hand-disconnecting a printer share disconnect(); the relay and smart-plug services had the same join on their shutdown path, where a wedged broker stopped the process from exiting at all. #1445 was this join too, from the add-printer probe, and its off-loop teardown stays as it is. disconnect() stays quiet on the way out, as it always effectively did. paho's callback used to land during the join, but it suppresses itself for a clean disconnect of a printer that reported in the last ten seconds, so a healthy printer disconnected by hand never announced itself offline. Announcing it now would tell the user their printer had gone offline a minute after they disconnected it on purpose (#1752). A retirement that takes more than five seconds logs which printer it was. The whole point is that the next one of these should not have to be diagnosed from a thread dump.