Files
bambuddy/backend/tests/unit
maziggy 8e9821dd2a fix(mqtt): never wait for a wedged paho network thread (issue #3068)
The reporter's A1 had been offline 38 hours and still answered on 8883, so
the connection watchdog did exactly what it exists for: rebuild the session
with a fresh client, since anything left in the old one's QoS 1 queue would
otherwise replay onto the next print (#1136). The rebuild ended in paho's
loop_stop(), which sets a terminate flag and then joins the network thread
with no timeout.

That thread only reads the flag between iterations of loop_forever, so it
cannot read it while parked inside reconnect() -> _ssl_wrap_socket() ->
do_handshake(). paho gives that handshake the keepalive as its socket
timeout -- 30s here -- and a socket timeout is per operation, renewed by
every byte the peer sends. A printer that answers TCP and then trickles
holds the join open for as long as it likes.

The join ran on the asyncio thread. Bambuddy stopped answering anything --
UI, API, /health -- while the process stayed up, which is why a
restart: unless-stopped container never restarted.

Retiring a client no longer waits for it. The replacement is built at once
and the old one is shut down on a thread of its own that nobody joins. Its
callbacks are detached first, inline: blocking until the network thread was
gone is what used to guarantee a client we had let go of could no longer
touch our state, and with the teardown detached a zombie that finishes its
handshake would otherwise auto-reconnect and report itself connected behind
its replacement's back. disconnect() still goes out, still promptly, because
that is what stops paho's auto-reconnect and the replay with it.

The reported watchdog is one of six callers. The queue's dispatch recovery
and check_staleness -- reached from an ordinary status poll -- share
_hard_reset_client; editing, deleting and hand-disconnecting a printer share
disconnect(); the relay and smart-plug services had the same join on their
shutdown path, where a wedged broker stopped the process from exiting at
all. #1445 was this join too, from the add-printer probe, and its off-loop
teardown stays as it is.

disconnect() stays quiet on the way out, as it always effectively did.
paho's callback used to land during the join, but it suppresses itself for a
clean disconnect of a printer that reported in the last ten seconds, so a
healthy printer disconnected by hand never announced itself offline.
Announcing it now would tell the user their printer had gone offline a
minute after they disconnected it on purpose (#1752).

A retirement that takes more than five seconds logs which printer it was.
The whole point is that the next one of these should not have to be
diagnosed from a thread dump.
2026-09-18 17:13:26 +02:00
..