Files
bambuddy/backend
maziggy 14b322d86a fix(mqtt): never wait for a wedged paho network thread (issue #3068)
The reporter's A1 had been offline 38 hours and still answered on 8883, so
    the connection watchdog did exactly what it exists for: rebuild the session
    with a fresh client, since anything left in the old one's QoS 1 queue would
    otherwise replay onto the next print (#1136). The rebuild ended in paho's
    loop_stop(), which sets a terminate flag and then joins the network thread
    with no timeout.

    That thread only reads the flag between iterations of loop_forever, so it
    cannot read it while parked inside reconnect() -> _ssl_wrap_socket() ->
    do_handshake(). paho gives that handshake the keepalive as its socket
    timeout -- 30s here -- and a socket timeout is per operation, renewed by
    every byte the peer sends. A printer that answers TCP and then trickles
    holds the join open for as long as it likes.

    The join ran on the asyncio thread. Bambuddy stopped answering anything --
    UI, API, /health -- while the process stayed up, which is why a
    restart: unless-stopped container never restarted.

    Retiring a client no longer waits for it. The replacement is built at once
    and the old one is shut down on a thread of its own that nobody joins. Its
    callbacks are detached first, inline: blocking until the network thread was
    gone is what used to guarantee a client we had let go of could no longer
    touch our state, and with the teardown detached a zombie that finishes its
    handshake would otherwise auto-reconnect and report itself connected behind
    its replacement's back. disconnect() still goes out, still promptly, because
    that is what stops paho's auto-reconnect and the replay with it.

    The reported watchdog is one of six callers. The queue's dispatch recovery
    and check_staleness -- reached from an ordinary status poll -- share
    _hard_reset_client; editing, deleting and hand-disconnecting a printer share
    disconnect(); the relay and smart-plug services had the same join on their
    shutdown path, where a wedged broker stopped the process from exiting at
    all. #1445 was this join too, from the add-printer probe, and its off-loop
    teardown stays as it is.

    disconnect() stays quiet on the way out, as it always effectively did.
    paho's callback used to land during the join, but it suppresses itself for a
    clean disconnect of a printer that reported in the last ten seconds, so a
    healthy printer disconnected by hand never announced itself offline.
    Announcing it now would tell the user their printer had gone offline a
    minute after they disconnected it on purpose (#1752).

    A retirement that takes more than five seconds logs which printer it was.
    The whole point is that the next one of these should not have to be
    diagnosed from a thread dump.
2026-09-20 13:34:55 +02:00
..
2025-11-28 10:23:59 +01:00