Files
bambuddy/backend/app/services/virtual_printer
maziggy ba6b1a8436 fix(vp): evict MQTT clients on drain timeout + tighten TCP keepalive (#1872)
Reporter (H2C + macOS 26.5.1 + BS 2.8.0.50): after every Mac sleep/wake
cycle, Bambu Studio couldn't see the VP or connect to it. Only fix was
quit BS + reboot Bambuddy. The physical printer's own cloud/LAN link
recovered in ~5 s from the same sleep — the delta was in VP session
handling.

Log evidence (bug-report-assets/logs/ddf1ede75df045cd94ad223d0f08f88a):

- 14:04:06 healthy `1Hz status push: 60 pushes/min to :54698`
- 14:04:06 → 14:09:16: five minutes of SSDP-only, no push summary for
  :54698, no OSError, no disconnect line
- 14:09:16: new source port :54861 connects and authenticates fine —
  the server was not rejecting reconnects
- 14:10:17 first DEBUG line: `MQTT drain timeout for
  device/…/report — client may be busy` — smoking gun

Root cause: `_publish_to_report:1149` caught `asyncio.wait_for(drain,
timeout=5)` TimeoutError at DEBUG and returned silently. TimeoutError
is not OSError, so the push loop's `except OSError` at :441 never saw
it — the zombie writer sat in self._clients until the kernel's default
TCP keepalive detected the dead peer (Linux default: ~2 h 11 min).

Two hunks:

1. `_publish_to_report`: on drain TimeoutError, close the writer (best
   effort, catch Exception so an already-broken close() doesn't mask
   the raise) and raise BrokenPipeError, which IS OSError. Push loop
   evicts on the same tick.

2. `_handle_client`: after SO_KEEPALIVE=1, set TCP_KEEPIDLE=60,
   TCP_KEEPINTVL=15, TCP_KEEPCNT=4 — dead-peer detection in ~2 min
   instead of ~2 h. `getattr(socket, ...)` guards keep it cross-
   platform (macOS uses TCP_KEEPALIVE not TCP_KEEPIDLE, other kernels
   may not expose all three — skip whichever is missing).

What I got wrong first pass and corrected on log-read: hypothesised
"missing MQTT session takeover on same client_id". Wrong. _handle_connect
parses the protocol client_id but discards it (assignment commented out
at :762), and self._clients is keyed on `f"{addr[0]}:{addr[1]}"` (socket
peer), so every reconnect gets a distinct key. No takeover race exists.
The log fixed this: the "not seen" symptom is BS-side (macOS UDP
receive after sleep + BS holding the pre-sleep socket state), but the
server-side amplifier was the zombie writer.
2026-07-02 10:13:01 +02:00
..