Reporter (H2C + macOS 26.5.1 + BS 2.8.0.50): after every Mac sleep/wake
cycle, Bambu Studio couldn't see the VP or connect to it. Only fix was
quit BS + reboot Bambuddy. The physical printer's own cloud/LAN link
recovered in ~5 s from the same sleep — the delta was in VP session
handling.
Log evidence (bug-report-assets/logs/ddf1ede75df045cd94ad223d0f08f88a):
- 14:04:06 healthy `1Hz status push: 60 pushes/min to :54698`
- 14:04:06 → 14:09:16: five minutes of SSDP-only, no push summary for
:54698, no OSError, no disconnect line
- 14:09:16: new source port :54861 connects and authenticates fine —
the server was not rejecting reconnects
- 14:10:17 first DEBUG line: `MQTT drain timeout for
device/…/report — client may be busy` — smoking gun
Root cause: `_publish_to_report:1149` caught `asyncio.wait_for(drain,
timeout=5)` TimeoutError at DEBUG and returned silently. TimeoutError
is not OSError, so the push loop's `except OSError` at :441 never saw
it — the zombie writer sat in self._clients until the kernel's default
TCP keepalive detected the dead peer (Linux default: ~2 h 11 min).
Two hunks:
1. `_publish_to_report`: on drain TimeoutError, close the writer (best
effort, catch Exception so an already-broken close() doesn't mask
the raise) and raise BrokenPipeError, which IS OSError. Push loop
evicts on the same tick.
2. `_handle_client`: after SO_KEEPALIVE=1, set TCP_KEEPIDLE=60,
TCP_KEEPINTVL=15, TCP_KEEPCNT=4 — dead-peer detection in ~2 min
instead of ~2 h. `getattr(socket, ...)` guards keep it cross-
platform (macOS uses TCP_KEEPALIVE not TCP_KEEPIDLE, other kernels
may not expose all three — skip whichever is missing).
What I got wrong first pass and corrected on log-read: hypothesised
"missing MQTT session takeover on same client_id". Wrong. _handle_connect
parses the protocol client_id but discards it (assignment commented out
at :762), and self._clients is keyed on `f"{addr[0]}:{addr[1]}"` (socket
peer), so every reconnect gets a distinct key. No takeover race exists.
The log fixed this: the "not seen" symptom is BS-side (macOS UDP
receive after sleep + BS holding the pre-sleep socket state), but the
server-side amplifier was the zombie writer.
Round 1 (b6636053 + 4ffefa60) shipped the keepalive parser, 1.5x idle
disconnect per MQTT spec section 4.4, and a per-minute status-push
diagnostic. Reporter's follow-up pcap showed the round-1 logic was
correct as designed, but the actual root cause sits one layer down:
the same OrcaSlicer install that stays connected to a real Bambu P1S
indefinitely sends zero MQTT packets after the initial CONNECT /
SUBSCRIBE / pushall / get_version burst - no PINGREQ at all - so any
spec-compliant server disconnects it at keep_alive x 1.5.
Real Bambu firmware does not enforce section 4.4. The reporter's
identical Orca install holds idle sessions against real hardware on
the same network. Spec compliance was itself the regression.
Fix: after CONNECT/auth, drop the application-level read timeout
entirely (read_timeout = None) and set SO_KEEPALIVE on the underlying
socket so the OS TCP stack reaps dead connections within a few
minutes. The 60s pre-CONNECT cap is preserved - a client that opens
TCP but never sends CONNECT still gets reaped. Negotiated keepalive
is still parsed and now logged at INFO ("MQTT client X authenticated
(negotiated keepalive=Ys, idle disconnect disabled)") for support-
bundle visibility.
After this ships, OrcaSlicer should stay connected to the VP
indefinitely while idle and reconnect cleanly on real network drops.
The publish_json code -4 and -6010 errors reported in the original
thread were downstream of this disconnect and should also clear.
OrcaSlicer connects, exchanges pushall + get_version, then sits idle waiting
for status pushes from the (virtual) printer. The VP MQTT server's read
loop used `asyncio.wait_for(reader.read(1), timeout=60)` regardless of what
the client negotiated, and `_handle_connect` explicitly skipped the
keepalive field in the CONNECT payload, so every idle slicer connection was
torn down at exactly 60s.
- Parse the 2-byte big-endian keepalive from CONNECT; return it from
_handle_connect alongside the auth bool.
- Use 1.5x the negotiated keepalive as the per-packet read timeout per
MQTT spec sec 4.4. Treat keep_alive == 0 as no timeout (spec sec 3.1.2.10).
- Retain the 60s default for the initial read before CONNECT arrives, so
a TCP-connect-without-CONNECT still gets reaped.
- 7 new tests: 4 unit-level for the parser (success, opt-out=0, auth-fail
tuple shape, malformed CONNECT) + 3 integration-style for the read loop
(long keepalive survives the old 60s mark, short keepalive closes idle
in ~3s, PINGREQ resets the window so DISCONNECT decides the exit).
OrcaSlicer's Linux BBLNetworkPlugin publishes MQTT payloads with the
C-string null terminator included in the length, so decoded messages
arrived as `{…}\x00`. The strict json.loads() raised JSONDecodeError
and the publish handler silently returned — pushall, get_version, and
project_file were never answered, and the slicer hit its 60 s sync
timeout. Print_queue mode only (proxy mode tunnels MQTT). The b069b521
serial-adaptation fix was correct but ran past this earlier silent
failure.
_handle_publish now strips trailing \x00/whitespace before parsing and
logs the raw payload on any remaining decode failure so future silent
variants are visible in support bundles.
OrcaSlicer's "Send job" flow sat on "Synchronizing device information…"
until it gave up, even though FTP upload worked when the user clicked
"Send job anyway". The virtual printer's MQTT server gated all incoming
command handling on `f"device/{self.serial}/request" in topic` — if the
slicer's cached serial for the VP didn't exactly equal the VP's computed
self.serial (model prefix + per-VP serial_suffix), every get_version,
pushall, and project_file publish was silently dropped. Nothing was
logged past the initial "MQTT publish to …" line, so the slicer never
received a push_status or get_version response on its subscribed
device/{serial}/report topic and hit its sync timeout. Responses were
also unconditionally published on device/{self.serial}/report, so even
when the inbound check happened to pass, replies targeted a topic the
slicer wasn't listening on if its serial had drifted.
Both directions are now serial-adaptive:
- `_handle_publish` accepts any authenticated publish on a
`device/*/request` topic and extracts the serial from the topic itself
rather than comparing against self.serial.
- A per-connection `_client_serials` dict tracks the serial the slicer
actually uses, populated from the first SUBSCRIBE or PUBLISH seen on
each connection and cleared on disconnect/stop.
- `_send_status_report`, `_send_version_response`, `_send_print_response`
now take an optional `serial` parameter (defaulting to self.serial)
so every outgoing publish — including the periodic 1-second status
push — targets the topic the slicer subscribed to.
- The version response's embedded `module[].sn` fields now also carry
the client's serial so the payload is internally consistent with the
topic.
- When the client's serial differs from self.serial an INFO log records
the adaptation so it's visible in future support bundles.
The working case (slicer's cached serial equals self.serial, as in my
own H2D-1 Proxy setup) is bit-for-bit identical to the old behavior —
the new check is strictly more permissive and only affects cases the
old code silently dropped.
Regression tests cover:
- `_extract_serial_from_topic` valid/invalid topic shapes
- mismatched-serial publish → handler runs, response topic and sn field
both use the client's serial
- non-`/request` topics → still rejected
- pushall → status_report routed to the client's subscribed topic
- `_client_serials` cleared on stop()
When running multiple virtual printers with different access codes on
separate bind IPs, FTP connections were always routed to the wrong VP.
Root cause: the iptables REDIRECT rule (990→9990) rewrites the
destination IP to the incoming interface's primary address. With Linux's
weak host model (arp_filter=0), packets for secondary IPs arrive on the
primary interface, and REDIRECT sends them all to the first VP's FTP
server. MQTT was unaffected because port 8883 had no redirect.
Fix: FTP server now binds directly to port 990 (standard implicit FTPS),
eliminating the iptables redirect entirely. Requires CAP_NET_BIND_SERVICE
(already set in the systemd service file and Docker image).
Also removed a global asyncio set_exception_handler() in the MQTT server
that was overwritten by each VP instance, causing spurious "Unhandled
exception in client_connected_cb" errors on startup.
Changes:
- FTP_PORT: 9990 → 990 (ftp_server.py)
- Removed set_exception_handler() from MQTT server
- Updated Dockerfile, docker-compose.yml port mappings
- Deprecated --redirect-990 in install script
- Updated wiki: removed iptables instructions for all platforms
- Added migration guide (docs/migration-vp-ftp-port.md)
- Added unit tests for port constant and no-global-state invariant