8 Commits
Author SHA1 Message Date
maziggy ba6b1a8436 fix(vp): evict MQTT clients on drain timeout + tighten TCP keepalive (#1872)
Reporter (H2C + macOS 26.5.1 + BS 2.8.0.50): after every Mac sleep/wake
cycle, Bambu Studio couldn't see the VP or connect to it. Only fix was
quit BS + reboot Bambuddy. The physical printer's own cloud/LAN link
recovered in ~5 s from the same sleep — the delta was in VP session
handling.

Log evidence (bug-report-assets/logs/ddf1ede75df045cd94ad223d0f08f88a):

- 14:04:06 healthy `1Hz status push: 60 pushes/min to :54698`
- 14:04:06 → 14:09:16: five minutes of SSDP-only, no push summary for
  :54698, no OSError, no disconnect line
- 14:09:16: new source port :54861 connects and authenticates fine —
  the server was not rejecting reconnects
- 14:10:17 first DEBUG line: `MQTT drain timeout for
  device/…/report — client may be busy` — smoking gun

Root cause: `_publish_to_report:1149` caught `asyncio.wait_for(drain,
timeout=5)` TimeoutError at DEBUG and returned silently. TimeoutError
is not OSError, so the push loop's `except OSError` at :441 never saw
it — the zombie writer sat in self._clients until the kernel's default
TCP keepalive detected the dead peer (Linux default: ~2 h 11 min).

Two hunks:

1. `_publish_to_report`: on drain TimeoutError, close the writer (best
   effort, catch Exception so an already-broken close() doesn't mask
   the raise) and raise BrokenPipeError, which IS OSError. Push loop
   evicts on the same tick.

2. `_handle_client`: after SO_KEEPALIVE=1, set TCP_KEEPIDLE=60,
   TCP_KEEPINTVL=15, TCP_KEEPCNT=4 — dead-peer detection in ~2 min
   instead of ~2 h. `getattr(socket, ...)` guards keep it cross-
   platform (macOS uses TCP_KEEPALIVE not TCP_KEEPIDLE, other kernels
   may not expose all three — skip whichever is missing).

What I got wrong first pass and corrected on log-read: hypothesised
"missing MQTT session takeover on same client_id". Wrong. _handle_connect
parses the protocol client_id but discards it (assignment commented out
at :762), and self._clients is keyed on `f"{addr[0]}:{addr[1]}"` (socket
peer), so every reconnect gets a distinct key. No takeover race exists.
The log fixed this: the "not seen" symptom is BS-side (macOS UDP
receive after sleep + BS holding the pre-sleep socket state), but the
server-side amplifier was the zombie writer.
2026-07-02 10:13:01 +02:00
maziggy 7a9c78c32c ● fix(vp): stop enforcing MQTT keepalive 1.5x to match real Bambu firmware (#1548)
Round 1 (b6636053 + 4ffefa60) shipped the keepalive parser, 1.5x idle
  disconnect per MQTT spec section 4.4, and a per-minute status-push
  diagnostic. Reporter's follow-up pcap showed the round-1 logic was
  correct as designed, but the actual root cause sits one layer down:
  the same OrcaSlicer install that stays connected to a real Bambu P1S
  indefinitely sends zero MQTT packets after the initial CONNECT /
  SUBSCRIBE / pushall / get_version burst - no PINGREQ at all - so any
  spec-compliant server disconnects it at keep_alive x 1.5.

  Real Bambu firmware does not enforce section 4.4. The reporter's
  identical Orca install holds idle sessions against real hardware on
  the same network. Spec compliance was itself the regression.

  Fix: after CONNECT/auth, drop the application-level read timeout
  entirely (read_timeout = None) and set SO_KEEPALIVE on the underlying
  socket so the OS TCP stack reaps dead connections within a few
  minutes. The 60s pre-CONNECT cap is preserved - a client that opens
  TCP but never sends CONNECT still gets reaped. Negotiated keepalive
  is still parsed and now logged at INFO ("MQTT client X authenticated
  (negotiated keepalive=Ys, idle disconnect disabled)") for support-
  bundle visibility.

  After this ships, OrcaSlicer should stay connected to the VP
  indefinitely while idle and reconnect cleanly on real network drops.
  The publish_json code -4 and -6010 errors reported in the original
  thread were downstream of this disconnect and should also clear.
2026-06-09 08:22:38 +02:00
maziggy 597762685c fix(virtual-printer): #1558 Send pre-flight + slicer-surface audit bundle
#1558: cached-as-base push_status only forced gcode_state=IDLE while letting
  the real printer's live-progress fields (mc_percent, stg_cur, layer_num, ...)
  leak through. Bambu Studio's Send pre-flight read them as busy and refused.
  The cached branch now overrides the activity-field set the same way it
  already overrode storage indicators (#1228) and protocol fields.

  Same bundle ships a multi-round VP audit that found adjacent bugs in the
  same family:

  - #1558: cached branch zeroes mc_print_stage / mc_percent / mc_remaining_time / stg / stg_cur / layer_num / total_layer_num / print_error
  - MQTT auth: per-IP rate-limit (5/60s lockout), hmac.compare_digest, access_code redacted in DEBUG log
  - FTP cmd_STOR streams chunks to disk + 4 GiB cap (was buffering whole upload)
  - Sticky-keys allowlist extended with upgrade_state / xcam / hw_switch_state / nozzle_diameter / nozzle_type / online / ams_status
  - _pending_files cleanup in finally for archive / queue / dispatch handlers
  - _add_to_print_queue position uses MAX+1 (was hardcoded 1)
  - DELETE VP removes orphan PendingUpload rows + upload_dir from disk
  - Per-VP cert regenerates on shared-CA rotation (real signature verification, not DN match)
  - DHCP target-IP refresh + queue_force_color_match toggle now restart proxy VPs
  - Per-slicer bridge-response routing (multi-slicer cross-leak fix via sequence_id map)
  - Child-service readiness barrier (FTP / MQTT / Bind / SSDP) — no false is_running before sockets bind
  - H2D Pro O1E / O2D model codes added (experimental, needs field confirmation)
  - FTP passive port range widened 50000-51000; docker-compose + wiki updated
  - VP refresh_loop crash now unbinds raw_message_handler; tailscale catches asyncio.TimeoutError; SlicerProxyManager lifecycle hardening
2026-05-30 13:34:10 +02:00
maziggy b663605318 fix(virtual-printer): honour client-negotiated MQTT keepalive instead of hardcoded 60s (#1548)
OrcaSlicer connects, exchanges pushall + get_version, then sits idle waiting
  for status pushes from the (virtual) printer. The VP MQTT server's read
  loop used `asyncio.wait_for(reader.read(1), timeout=60)` regardless of what
  the client negotiated, and `_handle_connect` explicitly skipped the
  keepalive field in the CONNECT payload, so every idle slicer connection was
  torn down at exactly 60s.

  - Parse the 2-byte big-endian keepalive from CONNECT; return it from
    _handle_connect alongside the auth bool.
  - Use 1.5x the negotiated keepalive as the per-packet read timeout per
    MQTT spec sec 4.4. Treat keep_alive == 0 as no timeout (spec sec 3.1.2.10).
  - Retain the 60s default for the initial read before CONNECT arrives, so
    a TCP-connect-without-CONNECT still gets reaped.
  - 7 new tests: 4 unit-level for the parser (success, opt-out=0, auth-fail
    tuple shape, malformed CONNECT) + 3 integration-style for the read loop
    (long keepalive survives the old 60s mark, short keepalive closes idle
    in ~3s, PINGREQ resets the window so DISCONNECT decides the exit).
2026-05-28 09:13:44 +02:00
maziggy 10c261dcf2 chore(tests): suppress B108 on dummy /tmp test fixtures 2026-04-19 09:48:10 +02:00
maziggy 68920f8c62 Fix virtual printer dropping null-terminated MQTT payloads from OrcaSlicer Linux (#927)
OrcaSlicer's Linux BBLNetworkPlugin publishes MQTT payloads with the
  C-string null terminator included in the length, so decoded messages
  arrived as `{…}\x00`. The strict json.loads() raised JSONDecodeError
  and the publish handler silently returned — pushall, get_version, and
  project_file were never answered, and the slicer hit its 60 s sync
  timeout. Print_queue mode only (proxy mode tunnels MQTT). The b069b521
  serial-adaptation fix was correct but ran past this earlier silent
  failure.

  _handle_publish now strips trailing \x00/whitespace before parsing and
  logs the raw payload on any remaining decode failure so future silent
  variants are visible in support bundles.
2026-04-19 08:04:05 +02:00
maziggy b069b5217c y Fix virtual printer "Synchronizing device information" timeout in Orca (#927)
OrcaSlicer's "Send job" flow sat on "Synchronizing device information…"
  until it gave up, even though FTP upload worked when the user clicked
  "Send job anyway". The virtual printer's MQTT server gated all incoming
  command handling on `f"device/{self.serial}/request" in topic` — if the
  slicer's cached serial for the VP didn't exactly equal the VP's computed
  self.serial (model prefix + per-VP serial_suffix), every get_version,
  pushall, and project_file publish was silently dropped. Nothing was
  logged past the initial "MQTT publish to …" line, so the slicer never
  received a push_status or get_version response on its subscribed
  device/{serial}/report topic and hit its sync timeout. Responses were
  also unconditionally published on device/{self.serial}/report, so even
  when the inbound check happened to pass, replies targeted a topic the
  slicer wasn't listening on if its serial had drifted.

  Both directions are now serial-adaptive:

  - `_handle_publish` accepts any authenticated publish on a
    `device/*/request` topic and extracts the serial from the topic itself
    rather than comparing against self.serial.
  - A per-connection `_client_serials` dict tracks the serial the slicer
    actually uses, populated from the first SUBSCRIBE or PUBLISH seen on
    each connection and cleared on disconnect/stop.
  - `_send_status_report`, `_send_version_response`, `_send_print_response`
    now take an optional `serial` parameter (defaulting to self.serial)
    so every outgoing publish — including the periodic 1-second status
    push — targets the topic the slicer subscribed to.
  - The version response's embedded `module[].sn` fields now also carry
    the client's serial so the payload is internally consistent with the
    topic.
  - When the client's serial differs from self.serial an INFO log records
    the adaptation so it's visible in future support bundles.

  The working case (slicer's cached serial equals self.serial, as in my
  own H2D-1 Proxy setup) is bit-for-bit identical to the old behavior —
  the new check is strictly more permissive and only affects cases the
  old code silently dropped.

  Regression tests cover:
  - `_extract_serial_from_topic` valid/invalid topic shapes
  - mismatched-serial publish → handler runs, response topic and sn field
    both use the client's serial
  - non-`/request` topics → still rejected
  - pushall → status_report routed to the client's subscribed topic
  - `_client_serials` cleared on stop()
2026-04-10 12:13:54 +02:00
maziggy 82d329d85c [Fix] Virtual Printer FTP routed to wrong VP with different access codes (#735)
When running multiple virtual printers with different access codes on
  separate bind IPs, FTP connections were always routed to the wrong VP.

  Root cause: the iptables REDIRECT rule (990→9990) rewrites the
  destination IP to the incoming interface's primary address. With Linux's
  weak host model (arp_filter=0), packets for secondary IPs arrive on the
  primary interface, and REDIRECT sends them all to the first VP's FTP
  server. MQTT was unaffected because port 8883 had no redirect.

  Fix: FTP server now binds directly to port 990 (standard implicit FTPS),
  eliminating the iptables redirect entirely. Requires CAP_NET_BIND_SERVICE
  (already set in the systemd service file and Docker image).

  Also removed a global asyncio set_exception_handler() in the MQTT server
  that was overwritten by each VP instance, causing spurious "Unhandled
  exception in client_connected_cb" errors on startup.

  Changes:
  - FTP_PORT: 9990 → 990 (ftp_server.py)
  - Removed set_exception_handler() from MQTT server
  - Updated Dockerfile, docker-compose.yml port mappings
  - Deprecated --redirect-990 in install script
  - Updated wiki: removed iptables instructions for all platforms
  - Added migration guide (docs/migration-vp-ftp-port.md)
  - Added unit tests for port constant and no-global-state invariant
2026-03-18 09:04:31 +01:00