2 Commits
Author SHA1 Message Date
maziggy 73afa95047 fix(camera): share one connection between concurrent one-shot captures (#2705)
Bambu firmware allows exactly one camera connection. The existing guards
(is_stream_active / try_get_active_buffered_frame, #1271 and #1348) only
stop a one-shot capturer from competing with the fan-out broadcaster.
Nothing coordinated the capturers with each other, so with no viewer
attached every consumer correctly concluded it was not competing with a
viewer and then collided with the others. On the reporter's P2S an Obico
poll and a snapshot opened two RTSP sockets 207 ms apart, which knocked
over the fan-out stream feeding the camera wall; it was then reaped for
having received no frames for 58s.

capture_camera_frame_bytes() now coalesces: the first caller opens the
connection, callers arriving while it is in flight await the same result.
Eight paths reach that function independently - Obico polling, the
snapshot route, the finish-photo moment and its disk-writing sibling,
plate detection, the camera test and the diagnose tool - so the
single-flight sits at the bottom of the stack and no call site changes.

Keyed by IP, since that is what the firmware's limit applies to and the
function never sees a printer_id. The key excludes the timeout on
purpose: the call sites disagree about it, from 10s to 30s, so keying on
it would mean the Obico-vs-snapshot pair from the report never coalesced
at all.

It coalesces, it does not cache. A call arriving after the previous
capture finished still captures fresh, because plate detection and the
finish-photo path judge a running print from these frames and a stale one
there is worse than a slow one - #1397 was a finish photo taken seconds
late showing the bed already lowered.

Each caller waits on its own deadline rather than inheriting whichever
one happened to open the connection, and shield() means giving up leaves
the capture running for whoever else is still waiting. A follower whose
leader fails takes a turn of its own instead of inheriting a failure it
never had a chance to avoid; the leader has finished by then, so there is
nothing left to compete with. Bounded at two rounds. That also covers the
follower whose timeout is longer than the leader's, which coalescing
alone cannot. Cancellation is disambiguated via leader.cancelled(), so a
follower's own cancellation propagates while a cancelled leader is
treated as a failed one.

The leader is deliberately not wrapped in a second wait_for: the
implementation already enforces the timeout internally, where it can also
kill the ffmpeg process, and an outer deadline would abandon the
subprocess instead of killing it.

The diagnose tool now marks a stage whose frame came from a capture
already in flight as coalesced_capture. The pass is real evidence the
camera works, but duration_ms is then mostly time spent queueing, and a
diagnostic must not report a connection it never opened - the same reason
that file declares its live_stream_active shortcut instead of quietly
passing. Failures are not annotated, since a follower whose leader fails
goes on to capture on its own.
2026-07-30 09:49:56 +02:00
maziggy 134847a3bd feat(camera): in-app diagnostic for "Connection lost" (#1395 follow-up)
Step 2 of the camera architecture overhaul agreed after #1395. When
  the camera viewer hits its error state OR before a print at any
  time, a Diagnose button runs a staged check against the printer and
  renders the result inline: which stage failed, how long it took,
  and a translated remediation hint. Cuts off the "user opens a
  'camera broken' ticket → ask for support bundle → triage" loop at
  the user's screen.

  Backend

  - New `backend/app/services/camera_diagnose.py` orchestrator with
    CameraDiagnoseResult / CameraDiagnoseStage dataclasses.
  - New POST /printers/{id}/camera/diagnose route in camera.py.
  - Stages:
      tcp_reachable — TCP socket open to 322 (RTSP) / 6000 (chamber)
        with 3 s timeout. Distinguishes timeout, refused, and host-
        unreachable into distinct summary codes so the frontend can
        show a precise remediation (firewall vs LAN-only off vs
        wrong IP).
      first_frame — captures one JPEG end-to-end via the existing
        capture_camera_frame_bytes pipeline. Auth + RTSP handshake +
        first keyframe collapse into one stage; the user-facing
        answer is the same regardless of which sub-layer failed.
  - Live-stream shortcut: when a viewer is currently watching the
    camera with a buffered frame < 10 s old, the diagnostic skips
    the real test and returns live_stream_active_healthy. Opening a
    fresh socket would kick the live viewer off on single-camera-
    connection firmwares (the #1348 reconnect-storm trigger), so we
    trust the real-world evidence instead.
  - Response surfaces protocol, port, and profile name for support
    triage — lets us ask "what does your modal say?" instead of
    "send the support bundle".

  Frontend

  - New CameraDiagnoseModal renders one row per stage with green-
    check / red-X / grey-skipped icons, the per-stage duration in
    ms, a remediation banner styled by overall status, and a Run
    again button.
  - Two entry points:
      1. The viewer's error overlay grows a Diagnose button next to
         Retry. Retry stays the primary action; Diagnose is the
         escape hatch for users who can't see what's wrong.
      2. A stethoscope icon in the viewer's always-visible control
         bar, between Refresh and Fullscreen. Pre-flight testing
         ("did my firmware update break the camera?", "is the
         camera up before I send a print?") doesn't require waiting
         for the stream to fail first.
  - Also lifted the previously-hard-coded "Camera unavailable" /
    "Retry" strings into camera.unavailable / camera.retry so the
    error UI is fully translated alongside the new keys.
2026-05-18 10:52:02 +02:00