Commit Graph
1638 Commits
Author SHA1 Message Date
maziggy 7096787be6 fix(printers): don't retract a fan kit on a partial airduct frame
device.airduct is pushed field by field - the modeCur handler reads it with
    an "in" check for that reason - so a frame can carry parts without carrying
    every fan. Absence in that list is what tells us a kit is not fitted, and
    taken from a truncated frame it made both accessory badges vanish mid-print
    and started rejecting fan=aux2 on a printer that has the fan.

    A parts list now counts as a full inventory only when it carries ids 1 (part
    cooling) and 2 (aux). Neither is optional on a machine that reports an
    airduct at all, and both appear in every layout in the support-package
    archive - P2S base 1,2 / P2S+kit 1,2,3 / X2D 1,2,3,10 / H2C,H2D,H2S 1,2,3,6.
    Anything narrower is a diff frame: its speeds are applied, presence is left
    alone. Presence can still be added from a partial frame; only retraction
    needs the full list, so a kit that really is removed still disappears.

    Also compose showChamberFan from both model lists rather than branching
    between them, so the P2S/X2D entries in MODELS_WITH_CHAMBER_FAN stay
    reachable instead of reading as dead, and note in the fan-speed docstring
    that the aux2 gate also rejects between connect and the first airduct push.
2026-08-02 09:42:26 +02:00
maziggy 3466195d60 fix(slice): give the slice modal one filament row per project slot (#2712)
The filament list is positional from the modal down to the CLI's
    filament_N.json parts, but for a source that already carries slice_info the
    requirements endpoint returns only the slots the plate consumes. A
    MakerWorld model declaring four filaments and painting with slot 4 alone
    therefore showed one dropdown, whose PETG pick the CLI bound to slot 1 —
    slot 4 sliced with the profile baked into the source, and the print came out
    PLA.

    The endpoint now takes full_slots, which widens that answer to every
    project slot with used_in_plate flags, and only the slice modal passes it.
    Print-time AMS matching shares the endpoint and keeps the used-only list, so
    it still asks for exactly the spools the job needs.
2026-08-02 09:42:07 +02:00
maziggy 89804b810d fix(slice): report a finished slice once, not once per queued poll
setInterval does not await an async callback. Slicing a large project
    blocks the backend for seconds, so poll ticks piled up behind one stalled
    request, each holding a snapshot taken while the job was still active.
    They resolved together, and every one of them ran the completion path —
    one toast and two query invalidations each. A 20s stall against the 1.5s
    interval produced 13 "Sliced X" toasts from a single slice.

    Only one poll round is now in flight at a time, which also stops queueing
    requests against a backend that is already saturated. Completion is
    recorded once per job id, and a round still awaiting a response when the
    effect tears down now returns instead of acting.
2026-08-02 09:41:46 +02:00
maziggy 4045ddbd1f fix(queue): withdraw an expected print when the command never goes out
feat(db): warn when the connection pool can outgrow the PostgreSQL server

    fix(mqtt): an unusable layer_num must not drop the printer connection

    test: patch settings.base_dir via monkeypatch so it unwinds on error

    test: restore the config module after reloading it
2026-08-02 09:41:32 +02:00
maziggy bdcb8a3fc7 fix(mqtt): keep the layer total that arrives with the print-start frame (#2702)
fix(support): redact push_status values, not the serialised JSON (#2702)
2026-08-02 09:41:19 +02:00
maziggy 216a8a3695 Security hardening (maziggy/bambuddy-security #7)
fix(settings): accept JSON booleans on the Spoolman settings endpoint
2026-08-02 09:41:04 +02:00
maziggy 8eb185b239 fix(camera): reuse the live view's frame for external-camera captures (#2707)
On a printer with an external camera, watching the live view while a print
    ran meant the layer timelapse recorded almost nothing and the finish photo
    went out with no image. The reporter measured 0 of 87 layer captures on one
    print and 0 of 105 on another, both watched throughout. A USB camera allows
    one V4L2 handle, so a capture during a live view fails outright.

    The built-in camera has had this rule since #1348 and #1271: reuse the
    viewer's buffered frame rather than opening a second connection. It was
    never extended to the external paths, and it could not have been -- the
    buffer it depends on was only ever populated by the built-in paths.
    generate_mjpeg_stream yields multipart-wrapped chunks, so the route layer
    could not recover the JPEG, and a guarded caller would have found an empty
    buffer and skipped every time.

    So the stream now publishes each raw frame through a new on_frame callback
    (parallel to on_process from #2675), and the six one-shot consumers reuse
    it: layer timelapse, the finish-photo moment and its background fallback,
    the notification snapshot, Obico polling, and the plate check. A viewer
    attached with nothing buffered yet skips that one attempt rather than
    competing -- kicking the viewer off is worse than missing a frame.

    on_frame exceptions are logged and swallowed, like iter_subscriber's
    on_unsubscribe: buffering is a side effect and must never be able to take
    the live stream down with it. The external stream's teardown now releases
    the buffered frame too, ownership-checked so a concurrent viewer of the
    same printer keeps its own.

    Two side effects on paths not touched here, both improvements: the snapshot
    endpoint and the finish-photo fallback chain consult get_buffered_frame and
    can now serve an external camera's live frame. plate_detection's docstring
    already claimed this behaviour while implementing it only for the built-in
    fallback; that drift is resolved.
2026-08-02 09:40:50 +02:00
maziggy 55f4aa61eb fix(camera): drain a streaming ffmpeg's stderr continuously (#2707)
ffmpeg is spawned with stderr=PIPE and it was only read on the error
    paths, so for the life of a working stream nobody read that pipe. ffmpeg
    writes its banner, the input analysis, then a progress line at a steady
    rate; a 64 KiB pipe fills eventually, ffmpeg blocks writing to it, frames
    stop, and the stream's own 30s timeout fires -- logged as "RTSP read
    timeout" with no hint that we starved it ourselves.

    How long that takes is unmeasured and evidently long: one H2D upstream
    ran 21m36s without stalling, and an earlier 512 B/s extrapolation of mine
    was mostly the one-off startup banner. So this is a bounded resource being
    treated as unbounded, not a fault anyone has reported.

    _FfmpegStderrTail drains the pipe continuously and keeps a 16 KiB rolling
    tail. That tail is what the error paths now report, which is better
    material than before: it holds what ffmpeg said as things went wrong,
    where the on-demand read returned whatever was printed first -- usually
    the banner, which the summariser strips anyway.

    Three readers wanted this one pipe, and asyncio rejects concurrent reads
    on a StreamReader, so the collector is authoritative: it registers by pid,
    _read_ffmpeg_stderr returns its tail when present and otherwise reads the
    pipe unchanged, and _terminate_ffmpeg skips its own stderr drain when the
    collector owns it (the collector keeps draining through teardown, which is
    all wait() needs). The generator starts it after the immediate-failure
    check, which reads the pipe directly because the process is already dead,
    and releases it after _terminate_ffmpeg.

    text() goes through _summarize_ffmpeg_stderr like every other stderr log
    here, so the access code ffmpeg echoes in its input URL stays masked.
    aclose() awaits the cancelled pump rather than firing and forgetting, so
    no pending task survives into loop teardown.
2026-08-02 09:40:34 +02:00
maziggy 097e10b344 fix(camera): one registry key per stream, not per printer (issue #2707)
Closing a camera view and reopening it immediately could leave the new
    stream unregistered while it was running and delivering frames. The
    damage was all indirect: is_stream_active() reported no viewer, so Obico
    polling and snapshots opened a second camera connection against the live
    view (the thing #1348 and #1271 exist to prevent); the janitor's /proc
    scan found an ffmpeg missing from _active_streams and killed the live
    stream as an orphan; and /camera/stop reported "Stopped 0" with a stream
    running.

    The fan-out stream id was f"{printer_id}-fanout" -- constant per printer,
    so every successive stream shared one registry key, and the departing
    generator's finally popped whatever was under it, including its
    successor's entry. The same finally also cleared the per-printer frame
    buffer unconditionally, discarding the new stream's frame. It needed the
    two streams to overlap, which the 4s teardown made easy.

    Each stream now gets its own key via _new_fanout_stream_id(), so a
    generator can only clean up after itself -- the external-camera path
    already does this (#2675) and this brings the fan-out path in line. The
    per-printer dicts are released through _release_printer_frame_state(),
    which checks that no other stream for the printer is still running; both
    the RTSP and chamber-image cleanups had the same unconditional pop.

    Also hoisted time and uuid to module level and dropped four
    function-local `import time` statements. A local import shadows the name
    for the whole function, so any use on a branch that doesn't reach the
    import raises UnboundLocalError -- a real hazard in camera_stream, whose
    external-camera branch imported both while the RTSP path needs them too.
    A test pins camera_stream as free of function-local imports.
2026-08-02 09:40:19 +02:00
maziggy 562e325972 fix(camera): drain ffmpeg's pipes during teardown (#NNNN)
Closing a camera view logged "ffmpeg didn't terminate gracefully,
    killing" followed by "ffmpeg did not exit within 2.0s of SIGKILL;
    abandoning wait", on every single close. Both waits expired every time,
    so teardown took a fixed 4.00s -- and since the firmware allows one
    camera connection, that was 4s in which nothing else could use it.

    ffmpeg is spawned with stdout and stderr as pipes and the teardown paths
    have stopped reading them, so it sits blocked in write() on a full 64 KiB
    pipe. SIGTERM cannot be acted on there: the handler only sets a flag that
    the main loop polls, and the loop never gets back to the check. SIGKILL
    does kill it, but asyncio resolves Process.wait()'s waiter through
    _try_finish(), which requires every pipe transport to report
    disconnected; paused, unread pipes never reach EOF, so wait() blocks with
    returncode already set. A negative-control test shows returncode=-9 at
    the instant the abandon fires.

    Draining both pipes while stopping the process fixes both halves: 4.00s
    becomes ~0.15s. The signal ladder and its bounds stay as backstops, so a
    genuinely wedged process still cannot hang a stream, a Stop request or
    the janitor.

    This corrects _FFMPEG_KILL_TIMEOUT's premise and #2580's conclusion. That
    12-hour hang was the unbounded form of this same self-inflicted stall, not
    an ffmpeg stuck in uninterruptible I/O -- the process observed doing it
    was in state S, which cannot survive a delivered SIGKILL. Bounding the
    wait capped the symptom without removing the cause.
2026-08-02 09:40:04 +02:00
maziggy 54908115a2 fix(camera): share one connection between concurrent one-shot captures (#2705)
Bambu firmware allows exactly one camera connection. The existing guards
    (is_stream_active / try_get_active_buffered_frame, #1271 and #1348) only
    stop a one-shot capturer from competing with the fan-out broadcaster.
    Nothing coordinated the capturers with each other, so with no viewer
    attached every consumer correctly concluded it was not competing with a
    viewer and then collided with the others. On the reporter's P2S an Obico
    poll and a snapshot opened two RTSP sockets 207 ms apart, which knocked
    over the fan-out stream feeding the camera wall; it was then reaped for
    having received no frames for 58s.

    capture_camera_frame_bytes() now coalesces: the first caller opens the
    connection, callers arriving while it is in flight await the same result.
    Eight paths reach that function independently - Obico polling, the
    snapshot route, the finish-photo moment and its disk-writing sibling,
    plate detection, the camera test and the diagnose tool - so the
    single-flight sits at the bottom of the stack and no call site changes.

    Keyed by IP, since that is what the firmware's limit applies to and the
    function never sees a printer_id. The key excludes the timeout on
    purpose: the call sites disagree about it, from 10s to 30s, so keying on
    it would mean the Obico-vs-snapshot pair from the report never coalesced
    at all.

    It coalesces, it does not cache. A call arriving after the previous
    capture finished still captures fresh, because plate detection and the
    finish-photo path judge a running print from these frames and a stale one
    there is worse than a slow one - #1397 was a finish photo taken seconds
    late showing the bed already lowered.

    Each caller waits on its own deadline rather than inheriting whichever
    one happened to open the connection, and shield() means giving up leaves
    the capture running for whoever else is still waiting. A follower whose
    leader fails takes a turn of its own instead of inheriting a failure it
    never had a chance to avoid; the leader has finished by then, so there is
    nothing left to compete with. Bounded at two rounds. That also covers the
    follower whose timeout is longer than the leader's, which coalescing
    alone cannot. Cancellation is disambiguated via leader.cancelled(), so a
    follower's own cancellation propagates while a cancelled leader is
    treated as a failed one.

    The leader is deliberately not wrapped in a second wait_for: the
    implementation already enforces the timeout internally, where it can also
    kill the ffmpeg process, and an outer deadline would abandon the
    subprocess instead of killing it.

    The diagnose tool now marks a stage whose frame came from a capture
    already in flight as coalesced_capture. The pass is real evidence the
    camera works, but duration_ms is then mostly time spent queueing, and a
    diagnostic must not report a connection it never opened - the same reason
    that file declares its live_stream_active shortcut instead of quietly
    passing. Failures are not annotated, since a follower whose leader fails
    goes on to capture on its own.
2026-08-02 09:39:40 +02:00
maziggy c655413971 fix(tests): add the new fan fields to the plate-clear status fixture
mqtt_relay reads state.left_aux_fan_speed, but the SimpleNamespace fixture in
    test_plate_clear_mqtt_notification enumerates its fields explicitly, so the two
    status-payload tests raised AttributeError. I updated the equivalent fixture in
    test_printer_manager_status_broadcast and missed this one — running only the
    touched suites is what hid it.

    Also addresses the round-2 review notes:

    - exhaust_fan_present: documented that the H2 series reports part 3 too, so the
      flag is not model-specific despite the name.
    - Mask the part id after shifting, matching get_flag_bits(id, 4, 8), for
      consistency with the state decode. No behaviour change for any observed id.
    - Noted the unmapped H2 id 6 beside the id branches.
    - Reject fan=aux2 when the printer reports no left_aux_fan_speed, so a POST
      against an A1 no longer sends M106 P10 for absent hardware. The UI already
      hid the badge; this closes the same hole on the API.
    - Added a test asserting EXHAUST_FAN_LABEL_MODELS and the frontend's
      MODELS_WITH_EXHAUST_LABEL cannot drift apart.
2026-08-02 09:38:51 +02:00
maziggy 3e326adce8 feat(notifications): optional Telegram forum topic via message_thread_id (#1518)
Telegram groups with Topics enabled always received notifications in the
    General topic, since only Bot Token and Chat ID were configurable. Splitting
    notifications per printer meant running a separate chat for each one.

    The Telegram provider now takes an optional Forum Topic ID - the last number
    in a topic's link, t.me/c/1234567890/25 - and routes its messages there. Left
    empty, nothing changes.

    The value is coerced to an int once in _send_telegram and attached to both the
    sendMessage JSON body and the sendPhoto form data. That ordering matters:
    Telegram rejects a string message_thread_id in the JSON body while accepting
    one in the multipart call, so passing the raw form value through would have
    worked for thumbnail notifications and 400'd for plain-text ones. A
    non-numeric value is rejected in the form and again server-side before any
    request goes out.

    No migration - provider config is a JSON blob.

    Adds Forum Topic ID plus help text to the Telegram section of the provider
    dialog, translated in all 13 locales. Backend tests cover omitted / blank /
    int-typed / non-numeric values and both send paths; frontend tests cover the
    field being optional, absent for other providers, round-tripping on save, and
    blocking save on a bad value.
2026-08-02 09:37:46 +02:00
maziggy 65b4c59175 Merge branch 'dev' into feature/p2s-x2d-accessory-fans 2026-08-02 09:36:55 +02:00
maziggy 6d8434b250 Security hardening (maziggy/bambuddy-security #N)
Subprocess output and user-supplied URLs are scrubbed of credentials
    before they reach the application log. Adds a shared redaction helper in
    core/logging_filters and routes the existing support-bundle sanitizer
    through the same pattern.
2026-08-02 09:36:38 +02:00
maziggy fe476dd98e fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse
    found nothing afterwards. Across 247 support bundles this was the norm, not an
    edge case: 457 automatic scans scheduled, 262 attached.

    The scan looked four times over ~65s. The attempt that found the video was #1
    272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
    files were still arriving when we stopped. What ran afterwards searched for the
    print name inside the filename; Bambu only writes "video_<timestamp>", so it
    fired 159 times and matched zero.

    The manual Scan had no baseline at all and matched on filename timestamp, FTP
    mtime, or "there is only one video" — all reading a clock a LAN-only printer
    cannot sync. The reporter's P1S was six and a half days out.

    - Poll for minutes instead of ~65s; drop the name-match fallback.
    - Persist the print-start baseline on the archive, so the diff survives a
      restart mid-print and the manual Scan runs the same comparison. With a
      baseline present the clock-based strategies are skipped entirely — they can
      only turn an honest "pick one" into a confident wrong answer.
    - When several files are new (a previous print's video landing late), exclude
      the ones already attached to another archive instead of ordering the
      candidates. Ordering could only be done on the printer's clock.
    - Delete the video from the printer once archived. Keeps /timelapse to
      unclaimed files, which is what makes the diff unambiguous, and stops P1S
      cards filling with AVIs.
    - Gate that delete on a verified transfer: download_file now compares against
      the size from the listing. An FTPS connection closing early does not always
      raise, so a partial buffer was being attached as a complete video — which
      would also have been the one case where deleting the source lost data.

    Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
    the deadline stops bounding the loop as soon as the sleeps are shortened.
    Per-round logging only speaks when the listing changed — 31 rounds of full
    listings would bury the interesting line in the support bundle.

    Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
    both dialects so a migrated database matches a fresh one.

    -----------

    fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)

    When a print records a timelapse, its last frame is the better finish photo:
    the firmware stops recording with the toolhead parked and before the end
    G-code drops the bed, where a live grab at that moment catches a lowered
    plate. Bambuddy waited 60s for the video and then gave up, because the
    print-complete notification blocks on that photo and holding a notification
    for minutes is worse than sending it with the live grab.

    P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
    Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
    worst 546s, while every other model finished inside 26s. So the printers that
    most needed the better framing were the ones that never got it.

    Keep the notification on the same bound, and keep waiting off to the side.
    _capture_finish_photo_from_timelapse now reports whether it ran out of time or
    concluded — a video that landed and failed extraction is not worth retrying,
    one that never arrived is. On the first, schedule a background task that waits
    up to 15 minutes and inserts the extracted frame at the front of the archive's
    photo list, where the gallery opens.

    The live grab stays on disk: the notification already links to that exact
    file, so removing it would leave a broken image in Discord or Telegram.

    The length check proves we received what the listing said, not that the file
    was finished. The first look happens ~5s after the print ends, while the
    printer may still be writing, so a growing file can be listed short, served
    short, and pass. Re-list after the download and only accept the video once its
    size has stopped changing — a failed re-list counts as not settled, since
    "could not check" must not mean "safe to delete".
2026-08-02 09:36:07 +02:00
maziggy d14c969197 feat(diagnostics): log end-of-print telemetry for finish-photo trigger research (#2547)
The finish photo needs a "printing done, toolhead parked, filament unload
    not started" moment. stg_cur=22 was meant to be it (#1721) and fires on no
    model in the field: across 247 support bundles there is not one
    FINISH PHOTO MOMENT (stage-22), including the 2026-06-13..07-08 window where
    it was the only pre-FINISH trigger in the code — 104 captures on A1, A1 Mini,
    H2C, H2D, P1S, P2S, X1C and X2D, all of them the FINISH fallback.

    A replacement can't be designed from the bundles we have. Out of that window
    Bambuddy parses only stg_cur and mc_print_sub_stage; every other stage/action
    field arrives and is dropped unread. The candidates that sound right
    (print_real_action, mc_action, mc_stage) are absent from A1/A1 Mini/P1S
    payloads, so none of them can be the universal answer alone.

    Dump the raw fields for the window between the last object layer and
    gcode_state=FINISH at DEBUG. Opens on the first end-of-print signal (last
    layer, progress >= 99, or no remaining time) so a dropped layer_num packet
    doesn't lose it, logs only what changed frame to frame, closes on the
    transition out of RUNNING, and arms once per print.

    Instrumentation only: gated on DEBUG being enabled, read-only against printer
    state, wrapped so it cannot break ingest, and capped at 400 frames per print.
    The probed fields are stage codes, counters and bitfields — nothing
    identifying, and no access code.
2026-08-02 09:35:49 +02:00
maziggy 1cda64c35a fix(mqtt): report why a printer refused the connection instead of looping silently
A printer with a wrong access code gave no explanation anywhere. The connect
    callback's failure branch was a bare `state.connected = False`, discarding the
    CONNACK reason code the printer had just sent, so the only trace was paho's
    follow-up disconnect -- logged every 30 seconds as "rc=Unspecified error",
    which is exactly what a powered-off printer produces. In the report behind this
    fix one of three printers had been in that loop for the whole capture, and
    neither the log nor the support bundle could say why.

    Bambu speaks MQTT 3.1.1, whose CONNACK return codes 4 and 5 paho maps onto
    reason codes 134 and 135. Both are now logged with the printer's own reason
    string and, for those two, the remedy: the access code is regenerated whenever
    LAN Only or Developer Mode is toggled, so it has to be re-read from the screen.
    The access code itself is never logged -- it would land in every bundle.

    The reason is kept on the client as a stable slug and plumbed through
    test_connection into the connection diagnostic, which now distinguishes two
    cases it previously conflated. "The printer refused our credentials" is
    asserted only when the printer said so; when all Bambuddy knows is that there
    is no session, the text hedges and names the alternatives (rebooting, or
    already at its limit of simultaneous connections). The old wording claimed the
    access code was most likely wrong in both cases.

    Frontend needed no change -- ConnectionDiagnostic already renders
    `<status>_<reason>` variants with fallback to the plain per-status text, so an
    unrecognised slug degrades to today's wording rather than a missing key.
2026-08-02 09:35:23 +02:00
maziggy ac3e3cc60f fix(printers): don't retract a fan kit on a partial airduct frame
device.airduct is pushed field by field - the modeCur handler reads it with
an "in" check for that reason - so a frame can carry parts without carrying
every fan. Absence in that list is what tells us a kit is not fitted, and
taken from a truncated frame it made both accessory badges vanish mid-print
and started rejecting fan=aux2 on a printer that has the fan.

A parts list now counts as a full inventory only when it carries ids 1 (part
cooling) and 2 (aux). Neither is optional on a machine that reports an
airduct at all, and both appear in every layout in the support-package
archive - P2S base 1,2 / P2S+kit 1,2,3 / X2D 1,2,3,10 / H2C,H2D,H2S 1,2,3,6.
Anything narrower is a diff frame: its speeds are applied, presence is left
alone. Presence can still be added from a partial frame; only retraction
needs the full list, so a kit that really is removed still disappears.

Also compose showChamberFan from both model lists rather than branching
between them, so the P2S/X2D entries in MODELS_WITH_CHAMBER_FAN stay
reachable instead of reading as dead, and note in the fan-speed docstring
that the aux2 gate also rejects between connect and the first airduct push.
2026-07-30 18:04:38 +02:00
MartinNYHC 9ebfcddbdb Merge branch 'dev' into feature/p2s-x2d-accessory-fans 2026-07-30 17:37:48 +02:00
maziggy db538e43f1 fix(slice): give the slice modal one filament row per project slot (#2712)
The filament list is positional from the modal down to the CLI's
filament_N.json parts, but for a source that already carries slice_info the
requirements endpoint returns only the slots the plate consumes. A
MakerWorld model declaring four filaments and painting with slot 4 alone
therefore showed one dropdown, whose PETG pick the CLI bound to slot 1 —
slot 4 sliced with the profile baked into the source, and the print came out
PLA.

The endpoint now takes full_slots, which widens that answer to every
project slot with used_in_plate flags, and only the slice modal passes it.
Print-time AMS matching shares the endpoint and keeps the used-only list, so
it still asks for exactly the spools the job needs.
2026-07-30 17:34:26 +02:00
maziggy 83142c726c fix(slice): report a finished slice once, not once per queued poll
setInterval does not await an async callback. Slicing a large project
blocks the backend for seconds, so poll ticks piled up behind one stalled
request, each holding a snapshot taken while the job was still active.
They resolved together, and every one of them ran the completion path —
one toast and two query invalidations each. A 20s stall against the 1.5s
interval produced 13 "Sliced X" toasts from a single slice.

Only one poll round is now in flight at a time, which also stops queueing
requests against a backend that is already saturated. Completion is
recorded once per job id, and a round still awaiting a response when the
effect tears down now returns instead of acting.
2026-07-30 17:14:07 +02:00
maziggy ad785a95cb fix(queue): withdraw an expected print when the command never goes out
feat(db): warn when the connection pool can outgrow the PostgreSQL server

fix(mqtt): an unusable layer_num must not drop the printer connection

test: patch settings.base_dir via monkeypatch so it unwinds on error

test: restore the config module after reloading it
2026-07-30 16:31:09 +02:00
maziggy beca3a8d73 fix(mqtt): keep the layer total that arrives with the print-start frame (#2702)
fix(support): redact push_status values, not the serialised JSON (#2702)
2026-07-30 15:11:08 +02:00
maziggy 88dc56d6e1 Security hardening (maziggy/bambuddy-security #7)
fix(settings): accept JSON booleans on the Spoolman settings endpoint
2026-07-30 13:52:56 +02:00
maziggy d9da60dd8d fix(camera): reuse the live view's frame for external-camera captures (#2707)
On a printer with an external camera, watching the live view while a print
ran meant the layer timelapse recorded almost nothing and the finish photo
went out with no image. The reporter measured 0 of 87 layer captures on one
print and 0 of 105 on another, both watched throughout. A USB camera allows
one V4L2 handle, so a capture during a live view fails outright.

The built-in camera has had this rule since #1348 and #1271: reuse the
viewer's buffered frame rather than opening a second connection. It was
never extended to the external paths, and it could not have been -- the
buffer it depends on was only ever populated by the built-in paths.
generate_mjpeg_stream yields multipart-wrapped chunks, so the route layer
could not recover the JPEG, and a guarded caller would have found an empty
buffer and skipped every time.

So the stream now publishes each raw frame through a new on_frame callback
(parallel to on_process from #2675), and the six one-shot consumers reuse
it: layer timelapse, the finish-photo moment and its background fallback,
the notification snapshot, Obico polling, and the plate check. A viewer
attached with nothing buffered yet skips that one attempt rather than
competing -- kicking the viewer off is worse than missing a frame.

on_frame exceptions are logged and swallowed, like iter_subscriber's
on_unsubscribe: buffering is a side effect and must never be able to take
the live stream down with it. The external stream's teardown now releases
the buffered frame too, ownership-checked so a concurrent viewer of the
same printer keeps its own.

Two side effects on paths not touched here, both improvements: the snapshot
endpoint and the finish-photo fallback chain consult get_buffered_frame and
can now serve an external camera's live frame. plate_detection's docstring
already claimed this behaviour while implementing it only for the built-in
fallback; that drift is resolved.
2026-07-30 13:05:15 +02:00
maziggy 40e7b60e8c fix(camera): drain a streaming ffmpeg's stderr continuously (#2707)
ffmpeg is spawned with stderr=PIPE and it was only read on the error
paths, so for the life of a working stream nobody read that pipe. ffmpeg
writes its banner, the input analysis, then a progress line at a steady
rate; a 64 KiB pipe fills eventually, ffmpeg blocks writing to it, frames
stop, and the stream's own 30s timeout fires -- logged as "RTSP read
timeout" with no hint that we starved it ourselves.

How long that takes is unmeasured and evidently long: one H2D upstream
ran 21m36s without stalling, and an earlier 512 B/s extrapolation of mine
was mostly the one-off startup banner. So this is a bounded resource being
treated as unbounded, not a fault anyone has reported.

_FfmpegStderrTail drains the pipe continuously and keeps a 16 KiB rolling
tail. That tail is what the error paths now report, which is better
material than before: it holds what ffmpeg said as things went wrong,
where the on-demand read returned whatever was printed first -- usually
the banner, which the summariser strips anyway.

Three readers wanted this one pipe, and asyncio rejects concurrent reads
on a StreamReader, so the collector is authoritative: it registers by pid,
_read_ffmpeg_stderr returns its tail when present and otherwise reads the
pipe unchanged, and _terminate_ffmpeg skips its own stderr drain when the
collector owns it (the collector keeps draining through teardown, which is
all wait() needs). The generator starts it after the immediate-failure
check, which reads the pipe directly because the process is already dead,
and releases it after _terminate_ffmpeg.

text() goes through _summarize_ffmpeg_stderr like every other stderr log
here, so the access code ffmpeg echoes in its input URL stays masked.
aclose() awaits the cancelled pump rather than firing and forgetting, so
no pending task survives into loop teardown.
2026-07-30 12:52:11 +02:00
maziggy f26bcbbcce fix(camera): one registry key per stream, not per printer (issue #2707)
Closing a camera view and reopening it immediately could leave the new
stream unregistered while it was running and delivering frames. The
damage was all indirect: is_stream_active() reported no viewer, so Obico
polling and snapshots opened a second camera connection against the live
view (the thing #1348 and #1271 exist to prevent); the janitor's /proc
scan found an ffmpeg missing from _active_streams and killed the live
stream as an orphan; and /camera/stop reported "Stopped 0" with a stream
running.

The fan-out stream id was f"{printer_id}-fanout" -- constant per printer,
so every successive stream shared one registry key, and the departing
generator's finally popped whatever was under it, including its
successor's entry. The same finally also cleared the per-printer frame
buffer unconditionally, discarding the new stream's frame. It needed the
two streams to overlap, which the 4s teardown made easy.

Each stream now gets its own key via _new_fanout_stream_id(), so a
generator can only clean up after itself -- the external-camera path
already does this (#2675) and this brings the fan-out path in line. The
per-printer dicts are released through _release_printer_frame_state(),
which checks that no other stream for the printer is still running; both
the RTSP and chamber-image cleanups had the same unconditional pop.

Also hoisted time and uuid to module level and dropped four
function-local `import time` statements. A local import shadows the name
for the whole function, so any use on a branch that doesn't reach the
import raises UnboundLocalError -- a real hazard in camera_stream, whose
external-camera branch imported both while the RTSP path needs them too.
A test pins camera_stream as free of function-local imports.
2026-07-30 12:35:27 +02:00
maziggy 18cc906fad fix(camera): drain ffmpeg's pipes during teardown (#NNNN)
Closing a camera view logged "ffmpeg didn't terminate gracefully,
killing" followed by "ffmpeg did not exit within 2.0s of SIGKILL;
abandoning wait", on every single close. Both waits expired every time,
so teardown took a fixed 4.00s -- and since the firmware allows one
camera connection, that was 4s in which nothing else could use it.

ffmpeg is spawned with stdout and stderr as pipes and the teardown paths
have stopped reading them, so it sits blocked in write() on a full 64 KiB
pipe. SIGTERM cannot be acted on there: the handler only sets a flag that
the main loop polls, and the loop never gets back to the check. SIGKILL
does kill it, but asyncio resolves Process.wait()'s waiter through
_try_finish(), which requires every pipe transport to report
disconnected; paused, unread pipes never reach EOF, so wait() blocks with
returncode already set. A negative-control test shows returncode=-9 at
the instant the abandon fires.

Draining both pipes while stopping the process fixes both halves: 4.00s
becomes ~0.15s. The signal ladder and its bounds stay as backstops, so a
genuinely wedged process still cannot hang a stream, a Stop request or
the janitor.

This corrects _FFMPEG_KILL_TIMEOUT's premise and #2580's conclusion. That
12-hour hang was the unbounded form of this same self-inflicted stall, not
an ffmpeg stuck in uninterruptible I/O -- the process observed doing it
was in state S, which cannot survive a delivered SIGKILL. Bounding the
wait capped the symptom without removing the cause.
2026-07-30 10:31:22 +02:00
maziggy 73afa95047 fix(camera): share one connection between concurrent one-shot captures (#2705)
Bambu firmware allows exactly one camera connection. The existing guards
(is_stream_active / try_get_active_buffered_frame, #1271 and #1348) only
stop a one-shot capturer from competing with the fan-out broadcaster.
Nothing coordinated the capturers with each other, so with no viewer
attached every consumer correctly concluded it was not competing with a
viewer and then collided with the others. On the reporter's P2S an Obico
poll and a snapshot opened two RTSP sockets 207 ms apart, which knocked
over the fan-out stream feeding the camera wall; it was then reaped for
having received no frames for 58s.

capture_camera_frame_bytes() now coalesces: the first caller opens the
connection, callers arriving while it is in flight await the same result.
Eight paths reach that function independently - Obico polling, the
snapshot route, the finish-photo moment and its disk-writing sibling,
plate detection, the camera test and the diagnose tool - so the
single-flight sits at the bottom of the stack and no call site changes.

Keyed by IP, since that is what the firmware's limit applies to and the
function never sees a printer_id. The key excludes the timeout on
purpose: the call sites disagree about it, from 10s to 30s, so keying on
it would mean the Obico-vs-snapshot pair from the report never coalesced
at all.

It coalesces, it does not cache. A call arriving after the previous
capture finished still captures fresh, because plate detection and the
finish-photo path judge a running print from these frames and a stale one
there is worse than a slow one - #1397 was a finish photo taken seconds
late showing the bed already lowered.

Each caller waits on its own deadline rather than inheriting whichever
one happened to open the connection, and shield() means giving up leaves
the capture running for whoever else is still waiting. A follower whose
leader fails takes a turn of its own instead of inheriting a failure it
never had a chance to avoid; the leader has finished by then, so there is
nothing left to compete with. Bounded at two rounds. That also covers the
follower whose timeout is longer than the leader's, which coalescing
alone cannot. Cancellation is disambiguated via leader.cancelled(), so a
follower's own cancellation propagates while a cancelled leader is
treated as a failed one.

The leader is deliberately not wrapped in a second wait_for: the
implementation already enforces the timeout internally, where it can also
kill the ffmpeg process, and an outer deadline would abandon the
subprocess instead of killing it.

The diagnose tool now marks a stage whose frame came from a capture
already in flight as coalesced_capture. The pass is real evidence the
camera works, but duration_ms is then mostly time spent queueing, and a
diagnostic must not report a connection it never opened - the same reason
that file declares its live_stream_active shortcut instead of quietly
passing. Failures are not annotated, since a follower whose leader fails
goes on to capture on its own.
2026-07-30 09:49:56 +02:00
gzimbric b938b83136 fix(tests): add the new fan fields to the plate-clear status fixture
mqtt_relay reads state.left_aux_fan_speed, but the SimpleNamespace fixture in
test_plate_clear_mqtt_notification enumerates its fields explicitly, so the two
status-payload tests raised AttributeError. I updated the equivalent fixture in
test_printer_manager_status_broadcast and missed this one — running only the
touched suites is what hid it.

Also addresses the round-2 review notes:

- exhaust_fan_present: documented that the H2 series reports part 3 too, so the
  flag is not model-specific despite the name.
- Mask the part id after shifting, matching get_flag_bits(id, 4, 8), for
  consistency with the state decode. No behaviour change for any observed id.
- Noted the unmapped H2 id 6 beside the id branches.
- Reject fan=aux2 when the printer reports no left_aux_fan_speed, so a POST
  against an A1 no longer sends M106 P10 for absent hardware. The UI already
  hid the badge; this closes the same hole on the API.
- Added a test asserting EXHAUST_FAN_LABEL_MODELS and the frontend's
  MODELS_WITH_EXHAUST_LABEL cannot drift apart.
2026-07-29 11:33:09 -05:00
maziggy 6cda236dce feat(notifications): optional Telegram forum topic via message_thread_id (#1518)
Telegram groups with Topics enabled always received notifications in the
General topic, since only Bot Token and Chat ID were configurable. Splitting
notifications per printer meant running a separate chat for each one.

The Telegram provider now takes an optional Forum Topic ID - the last number
in a topic's link, t.me/c/1234567890/25 - and routes its messages there. Left
empty, nothing changes.

The value is coerced to an int once in _send_telegram and attached to both the
sendMessage JSON body and the sendPhoto form data. That ordering matters:
Telegram rejects a string message_thread_id in the JSON body while accepting
one in the multipart call, so passing the raw form value through would have
worked for thumbnail notifications and 400'd for plain-text ones. A
non-numeric value is rejected in the form and again server-side before any
request goes out.

No migration - provider config is a JSON blob.

Adds Forum Topic ID plus help text to the Telegram section of the provider
dialog, translated in all 13 locales. Backend tests cover omitted / blank /
int-typed / non-numeric values and both send paths; frontend tests cover the
field being optional, absent for other providers, round-tripping on save, and
blocking save on a bad value.
2026-07-29 14:10:38 +02:00
MartinNYHC f2babc9027 Merge branch 'dev' into feature/p2s-x2d-accessory-fans 2026-07-29 12:36:45 +02:00
maziggy e325948dcc Security hardening (maziggy/bambuddy-security #N)
Subprocess output and user-supplied URLs are scrubbed of credentials
before they reach the application log. Adds a shared redaction helper in
core/logging_filters and routes the existing support-bundle sanitizer
through the same pattern.
2026-07-29 12:36:17 +02:00
maziggy 5a67dffe4f fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse
found nothing afterwards. Across 247 support bundles this was the norm, not an
edge case: 457 automatic scans scheduled, 262 attached.

The scan looked four times over ~65s. The attempt that found the video was #1
272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
files were still arriving when we stopped. What ran afterwards searched for the
print name inside the filename; Bambu only writes "video_<timestamp>", so it
fired 159 times and matched zero.

The manual Scan had no baseline at all and matched on filename timestamp, FTP
mtime, or "there is only one video" — all reading a clock a LAN-only printer
cannot sync. The reporter's P1S was six and a half days out.

- Poll for minutes instead of ~65s; drop the name-match fallback.
- Persist the print-start baseline on the archive, so the diff survives a
  restart mid-print and the manual Scan runs the same comparison. With a
  baseline present the clock-based strategies are skipped entirely — they can
  only turn an honest "pick one" into a confident wrong answer.
- When several files are new (a previous print's video landing late), exclude
  the ones already attached to another archive instead of ordering the
  candidates. Ordering could only be done on the printer's clock.
- Delete the video from the printer once archived. Keeps /timelapse to
  unclaimed files, which is what makes the diff unambiguous, and stops P1S
  cards filling with AVIs.
- Gate that delete on a verified transfer: download_file now compares against
  the size from the listing. An FTPS connection closing early does not always
  raise, so a partial buffer was being attached as a complete video — which
  would also have been the one case where deleting the source lost data.

Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
the deadline stops bounding the loop as soon as the sleeps are shortened.
Per-round logging only speaks when the listing changed — 31 rounds of full
listings would bury the interesting line in the support bundle.

Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
both dialects so a migrated database matches a fresh one.

-----------

fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)

When a print records a timelapse, its last frame is the better finish photo:
the firmware stops recording with the toolhead parked and before the end
G-code drops the bed, where a live grab at that moment catches a lowered
plate. Bambuddy waited 60s for the video and then gave up, because the
print-complete notification blocks on that photo and holding a notification
for minutes is worse than sending it with the live grab.

P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
worst 546s, while every other model finished inside 26s. So the printers that
most needed the better framing were the ones that never got it.

Keep the notification on the same bound, and keep waiting off to the side.
_capture_finish_photo_from_timelapse now reports whether it ran out of time or
concluded — a video that landed and failed extraction is not worth retrying,
one that never arrived is. On the first, schedule a background task that waits
up to 15 minutes and inserts the extracted frame at the front of the archive's
photo list, where the gallery opens.

The live grab stays on disk: the notification already links to that exact
file, so removing it would leave a broken image in Discord or Telegram.

The length check proves we received what the listing said, not that the file
was finished. The first look happens ~5s after the print ends, while the
printer may still be writing, so a growing file can be listed short, served
short, and pass. Re-list after the download and only accept the video once its
size has stopped changing — a failed re-list counts as not settled, since
"could not check" must not mean "safe to delete".
2026-07-29 11:51:30 +02:00
maziggy c1a4c99059 feat(diagnostics): log end-of-print telemetry for finish-photo trigger research (#2547)
The finish photo needs a "printing done, toolhead parked, filament unload
not started" moment. stg_cur=22 was meant to be it (#1721) and fires on no
model in the field: across 247 support bundles there is not one
FINISH PHOTO MOMENT (stage-22), including the 2026-06-13..07-08 window where
it was the only pre-FINISH trigger in the code — 104 captures on A1, A1 Mini,
H2C, H2D, P1S, P2S, X1C and X2D, all of them the FINISH fallback.

A replacement can't be designed from the bundles we have. Out of that window
Bambuddy parses only stg_cur and mc_print_sub_stage; every other stage/action
field arrives and is dropped unread. The candidates that sound right
(print_real_action, mc_action, mc_stage) are absent from A1/A1 Mini/P1S
payloads, so none of them can be the universal answer alone.

Dump the raw fields for the window between the last object layer and
gcode_state=FINISH at DEBUG. Opens on the first end-of-print signal (last
layer, progress >= 99, or no remaining time) so a dropped layer_num packet
doesn't lose it, logs only what changed frame to frame, closes on the
transition out of RUNNING, and arms once per print.

Instrumentation only: gated on DEBUG being enabled, read-only against printer
state, wrapped so it cannot break ingest, and capped at 400 frames per print.
The probed fields are stage codes, counters and bitfields — nothing
identifying, and no access code.
2026-07-29 09:48:18 +02:00
maziggy 91269f14fe fix(mqtt): report why a printer refused the connection instead of looping silently
A printer with a wrong access code gave no explanation anywhere. The connect
callback's failure branch was a bare `state.connected = False`, discarding the
CONNACK reason code the printer had just sent, so the only trace was paho's
follow-up disconnect -- logged every 30 seconds as "rc=Unspecified error",
which is exactly what a powered-off printer produces. In the report behind this
fix one of three printers had been in that loop for the whole capture, and
neither the log nor the support bundle could say why.

Bambu speaks MQTT 3.1.1, whose CONNACK return codes 4 and 5 paho maps onto
reason codes 134 and 135. Both are now logged with the printer's own reason
string and, for those two, the remedy: the access code is regenerated whenever
LAN Only or Developer Mode is toggled, so it has to be re-read from the screen.
The access code itself is never logged -- it would land in every bundle.

The reason is kept on the client as a stable slug and plumbed through
test_connection into the connection diagnostic, which now distinguishes two
cases it previously conflated. "The printer refused our credentials" is
asserted only when the printer said so; when all Bambuddy knows is that there
is no session, the text hedges and names the alternatives (rebooting, or
already at its limit of simultaneous connections). The old wording claimed the
access code was most likely wrong in both cases.

Frontend needed no change -- ConnectionDiagnostic already renders
`<status>_<reason>` variants with fallback to the plain per-status text, so an
unrecognised slug degrades to today's wording rather than a missing key.
2026-07-29 09:02:39 +02:00
maziggy 3a6cf04a7e fix(vp): relay A2L AMS filament to the slicer instead of blanking every slot
Every slot of the A2L's AMS Lite rendered as "?" in Bambu Studio through the
Virtual Printer while Bambuddy's own AMS card was correct, and a filament set
by hand in Studio reverted about a second later.

The A2L reports its AMS Lite as physical unit id 16 but packs the slot presence
bits at base 24, so bambu_mqtt normalises the id to 6 at the ingest boundary and
every internal reader gets the right bits. The VP bridge is not downstream of
that: BambuMQTTClient._on_message fans raw payload bytes out to raw-message
handlers before parsing, so mqtt_bridge._on_printer_raw does its own json.loads
and still holds id 16. It then called the shared apply_tray_exist_bits, which
computed 16*4 = bits 64-67 -- never set -- concluded all four slots were empty,
and wiped tray_type / tray_color / tray_info_idx / tag_uid / tray_uuid / remain
from the copy sent to the slicer. That runs on every push, which is why a manual
pick could not survive the next 1 Hz cached-as-base report.

apply_tray_exist_bits now folds the unit id through normalize_am_unit_id, so 16
and 6 land on the same bit base whichever id the caller holds. The bridge's
cached ids stay physical on purpose -- Studio addresses the Lite as 16, sending
ams_get_rfid {ams_id: 16} through the VP -- so normalising the cache instead
would have broken the slicer's own command path.

Confirmed from the reporter's debug log, which shows the cleanup clearing slots
at bits 64-67 under the VP's log label. Before #2670 added the
0 <= ams_id <= 15 range guard this wiped the slots; after it, unit 16 fell out
of the guard and the A2L got no empty-slot cleanup at all -- two different wrong
answers, both fixed here.
2026-07-29 08:41:33 +02:00
maziggy d0efb9db9e fix(vp): relay A2L AMS filament to the slicer instead of blanking every slot
Every slot of the A2L's AMS Lite rendered as "?" in Bambu Studio through the
Virtual Printer while Bambuddy's own AMS card was correct, and a filament set
by hand in Studio reverted about a second later.

The A2L reports its AMS Lite as physical unit id 16 but packs the slot presence
bits at base 24, so bambu_mqtt normalises the id to 6 at the ingest boundary and
every internal reader gets the right bits. The VP bridge is not downstream of
that: BambuMQTTClient._on_message fans raw payload bytes out to raw-message
handlers before parsing, so mqtt_bridge._on_printer_raw does its own json.loads
and still holds id 16. It then called the shared apply_tray_exist_bits, which
computed 16*4 = bits 64-67 -- never set -- concluded all four slots were empty,
and wiped tray_type / tray_color / tray_info_idx / tag_uid / tray_uuid / remain
from the copy sent to the slicer. That runs on every push, which is why a manual
pick could not survive the next 1 Hz cached-as-base report.

apply_tray_exist_bits now folds the unit id through normalize_am_unit_id, so 16
and 6 land on the same bit base whichever id the caller holds. The bridge's
cached ids stay physical on purpose -- Studio addresses the Lite as 16, sending
ams_get_rfid {ams_id: 16} through the VP -- so normalising the cache instead
would have broken the slicer's own command path.

Confirmed from the reporter's debug log, which shows the cleanup clearing slots
at bits 64-67 under the VP's log label. Before #2670 added the
0 <= ams_id <= 15 range guard this wiped the slots; after it, unit 16 fell out
of the guard and the A2L got no empty-slot cleanup at all -- two different wrong
answers, both fixed here.
2026-07-29 08:40:42 +02:00
Gabe 84a7b797cd fix(printers): decode airduct part state from its low 8 bits
Review feedback on #2691: `state` is bit-packed like its sibling `range`
(end << 16 | start), and Bambu Studio decodes it with
get_flag_bits(state, 0, 8). Masking with & 0xFF before clamping means a
packed value decodes to the real percentage instead of clamping to 100.

Also moves the uses_exhaust_fan_label import to the top of printers.py with
the other imports.

Tests: packed value (60 << 16 | 45) decodes to 45, and plain 0-100 values
round-trip unchanged.
2026-07-28 11:16:23 -05:00
Gabe 15ec0bf1c5 fix(printers): use the model-appropriate name in the fan-speed response
The fan-speed endpoint always reported 'Chamber fan set to N%', so on
P2S/X2D — where the printer card labels that fan 'Exhaust' — clicking
Exhaust produced a toast saying Chamber fan.

Adds uses_exhaust_fan_label() to printer_models so the badge label and the
API response share one source of truth, and uses it to pick 'Exhaust fan'
vs 'Chamber fan' in the response message.

Tests: helper coverage for P2S/X2D (incl. internal codes N7/N6), other
enclosed models, and unknown/missing model; API test asserting the message
matches the badge label per model.
2026-07-28 11:16:23 -05:00
Gabe 9ee162d51a feat(printers): expose P2S/X2D accessory fans (left aux + exhaust)
The P2S/X2D have two fans bambuddy didn't fully handle. On the P2S both are
add-on kits; on the X2D they ship from the factory.

1. Left auxiliary part cooling fan — not shown or controllable. It is reported
   ONLY as device.airduct part id 10 (raw id 160 >> 4; FAN_REMOTE_COOLING_1 in
   Bambu Studio's DevFan::ParseV3_0) and is never mirrored into a flat
   big_fanX_speed field, which is why it was invisible. This is the gap
   identified in #2576, where the single 'Auxiliary' fan (big_fan1 / M106 P2)
   only reaches the right-hand aux fan.

2. Chamber exhaust fan — shown on every P2S labelled 'Chamber Fan'. On P2S/X2D
   Bambu's firmware/UI (and Bambu Studio's FAN_CHAMBER_0_IDX) call it 'Exhaust',
   and it is a kit on the P2S rather than built in.

Both are now detected from device.airduct.parts, which lists only the fans that
physically exist, so each tile appears only when the hardware is present.

- bambu_mqtt: parse airduct part 10 -> left_aux_fan_speed (None when absent) and
  part 3 presence -> exhaust_fan_present; set_fan_speed() accepts index 10 plus a
  set_left_aux_fan() helper
- schema / status route / printer_manager broadcast / mqtt_relay expose both fields
- POST /printers/{id}/fan-speed accepts fan=aux2 -> M106 P10, the command Bambu's
  official P2S machine profiles use
- frontend: 'Left Auxiliary Fan' tile shown when reported; big_fan2 tile labelled
  'Exhaust' and presence-gated on P2S/X2D, unchanged 'Chamber Fan' elsewhere
- i18n: leftAuxiliary + exhaust for all 12 locales

Verified fan -> field map on a live P2S (fw 01.02.00.00), stable across cooling
and heating airduct modes:
  Part cooling -> cooling_fan_speed / airduct id 1  (built in)
  Aux          -> big_fan1_speed    / airduct id 2  (built in)
  Exhaust      -> big_fan2_speed    / airduct id 3  (kit)
  Left aux     -> airduct id 10 only (kit; forced off in heating by mode config)

Tests: airduct id-10 parsing (raw 160 -> id 10, not literal 160), id-3 presence,
base-P2S absence, diff-push survival, clamping, malformed entries, M106 P10
emission, invalid-index rejection; fan-speed API aux2->10 mapping; frontend tile
presence and labelling per model/kit.
2026-07-28 11:16:22 -05:00
maziggy dd171252dc fix(cloud): complete the CSRF handshake on Bambu Cloud TOTP sign-in (#2696)
Signing in to Bambu Cloud with an authenticator-app account failed every
time with "Invalid code", whatever the code was. Bambu Lab added double-
submit CSRF protection to the bambulab.com web origin - which is where,
and only where, this service posts the two-factor code. The endpoint
refused the request with 403 "CSRF error: missing_cookie" before it ever
evaluated the code, and Bambuddy reported that refusal as a bad code.

Verified against the live endpoint with a deliberately invalid key: a
bare POST returns missing_cookie; GET /api/csrf mints a bbl_csrf_token
cookie; a POST carrying only the cookie returns missing_header; a POST
carrying the cookie plus an x-bbl-csrf-token header reaches application
logic. Landing on the sign-in page first - the intuitive fix - does not
help, as that page sets only Cloudflare's __cf_bm. Of five header
spellings tried, only x-bbl-csrf-token is accepted, so the tests pin it.

verify_totp now performs that handshake against the same origin it will
post to (bambulab.cn for the China region - a token minted by the global
site is a cookie the .cn endpoint never issued), and declines to submit
the code at all when no token can be obtained rather than burning the
user's 30-second TOTP window on a request that is certain to be refused.
A CSRF refusal now also says the code was never checked instead of
masquerading as a wrong code, which is what sent the reporter chasing
clock drift and leading-zero parsing.

Only TOTP sign-ins were affected. Every other cloud call, the email-code
two-factor path included, goes to api.bambulab.com, which is not gated,
and existing stored tokens were unaffected throughout.

The region-routing test's MockTransport needed teaching about the
handshake: it returns one canned response for every request and set no
cookie, so the fix correctly refused to POST and the test lost the URL it
asserts on. It now mints a token for /api/csrf and additionally checks
the handshake stays on the .cn origin.
2026-07-28 15:08:00 +02:00
maziggy 4059d63373 Bumped version 2026-07-28 15:07:28 +02:00
maziggy 9d549050f7 fix(cloud): complete the CSRF handshake on Bambu Cloud TOTP sign-in (#2696)
Signing in to Bambu Cloud with an authenticator-app account failed every
time with "Invalid code", whatever the code was. Bambu Lab added double-
submit CSRF protection to the bambulab.com web origin - which is where,
and only where, this service posts the two-factor code. The endpoint
refused the request with 403 "CSRF error: missing_cookie" before it ever
evaluated the code, and Bambuddy reported that refusal as a bad code.

Verified against the live endpoint with a deliberately invalid key: a
bare POST returns missing_cookie; GET /api/csrf mints a bbl_csrf_token
cookie; a POST carrying only the cookie returns missing_header; a POST
carrying the cookie plus an x-bbl-csrf-token header reaches application
logic. Landing on the sign-in page first - the intuitive fix - does not
help, as that page sets only Cloudflare's __cf_bm. Of five header
spellings tried, only x-bbl-csrf-token is accepted, so the tests pin it.

verify_totp now performs that handshake against the same origin it will
post to (bambulab.cn for the China region - a token minted by the global
site is a cookie the .cn endpoint never issued), and declines to submit
the code at all when no token can be obtained rather than burning the
user's 30-second TOTP window on a request that is certain to be refused.
A CSRF refusal now also says the code was never checked instead of
masquerading as a wrong code, which is what sent the reporter chasing
clock drift and leading-zero parsing.

Only TOTP sign-ins were affected. Every other cloud call, the email-code
two-factor path included, goes to api.bambulab.com, which is not gated,
and existing stored tokens were unaffected throughout.

The region-routing test's MockTransport needed teaching about the
handshake: it returns one canned response for every request and set no
cookie, so the fix correctly refused to POST and the test lost the URL it
asserts on. It now mints a token for /api/csrf and additionally checks
the handshake stays on the .cn origin.
2026-07-28 15:05:28 +02:00
maziggy 8551e32f14 feat(slicer): keep the designer's print settings when re-slicing for another printer (#2622)
Published models often deviate from the stock Bambu profile on purpose -
five walls, 100% infill, a 0.1mm first layer. Re-slicing one for a
different printer discarded all of it: the picked process preset
overrides the file's embedded settings, and that override is precisely
what makes cross-printer re-slicing work, so it cannot just be dropped.
"Slice as designed" (#2611) does not help - it is all-or-nothing and
only offered when the picked printer already matches the design's
target.

The deviation list does not have to be computed. Bambu Studio writes it
into the 3MF as different_settings_to_system, laid out as
[process, *filaments, printer] - verified against real files at 2, 3 and
4 filament slots. The parser refuses any file whose array length
contradicts its own filament count rather than guessing an index, since
reading the printer slot as the process slot would carry the designer's
machine_start_gcode onto a foreign printer.

The slice dialog now lists exactly which print settings the author
changed and what each was set to, with a checkbox per setting. Design
intent - wall count, infill, layer and first-layer height, supports,
seam, brim, ironing - is ticked by default. Printer-specific values -
speeds, accelerations, jerk, fans, temperatures, prime-tower geometry -
are listed with a badge but start unticked: tuned for the author's
machine, they can be merely wrong on the target or outside the range its
profile accepts, which fails the slice outright.

Only ticked keys are sent, and only keys the source actually flags as
changed are applied. Values are written into the outgoing process JSON,
the same mechanism the support carry-over has used since #1881: for a
Standard preset pick that JSON is an inherits stub, so the patch is the
child in the chain and wins over the flattened parent. Process slot
only - filament picks are honoured as chosen.

The wiki's "this is not a settings merge" note under Slice as designed
described the gap this closes; rewritten to point at the new panel.

Translated in all locales; wiki updated. Covered by backend and frontend
tests.
2026-07-28 14:49:50 +02:00
maziggy 8fd1f884dc feat(mqtt): publish the plate-clear gate and add a notification for it (#2525)
When a print reaches a terminal state Bambuddy holds the queue until
someone confirms the build plate is clear. That gate was visible only in
the Web UI: the printer's own MQTT push reports nothing beyond RUNNING,
PAUSE, FAILED, FINISH and IDLE, so an external automation could not tell
"finished" from "finished and still waiting for a human".

The per-printer status topic now carries an awaiting_plate_clear field,
and every transition is additionally published on a new retained topic,
bambuddy/printers/{serial}/plate_clear. Retained, and published from the
flag itself rather than from printer telemetry: a subscriber learns the
state of every printer the moment it connects, and the state stays
correct after Auto Off powers a printer down - telemetry stops there,
which would otherwise leave the status topic frozen at false.

Publishing is edge-triggered. The queue clears the gate on every
dispatch whether or not it was up, and no subscriber should see a
"plate cleared" for a plate that was never dirty. Persistence and the
WebSocket broadcast stay unconditional; they are idempotent and predate
this.

A matching Plate Clear Required notification event was added, off by
default on every provider because it fires after every print at the
same moment as the print-complete alert. Only the rising edge notifies.
Acknowledging still goes through POST /printers/{id}/clear-plate.

Two tests in test_printer_manager_status_broadcast.py asserted
_schedule_async.call_count == 2 for the setter. The new emission makes
it three on a transition, so they now assert that the persist and
broadcast coroutines are actually scheduled - which is the contract

Translated in all locales; wiki updated. Covered by backend and
frontend tests.
2026-07-28 13:36:55 +02:00
maziggy af7874546a feat(projects): per-file print progress and complete-sets tracking (#1897)
Projects made of many distinct files that each need N prints (e.g. 13
plates x 10 sets = 130 prints) only had aggregate progress. Finding out
"how many times have I printed plate_7?" meant reading the Activity
Timeline line by line, unusable at 130 events.

Projects now take an optional Copies per File target. Every printable
file in the project's linked folders shows an X / N badge with a mini
progress bar (gray not started, amber in progress, green done), and the
progress card gains a Complete Sets bar - the minimum per-file count,
i.e. how many finished assemblies can be shipped right now. Without the
target, printable files show a plain printed-count badge.

Counting matches the aggregate project stats: completed runs only,
served by a new /projects/{id}/file-progress endpoint. Runs attribute
to a file via a new library_file_id stamp on queue-dispatched archives,
falling back to content hash and then filename for historical rows.

Also fixed: files queued from a project-linked File Manager folder now
inherit that project, so their prints count toward project statistics -
previously only prints started from the project page were attributed.

Test-harness fix along the way: the test suite's get_db override never
committed, unlike production get_db, so endpoints relying on the
request-scoped commit silently lost their writes in tests. The override
now mirrors production commit/rollback semantics.
2026-07-28 12:46:31 +02:00
maziggy 1fb6978ee1 feat(library): let users delete empty folders (#1781)
Library folders have no ownership tracking, so folder deletion was
gated entirely behind library:delete_all - a user with
library:delete_own could create folders and delete their own files,
but the emptied folder sat there until an admin removed it.

Users with library:delete_own can now delete folders that are truly
empty: no subfolders and no files, including trashed ones - folder
deletion cascades, so removing a folder that holds another user's
trashed file would silently break trash restore. External folders
(operator-configured mounts) and folders linked to a project or
archive still require library:delete_all even when empty, since
deleting them affects more than the folder itself. The bulk-delete
endpoint applies the same rule instead of skipping all folders for
non-admin users.

The folder tree's Delete entry enables accordingly and shows a
"You can only delete empty folders" hint on non-empty folders. The
backend stays authoritative - a folder that only contains trashed
files is invisible in the tree but still refuses deletion.
2026-07-28 12:05:00 +02:00
maziggy eae5359fbc feat(printers): show AI failure detection state on printer cards (#1546)
The live Obico classification was only visible under Settings ->
Failure Detection, so tracking how detection matched an ongoing print
meant flipping between the Printers screen and Settings.

Each printer card's badge row now shows an AI badge whenever detection
is enabled for that printer, like the other health badges: gray Idle
while no print is being watched, then green Safe, amber Warning, or
red Failure while a print is actively monitored. The tooltip carries
the current smoothed score; clicking jumps to the full detection
status and history in Settings. Printers excluded from the monitored
subset show no badge.

Served by a new lightweight /obico/printer-status endpoint readable
with printer permissions alone - it exposes only the enabled flag, the
monitored-printer set, and per-printer classification, keeping ML URL
and other configuration behind the existing settings-gated endpoint.
2026-07-28 11:37:36 +02:00