Commit Graph
2147 Commits
Author SHA1 Message Date
maziggy c204c79363 fix(jog): send the nozzle-bed gap the API promises on every model (issue #1334)
POST /printers/{id}/bed-jog takes a signed nozzle-bed gap, documented since it
    was written: positive asks for more room between the nozzle and the plate. On
    an A1 it did the opposite. The reporter sent distance=5 for clearance and
    watched the toolhead come down.

    The sign had been flipped on A1 models since the original report on this issue,
    where an A1 Mini owner clicked an arrow labelled "move the plate up" and watched
    the nozzle dive. That is a labelling problem -- a bed-slinger's plate does not
    move in Z at all, so closing the gap shows up as the toolhead descending -- and
    it was solved in the transport layer, which turned a parameter documented as
    model-independent into one that meant the opposite thing on part of the fleet.

    Z is the nozzle-to-bed distance on every Bambu model, by definition of the
    coordinate system rather than by convention: G1 Z+ opens the gap whether the bed
    drops away from a fixed nozzle (X1/P1/H2, whose end G-code parks with
    G1 Z{max_layer_z + 100}) or the nozzle rises off a fixed bed (A1/A2L). The
    finish-photo plate restore already relies on exactly that and carries no model
    branch. So distance goes onto the wire unchanged and one call means one physical
    outcome everywhere: positive is the safe direction on every printer.

    Which way an arrow points is a different question, about the machine in front of
    the user rather than about G-code, so the printer card answers it and asks for
    the gap it wants. The buttons move what you would expect them to move, exactly
    as before; on a bed-slinger they now say toolhead rather than plate.

    The A2L never had the old fix. It slings its bed the same way the A1 does, but
    the inversion listed the A1 names and the A2L was not among them, so its up
    arrow has been sending the toolhead at the plate for as long as the machine has
    been supported. The new classifier also covers the alternate internal codes
    A04 / A11 / A12, which LINEAR_RAIL_MODELS and SINGLE_NOZZLE_FLOW_MODELS both
    carry and the old gate did not.

    is_bed_slinger is gone from the backend rather than widened: with the route
    model-independent it had no caller, and a kinematics helper sitting unused in
    the service layer invites the next person to assume the backend handles
    direction. It does not, deliberately.

    Separately, the soft-endstop comments on both jog routes claimed the firmware
    clamps a bare move at the travel limit. It does not, and #2579 measured that:
    an H2D at its Z limit ran straight past a clean G91/G1 Z-1.00/G90, while its own
    touchscreen refuses the identical move. What #2579 removed was M211 S0, which
    disabled the limits globally and took the touchscreen's protection with them.
    The jog popover has warned about this correctly the whole time; only the code
    comments disagreed with it.
2026-09-20 13:36:46 +02:00
maziggy 1f88b9846f fix(queue): tell a pinned queue item why it is waiting (issue #3074)
A job queued as "Any X1C" explains itself when it cannot start: the
    model-based branch builds a reason for every candidate printer and puts it
    on the row, so the queue shows "Busy: X1C-01" or "Waiting for filament:
    X1C-02 (needs PETG)". The same job pinned to one printer showed nothing.
    It sat at Pending with waiting_reason NULL for as long as that printer was
    busy, which from the outside is indistinguishable from a queue that has
    stopped working -- the reporter watched fourteen minutes of it while his
    X1C ran a print he had started from its own screen.

    The fixed-printer branch had six ways out and none of them wrote the field.
    The sensor interlock (#1148) was its only writer, and it cleared the field
    up front on every pass where no sensor was holding the printer, so NULL was
    not an oversight on those paths but a guarantee.

    Every exit now writes, through one helper. The reasons reuse the
    model-based branch's vocabulary so _is_busy_only() keeps deciding what is
    worth a notification: a printer that is printing, drying, or working
    through the item ahead of this one reads as "Busy: <printer>" and stays
    silent, because it resolves itself. A printer that is off with no Auto On
    plug, and one whose plug could not switch it on, are worth saying.

    A finished plate nobody has acknowledged is split out from plain busy and
    named as itself. _is_printer_idle() returns the same plain False for that
    and for a running print, but they are not the same thing to the person
    looking at the queue: one clears itself and the other needs somebody to
    walk over to the printer.

    That notification fires on the transition into asking, where a busy-only
    reason counts as not asking. Testing whether the item was waiting at all --
    which is what the model-based branch does -- would never fire it here:
    nobody's queue goes straight from idle to an unconfirmed plate, it waits
    behind the print first. The cost is that a printer dropping offline,
    returning busy and dropping again asks twice rather than once.

    The interlock stays silent. It has never sent this notification, and a
    change about what the queue displays is not the place to start.

    Clearing the field up front is gone with it. It existed so a shut door
    could not leave "Waiting on Enclosure Door" standing while the printer
    stayed busy with something else, and the new rule carries that guarantee
    instead -- whichever exit runs next overwrites it, and the dispatch path
    clears it.

    Two paths clear it that the report did not mention. A staged item and a
    future-scheduled one skip before this branch and never reach it again, so
    anything written on an earlier pass would outlive its condition for the
    life of the row. That includes the filament-deficit check, which stages the
    item itself.

    The notification is wrapped: a queue that cannot say why it is waiting is
    the bug being fixed, and a queue that stops dispatching because a provider
    timed out would be a worse one.

    On the frontend, the queue timeline drops any pending item carrying a
    reason, on the grounds that such an item will not auto-dispatch. That held
    while only the model-based branch wrote the field; "Busy: <printer>" is
    now the commonest reason there is, and it describes the very chain the
    timeline forecasts, so the rule would have emptied the view for anyone
    whose queue is pinned. It now asks whether the reason needs the user, via
    a small shared reader of the same shape the scheduler encodes.

    Which job goes out, and when, is unchanged: running the previous scheduler
    and this one over the same 768 states dispatches the same items in the same
    order with the same statuses, across 1452 rows that now carry a reason.
2026-09-20 13:36:03 +02:00
maziggy 0b8cc823e7 fix(queue): print a plate whose filaments are all on the external spool (issue #3087)
The reporter's P1S heated up, sat at Heatbed preheating for ten and a half
    minutes, then paused with 07FF_8012, "Failed to get AMS mapping table".
    Resuming only reheated it. Prints that fed from the AMS were fine.

    The plate was one filament of a seven-filament MakerWorld project, mapped by
    hand to the external spool. slice_info.config numbers filaments across the
    whole project, so the mapping for that plate is [-1,-1,-1,-1,-1,-1,254]: six
    placeholders and the spool holder. The command builder decides whether a print
    needs the AMS by asking whether the mapping is entirely external, and six -1s
    answer no. So the print went out as use_ams=true carrying a flat mapping of
    nothing but -1 -- 254 is deliberately never sent raw, the firmware reads it as
    AMS tray 0 -- which is exactly the mapping table the firmware then could not
    find.

    The builder cannot fix this itself. Down there a -1 is either padding for a
    filament this plate does not print, which is BambuStudio's own convention, or a
    slot that never resolved to a tray, and sending the second one to the spool
    holder is what #2589 exists to prevent. They are the same byte.

    The scheduler knows. extract_filament_requirements drops every filament with
    used_g <= 0, so it names precisely the slots the plate prints. When all of those
    are an explicit 254/255, dispatch now sends use_ams=false and the print runs.
    When one of them resolved to nothing, the flag is left alone and the firmware
    rejects the print as it does today -- deliberately, because that is the case
    where guessing would print a filament in the wrong material without saying so.

    Single-nozzle only, mirroring the reconcile in the command builder: on a
    two-extruder printer use_ams selects which nozzle to feed rather than whether to
    use the AMS, so an H2D with a spool on each side must keep the flag it was
    given. Judged generously from the model name and from live telemetry -- a second
    nozzle reporting a diameter, an extruder map, or more than one external feed --
    because a wrong yes only preserves existing behaviour while a wrong no would
    reroute the print. H2S stays single-nozzle (#1386).

    Nothing in the command builder changed. Its own reconcile keeps the exact
    semantics #2589, #2595 and #797 gave it, and now usually agrees with a decision
    that was already made one layer up. Where the file has no parseable filament
    list the mapping is left exactly as before, the same evidence-only convention
    as #2771, and the parse itself sits behind a check for anything external at all
    so an AMS-only print never opens the file.

    Covered end to end at the dispatcher, including the reporter's seven-filament
    shape, a plate mixing the spool holder with an AMS tray, a consumed slot that
    never resolved, both external feeds on a dual-nozzle machine, and a 3MF with no
    filament list at all.
2026-09-20 13:35:25 +02:00
maziggy 14b322d86a fix(mqtt): never wait for a wedged paho network thread (issue #3068)
The reporter's A1 had been offline 38 hours and still answered on 8883, so
    the connection watchdog did exactly what it exists for: rebuild the session
    with a fresh client, since anything left in the old one's QoS 1 queue would
    otherwise replay onto the next print (#1136). The rebuild ended in paho's
    loop_stop(), which sets a terminate flag and then joins the network thread
    with no timeout.

    That thread only reads the flag between iterations of loop_forever, so it
    cannot read it while parked inside reconnect() -> _ssl_wrap_socket() ->
    do_handshake(). paho gives that handshake the keepalive as its socket
    timeout -- 30s here -- and a socket timeout is per operation, renewed by
    every byte the peer sends. A printer that answers TCP and then trickles
    holds the join open for as long as it likes.

    The join ran on the asyncio thread. Bambuddy stopped answering anything --
    UI, API, /health -- while the process stayed up, which is why a
    restart: unless-stopped container never restarted.

    Retiring a client no longer waits for it. The replacement is built at once
    and the old one is shut down on a thread of its own that nobody joins. Its
    callbacks are detached first, inline: blocking until the network thread was
    gone is what used to guarantee a client we had let go of could no longer
    touch our state, and with the teardown detached a zombie that finishes its
    handshake would otherwise auto-reconnect and report itself connected behind
    its replacement's back. disconnect() still goes out, still promptly, because
    that is what stops paho's auto-reconnect and the replay with it.

    The reported watchdog is one of six callers. The queue's dispatch recovery
    and check_staleness -- reached from an ordinary status poll -- share
    _hard_reset_client; editing, deleting and hand-disconnecting a printer share
    disconnect(); the relay and smart-plug services had the same join on their
    shutdown path, where a wedged broker stopped the process from exiting at
    all. #1445 was this join too, from the add-printer probe, and its off-loop
    teardown stays as it is.

    disconnect() stays quiet on the way out, as it always effectively did.
    paho's callback used to land during the join, but it suppresses itself for a
    clean disconnect of a printer that reported in the last ten seconds, so a
    healthy printer disconnected by hand never announced itself offline.
    Announcing it now would tell the user their printer had gone offline a
    minute after they disconnected it on purpose (#1752).

    A retirement that takes more than five seconds logs which printer it was.
    The whole point is that the next one of these should not have to be
    diagnosed from a thread dump.
2026-09-20 13:34:55 +02:00
maziggy 743d657b46 fix(drying): dry a composite spool as its base material (issue #3067)
The reporter's AMS-HT would not auto-dry PA6-CF, and drying the same spool by
    hand worked. The scheduler reduced a tray to a preset key by splitting on spaces
    only, so "PA6-CF" stayed "PA6-CF", matched none of the eight rows the preset
    table has, and the tray was read as holding nothing worth drying. The AMS was
    then passed over on every scheduler sweep, silently, because every caller reads
    "no row" as "nothing to do for this tray".

    It was never only nylon. Of the 41 types a printer can report, 33 had no row
    under that rule, and 20 of those have a base material sitting right there: every
    -CF, -GF and -AERO variant of PLA, PETG, ABS, ASA, PC and PA.

    Doing it by hand worked because the drying popover has resolved composites since
    exact key first, so a row the user added for the exact type still wins, then the
    suffix, then an alias map reading PA6, PA11, PA12, PAHT, PPA and Nylon as PA.

    PPA is the one alias that is a judgement rather than a spelling. Polyphthalamide
    is a distinct polymer, not a grade of nylon -- but it is an aromatic polyamide,
    it takes up moisture the same way, and PA's row is the hottest the table has.

    A material with no row and no alias is still skipped rather than dried at a
    number nothing here can source. That is where this parts company with the
    popover, which falls back to PLA because a dropdown has to show something.

    The table is user-editable JSON, so a preset row can be present and empty. That
    has always meant "skip this material" and still does: the temp and hours reads
    fall back per field to 55C/12h, which would dry a PLA spool at 55 degrees.

    Two more callers had the same line and move with it: per-filament humidity
    thresholds, where the override set for a material never applied to that
    material's composites, and the chamber preheat target, which had the suffix half
    of this from #2902 but not the aliases.

    The preheat test for that half reimplemented the lookup inline rather than
    calling it, so it would have passed whatever the function did. It calls it now.
2026-09-20 13:34:32 +02:00
maziggy 2b2617ecb7 fix(archives): come back for a 3MF whose transfer ran out of time (issue #3063)
The reporter's P1S had the sliced file on its card and was serving it. The 19MB
    transfer just did not finish inside the budget while the printer was also running
    its camera, its status messages and the upload of the job itself. Bambuddy wrote
    an empty fallback archive and never looked again -- then downloaded that same file
    successfully three times over the next two minutes and discarded every copy,
    because the only code that would have attached one had already run.

    The recovery machinery was there. It was armed for exactly one give-up, the FTPS
    cool-off, on the grounds that the three storage verdicts are settled: a job on
    internal eMMC never appears at any FTPS path, and sweeping for it again is what
    where the file is demonstrably still on the card.

    The sweep already had the signal and never used it. A file that is genuinely not
    there is answered with 550, which surfaces as FileNotOnPrinterError and is caught
    by name; a timeout returns falsy instead. So "the printer says no such file" and
    "we never got a straight answer" are distinguishable without guessing, and only
    the second schedules anything.

    Not scheduled either for a 3MF that downloaded fine and turned out to be another
    plate's. Recovery checks that a candidate is a readable 3MF but not which plate it
    holds, and the names a retry would use are the same stale ones that fetched the
    contradicted file -- so it would put back exactly what #2957 discards.

    The ladder follows the cause: a cool-off has to expire, so its first attempt sits
    past the 300s; nothing has to expire here, and this reporter's file completed 48
    seconds after the budget was spent.

    The archives banner gets its own wording for this, because the old text sends an
    owner whose card is working to switch on a setting that is already on. It names
    the Connection Timeout setting instead.
2026-09-20 13:34:10 +02:00
maziggy ad3b295312 fix(archives): let Items Printed go to 0 for a ruined plate (issue #3051)
A jam can destroy everything on the plate while the printer still reports the
    job as a success, so the honest count of usable parts is zero. The edit dialog
    floored the field at one, and a project's completed-items count sums that
    column, so there was no way to record that a job produced nothing.

    The floor was in the dialog only; the API stored whatever it was given, which
    also meant a negative count was accepted and would have subtracted from the
    project totals. The column is now bounded at zero instead.

    Filament Trends counted prints as `quantity || 1`, which would have read a
    deliberate 0 as "unset" and charged the ruined plate as one print while the
    project page counted none.
2026-09-20 13:33:46 +02:00
maziggy 2f7ec240fd fix(ams): stop reading the printer's command acks as status (issue #3040)
Every project_file carried "cfg": "0" — the device-config bitmask, which
    Bambu Studio has never sent and the firmware ignores. The printer echoes a
    command's fields back in its ack, and the ack was ingested as telemetry, so
    bit 18 read as "AMS Filament Backup off" 25 ms after every dispatch.

    Families that repeat cfg in their periodic status (P2S, H2C, X2D) corrected
    themselves a second later; the P1S, A1, A1 Mini and A2L send it only in a
    full status dump, so the wrong value stuck and silently disabled the
    prefer-lowest-remaining gate. The A1 family, which reports no cfg at all and
    is meant to stay "unknown", was pinned to a definite "off".

    Acks are no longer read as status, for the backup bit or the per-job
    timelapse flag they also echo, and cfg is gone from the print command.
2026-09-20 13:32:50 +02:00
maziggy 5ec99a7e06 fix(ams): resolve a slot's K profile by index when the printer does not file per hotend (issue #3044)
An X2D with two AMS 2 Pro, one per hotend, showed a K value on every slot
    of the first and nothing on any slot of the second. Configure Slot was
    worse than blank there: the picker offered no matching profile, the slot
    read as though nothing were bound, and choosing one changed nothing the
    user could see. Both symptoms are one rule.

    A calibration index can mean two different profiles on a dual-nozzle
    machine -- on the maintainer's H2C, index 16 is the left hotend's black
    PLA at K=0.018 and 15 is the right's at K=0.020 -- so the index is
    resolved against the slot's own hotend, and a miss shows nothing rather
    than the other nozzle's number. That is right whenever the printer files
    its calibrations per hotend. This one files them per filament: the second
    AMS's slots point at the same entries as the first, every entry tagged
    with one extruder, and requiring a match found nothing at all.

    The hotend now has to appear in the table the printer actually sent
    before it is used to narrow anything. Where it does not, the index stands
    on its own, which is what BambuStudio does for this same card --
    AMSItem.cpp resolves it through get_pa_k_n_value_by_cali_idx, matching
    cali_idx and nothing else. Where it does, nothing changes: the H2C case
    still blanks rather than borrowing, and the other hotend's profiles stay
    reachable under Other K profiles. The relaxed path still refuses an
    answer when the candidates disagree on a value.

    The premise that the table is always numbered per nozzle had been written
    into three comments and two layers of code; it is corrected where it
    appears.

    Alongside it, in the same picker: the K-profile options rendered the
    hotend suffix twice in the matching group and three times under Other, so
    every option on a dual-nozzle printer read "... . Left . Left".
2026-09-20 13:31:43 +02:00
maziggy 4440c95976 fix(queue): skip preheat entirely when no loaded filament wants a chamber (issue #3041)
Preheat & Heat Soak delayed every PLA print by five to seven minutes and
    gave nothing back. The filament map correctly derived a chamber target of
    0, and the chamber phase correctly skipped -- but the stage then heated
    the bed, waited for it, and held the full soak anyway, because the soak
    had no idea it was holding for a chamber nobody asked for. The print's
    own G-code sets the bed the moment it starts, so the bed phase only moved
    the warm-up ahead of the FTP upload instead of overlapping with it.

    A 0 that comes out of the filament map now skips the stage before any
    command goes out. The one thing the skip still does is put the airduct
    flap back to cooling on the models that have one -- an H2D left in
    heating mode by the ABS job before it would otherwise cook the PLA that
    follows, and that costs one MQTT command and no waiting.

    Explicit instructions are untouched. A chamber target of 0 typed into a
    print's own override still heats the bed and runs the soak, which is what
    the queue documentation has always promised it does, as does forcing a
    print's Preheat override to On. Prints that want chamber heat are
    unaffected, including the P1S/P1P/A1 tier where the bed and the soak
    timer are the whole mechanism.

    The existing unit tests all ran with soak_seconds=0, which is why the
    production default was never exercised; the PLA test now runs at the real
    default and asserts nothing is dispatched and nothing is slept.

    Surfaced in the UI on the way through: the Settings hint claimed the
    derived 0 skipped "the chamber phase", and the per-print chamber override
    field said nothing about a typed 0 meaning bed-only -- a user reaching
    for 0 to turn preheat off got the delay instead.
2026-09-20 13:30:36 +02:00
maziggy ad09406672 fix(slicer): strip zero-valued filament-index sentinels, and sanitise the preview slice too (issue #3030)
Bambu Studio writes 0 into wall_filament, sparse_infill_filament and
    solid_infill_filament to mean "use whichever filament the object is set
    to". Bambu Studio and OrcaSlicer 2.4 define these min 0 and accept it;
    OrcaSlicer 2.3 and earlier used the 1-based scheme (min 1, default 1)
    and reject it with "0 not in range [1.000000,...]". Sidecar images are
    version-tagged, so an install can be pinned to one of those builds.

    Same shape as the -1 inherit markers from #1201 with a different marker,
    so the allowlist becomes a key-to-marker map rather than one global
    constant. The buckets must not bleed: a -1 on a filament index is a real
    value, and a 0 on a raft field is a setting the user chose.

    The key is removed rather than rewritten, which is what makes it safe on
    every build. The CLI then uses its own default: 0 where 0 was legal
    (unchanged), 1 on the older builds, which is what "the active filament"
    means under that scheme.

    The preview slice never ran the sanitiser at all, so a file that sliced
    fine could still fail its automatic plate preview and fall back to the
    painted-face heuristic. It matters more there than in a real slice: the
    preview runs on the file's own embedded settings, so there is no
    --load-settings pass that could supply a replacement for a field the
    range validator has already rejected. That also explains the reported
    "same error on a later attempt of an unchanged file" without any second
    copy of the keys -- the validator that emits it reads the merged global
    config, which per-object model_settings.config overrides never reach.

    The sanitiser moves to utils/threemf_tools so the service can use it
    without importing a route module, and both preview callers pick it up
    from one place. Drops _strip_3mf_embedded_settings and its constant,
    which have had no callers since the strip-everything experiment was
    reverted.
2026-09-20 13:30:11 +02:00
maziggy 417d03d174 fix(slicer): keep protocol-handler download tokens valid for their whole TTL (issue #3029)
The Slice and Open in Slicer actions mint a short-lived token and put it in
    the URL, because a protocol handler cannot carry an Authorization header.
    That token was spent by the first request to reach the endpoint, which made
    the handoff depend on the slicer fetching the URL exactly once. Nothing
    guarantees that: Bambu Studio's downloader retries three times after a
    failed attempt, transfers get resumed, on-access scanners fetch. The first
    request won and the slicer was handed a 403.

    verify_slicer_download_token takes a keyword-only single_use flag. The
    default still consumes via DELETE...RETURNING; single_use=False verifies
    with a SELECT and leaves the row for the rest of its five-minute TTL. The
    stored row is the same either way, so the endpoint decides, not the mint.

    The three protocol-handler downloads pass single_use=False: a library file,
    an archive's sliced 3MF, an archive's source 3MF. Resource binding and
    expiry are untouched. The two browser downloads keep consuming, because
    what they hand over is itself consumed -- the prepared printer bundle is
    deleted the moment it has been streamed.

    Also: add "/source-dl/" to PUBLIC_API_PATTERNS. Those patterns match by
    substring and the source 3MF route's segment is source-dl, which does not
    contain "/dl/", so with auth enabled the middleware rejected the slicer's
    header-less request before the route's token check ran. Open source 3MF in
    slicer could never work on an install with authentication on.
2026-09-20 13:29:25 +02:00
maziggy e1fad9d68f fix(auth): decouple media routes from the camera stream token (issue #3025)
Thirteen routes with nothing to do with a camera took the camera stream
    token as their credential -- library and archive thumbnails, plate
    previews and plate thumbnails, timelapses, print photos, archive QR
    codes, project covers, print-log thumbnails, printer covers and
    external-link icons. A browser cannot put an Authorization header on an
    <img src>, so these need a credential that fits in the URL, and the
    camera token was the only one that existed. Minting one costs
    camera:view, so a user granted library access to their own files got a
    grid of broken images until they were also handed the live camera.

    Adds a media token: minted by POST /auth/media-token behind plain
    authentication, and identified -- it records the principal the way the
    websocket token does rather than being anonymous the way the camera
    token is. Each route now gates on the permission and ownership rules of
    the resource it serves, through the same _ensure_*_visible helpers its
    header-authenticated siblings already use. The three camera routes keep
    the camera token, and require_camera_stream_token_if_auth_enabled now
    documents that it is for those only.

    The media dependencies accept ordinary Authorization / X-API-Key headers
    as well as ?token=, delegating that path to the existing checkers, so
    API-key scope rules and the per-printer allowlist are unchanged.

    Long-lived camera_stream, camwall and overlay tokens are deliberately
    not accepted on the media routes -- those are handed to kiosks, walls
    and Home Assistant to display video. The cam wall, streaming overlay and
    kiosk views use only the three camera routes and are unaffected.

    Frontend: withMediaToken alongside withStreamToken, and
    useStreamTokenSync fetches a media token for every signed-in user while
    asking for a camera token only when the user can mint one, which also
    stops the 403 that fired on every page load for everyone else.

    Also fixed, same class:
    - /printers/{id}/files/plate-thumbnail/{i} is rendered in an <img> but
      had a header-only guard, so the file manager's plate thumbnails 401'd
      whenever auth was enabled. It now takes a media token too.
    - getProjectCoverImageUrl returned a URL ending in ?token=, and the
      project edit dialog appended its own ?v= cache-buster after it, so the
      second ? landed inside the token value. The version is now a parameter
      applied before the token.

    Tests: 15 integration tests for the token boundary, permission
    enforcement and per-row scoping; 10 frontend tests for the URL split and
    the two-query hook. test_cover_image_get_uses_stream_token_gate is
    renamed and repointed at the media gate -- what it pins, that the
    credential has to fit in a URL, is unchanged.
2026-09-20 13:28:50 +02:00
maziggy ef6446d30a fix(auth): let the sidebar read install flags without settings:read (issue #3023)
cost_centers:read_own exists so a non-admin can see their own wallet, balance
    and cost-centre spend, and the Finance page honoured it -- typing the URL
    worked and rendered their balance. The sidebar never offered the entry.

    It decides whether to show Finance by reading billing_enabled from
    GET /settings, which requires SETTINGS_READ. A non-admin gets 403 there, so
    the value arrived undefined, `undefined !== true` held, and the entry was
    hidden from precisely the users the permission was written for. The permission
    map and the route guard were both already right; only discovery was broken.

    Three more fields came from that same 403, and one of them failed the other way
    up. The Notifications gate tests `=== false`, which undefined never satisfies,
    so an administrator who switched user notifications off still left the entry
    showing to the non-admins it governs. Nobody reported that one, and no
    administrator could have reproduced either: administrators can read /settings.
    The remaining two were quieter -- the sponsor prompt fell back to EUR whatever
    the install uses, and the update check ran where it had been turned off.

    SETTINGS_READ cannot be the price of knowing whether billing is on. It also
    grants sight of the SMTP, LDAP and MQTT credentials, which is the reason
    /settings/ui-preferences exists at all.

    So: a second endpoint, GET /settings/ui-flags, carrying those four fields and
    asking only that the caller be signed in, via the existing
    require_auth_if_enabled. Layout drops its /settings query altogether, which
    closes the class rather than the two instances that happened to be visible.

    Deliberately not four more fields on /ui-preferences. That endpoint is served
    to anyone at all on the recorded grounds that its contents are "public defaults
    that ship with the app" (test_route_auth_coverage.py), and its field set is
    pinned by a test written to make anyone adding to it stop and think. These
    fields are not defaults -- they say how this deployment is configured -- so
    they get their own endpoint at their own trust level instead of stretching that
    charter to fit them. require_auth_if_enabled also keeps the auth-disabled case
    that /ui-preferences was ungated for: "works when there is no auth" and
    "readable by anyone" are different statements, and conflating them is what put
    a settings read in front of a permission that never needed one.

    Twelve tests. Backend pins that the operator can read the flags, that the same
    operator still gets 403 from /settings, that an anonymous caller is refused
    when auth is on, that it answers when auth is off, the exact field set, that no
    credential ever appears, and that the public endpoint did not quietly gain
    these fields. Frontend pins Finance visible for cost_centers:read_own with
    /settings returning 403, and Notifications hidden when the flag is off -- each
    waiting on a positive signal before asserting an absence, so the negative cases
    cannot pass before the query resolves.

    Reported by @lonix, who traced it to the queryKey and the route gate.
2026-09-20 13:28:17 +02:00
maziggy 23a6633f42 fix(queue): say when an unscheduled item runs instead of calling it ASAP (issue #3018)
The print dialog offers ASAP, Queue and Schedule. ASAP and Queue differ only
    in where the item is inserted, and neither is stored on the item -- scheduleType
    is a frontend-only concept, and grep finds no "asap" anywhere in the backend. So
    the queue's time column had nothing to read but scheduled_time, and labelled
    every unscheduled item "ASAP": the name of the one mode the user may well have
    chosen against.

    Someone who picked Queue then watched their row appear as ASAP and start
    immediately, and concluded Bambuddy had overridden them. Two reporters wrote
    that same sentence thirteen months apart, and #2557 was closed as A2L-specific
    after the first of them -- kilrah replied there with an X1C before filing this.

    The column answers when an item runs, so it now says that. The key is renamed
    whenFree rather than just retranslated: left called asap, the next translator
    puts ASAP back.

    The dispatch is unchanged, because it was right. A print scheduled for later
    does not reserve the printer until then; an unscheduled item behind it uses the
    idle printer rather than leaving an X1C dark until 6 AM. Two of the new tests
    pin that, so it does not get "fixed" later on the strength of a report like this
    one.

    What genuinely could not answer the question was the queue's own log. Its
    per-printer line called every entry in busy_printers "not available" -- but that
    set holds both printers that cannot take work and printers the pass has just
    claimed for some, which are opposite facts. It also read printer state at
    logging time rather than at the decision, so #3018's bundle carries

        Queue: printer 1 not available — connected=True, state=IDLE, ...
        Launching 1 upload(s) (pool 0/4 in flight)
        Starting queue item 18

    a printer reported unavailable, evidence that it was available, and a dispatch
    to it, in three consecutive lines. It is the first line anyone greps for "why
    did my item not go out".

    Each of the nine sites that removes a printer from a pass now records why, and
    the summary reports a claim as a reservation and everything else as an
    obstruction with its reason. The live fields stay, since a bundle reader wants
    them next, but are labelled as read now rather than offered as the cause.
    print_scheduler.py:1210 already documented that these two meanings differ -- the
    dispatching_printers snapshot exists for it. This carries that distinction into
    the log.
2026-09-20 13:27:43 +02:00
maziggy eab55cef75 fix(archives): report a refused FTPS handshake as the printer, not the slicer (issue #2780)
The Archives banner picks its wording from a priority list of the causes it
    knows. REASON_FTPS_COOLOFF was added by #2957 and never put in that list, so an
    install whose empty archives all came from a printer refusing the TLS handshake
    matched nothing, got reason: null, and fell to the original wording: the slicer
    did not leave the .gcode.3mf on the card, switch on "Store sent files on
    external storage", here is installation step 4.

    Every clause of that is wrong for this cause. The slicer did write the file --
    reason he read the whole thing as Bambuddy being broken. The setting was
    already on. And there is nothing on his side to change: the printer's file
    service answered port 990 with something that is not TLS, so no lookup ever
    ran and where the file went was never tested. It is #2899's mistake -- an
    error message describing a cause that was ruled out before it was printed --
    in a surface that did not get that pass.

    The slug now leads the list rather than joining the end of it. The other three
    describe an install working as configured and each ends in something the
    operator can change; this one reports a fault nobody can yet explain, which is
    both the more urgent thing to say and the thing that produces a useful report.
    The banner also dismisses one-shot into localStorage, so a reason ranked below
    another is not deferred to next time -- it is never shown to that user again.
    Ranking it first cannot bury a permanent cause in exchange: a successful
    recovery clears the row's markers (#2957), so a row still carrying this slug is
    one whose retry failed too, days after the print.

    New wording in all fourteen languages says the printer refused the connection,
    that this is not a slicer setting and not something the operator did, that
    Bambuddy comes back for the file when the five-minute pause clears so a brief
    episode fills itself in, and that a card still empty means the refusal outlasted
    the retry. It links to the handshake entry in the troubleshooting guide instead
    of to the installation guide.

    The client's getNo3MFWarning type still declared the old three-slug union, which
    made all three new comparisons provably dead -- caught by tsc, not by any test.

    Four tests. One pins the slug reaching the banner, one pins it outranking the
    three settled causes, one pins those three keeping their order behind it, and
    one asserts the rendered wording carries no slicer advice at all.

    Also corrects the wiki page these reports are pointed at. It said to power-cycle
    the printer; the reporter who prompted that advice power-cycled both of his and
    the failure continued unchanged, and bambu_ftp.py has carried the retraction in
    a comment since. The page now states what was actually measured -- that a
    version mismatch reports itself differently, that every printer probed refuses
    TLS 1.3 and completes on 1.2 so there is no version to fall back from, and that
    three P2S units failed while three more on the same switch never did -- says
    plainly that the trigger is unknown, and names the one cleartext-probe line
    worth collecting.
2026-09-20 13:27:10 +02:00
maziggy 58df1cb866 feat(ftp): log how every FTP session closes (issue #3009)
disconnect() and _abandon_connection() logged nothing, at any level. A
    session closed cleanly and a socket genuinely abandoned therefore produced
    identical output -- none -- and the only way to tell them apart was to read
    the source.

    That is how #3009 was filed. Its trace shows a print completion opening two
    FTP connections, deleting one file, and then nothing until the printer was
    powered off 21 minutes later, read as connections left open and offered as a
    mechanism for the 0500-C010 SD-card error that #645 has been chasing since
    April. The two connections are the post-print SD cleanup in main.py walking
    its candidate filenames, each through delete_file_async, which closes in a
    finally; running that against the mock FTPS server shows the server logging
    "FTP session closed (disconnect)" for both the 250 and the 550, holding zero
    sessions afterwards. Nothing in a support bundle could have shown that.

    Both close paths now log one DEBUG line: the printer, whether QUIT was
    acknowledged or the socket had to be dropped without it, why, and how long
    the session was held. Every connect in a debug log now has a matching close.

    The duration comes from a stamp taken when the control socket opens rather
    than after login, so a session that dies during login is accounted for too;
    where no socket was ever established the line says "held unknown" rather
    than claiming a number. The four connect() failure paths pass their own
    reason, so a close line stands on its own next to the warning above it.

    Nine tests, seven of which fail against the unlogged version. The other two
    assert silence -- a bare disconnect(), and a connect skipped by the handshake
    cool-off -- where no socket was opened and a close line would pair with no
    connect.
2026-09-20 13:26:42 +02:00
maziggy 8186ef817e fix(orca-cloud): close the HTTP client when an authenticated build fails
OrcaCloudService owns an httpx client from construction, and every path in
    _build_authenticated_service after that point can raise: no stored refresh
    token, a rejected refresh, an unreachable Orca, and the token-rotation write.
    On success the caller closes the client. On failure nobody is ever handed it,
    so all four paths leaked one into the connection pool.

    That went unnoticed while the only callers were routes, where the trigger is a
    person retrying a broken sign-in a handful of times. It stopped being harmless
    in 9434875f, which added a caller in spool assignment -- one build per
    Orca-referenced spool, failing on every assignment for as long as the stored
    credentials cannot be refreshed.

    The unwind guard catches BaseException rather than Exception: a cancelled
    request leaks the client just as surely as a failed refresh, and cancellation
    during shutdown is exactly when dangling sockets are least welcome. The close
    inside it is guarded in turn, so a failing cleanup cannot replace the error the
    caller needs to see -- least of all a CancelledError, which has to keep
    propagating for cancellation to work at all.

    Six tests. Four fail against the unguarded builder, verified by reverting the
    guard and re-running; the other two pin the surrounding contract (a failing
    close must not mask the real error, and a successful build must leave the
    client open for its caller) and pass either way. The shared _expired_service
    helper now gives the mock an awaitable close(), so the four pre-existing
    refresh tests exercise the same path.

    Also corrects two comments and the changelog entry from 9434875f, which
    overstated what the captures support. They claimed Bambu Cloud returns a
    preset's filament_id in either of two places and only one was read. The
    responses recorded in #1053 show something narrower: a Studio-created preset
    carries it on the envelope, and an Orca-created one has none at all -- the
    envelope says null and `setting` is a delta from the base. The `setting` lookup
    stays as belt-and-braces for a shape no captured response has needed yet, but
    it is not why a custom profile reached the slicer as its base. That is the
    OrcaSlicer preset format having no filament_id field, filed upstream as
    OrcaSlicer PR #13315.

    The eight-character truncation is now evidenced across three models rather than
    one -- an A1 storing PFUS9DDC of PFUS9DDC938FE3AB8F, a P1S storing PFUS7A65 of
    PFUS7A65290D3DADC4, and an H2D storing 8219C45D of an Orca profile UUID.
2026-09-20 13:26:02 +02:00
maziggy a3e6fae5cb fix(ams): resolve a custom filament's own id from every preset source (issue #3003)
A custom filament profile reaches an AMS slot as itself through exactly one
    field, tray_info_idx, and every source we can read that id from was reading it
    from the wrong place or not reading it at all.

    Bambu Cloud returns a preset's own filament_id either on the response envelope
    or inside the preset JSON under `setting`, and only the envelope was read.
    Presets of the second shape fell through to the base_id branch and reached the
    slicer as the Bambu filament they inherit from. filament_type next door already
    handled both spreads; filament_id now does too.

    Orca Cloud was absent from the resolver entirely. A spool stores the bare
    profile UUID, which matched no branch and fell through normalize_slicer_filament
    -- a function that passes anything it does not recognise straight through -- so
    a 36-character UUID went into the field. Orca profiles carry their own
    filament_id in the slicer JSON that OrcaProfileDetail already exposes under
    `setting`, so the lookup is the same one the Bambu branch does. It is
    best-effort: no pairing, a dead token or a missing orca_cloud:auth permission
    degrades to the fallback rather than failing the assignment, and it passes
    clear_on_auth_failure=False because a background caller cannot tell a real
    revocation from a lost refresh-rotation race.

    configure_ams_slot sent the cloud setting_id as tray_info_idx when it found no
    real filament id. That field is 8 characters on the printer -- exactly the width
    of a local preset id, less than half a cloud one. Measured on the reporter's A1:
    sent PFUS9ddc938fe3ab8f, the tray read back PFUS9DDC, acknowledged as a success.
    The slot then resolved to nothing, so the slicer showed Generic anyway and the
    calibration table, keyed by the same field, lost the slot. It now falls back to
    the slot's existing filament id or the generic for the material, and the route's
    guard was aligned with the resolver's so both refuse the same four shapes from
    one shared definition.

    This reverses the contract #1053 pinned. Six tests asserted that the PFUS
    belonged in tray_info_idx; the A1 capture shows it never worked, so they were
    rewritten with the measurement in their docstrings.

    Verified against 874 AMS trays across twelve models in the support archive: 92
    already carry a custom "P" + 7 hex filament id, which is what confirms the
    mechanism works and this is a lookup failure rather than a platform limit. No
    tray on any model carries a setting_id, so a profile with no filament_id of its
    own still cannot be told apart from its base.
2026-09-20 13:25:27 +02:00
maziggy 7490c93081 ci: balance the backend test shards by measured time, not test count
Backend Tests (shard 1/4) timed out after 10 minutes on e8b901f54
    ("Updated BACKERS"), a docs-only commit. The step was not hung: it
    reached 98%, every test passing, and was killed about 12 seconds short
    of finishing.

    pytest-split balances by duration only when it has a durations file,
    and there was none. Without one it splits by test COUNT -- 2926 / 2926
    / 2926 / 2925, exactly 25% each, which is what the comment here claimed
    was good enough. Count is not time. Measured over the full suite
    (11703 tests, 1045.6s):

        shard 1   2926 tests   658.2s
        shard 2   2926 tests   173.1s
        shard 3   2926 tests   113.3s
        shard 4   2925 tests   101.1s

    Shard 1 was carrying 63% of the suite's runtime -- a 6.5x spread -- and
    had been walking toward the cap for weeks: 475s, 404s, 445s, then 587s
    on 08-30, thirteen seconds under it, and 613s here. CI printed the
    reason in its own log on every run: "[pytest-split] No test durations
    found."

    Commit tests/.test_durations and pass --durations-path explicitly:

        shard 1   1273 tests   261.6s
        shard 2   1070 tests   261.5s
        shard 3   1522 tests   262.0s
        shard 4   7838 tests   260.6s

    A 1.01x spread. The lopsided test counts are the point: shard 4 takes
    thousands of fast unit tests, shard 1 keeps the slow integration ones.
    The four groups still partition the suite exactly -- union is 11703
    with nothing dropped or duplicated, and all four pass.

    --durations-path has to be explicit because pytest-split defaults it to
    $CWD/.test_durations, and the two matrices run from different
    directories: the native job from backend/, the Docker job from the
    image root. Left implicit, the Docker one finds nothing and silently
    falls back to the count split. Node IDs match across both because
    backend/tests/pytest.ini pins rootdir to backend/tests either way, so a
    single file serves them both; verified the file clears .dockerignore
    and lands at /app/backend/tests/.test_durations at full size.

    timeout-minutes 10 -> 15 for headroom, since the file goes stale as
    tests are added. Staleness degrades slowly rather than breaking --
    unknown tests are treated as average.
2026-09-20 13:22:17 +02:00
maziggy 5584dca898 Bumped version 2026-09-20 13:19:02 +02:00
maziggy 2c7c97c130 fix(jog): send the nozzle-bed gap the API promises on every model (issue #1334)
POST /printers/{id}/bed-jog takes a signed nozzle-bed gap, documented since it
was written: positive asks for more room between the nozzle and the plate. On
an A1 it did the opposite. The reporter sent distance=5 for clearance and
watched the toolhead come down.

The sign had been flipped on A1 models since the original report on this issue,
where an A1 Mini owner clicked an arrow labelled "move the plate up" and watched
the nozzle dive. That is a labelling problem -- a bed-slinger's plate does not
move in Z at all, so closing the gap shows up as the toolhead descending -- and
it was solved in the transport layer, which turned a parameter documented as
model-independent into one that meant the opposite thing on part of the fleet.

Z is the nozzle-to-bed distance on every Bambu model, by definition of the
coordinate system rather than by convention: G1 Z+ opens the gap whether the bed
drops away from a fixed nozzle (X1/P1/H2, whose end G-code parks with
G1 Z{max_layer_z + 100}) or the nozzle rises off a fixed bed (A1/A2L). The
finish-photo plate restore already relies on exactly that and carries no model
branch. So distance goes onto the wire unchanged and one call means one physical
outcome everywhere: positive is the safe direction on every printer.

Which way an arrow points is a different question, about the machine in front of
the user rather than about G-code, so the printer card answers it and asks for
the gap it wants. The buttons move what you would expect them to move, exactly
as before; on a bed-slinger they now say toolhead rather than plate.

The A2L never had the old fix. It slings its bed the same way the A1 does, but
the inversion listed the A1 names and the A2L was not among them, so its up
arrow has been sending the toolhead at the plate for as long as the machine has
been supported. The new classifier also covers the alternate internal codes
A04 / A11 / A12, which LINEAR_RAIL_MODELS and SINGLE_NOZZLE_FLOW_MODELS both
carry and the old gate did not.

is_bed_slinger is gone from the backend rather than widened: with the route
model-independent it had no caller, and a kinematics helper sitting unused in
the service layer invites the next person to assume the backend handles
direction. It does not, deliberately.

Separately, the soft-endstop comments on both jog routes claimed the firmware
clamps a bare move at the travel limit. It does not, and #2579 measured that:
an H2D at its Z limit ran straight past a clean G91/G1 Z-1.00/G90, while its own
touchscreen refuses the identical move. What #2579 removed was M211 S0, which
disabled the limits globally and took the touchscreen's protection with them.
The jog popover has warned about this correctly the whole time; only the code
comments disagreed with it.
2026-09-19 10:38:54 +02:00
maziggy 48244bcfcf fix(queue): tell a pinned queue item why it is waiting (issue #3074)
A job queued as "Any X1C" explains itself when it cannot start: the
model-based branch builds a reason for every candidate printer and puts it
on the row, so the queue shows "Busy: X1C-01" or "Waiting for filament:
X1C-02 (needs PETG)". The same job pinned to one printer showed nothing.
It sat at Pending with waiting_reason NULL for as long as that printer was
busy, which from the outside is indistinguishable from a queue that has
stopped working -- the reporter watched fourteen minutes of it while his
X1C ran a print he had started from its own screen.

The fixed-printer branch had six ways out and none of them wrote the field.
The sensor interlock (#1148) was its only writer, and it cleared the field
up front on every pass where no sensor was holding the printer, so NULL was
not an oversight on those paths but a guarantee.

Every exit now writes, through one helper. The reasons reuse the
model-based branch's vocabulary so _is_busy_only() keeps deciding what is
worth a notification: a printer that is printing, drying, or working
through the item ahead of this one reads as "Busy: <printer>" and stays
silent, because it resolves itself. A printer that is off with no Auto On
plug, and one whose plug could not switch it on, are worth saying.

A finished plate nobody has acknowledged is split out from plain busy and
named as itself. _is_printer_idle() returns the same plain False for that
and for a running print, but they are not the same thing to the person
looking at the queue: one clears itself and the other needs somebody to
walk over to the printer.

That notification fires on the transition into asking, where a busy-only
reason counts as not asking. Testing whether the item was waiting at all --
which is what the model-based branch does -- would never fire it here:
nobody's queue goes straight from idle to an unconfirmed plate, it waits
behind the print first. The cost is that a printer dropping offline,
returning busy and dropping again asks twice rather than once.

The interlock stays silent. It has never sent this notification, and a
change about what the queue displays is not the place to start.

Clearing the field up front is gone with it. It existed so a shut door
could not leave "Waiting on Enclosure Door" standing while the printer
stayed busy with something else, and the new rule carries that guarantee
instead -- whichever exit runs next overwrites it, and the dispatch path
clears it.

Two paths clear it that the report did not mention. A staged item and a
future-scheduled one skip before this branch and never reach it again, so
anything written on an earlier pass would outlive its condition for the
life of the row. That includes the filament-deficit check, which stages the
item itself.

The notification is wrapped: a queue that cannot say why it is waiting is
the bug being fixed, and a queue that stops dispatching because a provider
timed out would be a worse one.

On the frontend, the queue timeline drops any pending item carrying a
reason, on the grounds that such an item will not auto-dispatch. That held
while only the model-based branch wrote the field; "Busy: <printer>" is
now the commonest reason there is, and it describes the very chain the
timeline forecasts, so the rule would have emptied the view for anyone
whose queue is pinned. It now asks whether the reason needs the user, via
a small shared reader of the same shape the scheduler encodes.

Which job goes out, and when, is unchanged: running the previous scheduler
and this one over the same 768 states dispatches the same items in the same
order with the same statuses, across 1452 rows that now carry a reason.
2026-09-18 19:38:24 +02:00
maziggy 5be18ff57a fix(queue): print a plate whose filaments are all on the external spool (issue #3087)
The reporter's P1S heated up, sat at Heatbed preheating for ten and a half
minutes, then paused with 07FF_8012, "Failed to get AMS mapping table".
Resuming only reheated it. Prints that fed from the AMS were fine.

The plate was one filament of a seven-filament MakerWorld project, mapped by
hand to the external spool. slice_info.config numbers filaments across the
whole project, so the mapping for that plate is [-1,-1,-1,-1,-1,-1,254]: six
placeholders and the spool holder. The command builder decides whether a print
needs the AMS by asking whether the mapping is entirely external, and six -1s
answer no. So the print went out as use_ams=true carrying a flat mapping of
nothing but -1 -- 254 is deliberately never sent raw, the firmware reads it as
AMS tray 0 -- which is exactly the mapping table the firmware then could not
find.

The builder cannot fix this itself. Down there a -1 is either padding for a
filament this plate does not print, which is BambuStudio's own convention, or a
slot that never resolved to a tray, and sending the second one to the spool
holder is what #2589 exists to prevent. They are the same byte.

The scheduler knows. extract_filament_requirements drops every filament with
used_g <= 0, so it names precisely the slots the plate prints. When all of those
are an explicit 254/255, dispatch now sends use_ams=false and the print runs.
When one of them resolved to nothing, the flag is left alone and the firmware
rejects the print as it does today -- deliberately, because that is the case
where guessing would print a filament in the wrong material without saying so.

Single-nozzle only, mirroring the reconcile in the command builder: on a
two-extruder printer use_ams selects which nozzle to feed rather than whether to
use the AMS, so an H2D with a spool on each side must keep the flag it was
given. Judged generously from the model name and from live telemetry -- a second
nozzle reporting a diameter, an extruder map, or more than one external feed --
because a wrong yes only preserves existing behaviour while a wrong no would
reroute the print. H2S stays single-nozzle (#1386).

Nothing in the command builder changed. Its own reconcile keeps the exact
semantics #2589, #2595 and #797 gave it, and now usually agrees with a decision
that was already made one layer up. Where the file has no parseable filament
list the mapping is left exactly as before, the same evidence-only convention
as #2771, and the parse itself sits behind a check for anything external at all
so an AMS-only print never opens the file.

Covered end to end at the dispatcher, including the reporter's seven-filament
shape, a plate mixing the spool holder with an AMS tray, a consumed slot that
never resolved, both external feeds on a dual-nozzle machine, and a 3MF with no
filament list at all.
2026-09-18 17:55:06 +02:00
maziggy 8e9821dd2a fix(mqtt): never wait for a wedged paho network thread (issue #3068)
The reporter's A1 had been offline 38 hours and still answered on 8883, so
the connection watchdog did exactly what it exists for: rebuild the session
with a fresh client, since anything left in the old one's QoS 1 queue would
otherwise replay onto the next print (#1136). The rebuild ended in paho's
loop_stop(), which sets a terminate flag and then joins the network thread
with no timeout.

That thread only reads the flag between iterations of loop_forever, so it
cannot read it while parked inside reconnect() -> _ssl_wrap_socket() ->
do_handshake(). paho gives that handshake the keepalive as its socket
timeout -- 30s here -- and a socket timeout is per operation, renewed by
every byte the peer sends. A printer that answers TCP and then trickles
holds the join open for as long as it likes.

The join ran on the asyncio thread. Bambuddy stopped answering anything --
UI, API, /health -- while the process stayed up, which is why a
restart: unless-stopped container never restarted.

Retiring a client no longer waits for it. The replacement is built at once
and the old one is shut down on a thread of its own that nobody joins. Its
callbacks are detached first, inline: blocking until the network thread was
gone is what used to guarantee a client we had let go of could no longer
touch our state, and with the teardown detached a zombie that finishes its
handshake would otherwise auto-reconnect and report itself connected behind
its replacement's back. disconnect() still goes out, still promptly, because
that is what stops paho's auto-reconnect and the replay with it.

The reported watchdog is one of six callers. The queue's dispatch recovery
and check_staleness -- reached from an ordinary status poll -- share
_hard_reset_client; editing, deleting and hand-disconnecting a printer share
disconnect(); the relay and smart-plug services had the same join on their
shutdown path, where a wedged broker stopped the process from exiting at
all. #1445 was this join too, from the add-printer probe, and its off-loop
teardown stays as it is.

disconnect() stays quiet on the way out, as it always effectively did.
paho's callback used to land during the join, but it suppresses itself for a
clean disconnect of a printer that reported in the last ten seconds, so a
healthy printer disconnected by hand never announced itself offline.
Announcing it now would tell the user their printer had gone offline a
minute after they disconnected it on purpose (#1752).

A retirement that takes more than five seconds logs which printer it was.
The whole point is that the next one of these should not have to be
diagnosed from a thread dump.
2026-09-18 17:13:26 +02:00
maziggy 4440738904 fix(drying): dry a composite spool as its base material (issue #3067)
The reporter's AMS-HT would not auto-dry PA6-CF, and drying the same spool by
hand worked. The scheduler reduced a tray to a preset key by splitting on spaces
only, so "PA6-CF" stayed "PA6-CF", matched none of the eight rows the preset
table has, and the tray was read as holding nothing worth drying. The AMS was
then passed over on every scheduler sweep, silently, because every caller reads
"no row" as "nothing to do for this tray".

It was never only nylon. Of the 41 types a printer can report, 33 had no row
under that rule, and 20 of those have a base material sitting right there: every
-CF, -GF and -AERO variant of PLA, PETG, ABS, ASA, PC and PA.

Doing it by hand worked because the drying popover has resolved composites since
exact key first, so a row the user added for the exact type still wins, then the
suffix, then an alias map reading PA6, PA11, PA12, PAHT, PPA and Nylon as PA.

PPA is the one alias that is a judgement rather than a spelling. Polyphthalamide
is a distinct polymer, not a grade of nylon -- but it is an aromatic polyamide,
it takes up moisture the same way, and PA's row is the hottest the table has.

A material with no row and no alias is still skipped rather than dried at a
number nothing here can source. That is where this parts company with the
popover, which falls back to PLA because a dropdown has to show something.

The table is user-editable JSON, so a preset row can be present and empty. That
has always meant "skip this material" and still does: the temp and hours reads
fall back per field to 55C/12h, which would dry a PLA spool at 55 degrees.

Two more callers had the same line and move with it: per-filament humidity
thresholds, where the override set for a material never applied to that
material's composites, and the chamber preheat target, which had the suffix half
of this from #2902 but not the aliases.

The preheat test for that half reimplemented the lookup inline rather than
calling it, so it would have passed whatever the function did. It calls it now.
2026-09-18 14:01:36 +02:00
maziggy a4cfbd4212 fix(archives): come back for a 3MF whose transfer ran out of time (issue #3063)
The reporter's P1S had the sliced file on its card and was serving it. The 19MB
transfer just did not finish inside the budget while the printer was also running
its camera, its status messages and the upload of the job itself. Bambuddy wrote
an empty fallback archive and never looked again -- then downloaded that same file
successfully three times over the next two minutes and discarded every copy,
because the only code that would have attached one had already run.

The recovery machinery was there. It was armed for exactly one give-up, the FTPS
cool-off, on the grounds that the three storage verdicts are settled: a job on
internal eMMC never appears at any FTPS path, and sweeping for it again is what
where the file is demonstrably still on the card.

The sweep already had the signal and never used it. A file that is genuinely not
there is answered with 550, which surfaces as FileNotOnPrinterError and is caught
by name; a timeout returns falsy instead. So "the printer says no such file" and
"we never got a straight answer" are distinguishable without guessing, and only
the second schedules anything.

Not scheduled either for a 3MF that downloaded fine and turned out to be another
plate's. Recovery checks that a candidate is a readable 3MF but not which plate it
holds, and the names a retry would use are the same stale ones that fetched the
contradicted file -- so it would put back exactly what #2957 discards.

The ladder follows the cause: a cool-off has to expire, so its first attempt sits
past the 300s; nothing has to expire here, and this reporter's file completed 48
seconds after the budget was spent.

The archives banner gets its own wording for this, because the old text sends an
owner whose card is working to switch on a setting that is already on. It names
the Connection Timeout setting instead.
2026-09-18 13:38:02 +02:00
maziggy 30e530a881 fix(archives): let Items Printed go to 0 for a ruined plate (issue #3051)
A jam can destroy everything on the plate while the printer still reports the
job as a success, so the honest count of usable parts is zero. The edit dialog
floored the field at one, and a project's completed-items count sums that
column, so there was no way to record that a job produced nothing.

The floor was in the dialog only; the API stored whatever it was given, which
also meant a negative count was accepted and would have subtracted from the
project totals. The column is now bounded at zero instead.

Filament Trends counted prints as `quantity || 1`, which would have read a
deliberate 0 as "unset" and charged the ruined plate as one print while the
project page counted none.
2026-09-18 13:07:00 +02:00
maziggy 0b830ac35c fix(ams): stop reading the printer's command acks as status (issue #3040)
Every project_file carried "cfg": "0" — the device-config bitmask, which
Bambu Studio has never sent and the firmware ignores. The printer echoes a
command's fields back in its ack, and the ack was ingested as telemetry, so
bit 18 read as "AMS Filament Backup off" 25 ms after every dispatch.

Families that repeat cfg in their periodic status (P2S, H2C, X2D) corrected
themselves a second later; the P1S, A1, A1 Mini and A2L send it only in a
full status dump, so the wrong value stuck and silently disabled the
prefer-lowest-remaining gate. The A1 family, which reports no cfg at all and
is meant to stay "unknown", was pinned to a definite "off".

Acks are no longer read as status, for the backup bit or the per-job
timelapse flag they also echo, and cfg is gone from the print command.
2026-09-16 09:52:24 +02:00
maziggy 309e64b8a2 fix(ams): resolve a slot's K profile by index when the printer does not file per hotend (issue #3044)
An X2D with two AMS 2 Pro, one per hotend, showed a K value on every slot
of the first and nothing on any slot of the second. Configure Slot was
worse than blank there: the picker offered no matching profile, the slot
read as though nothing were bound, and choosing one changed nothing the
user could see. Both symptoms are one rule.

A calibration index can mean two different profiles on a dual-nozzle
machine -- on the maintainer's H2C, index 16 is the left hotend's black
PLA at K=0.018 and 15 is the right's at K=0.020 -- so the index is
resolved against the slot's own hotend, and a miss shows nothing rather
than the other nozzle's number. That is right whenever the printer files
its calibrations per hotend. This one files them per filament: the second
AMS's slots point at the same entries as the first, every entry tagged
with one extruder, and requiring a match found nothing at all.

The hotend now has to appear in the table the printer actually sent
before it is used to narrow anything. Where it does not, the index stands
on its own, which is what BambuStudio does for this same card --
AMSItem.cpp resolves it through get_pa_k_n_value_by_cali_idx, matching
cali_idx and nothing else. Where it does, nothing changes: the H2C case
still blanks rather than borrowing, and the other hotend's profiles stay
reachable under Other K profiles. The relaxed path still refuses an
answer when the candidates disagree on a value.

The premise that the table is always numbered per nozzle had been written
into three comments and two layers of code; it is corrected where it
appears.

Alongside it, in the same picker: the K-profile options rendered the
hotend suffix twice in the matching group and three times under Other, so
every option on a dual-nozzle printer read "... . Left . Left".
2026-09-07 19:23:48 +02:00
maziggy 069ee8fc87 fix(queue): skip preheat entirely when no loaded filament wants a chamber (issue #3041)
Preheat & Heat Soak delayed every PLA print by five to seven minutes and
gave nothing back. The filament map correctly derived a chamber target of
0, and the chamber phase correctly skipped -- but the stage then heated
the bed, waited for it, and held the full soak anyway, because the soak
had no idea it was holding for a chamber nobody asked for. The print's
own G-code sets the bed the moment it starts, so the bed phase only moved
the warm-up ahead of the FTP upload instead of overlapping with it.

A 0 that comes out of the filament map now skips the stage before any
command goes out. The one thing the skip still does is put the airduct
flap back to cooling on the models that have one -- an H2D left in
heating mode by the ABS job before it would otherwise cook the PLA that
follows, and that costs one MQTT command and no waiting.

Explicit instructions are untouched. A chamber target of 0 typed into a
print's own override still heats the bed and runs the soak, which is what
the queue documentation has always promised it does, as does forcing a
print's Preheat override to On. Prints that want chamber heat are
unaffected, including the P1S/P1P/A1 tier where the bed and the soak
timer are the whole mechanism.

The existing unit tests all ran with soak_seconds=0, which is why the
production default was never exercised; the PLA test now runs at the real
default and asserts nothing is dispatched and nothing is slept.

Surfaced in the UI on the way through: the Settings hint claimed the
derived 0 skipped "the chamber phase", and the per-print chamber override
field said nothing about a typed 0 meaning bed-only -- a user reaching
for 0 to turn preheat off got the delay instead.
2026-09-07 16:38:36 +02:00
maziggy 9a837d19a1 fix(slicer): strip zero-valued filament-index sentinels, and sanitise the preview slice too (issue #3030)
Bambu Studio writes 0 into wall_filament, sparse_infill_filament and
solid_infill_filament to mean "use whichever filament the object is set
to". Bambu Studio and OrcaSlicer 2.4 define these min 0 and accept it;
OrcaSlicer 2.3 and earlier used the 1-based scheme (min 1, default 1)
and reject it with "0 not in range [1.000000,...]". Sidecar images are
version-tagged, so an install can be pinned to one of those builds.

Same shape as the -1 inherit markers from #1201 with a different marker,
so the allowlist becomes a key-to-marker map rather than one global
constant. The buckets must not bleed: a -1 on a filament index is a real
value, and a 0 on a raft field is a setting the user chose.

The key is removed rather than rewritten, which is what makes it safe on
every build. The CLI then uses its own default: 0 where 0 was legal
(unchanged), 1 on the older builds, which is what "the active filament"
means under that scheme.

The preview slice never ran the sanitiser at all, so a file that sliced
fine could still fail its automatic plate preview and fall back to the
painted-face heuristic. It matters more there than in a real slice: the
preview runs on the file's own embedded settings, so there is no
--load-settings pass that could supply a replacement for a field the
range validator has already rejected. That also explains the reported
"same error on a later attempt of an unchanged file" without any second
copy of the keys -- the validator that emits it reads the merged global
config, which per-object model_settings.config overrides never reach.

The sanitiser moves to utils/threemf_tools so the service can use it
without importing a route module, and both preview callers pick it up
from one place. Drops _strip_3mf_embedded_settings and its constant,
which have had no callers since the strip-everything experiment was
reverted.
2026-09-07 14:37:49 +02:00
maziggy b9bd312826 fix(slicer): keep protocol-handler download tokens valid for their whole TTL (issue #3029)
The Slice and Open in Slicer actions mint a short-lived token and put it in
the URL, because a protocol handler cannot carry an Authorization header.
That token was spent by the first request to reach the endpoint, which made
the handoff depend on the slicer fetching the URL exactly once. Nothing
guarantees that: Bambu Studio's downloader retries three times after a
failed attempt, transfers get resumed, on-access scanners fetch. The first
request won and the slicer was handed a 403.

verify_slicer_download_token takes a keyword-only single_use flag. The
default still consumes via DELETE...RETURNING; single_use=False verifies
with a SELECT and leaves the row for the rest of its five-minute TTL. The
stored row is the same either way, so the endpoint decides, not the mint.

The three protocol-handler downloads pass single_use=False: a library file,
an archive's sliced 3MF, an archive's source 3MF. Resource binding and
expiry are untouched. The two browser downloads keep consuming, because
what they hand over is itself consumed -- the prepared printer bundle is
deleted the moment it has been streamed.

Also: add "/source-dl/" to PUBLIC_API_PATTERNS. Those patterns match by
substring and the source 3MF route's segment is source-dl, which does not
contain "/dl/", so with auth enabled the middleware rejected the slicer's
header-less request before the route's token check ran. Open source 3MF in
slicer could never work on an install with authentication on.
2026-09-07 14:05:25 +02:00
maziggy 816f073a9e fix(auth): decouple media routes from the camera stream token (issue #3025)
Thirteen routes with nothing to do with a camera took the camera stream
token as their credential -- library and archive thumbnails, plate
previews and plate thumbnails, timelapses, print photos, archive QR
codes, project covers, print-log thumbnails, printer covers and
external-link icons. A browser cannot put an Authorization header on an
<img src>, so these need a credential that fits in the URL, and the
camera token was the only one that existed. Minting one costs
camera:view, so a user granted library access to their own files got a
grid of broken images until they were also handed the live camera.

Adds a media token: minted by POST /auth/media-token behind plain
authentication, and identified -- it records the principal the way the
websocket token does rather than being anonymous the way the camera
token is. Each route now gates on the permission and ownership rules of
the resource it serves, through the same _ensure_*_visible helpers its
header-authenticated siblings already use. The three camera routes keep
the camera token, and require_camera_stream_token_if_auth_enabled now
documents that it is for those only.

The media dependencies accept ordinary Authorization / X-API-Key headers
as well as ?token=, delegating that path to the existing checkers, so
API-key scope rules and the per-printer allowlist are unchanged.

Long-lived camera_stream, camwall and overlay tokens are deliberately
not accepted on the media routes -- those are handed to kiosks, walls
and Home Assistant to display video. The cam wall, streaming overlay and
kiosk views use only the three camera routes and are unaffected.

Frontend: withMediaToken alongside withStreamToken, and
useStreamTokenSync fetches a media token for every signed-in user while
asking for a camera token only when the user can mint one, which also
stops the 403 that fired on every page load for everyone else.

Also fixed, same class:
- /printers/{id}/files/plate-thumbnail/{i} is rendered in an <img> but
  had a header-only guard, so the file manager's plate thumbnails 401'd
  whenever auth was enabled. It now takes a media token too.
- getProjectCoverImageUrl returned a URL ending in ?token=, and the
  project edit dialog appended its own ?v= cache-buster after it, so the
  second ? landed inside the token value. The version is now a parameter
  applied before the token.

Tests: 15 integration tests for the token boundary, permission
enforcement and per-row scoping; 10 frontend tests for the URL split and
the two-query hook. test_cover_image_get_uses_stream_token_gate is
renamed and repointed at the media gate -- what it pins, that the
credential has to fit in a URL, is unchanged.
2026-09-07 13:38:15 +02:00
maziggy 93eeb05264 fix(auth): let the sidebar read install flags without settings:read (issue #3023)
cost_centers:read_own exists so a non-admin can see their own wallet, balance
and cost-centre spend, and the Finance page honoured it -- typing the URL
worked and rendered their balance. The sidebar never offered the entry.

It decides whether to show Finance by reading billing_enabled from
GET /settings, which requires SETTINGS_READ. A non-admin gets 403 there, so
the value arrived undefined, `undefined !== true` held, and the entry was
hidden from precisely the users the permission was written for. The permission
map and the route guard were both already right; only discovery was broken.

Three more fields came from that same 403, and one of them failed the other way
up. The Notifications gate tests `=== false`, which undefined never satisfies,
so an administrator who switched user notifications off still left the entry
showing to the non-admins it governs. Nobody reported that one, and no
administrator could have reproduced either: administrators can read /settings.
The remaining two were quieter -- the sponsor prompt fell back to EUR whatever
the install uses, and the update check ran where it had been turned off.

SETTINGS_READ cannot be the price of knowing whether billing is on. It also
grants sight of the SMTP, LDAP and MQTT credentials, which is the reason
/settings/ui-preferences exists at all.

So: a second endpoint, GET /settings/ui-flags, carrying those four fields and
asking only that the caller be signed in, via the existing
require_auth_if_enabled. Layout drops its /settings query altogether, which
closes the class rather than the two instances that happened to be visible.

Deliberately not four more fields on /ui-preferences. That endpoint is served
to anyone at all on the recorded grounds that its contents are "public defaults
that ship with the app" (test_route_auth_coverage.py), and its field set is
pinned by a test written to make anyone adding to it stop and think. These
fields are not defaults -- they say how this deployment is configured -- so
they get their own endpoint at their own trust level instead of stretching that
charter to fit them. require_auth_if_enabled also keeps the auth-disabled case
that /ui-preferences was ungated for: "works when there is no auth" and
"readable by anyone" are different statements, and conflating them is what put
a settings read in front of a permission that never needed one.

Twelve tests. Backend pins that the operator can read the flags, that the same
operator still gets 403 from /settings, that an anonymous caller is refused
when auth is on, that it answers when auth is off, the exact field set, that no
credential ever appears, and that the public endpoint did not quietly gain
these fields. Frontend pins Finance visible for cost_centers:read_own with
/settings returning 403, and Notifications hidden when the flag is off -- each
waiting on a positive signal before asserting an absence, so the negative cases
cannot pass before the query resolves.

Reported by @lonix, who traced it to the queryKey and the route gate.
2026-09-07 12:48:31 +02:00
maziggy 09b4584d5f fix(queue): say when an unscheduled item runs instead of calling it ASAP (issue #3018)
The print dialog offers ASAP, Queue and Schedule. ASAP and Queue differ only
in where the item is inserted, and neither is stored on the item -- scheduleType
is a frontend-only concept, and grep finds no "asap" anywhere in the backend. So
the queue's time column had nothing to read but scheduled_time, and labelled
every unscheduled item "ASAP": the name of the one mode the user may well have
chosen against.

Someone who picked Queue then watched their row appear as ASAP and start
immediately, and concluded Bambuddy had overridden them. Two reporters wrote
that same sentence thirteen months apart, and #2557 was closed as A2L-specific
after the first of them -- kilrah replied there with an X1C before filing this.

The column answers when an item runs, so it now says that. The key is renamed
whenFree rather than just retranslated: left called asap, the next translator
puts ASAP back.

The dispatch is unchanged, because it was right. A print scheduled for later
does not reserve the printer until then; an unscheduled item behind it uses the
idle printer rather than leaving an X1C dark until 6 AM. Two of the new tests
pin that, so it does not get "fixed" later on the strength of a report like this
one.

What genuinely could not answer the question was the queue's own log. Its
per-printer line called every entry in busy_printers "not available" -- but that
set holds both printers that cannot take work and printers the pass has just
claimed for some, which are opposite facts. It also read printer state at
logging time rather than at the decision, so #3018's bundle carries

    Queue: printer 1 not available — connected=True, state=IDLE, ...
    Launching 1 upload(s) (pool 0/4 in flight)
    Starting queue item 18

a printer reported unavailable, evidence that it was available, and a dispatch
to it, in three consecutive lines. It is the first line anyone greps for "why
did my item not go out".

Each of the nine sites that removes a printer from a pass now records why, and
the summary reports a claim as a reservation and everything else as an
obstruction with its reason. The live fields stay, since a bundle reader wants
them next, but are labelled as read now rather than offered as the cause.
print_scheduler.py:1210 already documented that these two meanings differ -- the
dispatching_printers snapshot exists for it. This carries that distinction into
the log.
2026-09-07 12:22:26 +02:00
maziggy 6564c74071 fix(archives): report a refused FTPS handshake as the printer, not the slicer (issue #2780)
The Archives banner picks its wording from a priority list of the causes it
knows. REASON_FTPS_COOLOFF was added by #2957 and never put in that list, so an
install whose empty archives all came from a printer refusing the TLS handshake
matched nothing, got reason: null, and fell to the original wording: the slicer
did not leave the .gcode.3mf on the card, switch on "Store sent files on
external storage", here is installation step 4.

Every clause of that is wrong for this cause. The slicer did write the file --
reason he read the whole thing as Bambuddy being broken. The setting was
already on. And there is nothing on his side to change: the printer's file
service answered port 990 with something that is not TLS, so no lookup ever
ran and where the file went was never tested. It is #2899's mistake -- an
error message describing a cause that was ruled out before it was printed --
in a surface that did not get that pass.

The slug now leads the list rather than joining the end of it. The other three
describe an install working as configured and each ends in something the
operator can change; this one reports a fault nobody can yet explain, which is
both the more urgent thing to say and the thing that produces a useful report.
The banner also dismisses one-shot into localStorage, so a reason ranked below
another is not deferred to next time -- it is never shown to that user again.
Ranking it first cannot bury a permanent cause in exchange: a successful
recovery clears the row's markers (#2957), so a row still carrying this slug is
one whose retry failed too, days after the print.

New wording in all fourteen languages says the printer refused the connection,
that this is not a slicer setting and not something the operator did, that
Bambuddy comes back for the file when the five-minute pause clears so a brief
episode fills itself in, and that a card still empty means the refusal outlasted
the retry. It links to the handshake entry in the troubleshooting guide instead
of to the installation guide.

The client's getNo3MFWarning type still declared the old three-slug union, which
made all three new comparisons provably dead -- caught by tsc, not by any test.

Four tests. One pins the slug reaching the banner, one pins it outranking the
three settled causes, one pins those three keeping their order behind it, and
one asserts the rendered wording carries no slicer advice at all.

Also corrects the wiki page these reports are pointed at. It said to power-cycle
the printer; the reporter who prompted that advice power-cycled both of his and
the failure continued unchanged, and bambu_ftp.py has carried the retraction in
a comment since. The page now states what was actually measured -- that a
version mismatch reports itself differently, that every printer probed refuses
TLS 1.3 and completes on 1.2 so there is no version to fall back from, and that
three P2S units failed while three more on the same switch never did -- says
plainly that the trigger is unknown, and names the one cleartext-probe line
worth collecting.
2026-09-07 11:26:01 +02:00
maziggy 10f0900fbc feat(ftp): log how every FTP session closes (issue #3009)
disconnect() and _abandon_connection() logged nothing, at any level. A
session closed cleanly and a socket genuinely abandoned therefore produced
identical output -- none -- and the only way to tell them apart was to read
the source.

That is how #3009 was filed. Its trace shows a print completion opening two
FTP connections, deleting one file, and then nothing until the printer was
powered off 21 minutes later, read as connections left open and offered as a
mechanism for the 0500-C010 SD-card error that #645 has been chasing since
April. The two connections are the post-print SD cleanup in main.py walking
its candidate filenames, each through delete_file_async, which closes in a
finally; running that against the mock FTPS server shows the server logging
"FTP session closed (disconnect)" for both the 250 and the 550, holding zero
sessions afterwards. Nothing in a support bundle could have shown that.

Both close paths now log one DEBUG line: the printer, whether QUIT was
acknowledged or the socket had to be dropped without it, why, and how long
the session was held. Every connect in a debug log now has a matching close.

The duration comes from a stamp taken when the control socket opens rather
than after login, so a session that dies during login is accounted for too;
where no socket was ever established the line says "held unknown" rather
than claiming a number. The four connect() failure paths pass their own
reason, so a close line stands on its own next to the warning above it.

Nine tests, seven of which fail against the unlogged version. The other two
assert silence -- a bare disconnect(), and a connect skipped by the handshake
cool-off -- where no socket was opened and a close line would pair with no
connect.
2026-09-07 10:54:55 +02:00
maziggy f3b1c59169 fix(orca-cloud): close the HTTP client when an authenticated build fails
OrcaCloudService owns an httpx client from construction, and every path in
_build_authenticated_service after that point can raise: no stored refresh
token, a rejected refresh, an unreachable Orca, and the token-rotation write.
On success the caller closes the client. On failure nobody is ever handed it,
so all four paths leaked one into the connection pool.

That went unnoticed while the only callers were routes, where the trigger is a
person retrying a broken sign-in a handful of times. It stopped being harmless
in 9434875f, which added a caller in spool assignment -- one build per
Orca-referenced spool, failing on every assignment for as long as the stored
credentials cannot be refreshed.

The unwind guard catches BaseException rather than Exception: a cancelled
request leaks the client just as surely as a failed refresh, and cancellation
during shutdown is exactly when dangling sockets are least welcome. The close
inside it is guarded in turn, so a failing cleanup cannot replace the error the
caller needs to see -- least of all a CancelledError, which has to keep
propagating for cancellation to work at all.

Six tests. Four fail against the unguarded builder, verified by reverting the
guard and re-running; the other two pin the surrounding contract (a failing
close must not mask the real error, and a successful build must leave the
client open for its caller) and pass either way. The shared _expired_service
helper now gives the mock an awaitable close(), so the four pre-existing
refresh tests exercise the same path.

Also corrects two comments and the changelog entry from 9434875f, which
overstated what the captures support. They claimed Bambu Cloud returns a
preset's filament_id in either of two places and only one was read. The
responses recorded in #1053 show something narrower: a Studio-created preset
carries it on the envelope, and an Orca-created one has none at all -- the
envelope says null and `setting` is a delta from the base. The `setting` lookup
stays as belt-and-braces for a shape no captured response has needed yet, but
it is not why a custom profile reached the slicer as its base. That is the
OrcaSlicer preset format having no filament_id field, filed upstream as
OrcaSlicer PR #13315.

The eight-character truncation is now evidenced across three models rather than
one -- an A1 storing PFUS9DDC of PFUS9DDC938FE3AB8F, a P1S storing PFUS7A65 of
PFUS7A65290D3DADC4, and an H2D storing 8219C45D of an Orca profile UUID.
2026-09-07 10:39:10 +02:00
maziggy 9434875fa1 fix(ams): resolve a custom filament's own id from every preset source (issue #3003)
A custom filament profile reaches an AMS slot as itself through exactly one
field, tray_info_idx, and every source we can read that id from was reading it
from the wrong place or not reading it at all.

Bambu Cloud returns a preset's own filament_id either on the response envelope
or inside the preset JSON under `setting`, and only the envelope was read.
Presets of the second shape fell through to the base_id branch and reached the
slicer as the Bambu filament they inherit from. filament_type next door already
handled both spreads; filament_id now does too.

Orca Cloud was absent from the resolver entirely. A spool stores the bare
profile UUID, which matched no branch and fell through normalize_slicer_filament
-- a function that passes anything it does not recognise straight through -- so
a 36-character UUID went into the field. Orca profiles carry their own
filament_id in the slicer JSON that OrcaProfileDetail already exposes under
`setting`, so the lookup is the same one the Bambu branch does. It is
best-effort: no pairing, a dead token or a missing orca_cloud:auth permission
degrades to the fallback rather than failing the assignment, and it passes
clear_on_auth_failure=False because a background caller cannot tell a real
revocation from a lost refresh-rotation race.

configure_ams_slot sent the cloud setting_id as tray_info_idx when it found no
real filament id. That field is 8 characters on the printer -- exactly the width
of a local preset id, less than half a cloud one. Measured on the reporter's A1:
sent PFUS9ddc938fe3ab8f, the tray read back PFUS9DDC, acknowledged as a success.
The slot then resolved to nothing, so the slicer showed Generic anyway and the
calibration table, keyed by the same field, lost the slot. It now falls back to
the slot's existing filament id or the generic for the material, and the route's
guard was aligned with the resolver's so both refuse the same four shapes from
one shared definition.

This reverses the contract #1053 pinned. Six tests asserted that the PFUS
belonged in tray_info_idx; the A1 capture shows it never worked, so they were
rewritten with the measurement in their docstrings.

Verified against 874 AMS trays across twelve models in the support archive: 92
already carry a custom "P" + 7 hex filament id, which is what confirms the
mechanism works and this is a lookup failure rather than a platform limit. No
tray on any model carries a setting_id, so a profile with no filament_id of its
own still cannot be told apart from its base.
2026-09-07 10:15:06 +02:00
maziggy e0a2398783 ci: balance the backend test shards by measured time, not test count
Backend Tests (shard 1/4) timed out after 10 minutes on e8b901f54
("Updated BACKERS"), a docs-only commit. The step was not hung: it
reached 98%, every test passing, and was killed about 12 seconds short
of finishing.

pytest-split balances by duration only when it has a durations file,
and there was none. Without one it splits by test COUNT -- 2926 / 2926
/ 2926 / 2925, exactly 25% each, which is what the comment here claimed
was good enough. Count is not time. Measured over the full suite
(11703 tests, 1045.6s):

    shard 1   2926 tests   658.2s
    shard 2   2926 tests   173.1s
    shard 3   2926 tests   113.3s
    shard 4   2925 tests   101.1s

Shard 1 was carrying 63% of the suite's runtime -- a 6.5x spread -- and
had been walking toward the cap for weeks: 475s, 404s, 445s, then 587s
on 08-30, thirteen seconds under it, and 613s here. CI printed the
reason in its own log on every run: "[pytest-split] No test durations
found."

Commit tests/.test_durations and pass --durations-path explicitly:

    shard 1   1273 tests   261.6s
    shard 2   1070 tests   261.5s
    shard 3   1522 tests   262.0s
    shard 4   7838 tests   260.6s

A 1.01x spread. The lopsided test counts are the point: shard 4 takes
thousands of fast unit tests, shard 1 keeps the slow integration ones.
The four groups still partition the suite exactly -- union is 11703
with nothing dropped or duplicated, and all four pass.

--durations-path has to be explicit because pytest-split defaults it to
$CWD/.test_durations, and the two matrices run from different
directories: the native job from backend/, the Docker job from the
image root. Left implicit, the Docker one finds nothing and silently
falls back to the count split. Node IDs match across both because
backend/tests/pytest.ini pins rootdir to backend/tests either way, so a
single file serves them both; verified the file clears .dockerignore
and lands at /app/backend/tests/.test_durations at full size.

timeout-minutes 10 -> 15 for headroom, since the file goes stale as
tests are added. Staleness degrades slowly rather than breaking --
unknown tests are treated as average.
2026-08-31 17:13:09 +02:00
maziggy 0e0bea1aa7 Bumped version 2026-08-30 08:56:26 +02:00
maziggy 0dfcff5925 Keep the RTSPS proxy's handler set off the server object (issue #3001)
asyncio's Server has a __dict__ and uvloop's, a Cython cdef class, does
not, so the attribute added in 1.2.5.4 raised AttributeError under uvloop.
Every RTSP camera failed before opening a socket, which is the
diagnostic's capture_exception at 0 ms.

Our own unit files all pin --loop asyncio for #1896 and were never
affected. The reports come from units we do not write: the Proxmox VE
Helper-Scripts LXC pins no loop, and installs predating that fix never
gained the flag because update.sh does not rewrite unit files. The loop
is not ours to assume, so fix the code rather than add another flag.

The set moves to a module-level WeakKeyDictionary, keyed weakly so an
abandoned proxy retires its own entry rather than leaking one and later
handing a new server a dead one's handlers.

Pinned on a real uvloop loop and, for hosts without uvloop, against a
__slots__ server; conftest builds its loop from the default policy, so
nothing in the suite had ever run the branch that broke.

Also routes the two external-camera teardowns through close_tls_proxy,
which #2968 introduced and left them out of.

-----

Say so at startup when running on uvloop (issue #3001)

An install on the wrong loop had no way to find out it was. #3001 was
loud enough to notice; the #1896 upload truncation it is also exposed to
is silent, and shows up as a print failing from a file that was corrupt
on arrival.

One WARNING in the lifespan naming the loop, the risk and the flag to
add. A warning and not a refusal: uvicorn has already chosen its loop by
the time any application code runs, and a server that answers requests
beats one that will not boot.

Asks the running loop what it is rather than whether uvloop imports --
uvicorn[standard] installs uvloop everywhere, so its presence says
nothing -- and matches on the module name so the question never imports
uvloop on a host without it.

-----

Repair a service file written before the --loop asyncio pin (issue #3001)

install.sh has pinned the loop since #1896, but nothing has ever
rewritten an existing service file, so every native install created
between 2025-11-28 (when uvicorn[standard] brought uvloop into the venv)
and 2026-07-05 still runs on uvloop no matter how often it is updated.

Both update scripts now add the flag themselves while the service is
stopped, so it takes effect on the same restart -- systemd via sed,
launchd via PlistBuddy, each backing the file up first and inserting
nothing but the flag.

Refuses to edit and explains instead when the shape is not a plain
single-line uvicorn unit: a wrapper script, a continued ExecStart,
several of them, a read-only file, or a service with drop-ins, since a
drop-in may be what defines ExecStart and editing the fragment would
change nothing while reporting success. A deliberate --loop uvloop is
left alone. Reads the effective ExecStart from systemd rather than the
file, so it is idempotent.
2026-08-30 08:01:42 +02:00
maziggy 3f1ed85791 Housekeeping 2026-08-29 16:57:04 +02:00
maziggy 49c948d01a Pin the TLS floor on the cleartext-probe test's context
CodeQL reports the context as allowing TLS 1.0 and 1.1, and it is right
about the mechanism: create_default_context() leaves minimum_version at
MINIMUM_SUPPORTED, which is the build's floor rather than a guarantee.
That is the reason every context in backend/app pins it, the reason
bambu_ftp.py carries a comment saying so, and the reason this same file
already pins it for the TLS-1.3 case further down. Line 127 was the one
that did not.

The floor cannot change what the test measures. The fixture answers with
a plain FTP banner and speaks no TLS, so the handshake still fails as
WRONG_VERSION_NUMBER, which is the assertion this test exists to make.
2026-08-29 15:29:59 +02:00
maziggy 9d35a4e8c6 Mark four test-side bandit false positives
The PR gate reported four new alerts, all in test code. A fixture's
/tmp/x.3mf is a column value the migration under test UPDATEs, not a
path anything opens. Two f-string statements interpolate column names
from a literal list declared two lines above, with the id bound - a
column name cannot be a bind parameter, which is why it is written into
the string at all. The two joins move onto their own lines because a
trailing marker would have taken line 64 past the 120-character limit.

The fourth is a near miss rather than a finding: the line below it
already carries the marker, as do four other wildcard sites in the same
file. The wildcard is what that test exists to assert about.
2026-08-29 15:25:26 +02:00
maziggy 0eb8d4b22f Mark the failure-reason migration's table name for bandit too
The line carried a noqa for ruff's S608 but nothing bandit reads, so the
same rule was silent in one tool and reported as a medium SQL-injection
finding in the other.

Nothing is interpolated but `table`, which the loop takes from a literal
tuple on the next line; the key and the label list are both bound
parameters. A table name cannot be one, which is why it is written into
the string at all.
2026-08-29 14:47:43 +02:00
maziggy 9755e08077 Decide whether a 3MF is sliced by looking inside it (issue #2993)
An archive that showed the green GCODE badge could re-import into the
    File Manager as a source-only project with no Print button, seemingly at
    random.

    Nothing was ever lost from the file. The download serves the stored
    bytes verbatim and the G-code was still in the zip; the two sides simply
    asked different questions. Archives looked inside the file. The library
    looked at the filename. So a sliced 3MF stored as Foo.3mf rather than
    Foo.gcode.3mf earned the badge and lost the Print button, and which one
    you got depended on how the print had reached the printer -- a slicer's
    LAN send names it .gcode.3mf, a per-plate export or a cloud-dispatched
    print does not.

    Both sides now ask one shared predicate about the zip itself, and every
    route into the library classifies on content. Only the central directory
    is read, and only when the name has not already settled it, so ingest
    costs nothing extra -- the external scan opens each 3MF for its
    thumbnail regardless. Rows already stored are re-checked once, internal
    ones only: an external row points at a mount that may be slow or absent,
    and startup is the worst place to discover that.

    The Slice action moves with it. Its refusal to slice an output was as
    name-bound as the Print gate, and without that a file that correctly
    gained a Print button would have offered to re-slice its own G-code.
2026-08-29 14:21:46 +02:00
maziggy f616a6bca9 Show the spool that is in the AMS slot, not the one that was
Pull a Bambu ABS Orange out of A1, put a PLA Matte Dark Blue in, and the
    slot card still read "Bambu ABS" against the new colour until the page
    was reloaded.

    Three things stood between the swap and a correct card.

    The RFID auto-assign rewrites the slot's slot_preset_mappings row and
    then broadcast an event that refreshed everything except the query that
    reads it. Only the manual assign path invalidated that one.

    Those queries then sat behind the 3s cascade debounce, which exists for
    print completion, where one event fans out across half the app. A swap
    touches one slot and the user is standing at the printer looking at the
    card; worse, the timer restarts on every further event, so a busy moment
    could defer it indefinitely. Slot changes now invalidate immediately.

    And the card trusted the stored preset over live telemetry outright.
    That priority is why a hand-picked preset name stays on a slot, but it
    also let a cached row outrank what the printer was reporting. The row is
    now skipped when it names a different official Bambu filament than the
    tray does, so the card is right from the status push alone. User and
    local presets carry ids that genuinely cannot be compared and are left
    exactly as they were.

    Spoolman mode was the worse half of the same bug: its AMS sync writes
    the same row but announced nothing at all, so there was no event to
    refresh on. It now reports each slot it changed or cleared.
2026-08-29 14:21:08 +02:00
maziggy 049afea980 Refuse a same-named 3MF that contradicts the running print (issue #2957)
When a print's own 3MF cannot be fetched the usage tracker borrows one from the
    library or a previous archive, matching on the filename stem. That is far weaker
    evidence than it looks: Bambu Studio writes the printer-side filename from the
    project's Title metadata, so every plate of a project reaches the printer under
    one name however the file was renamed on disk.

    The reporter's single-filament job was handed a previous archive's three-filament
    plate. Three spools were debited for material that was never extruded, and
    nothing on the archive said the numbers were someone else's.

    A candidate is now rejected when it positively contradicts the print - a
    different plate, or a filament count the slicer's ams_mapping disagrees with -
    and the accepted one is logged with both expectations. Only on a contradiction:
    the plate needs firmware that echoes it and the count needs a print command
    Bambuddy saw, and refusing everything uncorroborated would retire the fallback
    recovery this same issue asked for.

    The count is scoped to one plate or not compared at all. Unscoped, the filament
    reader collects every <filament> in the file, and that sum against one plate's
    count would reject every multi-plate library upload on exactly the firmwares
    that cannot tell us the plate.

    -----

    Give a download the time the file needs, and one printer at a time (issue #2957)

    ftp_timeout is handed to every download as both the socket inactivity timeout
    and the whole-transfer deadline, which makes its 30s default a cap on how big a
    file a printer may serve. The reporter measured one 5.4 MB 3MF at 45s off a worn
    P1S SD card and 25s off a new one, and a 15.15 MB 3MF at 105s. None of those
    links were broken - they were slow, which is what the inactivity timeout exists
    to tell apart.

    download_to_file already asks for SIZE. It now reports it, and the total
    deadline follows the file at the 25 KB/s floor _upload_deadline has used since
    do, so #2572's cap on the executor queue wait is untouched. Capped at 300s for a
    reason that is not about FTP: on_print_start holds a pooled DB connection across
    its whole 3MF hunt.

    Downloads also take turns per printer now. He watched Bambu Studio lose its own
    connection while Bambuddy pulled a 12 MB 3MF, and a later log caught two
    Bambuddy downloads of the same file overlapping at print start. The gate is
    soft - whoever cannot have it within 30s goes anyway, because a print losing its
    3MF to queueing is worse than the contention, and a soft gate cannot deadlock.

    It is also meaningful for the first time. The 90s cap on a multi-path lookup
    returned while its worker kept walking the remaining paths, still on the
    printer's socket; that walk is now cancelled and waited out before the printer
    is handed on.

    -----

    Look in the shared 3MF cache again before each cover retry (issue #2957)

    The cover endpoint and the print-start archive flow share a cache so whichever
    fetches the 3MF first hands it to the other (#972). The cover consulted it once
    on the way in, then retried for up to two and a half minutes without looking
    again.

    In the reporter's log the archive flow published the file 42 seconds into that
    sequence and the cover's third attempt still pulled its own 5,250,969-byte copy,
    off a printer that was mid-print on the same SD card.

    A file picked up that way is left alone rather than re-registered under this
    endpoint's own name or deleted on the way out. It is the archive flow's.
2026-08-29 14:20:15 +02:00