Commit Graph
408 Commits
Author SHA1 Message Date
maziggy 268940573d Stop a drying cycle reporting itself finished a minute in (#2759)
Starting the dryer on an AMS 2 Pro holding two PETG and two PLA spools
    and picking PLA showed "PLA @ 45°C" for about a minute and then switched
    to "PETG @ 65°C" for the remaining twelve hours.

    Bambu never echoes back which filament or temperature a cycle is
    running, so the badge reads the target cached when the command went out,
    and that cache had been dropped. Between accepting the command and
    settling its countdown the firmware publishes one update with the
    remaining time at zero while the unit is still in its Checking phase --
    the reporter's log has 720, then 0, then 719, and four seconds later the
    same unit's info hex decodes to dry_status 2, Drying. The falling-edge
    detector read that zero as the cycle ending. Losing the cached target
    left the badge to guess from the first loaded slot, which was PETG, and
    its RFID-recommended 65°C. The same false ending fired
    on_drying_complete, so anyone with smart-plug auto-off-after-drying
    switched on had power scheduled to cut one minute into a twelve-hour
    dry; the reporter had it off, which is the only reason this reads as a
    cosmetic bug.

    A remaining time of zero now ends a cycle only when the unit is not also
    reporting an active phase. dry_status comes from the same info hex
    already parsed a few lines above, so this costs nothing to check.
    Stopping and Error are deliberately not treated as active -- those
    should end it -- and a unit that reports no phase at all still ends its
    cycles, so the gate can only ever suppress on positive evidence that the
    cycle is live. A suppressed edge leaves the remembered dry_time alone,
    exactly as the #1462 absent-value skip does, so the push that really
    ends the cycle still sees a non-zero previous.

    The fallback guess is tightened to match. On a mixed unit the first tray
    is evidence of nothing, and naming a temperature the cycle is not using
    is worse than naming none, so it now answers only when every loaded
    spool is the same filament and otherwise leaves the badge showing the
    countdown alone. Both the websocket and REST status builders carried
    their own copy of that loop; they now share one helper, which also takes
    the temperature from the first slot that carries an RFID one rather than
    giving up when slot 1 holds a third-party spool.
2026-08-15 14:07:01 +02:00
maziggy 67e8a78cb8 Add auto-orient and auto-arrange to server-side slicing (#2548)
Both are per-slice checkboxes, off by default, forwarded as the sidecar's
    orient / arrange form fields. An unticked box is sent by omission: the
    sidecar treats any present value as truthy, so a literal "false" would
    have arranged every slice.

    Arrange unions with the #1493 cross-class decision rather than replacing
    it, and the per-plate slice-all loop is now keyed on the arrange flag
    itself — the project-wide collapse belongs to --arrange, not to the
    cross-class case. The loop also covers the embedded-settings path, whose
    crash-retry is suppressed there since a single --slice 0 retry would
    return one consolidated plate.
2026-08-15 14:02:31 +02:00
maziggy a8378b0e0e Merge remote-tracking branch 'upstream/dev' into feature/upload-prefer-filename-for-name 2026-08-15 14:00:15 +02:00
maziggy 71a06f3638 Add batch orders with a quantity per plate (#342)
Printing a multi-plate file in different quantities per plate meant
queueing each plate separately and tracking the counts by hand: one
shared Quantity field cannot say "plate 1 once, plate 2 twice, plate 3
three times". Each selected plate now carries its own quantity, and the
submission becomes an order on a new Batches tab.

The point is the distinction the old flat batch could not express.
print_batch_plates stores how many runs of each plate were wanted,
separately from what was queued, so a run that fails, is cancelled or is
skipped does not satisfy a target -- the order goes on saying it owes a
print instead of quietly under-delivering. Queue remaining re-queues
exactly what is missing, for the whole order or one plate, by cloning
the most recent item for that plate: that inherits the printer or model
target, AMS mapping, filament overrides and print options along with the
validation they already passed, rather than re-serialising twenty fields
through a template that would drift from the model the first time
someone adds a column. Clones append to the end of the relevant
printer's queue and take the same advisory lock the add-to-queue route
does; positions are per-printer sequences, not global.

Cost is measured, not estimated. print_log_entries gains queue_item_id,
set where the queue item is already in scope, so each run's material and
energy are attributed through the item that produced them -- an
unrelated reprint of the same archive never lands in an order's total,
and a multi-plate order gets each plate's own cost rather than the whole
file's via the plate-scoped estimate from #2614. Before any run has
completed there is no honest figure, so cost reads as unknown instead of
a fabricated 0.00.

The Batches tab wires up GET /queue/batches, which has been unreferenced
since the batch MVP shipped, along with six locale keys that were
translated and never used. It is a separate tab because an order
outlives the queue that produced it: once its runs finish they leave the
active queue, so Queue and History each hold half the picture.

completed was not a reachable status before now, so every batch created
since April is still marked active however long ago its last print
finished -- 73 of them on the development install. A startup pass closes
out the finished ones: those whose runs all completed become completed,
and groupings whose items were all cancelled become cancelled, which is
what they are. Not applied to orders, which state their intent
independently of their runs and still owe the work. Only batches with
nothing queued or printing are considered, and repeating the pass also
catches an order whose last run landed while the process was down.
Batches with neither items nor targets are no longer listed at all --
empty shells left when a grouping's items went with their source
archive.

Dispatch applies the same source-file gates as POST /queue/. It creates
queue items, so without them it would be a weaker door to the same
outcome; the archive and library-file checks move into shared helpers
so a third route cannot drift from them.
2026-08-04 11:11:36 +02:00
maziggy a9b57ccd3c Add variant-group endpoints and cross-model queue creation (#671, #2570)
Adds /library/variant-groups for declaring that several sliced files are
the same job for different printers, and a variants payload on queue
creation that turns such a set into one queue item with a candidate per
file.

The candidate set is validated as a set: one file per printer model, each
file sliced for the model it is offered as, and at least one model with
an active printer. A cross-model item deliberately holds no file of its
own, because print_queue.library_file_id is ON DELETE CASCADE and would
destroy the whole job when a single alternative is deleted.

Fixes internal printer-model codes never being resolved on queue create
and update: normalize_printer_model returns unknown input unchanged, so
the or-chain never reached the code map and a "C13" target matched no
printer and waited forever.

Skips candidates whose file is trashed or missing. Library deletes are
soft, and SQLite runs with PRAGMA foreign_keys off, so neither case is
covered by the schema; the hard-delete paths now also drop the rows.

Adds library_files.variant_target_model so a user can say which printer
a file without slicer metadata is for, kept out of file_metadata so the
assertion is never mistaken for parsed data.
2026-08-03 11:10:11 +02:00
maziggy da07c5884b Add variant-group data model for cross-model queue alternatives (#671)
Adds file_variant_groups plus variant_group_id / variant_position on
library_files, so a set of files that are the same job sliced for
different printers can be resolved to whichever printer frees up first.

Backfills groups from the sliced_from_library_file_id provenance that
slice_and_persist and the pipeline runner have been writing into
file_metadata since they shipped, and which nothing has ever read.
Only sources with two or more children carrying distinct
sliced_for_model values are grouped: a single candidate is not a
choice, and two slices for the same printer give the resolver no basis
to prefer one.
2026-08-03 10:17:28 +02:00
maziggy ad375f6ca7 Housekeeping 2026-08-02 12:29:32 +02:00
maziggy bbbb9d35c7 Bound the scheme repetition in the log credential-redaction pattern. As an
unbounded repetition the match was quadratic in the subject length: on a run
of scheme-legal characters the engine restarted at every offset and consumed
to the end before failing to find "://". ffmpeg echoes the configured camera
URL into its stderr and the whole blob reaches the pattern before any
truncation, so the subject length is attacker-influenced.
2026-08-02 11:17:49 +02:00
maziggy 3da4eee16e Bound the scheme repetition in the log credential-redaction pattern. As an
unbounded repetition the match was quadratic in the subject length: on a run
of scheme-legal characters the engine restarted at every offset and consumed
to the end before failing to find "://". ffmpeg echoes the configured camera
URL into its stderr and the whole blob reaches the pattern before any
truncation, so the subject length is attacker-influenced.
2026-08-02 11:17:24 +02:00
MartinNYHC 8be3413fbd Merge branch 'main' into dev 2026-08-02 11:10:25 +02:00
maziggy 4ff6377050 Fix per-job queue ETA showing for jobs that cannot start now
The scheduler only writes waiting_reason on the model-based assignment
    path, so a job pinned to a specific printer sits behind a running print
    with no marker at all. Every such job rendered an identical "starts now"
    ETA that was wrong by the length of everything ahead of it.

    Decide eligibility on the page instead: an item gets an ETA only when its
    printer is idle and it is the item the scheduler would dispatch next,
    following the same ordering the scheduler uses. Staged and future-
    scheduled items do not block the item behind them, matching the
    scheduler, and items conditional on a previous print are excluded.

    The value also froze at first render, since react-query's structural
    sharing keeps the queue reference stable and nothing re-rendered the row.
    formatETA now accepts a base instant and the page drives it from a 30s
    clock shared by every visible row.

    Retire the borrowed printers.estimatedCompletion tooltip for a queue key
    that says what the number means, translated into all 13 locales.
2026-08-02 09:54:41 +02:00
maziggy a7b96ea9d6 feat(vp): per-VP "Save AMS mapping" toggle + reprint auto-apply
Lets a reprint reuse the AMS slot the slicer itself picked, instead of
    re-deriving one from the file's static type/color.

    When a Print Queue VP has "Save AMS mapping" on, the slicer's own
    live-resolved ams_mapping (from the project_file MQTT command) is
    persisted onto the archive as extra_data.slicer_ams_mapping. A later
    reprint can reuse it via a new "Mapping" button in the filament-mapping
    panel — one click snaps every slot to the saved pick, click again
    reverts to auto-match. Archive cards and queue rows get an "AMS mapping
    saved" badge so it's visible beforehand. add_to_queue also falls back
    to the saved mapping automatically when the caller sends no explicit
    ams_mapping (e.g. a plain reprint with no per-slot edits).

    The queue item's own ams_mapping (used for that dispatch) is still
    captured unconditionally whenever the slicer provides it — that part is
    a correctness fix, not gated behind the toggle. Only the archive
    persistence for future reprints is opt-in.

    Split out from the original combined PR per review: this half is
    genuinely opt-in and low-risk (#2684). The dispatch-time validation
    gate that keeps a stored mapping honest (#1308) changes behaviour for
    every existing user and will land as its own PR.

    Review fixes applied:
    - _extract_slicer_ams_mapping_json: dropped the unreachable `v is None`
      arm and rejected bool explicitly (isinstance(v, int) accepts bool).
    - Translated the Russian docstring text to English.
    - save_ams_mapping's model comment moved to a trailing comment on the
      column line, matching the file's convention.
    - usingArchiveMapping now resets when the plate or archive changes, so
      the Mapping button can't read ON against a mapping it never applied.
    - Translated "Click to change slot assignment" and "Re-read".
    - add_to_queue's fallback is now called out explicitly in code comments
      and covered by three new integration tests (fallback fires, explicit
      mapping wins, unrelated extra_data doesn't false-trigger).
2026-08-02 09:48:11 +02:00
maziggy 4045ddbd1f fix(queue): withdraw an expected print when the command never goes out
feat(db): warn when the connection pool can outgrow the PostgreSQL server

    fix(mqtt): an unusable layer_num must not drop the printer connection

    test: patch settings.base_dir via monkeypatch so it unwinds on error

    test: restore the config module after reloading it
2026-08-02 09:41:32 +02:00
maziggy 6d8434b250 Security hardening (maziggy/bambuddy-security #N)
Subprocess output and user-supplied URLs are scrubbed of credentials
    before they reach the application log. Adds a shared redaction helper in
    core/logging_filters and routes the existing support-bundle sanitizer
    through the same pattern.
2026-08-02 09:36:38 +02:00
maziggy fe476dd98e fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse
    found nothing afterwards. Across 247 support bundles this was the norm, not an
    edge case: 457 automatic scans scheduled, 262 attached.

    The scan looked four times over ~65s. The attempt that found the video was #1
    272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
    files were still arriving when we stopped. What ran afterwards searched for the
    print name inside the filename; Bambu only writes "video_<timestamp>", so it
    fired 159 times and matched zero.

    The manual Scan had no baseline at all and matched on filename timestamp, FTP
    mtime, or "there is only one video" — all reading a clock a LAN-only printer
    cannot sync. The reporter's P1S was six and a half days out.

    - Poll for minutes instead of ~65s; drop the name-match fallback.
    - Persist the print-start baseline on the archive, so the diff survives a
      restart mid-print and the manual Scan runs the same comparison. With a
      baseline present the clock-based strategies are skipped entirely — they can
      only turn an honest "pick one" into a confident wrong answer.
    - When several files are new (a previous print's video landing late), exclude
      the ones already attached to another archive instead of ordering the
      candidates. Ordering could only be done on the printer's clock.
    - Delete the video from the printer once archived. Keeps /timelapse to
      unclaimed files, which is what makes the diff unambiguous, and stops P1S
      cards filling with AVIs.
    - Gate that delete on a verified transfer: download_file now compares against
      the size from the listing. An FTPS connection closing early does not always
      raise, so a partial buffer was being attached as a complete video — which
      would also have been the one case where deleting the source lost data.

    Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
    the deadline stops bounding the loop as soon as the sleeps are shortened.
    Per-round logging only speaks when the listing changed — 31 rounds of full
    listings would bury the interesting line in the support bundle.

    Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
    both dialects so a migrated database matches a fresh one.

    -----------

    fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)

    When a print records a timelapse, its last frame is the better finish photo:
    the firmware stops recording with the toolhead parked and before the end
    G-code drops the bed, where a live grab at that moment catches a lowered
    plate. Bambuddy waited 60s for the video and then gave up, because the
    print-complete notification blocks on that photo and holding a notification
    for minutes is worse than sending it with the live grab.

    P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
    Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
    worst 546s, while every other model finished inside 26s. So the printers that
    most needed the better framing were the ones that never got it.

    Keep the notification on the same bound, and keep waiting off to the side.
    _capture_finish_photo_from_timelapse now reports whether it ran out of time or
    concluded — a video that landed and failed extraction is not worth retrying,
    one that never arrived is. On the first, schedule a background task that waits
    up to 15 minutes and inserts the extracted frame at the front of the archive's
    photo list, where the gallery opens.

    The live grab stays on disk: the notification already links to that exact
    file, so removing it would leave a broken image in Discord or Telegram.

    The length check proves we received what the listing said, not that the file
    was finished. The first look happens ~5s after the print ends, while the
    printer may still be writing, so a growing file can be listed short, served
    short, and pass. Re-list after the download and only accept the video once its
    size has stopped changing — a failed re-list counts as not settled, since
    "could not check" must not mean "safe to delete".
2026-08-02 09:36:07 +02:00
maziggy 1cda64c35a fix(mqtt): report why a printer refused the connection instead of looping silently
A printer with a wrong access code gave no explanation anywhere. The connect
    callback's failure branch was a bare `state.connected = False`, discarding the
    CONNACK reason code the printer had just sent, so the only trace was paho's
    follow-up disconnect -- logged every 30 seconds as "rc=Unspecified error",
    which is exactly what a powered-off printer produces. In the report behind this
    fix one of three printers had been in that loop for the whole capture, and
    neither the log nor the support bundle could say why.

    Bambu speaks MQTT 3.1.1, whose CONNACK return codes 4 and 5 paho maps onto
    reason codes 134 and 135. Both are now logged with the printer's own reason
    string and, for those two, the remedy: the access code is regenerated whenever
    LAN Only or Developer Mode is toggled, so it has to be re-read from the screen.
    The access code itself is never logged -- it would land in every bundle.

    The reason is kept on the client as a stable slug and plumbed through
    test_connection into the connection diagnostic, which now distinguishes two
    cases it previously conflated. "The printer refused our credentials" is
    asserted only when the printer said so; when all Bambuddy knows is that there
    is no session, the text hedges and names the alternatives (rebooting, or
    already at its limit of simultaneous connections). The old wording claimed the
    access code was most likely wrong in both cases.

    Frontend needed no change -- ConnectionDiagnostic already renders
    `<status>_<reason>` variants with fallback to the plain per-status text, so an
    unrecognised slug degrades to today's wording rather than a missing key.
2026-08-02 09:35:23 +02:00
maziggy aef4f3a3e9 fix(oidc): strip the required BAMBUDDY_OIDC_* values and register the local-login bypass
A Kubernetes Secret written as a block scalar carries a trailing newline, and
the schema bounds the four required variables by max_length only, so an
unstripped issuer_url was stored and enabled and then raised httpx.InvalidURL
on the first click of the SSO button -- the authorize-time failure the
all-or-nothing rule exists to prevent. Whitespace-only values got through the
same way, contradicting the reader's own "an empty required var counts as
unset". The optional variables have always treated blank as unset; the
required ones now do too.

Also registers BAMBUDDY_LOCAL_LOGIN (#1589) in the typo guard, which logged
"possible typo" for it on every boot while listing every BAMBUDDY_OIDC_*
variable as legitimate.
2026-08-01 11:37:01 +02:00
MarianandClaude Opus 4.8 9e783fbad6 fix(oidc): keep the local-login bypass lenient under strict env_bool
Promoting env_bool to strict rejection made BAMBUDDY_LOCAL_LOGIN=on raise
EnvOIDCConfigError uncaught on the login/forgot-password path -- a 500 on
the exact recovery endpoint the bypass exists to keep open. env_bool gains
a strict flag (default True for the startup OIDC reader); the local-login
caller opts out so an unrecognized value falls back to "off" instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016q8EAf9Rj7ZHL92sPnXYxy
2026-08-01 09:17:33 +00:00
Marian 77c9bdd694 fix(oidc): reject an unrecognized boolean instead of guessing
_env_bool returned the default for anything outside {true,1,yes}, so
BAMBUDDY_OIDC_REQUIRE_EMAIL_VERIFIED=on silently read as OFF and
BAMBUDDY_OIDC_ENABLED=on silently disabled the provider -- the exact
opposite of what .env.example claimed. Unrecognized values now raise
EnvOIDCConfigError, caught in _apply_env_oidc_provider the same way a
bad DEFAULT_GROUP or a ValidationError already is: logged and left
running, never released on a typo.

Also promotes _env_bool to env_bool now that it has a call site in
auth.py, and corrects the boolean-parsing sentence in .env.example.
2026-08-01 09:17:33 +00:00
Marian c547505c64 fix(oidc): treat a blank optional env var as unset, not a refusal
BAMBUDDY_OIDC_SCOPES, _EMAIL_CLAIM and _ICON_URL fell back to their
default only when the key was absent, so `BAMBUDDY_OIDC_ICON_URL=` in
a compose file (as .env.example ships it, commented) reached the
schema validator as an empty string and got the whole provider
refused. default_group already treated blank as unset; these three
now follow the same rule.
2026-08-01 09:17:33 +00:00
MartinNYHC 21d61c3535 Merge branch 'dev' into feature/oidc-env-config 2026-07-31 17:05:24 +02:00
MarianandClaude Opus 4.8 6eea61dc78 fix(oidc): survive a failing rollback in the never-raise handler too
The recovery rollback after a failed commit was itself unguarded, so a
rollback that raises on a wedged connection would still take the boot
down -- the exact failure the never-raise contract exists to prevent.
Suppressed; the caller's `async with` discards the session regardless.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016q8EAf9Rj7ZHL92sPnXYxy
2026-07-31 13:18:21 +00:00
Marian c4b5d42f48 fix(oidc): log distinctly when env config adopts a UI-created provider
A name collision with a provider that was NOT already env-managed
overwrites its issuer, client id and secret in place and locks it
behind the env-managed 409 -- but it logged the same routine "applied"
line as an ordinary re-apply, giving no signal a UI provider was just
taken over. Adoption is now a WARNING with its own wording; a routine
re-apply of an already env-managed provider keeps the INFO line.
2026-07-31 13:11:15 +00:00
Marian 6dff1e9644 fix(oidc): make apply_env_oidc_provider never raise on DB errors
The db.execute/db.commit calls in the upsert and release paths sat
outside the try/except that only wrapped OIDCProviderCreate, so a
commit failure at startup (connection blip, WAL lock) propagated out
of the lifespan and took the instance down -- the exact outcome this
module exists to avoid. The body now runs inside a private
_apply_env_oidc_provider(), with the public entry point catching,
logging and rolling back on any exception.
2026-07-31 13:10:14 +00:00
Sergey Dontsov bab1cfb906 feat(vp): per-VP "Save AMS mapping" toggle + reprint auto-apply
Lets a reprint reuse the AMS slot the slicer itself picked, instead of
re-deriving one from the file's static type/color.

When a Print Queue VP has "Save AMS mapping" on, the slicer's own
live-resolved ams_mapping (from the project_file MQTT command) is
persisted onto the archive as extra_data.slicer_ams_mapping. A later
reprint can reuse it via a new "Mapping" button in the filament-mapping
panel — one click snaps every slot to the saved pick, click again
reverts to auto-match. Archive cards and queue rows get an "AMS mapping
saved" badge so it's visible beforehand. add_to_queue also falls back
to the saved mapping automatically when the caller sends no explicit
ams_mapping (e.g. a plain reprint with no per-slot edits).

The queue item's own ams_mapping (used for that dispatch) is still
captured unconditionally whenever the slicer provides it — that part is
a correctness fix, not gated behind the toggle. Only the archive
persistence for future reprints is opt-in.

Split out from the original combined PR per review: this half is
genuinely opt-in and low-risk (#2684). The dispatch-time validation
gate that keeps a stored mapping honest (#1308) changes behaviour for
every existing user and will land as its own PR.

Review fixes applied:
- _extract_slicer_ams_mapping_json: dropped the unreachable `v is None`
  arm and rejected bool explicitly (isinstance(v, int) accepts bool).
- Translated the Russian docstring text to English.
- save_ams_mapping's model comment moved to a trailing comment on the
  column line, matching the file's convention.
- usingArchiveMapping now resets when the plate or archive changes, so
  the Mapping button can't read ON against a mapping it never applied.
- Translated "Click to change slot assignment" and "Re-read".
- add_to_queue's fallback is now called out explicitly in code comments
  and covered by three new integration tests (fallback fires, explicit
  mapping wins, unrelated extra_data doesn't false-trigger).

Closes #2684
2026-07-31 11:08:58 +03:00
MartinNYHC c3448dae91 Merge branch 'dev' into feature/oidc-env-config 2026-07-31 08:20:49 +02:00
Marian ed29319cf9 docs(oidc): drop the upgrade-path claim from the release-path rationale
The two-flagged-row state was never released, so no installation can carry it
in -- the reason to release every flagged row is not repair, it is that the
sweep's invariant is not enforced anywhere and scalar_one_or_none() turns a
broken one into a failed boot. Comment, test docstring and test name say that
instead.
2026-07-30 14:44:49 +00:00
Marian 3c679459e1 feat(oidc): set the default group from the environment, by name
Without it every account auto-created through the env provider fell back to
Viewers (routes/mfa.py), and because the provider is locked the UI could not
correct it either -- a real limitation for a declarative deployment running
BAMBUDDY_OIDC_AUTO_CREATE_USERS=true.

BAMBUDDY_OIDC_DEFAULT_GROUP names a group rather than an id: ids are handed out
per installation, so the same compose file would point at a different group on
the next deployment. The name is matched exactly, resolved against the database
before anything is written, and default_group_id joins _APPLIED_FIELDS so
dropping the variable clears the group again -- the environment is the whole
truth for this row.

A name that matches no group is refused rather than defaulted: silently landing
users in Viewers is the failure this variable exists to remove, and the API
already answers 422 for a default_group_id that does not exist. The refusal is
logged and survivable, and it says which of the two cases happened, because
they differ sharply -- an existing provider keeps running on its last good
config, while on a first boot nothing is created and no SSO button appears.

Raised by maziggy in review of #2625 as a scope decision; documented in
.env.example and in the companion wiki PR.
2026-07-30 14:43:23 +00:00
Marian d64f5d9651 fix(oidc): release the row the env config managed before a rename
apply_env_oidc_provider matches the provider by BAMBUDDY_OIDC_NAME but never
cleared is_env_managed from the row it managed previously. Renaming the
variable therefore left two flagged rows, and both consequences are reachable
by ordinary config edits: the old row stayed enabled with a stale issuer and
secret on the login page while _refuse_if_env_managed answered 409 to every
attempt to edit, disable or delete it -- the dead end reachable only through
the database that the release path exists to prevent -- and unsetting the
variables later hit scalar_one_or_none() on two rows, so MultipleResultsFound
propagated out of the lifespan and the app stopped booting.

The upsert now sweeps the flag off every other row, the same shape the
autologin sweep one block down already uses: disable and release rather than
delete, for the same cascade reason as everywhere else in this branch. The
release path releases every flagged row it finds instead of exactly one -- the
sweep should keep that at one, but a release path that dies with
MultipleResultsFound the moment that invariant breaks is a second way to lose
the boot, and the query costs the same either way.

Releasing now clears is_autologin as well. Without it a released row keeps a
latent autologin claim: update_oidc_provider only re-runs the exclusivity sweep
when a request sets is_autologin=True, so merely re-enabling the row in the UI
would silently make it the autologin target again.

Reported by maziggy in review of #2625, with the rename reproduction.
2026-07-30 14:43:23 +00:00
maziggy ad785a95cb fix(queue): withdraw an expected print when the command never goes out
feat(db): warn when the connection pool can outgrow the PostgreSQL server

fix(mqtt): an unusable layer_num must not drop the printer connection

test: patch settings.base_dir via monkeypatch so it unwinds on error

test: restore the config module after reloading it
2026-07-30 16:31:09 +02:00
MartinNYHC ea1869aeb3 Merge branch 'dev' into feature/oidc-env-config 2026-07-29 13:09:20 +02:00
maziggy e325948dcc Security hardening (maziggy/bambuddy-security #N)
Subprocess output and user-supplied URLs are scrubbed of credentials
before they reach the application log. Adds a shared redaction helper in
core/logging_filters and routes the existing support-bundle sanitizer
through the same pattern.
2026-07-29 12:36:17 +02:00
maziggy 5a67dffe4f fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse
found nothing afterwards. Across 247 support bundles this was the norm, not an
edge case: 457 automatic scans scheduled, 262 attached.

The scan looked four times over ~65s. The attempt that found the video was #1
272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
files were still arriving when we stopped. What ran afterwards searched for the
print name inside the filename; Bambu only writes "video_<timestamp>", so it
fired 159 times and matched zero.

The manual Scan had no baseline at all and matched on filename timestamp, FTP
mtime, or "there is only one video" — all reading a clock a LAN-only printer
cannot sync. The reporter's P1S was six and a half days out.

- Poll for minutes instead of ~65s; drop the name-match fallback.
- Persist the print-start baseline on the archive, so the diff survives a
  restart mid-print and the manual Scan runs the same comparison. With a
  baseline present the clock-based strategies are skipped entirely — they can
  only turn an honest "pick one" into a confident wrong answer.
- When several files are new (a previous print's video landing late), exclude
  the ones already attached to another archive instead of ordering the
  candidates. Ordering could only be done on the printer's clock.
- Delete the video from the printer once archived. Keeps /timelapse to
  unclaimed files, which is what makes the diff unambiguous, and stops P1S
  cards filling with AVIs.
- Gate that delete on a verified transfer: download_file now compares against
  the size from the listing. An FTPS connection closing early does not always
  raise, so a partial buffer was being attached as a complete video — which
  would also have been the one case where deleting the source lost data.

Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
the deadline stops bounding the loop as soon as the sleeps are shortened.
Per-round logging only speaks when the listing changed — 31 rounds of full
listings would bury the interesting line in the support bundle.

Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
both dialects so a migrated database matches a fresh one.

-----------

fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)

When a print records a timelapse, its last frame is the better finish photo:
the firmware stops recording with the toolhead parked and before the end
G-code drops the bed, where a live grab at that moment catches a lowered
plate. Bambuddy waited 60s for the video and then gave up, because the
print-complete notification blocks on that photo and holding a notification
for minutes is worse than sending it with the live grab.

P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
worst 546s, while every other model finished inside 26s. So the printers that
most needed the better framing were the ones that never got it.

Keep the notification on the same bound, and keep waiting off to the side.
_capture_finish_photo_from_timelapse now reports whether it ran out of time or
concluded — a video that landed and failed extraction is not worth retrying,
one that never arrived is. On the first, schedule a background task that waits
up to 15 minutes and inserts the extracted frame at the front of the archive's
photo list, where the gallery opens.

The live grab stays on disk: the notification already links to that exact
file, so removing it would leave a broken image in Discord or Telegram.

The length check proves we received what the listing said, not that the file
was finished. The first look happens ~5s after the print ends, while the
printer may still be writing, so a growing file can be listed short, served
short, and pass. Re-list after the download and only accept the video once its
size has stopped changing — a failed re-list counts as not settled, since
"could not check" must not mean "safe to delete".
2026-07-29 11:51:30 +02:00
MarianandClaude Opus 4.8 1116b43fbd fix(oidc): never log the client_secret when env config is rejected
apply_env_oidc_provider logged the raw Pydantic exception on rejection.
client_secret has max_length=512, so a longer value raises string_too_long
and str(exc) embeds input_value=..., leaking BAMBUDDY_OIDC_CLIENT_SECRET into
the logs (maziggy review, PR #2625).

Split the catch: ValidationError logs errors(include_input=False), which
strips submitted values; any other exception logs only its class name, never
str(exc). Rejection stays survivable — a bad config is still skipped and the
app still boots.

Adds two regression tests: an over-long secret is rejected without the value
reaching the log, and a non-ValidationError is survived without leaking its
message.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-28 21:05:29 +00:00
Marian c163b3524a fix(oidc): identify the env provider by name, and release it when unconfigured
Two problems, both from using is_env_managed as the provider's identity.

An operator who names the env provider after one that already exists hit the
unique constraint on `name` during the insert. That happens inside the
lifespan, so the app did not boot -- from a function whose docstring promises
it never raises. The lookup now matches on the name, which is unique, so an
existing provider is adopted and updated instead of duplicated.

And removing the config left the row disabled but still flagged, so the API
went on refusing every edit and delete while nothing managed it any more: a
dead end reachable only through the database. The flag is now cleared as well,
handing the provider back to the UI. Re-adding the config finds the same row
by name, so the account links it carries survive the round trip.

Falls out of the same change: the issuer URL and client id can be rotated
under an unchanged name without orphaning those links.

Found by Marian asking what happens when you want to change the provider --
the answer was "you cannot, ever again".

Refs #2593
2026-07-28 21:05:29 +00:00
Marian 8d1b2b9027 chore(config): register BAMBUDDY_OIDC_* in the typo-guard
Unknown BAMBUDDY_* vars log "possible typo" on every boot, so a correct OIDC
config would have told its operator it was wrong, once per restart.

The test asserts against the reader's own variable list rather than a copied
one, so a thirteenth variable added later fails here instead of surfacing in
somebody's logs.

Refs #2593
2026-07-28 21:05:29 +00:00
Marian 58602a3f1b feat(oidc): upsert the env-managed provider
The row is updated in place, never delete-recreated: user_oidc_links
references it with ON DELETE CASCADE, so recreating the provider would
silently unlink every account bound to it. For the same reason, removing the
variables disables the provider rather than deleting it -- the links would not
come back when the config does.

Config goes through OIDCProviderCreate, the schema the API already uses, so
the environment cannot reach a state the UI would have refused. That covers
the SEC-1 auto-link check: auto-link plus unverified email is an account
takeover, and it is rejected here exactly as it is in the UI.

Nothing raises. This runs during startup, so a typo in one variable must not
stop the app from booting -- a rejected config is logged and skipped, leaving
the previous provider untouched.

Refs #2593
2026-07-28 21:05:29 +00:00
Marian d6ecd92480 feat(oidc): read BAMBUDDY_OIDC_* env config
A declarative deployment has no way to click through the settings UI, so one
provider can be configured entirely from the environment. This reads and
defaults only -- validity is decided later by the same OIDCProviderCreate
schema the API uses, so env config cannot bypass a check the UI enforces.

All four required vars or nothing, and an empty one counts as unset: a
provider missing its secret would otherwise be written to the database and
fail at authorize time, far from the typo in the compose file that caused it.
Booleans follow the project's existing spelling convention (true/1/yes), so an
unrecognised value leaves the documented default rather than guessing.

Refs #2593
2026-07-28 21:05:29 +00:00
Marian e3cada51ac feat(oidc): add is_env_managed column to oidc_providers
Marks the single provider that BAMBUDDY_OIDC_* environment variables define,
so a later change can upsert it on startup and refuse UI/API writes to it. The
row is never delete-recreated: user_oidc_links.provider_id is FK ON DELETE
CASCADE, so dropping the provider would take every account link with it.

The migration carries its own test rather than relying on the model test. A
table created from metadata already has the column, so that path never
exercises the ALTER; an installed instance gets it only through
run_migrations, and that is the path an upgrade actually takes. Covered:
the column appears on a pre-existing table, rows created before the upgrade
read as not env-managed, and re-running is a no-op because every boot replays
the whole migration set.

Refs #2593
2026-07-28 21:05:29 +00:00
maziggy 4059d63373 Bumped version 2026-07-28 15:07:28 +02:00
maziggy 8fd1f884dc feat(mqtt): publish the plate-clear gate and add a notification for it (#2525)
When a print reaches a terminal state Bambuddy holds the queue until
someone confirms the build plate is clear. That gate was visible only in
the Web UI: the printer's own MQTT push reports nothing beyond RUNNING,
PAUSE, FAILED, FINISH and IDLE, so an external automation could not tell
"finished" from "finished and still waiting for a human".

The per-printer status topic now carries an awaiting_plate_clear field,
and every transition is additionally published on a new retained topic,
bambuddy/printers/{serial}/plate_clear. Retained, and published from the
flag itself rather than from printer telemetry: a subscriber learns the
state of every printer the moment it connects, and the state stays
correct after Auto Off powers a printer down - telemetry stops there,
which would otherwise leave the status topic frozen at false.

Publishing is edge-triggered. The queue clears the gate on every
dispatch whether or not it was up, and no subscriber should see a
"plate cleared" for a plate that was never dirty. Persistence and the
WebSocket broadcast stay unconditional; they are idempotent and predate
this.

A matching Plate Clear Required notification event was added, off by
default on every provider because it fires after every print at the
same moment as the print-complete alert. Only the rising edge notifies.
Acknowledging still goes through POST /printers/{id}/clear-plate.

Two tests in test_printer_manager_status_broadcast.py asserted
_schedule_async.call_count == 2 for the setter. The new emission makes
it three on a transition, so they now assert that the persist and
broadcast coroutines are actually scheduled - which is the contract

Translated in all locales; wiki updated. Covered by backend and
frontend tests.
2026-07-28 13:36:55 +02:00
maziggy af7874546a feat(projects): per-file print progress and complete-sets tracking (#1897)
Projects made of many distinct files that each need N prints (e.g. 13
plates x 10 sets = 130 prints) only had aggregate progress. Finding out
"how many times have I printed plate_7?" meant reading the Activity
Timeline line by line, unusable at 130 events.

Projects now take an optional Copies per File target. Every printable
file in the project's linked folders shows an X / N badge with a mini
progress bar (gray not started, amber in progress, green done), and the
progress card gains a Complete Sets bar - the minimum per-file count,
i.e. how many finished assemblies can be shipped right now. Without the
target, printable files show a plain printed-count badge.

Counting matches the aggregate project stats: completed runs only,
served by a new /projects/{id}/file-progress endpoint. Runs attribute
to a file via a new library_file_id stamp on queue-dispatched archives,
falling back to content hash and then filename for historical rows.

Also fixed: files queued from a project-linked File Manager folder now
inherit that project, so their prints count toward project statistics -
previously only prints started from the project page were attributed.

Test-harness fix along the way: the test suite's get_db override never
committed, unlike production get_db, so endpoints relying on the
request-scoped commit silently lost their writes in tests. The override
now mirrors production commit/rollback semantics.
2026-07-28 12:46:31 +02:00
maziggy f957fcc717 Bumped version 2026-07-27 15:03:02 +02:00
maziggy 1bdd7d224a fix(library): sort File Manager by real filesystem mtime, recursively (#2680)
The folder tree's "sort by recent activity" and the file pane's date sort
put external (mapped/NAS) files in a near-random order instead of ls -t's
newest-first. Nothing captured the files' on-disk mtime: the sort keyed off
the DB updated_at/created_at, which for a bulk external scan is the same
scan instant for every row, so a whole block tied and sorted arbitrarily;
only rows Bambuddy had later touched individually looked "partially right."
The tree also bubbled up only immediate child-file activity, so a file added
deep in a subtree never lifted its parent folders.

- Add nullable fs_modified_at to LibraryFile and LibraryFolder (dialect-
  branched migration, mirroring the #2615 dispatching_at pattern).
- External scan records each file's and directory's real os.stat().st_mtime
  and refreshes it on every re-scan, so a file edited over the mount
  re-sorts and existing installs backfill on the next scan.
- list_folders computes each folder's activity as a recursive newest-
  descendant roll-up (post-order), so a fresh deep file lifts every ancestor.
- Folder tree sort and the file pane's date sort now use the real mtime,
  falling back to created_at for managed uploads with none.
- New toolbar toggle shows/hides each item's last-modified date in the right
  pane (grid + list), with strings in all locales.

Store the mtime as naive UTC to match the other timestamp columns so activity
comparisons never mix naive and aware values on either dialect. Covered by
integration tests (mtime capture, re-scan refresh, deep-file recursive bubble,
folder mtime) and a frontend test proving fs_modified_at is preferred over
created_at.
2026-07-27 11:01:43 +02:00
maziggy 30e4577838 Bumped version 2026-07-24 12:32:36 +02:00
maziggy 559571a35f Bumped version 2026-07-24 12:08:46 +02:00
maziggy 7c875b88ff chore(bandit): annotate tri-state migration SQL as a B608 false positive
The bed_levelling/flow_cali/nozzle_offset_cali boolean->tristate migration
builds two UPDATE statements with an f-string interpolating the column
name. Bandit flags these as B608 (SQL injection) at medium severity, which
failed the release-gate scan in test_security.sh.

The interpolated _col only ever iterates the hardcoded _tristate_cols tuple,
never user input, and SQL identifiers can't be passed as bound parameters.
Suppress with `# nosec B608` (matching the existing settings.py convention)
plus an inline rationale. No behavior change.
2026-07-24 11:24:25 +02:00
maziggy a898de3bb7 Bumped version 2026-07-24 11:10:11 +02:00
maziggy 56accd24de fix(smart-plugs): don't blank printer state when an accessory plug switches off (#2629)
An end-of-print auto-off on a plug that powers a filter fan marked the linked
printer offline and forced its state to "unknown". The mark was unrecoverable:
connected heals on the next MQTT message but state does not (only frames
carrying gcode_state rewrite it, and steady-state push_status frames are
partial), so the printer stayed "unknown" until a manual Force Refresh and the
queue never dispatched to it again.

The offline mark is now an explicit presumption: mark_power_off records the
state it overwrites and _on_message undoes it as soon as the printer sends
another report on its own topic, since inbound traffic proves the power was
never cut. A reconnect discards the saved state, so a genuine power cut is
unaffected. Each plug also gains a controls_printer_power flag (default true,
backfilled) that gates all five power-off paths, and the queue's power-on step
now picks the flagged plug instead of whichever linked plug came first.
2026-07-22 09:57:27 +02:00
maziggy 2e45893dd5 feat(print-options): add "Auto" state to bed levelling, flow & nozzle-offset calibration
Bed levelling, flow calibration, and nozzle-offset calibration were on/off
only, so the sole way to run bed levelling was to force a full level before
every print. Bambu Studio has always offered a third "Auto" state that lets
the printer skip the calibration when it was done recently -- the state most
users actually want. Make these three options tri-state (off/on/auto),
defaulting to auto, and leave vibration/layer-inspect/timelapse as on/off
(Bambu Studio exposes no auto for those).

Wire encoding follows Bambu Studio's source exactly: each option sends a JSON
bool (true only for "on") plus a companion int -- off=0, on=1, auto=2. The
bool fields stay booleans (the #1478 H2S regression); only the companion int
widened from {0,1} to {0,1,2}. #1721's observation that stage 8/39 stays
queued when sending 2 is the auto contract (queued, skipped at runtime if
recent), not a broken "off".

- schemas: TriState = Literal[off/on/auto] with a BeforeValidator coercing
  legacy bool / 0-1 / true-false so old clients and un-migrated rows validate
- model + migration: boolean columns -> String; SQLite via column affinity +
  data backfill, PostgreSQL via ALTER COLUMN TYPE guarded on information_schema
  (verified on both dialects); settings rows normalised true/false -> on/off
- MQTT: start_print takes the tri-state strings and emits the paired bool+int
- Virtual Printer: reconstructs the slicer's auto/on/off from the int companion
  (auto_bed_leveling / extrude_cali_flag) in both capture paths
- frontend: CalibrationMode type; off/auto/on segmented controls in the print
  dialog, queue bulk-edit, and Settings -> Workflow; calibrationMode_* strings
  in all 11 locales
2026-07-20 17:55:56 +02:00