mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-10-04 13:11:35 +02:00
29311cab2bf0fcb58c5f175feeab1c7f96cd2d21
742
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bbd991510f |
fix(backup): refuse prometheus_enabled when the backup has no token either (#2656)
The companion-credential rule has five conditions, and the second one -- "the backup itself carried a usable credential" -- was applied to all five pairs. It should not be. It is what stops the rule over-refusing an anonymous MQTT broker or an anonymous LDAP bind, both of which are working configs: there, an empty credential in the backup means the restore is not producing anything weaker than what was backed up. For prometheus_enabled it does not transfer. An empty prometheus_token removes /api/v1/metrics' only gate, so the exposure is a property of the toggle, not of a downgrade relative to the backup -- and prometheus_token is optional, so a backup taken on an instance that enabled Prometheus without ever setting one carries the toggle and no usable token. That payload skipped the refusal entirely: not a candidate, so the local-state pass never ran, and the blocklist quietly dropped the token key. On a token-less target the result was prometheus_enabled=true, no token row anywhere, and a full unauthenticated metrics dump -- the same hole the rule was written to close, reached from the likelier of the two directions. So condition 2 is now per-pair: an exposure class (prometheus) that skips it and is judged on local state alone, and an availability class (mqtt, ldap, ha, virtual_printer) that keeps it. Nothing else changes -- the local-state pass already stands down when the instance has its own credential, when HA_TOKEN is in the environment, and when the toggle is already on locally, so "the exposure pre-dates this restore" still holds and refusals still get no tally increment. One wording consequence: an exposure toggle can now be refused on a payload with no credential-like key in it at all, where the shared caveat would have read "0 credential-like key(s) will be skipped". That case gets its own preview detail code, settingsCompanionOnlyWillSkip, in all 13 locales. Tests: 6 that fail pre-fix -- the token key absent and blank at the unit level, the new preview wording, and the integration test through the real endpoint for both payloads (200 with a full metrics body before this, 404 after). Plus 3 controls, because over-refusal is still the real risk: the exposure route must still stand down for a local token and for an already-on toggle, and the availability class must still let a credential-less mqtt/ldap/virtual_printer toggle through. The anonymous-broker and anonymous-bind controls are unchanged and still pass. |
||
|
|
cb6a4e6d88 |
fix(backup): stop two backed-up K-profiles claiming one live slot (#2656)
`_match_kprofile` ends in a single-candidate fallback, and the per-nozzle loop called it once per entry with no record of which live profiles were already taken. Two backup entries sharing a `filament_id` and matching on neither `setting_id` nor `name` both resolved to the same live profile, both got the same `cali_idx`, and both went into the batch — so the second overwrote the first on the printer while the tally counted two restored. Reachable in the ordinary way: the user deletes one of a pair after the backup, and the delete-then-add re-key this code already reasons about is exactly what strips the `setting_id` match. Fix: thread a `claimed` set of slot ids through the loop; a live profile can only stand in for one entry. A displaced entry falls through to `cali_idx: -1` — add-as-new is the safe outcome — and folds into the existing `kprofilesUnmatched` note rather than earning a new code. The single-candidate fallback is still judged against every candidate rather than the unclaimed ones. Two live profiles for one filament are ambiguous whether or not another entry has taken one, and narrowing to "available" would turn a guess the code deliberately refuses into a match. Tests: 2 regression (the displaced entry is added rather than aliased, and keeps its own setting_id) + 2 controls (two genuine matches keep their own slots; a claimed slot does not make an ambiguous pair matchable). Both regressions confirmed failing against the pre-fix service. |
||
|
|
582a6b18bc |
fix(backup): page Gitea's tree off what came back, not what we asked for (#2656)
The pager computed `seen = (page - 1) * 1000 + len(entries)`, taking the requested `per_page` as fact. Gitea clamps `per_page` to `MAX_RESPONSE_ITEMS`, which defaults to 50. So on a default install page 1 returns 50 entries and sets `seen` to 50, then page 2 sets it to 1050 — which clears any `total_count` under 1050. The loop returns `success: true` holding the first 100 entries of a much larger tree. The restore then reads every missing path as "category not present in this commit" and skips it silently, which is precisely the failure this override was written to prevent. Same class as the GitLab pager fix, in the one direction that got left behind. Fix: accumulate `seen += len(entries)`. A genuinely over-cap tree still hard-fails rather than truncating; the cap is a page count, not a file count, because the page size is the server's choice. Test: a 120-entry tree served 50 at a time reaches its last entry, in three requests. Confirmed failing against the pre-fix backend — it stopped after two pages and reported 100 entries as the whole tree. |
||
|
|
be8e8d545f |
fix(backup): don't let an old backup commit blank an archive's owner (#2656)
`created_by_id` and `deleted_at` both went into the archive `fields` dict unconditionally, via `entry.get(...)`. A backup commit taken before the collector wrote those keys carries neither, so `.get` yielded None for both and the overwrite branch — a blanket `setattr` over every key — wrote NULL over a live owner. That is exactly the failure carrying `created_by_id` was added to fix, only now inflicted on rows that were fine: `_ensure_archive_visible` fails closed on a NULL owner, so the archive 404s for the person who owns it. It emitted no note either, because `archivesOwnerCleared` only fires for an id that isn't in `valid_users`, not for an absent key — and the row still counted as restored. `deleted_at` had the mirror problem: an old commit silently un-deleted, since `archivesUndeleted` reads the same absent value. Absent is not the same as explicitly null. Both keys now only enter `fields` when the entry actually carries them, so an old commit leaves the column alone on overwrite and a current one can still say "this archive has no owner" or "this archive is live". Same shape as the tag-column rule: don't clear what the backup doesn't know about. Behaviour change to an existing test, called out deliberately: `test_overwrite_undeletes_a_locally_deleted_archive_and_says_so` now has to put `deleted_at: None` in the entry to mean it. Tests: 2 regression (owner and deleted_at both left alone by a key-less entry) + 2 controls (an explicit null is still honoured, with its note). Both regressions confirmed failing against the pre-fix service. |
||
|
|
158301ac8a |
fix(backup): halve the provider round-trips, and stop losing commit metadata (#2656)
The three remaining review items, all in the read path. E1 — four provider calls where two would do. preview() called list_commits twice: once inside _resolve_ref to turn HEAD into a SHA, once more at limit=20 purely to find the entry describing that same SHA. And list_tree's recursive tree GET was thrown away, so fetch_files immediately fetched the identical tree again to map path -> blob SHA. _resolve_ref now returns the entry it already has, and list_tree returns its blob_shas map for fetch_files to take as an optional argument. GitLab reads files by path and ignores it. E2 — `commit: null` for a ref outside the 20 most recent. Two causes, and the second is the one that actually bit: REF_PATTERN accepts a 7-character ref while providers return the full 40, so the exact `==` in the scan never matched an abbreviated SHA *even when the commit was in the window*. Fixed by prefix comparison, plus a get_commit(ref) on the GitHub and GitLab backends for the genuinely-outside-the-window case. Gitea and Forgejo inherit GitHub's. Still best-effort: it is a subject line and a date, so a failed lookup renders the preview without them rather than failing it. E7 — the two tree readers disagreed, and each was wrong in the other's direction. GitHub's recursive trees endpoint is not paginated and signals overflow with truncated=true, which _blob_shas_at hard-fails on. Gitea and Forgejo *do* page that endpoint, and inherited that single GET unchanged — so a large backup repo returned only the first page and every category beyond it looked absent from the commit. GiteaBackend now has its own paging _blob_shas_at. GitLab had the mirror-image bug the review did not name: at its 50-page cap it exited through the while condition and returned success: True with a silently partial path list. Both now fail loudly, which is what the GitHub version was always doing. Both halves of E7 are the same failure the module already refuses to allow: a restore that skips categories and calls it "not present in this backup commit". 24 new or changed tests, all failing against this commit's parent. |
||
|
|
945f50a1ab |
fix(backup): stop overwrite writing a spool's other tag key (#2656)
Four of the review's smaller items. E6, the substantive one. tag_uid and tray_uuid are both in the overwrite setattr loop, so a spool matched on one key got the backup's *other* key written onto it. Neither column has a unique constraint (models/spool.py, and no unique index in the migrations), so nothing errors — a duplicate tag simply appears, after which _find_spool's .scalars().first() is non-deterministic and an AMS tag lookup resolves to an arbitrary one of the two spools. The same loop could also clear a tag the user had scanned since the backup was taken, when the backup entry held None. _find_spool now reports which key matched, and _guard_tag_overwrite drops a tag column from the write when the incoming value is empty and the local row has one (the backup predates the scan, so the local tag is the newer fact) or when another local spool already holds it. Announced in the tally the way the archive un-delete case already announces itself, rather than done silently — the spoolTagKept locale key landed with the rest of the i18n block last commit. E5. The Restore button is hidden without github:restore. All three endpoints are gated on it server-side, so the modal 403s on its first preview; offering the button is offering an action that cannot work. Button only — the card stays visible, since configuring backups is a separate permission — and hasPermission returns true with auth off, so a single-user instance is unaffected. E3. models/github_backup.py: the trigger comment said manual/scheduled; this PR added a third value. E4. ha_token_from_env: recommending no change, with the reasoning recorded as a test rather than left in a review thread. It is built only in the settings GET response, is absent from AppSettingsUpdate, and is therefore never a Settings row — it cannot reach a backup, so an allowlist entry would be dead code. Worse, a name-shaped exception to a belt-and-braces denylist is a live hole: an attacker-authored settings/app_settings.json could get a *token*-named row written by choosing that name. 4 unit tests and 1 frontend test that fail against this commit's parent, plus 4 controls: a free tag is still written, an unchanged tag is not reported as kept, an insert is unaffected, and the button still shows with auth disabled. |
||
|
|
af8d14d796 |
i18n(backup): make the restore notes and preview details translatable (#2656)
A German user got a translated modal with "Not present in this backup commit" in
the middle of it. Every tally note and preview caveat was a server-built English
sentence rendered verbatim.
Follows the backup.pathCheck contract already in use one card down in the same
component, deliberately rather than inventing a second convention: the server
sends a `code` plus typed `params` and carries the English along as the
fallback, and the client renders
`t(`...${code}`, { ...params, defaultValue: message })`. The defaultValue arm is
what keeps a newer backend's unfamiliar code readable instead of printing the
raw key — covered by its own test.
Shapes:
- notes: list[str] -> list[GitHubRestoreNote] {code, params, message}. Breaking,
but the field is unreleased in this same PR.
- GitHubRestorePreviewCategory gains detail_code / detail_params; `detail` stays
as the English fallback.
- _CategoryTally.note(code, message, **params), deduped on (code, params) rather
than on the rendered text, so two offline printers both keep their names. The
20-note cap is unchanged.
28 new leaves across 13 locales: 20 notes.* and 8 details.*. `noData` collapses
the four per-category "No X data in this backup" strings into one, since the
category heading already renders beside it. Counts use single-form {{count}} in
the existing "N record(s)" style rather than i18next plural suffixes — nothing in
this block uses _one/_other and the parity script has extra rules for them.
Parity holds at 5737 leaves in all 13 locales.
Deliberately out of scope, and worth saying so rather than leaving it to look
like an oversight: result.message, the commit-picker subject lines and the HTTP
error strings stay English. Those also originate in the provider backends, so
code-ifying them widens the diff well past the restore service.
spoolTagKept is added here with the rest of the locale churn but is not emitted
until the next commit, so the 13-locale change lands once.
|
||
|
|
cfa82bfcfb |
fix(backup): carry the owner across, or restored archives are invisible (#2656)
Neither _collect_archives nor _restore_archives touched created_by_id, so every restored archive row landed NULL. That column is not attribution, it is what the access check runs on: _ensure_archive_visible (api/routes/archives.py) fails closed on NULL — a 404 for any caller without archives:read_all — and the list paths filter created_by_id == user.id. On a multi-user instance the tally therefore reported archives restored while the person who owns them could neither list nor open them. Same shape as the deleted_at fix, and the same remedy: the collector records the key next to deleted_at, the restore mirrors the printer_id/project_id pattern exactly — one hoisted select(User.id), a membership test per row, an unknown id coerced to None rather than failing the row, and one de-duplicated note. It is in the overwrite setattr loop too, so overwrite keeps meaning "make local match the backup". Additive on the backup side, so older backups still restore; they just cannot know the owner. Clearing the id is not silent-safe, so the note says what it costs: those archives are visible only to users with archives:read_all until an admin reassigns them. Caveat recorded in a comment and raised in the PR, not decided here: this is the one place the module reuses a raw backup id, against its own rule. Validating it means a *stale* id clears rather than pointing somewhere wrong, but a live id belonging to a different person on a different instance would still collide. Collecting username and resolving on that would close it. 6 unit tests and 1 integration test that all fail against the parent commit, plus 2 controls that pass either way — a backup with no created_by_id key still restores, and a second operator still gets a 404. |
||
|
|
ca93d4cf44 |
fix(backup): never restore a toggle whose credential can't come with it (#2656)
A settings restore refuses to write anything credential-shaped, but wrote the switches that depend on those credentials like any other key. Restoring the two halves apart is not a partial restore, it is a downgrade. The sharp case is Prometheus. /api/v1/metrics is on PUBLIC_API_ROUTES and its only gate is `if token:`, so an empty or absent token means no authentication at all. prometheus_token matches the `token` hint and is refused; prometheus_enabled is an ordinary key and was written. On an instance that never enabled Prometheus there is no local token row, so overwrite-*off* alone was enough to publish the whole metrics body to anyone who could reach the port. The new integration test shows exactly that: 200 with a full unauthenticated body before, 404 after. Four more pairs are the same shape and break an integration rather than open one: ldap_enabled/ldap_bind_password, mqtt_enabled/mqtt_password, ha_enabled/ha_token (with an HA_TOKEN env arm, since get_homeassistant_settings prefers the environment over the row), and virtual_printer_enabled/virtual_printer_access_code — the last largely vestigial post-migration, included for consistency. A toggle is refused only when all five hold: the payload value is truthy, the backup carried a non-empty companion credential, that credential is denylisted, this instance has no usable value for it, and the toggle is not already on locally. The second condition is what keeps the rule honest — an anonymous MQTT broker and an anonymous LDAP bind are legitimate configs that pass empty credentials straight through, and without it both would be false positives. With it, the rule fires only when the restore would produce a config weaker than both the backup and the local instance. A present-but-blank prometheus_token row counts as unusable, since that is precisely the `if token:` hole. The rule needs the payload *and* local database state, which the old static _count_items could not see, so preview and restore now share one classifier: _plan_settings() runs a single SELECT over both halves of every candidate pair before anything enters the session, and returns the three refusal buckets. preview() takes the session the route already has. _is_skipped_setting_key is gone rather than having its docstring corrected as asked: a name is no longer enough to decide, so the union predicate had no caller left. Also implements the review's third ruling — the tally counts what the preview counted, and refusals live in the notes. Two `skipped += 1` increments are dropped (blocked, protected) and the companion refusal adds none; the value-is- None and overwrite-off skips stay, because they depend on the run's flags, which the preview cannot see. restored + skipped + failed now equals the item count the user was shown — off by three before. Behaviour change called out for review: test_credential_keys_are_never_restored and test_auth_settings_are_never_restored asserted skipped == 2 and 4; both are now 0, which is the point of the ruling. 16 new unit tests plus 2 integration tests. Nine of them are controls, because over-refusal is the real risk of this change — the anonymous-broker and anonymous-bind guards are load-bearing, not decoration. |
||
|
|
07244b6a43 |
fix(backup): don't count a stale selection, and say when archive links are dropped (#2656)
Two smaller restore-path fixes from the same review. The modal's footer counted `selected` raw while the checkboxes rendered `selected && isAvailable`. Switching commits keeps `selected` on purpose (it is only pruned once the new preview lands), so for as long as the new commit's preview was in flight — with the category list replaced by its spinner — the footer still read "2 selected" over an enabled Restore button, and clicking it restored the newly-picked commit with the previous commit's categories, none of which the user had seen an item count for. The count and the POST body now come from one `selectedCategories` memo gated on availability, exactly as the checkboxes are, so both go empty until the preview lands. Restoring Spool inventory without Print archives leaves archive_id_map empty, so every usage -> archive link is nulled even where the archive exists locally. It can't be resolved here (the archives payload isn't fetched for a category that wasn't selected) and a later archives-only restore won't repair it either, since the usage dedupe key doesn't include archive_id and those rows read as already-present. So it gets a note naming the remedy while the user can still redo the run with both categories ticked. Carries the rebuilt bundle (index-C2LOlVCR.js -> index-C16HJNOV.js). |
||
|
|
e7495dd41b |
fix(backup): reconnect the MQTT relay after restoring mqtt_* settings (#2656)
The relay reads its broker config once, when configure() is called — which is
why the settings PUT handler reconfigures it after writing those rows
(api/routes/settings.py:246). The restore wrote the rows and stopped there, so
the relay stayed on the pre-restore broker until the next backend restart while
the UI showed the restored values: the one way a settings restore could look
applied without being applied.
_restore_settings now reports the keys it actually wrote, and run_restore
reconfigures the relay from the committed rows when any of them is an mqtt_ one.
Three details worth keeping:
* it runs after the commit, because configure() drops the connection and
rebuilds it — not something to do on values a later failure could roll back;
* it is keyed on written, not merely present: a key skipped for overwrite=off
or by the credential blocklist must not trigger a reconnect;
* mqtt_password is never restorable, so configure() gets the row already in the
database and an unchanged broker keeps working.
A broker that refuses the new config is noted on the settings tally ("restart
Bambuddy") rather than failing the restore, matching the PUT handler's
best-effort handling of the same call.
|
||
|
|
3ba89c60a1 |
fix(backup): release the SQLite writer before the K-profile MQTT phase (#2656)
_apply ran archives and spools first, which autoflushes their INSERTs and so opens SQLite's single write transaction, then called _restore_kprofiles — which awaits get_kprofiles per printer per nozzle at timeout=5.0 with max_retries=3, i.e. up to ~15 s each against a printer that ignores the request. The commit only came afterwards, in run_restore. busy_timeout is 15 s (core/database.py:21), so a restore covering a couple of unresponsive printers held the writer past it and unrelated writes elsewhere in the app failed with "database is locked". Commit the database categories before the MQTT phase starts. The K-profile work is not in that transaction anyway — it leaves over MQTT — so the only thing lost is rolling those categories back when a K-profile send fails, and that rollback was never the right behaviour: extrusion_cali_set has already reached the printer by then, so undoing the database half would just make the two disagree. |
||
|
|
737258202b |
fix(backup): never restore the auth policy settings from a backup (#2656)
_collect_settings exports every Settings row minus two credential keys, so auth_enabled / advanced_auth_enabled / local_login_enabled / setup_completed all travel in a backup, and none of them are credential-shaped enough for _SECRET_KEY_HINTS to catch. Writing them back was the one part of a settings restore that changed who can reach the instance rather than how it behaves: * auth_enabled=false — from any backup taken before auth was turned on — disabled authentication. core.auth caches only the enabled=True result, on a 30 s TTL, precisely so staleness fails closed; set_auth_enabled pairs its write with invalidate_auth_enabled_cache(). The restore did neither, so it left the stored value the open one. * local_login_enabled=false walked straight past the #1589 refusals in update_settings (no enabled OIDC provider / no OIDC link on the caller), which exist to stop exactly that lockout. * /github-backup/restore is gated on GITHUB_RESTORE alone, so honouring these keys made that permission a way to rewrite auth config without SETTINGS_UPDATE. Flipping an existing row needed overwrite_existing, so the odds were lower than the severity. Both refusals now share _is_skipped_setting_key so the preview's item count still matches what a restore writes, and the skipped keys get their own note pointing at the auth UI rather than being folded in with the credential ones. |
||
|
|
cbe412607f |
fix(backup): resolve K-profile cali_idx live instead of reusing the backup's (#2656)
Restoring K-profiles addressed extrusion_cali_set at the cali_idx recorded in the backup. If that slot no longer existed on the printer the write was silently dropped and the restore still reported the profile restored. Not an edge case: Bambuddy's own K-profile editor is what re-keys the slot. On a single-nozzle printer an edit is delete-then-add, so any edit between backup and restore reproduces it. Found testing on an X1E. Backup held cali_idx 8151; an edit through the UI re-keyed the profile to 4606; the restore published cali_idx 8151, the printer ignored it, and the tally read "1 restored" while the k-value stayed put. Resending the identical payload with cali_idx 4606 applied, isolating the stale index as the cause. Fix mirrors the natural-key matching spools and archives already use, which the module docstring already promised but scoped to spool.id and print_archives.id. Before writing, read the live profiles for the nozzle and match on filament_id + setting_id, falling back to filament_id + name, then to the sole candidate for that filament. Send that profile's current cali_idx; where nothing matches send -1 so the printer adds a new profile instead of addressing a dead slot, and say so in the tally. A failed read degrades to adding rather than aborting. Also corrects the tally note. The printer does acknowledge extrusion_cali_set -- it answers with a result/reason pair -- so "published without acknowledgement" was false. It reports "fail" on writes that land, though, so the note now says the acknowledgement is unreliable rather than absent. Consuming result is left to a follow-up. Re-verified on the same X1E: perturbed to k=0.061, restored from the commit carrying the stale slot, payload went out with cali_idx 4606 and the printer read back 0.027. |
||
|
|
6a239314dc |
feat(backup): restore selected categories from a Git backup commit (#2656)
The Git backup feature was push-only: there was no equivalent of the local backup's Restore button, so recovering meant hand-downloading JSON files from the repository. This adds the read side. Providers gain list_commits / list_tree / fetch_files on the GitProviderBackend ABC. GitHub implements them against the Git Data API and Gitea/Forgejo inherit that unchanged; GitLab overrides for its own REST shape, including tree pagination and subgroup path encoding. fetch_files is batched so the path -> blob SHA lookup happens once per restore rather than once per file, and uses the blobs API rather than contents because contents silently inlines only the first 1 MB. The new GitHubRestoreService resolves HEAD to a concrete SHA up front, so a preview and the restore that follows act on the same commit even if a scheduled backup lands in between. Categories are applied archives -> spools -> settings -> kprofiles: archives first because spool usage history references archive_id, K-profiles last because they leave the database and publish over MQTT. Restores never reuse the backup's primary keys. spool.id and print_archives.id are bare autoincrement columns, so ids from an old backup very likely belong to unrelated rows today; rows are matched on natural keys (tag_uid, then tray_uuid, then a descriptive composite for spools; content_hash or filename plus started_at for archives), inserted without an explicit id, and an old_id -> new_id map rewrites the foreign keys in spool usage history. created_at is carried across on insert so restoring the same backup twice matches instead of duplicating. Dangling printer/project links are cleared and reported rather than failing the row. Settings restore re-applies the collector's credential denylist on the read side, plus a pattern guard, because a backup taken before that denylist existed can still contain secrets. Restored archives are metadata-only: the 3MF and thumbnail bytes are not in a Git backup and print_archives.file_path is NOT NULL, so inserted rows get an empty path and the UI says so. Backup and restore take a mutex against each other; both write the same tables and talk to the same printers. Restores are logged as GitHubBackupLog rows with trigger="restore", which needs no migration and surfaces them in the existing History card. Cloud profiles are deliberately not a restore category. The collector never actually writes cloud_profiles/*.json - it reads a "setting" list key the Bambu Cloud API does not return - and the preset list it would write carries no setting payload. Filed separately. Permission github:restore already existed and is granted to Administrators, so no permission changes were needed. Tests: 125 new backend tests (provider reads across all four providers, the per-category appliers, the API endpoints) and 13 frontend tests. Full suites pass with no regressions; the 35 backend failures on Windows are byte-identical with and without this branch. |
||
|
|
33ab5f1ead |
Add temperatures to the streaming overlay and a URL builder (#1422)
The overlay at /overlay/{printer} draws live print data over a
full-screen camera view for OBS, a wall display or any browser source.
It has been tunable since it shipped -- which fields, what size, what
frame rate -- but only through query parameters documented in the wiki,
and temperatures were not among the fields on offer. The request asked
for temperatures first and for the field set to be selectable in the web
UI; this addresses both.
Nozzle, bed and chamber readings join the list. The target is drawn only
while the heater is still climbing, so a settled hotend reads "220°C"
for the rest of the print instead of the noisier "220 / 220°C" -- 219.6
against a target of 220 rounds to the same number, and repeating it says
nothing. Both nozzles appear on a dual-nozzle machine. They are drawn
whether or not a print is running, because a preheating printer is
exactly when they are worth watching, and each reading appears only when
the printer genuinely reports one: chamber temperature stays absent on
P1 and A1 models, which publish a chamber_temper with no sensor behind
it, so the overlay never puts a measurement on screen that does not
exist. Labels reuse the heater chart's strings rather than inventing a
second vocabulary for the same three things.
The feed sends an allow-list rather than the temperatures dict. That
dict doubles as the MQTT client's working memory -- derived heater flags
and private target-set timestamps live alongside the readings -- and an
overlay token is a narrower grant than a login, so it gets exactly what
the overlay draws and does not pick up fields as the dict grows. The
same chamber-sensor gate the full status payload already applies is
applied here. The integration test that asserts the payload's exact key
set, which exists to catch that surface widening silently, is updated
deliberately.
Temperatures are not in the default field set, so an overlay URL already
pasted into a scene renders identically after upgrading.
Settings -> API Keys -> Streaming Overlay now builds the URL: printer,
field checkboxes, size, frame rate, camera toggle, an optional token,
and a copy button. It persists nothing and calls nothing new -- the URL
is the configuration, which keeps a scene reproducible by copy-paste and
lets two displays show different fields off one token. Fields are
emitted in the overlay's own top-to-bottom order rather than click
order, and parameters left at their default are omitted, so the same
selection always produces the same URL. The preview alongside it stays
off until asked for: an always-live iframe would hold a subscriber on
the printer's single camera connection for as long as the settings tab
stayed open.
The preview needed one narrow security-header change. Every SPA route
sent frame-ancestors 'none', which is stricter than the SAMEORIGIN in
X-Frame-Options beside it and refuses even a same-origin frame, so the
preview showed Firefox's "another site has embedded it" page instead of
the overlay. The overlay path now sends 'self', mirroring /gcode-viewer,
which admits a framer only on this origin -- Bambuddy's own UI. Every
other path keeps 'none', and embedding the overlay from another host
still requires TRUSTED_FRAME_ORIGINS.
|
||
|
|
48c231d8ce |
Stop a drying cycle reporting itself finished a minute in (#2759)
Starting the dryer on an AMS 2 Pro holding two PETG and two PLA spools and picking PLA showed "PLA @ 45°C" for about a minute and then switched to "PETG @ 65°C" for the remaining twelve hours. Bambu never echoes back which filament or temperature a cycle is running, so the badge reads the target cached when the command went out, and that cache had been dropped. Between accepting the command and settling its countdown the firmware publishes one update with the remaining time at zero while the unit is still in its Checking phase -- the reporter's log has 720, then 0, then 719, and four seconds later the same unit's info hex decodes to dry_status 2, Drying. The falling-edge detector read that zero as the cycle ending. Losing the cached target left the badge to guess from the first loaded slot, which was PETG, and its RFID-recommended 65°C. The same false ending fired on_drying_complete, so anyone with smart-plug auto-off-after-drying switched on had power scheduled to cut one minute into a twelve-hour dry; the reporter had it off, which is the only reason this reads as a cosmetic bug. A remaining time of zero now ends a cycle only when the unit is not also reporting an active phase. dry_status comes from the same info hex already parsed a few lines above, so this costs nothing to check. Stopping and Error are deliberately not treated as active -- those should end it -- and a unit that reports no phase at all still ends its cycles, so the gate can only ever suppress on positive evidence that the cycle is live. A suppressed edge leaves the remembered dry_time alone, exactly as the #1462 absent-value skip does, so the push that really ends the cycle still sees a non-zero previous. The fallback guess is tightened to match. On a mixed unit the first tray is evidence of nothing, and naming a temperature the cycle is not using is worse than naming none, so it now answers only when every loaded spool is the same filament and otherwise leaves the badge showing the countdown alone. Both the websocket and REST status builders carried their own copy of that loop; they now share one helper, which also takes the temperature from the first slot that carries an RFID one rather than giving up when slot 1 holds a third-party spool. |
||
|
|
ebc73e46ba |
Say when AMS drying was running and a print never started (#2758)
Dispatching to an X2D with two AMS units mid-drying failed silently: the file uploaded, the printer accepted it and stayed idle. The watchdog waits for an active state or HMS_MQTT_VERIFY_FAILED, and a drying refusal is neither, so it timed out, re-uploaded the whole 3MF twice more, and closed with advice about the printer screen and the SD card. Studio, asked directly, said it could not start the job because of the drying. Latch the AMS dry_time telemetry across both watchdog phases and name the units in the give-up message, plus an INFO log on every failed window so the correlation reaches a support bundle from the first attempt. Detection only, no gate. These models support drying CONTINUING through a print (supports_drying_while_printing covers X2D from 01.01.00.00), so drying is not incompatible with printing and stopping it before every dispatch would tear down cycles the hardware is happy to run. One of the two units was also drying without its external PSU, which would make this a power budget problem at start-of-print calibration rather than a drying one -- dry_sf_reason 1/8 exist for exactly that. The message names both possibilities rather than asserting one. Also correct _sync_drying_state's docstring, which claimed to adopt drying it did not start; it only prunes. Behaviour unchanged -- populating it would let the scheduler stop a cycle the user started by hand. |
||
|
|
28a6ca6f4d |
Add CAP_NET_BIND_SERVICE everywhere the service is defined (#2549)
The Virtual Printer binds 990 and 322, below 1024, which a service running
as a normal user may not do without CAP_NET_BIND_SERVICE. Without it the
rest of Bambuddy works and only the VP is dead -- sockets never open, the
slicer never finds the printer, and the sole trace is one journal line.
|
||
|
|
4af782cc2f |
Report a refused AMS filament setting instead of discarding it (#2756)
Configuring a slot publishes ams_filament_setting and the printer answers with a verdict. The answer was received and dropped at DEBUG, so a refusal left no trace at the level support bundles are collected at: the reporter saw six Configure Slot attempts on an X1C all return success, all read back by the #2582 verification as holding the previous profile, and no record of what the printer said about any of them. Promote a non-success response to INFO with result, reason, ams_id and tray_id. Refusals only -- unlike extrusion_cali_set (#2718) and ams_filament_drying (#1447) this command is not rare, since every spool assignment and K-profile re-apply sends one, so promoting each ack would bury the interesting line. The developer-mode probe is excluded: it sends this same command to the external slot expecting a refusal on P1 firmware, so promoting it would put an alarming line in every P1 bundle on every reconnect. Matched by sequence id, which user commands cannot collide with -- they publish a hardcoded "0". Diagnostics only; no change to which commands are sent or how they are built. |
||
|
|
689f5276e4 |
Show the compose directory in the Docker update command (#2664)
The printed command only works from the directory holding the compose file, which is the thing the user came to the page not knowing. Adds a copy button, a saved Compose directory setting, BAMBUDDY_COMPOSE_DIR, and best-effort detection from a bind mount's host path. Compose records the directory on every container it creates, but reading that label needs the Docker socket mounted in — root-equivalent access for a convenience string. The mountinfo guess is a prefill only: its root field is relative to the mounted device, so a compose dir on its own mount loses that prefix, and nothing in the container can detect it. The field is restricted to path characters. It is the one setting whose purpose is to be pasted into a root shell, so "/opt/bambuddy; rm -rf /" would otherwise render as a plausible update command. |
||
|
|
a08d3e62f3 |
Show the Print Log's per-run cost and energy, and let users pick columns (#2636)
The list and update endpoints serialised field by field and never named cost / energy_kwh / energy_cost, so values Bambuddy had been recording all along went out as nulls. Both now validate from the ORM row, which removes the chance to omit a field rather than patching the three that were missing. Adds a Filament Used column plus a Columns picker for Cost, Energy, Energy Cost and Finished, persisted per browser. Also fixes the log view being unreachable with zero archives: the empty state ran before the view check, hiding a log that outlives the archives it refers to. --- Sort the Print Log by any column (#2636) Adds sort_by / sort_dir to the print-log endpoint, driven by clickable column headers. Server-side because paging is: ordering the rows the client holds would sort one page rather than the log. Empty values are held last in both directions — Postgres sorts NULLs high and SQLite low, so the same click would otherwise open on blanks on one backend and values on the other. id DESC breaks ties so paging through a low-cardinality sort can't repeat or skip a row. |
||
|
|
e95c42c021 |
Add auto-orient and auto-arrange to server-side slicing (#2548)
Both are per-slice checkboxes, off by default, forwarded as the sidecar's orient / arrange form fields. An unticked box is sent by omission: the sidecar treats any present value as truthy, so a literal "false" would have arranged every slice. Arrange unions with the #1493 cross-class decision rather than replacing it, and the per-plate slice-all loop is now keyed on the arrange flag itself — the project-wide collapse belongs to --arrange, not to the cross-class case. The loop also covers the embedded-settings path, whose crash-retry is suppressed there since a single --slice 0 retry would return one consolidated plate. |
||
|
|
a9b57ccd3c |
Add variant-group endpoints and cross-model queue creation (#671, #2570)
Adds /library/variant-groups for declaring that several sliced files are the same job for different printers, and a variants payload on queue creation that turns such a set into one queue item with a candidate per file. The candidate set is validated as a set: one file per printer model, each file sliced for the model it is offered as, and at least one model with an active printer. A cross-model item deliberately holds no file of its own, because print_queue.library_file_id is ON DELETE CASCADE and would destroy the whole job when a single alternative is deleted. Fixes internal printer-model codes never being resolved on queue create and update: normalize_printer_model returns unknown input unchanged, so the or-chain never reached the code map and a "C13" target matched no printer and waited forever. Skips candidates whose file is trashed or missing. Library deletes are soft, and SQLite runs with PRAGMA foreign_keys off, so neither case is covered by the schema; the hard-delete paths now also drop the rows. Adds library_files.variant_target_model so a user can say which printer a file without slicer metadata is for, kept out of file_metadata so the assertion is never mistaken for parsed data. |
||
|
|
752e345d1a |
Add cross-model variant resolution to the queue scheduler (#671)
Adds print_queue_variants: the candidate files a queue item may run, each with its own model, plate, AMS mapping and nozzle mapping. The scheduler walks them in priority order and takes the first whose model has an idle printer, then folds that candidate onto the queue row before the selection commit — so upload, archive creation, print history and reprint keep seeing an ordinary single-file item. Candidates are ordered least-attempted first, so a printer that accepts the file and never starts hands the job to the alternative on the next lap instead of spending the item's whole retry budget on the machine that is wedged. The item-level DISPATCH_MAX_ATTEMPTS bound is unchanged. An item whose candidate files have all been deleted is held pending with an actionable reason rather than failing deep in the upload, and waiting notifications name the job and every model it is waiting on. |
||
|
|
da07c5884b |
Add variant-group data model for cross-model queue alternatives (#671)
Adds file_variant_groups plus variant_group_id / variant_position on library_files, so a set of files that are the same job sliced for different printers can be resolved to whichever printer frees up first. Backfills groups from the sliced_from_library_file_id provenance that slice_and_persist and the pipeline runner have been writing into file_metadata since they shipped, and which nothing has ever read. Only sources with two or more children carrying distinct sliced_for_model values are grouped: a single candidate is not a choice, and two slices for the same printer give the resolver no basis to prefer one. |
||
|
|
b04664c64a |
Raise the chamber-temperature ceiling from 60 to 65 C
Every field that takes a chamber target stopped at 60: the per-filament chamber map and per-print override in Preheat & Heat Soak, the chamber quick-select presets, and the printer-card chamber control. 60 is the X1E's ceiling and the X1E was the only heated-chamber model when that limit was written; the H2 series and X2D heat to 65, so the top of their range was unreachable. The ceiling now lives in one constant per side (MAX_CHAMBER_TEMP_C in backend/app/utils/printer_models.py and frontend/src/utils/printer.ts) rather than as a literal at each call site. X1E firmware clamps a higher request to its own maximum, so a shared ceiling is safe. Also fixes a live bug at PrintersPage.tsx:7985: parsePresetTriple was bounded to 60 there, and it rejects the whole triple on any out-of-range entry, so a saved 65 preset would have silently reverted the printer card to the defaults while Settings still showed 65. |
||
|
|
e52b73e21f |
Stop the Bambu Cloud TOTP tests reaching the network
verify_totp fetches a CSRF token from the bambulab.com web origin before posting the code (#2696) and returns early when it cannot get one. These tests patch only post, so the pre-flight GET went out for real: it succeeded wherever bambulab.com was reachable and returned a tokenless 403 on a CI runner, where six tests then asserted on a post that never happened. Stub the handshake for the module. It is covered end to end, no-token path included, in tests/unit/test_cloud_totp_csrf.py. |
||
|
|
bbbb9d35c7 |
Bound the scheme repetition in the log credential-redaction pattern. As an
unbounded repetition the match was quadratic in the subject length: on a run of scheme-legal characters the engine restarted at every offset and consumed to the end before failing to find "://". ffmpeg echoes the configured camera URL into its stderr and the whole blob reaches the pattern before any truncation, so the subject length is attacker-influenced. |
||
|
|
bf424493ba |
Suppress Bandit B104 false positive in the SSRF guard tests
The parametrize list feeds "0.0.0.0" to TasmotaService._validate_ip and asserts it is refused. B104 matches the literal wherever it occurs and cannot distinguish a rejection fixture from a bind address. Split the list across lines so the token carries its own nosec with the reason; the single-line form was 117 chars against a 120 limit. |
||
|
|
77135aaf8f |
Fix unawaited coroutine warning in printer offline notification tests
on_printer_status_change builds reconcile_stale_active_prints(...) as a call argument, so the coroutine is constructed even when the spawn helper is mocked out. A bare MagicMock retained it in call_args and it finalised unawaited during a later test's GC, surfacing as a PytestUnraisableExceptionWarning attributed to test_printer_sensor_history. Patch spawn_background_task with a side_effect that closes the coroutine, and correct the _state() docstring, which claimed state="IDLE" kept the reconcile-edge branch quiescent when it does the opposite. |
||
|
|
aef4f3a3e9 |
fix(oidc): strip the required BAMBUDDY_OIDC_* values and register the local-login bypass
A Kubernetes Secret written as a block scalar carries a trailing newline, and the schema bounds the four required variables by max_length only, so an unstripped issuer_url was stored and enabled and then raised httpx.InvalidURL on the first click of the SSO button -- the authorize-time failure the all-or-nothing rule exists to prevent. Whitespace-only values got through the same way, contradicting the reader's own "an empty required var counts as unset". The optional variables have always treated blank as unset; the required ones now do too. Also registers BAMBUDDY_LOCAL_LOGIN (#1589) in the typo guard, which logged "possible typo" for it on every boot while listing every BAMBUDDY_OIDC_* variable as legitimate. |
||
|
|
8b46006644 | Merge branch 'dev' into feature/oidc-env-config | ||
|
|
9e783fbad6 |
fix(oidc): keep the local-login bypass lenient under strict env_bool
Promoting env_bool to strict rejection made BAMBUDDY_LOCAL_LOGIN=on raise EnvOIDCConfigError uncaught on the login/forgot-password path -- a 500 on the exact recovery endpoint the bypass exists to keep open. env_bool gains a strict flag (default True for the startup OIDC reader); the local-login caller opts out so an unrecognized value falls back to "off" instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016q8EAf9Rj7ZHL92sPnXYxy |
||
|
|
77c9bdd694 |
fix(oidc): reject an unrecognized boolean instead of guessing
_env_bool returned the default for anything outside {true,1,yes}, so
BAMBUDDY_OIDC_REQUIRE_EMAIL_VERIFIED=on silently read as OFF and
BAMBUDDY_OIDC_ENABLED=on silently disabled the provider -- the exact
opposite of what .env.example claimed. Unrecognized values now raise
EnvOIDCConfigError, caught in _apply_env_oidc_provider the same way a
bad DEFAULT_GROUP or a ValidationError already is: logged and left
running, never released on a typo.
Also promotes _env_bool to env_bool now that it has a call site in
auth.py, and corrects the boolean-parsing sentence in .env.example.
|
||
|
|
c547505c64 |
fix(oidc): treat a blank optional env var as unset, not a refusal
BAMBUDDY_OIDC_SCOPES, _EMAIL_CLAIM and _ICON_URL fell back to their default only when the key was absent, so `BAMBUDDY_OIDC_ICON_URL=` in a compose file (as .env.example ships it, commented) reached the schema validator as an empty string and got the whole provider refused. default_group already treated blank as unset; these three now follow the same rule. |
||
|
|
18938a10ee |
fix(kprofiles): stop reporting rejected K-profile writes as saved
Saving a K-profile was fire-and-forget. set_kprofiles_batch published and returned True, and the printer's extrusion_cali_set answer was logged at DEBUG and dropped, so a write the printer refused was reported to the user as saved (#2718, reporter @jmoore-skild). The reason it could not simply be gated on: the answer itself was wrong. Single-nozzle firmware returned result:"fail" with reason:"invalid tray_id" on writes that demonstrably applied. Measured against an X1C and an H2D over MQTT, the cause is the tray_id:-1 Bambuddy itself put in the payload. Sending three otherwise identical writes isolated it: tray_id:-1 fails, tray_id:0 succeeds, and cali_idx:-1 is accepted either way, so only that one field is at fault. The H2D ignores the value entirely; the X1C validates it, complains, and applies the write anyway. BambuStudio always sends a real tray_id and defaults it to 0 for a manually entered profile. With tray_id:0 the acknowledgement is honest, and the printer echoes back the sequence_id we sent -- confirmed for extrusion_cali_get, _set and _del on both printer classes -- so it can be matched to the write that caused it. Writes now return their sequence_id and the routes await the verdict, turning a real failure into an error that carries the printer's own reason. A printer that stays silent is still treated as success: no answer is not evidence of refusal, and firmware that never answers must not turn every save into an error. Raises the ack to INFO. It sat at DEBUG, so the one line that explains a failed save was absent from every support bundle -- the same reasoning that put ams_filament_drying at INFO for #1447. Also fixes extrusion_cali_set building its payload from str(self._sequence_id) without incrementing first, reusing the previous command's id. Harmless while nothing correlated on it, fatal now that the write path does. Adds supports_nozzle_flow_type() for the Standard / High Flow choice, which the K-Profiles UI previously showed as "Not reported by printer" -- not a value anyone can save. Most printers omit the nozzle identity from their calibration table entirely, and the slicer treats that as Standard rather than unknown; Bambuddy now does the same and keeps the choice editable. The field is hidden only where the model ships a single nozzle variant, using the slicer's own rule (len(nozzle_volume) // len(nozzle_diameter) > 1 over the machine preset) evaluated across every bundled Bambu profile. That puts only A1, A1 Mini and A2L on the hidden side -- it is not the single- versus-dual-nozzle split, since P1P, P1S, P2S, X1, X1C, X1E and H2S are all single-nozzle and all carry two variants. Editing a profile also no longer writes back an empty nozzle_id. Wiki records that on printers which omit the field the chosen flow type is discarded by the firmware and reads back as Standard, in Bambu Studio as well, so it does not get filed as a bug again. |
||
|
|
a35ba8fa5f |
fix(kprofiles): read the nozzle diameter the printer actually sent (issue #1748)
Every K-profile came back as 0.4mm on printers running any other nozzle (#1748, reporters @Liquidmasl and @jmoore-skild). The printer puts nozzle_diameter on the extrusion_cali_get envelope only; the per-filament entries carry setting_id, filament_id, name, k_value, n_coef and cali_idx, and nothing else. The parser read the field per entry with a hardcoded "0.4" fallback, so the fallback fired on every profile of every response. The envelope value was already in scope, read into response_nozzle and used only to match the request. This never reproduced on H2D because that firmware does include the field per entry. Both construction sites are in the same handler, so the code path is shared; what differs is the payload, and every single-nozzle model omits it. The display was the least of it. Editing is delete-and-re-add on single-nozzle printers, and the dialog rebuilt nozzle_id and nozzle_diameter from its own greyed-out selects, so saving an untouched 0.6mm profile rewrote it on the printer as HH00-0.4. Deleting aimed extrusion_cali_del at the wrong nozzle the same way. Both now pass through what the printer reported. The cali_idx cascade in inventory.py, spoolman_inventory.py and spoolman.py matches on nozzle_diameter, so on a 0.6 or 0.8 nozzle it never found the printer-side entry and the assignment silently failed to stick -- that is the "cannot auto-map a K-profile" half of the report, fixed at the source without touching those three call sites. nozzle_id has no source in the payload at all, and state.nozzles carries material (hardened_steel), not flow, so it cannot honestly produce HH/HS. Rather than keep inventing one, the UI now says the printer did not report it: the card shows the diameter alone, the dialog shows "Not reported by printer", and the High Flow / Standard filter is hidden instead of being offered as a control that can only ever empty the list. Import stops stamping HH00 on profiles whose source reported none. Also correlates K-profile requests by sequence_id. Responses were matched by nozzle diameter through a single shared expectation slot, so a second request overwrote the first's and the first's valid answer was discarded as a mismatch -- the "Failed to get K-profiles after 3 attempts" in the same logs, with the printer having answered correctly both times. Pending state is now one entry per request, keyed by the id we already send, with the nozzle match kept as a fallback for firmware that does not echo it back. Fixes the flow-type select naming a new profile with the opposite label, which contradicted the identical expression 44 lines above it. |
||
|
|
21d61c3535 | Merge branch 'dev' into feature/oidc-env-config | ||
|
|
455a9e4ba7 |
fix(backup): collect cloud profiles from every connected account (#2717)
Enabling Cloud Profiles for a Git backup produced nothing, and said it had
worked. Two independent faults, either one sufficient.
The collector looked for a "setting" list. The Bambu Cloud listing endpoint
is keyed by preset type instead, each key holding private and public arrays,
so the loop body never executed once — and the entries carry no type of
their own either, which routes/cloud.py already knew: it takes the type from
the outer key and maps Bambu's "print" to process. Two bugs on one line.
It also asked build_authenticated_cloud for the credential store used when
authentication is disabled. With auth on, tokens live on User rows, so the
collector returned at "Cloud not authenticated" before ever reaching the bad
key. Every multi-user install was collecting from zero accounts.
Neither failure surfaced. backup_metadata.json recorded the configured flag
rather than the outcome, so it claimed cloud_profiles: true on runs that
wrote nothing, and the log read "Collected cloud profiles: 0 filament, 0
printer, 0 process" at INFO — which is exactly what a successful backup of
an empty account looks like.
Cloud profiles now come from every connected account across both clouds. The
toggle predates Orca Cloud entirely, and Orca has the same three preset
types, so both are collected and grouped the same way:
cloud_profiles/bambu/user-3/{filament,printer,process}.json
cloud_profiles/orca/user-3/{filament,printer,process}.json
Accounts are keyed by Bambuddy user id, "global" when auth is off. Never by
email: a backup repository can be public, and the Bambu listing's user_id is
dropped for the same reason. Both credential stores are read on every run,
because a Settings row survives someone enabling auth later and dropping it
would silently stop backing that account up.
Bambu costs one get_setting_detail per private preset. The listing is
metadata only, and without base_id and setting the backup is a list of names
that create_setting cannot rebuild from. Public presets are skipped — Bambu's
bundled catalogue is the same hundreds of entries for everyone, always
re-downloadable, not recreatable under your account, and would rewrite the
repository on every run. Orca needs no second call; its sync-pull carries
each profile's content inline. Where the Orca route drops a profile whose
content.type it cannot map, the backup writes it to other.json instead:
silently omitting a profile because Orca added a type is the same class of
bug as this one.
Failures are contained per account and per preset, and counted rather than
swallowed. A partial backup that looks complete is how this stayed invisible.
The metadata now reports what was collected, per cloud and per account, and a
run that collects nothing while the category is enabled warns with the reason
instead of an INFO line that reads like success.
The checkbox gated on the viewer's own Bambu sign-in, which is not the same
question as whether there is anything to back up — with auth enabled the
accounts belong to individual users, and an administrator who never signed
in personally saw the category disabled with plenty in scope. It now gates
on the total across both clouds and shows the counts. That comes from its
own endpoint rather than a field on /config, since /config answers null
until the first save and would disable the toggle during the very setup it
belongs to. Counts only, never identities.
One deliberate restraint. _build_authenticated_service clears stored
credentials when a refresh is rejected, which is right for a route — the
user is on the page and can pair again — and wrong for a scheduled job.
Orca reports every rejection with one composite reason ("unknown, expired,
revoked, or already used"), so a genuine revocation cannot be told apart
from a lost token-rotation race, and acting destructively on a signal that
cannot be disambiguated is the #2562 mistake in a different cloud. It also
gains nothing: the Profiles route hits the same failure and clears it then,
with the user present. Background callers now pass clear_on_auth_failure=
False and skip the account. A successful refresh is still persisted either
way — by that point the old token is consumed, so dropping the new pair
would break a working pairing for real.
Restore is not part of this. Nothing reads cloud_profiles/* yet; the format
carries base_id/setting for Bambu and content for Orca so that it can.
|
||
|
|
4f2c073a34 |
fix(vp): gate the slicer's AMS pick behind the toggle and scope its badges (#2700)
Round-3 review of the "Save AMS mapping" PR. The queue item's ams_mapping was set unconditionally, on the reasoning that honouring the slicer's own pick is a correctness fix rather than a feature. It is both. Storing a resolved mapping makes _ensure_ams_mapping return early, so _compute_ams_mapping_for_printer never runs — and that function is where prefer_lowest_filament lives, along with the AMS-filament-backup gate that qualifies it (#1766), the inventory-remain overrides, and the per-slot force-colour overrides. Every existing queue-mode VP pointed at a printer would have quietly lost all of it on upgrade, without a setting to turn it back on. So save_ams_mapping now gates the queue item too, not just the archive persistence. Off is exactly the old behaviour. The correctness case the PR was written for — two spools of the same red PLA, and the slot the user picked in the slicer thrown away — is still fixed, for anyone who asks for it. Force color match wins over it when both are on. Its only effect on a fixed-printer item is the filament_overrides written onto the queue item, and those are read inside the function a stored mapping skips, so the two toggles sitting next to each other on the same card silently cancelled. The dispatch now matches strictly, as asked, while the slicer's pick is still saved onto the archive — that is what the toggle's name promises, and a later reprint is a separate decision from this print. The queue-add fallback applies the same rule to a request that carries force-colour overrides. A mapping shorter than a plate's highest slot id cannot address that plate's own slots, and _ensure_ams_mapping would have kept it anyway, since it only rejects an all-unresolved one. Each plate now checks the length it needs and falls back to a computed mapping if the array does not reach. Bambu Studio sends a file-global array, so this normally never fires; it also means a multi-plate Send All degrades safely if that ever stops being true. The badges claimed more than they delivered. Both rendered whenever a saved mapping existed, ignoring which printer it belonged to, while the tooltips promised the reprint would reuse those exact spools — true only on the printer the trays were resolved against. The queue row's flag is now computed against that row's own printer, which is precisely when dispatch reuses the mapping, and the archive card names the printer instead of implying any of them will do. It hides itself when that printer no longer exists. Retranslated in all 13 locales. Frontend tests, which the PR had none of. The printer-scoping rule is now a pure function rather than an inline expression, covered for the mismatched printer, the no-printer-selected case that would otherwise compare undefined against undefined, and malformed extra_data. The toggle's undo bookkeeping is covered for unresolved slots, short mappings, and hand-made picks — preserved when the toggle never wrote that slot, replaced when it did, which is behaviour worth pinning either way. Also reverts all three queue-mode switches when a save fails, not just the new one; without it the card shows a setting the server rejected. |
||
|
|
4b4cb18a64 | Merge branch 'dev' into feature/save-ams-mapping-toggle | ||
|
|
c457cf54bf |
feat(support): record process memory, threads and children in bundles (#2734)
A bundle described everything except the process it runs in. So a report of memory climbing over days until the OOM killer fires arrives with no way to act on it: the numbers that name the mechanism only exist while it is happening, and by the time anyone asks, the container has been restarted. The new `process` section carries what actually separates the candidates. Resident against virtual memory: 650MB RSS with 12.9GB VMS is address space — thread stacks or allocator arenas — not a heap full of live data, and that reading is the opposite of the one the reporter drew from the same figures. Thread count and child-process count then split those two apart, and a census of live objects by type names what a growing heap is filling up with. Open files, sockets and uptime round it out. Three constraints worth keeping: The heap census is skipped above 2GB. gc.get_objects() materialises every tracked object, so it costs most on exactly the process that can least afford it — a bundle generated to diagnose runaway memory must not be the allocation that tips the host over. Everything else is still collected, and the skip is recorded with its reason rather than silently omitted. Children are recorded by executable name only. An ffmpeg command line carries the camera URL, and with it the camera's password. Collection runs off the event loop and every metric is independently best-effort. psutil raises on hardened kernels and in restricted containers, and the bundle is how someone reports a problem in the first place — it has to be produced even when half the numbers are unavailable. This does not fix #2734, and nothing here should be read as having found its cause. The bundle's own evidence contradicts both proposed causes: the orphan janitor ran 7 times in 26 days over 725 stream-ends and killed no orphaned ffmpeg, which is not the #776 signature; and the 5 "database is locked" errors all fall between two OOM kills, making them a symptom of the memory pressure rather than a source of it. |
||
|
|
ce3e59884a |
fix(slicer): bound slices by silence, not by total slicing time (#2730)
A heavy MakerWorld model — one Bambu Studio also takes a long time over —
failed after five minutes with "Slicer sidecar unreachable". The sidecar
was reachable the whole time and still slicing when we hung up on it.
SlicerApiService carried a hardcoded 300s timeout, passed to httpx as a
bare float so it covered connect, read, write and pool alike. On a single
long request that is not a health check, it is a cap on how long a model is
allowed to take. And because httpx.ReadTimeout subclasses RequestError,
expiry landed in the same handler as a refused connection and was reported
as an unreachable sidecar — so the reporter went and updated their sidecar
container, which was never the problem.
The information to do better was already being collected. _poll_progress
polls /slice/progress/{id} once a second alongside the blocking POST to
drive the live progress toast, so at minute five Bambuddy had fresh
evidence the slicer was working. It killed the request anyway.
So the read timeout comes off the HTTP call and the poller supervises
instead: the deadline moves forward on every progress update, and only
genuine silence ends the wait. A model that keeps reporting runs to
completion however long it takes. Connect and pool keep short timeouts —
a sidecar that will not accept a connection is unreachable and should
still say so quickly.
Only a *changed* progress payload counts as alive. The sidecar re-serves
its last snapshot on every poll, so counting repeats would leave the
watchdog unable to detect a stall at all.
The window is floored at three poll intervals: liveness can only be
observed as fast as the poller ticks, so anything shorter would expire in
the gap between two polls and fail every slice instantly.
New setting slicer_stall_timeout_minutes (Settings > Workflow > Slicer),
default 15, range 1-240, alongside the sidecar URL and gated on
use_slicer_api like its neighbours. Sidecars too old to report progress
have no liveness signal, so for those the same number bounds total elapsed
time — the old behaviour, configurable and no longer 300s flat. The
message says which case applies and where to change it.
SlicerTimeoutError is its own type and maps to 504, not 502: the sidecar
answered throughout, we stopped waiting. Connection failures keep
SlicerApiUnavailableError. The preview slice path gets the same treatment.
|
||
|
|
3abab1fd45 |
fix(printers): recover MQTT sessions that stopped reconnecting (#2732)
The reporter's printer lost its session to a keep-alive timeout at 02:19 and did not come back until 11:24 — nine hours offline, with the web UI open throughout. check_staleness() was never going to catch it. Its first line is `if self.state.connected and self.is_stale()`, so it only ever handles the half-broken session that is still connected but has gone quiet. This client had connected=False from 02:19:42 (the offline notification fired a minute later), so every call returned immediately, and paho's own retry was the only thing left watching. When that stopped making progress nothing noticed. Adds a sweep every 60s that rebuilds a client when all four hold: it is disconnected, it had a working session before, it has been silent for five minutes, and its MQTT port still answers. The port check is what keeps this from becoming a nuisance — a switched-off printer is left to paho, so a farm powering down overnight causes no client churn and no log spam. The five-minute grace sits well past the 60s stale timeout and the 30s max reconnect backoff, so a session recovering on its own is never interrupted. The rebuild goes through force_reconnect_stale_session from async context, which takes the hard-reset path: fresh client_id and paho's QoS 1 queue dropped, so a project_file left unacked on the dead session cannot replay into the new one and trip 0500_4003 (#1136). Rate-limited per printer, cooldown cleared when the printer returns, and the sweep continues past a client that throws rather than abandoning the rest of the farm. The log line names how long the printer was gone and the last connect error, so a session that dies repeatedly leaves a trail. check_port gains a public alias in printer_diagnostic rather than having the watchdog reach for the private name. Also corrects the Developer Mode path added in the previous commit: the wiki documents it under Settings > Network, not Settings > General. The menu path is dropped from the translated string entirely, since it varies by model and firmware and the wiki carries the detail. |
||
|
|
5e2b7b53e6 |
fix(printers): surface the printer's own "command verification failed"
A P1S on firmware 01.10.00.00 rejected every control command and said so: HMS 0500-0500-0001-0007, "MQTT command verification failed". Bambuddy received that, dropped it, and reported a healthy printer instead. The frontend filtered it out. This code's meaning lives in attr's low half (0500) and code's high half (0001), both of which the MMMM_EEEE short form discards, so it collapsed to "0500_0007" — no catalog entry, no firmware actions, and filterKnownHMSErrors drops uncatalogued action-less errors. Catalog lookups now try full_code first, in both the description and the filter, and errors matched that way display the four-group code the printer's own screen shows. The remedy line is ours, not Bambu's: their wiki says to update Studio or Handy, which does not apply to a print sent from Bambuddy. The developer-mode probe made it worse. It read anything that was not an explicit refusal as confirmation, and this firmware answers the probe with an empty result while refusing everything else — so an inference drawn from a non-answer became "developer_mode: pass" in the support bundle of a printer that had not accepted a command all day. The probe now has three outcomes: explicit success enables, explicit verify-failure disables, anything else stays unknown and the diagnostic reports skip. The HMS is authoritative over that inference in both directions. It forces developer_mode False when present, and clears back to unknown when the printer stops reporting it, so enabling Developer Mode and restarting the printer is picked up without restarting Bambuddy. Dispatch no longer treats a refusal as a wedge. The watchdog latches the HMS across both phases and fails the item on the first attempt naming the code and the fix, rather than spending three uploads and 270s a lap to arrive at a message about SD cards. The check runs after the active-state exit in both phases, so a lingering HMS can never abort a print that is visibly running. Also: the "wrong or mis-cased serial number" hint no longer fires in the moment after a reconnect. _report_messages_since_connect is reset by _on_connect, so a reconnect landing microseconds before the staleness check leaves it at 0 for reasons that have nothing to do with the serial — this reporter's healthy printer was told to go check its serial 1 ms after reconnecting. |
||
|
|
11dc612bc4 |
feat(obico): authenticate to a token-protected ML API (#2733)
Obico's ml_api container takes an optional ML_API_TOKEN environment variable.
With it set, ml_api/auth.py answers a bare 401 to any request whose
Authorization header isn't "Bearer <token>"; with it unset it ignores the
header entirely. Bambuddy never sent one, so pointing it at a protected server
meant deleting the token there — which the reporter had set for their Home
Assistant integration and did not want to undo.
Settings -> Failure Detection gains an ML API Token field. When it is empty no
header is sent, so an unconfigured install's request stays byte-identical to
what shipped before the setting existed.
This failed in the worst possible way, and that is the more important half of
the change. Obico decorates /p/ with token_required but leaves /hc/ open. Test
Connection pinged /hc/, so it reported success against a server that was
rejecting every real detection call, the settings looked right, and detection
silently never ran. The only symptom was a generic "ML API call failed" buried
in the status card.
So the test now proves what it claims. After health passes it probes GET /p/
with no img parameter: the auth decorator runs before the handler, so 401 means
the token was rejected and 422 ("Invalid request params") means it was
accepted. No inference work is done either way. A probe that itself errors
reports the token as unknown rather than as working — the UI says it could not
be checked instead of claiming success.
The detection loop checks for 401 before raise_for_status, so a rejected token
is reported as a rejected token, naming the setting and the environment
variable, instead of surfacing "401 Unauthorized" with no hint of what to do.
The message never contains the token; a test pins that.
The setting name carries "token", so the support bundle's keyword redactor
masks it with no new rule. Resolving "field omitted" to the saved token is the
route's job, keeping test_connection a pure outbound call with no database
access.
Second fix, same issue: support bundles misreported which printers Obico
watches. The bundle split obico_enabled_printers on commas and read an empty
value as "no printers". The settings UI writes a JSON array, and empty means
*all* printers — the default — so a working Obico setup showed obico_enabled
false against every printer in its own bundle. That is the reporter's bundle
exactly, and it points anyone reading it at the wrong subsystem. The bundle now
parses the setting the way ObicoDetectionService does, keeps a comma fallback
for any install that stored the legacy shape, and factors in the global switch.
|
||
|
|
db6cdb0745 |
fix(camera): take the finish photo when the print ends, not when its last layer starts (#2547)
The photo fired the moment layer_num reached total_layer_num. That edge is where the printer *starts* its final layer, not where it finishes it: the reporter's H2C capture shows it arriving at 92% with mc_remaining_time=2, three minutes and seventeen seconds and one filament change before the print actually ended, so the frame caught the toolhead mid-print over the model. The trigger also latched _finish_photo_captured, which locked out both the stage-22 and FINISH triggers for the rest of the print — so on firmware that never reports an end-of-print filament unload (H2C and A1 Mini confirmed) nothing could replace the bad frame. Remove the last-layer trigger. The photo is now taken at the FINISH-state trigger, which every model sends and which lands after the toolhead parks. Since Bambu's end G-code drops the plate ~100mm just before that, restore the framing before capturing: absolute G90/G1 Z to max_z_height + 10mm clearance, settle, capture, then drop it back so the print is as reachable as the printer left it. Absolute is the safety argument — that Z is a height the toolhead occupied seconds earlier, so it is inside the travel limits by construction and leaves the nozzle above the part, and it is unambiguous across model families because Z is the nozzle-to-bed gap whether the bed moves or the toolhead does. M211 is never touched (#2579). This is what #1145, #1397 and #1565 asked for. The height is only trusted when two independent sources agree: the archive is matched by the finished print's subtask_name by equality (not LIKE, so "Cube" cannot resolve to "Cube v2"), and its layer count from the 3MF must match the layer count the printer reported over MQTT. Matching on "most recent archive for this printer" was not safe — on_print_complete pops the _active_prints binding concurrently, and a print Bambuddy failed to archive would have resolved to its predecessor. A wrong height is the one failure that could drive the nozzle into the model. The move is additionally skipped when the print height is unknown, when a queue item is pending for the printer, when the printer has left FINISH, and when the new finish_photo_restore_plate setting is off. for every FINISH-state capture — which is what shipped the mid-print photo — the bank is used only when the dispatcher recorded that it injected End G-code into this print, since a SwapMod snippet may have ejected the plate. The flag is handed over in two steps (mark_pending at dispatch, adopt at print start) so it can never outlive its print: a job started from the slicer or SD card adopts False rather than inheriting its predecessor's answer. Those prints also skip the plate move outright, bank or no bank. The bank now refreshes on mc_percent advances as well as layer changes, via a new on_print_progress callback. Layer changes stop the instant the final layer begins, which left the #1867 fallback frame stale by the whole length of that layer; progress keeps ticking there and freezes before the End G-code runs, so a swapped plate still cannot reach the bank. The last-layer throttle exemption is dropped, since it would now fire a grab on every percent tick. On the timelapse path the moment producer returns early, so the consumer does the restore itself before its live-grab fallback — the documented usual outcome on P1-series, where the video has not transferred by the time the notification goes out and the shipped photo was of an already-dropped plate. The two waits are now derived from the settle window and the video poll timeout rather than hardcoded; at the old flat 75s that fallback was guaranteed to be cut off mid-settle. extract_max_z_height_from_3mf reads only a bounded prefix of the plate G-code, since a sliced plate is routinely tens of megabytes and the header is ~40 lines. It returns None for missing, unparseable, zero and negative values so callers must treat "don't know" as such rather than defaulting. |
||
|
|
13c37ffe51 |
fix(vp): scope saved AMS mapping to the printer it was resolved against
Round-2 review fixes for #2700. Blocking: the toggle didn't actually gate the archive write. archive.py's promotion fired for any print_data carrying ams_mapping, but bambu_mqtt's request-topic interception captures ams_mapping unconditionally for every print source (slicer-direct LAN prints included). Since main.py's real-printer auto-archive path forwards the full MQTT payload as print_data, every archive on any install — VP or not — grew extra_data.slicer_ams_mapping. Fixed by replacing the print_data-sniffing with an explicit `slicer_ams_mapping` param on archive_print() that only the VP-queue path (already gated on save_ams_mapping) ever passes. Blocking: a saved mapping could get reused on a printer it was never resolved against — tray IDs only mean something relative to one printer's AMS layout. extra_data.slicer_ams_mapping is now stored as {mapping, printer_id} instead of a bare array: - add_to_queue's fallback only fires when the reprint's target printer_id matches the mapping's origin printer. - The frontend's archiveAmsMapping only surfaces (and the Mapping button only appears) when the print modal's selected printer matches too. - A model-based VP (target_printer_id=None, no MQTT bridge to any real printer) never stamps a mapping in the first place — there's no live AMS layout for the slicer to have resolved tray IDs against. Also from review: - Multi-plate archives now get the Mapping button too (the per-plate FilamentMapping loop was missing archiveAmsMapping entirely). - Added coverage for the previously-untested late-MQTT archive patch path (_restamp_recent_queue_item), including the model-based-VP skip case. - usingArchiveMapping now also resets on printer change, not just plate/archive (it already worked via the printer-scoping above, but is now an explicit dependency too). - The Mapping button's revert (OFF) now undoes only the slots it itself set, not every manual pick in scope — matches the comment above it. - Added a comment on why negative-value slots (external spool) are skipped rather than cleared when applying a saved mapping. |
||
|
|
bab1cfb906 |
feat(vp): per-VP "Save AMS mapping" toggle + reprint auto-apply
Lets a reprint reuse the AMS slot the slicer itself picked, instead of re-deriving one from the file's static type/color. When a Print Queue VP has "Save AMS mapping" on, the slicer's own live-resolved ams_mapping (from the project_file MQTT command) is persisted onto the archive as extra_data.slicer_ams_mapping. A later reprint can reuse it via a new "Mapping" button in the filament-mapping panel — one click snaps every slot to the saved pick, click again reverts to auto-match. Archive cards and queue rows get an "AMS mapping saved" badge so it's visible beforehand. add_to_queue also falls back to the saved mapping automatically when the caller sends no explicit ams_mapping (e.g. a plain reprint with no per-slot edits). The queue item's own ams_mapping (used for that dispatch) is still captured unconditionally whenever the slicer provides it — that part is a correctness fix, not gated behind the toggle. Only the archive persistence for future reprints is opt-in. Split out from the original combined PR per review: this half is genuinely opt-in and low-risk (#2684). The dispatch-time validation gate that keeps a stored mapping honest (#1308) changes behaviour for every existing user and will land as its own PR. Review fixes applied: - _extract_slicer_ams_mapping_json: dropped the unreachable `v is None` arm and rejected bool explicitly (isinstance(v, int) accepts bool). - Translated the Russian docstring text to English. - save_ams_mapping's model comment moved to a trailing comment on the column line, matching the file's convention. - usingArchiveMapping now resets when the plate or archive changes, so the Mapping button can't read ON against a mapping it never applied. - Translated "Click to change slot assignment" and "Re-read". - add_to_queue's fallback is now called out explicitly in code comments and covered by three new integration tests (fallback fires, explicit mapping wins, unrelated extra_data doesn't false-trigger). Closes #2684 |