mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-10-06 14:11:35 +02:00
171848a0fc34c8bac7c01c8a29e5ddc77dd84f2a
2752
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
171848a0fc |
fix(slicer-presets): #1581 SliceModal refresh + invalidate on local-profile delete/import
Two-part fix for the reporter's "removed profiles still show on the slice
menu" symptom.
Local half (real bug). LocalProfilesView's import and delete mutations
invalidated ['localPresets'] (the management view's own query) but not
['slicerPresets'] (the SliceModal's unified preset query, staleTime 60s).
A freshly-deleted preset kept rendering in the slice dropdown until that
staleTime elapsed plus a refocus/remount. Both mutations now also call
queryClient.invalidateQueries({queryKey: ['slicerPresets']}).
Cloud half (opt-in cache bypass). _fetch_cloud_presets keeps a 5-minute
per-(user, token) in-process cache (slicer_presets.py:69, balances
"users see freshly-saved presets quickly" against "busy install doesn't
hit Bambu Cloud once per modal open"). Users delete cloud presets in
Bambu Studio / Bambu Handy, not in Bambuddy, so there's no event hook
to invalidate on. Rather than shorten the TTL globally, the listing
endpoint gains an opt-in ?refresh=true query param that bypasses both
the cloud cache AND the 1-hour bundled-preset cache for that one call;
the fresh result is still written back so subsequent normal callers
keep hitting the cache.
New SliceModal "Refresh" button. Lives in the preset section header
next to the cloud-status banner. Calls getSlicerPresets({refresh: true})
and writes the fresh slots into the ['slicerPresets'] cache via
queryClient.setQueryData so the spinner stops immediately rather than
triggering a second refetch. RefreshCw icon spins while in-flight;
disabled during slice enqueue to prevent double-fire.
|
||
|
|
396e9aa09e |
security: harden path-traversal class across routes + services; fifth CI backstop
Two attacker-controlled strings were being joined to library_dir with no
resolve + containment check in the project ZIP import endpoint:
- linked_folders[*].name from the request's project.json
- per-entry zf.namelist() paths from the ZIP itself
An absolute path in either field collapsed the join (Path("/lib") / "/etc"
becomes Path("/etc") because pathlib discards the left side when the right
is absolute) and the next write_bytes landed wherever the attacker chose.
Adjacent finding from the routes audit: GET /archives/{id}/photos/{filename}
had NO validation on filename and FileResponse-served arbitrary paths -
the DELETE counterpart at least gated on the photos membership check.
Adjacent finding from the services audit: ArchiveService.attach_timelapse
wrote archive_dir / filename where filename ultimately came from a printer's
FTP listing (compromised-printer threat model) or the /timelapse/select
query param. A malicious printer that exposes a directory entry with ..
segments could write the timelapse outside the archive directory.
New backend/app/utils/safe_path.py::safe_join_under(parent, *parts) is the
single source of truth: rejects empty / null-byte / absolute parts up-front,
joins under parent, resolves both sides, asserts is_relative_to. Returns the
resolved canonical path on success, raises HTTPException(400) on escape, or
PathTraversalError when http=False (for service-layer callers that need to
match a non-HTTP return contract).
Wired into the import vectors, both archive photo handlers, and the
attach_timelapse service. The full audit sweep inspected every Path/Name
join in backend/app/api/routes/ AND backend/app/services/ - 25 route-layer
sites + 8 service-layer sites confirmed safe and tagged with
# SEC-PATH-OK: <reason> so future audits trust the inline guard at a glance.
Fifth CI backstop test_route_path_arithmetic_is_safe_joined_or_marked
AST-walks both layers and fails the build on any <dir-like>/<bare variable>
join that doesn't either route through safe_join_under or carry the marker.
The services layer is in scope because it receives values verbatim from the
routes AND from external sources Bambuddy has no control over (the printer
FTP-listing case above).
SECURITY.md gets a fifth rule + a fifth row in the CI test mapping table;
the rule now names the printer FTP-listing case explicitly so future
services-layer audits set the right expectation.
--------------
fix(library): suppress warning storm when bulk-uploading ZIPs of empty/stub STL files
Uploading a ZIP of stub or empty STL files (e.g. the 24-byte
"solid test\nendsolid test" shape) produced one WARNING per file in
stl_thumbnail.py::generate_stl_thumbnail. The warnings were technically
correct - trimesh returns a valid Mesh with zero vertices, the safeguard
matches, and the function returns None so the library entry is still
created without a thumbnail - but the volume turned a successful upload
into thousands of WARNING lines in the journal.
Two changes:
1. The per-file "Failed to load STL or empty mesh" message in
stl_thumbnail.py is now logger.debug instead of logger.warning. It's
a per-file content observation, not an actionable error; the caller
already handles None correctly. The branch now catches the rare
"large enough but trimesh still can't parse it" case, visible in
debug logs without spamming production.
2. New module constant MIN_USABLE_STL_BYTES = 200 (smallest binary STL
with one triangle is 134B, smallest ASCII ~150B; 200 is a safe floor
below any real STL). The three thumbnail call sites in library.py
(extract_zip_file, single-file upload, _backfill_external_stl_thumbnails)
pre-skip files below this size before calling generate_stl_thumbnail.
Stubs never enter the trimesh pipeline at all.
Behavior is unchanged for real STLs: any file >=200 bytes runs through
the existing pipeline, MAX_VERTICES still triggers simplification at
100k vertices for the 256x256 thumbnail render, large files still get
thumbnails.
------------
fix(stl-thumbnail): silence matplotlib first-import noise (writable cache + font_manager log level)
On first STL upload, three matplotlib-internal log lines surfaced:
WARNING [matplotlib] /opt/claude/.config/matplotlib is not a writable directory
INFO [matplotlib.font_manager] Failed to extract font properties from NotoColorEmoji.ttf
INFO [matplotlib.font_manager] generated new fontManager
The writable-dir warning fired because Bambuddy's $HOME isn't writable for
matplotlib's default config path; matplotlib fell back to /tmp/matplotlib-XXX
which lost the font cache on every host reboot, so font_manager rebuilt it
each cold start - producing another batch of INFO lines.
Fix is two small additions in stl_thumbnail.py before the matplotlib import:
1. New _configure_matplotlib_cache() sets MPLCONFIGDIR to
settings.base_dir/.cache/matplotlib (mkdir if missing) so the cache
persists across container restarts and the writable-dir warning never
fires. Respects an externally-set MPLCONFIGDIR so operators who chose
their own path aren't overridden. Best-effort with a debug fallback if
settings can't be imported or the mkdir fails.
2. logging.getLogger("matplotlib.font_manager").setLevel(WARNING) at module
import demotes the per-font INFO scan that fires when font_manager
builds its cache cold. Real font warnings (>= WARNING) still surface.
3 new tests: font_manager logger at WARNING after module import;
_configure_matplotlib_cache creates the directory under base_dir and sets
MPLCONFIGDIR; an externally-set MPLCONFIGDIR is preserved verbatim.
5516 backend tests green, frontend gates clean.
|
||
|
|
db9b20631c |
fix(cloud): #1575 surface actionable error when Bambu Cloud Cloudflare challenge swallows the JSON response
When Cloudflare in front of bambulab.com returns a "Just a moment..." interstitial instead of the JSON the API normally produces, the parse error in verify_totp / verify_code / login_request used to surface as the opaque "Invalid response from Bambu Cloud" or a generic 401 from BambuCloudAuthError. Reporter hit this with three back-to-back TOTP attempts; a curl from a different network with the same honest Bambuddy UA returns clean JSON, so the trigger is CF-side (per-IP / TLS-fingerprint / rate / transient mitigation), not our code. Add a small _detect_cloudflare_challenge() helper that inspects the response for four CF markers (body "Just a moment...", body "challenges.cloudflare.com", 403 with cf-mitigated header, 503 with cf-ray header) and returns a message that attributes the block to Bambu Lab's Cloudflare protection, suggests waiting a few minutes, and points the user at a same-network browser sign-in as the standard workaround. Wired into all three JSON-parse sites; verify_totp previously had a defensive catch, login_request and verify_code now do too. No header changes, no impersonation, no retry loop - pure diagnostics. Stays clearly on the right side of Bambu Lab's "no falsified client identity" line. |
||
|
|
be15a375a6 |
fix(oidc): #1569 populate User.email from standard 'email' claim when email_claim is preferred_username
When an operator configures `Email Claim = preferred_username` (e.g. Authentik) the primary `_resolve_provider_email` correctly rejects the identity value as non-email shaped and returns None, leaving auto-provisioned users with `email=None` even though the same token carries a valid standard `email` claim. Add a narrow fallback in the auto-create-users branch only: when `provider.email_claim != "email"` and the primary returned None, resolve the standard `email` claim with the same Fall A/B shape + email_verified enforcement and use it for `User.email` and `UserOIDCLink.provider_email`. The auto-link-existing-accounts gate is left on the primary `provider_email`, so the GHSA Fall-B / Fall-C guards remain intact - the fallback never feeds account matching. |
||
|
|
b7d7c82501 | fix(security): WebSocket auth gate + audit-driven hardening sweep | ||
|
|
9c8df1744d |
chore(deps): bump vitest 3.2.4 → 4.1.8 (GHSA-5xrq-8626-4rwp, CVSS 9.8)
The Vitest UI server's /__vitest_attachment__ handler bypasses
isFileServingAllowed via a path-traversal payload, allowing arbitrary file
read/execute on the host. Dev-scope only and not exploitable in
Bambuddy's CI/CLI usage (we don't start the Vitest UI server and
@vitest/ui is not installed), but bumping clears the Dependabot alert
and brings us onto the supported 4.x line.
Bumped:
vitest 3.2.4 → 4.1.8
@vitest/coverage-v8 3.2.4 → 4.1.8
Migration-required fix:
StreamOverlayPage.test.tsx mocked `WebSocket` via
vi.stubGlobal('WebSocket', vi.fn().mockImplementation(() => ({...})))
and the page does `new WebSocket(url)`. Vitest 4 dropped support for
arrow-function constructor mocks ("is not a constructor"). Rewrote
with a plain `function` so `new` resolves correctly.
All 2043 frontend tests pass; npm run build clean; npm audit shows 0
vulnerabilities.
|
||
|
|
ec51394196 |
fix(security): GHSA-r2qv-8222-hqg3 — allowlist API-key permissions (CVSS 9.9)
API-key permission gates went from a 17-entry admin denylist with the three
documented scope flags (can_read_status / can_queue / can_control_printer)
enforced only inside /api/v1/webhook/* to an explicit per-Permission
allowlist consulted by every dependency:
- core/auth.py: _APIKEY_SCOPE_BY_PERMISSION maps every non-admin
Permission to one scope flag on APIKey; unmapped = 403.
_check_apikey_permissions now takes the api_key and checks the flag.
- require_any_permission_if_auth_enabled + require_ownership_permission
were returning None for any valid key with zero scope check; both now
invoke _check_apikey_permissions and fail closed.
- Two new scope flags on api_keys: can_manage_library (LIBRARY_UPLOAD /
UPDATE_OWN / DELETE_OWN / MAKERWORLD_IMPORT) and can_manage_inventory
(INVENTORY_CREATE / UPDATE / DELETE / FORECAST_WRITE — required by
SpoolBuddy kiosks). Default TRUE, backfilled from can_queue so existing
"queue-only" keys keep working and hardened "read-only" keys do not
silently gain writes.
- CLOUD_AUTH now routed through can_access_cloud for defence-in-depth
alongside the existing _cloud_api_key_gate.
- Migration column-existence check (_api_keys_column_exists) gates the
backfill so user-edited values are never overwritten on restart.
Structural drift backstop: test_every_permission_has_a_classification fails
CI on any new Permission added without an explicit scope mapping —
prevents the denylist-shape regression that grew the prior surface.
Backend 5469 tests green; ruff clean. Frontend build green; i18n parity
green across 9 locales (5005 leaves each, +6 new keys). Wiki permissions
table + allowlist callout + upgrade notes updated.
|
||
|
|
5a93b33fd1 | Housekeeping | ||
|
|
4ffefa60f4 |
chore(vp): per-minute MQTT status-push counter for idle-disconnect triage (#1548 follow-up)
The 1 Hz status push was silent at INFO, so support bundles couldn't show
whether the push task was actually reaching a specific slicer connection.
Now ``_periodic_status_push`` emits one line per minute per connected
slicer ("1Hz status push: N pushes/min to <client>") and stays silent when
no slicer is attached. No behaviour change to the push itself — counters
are local to the task and reset every 60 ticks.
Motivated by the open follow-up on #1548: keepalive parser shipped (
|
||
|
|
f6f6f92e38 | test(pytest): silence upstream starlette httpx2 deprecation noise | ||
|
|
e10462678f |
fix(security): GHSA-6mf4-q26m-47pv — fail-closed on auth-probe DB errors (CVSS 9.8)
is_auth_enabled() and auth_middleware both caught every exception during
the auth-state probe and returned the "allow" answer instead of denying
the request. Reporter's PoC floods /api/v1/auth/login to exhaust file
descriptors, forcing the next SQLite connect to raise, then hits a
protected endpoint during the fail-open window with no token — granting
unauthenticated access to admin-account creation, API-key creation, DB
backup download, and printer control. CWE-636 / CWE-755. Affects >= 0.1.6.
Fix:
- is_auth_enabled (backend/app/core/auth.py): only returns False for the
legitimate "settings row absent" case; any actual exception propagates
so the caller can deny the request.
- auth_middleware (backend/app/main.py): returns 503 on any probe failure
instead of await call_next(request).
4 new regression tests in test_auth_fail_closed.py pin the contract
(propagates DB exceptions, returns False for no-row, True for "true",
False for "false"). 1 existing security test renamed and updated to
accept either 500 or 503 (both fail-closed) and to verify the
SQLAlchemy detail does not leak in the body.
Codebase grep confirmed no other auth-decision predicate has the same
fail-open shape: _validate_api_key returns None on catch (→ 401 fail-
closed downstream), is_advanced_auth_enabled propagates correctly,
permissions.py has no catch-alls.
Reported by @wondercrash via private advisory.
|
||
|
|
597762685c |
fix(virtual-printer): #1558 Send pre-flight + slicer-surface audit bundle
#1558: cached-as-base push_status only forced gcode_state=IDLE while letting the real printer's live-progress fields (mc_percent, stg_cur, layer_num, ...) leak through. Bambu Studio's Send pre-flight read them as busy and refused. The cached branch now overrides the activity-field set the same way it already overrode storage indicators (#1228) and protocol fields. Same bundle ships a multi-round VP audit that found adjacent bugs in the same family: - #1558: cached branch zeroes mc_print_stage / mc_percent / mc_remaining_time / stg / stg_cur / layer_num / total_layer_num / print_error - MQTT auth: per-IP rate-limit (5/60s lockout), hmac.compare_digest, access_code redacted in DEBUG log - FTP cmd_STOR streams chunks to disk + 4 GiB cap (was buffering whole upload) - Sticky-keys allowlist extended with upgrade_state / xcam / hw_switch_state / nozzle_diameter / nozzle_type / online / ams_status - _pending_files cleanup in finally for archive / queue / dispatch handlers - _add_to_print_queue position uses MAX+1 (was hardcoded 1) - DELETE VP removes orphan PendingUpload rows + upload_dir from disk - Per-VP cert regenerates on shared-CA rotation (real signature verification, not DN match) - DHCP target-IP refresh + queue_force_color_match toggle now restart proxy VPs - Per-slicer bridge-response routing (multi-slicer cross-leak fix via sequence_id map) - Child-service readiness barrier (FTP / MQTT / Bind / SSDP) — no false is_running before sockets bind - H2D Pro O1E / O2D model codes added (experimental, needs field confirmation) - FTP passive port range widened 50000-51000; docker-compose + wiki updated - VP refresh_loop crash now unbinds raw_message_handler; tailscale catches asyncio.TimeoutError; SlicerProxyManager lifecycle hardening |
||
|
|
40cd45e2d5 |
fix(library): render 3D preview for sliced .gcode.3mf files (#1543)
Reporter exported a multi-plate .gcode.3mf from Bambu Studio to the
shared folder Bambuddy watches and the 3D preview tab came up empty;
if he re-uploaded the same file via the file manager, the preview
worked. Root cause: two paths classify file_type differently. The
shared-folder scan at backend/app/api/routes/library.py:1343-1348
does a compound-extension check and tags the file `gcode.3mf`; the
upload path at the same file's :1588 does a single `ext[1:]` and tags
it `3mf`. Then frontend/src/components/ModelViewerModal.tsx:71-73 had
hasModel = normalizedType === '3mf' || 'stl'
hasGcode = normalizedType === 'gcode' || '3mf'
Neither matched `gcode.3mf`, so the capabilities object landed with
both flags false and the modal rendered an empty bed.
FileManagerPage.tsx:858 also gated the Preview-3D context action on
`file_type === '3mf' || 'gcode' || 'stl'`, so for shared-folder files
the menu entry didn't even appear, and the type pill at :765-770 had
no colour case for `gcode.3mf` so it fell through to the generic
gray.
Fix (frontend-only, no backend churn):
- ModelViewerModal.tsx introduces an
`isThreeMfFamily = normalizedType === '3mf' || normalizedType === 'gcode.3mf'`
predicate used in two branches — the capabilities check
(`hasModel = isThreeMfFamily || 'stl'`, `hasGcode = isThreeMfFamily
|| 'gcode'`) and the plates-loading branch that previously hard-
gated on `!== '3mf'` and would have returned setPlatesData(null)
for the shared-folder file.
- FileManagerPage.tsx adds `gcode.3mf` to the Preview-3D action gate
and shares the gcode blue type-pill colour so sliced-output files
are visually distinguishable from source 3MFs.
The compound `gcode.3mf` classification on the backend is intentionally
preserved — it carries useful "this is a sliced output" semantics that
other UI surfaces could use later. The `canOpenInSlicer` and
`sliceableType` checks at ModelViewerModal.tsx:269, 277-280 are
deliberately left alone — a sliced output isn't openable in the slicer,
and `sliceableType` already explicitly excludes `.gcode` and
`.gcode.3mf` per the comment.
Out of scope (separate Bambu-Studio format limitation, not a Bambuddy
bug): Vlado's secondary observation that the upload-path 3D preview
"shows only one plate" even though his project has 5 plates — Bambu
Studio's .gcode.3mf export contains the model data and g-code for the
active plate only, not the entire multi-plate project. The print
picker enumerates plates via gcode_*.gcode entries inside the zip
(a separate code path), which is why the user can still pick the
plate at print time. The empty-bed fix is the data point that closes
the user-visible bug.
|
||
|
|
ed232718f5 |
fix(prints): connected-edge reconciliation closes the missed-PRINT-COMPLETE loop behind smart-plug ghost prints (#1542 follow-up)
Reporter ran a fresh trace after the doubled-extension fix landed and
found a distinct second cause behind his ghost prints, hitting 4-of-4
of his A1s. Timeline:
22:50 PRINT START
...print runs all night...
23:13 / 00:47 / 09:35 MQTT disconnects (A1 keepalives are unstable)
print finishes during one of those disconnect windows → PRINT COMPLETE
is never observed
smart plug detects idle → cuts power
power resumes for the next scheduled print → firmware auto-replays the
leftover .3mf from the SD card
09:46 Bambuddy reconnects to a fresh PRINT START for the ghost
The existing IDLE-after-RUNNING completion check at
backend/app/services/bambu_mqtt.py:3022 was meant to catch the simple
disconnect-then-finish case via `_previous_gcode_state` preserved across
reconnects, but with multiple disconnect/reconnect cycles + a smart-plug
power-off that Bambuddy can't distinguish from any other transient drop,
the IDLE window that branch needs simply never reaches it. The SD .3mf
lingers, the firmware ghost-replays every power cycle, and the loop
repeats.
Fix: new connected-edge reconciliation pass.
* `_is_active_archive_stale(archive, state)` — pure decision function
with three triggers:
(1) printer state is terminal (IDLE / FINISH / FAILED)
(2) printer running with a different `subtask_id` than the archive —
Bambu firmware mints a fresh subtask_id for each print including
the ghost replay, so a mismatch is unambiguous
(3) printer running but `subtask_name` is empty — printer doesn't
know what it's running, archive reference is broken
Conservative on PAUSE / PREPARE / SLICING and on RUNNING with matching
subtask. False-positive cost = one misreported "aborted" status that
the next real PRINT COMPLETE would have overwritten anyway. False-
negative cost = the ghost-print loop.
* `reconcile_stale_active_prints(printer_id)` — queries archives in
`status="printing"` for the printer, runs the decision function, and
synthesises `on_print_complete(status="aborted", _reconciled=True)`
for each stale match. Reuses the existing PRINT COMPLETE chain (SD
cleanup, status update, usage tracker, notifications) — no reimplem-
entation. Per-archive try/except so one failure doesn't block the
rest. Returns 0 when status is None / disconnected — the connected
edge is the only legitimate trigger.
* `on_printer_status_change` now runs a connected-edge check at the
start. New `_printer_reconciled_since_connect: dict[int, bool]`
tracker flips False → True on the first connected status update for
this connection and back to False on disconnect, so reconciliation
fires exactly once per (re)connection. The flag is set BEFORE the
task is spawned so concurrent status updates within the same
connection don't re-trigger it (no await between check and set,
asyncio guarantees atomicity). The reconciliation runs as
`asyncio.create_task` so the hot WebSocket dedup / broadcast path
isn't blocked.
* Single handler covers both startup and reconnect — when the first
MQTT connection completes after startup the printer pushes status,
the connected edge fires, reconciliation runs. No separate startup
hook needed.
Idempotency: when the existing #3022 branch DOES fire on a clean
disconnect-then-IDLE-on-reconnect, it lands `archive.status` to terminal
synchronously. The async-scheduled reconciliation then queries
`WHERE status="printing"` and finds 0 rows → no-op. The narrow race
where reconcile's query lands between a real on_print_complete's
archive update and its archive lookup produces at most a duplicate
notification (no double SD-cleanup since FTP delete on a missing file
is a no-op 550).
Ghost-print collateral worth being explicit about: if the ghost is
already running when reconciliation fires, the synthesised SD-cleanup
hits 550-file-locked (same root cause as the #1542 first case). The
cleanup retries 3× then logs "lingering". The ghost runs to completion,
its own end-of-print cleanup deletes the file, the next power cycle
has nothing to replay, the loop breaks. A perfect cancel would require
a `print_stop` MQTT command to the printer mid-ghost — invasive,
explicitly out of scope.
|
||
|
|
7a7dfed81a |
fix(archives): resolve raw_data wrapper in fallback-archive filament extraction (#1533 follow-up)
Reporter updated to 0.2.5b1 expecting the #1533 fix to populate filament fields on his P2S virtual-printer prints when the .3mf is locked. His support bundle showed Bambuddy still creating fallback archives with NULL filament fields even though the print-start log line proved AMS-0-T0 had PETG loaded at the moment the helper should have read it (`AMS 0: T0(type=PETG, color=FFFFFFFF, ...)`). Cause: the #1533 helper `_extract_filament_data_from_mqtt(data)` in backend/app/main.py only looked at `data["ams"]`, but the dict that on_print_start actually receives at runtime is the wrapper shape `{"filename", "subtask_name", "remaining_time", "raw_data": <mqtt>, "ams_mapping"}` that backend/app/services/bambu_mqtt.py:2971-2980 constructs. So `data["ams"]` was undefined on every real call and the helper silently returned `{}`, leaving the fallback archive's filament_type / filament_color NULL — the exact regression the original fix was meant to close. The 15 unit tests that shipped with #1533 all passed the bare inner shape directly and never exercised the callback wiring, so the regression slipped through a green build. Surrounding code in the same fallback path (main.py:2064, 2674) already reads `data.get("raw_data")` — the original hunk was the outlier that forgot the wrapper. Fix: helper now resolves `data["raw_data"]["ams"]` first (the callback shape) and only falls back to `data["ams"]` when the wrapper isn't present, preserving the inner-shape callers from the existing tests. Defensive against a non-dict `raw_data` (e.g. partial MQTT decode failure) falling through to the inner lookup instead of crashing. Tests: 5 new in `TestOnPrintStartCallbackShape` in test_fallback_archive_mqtt_filament.py — wrapper payload with ams_mapping resolves to the inner AMS state; wrapper without ams_mapping lists all loaded slots; the existing inner-shape callers still work after the additive wrapper lookup; missing raw_data returns `{}` instead of raising; junk raw_data (string) doesn't shadow a present inner `ams`. Existing 15 inner-shape tests untouched and green. Full 5378-test backend suite green; backend ruff clean. What this does NOT fix: per-filament gram usage still needs the actual .3mf — the printer locks it during print (P-line firmware behaviour, not a Bambuddy bug), and the existing 19 FTP candidate paths + directory probes are expected to 550 in that window. Per-print filament type and colour are the data point the reporter explicitly called out as load-bearing for AMS-expansion planning at his maker space, so this is what moves the needle. |
||
|
|
9347921378 |
fix(inventory): drop profile-only mismatch popup and clarify reconfigure intent (#1552)
Reporter assigned a spool to a slot whose stored slicer profile differed
from the new spool's, got the Cancel / Assign Anyway popup, and was under
the impression that confirming the popup just linked the spool in
Bambuddy's DB without pushing the new profile to the AMS — i.e. that he
then had to manually open Configure AMS Slot to fix it. The auto-push has
actually been in place since the assign route existed:
backend/app/api/routes/inventory.py::assign_spool calls
apply_spool_to_slot_via_mqtt after upserting the SpoolAssignment row,
which publishes ams_filament_setting + extrusion_cali_sel over MQTT, and
backend/app/api/routes/spoolman_inventory.py::assign_spoolman_slot does
the same for the Spoolman backend. The only short-circuit is when
firmware reports the slot explicitly empty (tray_state in {9, 10}), in
which case main.py::on_ams_change deferred-replays the configure once a
spool appears. So the popup was friction without revealing what it did.
Two changes:
- frontend/src/components/AssignSpoolModal.tsx and
spoolbuddy/AssignToAmsModal.tsx: profile-only mismatch no longer
fires the popup. The condition becomes
`if (materialMatchResult !== 'exact')` instead of
`materialMatchResult !== 'exact' || !profileMatches`. The 'profile'
member is dropped from the mismatchType union and its standalone
branch in each popup render body is removed as dead code. Material
mismatch still warns — Bambu firmware can refuse the print when the
type is wrong.
- Every firing warning (material, partial, material+profile,
partial+profile) now appends one line via a new
inventory.assignReconfigureNote i18n key:
"The AMS slot will be reconfigured to use the spool's profile."
Real translations across all 9 locales per
feedback_translate_dont_fallback; parity script clean at 4999
leaves per locale.
|
||
|
|
632334953c |
fix(inventory): support transparent / clear filament end-to-end (#1545)
Reporter wanted to select a transparent filament colour in the spool
editor; CMW-ISS confirmed on v0.2.5b1 that AMS-detected transparent
spools were silently labelled "Black" in the filament-mapping dropdown
because the colour name resolver dropped the alpha byte and the underlying
RGB 000000 HSL-bucketed to "Black". Spoolman already supported 8-digit
hex; the built-in inventory didn't.
Eight collapsing sites fixed together so transparent reaches the user
intact:
- frontend/src/utils/colors.ts: hexToColorName / getColorName /
resolveSpoolColorName / isLightColor short-circuit to "Clear" for
alpha=00 before HSL bucketing or catalog lookup
- frontend/src/utils/amsHelpers.ts::normalizeColor preserves the alpha
byte when alpha < FF (normalizeColorForCompare unchanged so type/colour
matching is unaffected)
- frontend/src/components/spool-form/constants.ts: new
{ name: 'Clear', hex: '00000000' } preset in QUICK_COLORS
- frontend/src/components/spool-form/ColorSection.tsx: hex draft accepts
0-8 chars, commits at 6 (+FF) or 8 verbatim; blur pads 7-char to 8;
selectColor passes 6-char as +FF / 8-char verbatim; isSelected matches
on full rgba; swatch buttons paint a checkerboard for alpha=00
- backend/app/api/routes/printers.py::get_available_filaments preserves
the full rgba on both AMS and vt_tray branches (6-char dedup key
unchanged)
- backend/app/services/spoolman.py::parse_ams_tray drops the silent
00000000 -> F5E6D3FF cream rewrite — the swatch renderer paints a
checkerboard underlay for alpha < FF already (added in #1154), so the
rewrite was hidden technical debt that made every AMS-detected
transparent spool land in inventory as cream
- backend/app/services/spool_tag_matcher.py::create_spool_from_tray
short-circuits the colour-catalog lookup for alpha=00 and stores
color_name="Clear" directly — otherwise an RFID-tagged transparent
Bambu spool would resolve against the #000000 catalog row (or "Black"
via the HSL fallback) before the frontend's resolver ever saw it
- Two shared helpers in utils/colors.ts — getSwatchStyle(rgba) (style
object: checkerboard for alpha=00) and spoolColorString(rgba)
(8-char hex string for SVG fill) — applied to every simple-swatch
site that would otherwise have rendered Clear spools as solid black:
LabelTemplatePickerModal, SpoolBuddyInventoryPage (SpoolCircle + dot),
SpoolBuddyAmsPage (both branches), SpoolBuddyWriteTagPage (4 sites),
ForecastPanel, AssignToAmsModal, AssignSpoolModal (both branches),
InventorySpoolInfoCard, TagDetectedModal, SpoolInfoCard, LinkSpoolModal,
and the FilamentSwatch tooltip title fallback
Intentionally NOT changed: native <input type="color"> keeps 6-char hex
(can't pick alpha; onChange still emits +FF, correct); Spoolman's
_find_or_create_filament strips alpha (Spoolman catalog is 6-char only);
print_scheduler colour matching strips alpha (auto-mapping treats Clear
as Black for slot compatibility); label_renderer prints "#RRGGBB" on the
physical label (printers can't print transparency, swatch fill via
_color_from_hex still honours alpha).
|
||
|
|
b663605318 |
fix(virtual-printer): honour client-negotiated MQTT keepalive instead of hardcoded 60s (#1548)
OrcaSlicer connects, exchanges pushall + get_version, then sits idle waiting
for status pushes from the (virtual) printer. The VP MQTT server's read
loop used `asyncio.wait_for(reader.read(1), timeout=60)` regardless of what
the client negotiated, and `_handle_connect` explicitly skipped the
keepalive field in the CONNECT payload, so every idle slicer connection was
torn down at exactly 60s.
- Parse the 2-byte big-endian keepalive from CONNECT; return it from
_handle_connect alongside the auth bool.
- Use 1.5x the negotiated keepalive as the per-packet read timeout per
MQTT spec sec 4.4. Treat keep_alive == 0 as no timeout (spec sec 3.1.2.10).
- Retain the 60s default for the initial read before CONNECT arrives, so
a TCP-connect-without-CONNECT still gets reaped.
- 7 new tests: 4 unit-level for the parser (success, opt-out=0, auth-fail
tuple shape, malformed CONNECT) + 3 integration-style for the read loop
(long keepalive survives the old 60s mark, short keepalive closes idle
in ~3s, PINGREQ resets the window so DISCONNECT decides the exit).
|
||
|
|
6d316c5593 |
fix(dispatch): align SD cleanup with upload path so doubled-extension library rows don't leave ghost prints (#1542)
A library row with archive.filename "Cube (1).gcode.3mf.gcode.3mf" uploaded
to /Cube_(1).gcode.3mf.3mf (single-iteration strip + append). Post-print
cleanup looked at /Cube_(1).3mf and /Cube_(1).gcode (subtask_name + ext),
missed the on-card file, and A1 firmware re-ran it on next power-on.
- New derive_remote_filename() helper in backend/app/utils/filename.py:
iterative strip of .gcode.3mf/.3mf suffixes, append single .3mf,
space->underscore. isinstance() guard raises TypeError on non-str
input rather than entering the strip loop with a duck-typed
object that returns truthy sentinels from endswith.
- Three previously-duplicated upload sites (background_dispatch reprint +
library, print_scheduler queue) now share the helper.
- SD cleanup fetches archive.filename and tries the derived path first,
with the legacy subtask_name + ext paths kept as fallbacks.
- 10 new unit tests pin the reproducer + edge cases (doubled .gcode.3mf,
doubled .3mf, raw .gcode preserved, idempotence, Unicode, plus the
type guard against MagicMock / None / int inputs).
|
||
|
|
2241924312 |
fix(library): reject FAT32-illegal filename chars at rename/upload/queue time (#1540)
Bambu printer SD cards are FAT32/exFAT, which forbids < > : " / \ | ? * plus control chars and trailing dots/spaces. Library rename only blocked path separators, so a name like L|R.3mf was accepted and only failed later at FTP upload with 553 Could not create file - far from the rename action that caused it. Bambu Studio refuses these names in its save dialog; Bambuddy now does the same. New backend/app/utils/filename.py centralises validation. Wired into update_file, upload_file, print_library_file, and queue add. Existing rows with bad names are left alone (no silent rewrite of user data); users get an actionable 400 pointing at rename. Frontend rename modal mirrors the same set client-side with inline error. New fileManager.invalidFilenameChar i18n key translated across all 9 locales. 26 new tests in test_filename_validation.py. |
||
|
|
e01a67f587 | Updated BACKERS | ||
|
|
158a3836a1 |
ci: bump actions to Node-24-compatible majors
GitHub forces Node-20 actions to run on Node 24 starting 2026-06-02 and
removes Node 20 from the runner on 2026-09-16. Bumping each action to
its first Node-24 major now gets us ahead of both deadlines and silences
the deprecation warnings already firing in every CI run.
Bumps (across ci.yml, security.yml, codeql.yml, auto-label-area.yml,
issue-closed.yml, stale.yml):
- actions/checkout v4 -> v6
- actions/setup-python v5 -> v6
- actions/setup-node v4 -> v6
- actions/cache v4 -> v5
- actions/upload-artifact v4 -> v7
- actions/github-script v7 -> v9
- actions/stale v9 -> v10
- docker/setup-buildx v3 -> v4
- docker/build-push v5 -> v7
Verified each major's breaking-change notes against our usage:
- setup-node v6 limits auto-cache to npm only; we already pass
cache: 'npm' explicitly, so nothing changes.
- github-script v9 drops require('@actions/github'); none of our
scripts use it (only require('fs') and the injected github/context
globals).
- setup-buildx v4 removes deprecated inputs; we call it with no
inputs.
- build-push v6 enables build summaries by default; informational,
can disable via DOCKER_BUILD_SUMMARY=false env if it gets noisy.
codeql-action stays on v4 (already runs on Node 24). Trivy and
github-repo-stats are Docker actions and aren't affected by the
Node-20 deprecation.
|
||
|
|
eab08d97b6 |
chore(deps): pin fastapi<0.136.0 to dodge MAL-2026-4750
Amazon Inspector flagged fastapi 0.136.x for shipping an undocumented `fastar>=0.9.0` dep in its [standard] extras group. `fastar` is a Rust-tar binding package, no plausible reason for a web framework to depend on it. Even if `fastar` is benign today, the advisory's "namespace-abuse vector" framing is valid — whoever controls the fastar PyPI namespace gains code execution at install time across every fastapi[standard] install. Bambuddy doesn't request [standard] so we don't pull fastar in practice, but pip-audit flags the fastapi package itself and breaks CI. Hold to 0.135.x (last clean release line) until upstream removes the dep. |
||
|
|
b48dca5b4b | Updated BACKERS | ||
|
|
c42e923e4c |
fix(archives): MQTT-derived filament type/color on fallback archives (#1533)
When the source .3mf can't be downloaded at print start (P1S/A1/P2S
firmwares lock the file mid-print), main.py creates a fallback
PrintArchive with file_path="" and every filament field NULL — even
though the MQTT payload already has the AMS state and the slicer's
slot-per-print-filament mapping (data["ams"]["ams"] and
data["ams_mapping"]).
New _extract_filament_data_from_mqtt(data, ams_mapping) builds a
{global_tray_id: (type, color)} map from the AMS units, then narrows
to slots referenced by ams_mapping (slicer order preserved, -1 VT-tray
sentinels skipped) or falls back to every loaded slot when no mapping
is present. Returns comma-separated filament_type and filament_color
matching the 3MF-extraction shape, so the inventory page, Quick Stats
rollup, and len(filament_type.split(",")) per-print count behave
identically for fallback rows.
The constructor at the fallback site now passes the resulting values
into the PrintArchive row.
This does NOT recover per-filament gram usage — that needs the .3mf's
slice_info.config or a deeper layer-delta integration via usage_tracker.
The reporter (maker-space lead evaluating Bambuddy partly for AMS
expansion planning) asked specifically for "the number of filaments
used", which is what this gives them.
15 unit tests cover empty/malformed payloads, the no-mapping path,
mapping filtering and reordering, VT-tray sentinels, dual-AMS global
ids, column-limit truncation, and defensive garbage handling.
|
||
|
|
d0ff6f7dc1 |
fix(spoolbuddy): tare banner now resolves to complete or timed-out (#1536)
The TARE button on Settings → Scale set a "Tare command sent. Waiting for device..." banner with no mechanism to clear it. The daemon writes back through /calibration/set-tare which stamps last_calibrated_at on the device row, but handleTare was set-and-forget — the banner stayed forever. The "Calibration complete!" success banner had the same shape. Snapshot last_calibrated_at when TARE is pressed, set an awaiting state, invalidate the device-list query every 1s while waiting (so detection responds within ~1s, not the 10s background poll), and when the snapshot advances flip the banner to "Tare complete!" with a 3s auto-dismiss. A 15s timeout falls open to "Tare timed out — is the SpoolBuddy daemon running?" so a dead daemon doesn't trap the user on the spinner. The calibration-complete and calibration-failed banners now share the same auto-dismiss helper. |
||
|
|
4343bd60b1 |
fix(notifications): honest UA + Cloudflare-challenge detection on ntfy (#1534)
The notification service's httpx client was the only outbound client in the codebase still leaking python-httpx/<version> as User-Agent; all other clients identify as Bambuddy/1.0 since the May 2026 compliance pass. Bring it in line. The reporter's ntfy server was behind a Cloudflare Tunnel and CF returned its JS challenge page (Just a moment...) to every API request — confirmed by reproducing the same 403 with curl. Cloudflare can't be solved from a backend, so add detection for the challenge shape (Server: cloudflare or cf-mitigated header, or <!DOCTYPE html>...Just a moment... body) and return an actionable error message that points at the real fix on the user's CF side instead of dumping the raw HTML. Normal 403s (auth failures with plain text bodies) still surface the original body so genuine errors stay debuggable. |
||
|
|
e9beb1e8fc |
fix(archives): handle fallback archives in source-3MF upload (#1531)
Archives created from prints Bambuddy didn't archive (cloud / Handy / SD-card prints) carry file_path="". The two source-upload routes computed the destination as (base_dir / archive.file_path).parent / "source", which collapsed to base_dir.parent / "source" for fallback rows — sending the file to /app/source/ (outside the data volume, orphaned on container restart) and raising 500 on the final relative_to. Centralise the destination math in _resolve_source_3mf_path. Normal archives keep the <archive>/source/<filename> layout. Fallback archives land at <base_dir>/archive/no_source/<id>/<filename>, which stays inside the data volume and is addressable by every existing read site. The helper also asserts the resolved directory is under base_dir.resolve() so a corrupted row fails with a clear message instead of writing outside the volume. Both upload routes (upload_source_3mf and upload_source_3mf_by_name) now route through the helper. Two regression tests in TestUploadSourceThreeMF pin both branches. |
||
|
|
4387a09162 |
fix(spoolbuddy): route weight sync by inventory mode exclusively (#1530)
POST /spoolbuddy/scale/update-spool-weight tried the local DB first and only fell back to Spoolman on a local miss. Combined with nfc/tag-scanned's post-#1119 always-Spoolman routing, a stale local Spool row sharing a numeric id with a Spoolman spool would absorb the sync silently while the Spoolman row stayed unchanged. Mirror the routing already used by nfc/tag-scanned: pick the branch via _get_spoolman_client_or_none() and never cross. Local mode now returns 404 on a local miss instead of falling through. New TestUpdateSpoolWeightSpoolman.test_stale_local_row_does_not_shadow_spoolman asserts both directions: Spoolman gets the update, the colliding local row's weight_used and last_scale_weight are untouched. |
||
|
|
e34958c3fa | Post work PR #1501 | ||
|
|
052e928107 |
feat: add system theme detection (prefers-color-scheme) (#1501)
feat: add system theme detection (prefers-color-scheme) |
||
|
|
554a73070f |
fix(maintenance): paused prints no longer accumulate runtime hours (#1521)
PAUSE counted toward runtime_seconds equally with RUNNING, inflating hours-based maintenance thresholds (rod lube, belt check, nozzle clean) by however long overnight or extended pauses lasted. Maintenance items track mechanical wear, which is zero while paused, so the predicate now excludes PAUSE. Field-comment and docstring trail across main.py / models/printer.py / maintenance.py updated to match. Existing runtime_seconds values cannot be retroactively split — only future accumulation is fixed. Adds 3 regression tests pinning PAUSE non-accumulation, RUNNING accumulation, and the FINISH state's last_runtime_update clear (prevents idle-time back-bill when the printer next goes RUNNING). |
||
|
|
4cce575bcc |
fix(stats): cancelled bucket icon now uses semantic warning token (orange) (#1390 follow-up)
Reporter flagged that the new Cancelled row's Ban icon rendered colourless while Successful and Failed used green/red icons in the same widget. Switch the Cancelled icon to text-status-warning (amber-500) so all three rows now use semantic status tokens consistently, and the colour matches the orange Archives + notification badges already use for cancelled status. |
||
|
|
e5ebab7ab8 |
chore(triage): tighten bug-report template + add Area dropdown to cut invalid-issue load
170 issues have been closed with the `invalid` label (61 of them in
the last 30 days alone — ~1 in 5 of all closed issues), almost always
because the reporter hadn't run the in-app Connection Diagnostic or
checked the documented troubleshooting page. The Connection Diagnostic
shipped weeks ago but the bug-report form let people skip it: the
"I ran it" checkbox was `required: false` and the Support Package
field was optional. Tighten both.
Form changes (.github/ISSUE_TEMPLATE/bug_report.yml):
- Connection Diagnostic checkbox: required: false → true
- Support Package field: required: false → true ("drag the .zip
or explain why you cannot attach one")
- New required textarea "Troubleshooting steps already taken" —
forces the reporter to type WHAT they tried and WHICH wiki pages
they checked before submitting. Empty answers can't submit.
- Pre-form intro spells out the search → wiki → diagnostic →
support package sequence and cites the 1-in-5 stat
- Final-checks list grew from one to three required confirmations
(searched issues + checked troubleshooting wiki + ran Connection
Diagnostic for any connection/printing/camera issue)
Bug categorization (the gap that motivated this):
- Old `Component` dropdown was Bambuddy / SpoolBuddy / Both — no
area triage signal
- Replaced with two required dropdowns:
- Product: Bambuddy / SpoolBuddy
- Area: 15 options covering the actual feature surface +
Other / not sure
- Auto-label workflow (.github/workflows/auto-label-area.yml)
reads the Area dropdown from the rendered issue body on
open/edit and applies the matching area:* label. Tolerant of
CRLF and the _No response_ placeholder, won't re-add on edit
re-fires, warns on unknown Area values
Maintainer hand-off — labels must exist BEFORE the workflow runs,
since github-script's addLabels throws on missing labels. Create the
16 labels (15 area:* + 1 area:unsorted) once via the `gh label create`
commands captured in CHANGELOG / commit context.
OS dropdown left untouched (Docker stays — per Martin).
Printer Model dropdown verified against backend/app/utils/printer_models.py
PRINTER_MODEL_MAP: all 13 current models present (X1 Carbon, X1, X1E,
X2D, P1S, P1P, P2S, A1, A1 Mini, H2D, H2D Pro, H2C, H2S).
|
||
|
|
1e734fb7c6 |
fix(stats): cancelled prints get their own bucket; gauge denominator excludes them (#1390 follow-up)
Reporter (@IndividualGhost1905) saw Total: 20 / Success: 18 / Failed: 1 and asked where the 20th print went. The Quick Stats endpoint counted status == "completed" → Successful and status == "failed" → Failed, but used a raw count(*) for Total Prints, so the four other PrintLogEntry statuses (aborted, stopped, cancelled, skipped) silently inflated the total without showing up in any breakdown row. The earlier #1390 round had committed a test locking in this exact behaviour ("uses total_prints as denominator so cancelled/stopped events count"), which was wrong: it conflated user intent with print quality. Three-bucket classification, applied across the whole stats surface and matching how the rest of the codebase already groups statuses (main.py:430, 1729; failure_analysis status filter): successful = completed failed = failed + aborted (printer-detected quality failures) cancelled = stopped + cancelled + skipped (user/queue stopped) Quick Stats endpoint returns the new cancelled_prints field; ArchiveStats.cancelled_prints defaults to 0 so older fixtures still parse. SuccessRateWidget gauge now divides by successful + failed only — a cancelled roll no longer drags the gauge down — and a Cancelled row appears in the breakdown so the missing prints don't silently vanish from Total Prints. Failure Analysis service applies the same denominator fix to both the headline failure_rate and the per-week trend, so a week with several cancellations and zero failures reads as 0% rather than a misleading "failed / total". i18n: new stats.cancelled key in all 9 locales with real translations (no English fallback), parity script clean. Tests: the existing 'uses total_prints as denominator' assertion is inverted to assert the new behaviour (40 / 20 / 35 → 67% gauge, Cancelled: 35 visible). The unchanged-display path (140 / 10 / 0 → 93%) still holds since 140 / (140 + 10) = 93.33% rounds the same. 33 StatsPage tests + 6 backend stats/failure tests green. |
||
|
|
b9b06a7351 |
chore(docker): silence Trivy DS-0026 on Dockerfile.test via HEALTHCHECK NONE
Trivy raised DS-0026 ("No HEALTHCHECK defined") against Dockerfile.test
on every run of the security workflow. The test image is a one-shot
pytest runner — there's no service to probe, so any HEALTHCHECK we
invented would be cargo-cult noise that fires once and means nothing.
HEALTHCHECK NONE is the documented Docker directive to explicitly opt
out of any inherited HEALTHCHECK and is the way Trivy itself expects
projects to signal "this image is intentionally not a long-running
service." Adding it closes code-scanning alert #813 cleanly.
Note: the perl-base CVE-2026-8376 alert (#811) is left open for now
and dismissed in the GitHub UI as "Won't fix - no upstream patch"
because Debian Trixie has not yet shipped a fixed perl-base; the
patched build will land automatically on the next base-image refresh.
|
||
|
|
12d344cbfc |
ci(docker): full backend suite in Docker, 4-way matrix shard, GHA cache backend
Earlier patch trimmed the duplicate unit-test re-run from docker-test
to drop a 5-10 min job that wasn't adding coverage. But "wasn't adding
coverage" only holds for pure-logic tests — system-touching tests
(ffmpeg version probes, ftp clients, subprocess shell-outs, locale/
timezone-sensitive assertions, paths) genuinely can pass on the GHA
host and fail in python:3.13-slim. Curation via a `docker_env`
marker is fragile (new tests get forgotten); gating on `main` only
defers the cost without removing it.
Instead, run the full backend suite IN Docker on every PR but make
it fast:
- New docker-backend-tests job runs the same 4-way pytest-split
matrix as the host backend-tests, just inside the test image.
- docker/setup-buildx-action + docker/build-push-action@v5 with
cache-from/cache-to: type=gha,scope=backend-test persist the
BuildKit cache (pip-install layer included) across CI runs and
across the 4 sibling shards. Cold build is ~150s/shard; warm
build drops to ~10s/shard.
- fail-fast: false so a single failing shard surfaces the rest's
output too.
Total CI wall-clock for a PR push is now gated by docker-test (the
image-build + integration HTTP smoke + integration test suite job)
at ~3 min, not by the unit-test re-run anymore.
The earlier ci.yml step that ran `docker compose run --rm
backend-test` synchronously in the docker-test job stays removed —
the new docker-backend-tests matrix covers the same ground and is
much faster.
|
||
|
|
a4afc9c073 |
ci(docker): stop re-running unit tests inside the test image
The "Docker Build" job in ci.yml was running the same 5287 backend tests + 2022 frontend tests inside the bambuddy-backend-test / bambuddy-frontend-test images that the host-side backend-tests and frontend-tests jobs had already run. Same test code, same Python version (env.PYTHON_VERSION), same requirements.txt the test image installs. On 2-vCPU GHA runners that re-run added 5-10 min of wall-clock for zero new coverage — and "frontend tests in Docker" added another 2-3 min for the same reason. Drop both steps from the CI job. Keep everything that validates the Docker IMAGE specifically: production image build, backend module import verification, static-files-copied check, integration container bring-up + health/API/static HTTP smoke checks, and the integration test suite (which IS genuinely Docker-specific — it runs against the live container via BAMBUDDY_TEST_URL). test_docker.sh keeps the unit-test reruns because devs running it locally don't have a separate host-side pytest job to compare against. Combined with the earlier 4-way pytest-split shard on the host backend-tests job, expected PR-push wall-clock drops from ~10-12 min to ~3 min, gated on max(backend-tests shard, frontend tests, docker-image-build+integration). |
||
|
|
0aadce1a8a |
ci(docker): drop -v, -n auto instead of -n 30, pip cache mount
Three things were making the Docker test runs noisier and slower than
they needed to be:
1. -v was hardcoded in Dockerfile.test:35 CMD and in docker-compose.
test.yml's integration-test-runner command. The ci.yml change to
drop -v from the bare pytest call missed both — Docker runs use
the image's CMD, not the workflow's.
2. -n 30 was hardcoded as the xdist worker count. On a 2-vCPU CI box
that's 30 Python processes fighting over 2 cores — mostly IPC and
import-thrash overhead. -n auto adapts to the host: 2 on CI, 30
on a 30-core dev box. Same final-result throughput on the dev
box, much better on small runners.
3. pip install had --no-cache-dir and no BuildKit cache mount, so
every Docker build re-fetched ~50 packages from PyPI (~60-90s
on a cold pip cache). Adding `RUN --mount=type=cache,target=
/root/.cache/pip` (with the `# syntax=docker/dockerfile:1.7`
directive that enables it) makes subsequent builds re-use the
download cache so they only do install work, ~5s instead of
~90s. DOCKER_BUILDKIT=1 is already exported in test_docker.sh
and is the GHA default since runner image 2023, so the cache
mount is always honoured.
Verified locally: Docker build is 19s warm (was ~90s cold each
time), test run is 102s with 5287 passed / 1 skipped (the
by-design spoolbuddy importorskip) — clean output, no [gwN]
worker spam, no "created: 30/30 workers" startup line.
GHA-side per-run cold-build slowness still happens because GHA
runners are ephemeral; a follow-up using docker/build-push-action
with type=gha cache backend would persist the BuildKit cache
across CI runs but that's a bigger workflow change.
|
||
|
|
ca08f1f340 |
fix(test): stop sys.modules-deleting backend.app.main in test_code_quality
+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup Root cause of the 4 CI failures on PR #1514 (all in test_print_start_assigns_printer_id_to_vp_archive.py + test_timelapse_baseline_restart_recovery.py): test_all_modules_importable in test_code_quality.py was deleting backend.app.main from sys.modules and re-importing it via importlib.import_module. That created NEW module-level dicts (_timelapse_baselines, _expected_prints, _active_prints, …) and re-ran root_logger.addHandler — hence the duplicate log lines at the same microsecond in captured stderr. Any sibling test that bound those names via "from backend.app.main import _timelapse_baselines" before the reimport now held a reference to the OLD dict; production code (reached via "from backend.app.main import on_print_start") resolved the symbol through the NEW module instance. Production mutated the new dict, the test read the old one, the assertion saw None / un-mutated mock_archive. Locally with -n 30, xdist load-balanced test_code_quality.py to a different worker process so the collision never happened (which is why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest made the collision deterministic. Fix: drop the "del sys.modules[name]" step. importlib.import_module already returns the cached module if cached, or runs the import machinery if not — either way, any import-time error surfaces. The "fresh import" framing was theatre; in practice every module in the list is already imported by other tests/fixtures before this test runs, so we were never actually getting a fresh import anyway — just destruction. CI workflow tightening (separate concern, same PR since both touch the test infrastructure): - Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar" lines per worker were eating ~30-60s of stdout I/O on 2-vCPU runners. --tb=short is sufficient for failure context. - Sharded backend-tests into a 4-way matrix via pytest-split (new dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner; all 4 run in parallel so wall-clock drops from 362s -> ~100s. - fail-fast: false on the matrix so a single failing shard doesn't hide failures in the other three — PRs see the complete failure picture in one push. |
||
|
|
22f222e4ac |
fix(test): snapshot _timelapse_baselines inside the patch context to dodge CI race
test_running_observed_captures_baseline_on_restart_recovery was reading _timelapse_baselines.get(1) after the patch() with-block exited. Locally and under low parallelism this works fine — the dict still holds what _capture_timelapse_baseline_at_start wrote. CI under xdist's default load-balancing scheduling intermittently saw the dict empty by the time the top-level assert ran, even though the production code logged "Baseline at print start: 3 video files for printer 1" right before returning. The duplicate log line at the same microsecond in the captured stderr is the tell — module state is being re-touched between the handler completing and the test asserting, almost certainly via the session-scoped event_loop fixture in conftest.py interacting badly with the per-file autouse _clear_baselines teardown of a sibling test on the same worker. The test is verifying the handler captured the baseline at the moment it returned, so capture the relevant value at exactly that point — inside the with-block, immediately after the await. That's immune to whatever happens to the module-level dict afterward. |
||
|
|
eb98521e93 |
fix(test): use /nonexistent/ instead of /tmp/ to satisfy Bandit B108
The test_returns_empty_when_3mf_missing test sets a deliberately
non-existent file_path on a PrintArchive to verify
compute_deficit_for_queue_item handles the missing-3MF branch
gracefully. The path just needs to fail an existence check — the
/tmp/ prefix was incidental.
Bandit B108 ("insecure temp file usage") regex-matches /tmp/,
/var/tmp/, and /dev/shm/. Dropping /tmp/ in favour of /nonexistent/
keeps the test behaviour identical (still a guaranteed-missing
path, still triggers the missing-file branch) while clearing the
GitHub Advanced Security finding on PR #1514 without adding a
# nosec annotation.
|
||
|
|
7fbde1ea3d |
test(docker): include gcode_viewer/ in the test image so the packaging assertion actually runs
Dockerfile.test only COPYed backend/ and pyproject.toml, so the integration test at tests/integration/test_gcode_viewer.py:63 silently pytest.skip'd in every Docker run with "gcode_viewer/ index.html not present at /app/gcode_viewer/index.html". That was deliberate fallback behaviour for unit-test environments where the assets are intentionally absent, but in CI it meant the #1218 packaging regression (3D Preview returning {"detail":"Not Found"} because the embedded PrettyGCode viewer wasn't bundled into the prod image) had no test guarding against a recurrence — the test that was supposed to catch it was the one being skipped. Add COPY gcode_viewer/ ./gcode_viewer/ to the backend-test stage, matching the path the production Dockerfile uses (static_dir.parent / "gcode_viewer" = /app/gcode_viewer/) so the assertion runs against the same layout the app sees at runtime. Path-anchored comment in the Dockerfile so a future maintainer doesn't strip the COPY as unused. |
||
|
|
a64df5a922 | Updated CHANGELOG | ||
|
|
b9d51ffd80 |
chore(deps): floor-pin starlette>=1.0.1 against PYSEC-2026-161
pip-audit reported starlette 1.0.0 in the dev venv. starlette is transitive via fastapi, whose range still admits 1.0.0, so the resolver was silently picking the vulnerable build. Same floor-pin strategy as the existing idna/urllib3 entries — direct pin in requirements.txt with a why-comment so it isn't mistaken for an unused line and dropped later. Verified clean: pip-audit reports "No known vulnerabilities found" after the upgrade (starlette 1.0.0 → 1.1.0 locally). |
||
|
|
3b9633a178 |
● feat(support): include sanitized connection / VP / log-health diagnostics in support bundle and bug report (#1506 follow-up)
The three diagnostic surfaces shipped earlier this month ( |
||
|
|
17602774f3 | Updated BACKERS.md | ||
|
|
03896d1af0 |
fix(scheduler): use inventory weight for "Prefer Lowest Filament" sort (#1508)
Reporter has a P1S with an inventory spool cloned to slot 1 and the
original (much further used) in slot 4, the preference enabled, and
the dispatch picked slot 1 every time. The sort's been blind to
Bambuddy inventory weights — it reads MQTT `tray.remain`, the
printer firmware's RFID-decremented value, which has two limitations:
- Bambu RFID only. Non-RFID spools report -1 and get clamped to a
sentinel; multiple non-RFID trays then tie in the sort and Python's
stable sort collapses to AMS-slot insertion order, so slot 1 wins.
- Even when set, it's the printer's counter, not Bambuddy's
`label_weight - weight_used` (internal) or Spoolman's
`remaining_weight`. The two diverge whenever the user re-spools,
swaps cardboard, or runs a print outside Bambuddy.
The reporter is on internal-inventory mode with non-RFID spools — both
failure modes apply, hence slot 1 every time.
Fix: when a slot is bound to an inventory spool the inventory record's
remaining weight becomes the sort signal. New async helper
`_build_inventory_remain_overrides(db, printer_id, loaded)` returns
`{global_tray_id: remaining_grams}` for bound slots — internal mode
joins SpoolAssignment → Spool once per dispatch; Spoolman mode joins
SpoolmanSlotAssignment then reuses `_spoolman_remaining_grams` from
filament_deficit.py for parity.
New `_prefer_lowest_sort_key` does a two-tier comparison: inventory-
tracked spools sort BEFORE MQTT-only spools, then ascending by
remaining within each tier, then ascending by ams_id*4+tray_id as the
deterministic slot tie-breaker. The tier flag dominates so grams
(inventory) and percent (MQTT) never get cross-compared — no unit
conversion needed.
MQTT-only behaviour is preserved exactly: remain=-1 still maps to the
101 sentinel and slot order still decides on ties. Users who haven't
bound any inventory spool see no change. The DB lookup runs only when
prefer_lowest_filament is enabled.
External / VT slots are skipped (tracked separately from AMS bindings).
|
||
|
|
eae96da56e |
fix(camera): probe ffmpeg for the right RTSP socket-timeout flag (#1504)
A previous attempt swapped `-timeout` → `-stimeout` unconditionally to
fix EADDRINUSE on the reporter's transitional ffmpeg. That broke every
install on a modern ffmpeg (5+/6+/7+) — current Debian/Ubuntu/Homebrew
— where `-stimeout` was removed and `-timeout` is back to meaning
socket I/O. Verified locally: `ffmpeg -stimeout ...` errors
"Unrecognized option 'stimeout'" on ffmpeg 7.1.
install on a modern ffmpeg (5+/6+/7+) — current Debian/Ubuntu/Homebrew
— where `-stimeout` was removed and `-timeout` is back to meaning
socket I/O. Verified locally: `ffmpeg -stimeout ...` errors
"Unrecognized option 'stimeout'" on ffmpeg 7.1.
ffmpeg has shipped THREE arrangements of this option over time and
Bambuddy supports the full range:
- Pre-deprecation (early 4.x and earlier): `-timeout` is socket I/O.
- Transitional (~late-4.x, Jammy-era): `-timeout` is deprecated and
repurposed to RTSP listen-mode timeout; any non-zero value implies
`-listen`, which makes ffmpeg bind the TLS-proxy port and fail with
EADDRINUSE. `-stimeout` is the replacement socket I/O option.
- Modern (5.x / 6.x / 7.x): `-stimeout` REMOVED. `-timeout` is back to
socket I/O — the original meaning.
So no single literal is correct on all installs.
Fix: `rtsp_socket_timeout_flag()` in services/camera.py probes
`ffmpeg -h demuxer=rtsp` once and picks `-stimeout` when ffmpeg
advertises it (covers transitional + older builds that kept it as an
alias), else `-timeout` (modern + pre-deprecation). Cached at module
level for the process lifetime — ffmpeg doesn't swap mid-run.
The function returns the option name without a leading dash; callers
prepend it themselves so a formatting bug can't pass an empty flag.
Wired into both RTSP ffmpeg call sites in lockstep: routes/camera.py
(printer camera) and services/external_camera.py (external RTSP),
which use the same TLS-proxy + ffmpeg pattern and would hit the same
regression on either ffmpeg cohort.
Tests: 8 in test_ffmpeg_rtsp_timeout_flag.py — 6 probe unit tests
(prefers stimeout when advertised, falls back to timeout on modern,
defaults to timeout when ffmpeg missing or probe raises, caches across
calls, trailing-space substring guard against `-listen_timeout`
false-positives), 2 parametrised guards against either RTSP ffmpeg
argv re-hard-coding a literal instead of consuming the probe. 37
probe + existing external-camera tests green.
|
||
|
|
6591fc011f |
fix(slicer): filter process / filament presets by nozzle diameter too (#1325 follow-up #2)
After the @BBL name fallback landed, IndividualGhost1905 reported that
an X2D 0.4 selection still mixed 0.2 / 0.6 / 0.8 nozzle process variants
into the main dropdown. The fallback's two extractors —
extractPrinterPresetModel and extractBblToken — both ended their regex
with `\s+[\d.]+\s*nozzle\s*$` and discarded the match, reducing
"Bambu Lab X2D 0.4 nozzle" and "0.40mm Strength @BBL X2D 0.8 nozzle"
to the same "X2D" string. Match was model-only; nozzle ignored.
Bambu's naming convention: 0.4 is the default and DROPS the suffix; 0.2
/ 0.6 / 0.8 carry an explicit "<size> nozzle" segment. So a process
preset with no suffix is implicitly 0.4 — not "any nozzle".
Have both extractors return { model, nozzle }, parsing the suffix out
instead of stripping. classifyByBambuName then requires both model AND
nozzle to compare equal; a null process nozzle counts as "0.4" per the
convention above. Differing nozzles fall into the existing "Other
printers" group — no new group label.
The bundle path was already nozzle-correct: a .bbscfg is scoped to one
printer-preset-name including its nozzle, and the bundle-side exact
match is therefore nozzle-aware. Only the @BBL name fallback needed
fixing. The `compatible_printers` tier is also unaffected (Bambu's
bundled `compatible_printers` lists include the full printer-preset
name with nozzle, so disambiguation already works).
If the selected printer preset name has no parseable nozzle (non-Bambu
/ hand-typed), the matcher degrades to model-only. Bambu printer
presets always carry one in practice; this is defensive.
9 new tests cover the matrix:
- 0.4 printer ↔ no-suffix process: match
- 0.4 printer ↔ 0.6 / 0.8 process: mismatch
- 0.6 printer ↔ 0.6 process: match
- 0.6 printer ↔ no-suffix process (=0.4): mismatch
- same rule applied to filament presets
- explicit "0.4 nozzle" suffix on process still matches 0.4 printer
- wrong-model still mismatches even when nozzles agree
- no-nozzle printer name degrades to model-only match
One existing test ("handles a trailing nozzle-size suffix on the @BBL
tag") had asserted that a 0.6-nozzle process matched a 0.4 printer —
the exact reporter complaint. Reframed: matching 0.4-suffix still
matches, 0.6-suffix now mismatches.
|