From 8e241915aee8f51cda0d5e9ad42435bcd23398e7 Mon Sep 17 00:00:00 2001 From: maziggy Date: Sat, 16 May 2026 09:47:05 +0200 Subject: [PATCH] fix(docker): pin matplotlib cache to /tmp so the STL thumbnail generator stops logging EPERM MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Matplotlib (imported lazily by stl_thumbnail.py) tried to create its font/style cache at $HOME/.config/matplotlib on first STL upload. HOME=/app per the Dockerfile but /app is root-owned and not writable by the PUID:PGID the entrypoint drops to, so matplotlib logged "Permission denied" and fell back to /tmp/matplotlib-* — wiped on every restart, paying the font-scan cost again on the next STL. Add ENV MPLCONFIGDIR=/tmp/matplotlib to make the cache directory writable and persistent across the container's lifetime. /tmp is writable by any uid, so this works regardless of PUID. --- CHANGELOG.md | 2 ++ Dockerfile | 10 ++++++++++ 2 files changed, 12 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1e047204..bdc45a9e6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,8 @@ All notable changes to Bambuddy will be documented in this file. - **Slice modal: pick the build plate (#1337, reported by @digitalskies)** — Slicing a plain STL through the integrated slicer always defaulted to whatever `curr_bed_type` lived in the chosen process preset (typically `Cool Plate`), which the slicer CLI then rejected for high-temp filaments with `Plate 1: Cool Plate does not support filament 1`. The user had no way to switch plates short of cloning the process preset in BambuStudio, which defeats the point of the in-app slicer. The Slice modal now exposes a `Build plate` dropdown with the six canonical BambuStudio / OrcaSlicer plates (Cool Plate, Cool Plate SuperTack, Engineering Plate, High Temp Plate, Textured PEI Plate, Smooth PEI Plate) plus an explicit `Auto (use process preset)` option that preserves the previous behavior. The dropdown sits between Process profile and Filament rows so it stays visible regardless of how many filament slots the picked plate uses (a long filament list would otherwise push it off the modal's `max-h-[85vh]` scroll viewport) and is **always enabled** — including when the user picks a Printer Preset Bundle from the top BundlePicker. When the user picks a specific plate, the new `bed_type` field on `SliceRequest` ([`backend/app/schemas/slicer.py`](backend/app/schemas/slicer.py)) flows through the dispatcher via two paths: (1) **resolved-preset path** — the route helper `_patch_process_bed_type` in [`backend/app/api/routes/library.py`](backend/app/api/routes/library.py) overwrites `curr_bed_type` on the resolved process JSON before forwarding to the sidecar (no preset cloning required); (2) **bundle dispatch path** — `slice_with_bundle` in [`backend/app/services/slicer_api.py`](backend/app/services/slicer_api.py) adds a `bedType` form field to the sidecar multipart so the sidecar can pass `--curr_bed_type` through to the CLI, which lets the override take effect even though Bambuddy can't patch the bundle's process JSON locally (the sidecar materialises it from the stored .bbscfg). Sidecar versions that don't recognise the field silently no-op — the slice still runs, just with the bundle's default plate; the slicer-API fork at maziggy/orca-slicer-api will need the matching change for the bundle path to take full effect. **i18n parity:** 8 new keys (`slice.bedType.{label,auto,coolPlate,coolPlateSuperTack,engineering,highTemp,texturedPEI,smoothPEI}`) added to all 8 locales — full German translation, English fallbacks elsewhere per project convention. **Regression tests:** 4 in [`test_slice_request_bed_type.py`](backend/tests/unit/test_slice_request_bed_type.py) (`bed_type` defaults to None, accepts the six canonical strings, rejects overlong input via the schema's `max_length=64`; `_patch_process_bed_type` overwrites an existing value, adds the field when missing, and returns the input unchanged for malformed JSON or non-dict roots), 4 in [`test_library_slice_api.py`](backend/tests/integration/test_library_slice_api.py) (resolved-preset path: with `bed_type` set, the sidecar receives `"curr_bed_type": "Textured PEI Plate"` in the presetProfile multipart part; without it, `curr_bed_type` stays out of the body entirely. bundle dispatch path: `bedType` form field carries the override through to the sidecar; omitting `bed_type` keeps the form field out of the request so the bundle's own `curr_bed_type` is preserved), 2 in [`SliceModal.test.tsx`](frontend/src/__tests__/components/SliceModal.test.tsx) (dropdown selection puts `bed_type` on the request; leaving it on Auto omits the field). 59 backend slice tests + 34 SliceModal tests pass; build and i18n parity script clean. ### Fixed +- **Matplotlib no longer logs `Permission denied: /app/.config` on every container start** — Matplotlib (imported lazily by the STL thumbnail generator in [`backend/app/services/stl_thumbnail.py`](backend/app/services/stl_thumbnail.py) when a user uploads an `.stl` file to the library) tries to create its font/style cache at `$HOME/.config/matplotlib` on first import. `HOME` is pinned to `/app` in the Dockerfile so containers with `pwd.getpwuid()` failures (PUID-mapped uids without a local passwd entry) still have a writable home — but `/app` itself is root-owned and not writable by the PUID:PGID the entrypoint drops to, so the first STL upload after a container start logged an `EPERM` warning and matplotlib fell back to a fresh `/tmp/matplotlib-*` dir. Functionally harmless (thumbnails still rendered) but it cluttered every support bundle and forced matplotlib to re-scan system fonts on every restart (~1-2 s per first-STL-upload). **Fix:** added `ENV MPLCONFIGDIR=/tmp/matplotlib` to the Dockerfile so matplotlib uses a guaranteed-writable cache dir up front. `/tmp` is writable by any uid so this works regardless of PUID, and the cache survives the container's lifetime so the font scan only pays its cost once per container. + - **BambuStudio now sees AMS / vt_tray / net info from a virtual printer without requiring a printer power-cycle** ([#1371](https://github.com/maziggy/bambuddy/issues/1371), reported by @Andlar94) — Symptom on a non-proxy VP (the user's A1 in `print_queue` mode): connecting BambuStudio to the VP showed no AMS / external spool info on the Device page; the only workaround was to power the printer off and back on while BambuStudio was open, after which the info populated for one window. **Root cause** in `MQTTBridge._on_printer_raw` at [`backend/app/services/virtual_printer/mqtt_bridge.py`](backend/app/services/virtual_printer/mqtt_bridge.py): the bridge's cache of the real printer's `push_status` was `self._latest_print_state = copy.deepcopy(print_data)` — a wholesale replacement on every incoming push. Bambu firmware sends two shapes of `push_status`: full pushall responses (on `pushall` request / printer reconnect) include AMS / vt_tray / net.info / lights_report, and ~1 Hz incremental updates with just the fields that changed (temperatures, fan speeds, wifi signal). The first incremental push after a pushall therefore wiped AMS info from the bridge cache, and BambuStudio (which reads the cache via the VP's own 1 Hz status push) saw a stripped-down state with no AMS visible until the next pushall — typically only on a manual printer power-cycle, which forces Bambuddy to reconnect and re-issue `pushall`. **Fix:** in the cache update path, preserve a small set of "slicer-visible sticky" top-level keys from the previous cache when the incoming push doesn't include them: `ams`, `vt_tray`, `ams_extruder_map`, `mapping`, `net`, `ipcam`, `lights_report`. Mirrors the same preservation pattern Bambuddy itself already uses for its own internal `state.raw_data` at [`bambu_mqtt.py:2686-2711`](backend/app/services/bambu_mqtt.py); without that, even Bambuddy's own UI would have shown blank AMS in the same way after an incremental push. The new sticky-key set adds three entries (`net`, `ipcam`, `lights_report`) that the slicer specifically cares about: BambuStudio reads `net.info[*].ip` for the FTP destination IP (which the bridge then rewrites to the VP bind IP), uses `ipcam.rtsp_url` for the camera mirror, and renders `lights_report` for the chamber-light toggle. Note: the fix only covers the typical "incremental push omits the sticky key" case — if a future firmware sends a *partial* AMS list (e.g. only one unit's tray subset), that incoming partial push would still replace the cached AMS. That's a rarer scenario and would need full Bambuddy-style per-unit deep-merge; deferred until anyone hits it. **Regression tests** in [`backend/tests/unit/test_vp_mqtt_bridge.py`](backend/tests/unit/test_vp_mqtt_bridge.py): new `test_incremental_push_preserves_ams_from_previous_cache` seeds the cache with a full pushall payload (AMS + vt_tray + lights_report), fires a temps-only incremental push, and pins that all three sticky fields survive with their original values; new `test_incoming_ams_update_replaces_cached_ams` pins the counterpart — when an incoming push DOES include `ams`, the cached value is replaced, so the preservation doesn't shadow real AMS state changes. All 29 mqtt_bridge tests + 158 in the wider virtual-printer sweep pass; ruff clean. - **Queue items no longer get permanently stuck in `printing` status when the printer was in `FINISH` state at dispatch time, and direct-dispatch (Library → Print) no longer reports false success in the same scenario** ([#1370](https://github.com/maziggy/bambuddy/issues/1370), reported by @Martinnygaard) — Symptom: queue page shows `Busy: ` even though the printer is connected, idle, and `awaiting_plate_clear=False`; no new prints will dispatch to it until the user manually deletes or reassigns the queue row. Reproducible by queueing (or directly dispatching) onto a printer that still has the un-dismissed "Print complete" prompt from a prior job. **Root cause** in `_watchdog_print_start` at [`backend/app/services/print_scheduler.py`](backend/app/services/print_scheduler.py) **and the parallel `_verify_print_response` at [`backend/app/services/background_dispatch.py`](backend/app/services/background_dispatch.py)**: the post-dispatch verifiers both treated *any* `gcode_state` transition away from `pre_state` as proof that the printer had accepted the `project_file` command. In the reporter's bundle, item 6 dispatched while printer 3 was in `FINISH` (residual from item 3 earlier that day) — firmware silently rejected the new `project_file` because the previous-print prompt was still up, and ~2 minutes later the user manually dismissed the screen prompt, putting the printer into `IDLE`. The watchdog saw `state != pre_state` and returned early as "command landed", but `FINISH → IDLE` is the user dismissing a prompt, **not** the printer accepting our project_file — so the queue row stayed at `'printing'` indefinitely and the scheduler's busy-printer seed (`SELECT printer_id FROM print_queue WHERE status='printing'` in [`print_scheduler.py:166-171`](backend/app/services/print_scheduler.py)) permanently marked printer 3 as busy. The same broad-transition bug existed in `_verify_print_response`, which would have caused direct-dispatch (Library → Print) onto a FINISH-state printer to report false success — silently failing to print while the UI showed the dispatch as complete. **Fix:** in both verifiers, narrow the "command landed" check to an allow-list of active-print states (`PREPARE` / `SLICING` / `RUNNING` / `PAUSE`) instead of "any state that isn't `pre_state`". Inactive states (`IDLE`, `FINISH`, `FAILED`) no longer short-circuit the early return. The `subtask_id`-advance signal stays as-is in both verifiers — it remains the definitive "command landed" path for H2D firmware that sits at `FINISH` for ~50 s after accepting `project_file` before transitioning to `PREPARE` (#1078 stays green in both). **Resilience hardening alongside the fix:** the watchdog's revert commit and `printer_manager._persist_awaiting_plate_clear` now run through `run_with_retry` ([`backend/app/core/database.py`](backend/app/core/database.py)), so SQLite single-writer `database is locked` contention can't silently drop the queue-row revert or the plate-clear gate flag. The revert path returns a tristate sentinel (`"reverted"` / `"already_moved_on"` / `"revert_failed"`) so the post-revert MQTT session-recovery logic only runs when we actually reverted (or the commit failed) — never when `on_print_complete` had already cleared the row, where a forced reconnect could break a healthy concurrent print on the same printer. (Most other queue/archive writes already went through `run_with_retry`; these two were the holdouts that surfaced as repeated `Failed to persist awaiting_plate_clear` warnings in the reporter's bundle.) **Manual recovery for users on 0.2.4** who already have stuck rows: stop Bambuddy, then `sqlite3 /app/data/bambuddy.db "UPDATE print_queue SET status='cancelled', completed_at=datetime('now') WHERE status='printing';"` and restart. **Regression tests** — new `test_reverts_on_finish_to_idle_user_dismissed_prompt` (queue) and `test_returns_false_on_finish_to_idle_user_dismissed_prompt` (direct-dispatch) reproduce the exact reporter scenario on both code paths; new `test_does_not_revert_on_pickup_via_active_state` (queue) and `test_returns_true_on_each_active_print_state` (direct-dispatch) iterate all four active-print states (PREPARE/SLICING/RUNNING/PAUSE) and pin that each one is correctly treated as a valid "command landed" signal. Existing `test_no_revert_if_item_already_completed` was also hardened — it now uses a real client mock and asserts `force_reconnect_stale_session.assert_not_called()`, so the tristate-sentinel guard around the recovery path is pinned (catches the regression I introduced and then fixed during the audit pass). The pre-existing `test_exits_on_state_change` (uses `RUNNING`) and `test_exits_on_subtask_id_change_even_if_state_still_finish` (the #1078 H2D path) both still pass without modification. All 29 watchdog tests across both files + 411 in the scheduler/queue/dispatch/printer-manager sweep + 3161 in the full backend unit suite + 302 in the targeted integration sweep all pass; ruff clean. diff --git a/Dockerfile b/Dockerfile index a86662eff..0e77eef33 100644 --- a/Dockerfile +++ b/Dockerfile @@ -121,6 +121,16 @@ ENV HOME=/app ENV USER=bambuddy ENV LOGNAME=bambuddy +# Matplotlib (imported lazily by the STL thumbnail generator) tries to create +# its font/style cache at $HOME/.config/matplotlib on first import. /app is +# root-owned and not writable by the PUID:PGID the entrypoint drops to, +# which trips an EPERM warning in everyone's logs and forces matplotlib +# to fall back to a per-restart temp dir (paying the font-scan cost on +# every container restart). Pinning the cache dir to /tmp/matplotlib +# silences the warning and keeps the cache alive for the container's +# lifetime. /tmp is writable by any uid, so this works regardless of PUID. +ENV MPLCONFIGDIR=/tmp/matplotlib + EXPOSE 322 EXPOSE 990 EXPOSE 3000