Three more idle-in-transaction / thundering-herd paths from farm testing:
- print_scheduler: _start_print commits before the FTP delete/upload and
_preheat_and_soak commits before the heat-soak wait, so the per-item
session no longer sits idle-in-transaction across preheat + upload.
- cloud/filament-info: rollback the request transaction after the token
read and before the sequential Bambu Cloud calls; single-flight
concurrent misses for the same setting_id through one shared call.
- printers/cover: coalesce identical in-flight cover requests so followers
serve from the cache the leader fills instead of duplicating the
multi-path FTP + 3MF extraction.
Also adds pool_use_lifo (PostgreSQL default on, DB_POOL_USE_LIFO override,
shown in /system/db-pool) so a bursty farm keeps a small hot connection set.
Two or three concurrent UI logins exhausted the PostgreSQL pool on the
reporter's 93-printer farm: QueuePool limit of size 10 overflow 20 reached,
with all 30 sessions idle in transaction on the auth_enabled SELECT. Three
regressions had landed on dev after an earlier configurable-pool change was
reverted and never re-applied (only the route-by-route session fixes were).
- Pool sizing is env-configurable again (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); the PostgreSQL default returns to
20 + 80 with pool_pre_ping and pool_recycle=1800, and GET
/api/v1/system/db-pool reports resolved config + live gauges without
checking out a connection. SQLite unchanged (20 + 200).
- is_auth_enabled caches for 30s again. Only enabled=True is ever cached, so
a stale read can only fail closed (require auth), never open; set_auth_enabled
invalidates immediately. An autouse test fixture resets the module cache
between tests to keep ordering deterministic.
- Every authenticated request checked out two pooled connections: the
permission dependency held one and the revoked-jti check opened another.
is_jti_revoked now reuses the caller's session; the token dependencies and
the auth-middleware gateway were restructured to open one session and pass
it in, so each request makes a single checkout.
Large PostgreSQL farms exhausted the fixed pool (pool_size=10 +
max_overflow=20): with ~93 printers every connection sat idle in
transaction and unrelated requests waited out the 30s pool timeout or
failed in the auth middleware.
- Make pool sizing env-configurable (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); raise the Postgres default to
20 + 80 with pool_pre_ping + pool_recycle=1800.
- Cache the auth_enabled probe (30s) to drop a per-request DB round-trip.
Only enabled=True is cached, so staleness fails closed; set_auth_enabled
invalidates immediately.
- Add GET /api/v1/system/db-pool exposing resolved config + live
checked_out/checked_in/overflow gauges without consuming a connection.
Session-hygiene (connections held across MQTT/FTP/camera/3MF I/O) is a
separate follow-up.
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.
Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.
FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.
Two problems the same bundle exposed:
A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.
The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.
Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).
Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
Administrators system group sync
- Fresh installs already bootstrap with ALL_PERMISSIONS, so they always have
every permission. Upgrades previously only got what one-off backfill blocks
in seed_default_groups() explicitly listed (library:purge, archives:purge,
the OWN/ALL read-flag block, orca_cloud:auth, pipelines:*). Any Permission
enum member added without a matching block silently stayed missing on
existing admin rows. The most recent gap was printer_sensor_history:read
(Sensor History charts returned 403 for upgraded admins).
- seed_default_groups() now syncs Administrators to ALL_PERMISSIONS on every
startup: append every Permission value that isn't already on the row.
Additive only -- hand-added custom permissions are preserved.
- The pure-admin one-off backfills (library:purge / archives:purge block,
the OWN/ALL + orca_cloud:auth + legacy-read-flag block, the Administrators
branch of the pipeline backfill) are retired since the sync subsumes
them. Non-admin backfills (Operators / Viewers OWN-tier reads, Operators
orca_cloud:auth, pipelines for non-admin groups, makerworld:*, clear_plate
cross-group adders) are untouched.
- Tests: test_administrators_printer_sensor_history_read_backfilled
(regression for the reported gap),
test_administrators_sync_covers_every_current_permission (generic
invariant -- any future new permission lands on admin without needing
a one-off test), test_administrators_sync_is_additive_only (custom
permissions preserved). 12/12 backfill-migration + 102/102 broader
permission tests green; ruff clean.
Pipelines runs dashboard
- PipelineRunsPage.tsx: the Pipeline / Status / Target filter row's three
native <select> elements are replaced with a bambu-themed FilterDropdown
(button trigger, floating menu, optgroup-style headers for the Target
picker, hover + selected states with a check mark, closes on outside
click and Escape). Same value/onChange contract -- visual only.
- SlicerPipelinesPanel.tsx: wrap list?.pipelines ?? [] in useMemo so the
reference is stable when the data is stable. Fixes the
react-hooks/exhaustive-deps warning where the inline fallback returned
a fresh empty array every render, invalidating both downstream useMemo
caches (target-options + filtered-pipelines list).
chore(i18n): extend parity gate to all locales with strict/info tiers
Previously the script only inspected en/zh-CN/zh-TW, leaving de/fr/it/ja/pt-BR
drift invisible. Now locales are auto-discovered from src/i18n/locales/, and a
STRICT list (de, zh-CN, zh-TW — currently in parity) gates CI while the rest
report informationally until their drift is caught up. ja notably has 27 real
placeholder bugs worth fixing before promotion to strict.
Native-install upgrade via the in-app Apply Update button got the new
code in via `git reset --hard origin/main` but then logged
ERROR: Could not open requirements file:
[Errno 2] No such file or directory: 'requirements.txt'
and continued. The new deps never installed, leaving the user with
new code but stale dependencies — surfaces as cryptic import errors
on the next restart.
Root cause: `pip install -r requirements.txt` ran with
`cwd=settings.base_dir`. On a native install, systemd sets
DATA_DIR=$INSTALL_PATH/data so base_dir resolves to the data dir
(e.g. /opt/bambuddy/data), not the source tree. Pip doesn't walk up
looking for the requirements file the way git walks up looking for
.git, so it fails. Same bug affected the optional npm step
(`frontend_dir = base_dir / "frontend"` doesn't exist).
Fix: introduce `settings.app_dir` pointing at the source-tree root
(distinct from `base_dir` only on native installs) and run pip +
npm with `cwd=settings.app_dir`. Git ops keep using `base_dir`
because they already work (git walks up).
Docker users were unaffected — Docker doesn't use the in-app updater
(image pull replaces it).
Regression test in test_updates_api.py mocks every subprocess in
_perform_update, captures their cwd, and asserts the pip step runs
in app_dir and that requirements.txt actually exists there. Any
future refactor that re-introduces cwd=base_dir for the pip step
fails CI before another user trips over it.
Adds an optional slicer-api/ Compose stack and wires Bambuddy's File
Manager, Archives, and MakerWorld pages to a new server-side Slice flow.
Slicing runs as an in-memory background job (POST returns 202 + job_id,
polled via GET /api/v1/slice-jobs/{id}) so a multi-minute slice no
longer pins the modal; result lands as a new .gcode.3mf in the same
folder (or new archive for archive sources) with the embedded
thumbnail extracted.
Backend
- New services: slice_dispatch (in-memory dispatcher, 30min retention
sweep) and slicer_api (HTTP bridge with 4xx/5xx/connection error
split that drives the 3MF embedded-settings fallback retry path).
- New schemas: SliceRequest, SliceResponse, SliceArchiveResponse,
SliceJobEnqueueResponse.
- New routes: POST /library/files/{id}/slice,
POST /archives/{id}/slice, GET /api/v1/slice-jobs/{id} (gated on
LIBRARY_READ since job IDs are sequential and the body leaks source
filenames and result IDs).
- AppSettings + env defaults: use_slicer_api, orcaslicer_api_url,
bambu_studio_api_url. DB-stored values override env defaults.
Frontend
- New SliceModal handles preset gating; enqueues then closes
immediately.
- New SliceJobTrackerProvider polls active jobs at app level, surfaces
a single toast per job (queued -> running -> completed / failed)
and invalidates library/archives queries on terminal status.
- Settings -> Workflow -> Slicer card: preferred slicer dropdown,
Use Slicer API toggle, contextual sidecar URL field.
- File Manager / Archives / MakerWorld get a Slice button gated on
the Use Slicer API setting.
- gcode-viewer adapter learns ?library_file=<id> so sliced library
files preview inline.
i18n
- New slice.* and settings.{useSlicerApi,slicerCard,orcaslicerApiUrl,
bambuStudioApiUrl,slicerApiUrlDescription,useSlicerApiDescription}
+ fileManager.noPermissionSlice keys across all 8 locales (en, de,
fr, it, ja, pt-BR, zh-CN, zh-TW). English fully translated, German
fully translated, the other six seeded with English fallbacks
pending native translation.
Tests
- 10 backend integration tests in test_library_slice_api.py covering
validation (404/400), happy-path enqueue, sidecar-down, 3MF
embedded-settings fallback, STL no-fallback, and preset-error ->
failed job paths.
- New unit tests in test_slicer_api.py for the HTTP bridge.
- 5 new SliceModal frontend tests covering preset gating, library +
archive enqueue paths, error surface, and preset-load failure.
- Existing SettingsPage tests adjusted: slicer dropdown asserts now
switch to the Workflow tab first; added a beforeEach URL reset so
one test's tab click doesn't bleed into sibling tests.
Sidecar
- New slicer-api/ folder is self-contained and optional. Two services
(orca-slicer-api on 3003, bambu-studio-api on 3001 behind --profile
bambu) build via Docker git-build-context from
maziggy/orca-slicer-api@bambuddy/profile-resolver. The fork patches
the OrcaSlicer CLI's profile compatibility quirks (inherits-chain
resolver, from:User -> system rewrite, '# ' clone-prefix strip,
sentinel-value strip) empirically required to slice real GUI
exports without segfaulting the CLI.
Docs
- CHANGELOG entry under [0.2.4b1] - Unreleased Added.
- README File Manager bullet for the new server-side Slice button.
- bambuddy-website features.html: new card under "Configurable Slicer".
- bambuddy-wiki: new page features/slicer-api.md + nav entry +
features index card.
Notes
- Opt-in: with Use Slicer API off, the existing "open in desktop
slicer via URI" flow is the default and unchanged.
- 3MF inputs that segfault the CLI on --load-settings transparently
retry with embedded settings; the resulting job carries
used_embedded_settings: true.
- Sliced files always export as .gcode.3mf so File Manager picks up
the embedded thumbnail; file_type is set to "gcode" (blue badge).
Bambuddy can now use an external PostgreSQL database via the
DATABASE_URL environment variable. SQLite remains the default.
Dialect-aware helpers handle upserts, PRAGMAs, FTS (FTS5 vs
tsvector+GIN), backup/restore, and health checks. All migration
blocks use savepoints to prevent Postgres transaction poisoning.
Backups are always portable SQLite format regardless of backend.
Cross-database restore imports SQLite backups into PostgreSQL
with automatic boolean/datetime conversion, NOT NULL default
filling, and FK constraint handling.
The kiosk touchscreen has no way to hard-refresh, and the service worker
served stale cached JS after updates. SpoolBuddy pages now unregister
any existing SW and skip registration entirely. Regular desktop/mobile
users still get the SW. Restored kiosk restart in SSH update flow since
SW is no longer an obstacle.
Merged version, status, and update check into a single card. Buttons
are side by side when up to date. Reduced padding and font sizes.
Removed separate "complete" banner since status is now cleared on
re-registration. SSH Setup section is smaller and more compact.
The SSH update set status to "complete" after the daemon had already
restarted and re-registered, overwriting the cleared state so it stuck
forever. Removed the post-restart "complete" write — daemon
re-registration is now the completion signal, clearing any update status.
Sends persistent notifications to the HA dashboard using the existing
HA connection from Settings. Zero config — just select "Home Assistant"
as provider type. Users can forward notifications to mobile via HA
automations.
Floating bug report button submits issues via bambuddy.cool relay (no GitHub
token needed locally). Collects 30s debug logs with printer push_all, sanitizes
all sensitive data, uploads logs as files to GitHub. Screenshot upload/paste/drag
with JPEG compression. Translated into all 7 languages. Includes 21 tests.
Three independent code paths wrote to archive.cost with conflicting
strategies, causing the same model to produce different prices on each
reprint (e.g. £0.77, £1.54, £2.03).
- Remove add_reprint_cost (Path 3) — redundant cost accumulation that
double-counted on top of usage tracker
- Fix usage tracker (Path 2) — compute cost from current session's
results only, not a SUM of all historical SpoolUsageHistory rows
- Remove _reprint_archives tracking set from main.py
- Update test mocks to match simplified query pattern
archive.cost now always reflects the cost of a single print.