Two or three concurrent UI logins exhausted the PostgreSQL pool on the
reporter's 93-printer farm: QueuePool limit of size 10 overflow 20 reached,
with all 30 sessions idle in transaction on the auth_enabled SELECT. Three
regressions had landed on dev after an earlier configurable-pool change was
reverted and never re-applied (only the route-by-route session fixes were).
- Pool sizing is env-configurable again (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); the PostgreSQL default returns to
20 + 80 with pool_pre_ping and pool_recycle=1800, and GET
/api/v1/system/db-pool reports resolved config + live gauges without
checking out a connection. SQLite unchanged (20 + 200).
- is_auth_enabled caches for 30s again. Only enabled=True is ever cached, so
a stale read can only fail closed (require auth), never open; set_auth_enabled
invalidates immediately. An autouse test fixture resets the module cache
between tests to keep ordering deterministic.
- Every authenticated request checked out two pooled connections: the
permission dependency held one and the revoked-jti check opened another.
is_jti_revoked now reuses the caller's session; the token dependencies and
the auth-middleware gateway were restructured to open one session and pass
it in, so each request makes a single checkout.
Large PostgreSQL farms exhausted the fixed pool (pool_size=10 +
max_overflow=20): with ~93 printers every connection sat idle in
transaction and unrelated requests waited out the 30s pool timeout or
failed in the auth middleware.
- Make pool sizing env-configurable (DB_POOL_SIZE / DB_MAX_OVERFLOW /
DB_POOL_TIMEOUT / DB_POOL_RECYCLE); raise the Postgres default to
20 + 80 with pool_pre_ping + pool_recycle=1800.
- Cache the auth_enabled probe (30s) to drop a per-request DB round-trip.
Only enabled=True is cached, so staleness fails closed; set_auth_enabled
invalidates immediately.
- Add GET /api/v1/system/db-pool exposing resolved config + live
checked_out/checked_in/overflow gauges without consuming a connection.
Session-hygiene (connections held across MQTT/FTP/camera/3MF I/O) is a
separate follow-up.
Extends the appliance endpoint that landed in the previous commit with a
time_synced field, sourced from /run/bambuddy/time-synced (the appliance's
ntp-gate.sh writes this once chronyd reports sync, or with a "warning"
marker after the 3-minute timeout). The RPi 5 has no battery-backed RTC,
so on a fresh boot the system clock is wrong until NTP catches up -- JWT
expiries and TLS certificate validity windows depend on this being right.
Exposing the gate lets the SPA render a "time not synced" indicator while
that's still true and clear it once "ok" comes through.
backend/app/core/local_config.py
New read_ntp_gate(path) function alongside read_local_toml. Three states:
"ok" chrony reported sync within the 3-minute window
"warning" 3-minute timeout elapsed without sync; user already waited
and the wizard proceeded with a degraded clock
None file absent (non-appliance install), OSError, empty content,
unknown marker, or binary garbage -- "unknown / don't gate"
Defensive read mode (errors="replace") survives non-utf8 content without
crashing. Module docstring broadened from "local.toml reader" to "small
readers for appliance-set state files".
backend/app/api/routes/system.py
/system/appliance now returns:
{hostname, timezone, locale, time_synced}
with the same no-auth posture: bootstrap surfaces (i18n init, time-sync
banner) read this before auth might be set up, and the contents are
non-secret (user-set defaults + a public sync flag). The endpoint
docstring expands to explain the RTC motivation -- otherwise the
time_synced field reads like a leftover.
Closes the cross-repo contract started in bambuddy-appliance: the firstboot
wizard writes /etc/bambuddy/local.toml with the user's hostname / timezone /
locale, but nothing on the main app side read it. Hostname + timezone are
already applied by the appliance's firstboot.sh via hostnamectl /
timedatectl. This PR closes the loop for the third field — locale — so the
language the user picked in the wizard actually shows up on first SPA load.
backend/app/core/local_config.py
New module. read_local_toml(path) returns a LocalConfig TypedDict
({hostname?, timezone?, locale?}) parsed from /etc/bambuddy/local.toml.
Defensive on every failure mode -- missing file returns {}, invalid TOML
returns {} + log warning, non-string values dropped with warning. The
reader never raises; a malformed config never blocks startup.
backend/app/api/routes/system.py
New endpoint GET /system/appliance. Returns {hostname, timezone, locale}
with null for any field not present in the TOML. No auth required: the
frontend i18n bootstrap reads this before auth might be set up, and the
contents are user-set defaults, not secrets. The function calls
read_local_toml() with no args (default path) so tests can monkeypatch
the module's read_local_toml reference to inject fixtures.
frontend/src/i18n/index.ts
One-shot applyApplianceLocale() runs after i18n.init(). Gated by a
bambuddy_appliance_locale_consumed localStorage flag so it runs at most
once per appliance. Fetches /api/v1/system/appliance, validates the
returned locale against supportedLngs, calls i18n.changeLanguage if
valid. Silent .catch() because the endpoint absent / unreachable means
non-appliance install or dev environment -- we leave the LanguageDetector's
choice in place. The consumed flag is set on success; future loads skip
the fetch entirely. Won't override a user's explicit language pick (the
language picker writes to a separate localStorage key, bambutrack_language).
datetime.fromtimestamp(ts) and datetime.now() return naive local
datetimes; .isoformat() then emits no tz marker. The frontend's
parseUTCDate helper appends 'Z' to bare strings, treats the value
as UTC, then converts to local for display — applying the local
offset twice. Reporter on UTC+3 saw boot_time +3h ahead while
uptime was correct (uptime is a backend-side delta of two
naive-local values, so the missing tz info cancels out).
Fix: pass tz=timezone.utc to datetime.fromtimestamp and
datetime.now in system.py's boot_time / uptime path, plus the two
adjacent generated_at sites in system.py and support.py.
System -> Uptime / Boot Time read psutil.boot_time(), which on shared-kernel
containers (Docker, LXC, Proxmox containers) is /proc/stat:btime - the host
kernel's boot time, not the container's. Reporter on Proxmox LXC saw the
Proxmox node's uptime instead of the Bambuddy container's.
PID 1 is the container's entrypoint (or the host init on bare metal), and
its create_time is the POSIX wall-clock timestamp of when it started.
Switching to psutil.Process(1).create_time() reports the right value on
containers and matches host boot within a sub-second on bare metal.
Defensive fallback to psutil.boot_time() on psutil.Error / OSError so the
endpoint still returns 200 with the best-available answer if /proc/1/stat
is unreadable (locked-down container, custom seccomp policy).
Two stacked causes under-reported multi-plate prints in the project
rollup and the archive card.
Root cause 1 - parser only read plate 1.
ThreeMFParser._parse_slice_info used root.find(".//plate") and pulled
prediction / weight from that one element. Any multi-plate file's
archive-level print_time_seconds / filament_used_grams reflected
plate 1 alone. The /plates endpoint already looped findall and was
correct, which is why the plate carousel showed the right numbers
while the archive card was wrong.
Fix: loop every <plate> and sum prediction + weight. Per-plate
concepts (plate_number, _plate_index, printable_objects) only set
when there's exactly one plate - for multi-plate exports the
archive represents all plates and a single index doesn't apply at
the file level. bed_type keeps the first plate's value as a
best-effort default. Malformed prediction / weight on individual
plates skip cleanly rather than poison the sum.
Root cause 2 - project rollup aggregated PrintArchive, not the
per-run log.
compute_project_stats and list_projects quick-stats summed
PrintArchive.print_time_seconds / filament_used_grams / cost /
energy_* WHERE project_id. A reprint reuses the source archive row
and writes a new PrintLogEntry, so 3 sequential runs collapsed to 1
archive - and that archive's numbers were already plate-1-only from
cause 1. The Archive Print Log path was already correct because it
drove off print_log_entries (archives.py:420 comment).
Fix: both compute_project_stats and the list_projects quick-stats
block inner-join print_log_entries -> print_archives WHERE
archives.project_id. total_archives becomes COUNT(PrintLogEntry.id),
failed_prints counts runs in failed/aborted/cancelled/stopped,
completed_items is SUM(PrintArchive.quantity) for runs where
status='completed', time/filament/cost/energy from PrintLogEntry.
Orphan log rows (archive_id IS NULL post archive deletion) are
excluded by the inner join.
Same-shape fixes carried forward (no follow-ups per project rule):
system.py system-info totals: total_print_time / total_filament
had the same bug shape - summed PrintArchive directly so reprints
collapsed to one row. Now sums PrintLogEntry.duration_seconds /
filament_used_grams. The semantic shift is also a correctness
improvement: the field now reflects time the printer actually spent
printing, not slicer-estimated time.
archives.py time-accuracy metric: estimate / actual per run where
estimate = PrintArchive.print_time_seconds. Post-parser-fix
multi-plate archives have file-level estimate but per-run actual =
one plate, so ratio = N x 100% for an N-plate file. The calc now
clamps each row to the [50%, 200%] plausibility band before
contributing to the printer-level average; single-plate accuracy
(the case the metric is designed for) stays fully included.
Backfill: users with AMS spool tracking - the reporter's case - have
per-run filament_used_grams from the tracked spool delta, so stats
become correct immediately. Users without tracking fall back to the
archive estimate and undercount until they reprint. Archive card
still reads PrintArchive.filament_used_grams directly so old
multi-plate archives keep plate-1-only numbers until reslice -
forward-only as the reporter accepted.
Adds a passive log-health check that complements the active Connection
Diagnostic. Scans Bambuddy's recent app log against a curated allowlist
catalog of known failure signatures (rejected access code, FTPS :990
timeout, FTPS TLS failure, flapping MQTT, unreachable camera, SQLite
"database is locked" contention), dedupes and classifies each finding
as layer8/environment/bug, and deep-links to the troubleshooting wiki.
Sample log lines are sanitized before they leave the process. Exposed
via GET /system/health and surfaced on two surfaces sharing one
SystemHealthPanel component: a System Health section on the System
page, and inline in the bug reporter when the form opens.
The Add-Printer and Edit-Printer dialogs gained a setup-time pre-flight:
saving runs the connection diagnostic and, on a failed check, warns with
a "save anyway" escape hatch instead of silently saving a printer that
will immediately show offline.
Log read/parse/sanitize primitives extracted from routes/support.py into
a shared services/log_reader.py (behaviour-preserving); affected support
tests repointed accordingly.
Tests: test_log_health.py (11), test_system_api.py (2 new),
SystemHealthPanel + BugReportBubble + AddPrinterPreflight +
EditPrinterPreflight (8 frontend). All strings translated across the 9
locales. Backend ruff clean, full unit suite green, frontend build +
eslint clean, i18n parity green.
Display the active database backend (SQLite or PostgreSQL) and its
version in the Database section of the System Info page. SQLite queries
sqlite_version(), PostgreSQL queries SELECT version() and
pg_database_size(). Helps users verify which database is in use.
Bambuddy can now use an external PostgreSQL database via the
DATABASE_URL environment variable. SQLite remains the default.
Dialect-aware helpers handle upserts, PRAGMAs, FTS (FTS5 vs
tsvector+GIN), backup/restore, and health checks. All migration
blocks use savepoints to prevent Postgres transaction poisoning.
Backups are always portable SQLite format regardless of backend.
Cross-database restore imports SQLite backups into PostgreSQL
with automatic boolean/datetime conversion, NOT NULL default
filling, and FK constraint handling.
- New APIBrowser component with full OpenAPI schema integration
- Fetches and parses /openapi.json automatically
- Groups endpoints by API tags (printers, archives, settings, etc.)
- Expandable endpoint sections with color-coded method badges
- Path parameter, query parameter, and JSON body editors
- Auto-populates request body with schema examples
- Live API request execution with response display
- Response shows status code, timing, and formatted JSON
- Copy response button with clipboard fallback
- Search to filter endpoints across all categories
- Expand All / Collapse All buttons
- Link to Swagger UI (/docs)
- Two-column layout for API Keys tab
- Left: API key management + webhook documentation
- Right: API Browser with dedicated test key input
- Parameter validation
- Shows warning for missing required parameters
- Validates before sending requests to avoid 422 errors
- UX improvements
- "Use in API Browser" button on newly created keys
- Responsive layout (stacked on mobile, side-by-side on xl+)