A bundle described everything except the process it runs in. So a report of
memory climbing over days until the OOM killer fires arrives with no way to
act on it: the numbers that name the mechanism only exist while it is
happening, and by the time anyone asks, the container has been restarted.
The new `process` section carries what actually separates the candidates.
Resident against virtual memory: 650MB RSS with 12.9GB VMS is address
space — thread stacks or allocator arenas — not a heap full of live data,
and that reading is the opposite of the one the reporter drew from the same
figures. Thread count and child-process count then split those two apart,
and a census of live objects by type names what a growing heap is filling
up with. Open files, sockets and uptime round it out.
Three constraints worth keeping:
The heap census is skipped above 2GB. gc.get_objects() materialises every
tracked object, so it costs most on exactly the process that can least
afford it — a bundle generated to diagnose runaway memory must not be the
allocation that tips the host over. Everything else is still collected, and
the skip is recorded with its reason rather than silently omitted.
Children are recorded by executable name only. An ffmpeg command line
carries the camera URL, and with it the camera's password.
Collection runs off the event loop and every metric is independently
best-effort. psutil raises on hardened kernels and in restricted
containers, and the bundle is how someone reports a problem in the first
place — it has to be produced even when half the numbers are unavailable.
This does not fix#2734, and nothing here should be read as having found
its cause. The bundle's own evidence contradicts both proposed causes: the
orphan janitor ran 7 times in 26 days over 725 stream-ends and killed no
orphaned ffmpeg, which is not the #776 signature; and the 5 "database is
locked" errors all fall between two OOM kills, making them a symptom of the
memory pressure rather than a source of it.
Obico's ml_api container takes an optional ML_API_TOKEN environment variable.
With it set, ml_api/auth.py answers a bare 401 to any request whose
Authorization header isn't "Bearer <token>"; with it unset it ignores the
header entirely. Bambuddy never sent one, so pointing it at a protected server
meant deleting the token there — which the reporter had set for their Home
Assistant integration and did not want to undo.
Settings -> Failure Detection gains an ML API Token field. When it is empty no
header is sent, so an unconfigured install's request stays byte-identical to
what shipped before the setting existed.
This failed in the worst possible way, and that is the more important half of
the change. Obico decorates /p/ with token_required but leaves /hc/ open. Test
Connection pinged /hc/, so it reported success against a server that was
rejecting every real detection call, the settings looked right, and detection
silently never ran. The only symptom was a generic "ML API call failed" buried
in the status card.
So the test now proves what it claims. After health passes it probes GET /p/
with no img parameter: the auth decorator runs before the handler, so 401 means
the token was rejected and 422 ("Invalid request params") means it was
accepted. No inference work is done either way. A probe that itself errors
reports the token as unknown rather than as working — the UI says it could not
be checked instead of claiming success.
The detection loop checks for 401 before raise_for_status, so a rejected token
is reported as a rejected token, naming the setting and the environment
variable, instead of surfacing "401 Unauthorized" with no hint of what to do.
The message never contains the token; a test pins that.
The setting name carries "token", so the support bundle's keyword redactor
masks it with no new rule. Resolving "field omitted" to the saved token is the
route's job, keeping test_connection a pure outbound call with no database
access.
Second fix, same issue: support bundles misreported which printers Obico
watches. The bundle split obico_enabled_printers on commas and read an empty
value as "no printers". The settings UI writes a JSON array, and empty means
*all* printers — the default — so a working Obico setup showed obico_enabled
false against every printer in its own bundle. That is the reporter's bundle
exactly, and it points anyone reading it at the wrong subsystem. The bundle now
parses the setting the way ObicoDetectionService does, keeps a comma fallback
for any install that stored the legacy shape, and factors in the global switch.
The reporter's 19-printer farm started prints "one by one", up to an hour apart.
check_queue awaited each dispatch inline, and a dispatch includes the FTP upload,
so every printer queued behind every other printer's transfer despite being an
independent machine. His logs give the arithmetic: 40978500 bytes in 254.1s,
157 KB/s - a Bambu printer's SD write, not the network, is the bottleneck. Nineteen
of those in series is ~80 minutes, and the next upload started 131 ms after the
previous one finished. The delay is linear in fleet size, which is why it got worse
the more printers he selected.
Dispatch is now collected during the (still sequential) selection loop and run
concurrently afterwards, capped by queue_max_concurrent_uploads - Settings ->
Workflow -> Queue & Dispatch, default 4, 1 restores the old behaviour. Every gate
is untouched; only the transfers overlap. The pass still awaits its uploads before
returning: _start_print flips the row pending -> printing only after the upload,
so an early return would let the next tick re-dispatch the same rows.
FTP work moves to its own thread pool. It was on asyncio's default executor -
min(32, cpu+4), six threads on a 2-core NAS, shared with everything else - which
was survivable only while uploads were serial.
Two problems the same bundle exposed:
A printer that accepts project_file but never starts (#1678) was retried forever:
270s watchdog, revert to pending, re-upload the whole file, repeat. Hence his
"printer who, since the morning, still not launch" - and on a farm each lap also
eats an upload slot the other printers are waiting on. Attempts are now counted on
the queue item; after three it fails with a message pointing at the printer instead
of queueing a fourth re-upload.
The debug bundle we asked him for held 4m49s of history. The push_status dumps fired
on every frame rather than on change - several while their own comment claimed
otherwise - which is 27,727 of the bundle's 29,830 lines and rolls 5 MB in under five
minutes on 19 printers. They now log transitions only. The bundle also read just the
live log while three rotated backups sat next to it, under a byte budget four times
larger than the file it was reading.
Migration verified on SQLite and Postgres: idempotent, backfills legacy NULLs
(dispatch_attempts + 1 is NULL for a NULL row, which would silently disable the cap).
Tests: 6 on concurrent dispatch (overlap, cap honoured, 1 == serial, default applies
with no settings row, a failed printer does not cancel its siblings, no early return),
4 on the retry budget, 6 on the bundle's rotated-log span, 7 on the debug gating.
Each verified to fail against the unfixed code - the first end-to-end log assertion I
wrote passed without the fix and had to be tightened.
The "0.0.0.0" written into the support bundle is a JSON sentinel that
scrubs the printer's local IP plus the gateway/peers it sees — not a
socket bind address. Annotate inline so bandit stops flagging it.
The support bundle shipped support-info.json + bambuddy.log, but the raw
shape of the printer's MQTT push_status — the field that blocks per-model
work like AMS Backup detection (deferred in 85fbd7fc) and every vt_tray /
vir_slot / mapping shape regression — was never captured.
Each connected printer now contributes push-status/printer-{i}.json with
{model, firmware_version, captured_at, raw_data}, indexed against
support-info.json["printers"]. Two-pass redaction: a structural walk
drops user-private keys (subtask_name, gcode_file, subtask_id, task_id,
project_id, design_id, profile_id, model_id, gcode_state,
gcode_file_prepare_percent) and rewrites net.info[*].ip to 0.0.0.0
(matches the #1429 VP bridge fix); then the JSON runs through the same
DB-derived sensitive_strings sanitizer the log path uses, catching any
printer name / serial / access code / cloud email that leaked into a
nested string field.
print.cfg, print.option, ams, vt_tray, vir_slot, mapping,
ams_extruder_map, and hardware fields are all preserved — those are the
fields per-model work needs.
Always-on inside the existing debug-logging-required gate; no opt-in
toggle (the bundle is already user-initiated and downloads locally
before the user chooses to send).
datetime.fromtimestamp(ts) and datetime.now() return naive local
datetimes; .isoformat() then emits no tz marker. The frontend's
parseUTCDate helper appends 'Z' to bare strings, treats the value
as UTC, then converts to local for display — applying the local
offset twice. Reporter on UTC+3 saw boot_time +3h ahead while
uptime was correct (uptime is a backend-side delta of two
naive-local values, so the missing tz info cancels out).
Fix: pass tz=timezone.utc to datetime.fromtimestamp and
datetime.now in system.py's boot_time / uptime path, plus the two
adjacent generated_at sites in system.py and support.py.
The three diagnostic surfaces shipped earlier this month
(6bc6a1d6 VP setup diagnostic, e222a0ef log-health scanner,
ed31b8f4 connection diagnostic in the bug-report bubble) were
only ever shown to the *user*. A bug report arriving in the
maintainer's inbox carried raw logs but no diagnostic results —
the user-visible "your X1C can't reach MQTT" finding never made
it into the issue body, so the maintainer had to ask the user
to re-run and paste.
New `services/diagnostic_snapshot.collect_diagnostic_snapshot`
runs all three concurrently with a per-probe 15 s wall-clock cap
(so total ≈ max(per-cap), not sum — fleet size doesn't matter)
and is fail-soft per probe: a crash inside one printer's check
emits `{"printer_id": N, "error": "..."}` for that entry rather
than nuking the whole snapshot. The snapshot is then added as a
`diagnostics` top-level key by `_collect_support_info()`, so both
flows (POST /support/bundle and POST /bug-report/submit via
`support_info=...`) pick it up without their own changes.
Private-data sanitization
-------------------------
The diagnostic schemas embed raw IPv4 in five field shapes that
must not land in a submitted GitHub issue or a shared support ZIP:
- PrinterDiagnosticResult.ip_address (top-level)
- DiagnosticCheck.params.printer_ip (network-mode check)
- DiagnosticCheck.params.host_ip (network-mode check)
- VPDiagnosticResult per-check params.bind_ip (VP setup)
- IPs embedded in log-health sample lines
The first two carry the printer's own IP (already in the
existing `collect_sensitive_strings` table via the Printer rows);
host_ip and bind_ip are NOT in the DB so a sensitive_strings-only
pass missed them.
Fix: `_sanitize_recursive` walks the full snapshot tree, masks
DB-known values with the same `[PRINTER]/[IP]/[SERIAL]/[ACCESS_CODE]`
labels the log sanitizer applies (via the shared
`collect_sensitive_strings`), then an IPv4-regex pass catches any
IP the DB didn't cover — most importantly the Bambuddy host IP
returned by `_get_host_ip()` and the VP `bind_ip` the user picked
at setup. Recursive walk so arbitrary nested dicts/lists don't
slip through future schema additions.
Live-DB smoke test against the dev fleet: zero raw IPv4 instances
in the serialized snapshot output; all five field shapes plus the
embedded log samples render as `[IP]`.
Progress indicators
-------------------
The bubble's "submitting" view and the System page's Download
button now render a static four-line checklist showing what's
running (printer connectivity → VP setup → log scan →
submit / build ZIP). Static, not faked phase progress — we can't
actually track server-side phases without SSE and the honest
"here's what's happening" list communicates the longer wait
without lying about percentage complete.
9 new i18n keys, real translations in all 9 locales (no English
fallback). parity script clean at 4993 leaves per locale.
Tests: 6 new in test_diagnostic_snapshot.py
- empty-input shape stable (the three top-level keys always present)
- per-printer / per-VP result coverage (lists match input lengths)
- fail-soft on a single-probe crash (other entries + log-health
still complete)
- timed_out marker when a probe exceeds the per-probe cap
(test patches the cap to 0.05 s)
- end-to-end IP sanitization across all five field shapes plus
log-sample IPs, with a final JSON-serialize-and-regex sweep
asserting zero raw IPv4 escapes anywhere in the result
- concurrent execution proof (4 × 0.2 s probes complete in
< 0.5 s; sequential would be 0.8 s)
Adds a passive log-health check that complements the active Connection
Diagnostic. Scans Bambuddy's recent app log against a curated allowlist
catalog of known failure signatures (rejected access code, FTPS :990
timeout, FTPS TLS failure, flapping MQTT, unreachable camera, SQLite
"database is locked" contention), dedupes and classifies each finding
as layer8/environment/bug, and deep-links to the troubleshooting wiki.
Sample log lines are sanitized before they leave the process. Exposed
via GET /system/health and surfaced on two surfaces sharing one
SystemHealthPanel component: a System Health section on the System
page, and inline in the bug reporter when the form opens.
The Add-Printer and Edit-Printer dialogs gained a setup-time pre-flight:
saving runs the connection diagnostic and, on a failed check, warns with
a "save anyway" escape hatch instead of silently saving a printer that
will immediately show offline.
Log read/parse/sanitize primitives extracted from routes/support.py into
a shared services/log_reader.py (behaviour-preserving); affected support
tests repointed accordingly.
Tests: test_log_health.py (11), test_system_api.py (2 new),
SystemHealthPanel + BugReportBubble + AddPrinterPreflight +
EditPrinterPreflight (8 frontend). All strings translated across the 9
locales. Backend ruff clean, full unit suite green, frontend build +
eslint clean, i18n parity green.
The two # noqa: S501 comments on the local-sidecar reachability probes
were using ruff/flake8 suppression syntax; bandit only honors # nosec,
so the scan flagged both calls as high-severity. Switched to
# nosec B501 with strengthened reasoning (reachability/health probe
only, no secrets in the request). No behavioural change.
* feat(auth): proxy OIDC provider icons server-side (#1333)
Strict img-src CSP blocked external OIDC icon hosts on the login page.
Loosening CSP was rejected via the MakerWorld precedent, so icons are
proxied: admin sets icon_url, backend fetches and caches the bytes in a
deferred BLOB column, the SPA renders from a same-origin
/api/v1/auth/oidc/providers/{id}/icon endpoint.
Issue #1312 follow-up. Investigation traced the "Name cannot be empty"
report to a sidecar image pre-dating the /profiles/bundle endpoint
addition. Two changes so the next occurrence is self-diagnosable from
the support bundle without a manual curl.
Backend: new _fetch_slicer_health(url) helper does a 2s GET on /health,
walks every non-dataPath key under checks looking for a version field
(the wrapper labels both sidecars as checks.orcaslicer regardless of
which CLI is bundled). _collect_slicer_api_info now exposes
bambu_studio_version and orcaslicer_version. Strips trailing slash
before appending /health to avoid double-slash 404s.
Docs: bambuddy-wiki/docs/features/slicer-api.md gains a Quick Start
callout that branch-built sidecars don't auto-update, a corrected
/health troubleshooting entry (both "unknown" version and "orcaslicer"
field name on bambu-studio-api are cosmetic wrapper bugs, not stale-
image indicators), a new "Name cannot be empty" troubleshooting entry,
and an Updating section that requires --no-cache --pull together
(BuildKit caches the git context separately from layers, so --no-cache
alone silently reuses the old checkout).
The settings-table passthrough auto-captured everything in `settings` (with
sensitive-key redaction), but features storing config in dedicated tables
were invisible. Triaging recent OIDC / 2FA / group bugs and the X1C slicer
investigation needed data that wasn't in the bundle.
New blocks in _collect_support_info:
- auth: OIDC providers (cleartext names, no secrets), TOTP / OTP /
API-key / long-lived-token / group counts
- library: file / folder / external / trash / makerworld totals
- inventory: spool + k-profile counts
- queue: pending count, oldest pending age
- maintenance: items total + enabled
- integrations.github_backup: providers used + recent failures
- integrations.slicer_api: enabled, URL source, reachability ping
- per-printer obico_enabled flag
Plus three smaller fixes caught testing against a real bundle:
- mqtt_broker no longer leaks (broker keyword added)
- virtual_printer_tailscale_auth_key no longer leaks (auth_key keyword
+ tskey- value-prefix safety net for future Tailscale settings)
- slicer-API reachability check now mirrors the route's three-level URL
precedence (DB → env var → default), instead of only looking at the
DB setting. Previously returned null for every installation running
the sidecar via env var or default port — i.e. most of them.
Settings dump now retains every key from the Settings table and replaces
sensitive values with [REDACTED] instead of dropping the row. New config
flags automatically surface in future bundles without a code change.
Adds integrations.spoolbuddy with per-device firmware, NFC/scale hardware,
calibration, online state and uptime — anonymized (no hostnames, IPs or
device IDs). Both /support/bundle and the bug-report bubble benefit, since
they share _collect_support_info().
Settings → Support → Debug Logging elevated httpx/httpcore to DEBUG,
which makes httpx log every outbound request URL. For Discord and
generic webhook notifications the bearer token is embedded in the URL
path, so users who turned on debug logging to capture a support bundle
were writing their webhook tokens straight into bambuddy.log.
Pin httpx/httpcore to WARNING regardless of the debug toggle. paho.mqtt
still honours debug. Users who enabled debug logging while notifications
were sending must rotate any exposed Discord/webhook URLs — the token
is the path, so the whole URL has to be regenerated in the provider UI.
The debug support bundle included virtual_printer_remote_interface_ip
unmasked in support-info.json. The setting key didn't match any
sensitive-key filter substring. Added "_ip" to the filter set so IP
address settings are excluded. Log file content was already redacted
by the existing IPv4 regex.
Bambuddy can now use an external PostgreSQL database via the
DATABASE_URL environment variable. SQLite remains the default.
Dialect-aware helpers handle upserts, PRAGMAs, FTS (FTS5 vs
tsvector+GIN), backup/restore, and health checks. All migration
blocks use savepoints to prevent Postgres transaction poisoning.
Backups are always portable SQLite format regardless of backend.
Cross-database restore imports SQLite backups into PostgreSQL
with automatic boolean/datetime conversion, NOT NULL default
filling, and FK constraint handling.
RTSP stream URLs (rtsps://bblp:<code>@<ip>:322/...) were not covered
by the credential sanitizer, leaking access codes in support bundles
and bug report logs. Extended the URL regex to match rtsps:// and added
access codes to the sensitive string collection for exact-match
redaction in both export paths.
The debug logging banner displayed a negative elapsed time (e.g. "-60m -59s")
equal to the server's UTC offset. datetime.now() stored local time without a
timezone indicator, but the frontend's parseUTCDate() interpreted it as UTC.
Use datetime.now(tz=timezone.utc) consistently for storing, parsing, and
comparing the enabled_at timestamp.
Floating bug report button submits issues via bambuddy.cool relay (no GitHub
token needed locally). Collects 30s debug logs with printer push_all, sanitizes
all sensitive data, uploads logs as files to GitHub. Screenshot upload/paste/drag
with JPEG compression. Translated into all 7 languages. Includes 21 tests.
Four support package improvements:
1. Mask first two octets of subnet IPs in support info
(192.168.1.0/24 → x.x.1.0/24) to avoid leaking private network
addresses.
2. Fix Docker network_mode_hint detection. The old heuristic
(interface count > 2) always reported "bridge" on single-NIC
hosts because get_network_interfaces() excludes Docker
interfaces. Now checks for docker0/br-*/veth* visibility via
socket.if_nameindex() — these are only visible in host mode.
3. Parse MQTT "fun" field at top level of payload (not just inside
"print" key). Some firmware versions send it there, which
explains why developer_mode was null for most users.
4. Add virtual_printers section to support info with mode, model,
enabled/running status, and pending file count.
Parse the MQTT "fun" field bit 0x20000000 to detect whether connected
printers have Developer LAN Mode enabled. Show a persistent orange
warning banner when any printer lacks it, since newer firmware silently
rejects MQTT write commands without developer mode.
- Parse fun field into developer_mode on PrinterState
- Add /printers/developer-mode-warnings lightweight polling endpoint
- Include developer_mode in printer status API and support bundle
- Orange banner with affected printer names and wiki link
- Translations for all 6 locales (en, de, fr, it, ja, pt-BR)
- 7 backend + 4 frontend tests
The log sanitizer only used regex patterns, missing arbitrary user-chosen
strings (printer names, usernames). Tasmota smart plug credentials were
logged verbatim in URLs by httpx.
- Make _sanitize_log_content() database-aware: query Printer names/serials,
User usernames, and Bambu Cloud email for exact-string replacement
(longest-first, skip <3 chars to prevent over-redaction)
- Fix serial regex leaking first 3 chars (remove capture group partial
redaction), add case-insensitive flag
- Move Tasmota credentials from URL-embedded (http://user:pass@host) to
httpx auth= parameter so they never appear in logs
- Add URL credentials regex as defense-in-depth for user:pass@ in logs
- Add 'username' and 'path' to settings sensitive_keys filter (catches
smtp_username, slicer_binary_path in support-info.json)
Suppress SQL/aiosqlite debug noise (~90% log volume reduction), add
caller-traced PRINT COMMAND logging to start_print(), log scheduler
queue checks, tighten stale archive ilike match to exact match, and
warn on multiple queue items in "printing" status. Includes 18 new
unit tests.
Downgrade 58 diagnostic logger.info calls to logger.debug in
bambu_mqtt.py — payload dumps, detector state changes, field
discovery, H2D disambiguation, and periodic status updates no longer
flood logs at the default INFO level. User-initiated actions (print,
stop, calibration, AMS load/unload) remain at INFO. Also suppress
paho-mqtt library INFO messages in production mode.
Backend:
- Add dual external spool support for H2D (vt_tray as list: Ext-L/Ext-R)
- Add cloud filament ID map endpoint (/cloud/filament-id-map)
- Fix RFID spool data erased by periodic AMS updates (skip tag matcher
for RFID-tagged trays)
- Fix AMS slot config overwrites RFID spool state
- Fix K-profile selection corrupts existing profiles on X1C/P1S
- Resolve K-profiles filament name via cloud filament ID map
- Update print scheduler and usage tracker for dual external spools
Frontend:
- Add printer model filtering to ConfigureAmsSlotModal (cloud/local/builtin
presets filtered by @BBL model suffix and compatible_printers)
- Add pre-population for configured slots (preset, color, K-profile)
- Add K-Profiles view with accurate filament name resolution
- Internationalize all ConfigureAmsSlotModal strings (en/de/fr/it/ja — 21 keys)
- Add 5 new ConfigureAmsSlotModal tests (model filtering, pre-selection,
color pre-population, i18n)
- Update PrintersPage for dual external spool rendering
Docs:
- Update CHANGELOG, README, website features, and wiki AMS docs
raw_data["ams"] is stored as a list by the MQTT handler, but the
support info code only checked for a nested dict format. AMS unit
and tray counts were always 0.
Remove vestigial _debug_logging_enabled and _debug_logging_enabled_at
globals from support.py (written but never read; DB is queried directly).
Simplify hue classification in PrintersPage.tsx and colors.ts by removing
always-true h < 345 checks and dead 'Unknown' fallbacks. Narrow
getWifiStrength param type to remove always-false null guard.
These modules were already imported at the top of each file.
Removes re-imports of re, json, zipfile, and logging from
inside functions in archive.py, library.py, main.py,
printers.py, support.py, and test_library_api.py.
Resolves all 30 CodeQL py/repeated-import findings.
The support bundle states that printer serial numbers are NOT collected,
but they were appearing in debug logs. Added regex to sanitize Bambu Lab
serial numbers (00M/01D/01S/01P/03W prefix + alphanumeric) while keeping
the prefix for debugging context.
Example: [01D00A12345678] -> [01D[SERIAL]]
Closes#216
that allows viewing and filtering application logs in real-time.
Features:
- Start/Stop live streaming with 2-second auto-refresh
- Filter by log level (DEBUG, INFO, WARNING, ERROR)
- Text search across messages and logger names
- Clear logs with one click
- Expandable multi-line log entries (stack traces, etc.)
- Auto-scroll to follow new entries
Closes#87
- Add /api/v1/support/debug-logging endpoints to toggle debug log level
- Add /api/v1/support/bundle endpoint to generate ZIP with system info and logs
- Debug logging state persists across restarts via Settings database
- Add debug logging indicator banner in Layout with real-time duration timer
- Add Support & Troubleshooting section to System Information page
- Privacy protection:
- Filter sensitive settings (emails, keys, tokens, URLs, configs)
- Sanitize paths to remove usernames
- Remove hostname from collected data
- Replace IP addresses with [IP] and emails with [EMAIL] in logs
- Add privacy info panel explaining what data is/isn't collected
- Require debug logging to be enabled before downloading support bundle