6 Commits
Author SHA1 Message Date
maziggy 5bbfeefa65 fix(backup): diagnose an unwritable backup path instead of quoting errno 30 (#2544)
Nightly backups to a mounted NAS share ran from May and then stopped, failing
with [Errno 30] Read-only file system. The reporter checked folder permissions
-- correctly: the mount is gid=backup,dir_mode=0775, the service user is in that
group, and his own shell writes to the share fine.

Errno 30 is EROFS. A permission problem is errno 13. EROFS means the filesystem
refused the write, and it refused because we told it to: our systemd unit ships
ProtectSystem=strict, which mounts everything read-only inside the service's
mount namespace and carves back out only ReadWritePaths=<install> <data> <logs>.
A NAS share is not one of those three. Reads are unaffected -- which is why the
UI happily listed his existing backups from the share while being unable to
write a new one -- and his shell is outside the namespace entirely, so every
check he could think to run said the directory was fine.

Both installers write the unit file wholesale, so a ReadWritePaths line added by
hand disappeared on the next install, taking the backups with it. They now back
the old unit up (.bak-<timestamp>) and carry the operator's extra writable paths
forward, reporting which ones they kept. The unit template documents the
carve-out.

The output directory is probed with a real write when it is saved and when the
backup card loads, so an unwritable path is caught there rather than at 03:00
for a week. On failure the card names the cause and hands over the fix with the
operator's path already in it (systemctl edit bambuddy -> ReadWritePaths=...),
and a failed run reports the same diagnosis rather than the raw OSError. EROFS
outside systemd, permission-denied, out-of-space, not-a-directory and missing are
told apart, in all 11 locales.

Docker: a backup path that is not bind-mounted is writable -- the write lands in
the container's ephemeral layer and is lost on the next compose up. The probe
compares the directory's device against the container root and warns, with the
compose snippet that mounts it properly.
2026-07-12 08:44:53 +02:00
maziggy aba00598bb fix(smart-plugs): read a REST plug's lifetime counter, and derive Today/Yesterday from it (issue #2539)
A Shelly reports one energy figure — aenergy.total, a lifetime counter in Wh
that never resets. Bambuddy had a single REST energy field and filed whatever
it found under "today", so the value never reset at midnight, and Yesterday
and Total stayed at zero: get_energy() simply never set those keys.

With `total` unpopulated, the hourly snapshot recorder skipped the plug, so
the Statistics page's energy figure was zero as well, not just the Settings
card.

Split the REST energy config in two: rest_energy_path still means "used
today", rest_energy_total_path means "lifetime counter". A Shelly has only
the latter; a Tasmota behind a REST bridge has both; sharing a URL costs one
fetch, not two.

Then derive Today and Yesterday from that counter using the snapshots we were
already taking: today = counter now - counter at the last local midnight;
yesterday = the gap between the two previous midnights. Local midnight, not
UTC — a UTC boundary rolls Today over at 02:00 in Berlin. The snapshot loop
now ticks on the local hour so a reading lands on the boundary instead of up
to an hour early. A counter that goes backwards (factory reset) reports
nothing rather than a negative.

Collateral, found while verifying on both engines: the smart-plug DateTime
columns are naive UTC but the code wrote aware datetimes into them. SQLite
drops the offset; asyncpg raises DataError. So on Postgres every snapshot
capture raised inside the loop's except, and every status poll raised on
last_checked — the whole subsystem was dead on the database we recommend for
multi-printer installs. All plug timestamps are naive UTC now.

Existing REST users with a cumulative path in the today field must move it to
the new lifetime field; the form and wiki now name which counter each wants.
2026-07-11 14:22:31 +02:00
maziggy f15e54c383 fix(windows): _local_zone falls back to stdlib utc when zoneinfo DB is missing
The Windows installer's embedded Python doesn't carry an IANA tz
  database, and the stdlib zoneinfo has no system DB to read on Windows.
  ZoneInfo("UTC") raises ZoneInfoNotFoundError on those installs, and
  the new /api/local-backup/status endpoint 500s on the resulting
  uncaught exception. Surfaced via a Windows traceback from a user's log:

    File "...\backend\app\services\local_backup.py", line 32, in _local_zone
      return ZoneInfo("UTC")
    zoneinfo._common.ZoneInfoNotFoundError: 'No time zone found with key UTC'

  _local_zone()'s try/except only covered the TZ-env branch — both
  fallbacks unconditionally called ZoneInfo("UTC") and re-raised.

  Fix (two parts):

  1. services/local_backup.py — return type widened from ZoneInfo to
     tzinfo, the UTC fallback is wrapped in its own try, and the
     last-resort fallback returns datetime.timezone.utc (stdlib, no
     IANA DB needed). str(timezone.utc) == "UTC" so the response shape
     on /api/local-backup/status is unchanged. The astimezone call in
     _calculate_next_run accepts any tzinfo — no other call sites
     affected.

  2. requirements.txt — pin tzdata>=2024.1; sys_platform == "win32" so
     the next Windows installer build ships the IANA DB, and any non-
     UTC TZ value (e.g. Europe/Berlin) resolves correctly. The stdlib
     fallback can only ever give UTC. Linux/macOS unaffected by the
     platform marker — they already have the system tz database.
2026-06-13 07:49:31 +02:00
maziggy a1cb5d5b4d fix(backup): interpret scheduled-backup HH:MM as local time, not UTC (#1602 follow-up)
The Scheduled Local Backups time-of-day picker was interpreted as UTC by
  _calculate_next_run, so a UTC+3 user had to enter 18:00 to get a 21:00
  local backup. The UI labeled the field "UTC" but it was still surprising.

  Picker is now interpreted in the container's local timezone, resolved
  from the TZ env var via zoneinfo.ZoneInfo (same source the Support page's
  environment.timezone shows). UTC fallback when TZ is unset or
  unrecognised. The /local-backup/status endpoint exposes the resolved
  zone, and the UI renders it next to the field via a new
  backup.localTimeHint i18n key with real translations in all 10
  non-English locales.

  One-time behaviour change for users who entered a UTC time as a
  workaround: the first scheduled cycle after upgrade will run at their
  local TZ offset earlier than expected. Re-enter the time as local once
  and it is correct from then on. No migration is shipped; migrating
  around a DST boundary would be ambiguous.
2026-06-04 10:17:10 +02:00
maziggy 396e9aa09e security: harden path-traversal class across routes + services; fifth CI backstop
Two attacker-controlled strings were being joined to library_dir with no
  resolve + containment check in the project ZIP import endpoint:

    - linked_folders[*].name from the request's project.json
    - per-entry zf.namelist() paths from the ZIP itself

  An absolute path in either field collapsed the join (Path("/lib") / "/etc"
  becomes Path("/etc") because pathlib discards the left side when the right
  is absolute) and the next write_bytes landed wherever the attacker chose.

  Adjacent finding from the routes audit: GET /archives/{id}/photos/{filename}
  had NO validation on filename and FileResponse-served arbitrary paths -
  the DELETE counterpart at least gated on the photos membership check.

  Adjacent finding from the services audit: ArchiveService.attach_timelapse
  wrote archive_dir / filename where filename ultimately came from a printer's
  FTP listing (compromised-printer threat model) or the /timelapse/select
  query param. A malicious printer that exposes a directory entry with ..
  segments could write the timelapse outside the archive directory.

  New backend/app/utils/safe_path.py::safe_join_under(parent, *parts) is the
  single source of truth: rejects empty / null-byte / absolute parts up-front,
  joins under parent, resolves both sides, asserts is_relative_to. Returns the
  resolved canonical path on success, raises HTTPException(400) on escape, or
  PathTraversalError when http=False (for service-layer callers that need to
  match a non-HTTP return contract).

  Wired into the import vectors, both archive photo handlers, and the
  attach_timelapse service. The full audit sweep inspected every Path/Name
  join in backend/app/api/routes/ AND backend/app/services/ - 25 route-layer
  sites + 8 service-layer sites confirmed safe and tagged with
  # SEC-PATH-OK: <reason> so future audits trust the inline guard at a glance.

  Fifth CI backstop test_route_path_arithmetic_is_safe_joined_or_marked
  AST-walks both layers and fails the build on any <dir-like>/<bare variable>
  join that doesn't either route through safe_join_under or carry the marker.
  The services layer is in scope because it receives values verbatim from the
  routes AND from external sources Bambuddy has no control over (the printer
  FTP-listing case above).

  SECURITY.md gets a fifth rule + a fifth row in the CI test mapping table;
  the rule now names the printer FTP-listing case explicitly so future
  services-layer audits set the right expectation.

--------------

  fix(library): suppress warning storm when bulk-uploading ZIPs of empty/stub STL files

  Uploading a ZIP of stub or empty STL files (e.g. the 24-byte
  "solid test\nendsolid test" shape) produced one WARNING per file in
  stl_thumbnail.py::generate_stl_thumbnail. The warnings were technically
  correct - trimesh returns a valid Mesh with zero vertices, the safeguard
  matches, and the function returns None so the library entry is still
  created without a thumbnail - but the volume turned a successful upload
  into thousands of WARNING lines in the journal.

  Two changes:

  1. The per-file "Failed to load STL or empty mesh" message in
     stl_thumbnail.py is now logger.debug instead of logger.warning. It's
     a per-file content observation, not an actionable error; the caller
     already handles None correctly. The branch now catches the rare
     "large enough but trimesh still can't parse it" case, visible in
     debug logs without spamming production.

  2. New module constant MIN_USABLE_STL_BYTES = 200 (smallest binary STL
     with one triangle is 134B, smallest ASCII ~150B; 200 is a safe floor
     below any real STL). The three thumbnail call sites in library.py
     (extract_zip_file, single-file upload, _backfill_external_stl_thumbnails)
     pre-skip files below this size before calling generate_stl_thumbnail.
     Stubs never enter the trimesh pipeline at all.

  Behavior is unchanged for real STLs: any file >=200 bytes runs through
  the existing pipeline, MAX_VERTICES still triggers simplification at
  100k vertices for the 256x256 thumbnail render, large files still get
  thumbnails.

------------

  fix(stl-thumbnail): silence matplotlib first-import noise (writable cache + font_manager log level)

  On first STL upload, three matplotlib-internal log lines surfaced:

    WARNING [matplotlib] /opt/claude/.config/matplotlib is not a writable directory
    INFO    [matplotlib.font_manager] Failed to extract font properties from NotoColorEmoji.ttf
    INFO    [matplotlib.font_manager] generated new fontManager

  The writable-dir warning fired because Bambuddy's $HOME isn't writable for
  matplotlib's default config path; matplotlib fell back to /tmp/matplotlib-XXX
  which lost the font cache on every host reboot, so font_manager rebuilt it
  each cold start - producing another batch of INFO lines.

  Fix is two small additions in stl_thumbnail.py before the matplotlib import:

  1. New _configure_matplotlib_cache() sets MPLCONFIGDIR to
     settings.base_dir/.cache/matplotlib (mkdir if missing) so the cache
     persists across container restarts and the writable-dir warning never
     fires. Respects an externally-set MPLCONFIGDIR so operators who chose
     their own path aren't overridden. Best-effort with a debug fallback if
     settings can't be imported or the mkdir fails.

  2. logging.getLogger("matplotlib.font_manager").setLevel(WARNING) at module
     import demotes the per-font INFO scan that fires when font_manager
     builds its cache cold. Real font warnings (>= WARNING) still surface.

  3 new tests: font_manager logger at WARNING after module import;
  _configure_matplotlib_cache creates the directory under base_dir and sets
  MPLCONFIGDIR; an externally-set MPLCONFIGDIR is preserved verbatim.
  5516 backend tests green, frontend gates clean.
2026-06-02 12:12:54 +02:00
maziggy 774a639e9a . 2026-04-12 14:16:17 +02:00