fix(spoolbuddy): sync SSH key over heartbeat to survive Bambuddy keypair rotation

Bambuddy's SSH keypair under <DATA_DIR>/spoolbuddy/ssh/ regenerates whenever
  the data dir is recreated (volume remount, container recreate, fresh deploy).
  The daemon previously only fetched the pubkey at registration, so any
  rotation after a successful boot left ~/.ssh/authorized_keys pointing at
  a stale public half — every Update click then failed with "Connection
  closed by authenticating user spoolbuddy [preauth]" until the daemon was
  restarted by hand. Each prior registration also appended a fresh entry
  without pruning, accumulating stale Bambuddy-tagged keys indefinitely.

  - HeartbeatResponse now carries ssh_public_key; the heartbeat route reads
    it via the same try/except shape as the register route so a missing or
    unreadable backend key doesn't break telemetry.
  - _deploy_ssh_key() strips lines tagged bambuddy-spoolbuddy and writes
    the current key once. No-op when already in sync (no mtime churn on
    every heartbeat). User-managed entries are preserved.
  - Daemon heartbeat handler calls _deploy_ssh_key when the response
    carries a key, so rotations propagate within one heartbeat instead
    of requiring a service restart.

  Tests: 5 unit (creates-when-missing, replace-stale-pileup, preserve-user-keys,
  idempotent, swallows-write-errors) + 2 backend integration (heartbeat carries
  the key; backend key-read failure leaves ssh_public_key None but the
  heartbeat still 200s).
This commit is contained in:
maziggy
2026-05-01 10:44:01 +02:00
parent 44e4b6b5d7
commit 2aabbe5d37
6 changed files with 179 additions and 9 deletions
+2
View File
@@ -10,6 +10,8 @@ All notable changes to Bambuddy will be documented in this file.
- **Filament Track Switch (FTS) support — print modal filament dropdown is no longer empty when an X2D / H2D has the FTS accessory installed** ([#1162](https://github.com/maziggy/bambuddy/issues/1162), reported by @mkavalecz) — When the FTS accessory is installed the printer's MQTT changes one nibble of the per-AMS `info` bitmask: bits 8-11 flip from a fixed extruder ID (0x0 / 0x1) to `0xE` ("uninitialized"), because the AMS is no longer wired to a single nozzle — the FTS dynamically routes any slot to either extruder. Bambuddy's MQTT parser already skipped 0xE entries when building `ams_extruder_map` (matching BambuStudio's reading for boot-time transient state), so with the FTS installed the map ended up empty and the print modal's filament dropdown — which filters by `extruderId === nozzle_id` to prevent cross-nozzle assignment ("position of left hotend is abnormal" failures) — filtered out *every* loaded slot. Net effect: empty Filament Mapping dropdown on every dual-nozzle print with the FTS, even when the AMS was fully loaded with the right material. Detection comes from a new MQTT field — `print.device.fila_switch` — which is non-null only when the accessory is installed; it carries the routing topology as two arrays: `in[track] = currently fed slot (-1 = empty)` and `out[track] = extruder this track terminates at`. The fix surfaces this through a new `FilaSwitchState` dataclass on `PrinterState` (`installed`, `in_slots`, `out_extruders`, `stat`, `info`) and the equivalent `FilaSwitchResponse` Pydantic schema on the `GET /printers/{id}/status` route. Frontend (`useFilamentMapping.ts` + `FilamentMapping.tsx`) skips the per-extruder filter when `printerStatus.fila_switch?.installed === true` so any compatible AMS slot can satisfy any nozzle's filament requirement, since the FTS handles the routing. Slots currently fed into a track also get a routing badge in the dropdown — `[L]` or `[R]` — so the user can tell at a glance which slot the FTS is currently routing where (idle slots get no badge: they can be routed to either extruder on demand). The hard "no cross-nozzle assignment" filter on real dual-nozzle printers without the FTS stays untouched (still trips the same way it always has — `fila_switch == null` keeps the existing behaviour). 4 backend tests in `test_bambu_mqtt.py::TestFilamentTrackSwitchDetection` (default-not-installed, detect-from-MQTT-using-the-reporter's-bundle, no-fila_switch-field-stays-not-installed, missing-in-out-arrays-don't-crash) and 2 frontend tests in `useFilamentMapping.test.ts` (FTS-active drops the nozzle filter; explicit `fila_switch: null` keeps the filter applied). Upstream fila_switch payloads with anything other than the documented shape are tolerated — `installed` flips on the *presence* of the field, the routing arrays default to empty lists if missing, and the dropdown skips the badge for slots not currently in `in_slots`.
### Fixed
- **SpoolBuddy SSH update fails with "permission denied for user spoolbuddy" after Bambuddy keypair rotation** (reported during user testing) — Bambuddy's data dir at `<DATA_DIR>/spoolbuddy/ssh/` can get recreated outside the daemon's control (volume remount, container recreate, fresh deploy), at which point `get_or_create_keypair()` generates a new ed25519 keypair. The SpoolBuddy daemon previously only fetched and deployed Bambuddy's public key at registration time (`/devices/register`), so any rotation after a successful registration left the device's `~/.ssh/authorized_keys` pointing at a defunct public half — every "Update" click from the Bambuddy UI then failed with `Connection closed by authenticating user spoolbuddy [preauth]` until the daemon was restarted manually. Worse, every prior successful registration appended a fresh entry to `authorized_keys` without ever pruning the old one, so a typical device accumulated 5+ stale Bambuddy-tagged keys (each one a permanent backdoor for whichever Bambuddy keypair held the matching private half at the time it was deployed). Two-pronged fix: **(1)** the heartbeat response (`HeartbeatResponse`, `routes/spoolbuddy.py:282`) now carries the current `ssh_public_key` alongside the existing `pending_command` / calibration fields, so the daemon's heartbeat picks up a key rotation within one cycle instead of needing a service restart; the same `try/except Exception: pass` pattern as the registration response keeps a missing/unreadable backend key from breaking telemetry. **(2)** `_deploy_ssh_key()` in `daemon/main.py` now syncs rather than appends — it strips every line tagged `bambuddy-spoolbuddy`, writes the current key once, and is a no-op when already in sync (so it doesn't churn the file every heartbeat). User-managed entries (any line not tagged `bambuddy-spoolbuddy`) are preserved untouched. 5 new unit tests in `spoolbuddy/tests/test_deploy_ssh_key.py` (creates-when-missing → mode-600 file with the current key; pile-up-of-stale-keys → only current key remains, no growth; preserves-unrelated-user-keys → user's own SSH access untouched; idempotent-when-in-sync → no mtime change so heartbeat doesn't churn the file; swallows-write-errors → readonly-fs PermissionError doesn't crash the heartbeat loop). 2 new backend integration tests in `test_spoolbuddy.py::TestDeviceEndpoints` — `test_heartbeat_returns_ssh_public_key` (response carries the key on every heartbeat) and `test_heartbeat_ssh_key_failure_does_not_break_heartbeat` (backend key-read failure leaves `ssh_public_key: None` but the heartbeat still 200s).
- **External-camera frames returned as black on go2rtc and other MJPEG sources** ([#1177](https://github.com/maziggy/bambuddy/issues/1177), reported by @nkm8) — `_capture_mjpeg_frame` returned the very first JPEG it found in the stream's bytes (`backend/app/services/external_camera.py:282`), but many MJPEG sources — go2rtc most notably, and several IP cameras — emit a "warm-up" frame on the byte that follows connection accept: usually the last keyframe held in the encoder, which is often black or stale until the encoder catches up to live content. Subsequent frames on the same connection are fine. The reporter saw it across snapshot UX, finish photos in notifications, and timelapse — every code path that opens a fresh capture connection (snapshot endpoint, `[PHOTO-BG]` finish photo, plate-detection CV, Obico ML inference, layer timelapse, Settings → Test). His own observation that go2rtc's `/api/frame.jpeg` (single-frame, internally already warmed) is never black while the first frame off `/api/stream.mjpeg` is, matched the hypothesis exactly. Support-bundle evidence was clean: every black notification frame in his log was 11095 bytes (a pure-black 1280×720 JPEG encodes to ~10–15 KB on standard libjpeg quality settings), while every captured-after-warm-up frame from the same source was 30–45 KB. Fix: read past the first frame and return the second; if the connection closes / times out / hits the 5 MB buffer cap before a second frame ever arrives, fall back to the first so callers still get *something* (degrading slow / single-frame streams to None would regress every code path that relied on pre-fix behaviour). The inner-loop now drains every complete frame already in the buffer before pulling the next chunk so high-FPS sources that pack multiple frames per chunk are handled correctly. The `snapshot` / `rtsp` / `usb` capture paths and the live-view streaming endpoint (`generate_mjpeg_stream`) are untouched. 7 new regression tests in `test_external_camera.py::TestCaptureMjpegFrameWarmupSkip` cover (a) two-frames-in-two-chunks → second returned, (b) two-frames-in-one-chunk → second returned, (c) frame split across chunk boundary → assembled correctly, (d) single-frame stream → first returned via fallback (no None regression), (e) timeout after first frame → first returned via fallback, (f) zero-frame stream → None, (g) non-200 status → None. Latency penalty: at most one frame interval (typically 50 ms – 1 s on a steady stream).
- **MakerWorld sidebar entry visible to every user regardless of group permissions** ([#1175](https://github.com/maziggy/bambuddy/issues/1175)) — Backend already enforced `makerworld:view` on every `/makerworld/*` route (`backend/app/api/routes/makerworld.py:145, 157, 242, 406`), the permission was correctly granted to the admin and standard-user role defaults (`permissions.py:298, 364, 454`), and the frontend `Permission` type union already included `'makerworld:view' | 'makerworld:import'` (`client.ts:2498`) — but the sidebar's hand-maintained `navPermissions` map in `Layout.tsx:278` had no entry for `makerworld`, so `isHidden('makerworld')` always returned false and the entry rendered for every authenticated user. Users without the permission saw the entry, clicked, and the page rendered while every API call inside it 403'd. Two-line fix: (1) `Layout.tsx:278` — add `makerworld: 'makerworld:view'` to the map, matching every other sidebar entry's gating shape; (2) `App.tsx:200` — wrap the route in `<PermissionRoute permission="makerworld:view">` for defence in depth, so a user who knows the URL can no longer reach the page directly (matches the existing pattern on `settings`, `groups/new`, `groups/:id/edit` two lines below). 2 new Layout tests pin the contract: with auth enabled and a user lacking `makerworld:view`, the sidebar `<a href="/makerworld">` link is absent (other links like `/files` still render); with the permission granted, the link renders.
+12
View File
@@ -279,6 +279,17 @@ async def device_heartbeat(
if was_offline:
logger.info("SpoolBuddy device back online: %s", device.device_id)
# Include current SSH public key so the daemon can re-deploy it whenever
# Bambuddy's keypair rotates (data dir wiped, container recreated, etc.) —
# otherwise SSH updates fail until the daemon restarts.
ssh_public_key: str | None = None
try:
from backend.app.services.spoolbuddy_ssh import get_public_key
ssh_public_key = await get_public_key()
except Exception:
pass
return HeartbeatResponse(
pending_command=pending,
pending_write_payload=pending_write,
@@ -287,6 +298,7 @@ async def device_heartbeat(
calibration_factor=device.calibration_factor,
display_brightness=device.display_brightness,
display_blank_timeout=device.display_blank_timeout,
ssh_public_key=ssh_public_key,
)
+1
View File
@@ -74,6 +74,7 @@ class HeartbeatResponse(BaseModel):
calibration_factor: float
display_brightness: int = 100
display_blank_timeout: int = 0
ssh_public_key: str | None = None
# --- NFC schemas ---
@@ -203,6 +203,54 @@ class TestDeviceEndpoints:
assert msg["type"] == "spoolbuddy_online"
assert msg["device_id"] == "sb-hb"
@pytest.mark.asyncio
@pytest.mark.integration
async def test_heartbeat_returns_ssh_public_key(self, async_client: AsyncClient, device_factory):
"""Heartbeat response carries the current SSH public key so the daemon
can re-deploy it whenever Bambuddy's keypair rotates without waiting
for a service restart."""
await device_factory(device_id="sb-ssh-hb")
fake_key = "ssh-ed25519 AAAATESTKEY bambuddy-spoolbuddy"
with (
patch("backend.app.api.routes.spoolbuddy.ws_manager") as mock_ws,
patch(
"backend.app.services.spoolbuddy_ssh.get_public_key",
AsyncMock(return_value=fake_key),
),
):
mock_ws.broadcast = AsyncMock()
resp = await async_client.post(
f"{API}/devices/sb-ssh-hb/heartbeat",
json={"nfc_ok": True, "scale_ok": True, "uptime_s": 5},
)
assert resp.status_code == 200
assert resp.json()["ssh_public_key"] == fake_key
@pytest.mark.asyncio
@pytest.mark.integration
async def test_heartbeat_ssh_key_failure_does_not_break_heartbeat(self, async_client: AsyncClient, device_factory):
"""If the backend can't read its own SSH key, the heartbeat must still
succeed — telemetry/commands are far more critical than key sync."""
await device_factory(device_id="sb-ssh-fail")
with (
patch("backend.app.api.routes.spoolbuddy.ws_manager") as mock_ws,
patch(
"backend.app.services.spoolbuddy_ssh.get_public_key",
AsyncMock(side_effect=OSError("disk full")),
),
):
mock_ws.broadcast = AsyncMock()
resp = await async_client.post(
f"{API}/devices/sb-ssh-fail/heartbeat",
json={"nfc_ok": True, "scale_ok": True, "uptime_s": 5},
)
assert resp.status_code == 200
assert resp.json()["ssh_public_key"] is None
@pytest.mark.asyncio
@pytest.mark.integration
async def test_heartbeat_returns_pending_command(self, async_client: AsyncClient, device_factory):
+30 -9
View File
@@ -66,8 +66,18 @@ def _get_ip() -> str:
return "unknown"
SSH_KEY_TAG = "bambuddy-spoolbuddy"
def _deploy_ssh_key(public_key: str) -> None:
"""Write Bambuddy's SSH public key to authorized_keys if not already present."""
"""Sync Bambuddy's SSH public key into authorized_keys.
Replaces any prior key tagged ``bambuddy-spoolbuddy`` so the file always
reflects Bambuddy's *current* keypair. Without this, every Bambuddy key
rotation (data dir wipe, container recreate, etc.) leaves a stale entry
behind and the file grows unbounded.
"""
target = public_key.strip()
home = Path.home()
ssh_dir = home / ".ssh"
auth_keys = ssh_dir / "authorized_keys"
@@ -75,17 +85,24 @@ def _deploy_ssh_key(public_key: str) -> None:
try:
ssh_dir.mkdir(mode=0o700, exist_ok=True)
# Check if key already deployed
existing_lines: list[str] = []
if auth_keys.exists():
existing = auth_keys.read_text()
if public_key.strip() in existing:
return
existing_lines = auth_keys.read_text().splitlines()
# Append key
with auth_keys.open("a") as f:
f.write(public_key.strip() + "\n")
kept = [line for line in existing_lines if SSH_KEY_TAG not in line]
new_lines = kept + [target]
# Already in sync — current key present and no stale Bambuddy entries.
if existing_lines == new_lines:
return
auth_keys.write_text("\n".join(new_lines) + "\n")
auth_keys.chmod(0o600)
logger.info("SSH public key deployed to %s", auth_keys)
removed = len(existing_lines) - len(kept)
if removed:
logger.info("SSH public key updated in %s (replaced %d stale entries)", auth_keys, removed)
else:
logger.info("SSH public key deployed to %s", auth_keys)
except Exception as e:
logger.warning("Failed to deploy SSH key: %s", e)
@@ -233,6 +250,10 @@ async def heartbeat_loop(config: Config, api: APIClient, start_time: float, shar
)
if result:
ssh_key = result.get("ssh_public_key")
if ssh_key:
_deploy_ssh_key(ssh_key)
cmd = result.get("pending_command")
if cmd == "tare":
scale = shared.get("scale")
+86
View File
@@ -0,0 +1,86 @@
"""Tests for daemon.main._deploy_ssh_key — Bambuddy key sync.
Background: Bambuddy generates an ed25519 keypair under its data dir and ships
the public half to the SpoolBuddy daemon over the registration/heartbeat
response. The daemon writes that key into ~/.ssh/authorized_keys so Bambuddy
can SSH in to drive remote updates. Whenever Bambuddy's keypair rotates (data
volume wiped, container recreated, fresh deploy) the device's authorized_keys
must drop the old entries and pick up the new one — otherwise:
1. SSH updates start failing silently with permission-denied
2. Stale Bambuddy-tagged keys pile up over time, eroding the security
boundary (any prior keypair Bambuddy held is permanently authorized).
These tests pin the replace-not-append semantics of the deploy helper.
"""
from unittest.mock import patch
from daemon.main import _deploy_ssh_key
CURRENT_KEY = "ssh-ed25519 AAAACURRENT bambuddy-spoolbuddy"
STALE_KEY_1 = "ssh-ed25519 AAAASTALE1 bambuddy-spoolbuddy"
STALE_KEY_2 = "ssh-ed25519 AAAASTALE2 bambuddy-spoolbuddy"
USER_KEY = "ssh-ed25519 AAAAUSER alice@laptop"
class TestDeploySshKey:
def test_creates_authorized_keys_when_missing(self, tmp_path):
with patch("daemon.main.Path.home", return_value=tmp_path):
_deploy_ssh_key(CURRENT_KEY)
auth_keys = tmp_path / ".ssh" / "authorized_keys"
assert auth_keys.exists()
assert auth_keys.read_text().strip() == CURRENT_KEY
assert auth_keys.stat().st_mode & 0o777 == 0o600
def test_replaces_all_prior_bambuddy_tagged_keys(self, tmp_path):
"""The pile-up scenario: 6+ stale keys accumulated over rotations.
After deploy, only the current key remains — no growth."""
ssh_dir = tmp_path / ".ssh"
ssh_dir.mkdir()
auth_keys = ssh_dir / "authorized_keys"
auth_keys.write_text(f"{STALE_KEY_1}\n{STALE_KEY_2}\n")
with patch("daemon.main.Path.home", return_value=tmp_path):
_deploy_ssh_key(CURRENT_KEY)
lines = auth_keys.read_text().strip().splitlines()
assert lines == [CURRENT_KEY]
def test_preserves_unrelated_user_keys(self, tmp_path):
"""Only Bambuddy-tagged keys get replaced — user's own keys stay."""
ssh_dir = tmp_path / ".ssh"
ssh_dir.mkdir()
auth_keys = ssh_dir / "authorized_keys"
auth_keys.write_text(f"{USER_KEY}\n{STALE_KEY_1}\n")
with patch("daemon.main.Path.home", return_value=tmp_path):
_deploy_ssh_key(CURRENT_KEY)
lines = auth_keys.read_text().strip().splitlines()
assert USER_KEY in lines
assert STALE_KEY_1 not in lines
assert CURRENT_KEY in lines
def test_idempotent_when_already_in_sync(self, tmp_path):
"""No-op when authorized_keys already matches the desired state —
avoids needless writes on every heartbeat."""
ssh_dir = tmp_path / ".ssh"
ssh_dir.mkdir()
auth_keys = ssh_dir / "authorized_keys"
auth_keys.write_text(f"{USER_KEY}\n{CURRENT_KEY}\n")
original_mtime = auth_keys.stat().st_mtime_ns
with patch("daemon.main.Path.home", return_value=tmp_path):
_deploy_ssh_key(CURRENT_KEY)
assert auth_keys.stat().st_mtime_ns == original_mtime
def test_swallows_write_errors(self, tmp_path):
"""A failed deploy must not crash the heartbeat loop."""
with (
patch("daemon.main.Path.home", return_value=tmp_path),
patch("daemon.main.Path.mkdir", side_effect=PermissionError("readonly fs")),
):
_deploy_ssh_key(CURRENT_KEY) # should not raise