fix(spoolbuddy): tolerate SPI_NO_CS rejection on Pi 5 (#1424)

Reporter on a Raspberry Pi 5 couldn't read NFC tags — gauge worked,
  SPI bus and wiring fine, but PN5180 transfers didn't complete.
  Manually commenting out `self._spi.no_cs = True` restored
  communication. Root cause: Pi 5's RP1 southbridge SPI driver
  (spi-rp1) doesn't honour the SPI_NO_CS ioctl the way the historical
  Broadcom driver on Pi 4 did.

  Safe to relax Pi-wide. SpoolBuddy's PN5180 NSS line is wired to
  GPIO23 (manual CS in _cs_low / _cs_high — the kernel's default
  auto-CS timing doesn't meet the PN5180's 5µs setup / 100µs hold
  spec). The hardware CE0 line (GPIO8) is not connected to the
  reader, so whether the kernel auto-toggles it is electrically
  invisible. The no_cs = True call was always cosmetic on this
  hardware.

  Wraps the assignment in try/except OSError in both the daemon and
  the diagnostic script; the daemon logs at debug level so future
  Pi-5-specific triage is greppable. README updated to drop the
  "spidev.no_cs = True resolves this" sentence and explain the
  manual GPIO23 CS scheme carries the timing on its own.
This commit is contained in:
maziggy
2026-05-19 12:33:33 +02:00
parent 18975b0dc2
commit c1123365da
4 changed files with 23 additions and 4 deletions
+2
View File
@@ -10,6 +10,8 @@ All notable changes to Bambuddy will be documented in this file.
- **Camera: in-app diagnostic for "Connection lost" (#1395 follow-up)** — Second step of the camera architecture overhaul. When the camera viewer hits its error state, a new **Diagnose** button next to **Retry** runs a staged check against the printer and renders the result inline: which stage failed, how long it took, and a translated remediation hint. Cuts off the "user opens a 'camera broken' ticket → wait days → ask for the support bundle → finally figure out it was their reverse proxy / LAN-only toggle / wrong access code" loop at the user's screen. **Backend** ships `backend/app/services/camera_diagnose.py` (orchestrator) and a new `POST /printers/{id}/camera/diagnose` route in `camera.py`. Stages: (1) `tcp_reachable` — opens a TCP socket to the camera port (322 RTSPS / 6000 chamber image) with a 3-second timeout; distinguishes timeout (`tcp_timeout` → "printer not reachable, check IP/network/power") from refused (`tcp_refused` → "camera port closed, check LAN-only and developer mode") from host-unreachable (`tcp_unreachable` → "printer not reachable"). (2) `first_frame` — captures one JPEG end-to-end via the existing `capture_camera_frame_bytes` pipeline (15-second timeout, same code that powers `/camera/snapshot`); auth, RTSP handshake, and first keyframe collapse into one stage because the user-facing answer is the same regardless of which sub-layer failed. **Live-stream shortcut**: when a viewer is currently watching the printer's camera AND the buffered last-frame timestamp is fresher than 10 seconds, the diagnostic skips the real test and returns `live_stream_active_healthy` — opening a fresh socket would kick the live viewer off on single-camera-connection firmwares (the #1348 reconnect-storm trigger), so we trust the real-world evidence instead. Response includes structured metadata for support triage: `protocol` (rtsp / chamber_image), `port`, `profile` (`default` or the model name with an override — currently only `P2S`), per-stage duration in ms, and the machine-readable summary code. **Frontend** adds `CameraDiagnoseModal.tsx` that fires the API call on mount, renders one row per stage with green-check / red-X / grey-skipped icons, and shows the summary remediation message in a bordered banner styled by overall status. The metadata line at the bottom (protocol / port / profile) lets support triage ask "what does your modal say?" instead of "send the support bundle". A **Run again** button re-runs the diagnostic without dismissing the modal. **EmbeddedCameraViewer** error state grows the Diagnose button (kept "Retry" as the primary action; Diagnose is the escape hatch for users who can't see what's wrong). A small stethoscope icon also lives in the viewer's always-visible control bar between **Refresh** and **Fullscreen**, so pre-flight testing ("did my firmware update break the camera?", "is the camera up before I send a print?") doesn't require waiting for the stream to fail first. Also lifted the previously-hard-coded "Camera unavailable" / "Retry" strings into `camera.unavailable` / `camera.retry` so the error UI is properly translated alongside the new keys. **i18n**: 16 new keys (`unavailable`, `retry`, plus `diagnose.{button,modalTitle,running,runFailed,retry,stage.*,summary.*,meta.*}`) translated across all 8 locales (en/de/fr/it/ja/pt-BR/zh-CN/zh-TW). German "Diagnose" is a real cognate — added to `IDENTICAL_TO_EN_ALLOWED.de` rather than translated to a synthetic. Parity check holds at 4849 leaves per locale. **Tests**: 11 backend unit tests in `test_camera_diagnose.py` cover the live-stream shortcut (skip when fresh, run when stale), the three TCP failure modes (timeout / refused / OSError) → distinct summary codes, the first-frame stage (no-frame and capture-exception cases), the all-OK path, and the result metadata (P2S → P2S profile / rtsp / 322; A1 → default / chamber_image / 6000; X1C → default / rtsp / 322). 1 backend integration test pins the route's response shape end-to-end. 3 frontend tests in `CameraDiagnoseModal.test.tsx` (mounted → API call, failure → translated remediation, Run again → re-call). 5021 backend tests + 1905 frontend tests green; ruff clean; build clean; i18n parity clean.
### Fixed
- **SpoolBuddy: NFC reader works again on Raspberry Pi 5 (#1424, reported by @flom89)** — Reporter on a Pi 5 installed SpoolBuddy successfully but couldn't talk to the PN5180 NFC module (the gauge worked, so SPI hardware and wiring were fine). Manually commenting out `self._spi.no_cs = True` in the daemon restored communication; reporter wasn't sure whether removing it would regress Pi 4 installs. Root cause: Pi 5 uses the new RP1 southbridge and its kernel SPI driver (`spi-rp1`) doesn't accept the `SPI_NO_CS` ioctl the same way the historical Broadcom driver on Pi 4 did — setting `no_cs = True` on Pi 5 either errors out or silently leaves the bus in a state where transfers don't complete. **Safe to drop Pi-wide, not just Pi 5** — SpoolBuddy's PN5180 NSS line is wired to GPIO23 (manual chip-select handled by `_cs_low()` / `_cs_high()` around every transfer, because the kernel's default 5µs setup / 100µs hold timing doesn't meet the PN5180's spec). The hardware CE0 line (GPIO8) is not connected to the reader, so whether the kernel auto-toggles it during `xfer2()` is electrically invisible to the PN5180. The `no_cs = True` call was a "be polite to the bus" gesture that was always cosmetic on this hardware. Fix wraps the assignment in `try / except OSError` in both `spoolbuddy/daemon/pn5180.py` (logs at debug level) and `spoolbuddy/scripts/read_tag.py` (silent — it's a diagnostic script with no logger). Try/except over a hard delete because Pi 4 installs that work today shouldn't see any behaviour change. README updated at `spoolbuddy/README.md:23-28` to drop the "spidev.no_cs = True resolves this" sentence in favour of explaining that manual CS via GPIO23 carries the timing on its own and that Pi 4 + Pi 5 are both supported. Hardware-only path so no automated test — verified by the reporter's bench test that commenting the line out restores reads. Ruff clean.
- **Cover thumbnails: stop hammering FTP and GitHub when a print's 3MF isn't on the printer (#1420, reported by reporter)** — Reporter on a P2S running 0.2.4.1 saw two log-flooding bugs trigger together once they started a print whose 3MF wasn't on the printer's FTP storage (typical SD-card-only print). **(1) Cover endpoint had no negative cache.** `GET /printers/{id}/cover` cached successful 3MF thumbnail downloads in `_cover_cache` keyed by `(subtask_name, view_key)`, but never recorded failures. When all 8 candidate FTP paths returned `550 Failed to open file`, the endpoint raised 404 without remembering that it just tried — and since `cover_url` stays populated on every `PrinterStatus` response while state is RUNNING/PAUSE, every React-Query refetch and every component remount drove the frontend to re-fetch, replaying the same 8-path FTP fan-out roughly every few seconds. On the user's hardware the printer's single FTP socket was so busy with these doomed retries that it surfaced as camera-stream symptoms ("ffmpeg didn't terminate gracefully"). Fix adds a parallel `_cover_404_cache: dict[int, set[tuple[str, str]]]` that records the same `(subtask_name, view_key)` key on every 404 path — both the all-FTP-paths-failed branch and the 3MF-has-no-thumbnail-inside branch. On the next call for the same key, the endpoint short-circuits to 404 before even consulting FTP. The negative cache is cleared in `clear_cover_cache()` alongside the positive cache, which `main.py::on_print_start` already calls — so when the next print starts (different subtask, or same subtask after a re-upload of a new file) Bambuddy retries fresh. **(2) GitHub update-check had no backoff on 403 rate-limit.** Once `api.github.com` returned `403 rate limit exceeded` (typical when multiple Bambuddy instances or other tools share a NAT'd source IP and exhaust the unauthenticated 60-req/hr quota), the next call hit GitHub again immediately. Fix adds module-level `_github_rate_limit_until` epoch-seconds plus three helpers in `updates.py`: `_seconds_until_github_unblocked()`, `_record_github_rate_limit(response)` (reads `X-RateLimit-Reset` from the 403, falls back to a 1-hour pause when the header is absent or unparseable, and only extends the window — never shortens it via an out-of-order response), and `_is_github_rate_limit_response(response)` (status 403/429 with `X-RateLimit-Remaining: 0`, body-text fallback when proxies strip the header). Both call sites — `GET /updates/check` and the in-app updater's `_discover_target_release` — short-circuit when the window is active; the route surfaces a structured `{error: "GitHub rate limit reached...", retry_after_seconds: <int>}` response so the SettingsPage UI can show a real wait time instead of an opaque "failed to check for updates". The "ffmpeg didn't terminate gracefully" warning line the reporter quoted is the standard SIGTERM → 2s wait → SIGKILL pattern in `camera.py::_terminate_ffmpeg` — RTSP/TLS streams routinely take >2s to drain and that warning fires for many users with no FTP issues; once the cover loop is silenced the resource pressure is gone, and the warning itself is cosmetic. **Tests**: `test_cover_negative_cache_skips_repeat_ftp_fanout` in `test_printers_api.py` mocks `download_file_try_paths_async` to return False, calls the endpoint twice, and asserts the second call's FTP mock count is unchanged (the negative cache held); `test_check_backs_off_after_github_rate_limit` in `test_updates_api.py` patches `httpx.AsyncClient` to return a 403 with `X-RateLimit-Reset` set 10 minutes ahead and asserts the second `/updates/check` request never reaches httpx and surfaces `retry_after_seconds > 0`. 134 printers + updates integration tests green; ruff clean.
- **Assign Spool: printer card refreshes immediately, no Force-refresh needed (#1414 follow-up, reported by @snozzlebert)** — After assigning a spool via the modal, the Filament page Location column updated correctly but the Printer card kept showing "Empty slot (External Slot 1)" until the user manually pressed Force-refresh. The MQTT command itself was going through fine; the gap was on the client side. `AssignSpoolModal`'s two `useMutation.onSuccess` callbacks invalidated the inventory / slot-assignment queries (Filament page reads from those — correct) but never invalidated `['printerStatus', printerId]` and never issued a `pushall` to make the printer republish its state. For Bambu RFID-tagged spools the printer echoes the new `tray_type` over MQTT on its own and the websocket push would eventually surface it, but for non-RFID spools and A1 mini external slots (reporter's case) the firmware doesn't volunteer that state change, so the card sat on stale `tray_type: ""` and `getEmptySlotKind()` rendered the "Empty slot" path. Fix adds a `nudgePrinterRepublish()` helper called from both `onSuccess` paths (internal-inventory `assignMutation` and `assignSpoolmanMutation`): calls `api.refreshPrinterStatus(printerId)` to issue the pushall (same call the Force-refresh button uses at `PrintersPage.tsx:1844`) and invalidates `['printerStatus', printerId]` so the refetch lands. Failures from `refreshPrinterStatus` are deliberately swallowed — the assignment itself already succeeded, and if the refresh nudge is offline the next regular poll / websocket update will catch up; we don't want to surface a misleading "assign failed" toast for a stale-cache cleanup that didn't go through. Mirrors the pattern `ConfigureAmsSlotModal` has used since #1235 (line 5346) but with the extra pushall step because assign-spool affects firmware-side state where configure-slot affects only client-side preset mapping. Same fix covers both inventory modes thanks to the shared helper. **Tests:** new `nudges the printer to republish after successful assignment (#1414)` in `AssignSpoolModal.test.tsx` — picks a material-matching spool to bypass the mismatch-confirm dialog, clicks the Polymaker spool card to select it, clicks "Assign Spool", asserts `api.refreshPrinterStatus(7)` was called. `printerId=7` (not the default 1) verifies the helper threads the prop value through correctly rather than hardcoding. The api mock at the top of the file gains `assignSpoolmanSlot`, `getSpoolmanSlotAssignments`, and `refreshPrinterStatus` so the existing 13 tests still pass alongside the new one. 14 modal tests + build clean.
+5 -2
View File
@@ -22,8 +22,11 @@
> **NSS:** We use GPIO23 for manual chip-select instead of the default SPI CE0
> (GPIO8) because the kernel SPI driver's automatic CS timing does not meet the
> PN5180's requirements (5µs setup, 100µs hold). Manual CS via GPIO23 with
> `spidev.no_cs = True` resolves this.
> PN5180's requirements (5µs setup, 100µs hold). The reader's NSS line is wired
> to GPIO23 only, so whether the kernel auto-toggles CE0 is electrically
> invisible to the PN5180. Pi 4 and Pi 5 are both supported — the code asks
> the driver to disable CE0 toggling but tolerates Pi 5's RP1 driver rejecting
> that request (#1424).
### Setup Steps
+8 -1
View File
@@ -121,7 +121,14 @@ class PN5180:
self._spi.open(SPI_BUS, SPI_DEVICE)
self._spi.max_speed_hz = SPI_SPEED_HZ
self._spi.mode = 0b00
self._spi.no_cs = True
# #1424: Pi 5's RP1 spi-rp1 driver rejects SPI_NO_CS, which used to
# work on Pi 4. Harmless either way on this hardware — NSS is wired
# to GPIO23 (manual CS in _cs_low/_cs_high), so the kernel's CE0
# toggling has no electrical effect on the reader.
try:
self._spi.no_cs = True
except OSError as e:
logger.debug("spidev.no_cs not supported (likely Pi 5 RP1 driver): %s", e)
def close(self):
self._spi.close()
+8 -1
View File
@@ -120,7 +120,14 @@ class PN5180:
self._spi.open(SPI_BUS, SPI_DEVICE)
self._spi.max_speed_hz = SPI_SPEED_HZ
self._spi.mode = 0b00
self._spi.no_cs = True
# #1424: Pi 5's RP1 spi-rp1 driver rejects SPI_NO_CS, which used to
# work on Pi 4. Harmless either way on this hardware — NSS is wired
# to GPIO23 (manual CS in _cs_low/_cs_high), so the kernel's CE0
# toggling has no electrical effect on the reader.
try:
self._spi.no_cs = True
except OSError:
pass
def close(self):
self._spi.close()