4 Commits
Author SHA1 Message Date
maziggy 6488a33488 Stop a print with no 3MF borrowing another model's data (#2843)
H2-series and P2S firmware keeps a slicer-sent file on internal eMMC.
Port 990 serves external storage only, so there is no file to fetch, and
the print becomes an archive with no 3MF -- the ordinary outcome for
anyone who sends from Bambu Studio rather than through Bambuddy.
Confirmed on the maintainer's own machines: an H2C and an H2D both
dispatched brtc://emmc for the same model one minute apart, while an X1C
sent ftp:// for it. Three separate defects live in what happens next.

The first is the serious one. An archive with no 3MF keeps the path the
printer is executing as its filename, and on a sliced job that is always
Metadata/plate_1.gcode. The fallback that looks for the same model in the
Library or among earlier prints took its search term from there, so it
searched for `plate_1` -- a name every Bambu print in existence has --
and matched on a substring, so it also matched any name merely ending
that way. On the H2D a 1.6 g Cube resolved to lid_plate_1.gcode.3mf and
was costed at 207 g across three real spools. It was not confined to
plate names either: in the same database `Bank.3mf` matched "Piggo the
piggy bank", and `x1c.gcode.3mf` matched "slice-test-x1c". The matcher
now takes the model name the printer reports when the filename is only a
plate path, refuses a bare plate stem rather than searching for it, and
anchors to a whole filename with LIKE metacharacters escaped, because `_`
is a wildcard and model names are full of them. A print that cannot be
identified is now left untracked, which is the honest answer -- the
previous behaviour was to charge the operator's spools for a model they
had not printed.

Checked against every row rather than argued from the code. Across 273
library stems the result sets are identical. Across 241 archive stems 14
differ, all of them strictly narrower, and every dropped match is one of
the false positives above; all 233 archives still match their own
filename, so no legitimate donor was lost. Of the eight no-3MF archives
on that install the old matcher picked a wrong donor for two -- one of
them a calibration run that would have been charged the 207 g -- and the
new one picks none.

The second defect is that those archives could not receive a timelapse at
all. attach_timelapse derived its destination from the missing file's
path, and (base_dir / "").parent is the parent of base_dir, one level
outside the data directory. In Docker that is /app, so every attempt
failed EACCES and the scan retried and discarded the video 25 times over
twelve minutes, roughly a hundred FTPS connections for bytes that had
already downloaded successfully. Where that location happened to be
writable it was worse: the file landed beside the installation and the
attach then failed anyway, because the path could not be made relative to
base_dir. #1820 introduced a shared helper precisely so these derivations
could not drift apart, and this was the one site still doing it by hand.
The directory is created only after the filename has passed the traversal
check, so a rejected name still leaves nothing behind.

The third is silence. When a print's filament cannot be read from a 3MF,
the remaining-percentage delta is the fallback, and that needs a reading
at print start -- which a spool without RFID does not have until someone
sets a remaining amount by hand. Those slots were skipped with a bare
continue. Every other reason for skipping a slot in that loop is logged,
and the comment a few lines below argues the case explicitly: charging
nothing silently is indistinguishable from having nothing to charge. It
now says so, for slots the print actually used.

Four existing tests needed updating rather than the production path.
They patch backend.app.services.archive.settings by name, and the shared
helper reads its own module-level binding, so they kept the real data
directory and wrote outside tmp_path -- which is how the first draft of
this change littered a working tree. They now patch both bindings.
2026-08-17 09:26:13 +02:00
maziggy 91acac2b35 Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing
but a filename. Bambuddy opened port 990, the printer accepted the connection
and answered with something that was not TLS, and connect() logged a warning
and returned False -- indistinguishable, to every caller, from "the file is
not at this path". So the 3MF lookup walked all six filename variants across
five directories with four retries each, the cover endpoint ran its own
sixteen-path sweep, and the timelapse scan added four more, all against a
sixteen-path sweep, and the timelapse scan added four more, all against a
printer that could not have answered any of them. One reporter's log carried
1813 identical handshake failures, another's 3511.

The evidence says this is the printer's own file service getting stuck, not a
model, firmware or TLS-configuration problem. In #2780's bundle the same two
printers ran clean from 22 July to 4 August and failed again from the 5th; a
second bundle shows an X2D serving files for five days, flipping on 19 July,
then failing every connection for eight days with zero successes. The same
models and firmware appear in roughly twenty other bundles with no occurrences
at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D
and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly
this symptom, and the reporter's own debug line proves they do not.

An ssl.SSLError from connect() now opens a five-minute cool-off for that
printer. Subsequent connects return False without touching the network, so a
wedged printer is contacted twice an hour instead of hundreds of times a
minute, and the single warning that is logged names the remedy. The cool-off
is dropped on expiry rather than kept, so the map holds one key per currently
wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup
abandons the remaining paths and skips the directory-walk fallback, the cover
endpoint returns 503 naming the file service instead of a 404 that reads as
"this print has no thumbnail", and the timelapse scan separates 503 (cannot
reach the printer) from 404 (no timelapse directory) -- one 500 used to cover
both, which is what the reporter hit when reproducing.

The Connection Diagnostic completed a bare TCP connect to 990, which is why it
reported the port green throughout: the port is open, it is what is behind it
that is broken. It now completes a real implicit-TLS handshake using the
model's own ftp_profiles cap, so a pass means the FTP client would also get
through. An open port that cannot negotiate reports warn with reason no_tls,
selecting a new message in all 13 locales that points at a printer restart
rather than at the firewall. No login is attempted, so this stays valid in the
pre-save Add Printer flow.

The cool-off tests run against a real socket that accepts on 990 and replies
with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than
mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off
map too -- every test here talks to 127.0.0.1, so one left behind would make
the next test's connect() a no-op.
2026-08-08 09:00:03 +02:00
maziggy 5a67dffe4f fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse
found nothing afterwards. Across 247 support bundles this was the norm, not an
edge case: 457 automatic scans scheduled, 262 attached.

The scan looked four times over ~65s. The attempt that found the video was #1
272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
files were still arriving when we stopped. What ran afterwards searched for the
print name inside the filename; Bambu only writes "video_<timestamp>", so it
fired 159 times and matched zero.

The manual Scan had no baseline at all and matched on filename timestamp, FTP
mtime, or "there is only one video" — all reading a clock a LAN-only printer
cannot sync. The reporter's P1S was six and a half days out.

- Poll for minutes instead of ~65s; drop the name-match fallback.
- Persist the print-start baseline on the archive, so the diff survives a
  restart mid-print and the manual Scan runs the same comparison. With a
  baseline present the clock-based strategies are skipped entirely — they can
  only turn an honest "pick one" into a confident wrong answer.
- When several files are new (a previous print's video landing late), exclude
  the ones already attached to another archive instead of ordering the
  candidates. Ordering could only be done on the printer's clock.
- Delete the video from the printer once archived. Keeps /timelapse to
  unclaimed files, which is what makes the diff unambiguous, and stops P1S
  cards filling with AVIs.
- Gate that delete on a verified transfer: download_file now compares against
  the size from the listing. An FTPS connection closing early does not always
  raise, so a partial buffer was being attached as a complete video — which
  would also have been the one case where deleting the source lost data.

Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
the deadline stops bounding the loop as soon as the sleeps are shortened.
Per-round logging only speaks when the listing changed — 31 rounds of full
listings would bury the interesting line in the support bundle.

Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
both dialects so a migrated database matches a fresh one.

-----------

fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)

When a print records a timelapse, its last frame is the better finish photo:
the firmware stops recording with the toolhead parked and before the end
G-code drops the bed, where a live grab at that moment catches a lowered
plate. Bambuddy waited 60s for the video and then gave up, because the
print-complete notification blocks on that photo and holding a notification
for minutes is worse than sending it with the live grab.

P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
worst 546s, while every other model finished inside 26s. So the printers that
most needed the better framing were the ones that never got it.

Keep the notification on the same bound, and keep waiting off to the side.
_capture_finish_photo_from_timelapse now reports whether it ran out of time or
concluded — a video that landed and failed extraction is not worth retrying,
one that never arrived is. On the first, schedule a background task that waits
up to 15 minutes and inserts the extracted frame at the front of the archive's
photo list, where the gallery opens.

The live grab stays on disk: the notification already links to that exact
file, so removing it would leave a broken image in Discord or Telegram.

The length check proves we received what the listing said, not that the file
was finished. The first look happens ~5s after the print ends, while the
printer may still be writing, so a growing file can be listed short, served
short, and pass. Re-list after the download and only accept the video once its
size has stopped changing — a failed re-list counts as not settled, since
"could not check" must not mean "safe to delete".
2026-07-29 11:51:30 +02:00
maziggy cc75a24371 fix(db): stop holding pooled connections across FTP/camera/SMTP work (#2572)
The remaining routes of the idle-in-transaction class: the file-manager,
storage, camera-snapshot and timelapse routes each took their printer row
via Depends(get_db) and then talked FTP/camera on the same held session, so
a farm dashboard polling cover/snapshot tiles (offline printers included)
crept the pool to exhaustion over ~23h. They now read in a short session and
release before the I/O; timelapse re-opens a fresh session only for the write.

Also caps the four bare-executor FTP helpers with asyncio.wait_for so a
saturated 48-worker pool can't pin a caller (and its DB connection)
indefinitely, and runs the synchronous smtplib send off the event loop with
an explicit timeout so a wedged relay can't freeze the loop.
2026-07-18 09:10:08 +02:00