mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-09-30 11:12:35 +02:00
dev
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6488a33488 |
Stop a print with no 3MF borrowing another model's data (#2843)
H2-series and P2S firmware keeps a slicer-sent file on internal eMMC. Port 990 serves external storage only, so there is no file to fetch, and the print becomes an archive with no 3MF -- the ordinary outcome for anyone who sends from Bambu Studio rather than through Bambuddy. Confirmed on the maintainer's own machines: an H2C and an H2D both dispatched brtc://emmc for the same model one minute apart, while an X1C sent ftp:// for it. Three separate defects live in what happens next. The first is the serious one. An archive with no 3MF keeps the path the printer is executing as its filename, and on a sliced job that is always Metadata/plate_1.gcode. The fallback that looks for the same model in the Library or among earlier prints took its search term from there, so it searched for `plate_1` -- a name every Bambu print in existence has -- and matched on a substring, so it also matched any name merely ending that way. On the H2D a 1.6 g Cube resolved to lid_plate_1.gcode.3mf and was costed at 207 g across three real spools. It was not confined to plate names either: in the same database `Bank.3mf` matched "Piggo the piggy bank", and `x1c.gcode.3mf` matched "slice-test-x1c". The matcher now takes the model name the printer reports when the filename is only a plate path, refuses a bare plate stem rather than searching for it, and anchors to a whole filename with LIKE metacharacters escaped, because `_` is a wildcard and model names are full of them. A print that cannot be identified is now left untracked, which is the honest answer -- the previous behaviour was to charge the operator's spools for a model they had not printed. Checked against every row rather than argued from the code. Across 273 library stems the result sets are identical. Across 241 archive stems 14 differ, all of them strictly narrower, and every dropped match is one of the false positives above; all 233 archives still match their own filename, so no legitimate donor was lost. Of the eight no-3MF archives on that install the old matcher picked a wrong donor for two -- one of them a calibration run that would have been charged the 207 g -- and the new one picks none. The second defect is that those archives could not receive a timelapse at all. attach_timelapse derived its destination from the missing file's path, and (base_dir / "").parent is the parent of base_dir, one level outside the data directory. In Docker that is /app, so every attempt failed EACCES and the scan retried and discarded the video 25 times over twelve minutes, roughly a hundred FTPS connections for bytes that had already downloaded successfully. Where that location happened to be writable it was worse: the file landed beside the installation and the attach then failed anyway, because the path could not be made relative to base_dir. #1820 introduced a shared helper precisely so these derivations could not drift apart, and this was the one site still doing it by hand. The directory is created only after the filename has passed the traversal check, so a rejected name still leaves nothing behind. The third is silence. When a print's filament cannot be read from a 3MF, the remaining-percentage delta is the fallback, and that needs a reading at print start -- which a spool without RFID does not have until someone sets a remaining amount by hand. Those slots were skipped with a bare continue. Every other reason for skipping a slot in that loop is logged, and the comment a few lines below argues the case explicitly: charging nothing silently is indistinguishable from having nothing to charge. It now says so, for slots the print actually used. Four existing tests needed updating rather than the production path. They patch backend.app.services.archive.settings by name, and the shared helper reads its own module-level binding, so they kept the real data directory and wrote outside tmp_path -- which is how the first draft of this change littered a working tree. They now patch both bindings. |
||
|
|
91acac2b35 |
Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing but a filename. Bambuddy opened port 990, the printer accepted the connection and answered with something that was not TLS, and connect() logged a warning and returned False -- indistinguishable, to every caller, from "the file is not at this path". So the 3MF lookup walked all six filename variants across five directories with four retries each, the cover endpoint ran its own sixteen-path sweep, and the timelapse scan added four more, all against a sixteen-path sweep, and the timelapse scan added four more, all against a printer that could not have answered any of them. One reporter's log carried 1813 identical handshake failures, another's 3511. The evidence says this is the printer's own file service getting stuck, not a model, firmware or TLS-configuration problem. In #2780's bundle the same two printers ran clean from 22 July to 4 August and failed again from the 5th; a second bundle shows an X2D serving files for five days, flipping on 19 July, then failing every connection for eight days with zero successes. The same models and firmware appear in roughly twenty other bundles with no occurrences at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly this symptom, and the reporter's own debug line proves they do not. An ssl.SSLError from connect() now opens a five-minute cool-off for that printer. Subsequent connects return False without touching the network, so a wedged printer is contacted twice an hour instead of hundreds of times a minute, and the single warning that is logged names the remedy. The cool-off is dropped on expiry rather than kept, so the map holds one key per currently wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup abandons the remaining paths and skips the directory-walk fallback, the cover endpoint returns 503 naming the file service instead of a 404 that reads as "this print has no thumbnail", and the timelapse scan separates 503 (cannot reach the printer) from 404 (no timelapse directory) -- one 500 used to cover both, which is what the reporter hit when reproducing. The Connection Diagnostic completed a bare TCP connect to 990, which is why it reported the port green throughout: the port is open, it is what is behind it that is broken. It now completes a real implicit-TLS handshake using the model's own ftp_profiles cap, so a pass means the FTP client would also get through. An open port that cannot negotiate reports warn with reason no_tls, selecting a new message in all 13 locales that points at a printer restart rather than at the firewall. No login is attempted, so this stays valid in the pre-save Add Printer flow. The cool-off tests run against a real socket that accepts on 990 and replies with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off map too -- every test here talks to 127.0.0.1, so one left behind would make the next test's connect() a no-op. |
||
|
|
5a67dffe4f |
fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse found nothing afterwards. Across 247 support bundles this was the norm, not an edge case: 457 automatic scans scheduled, 262 attached. The scan looked four times over ~65s. The attempt that found the video was #1 272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e. files were still arriving when we stopped. What ran afterwards searched for the print name inside the filename; Bambu only writes "video_<timestamp>", so it fired 159 times and matched zero. The manual Scan had no baseline at all and matched on filename timestamp, FTP mtime, or "there is only one video" — all reading a clock a LAN-only printer cannot sync. The reporter's P1S was six and a half days out. - Poll for minutes instead of ~65s; drop the name-match fallback. - Persist the print-start baseline on the archive, so the diff survives a restart mid-print and the manual Scan runs the same comparison. With a baseline present the clock-based strategies are skipped entirely — they can only turn an honest "pick one" into a confident wrong answer. - When several files are new (a previous print's video landing late), exclude the ones already attached to another archive instead of ordering the candidates. Ordering could only be done on the printer's clock. - Delete the video from the printer once archived. Keeps /timelapse to unclaimed files, which is what makes the diff unambiguous, and stops P1S cards filling with AVIs. - Gate that delete on a verified transfer: download_file now compares against the size from the listing. An FTPS connection closing early does not always raise, so a partial buffer was being attached as a complete video — which would also have been the one case where deleting the source lost data. Bounded twice on purpose: wall-clock deadline plus a derived round cap, since the deadline stops bounding the loop as soon as the sleeps are shortened. Per-round logging only speaks when the listing changed — 31 rounds of full listings would bury the interesting line in the support bundle. Migration adds print_archives.timelapse_baseline as JSON, spelled the same on both dialects so a migrated database matches a fresh one. ----------- fix(finish-photo): add the timelapse frame to the archive after the notification (#2704) When a print records a timelapse, its last frame is the better finish photo: the firmware stops recording with the toolhead parked and before the end G-code drops the bed, where a live grab at that moment catches a lowered plate. Bambuddy waited 60s for the video and then gave up, because the print-complete notification blocks on that photo and holding a notification for minutes is worse than sending it with the live grab. P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly. Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s, worst 546s, while every other model finished inside 26s. So the printers that most needed the better framing were the ones that never got it. Keep the notification on the same bound, and keep waiting off to the side. _capture_finish_photo_from_timelapse now reports whether it ran out of time or concluded — a video that landed and failed extraction is not worth retrying, one that never arrived is. On the first, schedule a background task that waits up to 15 minutes and inserts the extracted frame at the front of the archive's photo list, where the gallery opens. The live grab stays on disk: the notification already links to that exact file, so removing it would leave a broken image in Discord or Telegram. The length check proves we received what the listing said, not that the file was finished. The first look happens ~5s after the print ends, while the printer may still be writing, so a growing file can be listed short, served short, and pass. Re-list after the download and only accept the video once its size has stopped changing — a failed re-list counts as not settled, since "could not check" must not mean "safe to delete". |
||
|
|
cc75a24371 |
fix(db): stop holding pooled connections across FTP/camera/SMTP work (#2572)
The remaining routes of the idle-in-transaction class: the file-manager, storage, camera-snapshot and timelapse routes each took their printer row via Depends(get_db) and then talked FTP/camera on the same held session, so a farm dashboard polling cover/snapshot tiles (offline printers included) crept the pool to exhaustion over ~23h. They now read in a short session and release before the I/O; timelapse re-opens a fresh session only for the write. Also caps the four bare-executor FTP helpers with asyncio.wait_for so a saturated 48-worker pool can't pin a caller (and its DB connection) indefinitely, and runs the synchronous smtplib send off the event loop with an explicit timeout so a wedged relay can't freeze the loop. |