mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-10-04 13:11:35 +02:00
6e6c0e06449775565ea7725a264b08806fd232b3
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
91acac2b35 |
Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing but a filename. Bambuddy opened port 990, the printer accepted the connection and answered with something that was not TLS, and connect() logged a warning and returned False -- indistinguishable, to every caller, from "the file is not at this path". So the 3MF lookup walked all six filename variants across five directories with four retries each, the cover endpoint ran its own sixteen-path sweep, and the timelapse scan added four more, all against a sixteen-path sweep, and the timelapse scan added four more, all against a printer that could not have answered any of them. One reporter's log carried 1813 identical handshake failures, another's 3511. The evidence says this is the printer's own file service getting stuck, not a model, firmware or TLS-configuration problem. In #2780's bundle the same two printers ran clean from 22 July to 4 August and failed again from the 5th; a second bundle shows an X2D serving files for five days, flipping on 19 July, then failing every connection for eight days with zero successes. The same models and firmware appear in roughly twenty other bundles with no occurrences at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly this symptom, and the reporter's own debug line proves they do not. An ssl.SSLError from connect() now opens a five-minute cool-off for that printer. Subsequent connects return False without touching the network, so a wedged printer is contacted twice an hour instead of hundreds of times a minute, and the single warning that is logged names the remedy. The cool-off is dropped on expiry rather than kept, so the map holds one key per currently wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup abandons the remaining paths and skips the directory-walk fallback, the cover endpoint returns 503 naming the file service instead of a 404 that reads as "this print has no thumbnail", and the timelapse scan separates 503 (cannot reach the printer) from 404 (no timelapse directory) -- one 500 used to cover both, which is what the reporter hit when reproducing. The Connection Diagnostic completed a bare TCP connect to 990, which is why it reported the port green throughout: the port is open, it is what is behind it that is broken. It now completes a real implicit-TLS handshake using the model's own ftp_profiles cap, so a pass means the FTP client would also get through. An open port that cannot negotiate reports warn with reason no_tls, selecting a new message in all 13 locales that points at a printer restart rather than at the firewall. No login is attempted, so this stays valid in the pre-save Add Printer flow. The cool-off tests run against a real socket that accepts on 990 and replies with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off map too -- every test here talks to 127.0.0.1, so one left behind would make the next test's connect() a no-op. |
||
|
|
5a67dffe4f |
fix(timelapse): poll longer, diff without a clock, delete once archived (#2704)
Timelapse was on, the video never reached the archive, and Scan for Timelapse found nothing afterwards. Across 247 support bundles this was the norm, not an edge case: 457 automatic scans scheduled, 262 attached. The scan looked four times over ~65s. The attempt that found the video was #1 272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e. files were still arriving when we stopped. What ran afterwards searched for the print name inside the filename; Bambu only writes "video_<timestamp>", so it fired 159 times and matched zero. The manual Scan had no baseline at all and matched on filename timestamp, FTP mtime, or "there is only one video" — all reading a clock a LAN-only printer cannot sync. The reporter's P1S was six and a half days out. - Poll for minutes instead of ~65s; drop the name-match fallback. - Persist the print-start baseline on the archive, so the diff survives a restart mid-print and the manual Scan runs the same comparison. With a baseline present the clock-based strategies are skipped entirely — they can only turn an honest "pick one" into a confident wrong answer. - When several files are new (a previous print's video landing late), exclude the ones already attached to another archive instead of ordering the candidates. Ordering could only be done on the printer's clock. - Delete the video from the printer once archived. Keeps /timelapse to unclaimed files, which is what makes the diff unambiguous, and stops P1S cards filling with AVIs. - Gate that delete on a verified transfer: download_file now compares against the size from the listing. An FTPS connection closing early does not always raise, so a partial buffer was being attached as a complete video — which would also have been the one case where deleting the source lost data. Bounded twice on purpose: wall-clock deadline plus a derived round cap, since the deadline stops bounding the loop as soon as the sleeps are shortened. Per-round logging only speaks when the listing changed — 31 rounds of full listings would bury the interesting line in the support bundle. Migration adds print_archives.timelapse_baseline as JSON, spelled the same on both dialects so a migrated database matches a fresh one. ----------- fix(finish-photo): add the timelapse frame to the archive after the notification (#2704) When a print records a timelapse, its last frame is the better finish photo: the firmware stops recording with the toolhead parked and before the end G-code drops the bed, where a live grab at that moment catches a lowered plate. Bambuddy waited 60s for the video and then gave up, because the print-complete notification blocks on that photo and holding a notification for minutes is worse than sending it with the live grab. P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly. Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s, worst 546s, while every other model finished inside 26s. So the printers that most needed the better framing were the ones that never got it. Keep the notification on the same bound, and keep waiting off to the side. _capture_finish_photo_from_timelapse now reports whether it ran out of time or concluded — a video that landed and failed extraction is not worth retrying, one that never arrived is. On the first, schedule a background task that waits up to 15 minutes and inserts the extracted frame at the front of the archive's photo list, where the gallery opens. The live grab stays on disk: the notification already links to that exact file, so removing it would leave a broken image in Discord or Telegram. The length check proves we received what the listing said, not that the file was finished. The first look happens ~5s after the print ends, while the printer may still be writing, so a growing file can be listed short, served short, and pass. Re-list after the download and only accept the video once its size has stopped changing — a failed re-list counts as not settled, since "could not check" must not mean "safe to delete". |
||
|
|
cc75a24371 |
fix(db): stop holding pooled connections across FTP/camera/SMTP work (#2572)
The remaining routes of the idle-in-transaction class: the file-manager, storage, camera-snapshot and timelapse routes each took their printer row via Depends(get_db) and then talked FTP/camera on the same held session, so a farm dashboard polling cover/snapshot tiles (offline printers included) crept the pool to exhaustion over ~23h. They now read in a short session and release before the I/O; timelapse re-opens a fresh session only for the write. Also caps the four bare-executor FTP helpers with asyncio.wait_for so a saturated 48-worker pool can't pin a caller (and its DB connection) indefinitely, and runs the synchronous smtplib send off the event loop with an explicit timeout so a wedged relay can't freeze the loop. |