5 Commits
Author SHA1 Message Date
Maksim Sadontsev ed85677913 Light the generated thumbnails so one model differs from another (#2816) (#2861) 2026-08-22 14:16:56 +02:00
maziggy 7c8f1f9435 Keep a dispatch's retries out of the FTPS cool-off (issue #2898)
A failed TLS handshake arms a 300s per-IP cool-off, and connect()
consulted it for every caller. A print dispatch retries after 2s, so
once the cool-off was armed all four attempts were answered from the
gate rather than the network, and every further job queued for that
printer failed the same way for the rest of the window. The reporter's
farm lost three jobs to one handshake error, with the retry budget
contributing nothing to any of them.

The gate was serving two callers that want opposite things from it. The
background sweeps -- the post-print 3MF, cover and timelapse fetches --
walk ~110 candidate paths against one wedged printer with nobody
waiting, and backing off for minutes is right for them. A dispatch is
one delete plus at most four upload attempts with someone watching a
progress bar. So the split is by caller: a client built with
respect_handshake_cooloff=False goes to the printer regardless, and the
dispatch's delete and upload -- and a firmware upload, same shape --
opt out. Everything else keeps #2780's behaviour untouched.

In the reported trace it is the pre-upload delete that takes the SSL
error and arms the cool-off, 8ms before the upload's first attempt, so
exempting the upload alone would have left one dispatch's worth of the
problem in place.

Callers that do respect the cool-off no longer sleep out a retry loop
against it: with_ftp_retry takes the printer's IP and stops at the
attempt that armed the gate, instead of spending three more attempts
and six seconds on connections that cannot happen. It also reports the
attempts it really made -- "failed after 4 attempts" for one attempt is
part of how this read as a network problem.

Two diagnosis fixes go with it. The cool-off skip was the one connect()
failure path that reported without naming its cause, and at DEBUG, so
four identical reason-free warnings were all the operator saw. It now
says at WARNING that nothing was sent and how long the printer has
left, once per cool-off rather than once per attempt -- not every
caller is gated, and a download-zip of 200 files would otherwise repeat
the sentence 200 times, which is the flood #2780 set out to stop.

And a dispatch that fails this way no longer tells anyone to check
whether the SD card is inserted and formatted -- nothing reached the
printer's filesystem, so the card is the one part of the machine that
was working. The message names the file service and rules the card out.
It is used only when a handshake failed during the dispatch itself,
read from the cool-off deadline MOVING rather than merely being armed:
the dispatch ignores the gate, so it can be running underneath one an
unrelated background fetch left behind, and blaming TLS for an upload
that really hit a full disk would repeat the mistake in the other
direction.

Tests count sockets rather than return values, since "returned False"
looks identical whether or not anything was attempted -- which is what
made the original report a log dive. Reverting any one of the five
behaviours above fails a distinct test.
2026-08-22 10:05:05 +02:00
maziggy 91acac2b35 Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing
but a filename. Bambuddy opened port 990, the printer accepted the connection
and answered with something that was not TLS, and connect() logged a warning
and returned False -- indistinguishable, to every caller, from "the file is
not at this path". So the 3MF lookup walked all six filename variants across
five directories with four retries each, the cover endpoint ran its own
sixteen-path sweep, and the timelapse scan added four more, all against a
sixteen-path sweep, and the timelapse scan added four more, all against a
printer that could not have answered any of them. One reporter's log carried
1813 identical handshake failures, another's 3511.

The evidence says this is the printer's own file service getting stuck, not a
model, firmware or TLS-configuration problem. In #2780's bundle the same two
printers ran clean from 22 July to 4 August and failed again from the 5th; a
second bundle shows an X2D serving files for five days, flipping on 19 July,
then failing every connection for eight days with zero successes. The same
models and firmware appear in roughly twenty other bundles with no occurrences
at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D
and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly
this symptom, and the reporter's own debug line proves they do not.

An ssl.SSLError from connect() now opens a five-minute cool-off for that
printer. Subsequent connects return False without touching the network, so a
wedged printer is contacted twice an hour instead of hundreds of times a
minute, and the single warning that is logged names the remedy. The cool-off
is dropped on expiry rather than kept, so the map holds one key per currently
wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup
abandons the remaining paths and skips the directory-walk fallback, the cover
endpoint returns 503 naming the file service instead of a 404 that reads as
"this print has no thumbnail", and the timelapse scan separates 503 (cannot
reach the printer) from 404 (no timelapse directory) -- one 500 used to cover
both, which is what the reporter hit when reproducing.

The Connection Diagnostic completed a bare TCP connect to 990, which is why it
reported the port green throughout: the port is open, it is what is behind it
that is broken. It now completes a real implicit-TLS handshake using the
model's own ftp_profiles cap, so a pass means the FTP client would also get
through. An open port that cannot negotiate reports warn with reason no_tls,
selecting a new message in all 13 locales that points at a printer restart
rather than at the firewall. No login is attempted, so this stays valid in the
pre-save Add Printer flow.

The cool-off tests run against a real socket that accepts on 990 and replies
with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than
mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off
map too -- every test here talks to 127.0.0.1, so one left behind would make
the next test's connect() a no-op.
2026-08-08 09:00:03 +02:00
maziggy b886590339 Speed up FTP test suite with class-scoped server fixtures
Function-scoped ftp_server meant 67 TLS server start/stop cycles.
Class-scoped reduces to 10. Adds per-test cleanup fixture to reset
failure injections and filesystem between tests. Isolates
test_disconnect_after_server_gone into its own class to prevent
close_all() from nuking other servers' asyncore sockets.
2026-02-10 18:11:25 +01:00
maziggy f2468077fe Add mock FTPS server and comprehensive FTP test suite (67 tests)
FTP bugs have been the #1 recurring issue across releases (0.1.8+).
This adds a real implicit FTPS mock server and 67 test cases covering
every known failure mode — connection, upload, download, delete, storage
info, model-specific SSL behavior, async wrappers, and failure injection.

New files:
- mock_ftp_server.py: implicit FTPS server on pyftpdlib with failure injection
- conftest.py: FTP test fixtures (certs, server, client factory)
- test_bambu_ftp.py: 67 tests across 10 test classes

Also adds pyOpenSSL to requirements-dev.txt (needed by pyftpdlib
TLS_FTPHandler in the Docker test image).
2026-02-07 10:55:05 +01:00