Files
bambuddy/backend
maziggy a1e5afbd2d Keep a dispatch's retries out of the FTPS cool-off (issue #2898)
A failed TLS handshake arms a 300s per-IP cool-off, and connect()
    consulted it for every caller. A print dispatch retries after 2s, so
    once the cool-off was armed all four attempts were answered from the
    gate rather than the network, and every further job queued for that
    printer failed the same way for the rest of the window. The reporter's
    farm lost three jobs to one handshake error, with the retry budget
    contributing nothing to any of them.

    The gate was serving two callers that want opposite things from it. The
    background sweeps -- the post-print 3MF, cover and timelapse fetches --
    walk ~110 candidate paths against one wedged printer with nobody
    waiting, and backing off for minutes is right for them. A dispatch is
    one delete plus at most four upload attempts with someone watching a
    progress bar. So the split is by caller: a client built with
    respect_handshake_cooloff=False goes to the printer regardless, and the
    dispatch's delete and upload -- and a firmware upload, same shape --
    opt out. Everything else keeps #2780's behaviour untouched.

    In the reported trace it is the pre-upload delete that takes the SSL
    error and arms the cool-off, 8ms before the upload's first attempt, so
    exempting the upload alone would have left one dispatch's worth of the
    problem in place.

    Callers that do respect the cool-off no longer sleep out a retry loop
    against it: with_ftp_retry takes the printer's IP and stops at the
    attempt that armed the gate, instead of spending three more attempts
    and six seconds on connections that cannot happen. It also reports the
    attempts it really made -- "failed after 4 attempts" for one attempt is
    part of how this read as a network problem.

    Two diagnosis fixes go with it. The cool-off skip was the one connect()
    failure path that reported without naming its cause, and at DEBUG, so
    four identical reason-free warnings were all the operator saw. It now
    says at WARNING that nothing was sent and how long the printer has
    left, once per cool-off rather than once per attempt -- not every
    caller is gated, and a download-zip of 200 files would otherwise repeat
    the sentence 200 times, which is the flood #2780 set out to stop.

    And a dispatch that fails this way no longer tells anyone to check
    whether the SD card is inserted and formatted -- nothing reached the
    printer's filesystem, so the card is the one part of the machine that
    was working. The message names the file service and rules the card out.
    It is used only when a handshake failed during the dispatch itself,
    read from the cool-off deadline MOVING rather than merely being armed:
    the dispatch ignores the gate, so it can be running underneath one an
    unrelated background fetch left behind, and blaming TLS for an upload
    that really hit a full disk would repeat the mistake in the other
    direction.

    Tests count sockets rather than return values, since "returned False"
    looks identical whether or not anything was attempted -- which is what
    made the original report a log dive. Reverting any one of the five
    behaviours above fails a distinct test.
2026-08-23 12:43:00 +02:00
..