Say why a drying cycle ended when the firmware cuts it short (#2770)

An H2D started a twelve-hour PETG dry at 65 degC and the AMS gave up on it
twenty minutes in, with 700 of the 720 minutes still on the clock. It cooled,
humidity climbed back over the threshold, auto-drying started another
twelve-hour cycle, and that one went the same way; the reporter's AMS
temperature history shows the loop running all morning.

The log had one line for it: "AMS 0 drying complete", which is exactly what it
says for a dry that ran its full twelve hours. Nothing in a support bundle told
the two apart, and the one number that does -- the time still remaining -- was
written into that line as the previous value, where it reads like a duration
rather than a shortfall. The reporter took 700 for seconds and concluded the
cycle had lasted twelve minutes.

Bambuddy did not stop that cycle; every stop it sends is logged with the full
outgoing command and there was none. So ending it was the printer's decision,
and the account of why lives in three things already received and parsed and
never written down: the drying phase and sub-phase from the AMS info hex, the
per-unit dry_sf_reason constraint codes, and the live HMS errors.

A cycle that ends with most of its countdown left now logs all three alongside
how much of the requested duration ran. One that reaches its duration keeps the
single line it has always had. A stop Bambuddy sent is named as ours -- it is
short of its duration too, and on the telemetry alone is indistinguishable from
the firmware abandoning the cycle, so without tracking it the print-takes-priority
stop and the Stop button would both have been blamed on the printer.

Diagnostics only. Nothing about when drying starts or stops has changed, and the
restart loop is not addressed: what the firmware objects to has to be established
before Bambuddy can sensibly decide how long to wait before trying again.
This commit is contained in:
maziggy
2026-08-05 12:38:45 +02:00
parent 0dc911fce1
commit 0596ff424e
3 changed files with 177 additions and 10 deletions
+1
View File
@@ -17,6 +17,7 @@ All notable changes to Bambuddy will be documented in this file.
### Fixed
- **LDAP login works again on directories that define no POSIX group class (#2769, reporter @peterskotte)** — Every LDAP user on an lldap directory was rejected with "Incorrect username or password", including users whose credentials, search filter and group membership all checked out when tested by hand with `ldapsearch`, and on an install where **Test Connection** reported success. The password was never the problem and the directory never saw the request. When resolving a user's groups Bambuddy looks for POSIX groups alongside the usual `memberOf` ones, and both of those searches name the `posixGroup` object class. The LDAP client validates class names in a filter against the schema the server publishes, and rejects an unknown one while building the request, before anything is sent. lldap marks every account it creates as `posixAccount`, which is what makes Bambuddy look for POSIX groups in the first place, but defines no group class beyond `groupOfNames` — so the search was refused, the refusal travelled all the way out of the login routine, and the login route reports any LDAP failure as bad credentials. A directory with no `posixGroup` class has no `posixGroup` entries, which is precisely the answer those searches would have returned, so Bambuddy now treats the refusal as the empty result it stands for, notes it once in the log and carries on with the `memberOf` groups. The reporter's mapped group is one of those, so it resolves as configured. This is not a regression from the recent primary-group work, though that is the natural suspect: the `memberUid` search has named the same class since LDAP support first shipped, and it runs for every user whether or not they have a `gidNumber`, so login has never worked against a directory of this shape. **Test Connection** passed throughout because it asks only whether any entry exists, a form of filter that carries no class name to validate. Nothing changes for Active Directory or for an OpenLDAP that loads the standard NIS schema — both define the class, and their POSIX groups are still read. Wiki updated. Covered by backend tests.
- **Spoolman no longer charges a Bambu Studio print to the wrong spool (#2768)** — A sliced file numbers its filaments 1, 2, 3, 4, and which AMS tray each of those came from is a separate decision made when the job is sent. Bambuddy learns that decision one of two ways: it made the choice itself, for a print started from Bambuddy, or it read the print command as it crossed the local network, for a print sent from a slicer. A job dispatched from Bambu Studio while the printer is signed in to Bambu's cloud satisfies neither — the command travels through Bambu's own broker and never appears on the network Bambuddy is listening to. With nothing recorded, the Spoolman writer fell back to assuming the AMS was loaded in slicer order: filament 1 from the first loaded tray, filament 2 from the second. The reporter's X1C was loaded in the order 2, 4, 1, AMS-HT, so every one of the four was deducted from the wrong spool. It also changed what the print looked like afterwards: on completion Bambuddy stamps the archive with the material and colour of the spools it charged, so the print showed the right filament while it ran and switched to a different one the moment it finished — which is how the reporter noticed. The printer knew the answer all along. It publishes the running job's slot-to-tray assignment in its own status, and Bambuddy's built-in filament inventory has read that field for as long as it has resolved mappings at completion; only the Spoolman writer, which resolves at print start instead, never learned to. It now consults the same two fallbacks at the same moment: the printer's report first, and failing that a colour match of the sliced filaments against the loaded trays, which covers the A1, A1 Mini, P1S and P2S — those models publish no such field, so their owners were on the positional guess no matter how the print was sent. Reading the field at completion rather than at print start is deliberate: a printer keeps publishing the last job's mapping while it sits idle, so consulting it early risks stamping the previous print's mapping onto this one. A mapping Bambuddy or the slicer actually recorded is never second-guessed, so nothing changes for prints started from Bambuddy, from the queue, or over LAN. Cancelled and failed prints take the same correction, since partial usage is charged through the same mapping. The resolved mapping and where it came from are now logged at both print start and completion, so the next report of a wrong deduction can be read straight out of a support bundle. Wiki updated. Covered by backend tests.
- **A drying cycle the printer abandons now says so, and says what the printer reported (#2770, reporter @tchavei)** — An H2D started a twelve-hour PETG dry at 65°C and the AMS gave up on it twenty minutes in, with 700 of the 720 minutes still on the clock. It then cooled off, humidity climbed back over the threshold, auto-drying started another twelve-hour cycle, and that one was abandoned the same way — a loop the reporter's AMS temperature history shows running all morning. The log had one line to say for it: `AMS 0 drying complete`, which is exactly what it says for a dry that ran its full twelve hours. Nothing in a support bundle told the two apart, and the remaining time — the one number that does — was written into the line as the *previous* value, where it reads like a duration rather than a shortfall. Bambuddy did not stop that cycle. Every stop it sends is logged with the full command as it goes out, and there was none, so ending it was the printer's decision — and the only account of why lives in three things Bambuddy already receives and parses but has never written down: the drying phase and sub-phase the AMS reports in its status word, the firmware's own cannot-dry reason codes (which distinguish an overheating unit from one being starved of power by a missing external supply), and whatever HMS errors are live at that moment. A cycle that ends with most of its countdown left now logs all three, alongside how much of the requested duration actually ran and how much was asked for. A cycle that reaches its configured duration keeps the single line it has always had, so a normal dry does not start reporting diagnostics nobody needs, and a cycle Bambuddy itself ends — the print-takes-priority stop, or the **Stop** button — now says so by name rather than being reported as an unexplained early end, since a stop is short of its duration too and looks identical in the telemetry. This is diagnostics only: nothing about when drying starts or stops has changed, and the repeated restart itself is not addressed here — what the firmware objects to has to be established before Bambuddy can sensibly decide how long to wait before trying again. Covered by backend tests.
- **A drying cycle no longer reports itself finished a minute after it starts (#2759)** — Starting the dryer on an AMS 2 Pro holding two PETG and two PLA spools and picking PLA showed "PLA @ 45°C" for about a minute, then switched to "PETG @ 65°C" for the remaining twelve hours. Bambu never echoes back which filament or temperature a cycle is running, so the badge reads the target Bambuddy cached when it sent the command — and that cache had been thrown away. Between accepting the command and settling its countdown the firmware publishes one update with the remaining time at zero while the unit is still in its Checking phase; the reporter's log caught 720 minutes, then 0, then 719. Bambuddy read the zero as the cycle ending. Losing the cached target left the badge to guess the filament from the first loaded slot, which happened to be PETG, and its RFID-recommended 65°C — a confident wrong answer for a cycle running PLA at 45. The same false ending also armed smart-plug auto-off-after-drying, so anyone with that switched on had power scheduled to cut one minute into a twelve-hour dry. A remaining time of zero is now only treated as the end of a cycle when the AMS also reports an idle phase, which the firmware already publishes alongside it; stopping a dry early still ends it immediately, and a unit that reports no phase at all still ends its cycles as before. The fallback guess has been tightened to match, in both directions. It names a filament only when every loaded spool agrees on one — on a mixed unit the badge shows the countdown alone rather than naming a spool the cycle isn't drying — and it no longer guesses a temperature at all. A unit loaded entirely with PLA does tell you what is being dried, but not at what temperature: that is picked freely when the cycle is started, so the spools' RFID-recommended value is never evidence of it, and a second AMS loaded only with PLA and drying at 45°C still read "PLA @ 55°C" whenever the cached target went missing. The badge now names a temperature only when Bambuddy sent it, and shows the filament and countdown without one otherwise. Covered by backend and frontend tests.
- **A print that never starts now says AMS drying was running, instead of blaming the SD card (#2758)** — Sending a job to an X2D with two AMS units mid-drying failed silently: the file uploaded, the printer accepted it and then simply stayed idle. Bambuddy waited out the start watchdog, re-uploaded the whole 3MF, waited again, and after three attempts gave up with advice to check the printer's screen and the SD card — while Bambu Studio, asked directly, said it could not start the job because of the drying. Bambuddy now watches the AMS drying telemetry it already receives across the dispatch window and, when a job never starts while a unit was drying, names the units in the failure message and records the correlation in the log from the first attempt rather than only after the retries are spent. This is deliberately a diagnosis and not a rule: the printers concerned support drying *continuing* through a print, so drying and printing are not in conflict as such, and the report also involved one AMS drying without its external power supply — which would make the start-of-print calibration a power problem rather than a drying one. Stopping the cycle automatically would therefore be acting on a guess, and could tear down drying the hardware was happy to continue. Until it is known which of the two is the real obstacle, Bambuddy tells you what it saw and leaves the call to you. The message for a stalled dispatch with no drying involved is unchanged. Wiki updated. Covered by backend tests.
- **A hand-written systemd service left the Virtual Printer unable to start, with nothing obvious to blame (#2549, reporter @Ru3ck3)** — The Virtual Printer binds ports 990 and 322, both below 1024, which a service running as a normal user may not do without the `CAP_NET_BIND_SERVICE` capability. Without it the rest of Bambuddy works perfectly and only the Virtual Printer is dead: its sockets never open, the slicer never finds the printer, and the sole trace is one line in the journal. The reporter lost days to this before someone on Discord spotted the missing line. The install script has carried it since March, but the three other places that define the same service did not — the manual-install template, the combined Bambuddy plus SpoolBuddy installer, and the unit the wiki tells you to paste. All three have it now, and the wiki no longer claims the capability is always included when its own instructions omitted it. Bambuddy also diagnoses this itself: **Diagnose** on the virtual printer card previously reported only that nothing was listening on port 990, which reads identically to an ordinary port conflict. It now checks whether the process actually holds the capability and, when that is what is wrong, says so and gives the line to add. The check stays quiet when the port is answering, since fronting it another way (an iptables redirect is the documented alternative) is a legitimate setup, and it stays quiet when the capability is held, so a port that failed for some other reason is not misattributed. Existing installs are unaffected until reinstalled; the diagnostic tells you whether yours needs the line. Translated in all locales; wiki updated. Covered by backend tests.
+90 -7
View File
@@ -46,6 +46,16 @@ _ACTIVE_PRINT_STATES = frozenset({"PREPARE", "SLICING", "RUNNING", "PAUSE"})
# are deliberately excluded — those SHOULD end it.
_ACTIVE_DRY_STATUSES = frozenset({1, 2, 3}) # Checking, Drying, Cooling
# A drying cycle that runs to term ends with its countdown all but exhausted, so
# the last dry_time we saw before the drop to 0 tells us whether the firmware
# ended the cycle on schedule or aborted it. More than this many minutes still on
# the clock means it was cut short, and the firmware's own reason codes are worth
# capturing at INFO — #2770 aborted a 12-hour cycle 20 minutes in (700 minutes
# left), and the log said only "drying complete", so the report carried no
# evidence of why. The margin absorbs a stale last observation between AMS
# pushes; it is not a judgement about how short "short" is.
_EARLY_DRY_END_MINUTES = 5
# CONNACK reason codes that mean the printer actively refused our credentials,
# as opposed to being unreachable or busy. Bambu speaks MQTT 3.1.1, whose
# single-byte CONNACK return codes paho maps onto the v5 reason-code space:
@@ -750,6 +760,11 @@ class BambuMQTTClient:
# — only the dry_time countdown — so we cache what we sent to drive
# the UI badge. Cleared on stop or on the dry_time falling edge to 0.
self._drying_targets: dict[int, dict[str, object]] = {}
# AMS ids we have sent a stop for and not yet seen end. A stop always
# ends a cycle far short of its duration, which on the telemetry alone
# is indistinguishable from the firmware abandoning it — so the cycle-end
# log would otherwise blame the printer for our own decision (#2770).
self._drying_stops_sent: set[int] = set()
self.state = PrinterState()
self._client: mqtt.Client | None = None
@@ -2819,13 +2834,7 @@ class BambuMQTTClient:
previous = self._previous_dry_times.get(ams_id, 0)
self._previous_dry_times[ams_id] = current
if previous > 0 and current == 0:
logger.info(
"[%s] AMS %d drying complete (dry_time %d → 0)",
self.serial_number,
ams_id,
previous,
)
self._drying_targets.pop(ams_id, None)
self._log_drying_cycle_end(ams_id, previous, ams_unit, self._drying_targets.pop(ams_id, None))
if self.on_drying_complete:
self.on_drying_complete(ams_id)
@@ -2865,6 +2874,71 @@ class BambuMQTTClient:
if self._pending_assignments:
self._check_assignment_verifications()
def _log_drying_cycle_end(
self,
ams_id: int,
remaining: int,
ams_unit: dict,
target: dict[str, object] | None,
) -> None:
"""Report a finished drying cycle, with the firmware's reason when it was
cut short (#2770).
A cycle that reaches its configured duration needs no explanation and
keeps the one-line "drying complete" it has always had. One that ends
with most of its countdown left was ended by somebody, and there are
only two candidates: a stop Bambuddy sent — the print-takes-priority
stop, or the user's Stop button — which is named as such, or the
firmware.
For the firmware case the only account of why lives in fields we already
parse but have never written down: the ``dry_status`` /
``dry_sub_status`` phase from the info hex, the per-unit
``dry_sf_reason`` constraint codes, and whatever HMS errors are live at
that moment. Logging them at INFO puts them in every support bundle by
default, which is what a report like #2770 needs before its cause can be
argued about at all.
"""
if ams_id in self._drying_stops_sent:
self._drying_stops_sent.discard(ams_id)
logger.info(
"[%s] AMS %d drying stopped by Bambuddy (dry_time %d → 0)",
self.serial_number,
ams_id,
remaining,
)
return
if remaining <= _EARLY_DRY_END_MINUTES:
logger.info(
"[%s] AMS %d drying complete (dry_time %d → 0)",
self.serial_number,
ams_id,
remaining,
)
return
requested_minutes: int | None = None
if target is not None:
try:
requested_minutes = int(target.get("duration_hours") or 0) * 60 or None
except (TypeError, ValueError):
requested_minutes = None
logger.info(
"[%s] AMS %d drying ended early — %d of %s minutes still on the clock. "
"Bambuddy sent no stop command, so the firmware ended this cycle: "
"dry_status=%s dry_sub_status=%s dry_sf_reason=%s hms=%s",
self.serial_number,
ams_id,
remaining,
requested_minutes if requested_minutes is not None else "?",
ams_unit.get("dry_status"),
ams_unit.get("dry_sub_status"),
ams_unit.get("dry_sf_reason") or [],
[e.full_code for e in self.state.hms_errors] or "none",
)
def register_assignment_verification(
self,
ams_id: int,
@@ -5492,13 +5566,22 @@ class BambuMQTTClient:
)
# Track the active-cycle target so the badge can show "PETG @ 65°C"
# while drying. Bambu only echoes dry_time on subsequent pushes.
# duration_hours is not shown anywhere; it is what lets the cycle-end log
# say how much of the requested time the firmware actually ran (#2770).
if mode == 1:
self._drying_targets[ams_id] = {
"filament": filament or "",
"temp": int(temp),
"duration_hours": int(duration),
}
self._drying_stops_sent.discard(ams_id)
else:
self._drying_targets.pop(ams_id, None)
# Remember that this cycle's end is ours, so the cycle-end log
# attributes it to Bambuddy instead of to the firmware (#2770). A
# stop always ends the cycle far short of its duration, which is
# otherwise indistinguishable from the firmware abandoning it.
self._drying_stops_sent.add(ams_id)
return True
@staticmethod
+86 -3
View File
@@ -4044,13 +4044,13 @@ class TestSendDryingCommand:
def test_start_caches_target_for_badge(self, mqtt_client):
"""mode=1 send populates _drying_targets so the badge can render it."""
mqtt_client.send_drying_command(ams_id=2, temp=65, duration=12, mode=1, filament="PETG")
assert mqtt_client._drying_targets[2] == {"filament": "PETG", "temp": 65}
assert mqtt_client._drying_targets[2] == {"filament": "PETG", "temp": 65, "duration_hours": 12}
def test_start_overwrites_prior_target_for_same_ams(self, mqtt_client):
"""A second start on the same AMS replaces the cached target."""
mqtt_client.send_drying_command(ams_id=0, temp=55, duration=4, mode=1, filament="PLA")
mqtt_client.send_drying_command(ams_id=0, temp=70, duration=6, mode=1, filament="ABS")
assert mqtt_client._drying_targets[0] == {"filament": "ABS", "temp": 70}
assert mqtt_client._drying_targets[0] == {"filament": "ABS", "temp": 70, "duration_hours": 6}
def test_stop_clears_target(self, mqtt_client):
"""mode=0 send drops the cache so the badge stops showing the target."""
@@ -4065,7 +4065,7 @@ class TestSendDryingCommand:
mqtt_client.send_drying_command(ams_id=128, temp=80, duration=6, mode=1, filament="PA-CF")
mqtt_client.send_drying_command(ams_id=0, temp=0, duration=0, mode=0)
assert 0 not in mqtt_client._drying_targets
assert mqtt_client._drying_targets[128] == {"filament": "PA-CF", "temp": 80}
assert mqtt_client._drying_targets[128] == {"filament": "PA-CF", "temp": 80, "duration_hours": 6}
class TestStartPrintAmsMapping:
@@ -6120,6 +6120,89 @@ class TestDryingCompleteCallback:
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 0, "tray": []}]})
assert mqtt_client._drying_events == [0]
def test_early_end_logs_firmware_reason_codes(self, mqtt_client, caplog):
"""#2770 — a 12-hour cycle the firmware abandoned 20 minutes in logged
only 'drying complete', so the report carried no evidence of why. An
early end now names the shortfall and the reason fields we already
parse: phase, sub-phase, cannot-dry codes and live HMS."""
from backend.app.services.bambu_mqtt import HMSError
mqtt_client.state.hms_errors = [
HMSError(code="0x2000003", attr=0x07008000, module=7, severity=2, full_code="0700800002000003")
]
mqtt_client._client = MagicMock()
mqtt_client.send_drying_command(ams_id=0, temp=65, duration=12, mode=1, filament="PETG")
mqtt_client._handle_ams_data(
{"ams": [{"id": "0", "dry_time": 700, "info": "10002123", "dry_sf_reason": [1], "tray": []}]}
)
with caplog.at_level(logging.INFO, logger="backend.app.services.bambu_mqtt"):
mqtt_client._handle_ams_data(
{"ams": [{"id": "0", "dry_time": 0, "info": "10002103", "dry_sf_reason": [1], "tray": []}]}
)
assert mqtt_client._drying_events == [0]
message = "\n".join(r.getMessage() for r in caplog.records)
assert "drying ended early" in message
# The shortfall, against the duration we asked the firmware for.
assert "700 of 720 minutes" in message
# dry_status 0 (Off) and dry_sub_status 0 from info hex 10002103.
assert "dry_status=0" in message
assert "dry_sub_status=0" in message
# InsufficientPower, and the AMS heater-fan HMS that goes with it.
assert "dry_sf_reason=[1]" in message
assert "0700800002000003" in message
def test_early_end_without_a_cached_target_still_logs(self, mqtt_client, caplog):
"""A cycle Bambuddy did not start — from the printer's screen, from
Studio, or from before a restart — has no cached duration to compare
against. The remaining time alone still proves it was cut short, so the
reason codes must be logged rather than withheld for lack of a target."""
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 480, "tray": []}]})
with caplog.at_level(logging.INFO, logger="backend.app.services.bambu_mqtt"):
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 0, "tray": []}]})
message = "\n".join(r.getMessage() for r in caplog.records)
assert "drying ended early" in message
assert "480 of ? minutes" in message
assert "hms=none" in message
def test_stop_we_sent_is_not_blamed_on_the_firmware(self, mqtt_client, caplog):
"""A stop Bambuddy sends — print takes priority, or the user's Stop
button — also ends the cycle far short of its duration, which on the
telemetry alone looks exactly like the firmware abandoning it. It must
be named as ours rather than reported as an unexplained early end."""
mqtt_client._client = MagicMock()
mqtt_client.send_drying_command(ams_id=0, temp=65, duration=12, mode=1, filament="PETG")
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 700, "tray": []}]})
mqtt_client.send_drying_command(ams_id=0, temp=0, duration=0, mode=0)
with caplog.at_level(logging.INFO, logger="backend.app.services.bambu_mqtt"):
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 0, "tray": []}]})
message = "\n".join(r.getMessage() for r in caplog.records)
assert "drying stopped by Bambuddy" in message
assert "ended early" not in message
# And the attribution is consumed, so a later firmware-ended cycle on
# the same unit is not credited to a stop we sent hours earlier.
mqtt_client.send_drying_command(ams_id=0, temp=65, duration=12, mode=1, filament="PETG")
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 700, "tray": []}]})
caplog.clear()
with caplog.at_level(logging.INFO, logger="backend.app.services.bambu_mqtt"):
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 0, "tray": []}]})
assert "drying ended early" in "\n".join(r.getMessage() for r in caplog.records)
def test_cycle_that_runs_to_term_keeps_the_plain_completion_log(self, mqtt_client, caplog):
"""The countdown of a cycle that finishes normally is all but exhausted
when it drops to 0. Nothing needs explaining, so it keeps the one-line
message it has always had — the early-end diagnostics must not become
noise on every completed dry."""
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 1, "tray": []}]})
with caplog.at_level(logging.INFO, logger="backend.app.services.bambu_mqtt"):
mqtt_client._handle_ams_data({"ams": [{"id": "0", "dry_time": 0, "tray": []}]})
message = "\n".join(r.getMessage() for r in caplog.records)
assert "drying complete (dry_time 1 → 0)" in message
assert "ended early" not in message
class TestPrintRunningObservedCallback:
"""#1485 follow-up: on_print_running_observed fires the FIRST time we