Commit Graph
932 Commits
Author SHA1 Message Date
maziggy f3b6a503bd fix(profiles): read the companion files that hold a preset's real gcode
A bundled preset can keep a setting in `<preset> template <key>.json`, a
file the preset itself does not reference -- the desktop slicer finds it
by name. Walking only `inherits` never reached it, so every one of the 56
instantiable BBL machine presets resolved `machine_start_gcode` to the
577-character generic block on fdm_machine_common instead of its own
6.5-21 KB one. That block holds the M620 AMS load and the M1002
gcode_claim_action calls, so a print sliced from it heats the bed, moves
the toolhead and extrudes nothing (bambuddy#2838).

Companions are now folded into each ancestor as the chain is walked, at
that ancestor's precedence, so a caller's own value still wins and the
0.2/0.6/0.8 variants reach their 0.4 sibling's companion. They are found
by listing rather than by a fixed set of keys.

Covered against the shipped bundle, not fixtures: a new e2e spec resolves
all 56 presets inside the image and fails on any that still lands on the
generic block.
2026-08-15 11:37:01 +02:00
maziggy aff737999f Build the slice output's path from a name a folder can have (#2832)
A print's display name comes from inside the 3MF, not from the filename,
so a MakerWorld title arrives with its punctuation: "Planter Pot with
Drip Tray, 12 cm / 5 inches". The slice-to-archive sink used it verbatim
for the output folder and the output file, and a slash in a folder name
is not a character -- it is another folder. mkdir(parents=True) created
the level it implied and the file's own join added a third that nobody
had made, so the slice failed with ENOENT on a path that half existed.
Renaming the print first was the only way through.

Reduce a display name to a single path component before it becomes one.
Characters a name cannot hold are replaced rather than dropped, so the
folder still reads like the model's title, and the set is the one the SD
card already rejects -- which covers a Windows install too, where the
colon in "Model: v2" fails the same way. The name shown in Bambuddy is
untouched: a title is allowed its punctuation, and refusing the slash
would reject the name this was reported about.

The joins are asserted to stay under the archive directory. That was
already claimed by a SEC-PATH-OK marker on both lines, citing a
sanitiser that is defined in another module and was never called here;
without the marker the path-join backstop flags them both. The claim is
now true, and a future edit that reaches around the reduction is caught
rather than trusted.

The library sink takes the same embedded name, so it gets the same
reduction: managed storage names the file after a UUID and never saw
this, but an external folder writes the name as given.

Display names are also stripped of control characters on the way into
the database, in the schema and in the archive service. The validator
hands back anything that is not a string rather than iterating it, so
the field still answers a list or a bare int with a 422 instead of
accepting the one and failing on the other.

Display names are also stripped of control characters on the way into
the database, in the schema and in the archive service, cleaned before
the filename fallback rather than after it so a whitespace-only embedded
name still falls through to the filename. The validator hands back
anything that is not a string rather than iterating it, so the field
still answers a list or a bare int with a 422 instead of accepting the
one and failing on the other.
2026-08-14 16:02:53 +02:00
maziggy 0623cc46df Repair no-3MF archives' photos and their silent filament writes (#1820)
Two faults behind the same kind of print: one that arrives without a
retrievable 3MF, which on an H2S is any job started from the printer's
own internal library.

Such an archive has no file_path, and Path("").parent is Path("."), so
every site that derived the archive's folder from it landed on the data
directory itself. The finish-photo capture spotted that and wrote to
<archive_dir>/<id>/photos instead. Nothing else did. The photo was
written in one place and looked for in another: reads 404'd, deletes
dropped the name and left the file, and the notification attachment
never found the image. Hand-uploaded photos worked only because upload
and read agreed with each other rather than with the capture. Give the
question one owner in utils/archive_paths and have all four sites ask
it. Lookups check the old shared location too, so photos already
uploaded there stay reachable; uploads now go where captures go.

Separately, the remain%-delta fallback that stands in for a missing 3MF
can charge nothing for several reasons, and did so without a word. The
AMS reading is coarse and, on the reporter's printer, noisy: it rises
mid-print, swings five points over a job, sits at 100% through a
36-minute print on a fresh spool, and goes negative on a nearly empty
one -- which the start-of-print gate rejects, dropping the only slot
that was printing. Two of their prints went uncounted for two different
reasons and both read as "no spools updated", which is also what a print
with nothing to charge prints. Name the slot and the two readings in
each case, on the Spoolman path and on the internal-inventory path,
which has carried the same gates since #1119.

The Spoolman path also had no notion of which slots the print used, so a
spool swapped into an idle slot mid-print reads as consumption and is
billed to whoever that slot is assigned to -- the fault #1269 fixed for
the internal tracker, still open here, and likeliest on exactly the
prints this fallback serves, where nothing else narrows the field. Use
the same three pieces of evidence it does: the print's mapping, its
mid-print tray changes, and the tray it started on. The last needs
storing, because the internal tracker's row is deleted before this runs
and a screen-started print has no mapping to fall back on -- hence a new
nullable column, and no backfill, since a row from before it existed has
nothing to say. Where no evidence exists at all, every slot is still
considered.

Both paths also treated tray_now == 255 as naming a slot. It does not:
it is the field's initial value, the fallback for an unparseable
reading, and what it reports with nothing loaded. Mapped as a tray id it
becomes (255, 1), so as the only evidence it excluded every real slot
and charged nothing at all -- this issue's own bug, arriving by a new
route. On the internal path that is live today; on the Spoolman path it
would have shipped with the guard above. The external holder reports 254
when it is genuinely in use.

The arithmetic is untouched: at one percent per step this cannot resolve
a small print, and pretending otherwise would be worse than saying so.
2026-08-14 15:32:40 +02:00
maziggy a9624d3887 Match a print completion to its queue job the way the printer names it (#2829)
Bambuddy has no run identifier to tie a completion to a queue row, so it
finds the row by printer and status='printing' alone. b5a34b7ba added a
check that the completion's subtask name agrees with the file the row was
dispatched with, so the printer's own calibration runs cannot close
someone's job early. It compared the two names verbatim.

The printer does not echo them verbatim. It substitutes underscores for
spaces, so 'H2D_Carbon_Filter_(V2)_Body & Solid Lid' came back as
'H2D_Carbon_Filter_(V2)_Body_&_Solid_Lid', the check refused it, and the
row stayed printing. check_queue counts every printing row as a busy
printer and nothing else ever closes one, so the printer's queue stopped
until someone cancelled by hand. It also truncates long names and marks
the cut with '...', which would have done the same to any long title.

Compare on the canonical form instead -- case, spaces and underscores --
which is the rule the 3MF lookup in this module has always used for the
same names, and treat a truncation marker as a prefix match. The check
keeps its purpose: the same printer the same day correctly refused a
completion for auto_pa_line_calib_mode.

One comparison being stricter than reality should not be able to stop a
queue indefinitely, so the scheduler now closes a row itself when it has
been printing for five minutes after its printer went terminal, with the
status that state implies. A real completion arrives within seconds, so
this only sees rows that were already stranded, and a disconnected
printer never qualifies. It restores the queue only -- notifications,
billing and auto-off are not replayed minutes late.
2026-08-14 14:26:30 +02:00
maziggy fffa68ec55 Explain a print that never reached the printer's card, instead of sweeping for it (#2780)
Bambuddy reads a print's 3MF, cover and timelapse over FTPS on port 990,
which on every Bambu model serves external storage only. Under some
configurations H2-series and P2S firmware keeps the sliced file on
internal storage, where Bambu Studio put it over the port-6000 service,
and then no path on 990 can find it.

The print command has always said which of the two it used -- `url` reads
ftp://<name> or brtc://emmc/<name>. We discarded it and swept anyway:
~110 connections per print, all certain to fail, ending in an archive
card with nothing on it and no stated reason. In the reporter's bundle
all 35 dispatches to their H2C and P2S said internal storage, all 25 to
their X1C said external, and all 44 empty cards belonged to the first two.

Read the field, skip the sweep when it cannot succeed, and record which
reason applied. A printer that uses the card is unaffected, and so is one
we have no answer for -- silence is not evidence, and reading it as bad
news would break archives that work today.

The answer is held per print and dropped when that print ends, rather
than kept as a standing fact about the printer. Plenty of prints never
announce themselves: 14 of the 79 print starts in that bundle arrived
with nothing on the request topic, started from the printer's own screen
or picked up after a restart. Left standing, one slicer print to internal
storage would suppress the lookup for every screen-started print after
it, on a printer whose files really are on the card. The sticky reading
is kept for the connection diagnostic alone, which is run after the print
that prompted it and would otherwise have nothing to report.

Two things that pointed the wrong way go with it. The archives banner
told everyone to enable "Store sent files on external storage"; the
reporter had it on for the whole three weeks and it would not have
helped. The diagnostic passed a printer whose slot was empty, because it
read only the toggle -- an empty slot is now a failure naming the slot,
and a printer that has storage and still used its own is a warning. On
P1-series that empty-slot failure yields to the existing unsupported-model
skip: the toggle cannot be switched on there at all, so telling the
operator to insert a card would promise a fix inserting a card does not
deliver (#2524).

Also close FTP sockets on the failure paths, which dropped them for the
garbage collector -- 1813 in a day in that bundle -- and drop the advice
to restart the printer, which the reporter tried twice while a single
manual connection to the same printer handshook cleanly.

This does not make the affected prints archive in full; that needs the
port-6000 protocol tracked in #2762.
2026-08-14 13:40:12 +02:00
maziggy 3954d3a7e6 Choose which rack nozzle each filament prints from on an H2C (#1784)
The Vortek rack holds six hotends, and a multi-colour plate is sliced to
use a different one per colour so it can skip the purge. Which of the six
each colour takes is not in the 3MF. The same plate, sliced and sent twice
from Bambu Studio with a different choice each time, produces two files
that differ only in rounding in the last digit of a few extrusion figures
-- the filament grouping, the toolchange stream, the 120 nozzle-change
markers and project_settings.config are all identical. The choice travels
only in the dispatched nozzle_mapping.

Bambuddy had no way to state it, so those plates went out with no nozzle
assignment at all and the printer chose for itself. That is what levelled
on one hotend and printed with another, millimetres above the plate.

Every rack-bound filament now carries a position picker beside its AMS
slot dropdown, listing all six with the nozzle each holds. An empty
position, or one holding the wrong diameter or flow type, is shown greyed
out with the reason rather than hidden, so someone looking for position 4
finds it. The choice is per filament *group* rather than per slot, because
a group is one hotend: two filaments the slicer grouped together share it
and cannot point at different positions.

Nothing has to be picked. Positions are assigned automatically, preferring
one already loaded with that colour, which on the plate this was built
against reproduces Bambu Studio's own pick exactly.

A nozzle currently picked up onto the carriage is offered too. The
firmware drops its rack position from the report entirely rather than
sending a placeholder (#943), and refusing it would rule out the position
most likely to be wanted -- the one the last print left mounted. Only
recoverable when exactly one position is missing; two gaps are genuinely
ambiguous and stay unavailable.

Positions are re-checked at dispatch, not just when queued, because the
rack can be re-loaded in between. The two failure modes differ on purpose:
an explicitly chosen position that no longer fits stops the print, names
what the position now holds, and deletes the uploaded file from the SD
card so it cannot be started by hand either -- an operator who named a
hotend must not silently get a different one. An automatic assignment that
cannot be made instead falls back to letting the firmware choose, which is
what happened before any of this existed.

The pick is stored as {group: position} rather than as the expanded
nozzle_mapping, though that is what goes on the wire. That column means
"Bambu Studio decided, forward verbatim", and only the group-and-position
form can be re-checked against what is actually mounted at dispatch.

The existing multi-rack refusal in extract_nozzle_mapping_from_3mf stays.
It still guards the #2800 fallback, which can only ever name one rack id.

Measured on the maintainer's H2C: rack position n is physical nozzle id
15 + n, confirmed by cross-referencing two captured dispatches against
Bambu Studio's own dialog. extruder_max_nozzle_count names which carriage
is the rack straight from the file, and is read rather than assumed -- a
fourth independent confirmation of the carriage indices fixed in 45dc139.

The print dialog is also wider, on every printer. Its filament rows carry
the most horizontal content in it and adding a picker truncated names to
"Bamb...". The column widths themselves only change on a rack machine.

Tests: 44 unit covering the plan, the resolver, the mounted-nozzle
recovery and every refusal; 9 dispatch integration asserting the two real
captures end to end; 7 API round-trip; 33 frontend. The API ones exist
because two integration bugs got through a green suite that tested the
pieces and not the seams -- the group data reached only one of the three
filament-requirements routes, and the field was declared on every schema
except the create one, where Pydantic dropped it in silence.
2026-08-14 11:27:50 +02:00
maziggy 45dc139c41 Print an H2C two-nozzle plate from the carriage it was levelled on
The two carriages were the wrong way round: extruder index 0 was treated as
the fixed hotend and index 1 as the swappable rack, and it is the other way
about. A plate using both was levelled with one nozzle and printed with the
other, several millimetres off the plate.

Three sources agree, and disagreed with the code. Telemetry reports
ams_extruder_map {'0': 1, '1': 0, '2': 0}. BambuStudio, dispatching a plate
that used all three of those AMS units, sent the filament from the unit on
extruder 1 to physical nozzle 1 and the ones on extruder 0 to rack positions
16 and 18, and that print completed. And the two constants could not both
have been right: _FIXED_NOZZLE_ID is 1 while the fixed extruder was 0, in a
scheme where physical nozzle id N sits on extruder N.

The old value came from the #2800 hardware A/B, where [17, -1, -1, 1] printed
in mid-air and [1, -1, -1, 17] printed correctly. That result stands -- it
established which wire worked. The extruder indices were not measured by it;
they were inferred by pairing the working wire with a slot_extruders list
produced by the 3MF reader that has since turned out to mis-read exactly
these files. The reasoning is recorded at the constants so a future
regression report is not re-litigated from scratch.

Also withholds nozzle_mapping entirely when more than one filament group
needs the rack. Two groups on one extruder means that extruder is a rack and
the plate wants a different hotend per group -- which physical slot each
takes is the slicer's choice against the live rack and is stated nowhere in
the file, since both groups can carry identical diameter and nozzle type.
Studio dispatched such a plate to 16 and 18; nothing here can reproduce that,
and answering anyway is what printed in mid-air, so the firmware picks.

Restores the fixed 32-entry padding that dfeac792f replaced with the plate's
slot count. That was derived from a single 3-entry capture which turned out
to be a calibration job; Studio's dispatch of a real project print on the
same machine is 32 entries.

Both constants are read in one function, on the nozzle-rack path, so no other
model is affected. Verified on hardware: the plate that printed in mid-air
now prints.
2026-08-14 09:18:55 +02:00
maziggy d1c65a6659 Report a print stage we cannot name at the default log level
STAGE_NAMES is hand-maintained and every new model adds to it, so a printer
occasionally reports a number that is not in it and the card reads "Unknown
stage (72)" -- which an H2C did, where the table runs to 66 and then jumps
to 74. Stage transitions were logged only at DEBUG, off in normal running,
so the sole record that it had happened was the card itself, and by the
time anyone looked the printer had moved on.

The asymmetry is the point: a stage we can name is worth DEBUG, and the one
we cannot is the interesting one. An unnamed stage is now logged at INFO,
once per stage number per session, with the model, the stage it came from
and the print state at the time -- which is what naming it afterwards
needs. Named stages are unchanged, so a normal print logs nothing new. -1
is excluded: it is Bambuddy's own "not in a stage" sentinel and the field's
initial value, so every print would otherwise report it on the way out of
its last real stage.

Fixes a latent crash found while testing this. The stage-change log line
builds its text before the log level is consulted, so get_stage_name runs
on every transition whatever the level is set to; a stg_cur that was not
hashable -- malformed telemetry rather than an unknown stage -- raised
TypeError out of STAGE_NAMES.get and aborted the whole state update.
Labelling a value can no longer do that.
2026-08-13 17:23:41 +02:00
maziggy dfeac792fb Stop an H2C refusing a multi-colour print as a hotend mismatch
The print uploaded, the printer took the command, and stopped at once with
HMS 0500-4047 -- "the available hotend quantity or model does not match the
sliced file". nozzle_mapping told the printer one of the plate's filaments
went to no hotend while ams_mapping named the tray it comes from, and the
firmware will not start a job on that contradiction.

Each filament in a 3MF names the group it belongs to, and on every other
dual-nozzle Bambu the group number is also the extruder index, so it was
read as one. On a rack machine it is not: the rack carriage holds six
hotends to the fixed carriage's one, so the slicer writes a group per
nozzle rather than per carriage. The failing plate carried groups 0, 1 and
2 against a two-entry physical_extruder_map, and the filament in group 2
was dropped -- indistinguishable downstream from a slot the plate does not
print, which is what reached the wire as -1.

extract_nozzle_mapping_from_3mf now resolves the group through the table
the file states for itself, the <nozzle id extruder_id> elements in
slice_info.config. Files carrying no such table keep the direct index, so
H2D slices are unaffected. A filament that still cannot be placed drops
the whole mapping with a logged reason instead of half an answer: the
firmware then picks its own nozzle, which is the pre-existing behaviour
and far better than an answer that contradicts itself.

Two related faults fixed in the same pass. The mapping was read across
every plate in the file, so on a multi-plate project a slot took its
extruder from whichever plate came last; it is now scoped to the plate
being dispatched, in extract_filament_requirements as well. And the array
is now one entry per filament slot, matching BambuStudio's own dispatch of
[1, 16, 16] for a three-filament plate, rather than padded to a fixed 32.

Verified against the file that failed: slot extruders [-1, 1, 0] became
[0, 1, 0], and the wire [-1, 16, 1, -1 x29] became [1, 16, 1].
2026-08-13 17:05:08 +02:00
maziggy e6842e1d3c Rank near-colour matches by how they look, and share one filament type table (#2804)
Three follow-ups to #2804, all bearing on one decision: which spool a print
uses when the exact colour is not loaded.

Colour ranking is now perceptual. The ranking added in #2804 measured RGB
distance, which rates a colour by how far apart the numbers are rather than
how far apart they look, and it overweights blue badly enough to invert the
answer: against a required #1E4821 green, a purple #38202F is the nearer of
two eligible spools by RGB and four times the further once measured properly.
Both sides now use CIEDE2000 -- perceptual_color_distance in
backend/app/utils/color_utils.py and colorDistance in amsHelpers.ts, kept
structurally identical so they can be read side by side. Verified against the
Sharma/Wu/Dalal published reference set, all 31 pairs to 1e-4, and the two
implementations agree to within 1e-9 across 800 sampled pairs. Eligibility is
untouched, still the per-channel RGB box, so this only reorders spools that
already qualified.

Type matching now agrees between the interface and the scheduler. Bambu
firmware treats PA-CF, PA12-CF and PAHT-CF as one material and the scheduler
has always matched them accordingly, but the interface compared raw type
strings and called that same pairing a mismatch. The badge contradicted what
the printer was about to do, and the manual override picker, which groups by
canonical type, offered the very spool the badge then rejected. The fifteen
comparison sites in useFilamentMapping.ts, useMultiPrinterFilamentMapping.ts
and PrinterSelector.tsx now call filamentTypesCompatible.

The pipeline pre-flight reads the matcher's table instead of its own copy.
That copy had drifted into disagreeing in both directions: it aliased PLA
Basic to PLA where the matcher never has, so a run could clear the check and
then fail to map its slots, and it lacked the nylon grouping, so it flagged
runs the matcher handles without complaint. A check whose job is to predict
dispatch is wrong whenever it disagrees with dispatch, whichever way it leans,
so it and the scheduler now both read backend/app/utils/filament_types.py.

That canonicaliser deliberately does not strip surrounding whitespace. It
looks like a free improvement, but it would collapse a junk tray_type to ""
just as a 3MF declaring no filament type yields "", and a typeless requirement
would start matching a junk-typed tray instead of reporting the slot unmapped.
Padded type strings are worth handling on their own terms, with that case
addressed.

One behaviour change outside the ranking: the pre-flight is stricter for a
printer reporting a product name such as "PLA Basic" where the generic
material belongs, which it now flags rather than passes. Rare in practice,
since the printer reports material and product name in separate fields, and it
is the answer the matcher would give. Nothing about which spool a print
actually uses changed outside the colour ranking itself.

Adds 203 backend and 6 frontend tests. The #2804 tie-break test now uses
identical colours: two colours at equal RGB distance are not perceptually
tied, which is rather the point.
2026-08-13 12:04:56 +02:00
Martin Grolmus 4f7a02b393 Pick the nearest eligible filament colour instead of the first one in tray order (#2804) (#2823) 2026-08-13 11:19:00 +02:00
maziggy 02616f0c91 fix(queue): stop a library-file delete from destroying the jobs queued against it (#2819)
Nothing tied a library file to the queue rows pointing at it, and the FK
that describes the relationship is ON DELETE CASCADE -- which SQLite does
not enforce and PostgreSQL does. So the same fault had two faces: rows
left pointing at a file that no longer existed, failing at the printer
with "Library file not found" days later, or rows deleted outright with
no error and no history.

Two routes into it, both fixed by taking the queue off the file before
the row goes.

Dispatch (the reported case): quantity>1 on the printer-card
upload-and-print flow puts cleanup_library_after_dispatch on every copy,
and _clone_queue_item copies library_file_id onto batch clones, so the
first dispatch consumed the file the rest were waiting on. The copies are
now pointed at the archive that dispatch just created -- it holds its own
copy of the 3MF -- and the consume flag is cleared on them. A copy already
printing from its own archive keeps it, a finished one keeps its outcome,
and a cross-model item (#671) keeps any candidate this does not consume.

Deletion: the File Manager, bulk delete, folder delete, emptying the trash
and the retention sweeper all removed rows with queued work against them.
Folder delete did not even clear the cross-model candidates, because the
file-id walk it already performs threw its result away. Jobs waiting on a
deleted file are now cancelled at that moment, naming the file, and every
other row referring to it is detached rather than destroyed -- print
history and batch progress are counted from those rows. A job that is
printing is left alone: what is deleted is the library copy, not the copy
on the machine. The trash is reversible so it still changes nothing about
the queue, and a job dispatched while its file is in the trash now says so
instead of "not found".

Verified row for row on PostgreSQL 16 as well as SQLite: without this,
PostgreSQL deletes every queue row referencing the file.
2026-08-13 10:26:59 +02:00
maziggy 454457a0af Attribute filament correctly when AMS backup swaps spools mid-print
Everything the completion path needs to split a print's filament across
the trays it fed from lived only in memory: the dispatched plate and
slot-to-tray mapping, the spool-assignment snapshot, and the tray-change
log. A print that outlived a restart lost all of it and fell back to
what the printer reports at completion -- which, with AMS Filament
Backup on, is the substitute tray. The whole print was charged to the
spool that only finished it while the spool that ran dry was charged
nothing.

Persist that context in a new active_print_sessions row, append tray
changes as they happen, and restore both the session and the printer's
tray-change log at restart recovery. Seed the log from the current tray
when there is nothing to restore, since last_loaded_tray advances even
when no change is logged.

Rank the queue item's stored ams_mapping above the printer's live
mapping field, which is what backup rewrites. Recover plate_id from the
archive or queue item, and give extract_layer_filament_usage_from_3mf a
plate_id instead of taking the first .gcode member -- a Bambu Studio
export stores plate 2 first, so per-layer figures were measured against
the wrong plate for both inventory backends.

Stop auto-unlinking a spool assignment when its slot reports empty
during a running print. At a runout the spool is still in the AMS, and
dropping the link leaves the completion path nothing to charge.

Capture the print-start context for both inventory backends. Spoolman's
own durable row (#1820) carries its plate-scoped figures and dispatched
mapping but not the tray-change log, and its slot assignments -- the
way. Registration in _active_sessions stays gated, since on_ams_change
reads it to decide whether to skip the remain%-based weight sync (#880).
2026-08-13 08:37:16 +02:00
maziggy df5aa04df1 Pool AMS backup spools in the print dialog's filament check
The dialog weighed each plate against the spool in the slot it mapped to
and knew nothing about AMS Filament Backup, so a two-plate job needing
1441 g of ABS was refused against a 1000 g spool while the identical full
spool in the next slot went uncounted. The dispatcher has pooled matching
spools since #1762 and would have run the print -- "Print anyway" was
always the right answer to this warning.

The rule for which spools back each other up now lives in one place:
build_slot_materials() in filament_deficit, which the dispatcher's pool
and the new slot_materials half of GET /printers/{id}/inventory-remain
both draw on. The dialog groups on the keys it is handed rather than
resolving spools a second time, which is what let the two answers drift
apart, and which also gives the check to Spoolman users -- it read the
internal inventory only, so in Spoolman mode it approved everything.

Where a pool really is short the warning quotes the pooled totals, since
the per-slot figure reads as a contradiction next to a full peer spool.
2026-08-11 11:32:17 +02:00
maziggy e9d9e51a81 fix(h2c): send physical nozzle IDs for both carriages, not extruder indices (#2800)
The first pass at the H2C rack mapping had both of its hardware-derived
values wrong, and the reporter's follow-up A/B on real hardware settled
them.

The rack does not feed extruder 0. A mixed-nozzle plate extracted as
slots=[0, -1, -1, 1] with rack position 17 dispatched as
[17, -1, -1, 1], and the rack nozzle printed several millimetres above
the bed. The rack is extruder 1.

Correcting only that is not enough. The fixed hotend answers to physical
ID 1, not to its extruder index of 0, and forwarding the index produced
[0, -1, -1, 17] -- a command the printer rejected outright rather than
mis-printing. Translating both carriages gives [1, -1, -1, 17], and the
same sliced file then cleaned, levelled and printed on the correct
nozzle at the correct Z through to completion.

Both values agree with three native Bambu Studio captures from the same
machine, which carry [1, 17, ...] and [17, 1, ...] depending on filament
slot order.

An extruder index naming neither carriage now omits the field instead of
reaching the wire as a physical ID that identifies no nozzle. The
fixed-hotend-only path is unchanged but no longer a guess: Bambu Studio
sends no nozzle_mapping at all for such a plate, which is what Bambuddy
already did.

Reported, diagnosed and hardware-verified by @tru3l3gend, who ran the
mixed-nozzle A/B on both nozzles and captured what Bambu Studio sends
for fixed-only and mixed plates.
2026-08-11 09:37:03 +02:00
maziggy d120a804ff fix(slicer): name the sidecar service in the update command (#2802)
The "update your sidecar image" advice told users to run a bare
`docker compose pull`. bambu-studio-api is declared with
`profiles: [bambu]`, and compose skips profile-gated services silently,
so the pull was a no-op for exactly the users the message was written
for -- and `restart: unless-stopped` kept the old container serving.
The reporter pulled, restarted, set MAX_MODEL_UPLOAD_MB and got the same
100 MB rejection, because the image never changed.

Name the service in both commands instead. Naming enables the profile
implicitly, for pull and up alike. `--profile bambu` would also work but
downloads the 220 MB Bambu image on an OrcaSlicer-only host and then
starts a sidecar the user never asked for.

Same correction in the sidecar README, the compose header and the
changelog entry, which all carried the bare form.
2026-08-11 08:35:23 +02:00
maziggy 15ea11e10b Recover the preview slice from custom G-code the sidecar cannot parse
Opening the slice dialog on an unsliced project runs a preview slice purely
to ask the slicer which AMS slots the chosen plate consumes. Bambu Studio 2.8
writes {if timelapse_inline_photo} into the machine's time_lapse_gcode but
exports no definition for that variable, so the template is unresolvable the
moment it leaves Studio: an older sidecar stops with a placeholder parse error
before producing any slice_info. The preview returned nothing and the caller
fell back to guessing from painted faces, silently. On the H2D project this
was found with, the guess dropped the support material -- a whole slot off a
four-filament plate.

Retry the preview once with just the named template emptied, still on the
file's own settings. Keeping the embedded settings is what keeps the answer
honest: overriding the process preset instead discards the project's support
configuration, which loses that slot and moves used_g by up to 2x. Measured
against the same file: retry reproduces all four slots gram for gram, a
printer+process override returns three.

Only templates that cannot extrude are eligible -- a start or filament-change
template lays a prime line or purges, so emptying one would move the very
grams the preview reports, and returning nothing beats a confident wrong
number. Verified on a working H2D slice that emptying time_lapse_gcode leaves
every used_g/used_m in slice_info identical.

Match on a normalised option name: the slicer reports timelapse_gcode while
the 3MF stores time_lapse_gcode, so a literal comparison finds nothing.

Decide whether a retry applies before logging, so a slice that recovers does
not announce itself at WARNING twenty seconds before it succeeds.
2026-08-10 15:44:55 +02:00
maziggy b03603c4e2 Release the keep-warm bed on the dispatch paths that skip the rollback (#2727)
Selecting an item hands any keep-warm hold on its printer over to the preheat
pin: `_sweep_keep_warm` drops the `_keep_warm` entry and records "bed" instead,
on the promise that `_dispatch_one` will unwind it on any non-success exit. Two
of that function's exits never reached the `finally` that keeps the promise --
the claim failure returns before the `try` opens, and the vanished-row return
sat inside it but left `item_printer_id` at None, which the rollback guards on.

Either one left the bed hot with nothing tracking it. The keep-warm entry was
already gone, so the max-duration cap no longer applied and `_release_keep_warm`
had nothing to act on; if the cancelled item was that printer's last pending
one, the printer also dropped out of the candidate set, and nothing would ever
switch the bed off. Reachable whenever a cancel or delete lands between
selection and the claim -- narrow, but the outcome is exactly what the cap was
added to prevent.

`_dispatch_one` now takes the printer it was selected for. `_launch_uploads`
already had it (it stores the same value in `_inflight`), so nothing new is
plumbed, and the parameter is optional so the tests that call `_dispatch_one`
directly keep their existing behaviour.

Two pieces of hardening found while tracing that:

  * The rollback switched the bed off unconditionally, where the keep-warm
    release deliberately checks first that firmware still reports the target it
    set. It now records what it pinned and declines when someone else owns the
    bed. Every uncertain case still switches off -- no recorded target, or a
    status that cannot be read -- because a bed left hot with no owner is the
    worse failure, and this runs in a `finally` where raising would mask the
    real exception. That is why the status read is factored out into a total
    helper returning None for "no evidence" rather than 0.

  * `_apply_keep_warm` ran unguarded between selection and `_launch_uploads`, so
    anything raising there discarded the tick's selections, computed AMS
    mappings included, and on a persistent fault stopped the queue dispatching
    altogether. Wrapped, for the same reason the deficit check is: an auxiliary
    comfort feature must never wedge dispatch.

Also documents why the max-duration check sits behind the FINISH and client
guards rather than ahead of them, since the ordering looks like a hole and is
not: with no client there is no M140 to send and the elapsed check fires on the
first tick after the printer returns, and leaving FINISH means the plate was
cleared, which routes the printer to `_release_keep_warm` instead. The invariant
to preserve if that is ever reordered is that every path out of an engaged hold
ends in a bed-off.

Six tests: the handover recording its target, both early returns releasing, a
call with no printer id staying a no-op, the reassigned-bed skip, the matching
and unreadable cases switching off, and eviction on deregistration.
2026-08-10 13:54:44 +02:00
MartinNYHC 37b0e25a2b Merge branch 'dev' into feature/queue-keep-warm-chamber-history 2026-08-10 13:31:28 +02:00
maziggy a849759469 Explain an oversized model instead of reporting a slicer crash (#2802)
The sidecar caps model uploads and reports a rejection as a bare
HTTP 500 "File too large" -- multer's MulterError is not the sidecar's
AppError, so its handler falls through to the default status. A 500
reads as a crash inside the slicer, and the one message Bambuddy had
about request size was written for the 413 a reverse proxy sends, so
it never appeared. The reporter tried MAX_FILE_SIZE, BODY_PARSER_LIMIT
and EXPRESS_PAYLOAD_LIMIT, stopped nginx, and moved from Windows to
Docker -- none of which the sidecar reads.

Match the rejection by what it says rather than by its status, so an
installation still on an older sidecar image gets the same explanation.
The 500 match is strict -- the body must be only multer's message --
because a genuine CLI failure is also a 500 and has to keep reaching
the embedded-settings fallback. Old images are told to update, since
they have no setting to change; current ones are told which one to set.

Raising SlicerInputError rather than SlicerApiServerError is also what
skips the fallback retry, which had been re-uploading the identical
oversized file after a second 25-second 3MF conversion.

Log the model size on every slice. Nothing recorded it, so a support
package from a slice that died on an upload cap looked exactly like one
that died on a bad profile, and this had to be sized by hand.

Fall back to the exception class name when a transport error stringifies
empty -- three lines of the reporter's log read "Slicer sidecar
unreachable: " and stopped there.

Needs a sidecar image update to take full effect; MAX_MODEL_UPLOAD_MB is
documented in slicer-api/.env.example.
2026-08-10 12:18:29 +02:00
maziggy 71709c16aa Stop tearing down AMS drying for a print that cannot start (#2801)
A printer in FINISH with an unacknowledged plate and something pending in
its queue stopped and restarted drying once per scheduler tick, for as
long as the plate stayed unacknowledged. The reporter's Home Assistant
history recorded about 2000 state changes over ten days. No cycle ever
ran long enough to remove moisture, and cycles the user had started by
hand on other AMS units of the same printer were torn down with it.

Two concerns had become tangled. Plate-clear answers "is the bed ready
for the next job" and says nothing about whether the AMS may heat. The
gap between a finished print and the acknowledgment is when drying is
most useful -- the printer is free and nobody is waiting on it -- and
leaving the plate unacknowledged is also how people hold the queue by
hand, so the hold was costing them the drying it should have enabled.

Four faults, all in print_scheduler.

The "print takes priority" stop sat inside the not-idle branch. Drying is
not one of the things _is_printer_idle looks at, so stopping a cycle can
never turn a non-idle printer into an idle one: the stop was futile every
time it fired, and never fired on the dispatches where it was supposed to
mean something. It now runs when the printer is actually dispatchable,
and only where the model cannot dry through a print -- #2758 settled that
capable hardware should keep its cycle.

mid_print was inferred from busy_printers, which means "the queue could
not dispatch here this pass", not "is printing". A plate-held printer was
therefore treated as printing: the mid-print spool-protection cap
silently lowered its drying temperature, the cycle was logged as
(mid-print) in FINISH, and it bypassed the very gate meant to hold it.
busy_printers keeps its dispatch role; auto-drying now gets a narrow set
snapshotted before the item loop -- running, held post-dispatch, or
mid-upload -- and mid_print comes from the printer's own state. The
interlock comment at the seed already documented this hazard and worked
around it by staying out of the set; this generalises that instead of
adding a third special case. The other call site was already passing the
narrow set, so the wide one was the inconsistency.

_stop_drying sent a stop to every AMS reporting dry_time > 0. One
auto-dried unit was enough to kill a manual cycle on a different unit of
the same printer, contradicting the contract _sync_drying_state already
documents: the entry gate only knows about cycles Bambuddy began, so the
action must not reach past them. Consequence worth stating -- after a
restart Bambuddy cannot prove a running cycle is its own, so it leaves it
alone rather than risk stopping somebody's manual dry.

Fourth, and the reason #2770's guard did not catch this: a reading at or
below the threshold popped the unit's whole entry, ended_at included, so
the 30-minute re-arm cooldown went with it. An AMS reads higher warm than
cool, which is #2770's own finding, so a unit a point or two above the
threshold dipped below it as it cooled, wiped its history, and re-armed
immediately. Lifting a suspension now clears the judgement and keeps the
clock.

queue_drying_block changes behaviour as a result. It previously had no
effect on dispatch at all -- both branches skipped anyway, and it only
decided whether drying was needlessly killed. With the stop on the
dispatch path it now does what it says: a queued print waits for a
running cycle. Off by default.

Reported by @superflyer11, who traced both defects to the line and
brought ten days of external sensor history to date the cadence.
2026-08-10 09:31:25 +02:00
maziggy ec26cba927 Resolve the H2C rack nozzle at dispatch instead of letting firmware pick (#2800)
An H2C ran its startup clean and bed levelling on one hotend, switched,
and then printed several millimetres above the plate. The same job from
Bambu Studio was fine.

The H2C is the only model that mounts its nozzle from a rack of six, and
a print command names that nozzle by physical rack position -- the
firmware reports those as 16 to 21 -- not by the extruder index, 0 or 1,
every other dual-nozzle printer uses. Bambuddy only ever had a rack
position when a job arrived through the Virtual Printer, which captures
Bambu Studio's pick and replays it (#1780). Anything queued from the
library, an archive, the webhook or a slicer pipeline carried none, so
the field was omitted and the firmware chose -- and its choice need not
match what the file was sliced for.

The scheduler now derives the per-slot extruder assignment from the file
it is about to send, and the MQTT layer resolves it against the rack
position the printer reports live. Both are needed: the file knows which
side a slot prints from, only the printer knows which hotend is in the
carriage, and it can be swapped from the touchscreen between queueing a
job and printing it.

Derived at dispatch rather than at creation because that is the first
point knowing both the real printer and the real file -- an item can be
created unassigned, reassigned later, or have its file swapped for a
G-code-injected copy. One call therefore covers the print dialog, bulk
library adds, the webhook and pipeline runs, and no column is needed.

extract_nozzle_mapping_from_3mf is deliberately untouched. Its output
feeds the AMS matcher, where nozzle_id is compared against a tray's
extruder_id as a hard filter, and physical_extruder_map is what makes
that comparison correct -- on an H2D it is [1, 0] and flips the two.
Dropping the translation to suit the rack would send every dual-nozzle
AMS match to the wrong extruder. The dense per-slot form is a separate
function reusing the same output.

Nothing here can fail a dispatch. The command is built and published with
no exception handler above it, and the queue item is already committed as
printing by then, so a bad input has to degrade to "firmware picks"
rather than wedge the item. resolve_rack_nozzle_mapping validates every
input and raises nothing; an unresolvable mapping, an unparseable value
or an unknown rack position all omit the field, which is the behaviour
that existed before. Slot IDs are bounded before the dense list is built:
they come from the file, and one declaring filament id="50000000" would
otherwise allocate a fifty-million-entry list on the dispatch path.

Two things are not guessed. A job printing only from the fixed hotend is
still left to the firmware, because that nozzle's physical ID is not
confirmed by a known-good capture. And the rack is taken to feed extruder
0 from a single hardware observation -- if that is flipped, a one-sided
job matches nothing and falls back to the old behaviour, so only a job
using both nozzles at once could be harmed, which is what a second
capture needs to confirm.

Confined to the H2C throughout. Building the print command for 21 model
spellings with and without the new argument changes exactly three of them
-- H2C, O1C and O1C2. The other 18, including H2D and X2D, are identical.

Reported by @tru3l3gend, who diagnosed it on real hardware against a
working Bambu Studio dispatch, established the rack ID range and supplied
a patch.
2026-08-10 08:54:44 +02:00
Thomas Scott Williams f86f5a7c34 feat(queue): keep the chamber warm between prints and skip redundant soak
Back-to-back prints in chamber-heated materials (ASA, ABS, PA, PC) each paid
a full heat-soak from cold, even when the print that just finished had left
the chamber at temperature. Two changes remove that cost.

Keep bed warm between prints
  While a printer sits in FINISH awaiting plate-clear and the next queued item
  needs chamber heat, hold the bed hot so the chamber does not cool during the
  bed-clearing window. The bed is the chamber's heating element here, not a
  print surface, so the hold runs at the new `queue_keep_warm_bed_temp`
  (default 90C, which also satisfies bed-threshold-linked aftermarket chamber
  heaters), raised to the item's own bed temperature when that is higher.

  Gated on `queue_keep_bed_warm` AND `require_plate_clear` AND
  `preheat_enabled`, all re-checked in the backend so a stale UI cannot leave
  the feature running. `queue_keep_warm_max_minutes` (default 120) bounds the
  hold: when it elapses the bed is switched off and the hold latches until the
  printer is next a candidate, so a plate nobody clears cannot leave the bed
  hot indefinitely. The hold is also released when the item is deleted, the
  queue empties, or a gate is toggled off mid-hold, and never when firmware
  reports a target other than the one it set — a temperature the user or a
  print changed is left alone. Publishing is idempotent.

Smart soak reduction from chamber history
  The scheduler samples each connected printer's chamber temperature every tick
  into a 2h rolling history. Preheat credits time the chamber has already spent
  at temperature against the configured soak, shortening or skipping it.

  Credit starts no earlier than the newest sample, the most recent unbroken run
  of samples, or the end of the last real dip below target. A dip only counts
  once it outlasts a grace period: an enclosed chamber cannot lose and regain
  several degrees quickly (measured on an X1C, cooling from 55C to below 48C
  takes 23-73 minutes, ~0.2 C/min), so a brief low reading is a door opening or
  sensor noise rather than lost soak — and a plate swap, which is exactly when
  keep-warm runs, produces one. A stale history credits nothing: at that
  cooling rate the chamber can cross the threshold unobserved, so the full soak
  runs instead.

Three supporting changes to preheat itself:

  * Cancelling or deleting a queued item now stops a preheat already running
    for it. Those routes only write `status` to the database, which a dispatch
    coroutine parked in `asyncio.sleep` cannot observe, so the heaters ran for
    the rest of max_wait + soak — 45 minutes at the default settings — and the
    printer stayed in `busy_printers`, blocking every other queued item behind
    a print that was not happening. The routes now signal the scheduler
    directly, and the stage sleeps in slices so it notices promptly and
    abandons the dispatch, letting the existing rollback shut the heaters off.

  * A chamber-heated print whose slicer metadata carries no bed temperature
    (common for Orca-exported 3MFs) used to skip preheat entirely and start
    with a cold chamber. It now heats the bed to `queue_keep_warm_bed_temp`.
    A parsed bed temperature still wins, and a print with no chamber
    requirement still skips — no bed temperature is invented for the print
    itself. Preheat's bed target is transient regardless: the print's own
    gcode issues its M140/M190 at start.

  * Preheat records which commands it sent (bed, chamber, airduct) and unwinds
    them if the dispatch aborts before the print starts — a failed upload, a
    cancelled item, an exception — instead of leaving the printer heating for
    a job that is not happening.
2026-08-09 17:07:34 -04:00
maziggy b6027b7138 Say why a preset's values are unavailable, not just that they are
The settings panel collapsed four causes into one message -- "the picked
preset's own values could not be read" -- with no indication of what to do
about it.

The overwhelmingly common cause has an obvious fix, and it isn't an edge
case: an install pulls its sidecar as SIDECAR_TAG:-latest regardless of
which Bambuddy channel it is on, so a current Bambuddy talking to a
sidecar that predates POST /profiles/resolve is the normal state, not a
misconfiguration. Those users would have seen an amber warning on every
slice with nothing pointing at the sidecar image.

resolve_profile now returns ResolvedProfile(values, reason) instead of
None for everything, the route passes the reason through, and the panel
picks its message from it:

  sidecar_outdated     -> name the fix: update the sidecar image
  sidecar_unavailable  -> the sidecar did not answer
  not_configured       -> no sidecar is configured
  preset_unresolved    -> the previous generic wording

A request that fails outright maps to sidecar_unavailable, since a
backend we cannot reach and a sidecar that will not answer are the same
thing from the dialog.

Every variant still ends with "anything you don't change still uses the
preset" -- that reassurance is the point of the notice, and it is true
whichever way the lookup failed.

Tests pin the distinction rather than just the happy path: a 404 and a
500 must produce different reasons, and each panel case asserts both that
its own message appears and that the "update the sidecar image" line does
not leak into the others.
2026-08-09 11:25:09 +02:00
maziggy f421bb8160 Show the picked preset's real values in the process-settings panel
The panel baselined every field on the option schema's compiled-in
defaults, so a preset setting a 0.42mm line width displayed 0 -- the C++
default meaning "derive from the nozzle". Every field was affected; the
Line width group just made it obvious.

Bambuddy cannot answer this itself. A standard-tier pick is only an
{inherits: ...} stub on our side, and local/cloud presets are deltas whose
remainder lives in the profile tree bundled inside the running sidecar.
The values now come from the sidecar's POST /profiles/resolve, which runs
the same resolver /slice does against the same profiles, so what the panel
shows cannot disagree with what a slice produces. Deliberately not the
local orca_profiles resolver: it walks OrcaSlicer's published tree, which
can differ from the image actually installed.

An untouched field shows the preset's value and reverting returns to it.
isModified compares against that baseline too, so fields the preset moved
off the C++ default are no longer flagged as user edits, and values nobody
typed are no longer sent. When the values can't be read -- sidecar offline
or older than the endpoint -- the panel falls back to schema defaults and
says so rather than presenting them as the preset's.

Row layout, from screenshots:

- The control column is anchored to the right edge at a fixed width. It
  had been packed left after a fixed label column, leaving the values
  stranded mid-container with dead space beside them.
- Units are no longer truncated to "mm o...". The cap fitted the common
  "mm" but not "mm or %" or "mm/s² or %".
- The "from file" tick moved ahead of the control it qualifies; it used to
  sit past the unit at the row's right edge, reading as unrelated.

Both the unit and the control keep fixed widths, and the tick's slot is
reserved on rows without one -- sizing any of them to content makes each
row's input land at a different x and the column comes out ragged.

Also fixes a field that could not be cleared: emptying a free-text input
dropped the key, so it snapped back to the baseline and retyping appended
to it ("0.42" + "0.5" = "0.420.5"). The number branch was fixed earlier;
the text branch -- coFloatOrPercent, coString, the vector types -- was
not, and the regression test used a number input so it never caught it.

Requires a sidecar built from orca-slicer-api 4b664b7 or later. Older
images 404 the endpoint, which is handled as the fallback above.
2026-08-09 11:02:03 +02:00
maziggy f0500578bd Edit the full print-parameter set from the slice dialog
Slicing from Bambuddy meant taking a process preset as-is; any change
meant a round trip through Bambu Studio. The slice dialog now carries
OrcaSlicer's full process tree -- pages, groups, labels, tooltips,
ranges and defaults extracted from the slicer's own sources.

Enable/disable rules are evaluated from the slicer's own enable_if
expressions via a recursive-descent interpreter (no eval, CSP), with
enum comparisons validated against each option's declared values.
Anything undecidable leaves the field editable rather than greyed.

Overrides apply after the source's support config (#1881) and the
designer's carried tweaks (#2622), so an explicit choice always wins;
an untouched panel sends the same request as before.

Adds slice_engine as a separate setting from preferred_slicer -- where
slicing runs is a different axis from which binary the sidecar drives.
Only the sidecar engine is registered, so no picker renders yet.
2026-08-08 16:33:41 +02:00
maziggy 328bac450a Stop auto-drying re-arming into a threshold it can never reach (#2770)
An H2D armed five 12-hour drying cycles inside four hours, one of them six
seconds after the previous one ended, and none ran more than a couple of
hours.

Two things combine. The firmware ends a cycle when it decides the filament
is dry rather than when the clock runs out, and reports no fault doing it --
across this printer's history the run length tracks how wet the spools were,
from nearly the full 12 hours starting at 32% down to minutes once the unit
sat at 10-13%. That part is the AMS doing its job.

The loop is ours. An AMS reports higher relative humidity while it is warm
than once it has cooled: the same unit read 10-13% cold and 15-20% through
every cycle. With the threshold at 14% the reading at the moment a cycle
ended was always still above it, so the next 30-second pass armed another
12-hour cycle. Nothing counted, nothing waited, and it only stopped when the
box finally cooled enough to read 13%.

Auto-drying now waits 30 minutes after a cycle ends before arming another on
the same unit, and gives up on a unit after two consecutive cycles that
bring the reading no lower -- logging why and sending a new notification,
on by default because it reports that Bambuddy has stopped acting. Progress
is judged against the lowest reading any cycle on that unit has ended at,
not against the threshold, so a genuinely wet spool in a humid room coming
down 40-37-35 keeps drying however far it still is from the target;
comparing against the best so far rather than the previous end stops a
sensor wobbling by one point reading as progress every other cycle. The
suspension lifts by itself once the reading falls below the threshold.

Neither guard can stop a running cycle, and a cycle Bambuddy cut short for a
print, or that the user stopped by hand, is not counted against the unit --
so a farm that dries between queue jobs is unaffected. The threshold field
now warns below 20%, and every cycle end logs the unit's temperature and
humidity, which is what made this diagnosable.

The same bundle showed unrelated tasks failing with "database is locked",
each inside a 30.000-second Discord connect timeout. Alarms are raised from
inside the loop that records sensor history, at a point where the new rows
are added but not committed; the first read in the notification path flushed
them to satisfy itself, opening a write transaction, and the provider was
then contacted over the network with that transaction still open. SQLite
allows one writer and 30 seconds outlives the 15-second busy timeout, so
every other write in that window failed. The two reads that run before a
provider is contacted no longer flush the caller's pending work, and the
connect timeout is 5 seconds rather than 30 -- the body keeps the full 30,
so image uploads on a slow uplink are unaffected. SQLite only; Postgres has
no single-writer limit.
2026-08-08 12:40:17 +02:00
MartinNYHC 406cf71149 Merge branch 'dev' into feature/billing 2026-08-08 10:56:09 +02:00
maziggy 604fa44593 Explain Bambu Cloud's CAPTCHA challenge instead of repeating it (#2790)
A reporter tried to connect to Bambu Cloud and got "We need you to confirm you
are not a robot" as an error toast, with no CAPTCHA anywhere to answer and
nothing to click. That sentence is Bambu's, not ours. Their anti-abuse layer had
flagged the network and was answering the sign-in with HTTP 418 and a challenge
body: {"captchaId": "...", "error": "We need you to confirm you are not a
robot"}.

Bambuddy had no idea what that was. The reply is well-formed JSON, so
_detect_cloudflare_challenge -- which triggers on an unparseable body, CF
markers, 403+cf-mitigated or 503+cf-ray -- never fired on it, and login_request
fell through to its generic error path, which lifts data["message"] or
data["error"] out and hands it to the UI verbatim. The user was left to conclude
their password was wrong or that Bambuddy was broken. Four sign-in attempts
inside eighteen seconds appear in their log, each one more evidence for the
thing that had flagged them.

is_captcha_challenge matches on the 418 status plus a challenge marker in the
body -- captchaId is the reliable one, the wording is matched too because Bambu
has shipped it under more than one phrasing. A bare 418 with no marker is
has shipped it under more than one phrasing. A bare 418 with no marker is
deliberately NOT reported as a CAPTCHA: telling someone to solve a challenge
that was never offered is the exact confusion this issue is about.

login_request, verify_code and verify_totp now return reason="captcha" with an
explanation covering the three things the reporter had no way to find out: the
credentials are not the problem, the block is keyed to the public IP address
rather than the account, and it clears by itself within a few hours.

Sign-in requests are then held back for 300s so Bambuddy stops deepening the
block. Keyed per origin, not per service: TOTP verification posts to
bambulab.com while everything else posts to api.bambulab.com, and a challenge
seen on one must not strand somebody halfway through a two-factor sign-in on the
other. Entries expire on read, so the map cannot grow past one per region. The
token endpoint is deliberately left ungated -- it is the way out.

The UI shows a persistent panel rather than a toast. A toast names a problem the
user cannot act on and then vanishes; this one stays put and carries a one-click
route to "Use access token instead", which is the only thing that works while
the challenge lasts, since that path does not touch the challenged endpoint.

MakerWorld meets the same challenge from the same edge and now shares the
detection. It used to require the literal word "robot" in the error text and
reported any other wording as an unexplained block.

The System Health scanner gets a bambu-cloud-captcha signature. The reporter's
bundle came back with zero findings while their log was full of the failure.

Its advice for a failed FTPS handshake was corrected at the same time: it still
blamed firewalls and outdated firmware, which the #2780 investigation ruled out
last release -- it is the printer's own file service wedging, and the fix is to
restart the printer. The wiki said so already; the health panel did not.
2026-08-08 10:04:55 +02:00
maziggy 14d0d14365 Power on a printer for jobs queued to a printer class (#2786)
Queue a print against a printer class -- "Any X1C", or a Slicer Pipeline whose
target type is Printer class -- with every printer of that class switched off,
and nothing happened. The job sat pending and no smart plug was touched, while
the same file pinned to a specific printer powered that printer on within one
queue check. The reporter's log holds both halves: thirteen minutes of the item
being polled as (133, None, ...) and passed over, then a PATCH onto printer 2,
then "Printer 2 offline, attempting to power on via smart plug(s)" on the very
next tick. Same item, same plug, same Auto On setting.

Powering a printer on had only ever been written inside `if item.printer_id:`.
The model-based branch below it walks the same queue but its matcher classes an
offline printer as a reason to keep waiting -- printers_offline collects the
*name*, for the waiting reason -- and nothing on that path ever looks at plugs.

_wake_printer_for_model adds it. The model query moves into _printers_for_model
so the matcher and the wake step answer "which printers can this job run on"
from one place: a job can only be woken onto a printer the matcher would also
have considered. Candidates that failed the cross-model gate are excluded --
switching a printer on for a file that can never legally run on it leaves the
job just as stuck, with the printer now drawing power.

Two things it does that the fixed-printer branch does not:

A printer awaiting plate-clear acknowledgment is skipped. Waking it buys
nothing; it boots into IDLE and is held by the gate. That is what the reporter's
log shows for the eighty minutes after their manual edit -- "printer 2 not
available -- connected=True, state=IDLE, awaiting_plate_clear=True" every thirty
seconds to the end of the capture. The flag is Bambuddy-side and persisted, so
it is readable while the printer is still off.

At most one printer per pass, because each wake blocks the queue loop for the
boot wait. Several queued jobs bring several printers up over the following
minutes rather than a whole shelf at once.

A failed power-on opens a 600s per-printer cool-off. Without it the walk is by
id, the pass spends its single attempt on the same broken printer every time,
and a healthy sibling two slots down is never reached -- one unreachable plug
starves its whole model, and costs a 180s boot timeout out of every 30s pass.
Entries expire on read: a printer inside its cool-off is skipped before the
power-on is reached, so a live entry can never be overwritten by a success.

The failed printer is deliberately NOT added to busy_printers. It is off, not
busy; labelling it busy would misdescribe it in every later item's waiting
reason and, because an all-busy reason is treated as needing no user action,
suppress the notification too.

Assignment is left to the next pass. AMS trays arrive with the first status push
after connect, so matching filament against a printer that booted five seconds
ago can reject the printer we just woke.

Finally, the waiting reason separates "Offline: X1C-1" from "Offline, no Auto On
smart plug: X1C-2". Those are different problems and only the second is one the
user has to go and fix -- it was also the first question the reporter had to be
asked, and the queue could not answer it.

Tests cover the wake, the plate-clear skip in both gate states, all-candidates-
awaiting-plate-clear waking nothing, one wake per pass, the starvation case over
two passes, cool-off expiry, no-Auto-On-plug being left alone and named, an
incompatible sliced model waking nothing, connected printers being left alone,
scheduled-for-later and manual-start jobs switching nothing on, and a regression
pin on the fixed-printer branch.
2026-08-08 09:35:43 +02:00
maziggy 91acac2b35 Stop retrying a printer whose FTPS handshake fails, and name the cause (#2780)
Two printers went on printing while every archive they produced held nothing
but a filename. Bambuddy opened port 990, the printer accepted the connection
and answered with something that was not TLS, and connect() logged a warning
and returned False -- indistinguishable, to every caller, from "the file is
not at this path". So the 3MF lookup walked all six filename variants across
five directories with four retries each, the cover endpoint ran its own
sixteen-path sweep, and the timelapse scan added four more, all against a
sixteen-path sweep, and the timelapse scan added four more, all against a
printer that could not have answered any of them. One reporter's log carried
1813 identical handshake failures, another's 3511.

The evidence says this is the printer's own file service getting stuck, not a
model, firmware or TLS-configuration problem. In #2780's bundle the same two
printers ran clean from 22 July to 4 August and failed again from the 5th; a
second bundle shows an X2D serving files for five days, flipping on 19 July,
then failing every connection for eight days with zero successes. The same
models and firmware appear in roughly twenty other bundles with no occurrences
at all. Both bundles show it happening with cap_tls_v1_2 in effect -- the X2D
and H2C entries in ftp_profiles were added on analogy with P2S to fix exactly
this symptom, and the reporter's own debug line proves they do not.

An ssl.SSLError from connect() now opens a five-minute cool-off for that
printer. Subsequent connects return False without touching the network, so a
wedged printer is contacted twice an hour instead of hundreds of times a
minute, and the single warning that is logged names the remedy. The cool-off
is dropped on expiry rather than kept, so the map holds one key per currently
wedged printer. ftps_handshake_blocked() lets the sweeps stop: the 3MF lookup
abandons the remaining paths and skips the directory-walk fallback, the cover
endpoint returns 503 naming the file service instead of a 404 that reads as
"this print has no thumbnail", and the timelapse scan separates 503 (cannot
reach the printer) from 404 (no timelapse directory) -- one 500 used to cover
both, which is what the reporter hit when reproducing.

The Connection Diagnostic completed a bare TCP connect to 990, which is why it
reported the port green throughout: the port is open, it is what is behind it
that is broken. It now completes a real implicit-TLS handshake using the
model's own ftp_profiles cap, so a pass means the FTP client would also get
through. An open port that cannot negotiate reports warn with reason no_tls,
selecting a new message in all 13 locales that points at a printer restart
rather than at the firewall. No login is attempted, so this stays valid in the
pre-save Add Printer flow.

The cool-off tests run against a real socket that accepts on 990 and replies
with a plaintext FTP banner, reproducing WRONG_VERSION_NUMBER rather than
mocking ssl. The autouse fixture clearing _mode_cache now clears the cool-off
map too -- every test here talks to 127.0.0.1, so one left behind would make
the next test's connect() a no-op.
2026-08-08 09:00:03 +02:00
behrinml a9e23910fc Merge branch 'dev' into feature/billing 2026-08-06 20:32:59 +02:00
behrinml dd1d40b0d4 implemented pr (worth fixing) feedback
update commit
2026-08-06 20:27:25 +02:00
maziggy 9c86a05657 Check filament deficit for Library-backed queue items (#2779)
A job needing 20.5 g was dispatched onto a spool holding 9 g and the printer
started. _resolve_source_3mf returned LibraryFile.file_path verbatim, but that
column stores a path relative to base_dir -- so it resolved against the process
working directory, found nothing, and compute_deficit_for_queue_item treated a
missing source as "nothing to verify" and returned no deficit.

Every library-backed queue item was affected: Slicer Pipeline jobs, which are
always library-backed, and everything added through the Library's bulk Add to
queue. Both callers share the resolver, so the Play button on the queue was as
blind as the auto-dispatcher. Archive-backed items (print history, VP intake)
resolved correctly and were never affected, and neither was PrintModal, which
resolves the file on its own path.

The library branch now uses the same idiom as the eleven other readers of
file_path -- absolute stays, relative joins base_dir. The join carries a
SEC-PATH-OK marker: the value is DB-stored and generated by the Library ingest,
and it is already what resolves the file for upload, so the check has to
resolve it identically or it is not checking what gets printed.

A source that is configured but absent now logs a warning naming the item and
the resolved path. It still dispatches, because the upload needs the same file
seconds later and fails there, where blocking would strand a queue on a moved
file -- but a safety check that skips itself must not do so in silence, which
is what hid this for every library-backed item.

Tests cover the relative path (the reporter's 20.5 g against 9 g), the absolute
path against a base_dir the file is not under, and the missing-source warning.
The existing cases all used archives with absolute paths, which is the gap the
bug lived in.
2026-08-06 12:21:50 +02:00
maziggy 306b9ba7fd Accept Forgejo tokens scoped to a single repository (#2775)
ForgejoBackend.test_connection asked GET /user who the token belonged to
before asking whether the token could reach the repository, and treated a 403
there as fatal. A Forgejo v15 repository-scoped token may only carry
read/write on issues and repositories, so it 403s on /user -- and was rejected
despite reaching its own repository fine, which is all a backup needs: the push
path uses the Contents API and restore reads commits, trees and blobs, all
under /repos/{owner}/{repo}. That /user call was the only one in the whole
provider layer.

The probe stays, because a 401 from it is genuinely conclusive and names a bad
token before the repo call has to guess -- Forgejo v15+ hides a private repo
behind 404 rather than 403, so the repo call cannot always tell those apart.
Every other status now falls through to the repo check.

Two additions keep the messages as sharp as before: the repo call's own 401 is
mapped to "Invalid access token" instead of a generic API error, and the 404
names write:repository and the scoped-to-another-repository case, mentioning a
possibly-invalid token only when /user did not confirm the identity.

The token hint under the field was one shared string reading "fine-grained
token with Contents read/write" -- GitHub's advice, shown to Gitea, Forgejo and
GitLab users too. It is now per provider via PROVIDER_TOKEN_HINT_I18N_KEY,
following the existing repo-URL placeholder map, translated in all 13 locales.

Tests pin the repository-scoped token connecting, a transient /user status not
blocking the repo call, both 404 wordings, and the repo-call 401; a frontend
test switches providers and asserts the hint follows.
2026-08-06 12:08:36 +02:00
maziggy b5163b94f8 fix(backup): report the categories a failed restore already committed (#2656)
The service reports what landed on a part-way failure -- categories commit as
they finish, so results names the ones on disk -- and the modal gated the
whole result panel on success, so it showed the failure message and dropped
them.

The cache invalidation was inside that same branch, which is the half that
mattered: a run that committed the settings category and then failed left the
app rendering pre-restore settings, with no reload and no re-read, which is
the failure the modal's own reload-on-close exists to prevent.

Gate on what was written instead. A refusal that never reached a category
still carries an empty results and still keeps the form, so the mutex and
backup-in-flight cases are unchanged. A partial does not read as a success:
the tick becomes a warning and a line says the listed categories are the ones
on disk.

---

fix(backup): keep the local owner when the backup names one we cannot resolve (#2656)

An owner the backup names but this instance has no user for was written as
NULL, and overwrite is a blanket setattr -- so restoring over a local archive
that had a perfectly good owner took it away, which is the 404-for-its-own-
owner failure this column is carried across to fix. Resolving by username
widened the trigger from a stale id to any user renamed since the backup.

It is the same state as an absent key: the backup has not told us who owns
this. So it takes the same action -- the column is not written at all.
Overwrite keeps the local owner, insert lands ownerless with the note, and an
explicit null still writes, so overwrite still means "match the backup".

The notes move to the insert path with it. On overwrite nothing was taken
away, so there is nothing to warn about, which is the rule the absent-key
case already follows.
2026-08-06 08:46:02 +02:00
MartinNYHC 6cd81fcd85 Merge branch 'dev' into feature/2656-restore-from-github 2026-08-06 08:23:37 +02:00
maziggy 1eea194953 Resolve a spool's material to a known drying preset before starting a cycle (#2774)
The drying popover prefilled its material from the loaded spool without
checking the preset table had that material. An AMS-HT holding Support for
PLA/PETG (tray_type PLA-S) fell back to PLA's temperature but kept PLA-S as
the material, and the dropdown displays its first option when handed a value
outside its list -- so it read PLA while PLA-S was sent. Same gap for every
composite: PETG-CF prefilled at PLA's 45C.

Resolve the tray_type to a key the table has before setting either value.
Support materials and composites resolve to their base, nylon is aliased
under its several spellings, and anything unrecognised falls back to PLA --
the coolest row, so an unknown material under-dries rather than deforming a
PLA spool.

Also record request-topic messages in the MQTT debug log. That topic carries
every command a printer is given, including Bambu Studio's, and returned
before the logging block -- so a capture could show only what the printer
said, never what it was told.
2026-08-06 08:19:01 +02:00
jmoore-skild f62e907e9f fix(backup): refuse the whole LDAP family on restore, not just its password (#2656)
A settings restore could substitute the instance'"'"'s authentication source.
auth.py reads the LDAP config live from the settings table on every
login, and none of ldap_server_url, ldap_user_filter, ldap_auto_provision
or ldap_default_group is credential-shaped, so the secret-key hints never
saw them and only the four auth-policy keys were protected.

ldap_enabled was covered by the companion-credential rule instead, and
that rule asks the wrong question. It judges availability - "will the
integration still work?" - and an anonymous bind works, so a payload that
simply OMITS ldap_bind_password skips the refusal and has its toggle
written. Omitting the credential is exactly what an attacker authoring
the file would do: they own the directory being pointed at, so they need
no bind credential from us.

Left unrefused, a backup repository anyone can write to yields admin:
point ldap_server_url at your own directory, set ldap_auto_provision and
ldap_default_group=Administrators, and the next login on a fresh username
is provisioned into the admin group. Overwrite-off is enough on an
instance that never configured LDAP - there are no rows to skip.

Refused by prefix so a key added to the LDAP schema later is refused by
default, and matched case-insensitively because the key comes from the
backup JSON rather than from our own writer. ldap_enabled leaves
_COMPANION_CREDENTIALS rather than sitting there as dead code, since
_is_protected_setting_key runs first.

The two tests asserting an anonymous bind was a false positive are
inverted - they encoded the hole - and the refusal reuses the existing
settingsAuthSkipped note, which already points at Settings >
Authentication.
2026-08-05 20:05:23 -04:00
jmoore-skild bb25e36510 Merge branch 'dev' into feature/2656-restore-from-github 2026-08-05 16:47:49 -04:00
jmoore-skild f6fca4b927 fix(backup): keep both restore tallies equal to the number the preview showed (#2656)
Two ways the K-profile and spool categories broke the
restored + skipped + failed == item_count invariant the settings count
holds:

* The spools preview counted only the spools and put the usage records
  in the detail, but _restore_spool_usage increments the same tally, so
  any backup with usage history reported a total larger than the number
  the user was shown. The preview now counts both and the detail breaks
  the total down instead of adding to it.
* A K-profile entry that is not a dict was dropped silently on the
  connected path. _kprofile_profile_count includes it, so the offline,
  printer-missing and step-failed paths all account for it; only the one
  path that talks to a printer let it leave the tally. It now counts
  failed.
2026-08-05 16:17:42 -04:00
jmoore-skild fc2dcf76b5 fix(backup): commit each database category so SQLite's writer is not held (#2656)
The database phase had the same shape the K-profile phase did: _find_
archive, _find_spool, the usage dedupe and _restore_settings are all one
SELECT per row or per key, interleaved with autoflushed INSERTs, inside a
single open write transaction. A few thousand archives plus a full usage
history plausibly passes the 15 s busy_timeout, and every concurrent
writer in the app fails with "database is locked" until it finishes.

Each category now commits before the next starts. The id maps are plain
dicts in memory and the session is expire_on_commit=False, so the
ordering tolerates it.

The cost is that a later failure no longer rolls back an earlier
category, so a tally is recorded only after its category commits and
run_restore reports the categories already on disk instead of an empty
result - the same correction the K-profile split needed.
2026-08-05 16:14:23 -04:00
jmoore-skild 8602c54c1f docs(backup): say what the secret-key hints actually refuse (#2656)
The comment called the hint list belt-and-braces over keys the collector
already refuses to write. It is not: _collect_settings filters exactly
bambu_cloud_token and auth_secret_key, so a current backup really does
carry mqtt_password, ldap_bind_password, ha_token and prometheus_token,
and the hints are the only thing that refuses them. The companion-
credential rule sits downstream of that, so reading the list as redundant
and shortening it would write a stale credential and make that rule inert
at the same time.

Comment and test docstring only - no behaviour change.
2026-08-05 16:09:50 -04:00
jmoore-skild 4ec0f3f9f2 fix(backup): resolve a restored archive's owner by username, not by id (#2656)
created_by_id is only meaningful on the instance that wrote it. Restoring
onto a rebuilt instance - this feature's main use case - renumbers the
users table, so a live id can land on a different person and hand one
user's print history to another under archives:read_own. The id path
cannot even detect that: archivesOwnerCleared fires only for an id that
is absent, so a valid-but-wrong id produced no note at all.

The collector now records created_by_username alongside the id, and the
restore prefers it. username is unique on users, so a match is the same
person; the one case it cannot resolve - a user renamed since the backup
- falls through to ownerless with a note rather than guessing from the
id. The id stays as the fallback for backups taken before this change.
2026-08-05 16:09:20 -04:00
behrinml 9a397f46e9 implemented pr feedback #2 2026-08-05 21:25:46 +02:00
behrinml 43bf854bd5 Merge remote-tracking branch 'upstream/dev' into feature/billing 2026-08-05 20:19:32 +02:00
maziggy 16c8c6f2ea Security hardening (maziggy/bambuddy-security #9) 2026-08-05 15:21:42 +02:00
maziggy cd004df817 Show Home Assistant sensors on the printer card (#1148, #448)
Binds binary_sensor and reading-carrying sensor entities to a printer and
renders their state on its card, worded by Home Assistant's device_class.
Optional per-sensor alert condition drives a notification on the transition
into the alert state and an opt-in interlock that holds queued prints while
alerting -- a hold with a readable waiting_reason, never a failure, and only
ever on a sensor that was read successfully.

Sensors get their own table rather than a wider entity pattern on SmartPlug:
get_smart_plug_by_printer would otherwise hand the card's power button a door
contact to switch.

The hold is passed to the model matcher directly rather than merged into
busy_printers: _check_auto_drying reads that set as "is currently printing"
and would put an idle-but-held printer down the mid-print drying path.

The notification_providers migration spells its default FALSE, not 0 --
Postgres rejects an integer default for a boolean and _safe_execute swallows
the error.
2026-08-05 14:26:38 +02:00
maziggy 945d4ca6eb Route external spools to a nozzle when the printer has no AMS (#2771)
Five X2Ds with no AMS, each printing from its external spool holder,
took a job sent to a named printer and refused the same job sent to
"Any X2D": the file uploaded, the firmware answered 0700_8012 "Failed
to get AMS mapping table", and the item failed after three attempts.

A named-printer job carries a mapping the frontend resolved at queue
time, so the scheduler's matcher never runs. A model-based job has no
printer until dispatch, so the matcher does run -- and could not see an
external spool on a dual-nozzle printer. _build_loaded_filaments derived
dual-nozzle status from ams_extruder_map, which is built from AMS info
bits, so a printer with zero AMS units reported an empty map; every
external spool got extruder_id=None, and the nozzle-aware hard filter in
_match_filaments_to_slots discarded it because None equals neither 0 nor
1. The mapping came back all -1, was cleared to None, and the print
command went out as use_ams:true with no ams_mapping and no
ams_mapping2 at all.

This is the backend half of #1257, which fixed the same logic in
useFilamentMapping.ts and left this copy behind. Mirror its inference:
a populated nozzles[1].nozzle_diameter, a non-empty ams_extruder_map, or
more than one vt_tray entry. Replaying the reporter's own push-status
now yields extruder 1 for Ext-L and 0 for Ext-R, and a nozzle-1
requirement resolves to [254] -- what their working named-printer
dispatch sent. Single-nozzle printers keep extruder_id=None; nozzles
always has two entries, so its length alone must not be the signal.

Also stop dispatching a job the firmware is certain to reject. When the
matcher ran, matched nothing, and the printer has no AMS, fail the item
with the filament and nozzle it wants instead of spending an upload and
two retries on it -- that path already ended in a failed item, just an
opaque one. With an AMS attached the firmware error still stands, since
there the user can load a spool and press Resume. Fail-safe like the
nozzle-diameter guard (#1899): every branch short of a positive finding
returns None and dispatches as before.

_apply_filament_overrides is extracted from _compute_ams_mapping_for_printer
so the message names the filament the matcher looked for rather than the one
the 3MF was sliced with.
2026-08-05 13:15:04 +02:00
maziggy 0596ff424e Say why a drying cycle ended when the firmware cuts it short (#2770)
An H2D started a twelve-hour PETG dry at 65 degC and the AMS gave up on it
twenty minutes in, with 700 of the 720 minutes still on the clock. It cooled,
humidity climbed back over the threshold, auto-drying started another
twelve-hour cycle, and that one went the same way; the reporter's AMS
temperature history shows the loop running all morning.

The log had one line for it: "AMS 0 drying complete", which is exactly what it
says for a dry that ran its full twelve hours. Nothing in a support bundle told
the two apart, and the one number that does -- the time still remaining -- was
written into that line as the previous value, where it reads like a duration
rather than a shortfall. The reporter took 700 for seconds and concluded the
cycle had lasted twelve minutes.

Bambuddy did not stop that cycle; every stop it sends is logged with the full
outgoing command and there was none. So ending it was the printer's decision,
and the account of why lives in three things already received and parsed and
never written down: the drying phase and sub-phase from the AMS info hex, the
per-unit dry_sf_reason constraint codes, and the live HMS errors.

A cycle that ends with most of its countdown left now logs all three alongside
how much of the requested duration ran. One that reaches its duration keeps the
single line it has always had. A stop Bambuddy sent is named as ours -- it is
short of its duration too, and on the telemetry alone is indistinguishable from
the firmware abandoning the cycle, so without tracking it the print-takes-priority
stop and the Stop button would both have been blamed on the printer.

Diagnostics only. Nothing about when drying starts or stops has changed, and the
restart loop is not addressed: what the firmware objects to has to be established
before Bambuddy can sensibly decide how long to wait before trying again.
2026-08-05 12:38:45 +02:00