failure_reason was written three different ways and nothing reconciled
them. derive_failure_reason wrote English display labels ("Layer shift"),
older builds of the archive editor wrote the translated label in whatever
locale that user was running, and the two stale-archive paths wrote
English prose sentences. All three reach one column -- the archive PATCH
has mirrored the field onto the latest print-log entry since #1444 -- and
the Failure Analysis widget groups on the raw value, so one real cause
occupied several buckets. On a live install before this landed:
print_log_entries held 91 rows reading "User cancelled" beside 1 reading
"userCancelled".
In an English UI those two render as the same words twice with different
counts, which is why nobody spotted it. In any other locale one of them
stays English, because a stored label has no key for t() to resolve. The
editor was worse than cosmetic about it: its reverse lookup compared the
stored value against t() in the current locale, so for a non-English user
nothing matched and the dropdown opened empty over an archive that
plainly showed a reason.
The keys were already canonical and already enforced.
_FAILURE_REASON_KEYS in api/routes/print_log.py rejects anything else
with a 400 and explains why in its own comment -- the widget renders
values back through t(), so an unrecognised one surfaces as a raw string.
derive_failure_reason had simply never been held to that rule. It now
produces keys, and the cancel branch returns userCancelled.
The two "Stale - ..." sentences become one new noStatusUpdate key. Both
describe the same observation, that no end-of-print status ever arrived;
which of the two situations occurred is already carried by status --
cancelled at the stale-cleanup site, the reconciled outcome at the
reconnect site -- so collapsing them loses nothing and gives Statistics
one bucket instead of two sentences that could never be translated. It
had to enter the vocabulary rather than merely be tolerated, because the
editor discards any value it does not recognise.
Existing rows are converted by a startup migration folding 168 historical
labels onto the 12 keys across both columns. It is exact rather than a
guess: every label across all 14 locales resolves to exactly one key,
with no collisions. The map is a frozen snapshot rather than something
read from the locale files at run time -- it maps what was written
historically, so regenerating it from the current translations would
silently stop recognising the very rows it exists to convert. A value
outside the map is left alone; guessing would be worse than leaving one
honest string in its own bucket. There is no one-shot settings flag, on
purpose: the statement only matches values in the map and a key is never
a label, so it is self-terminating, and a flag would permanently skip
anyone who restores an older database.
The last part is a data-loss bug that was not in the report. The editor's
fallback to '' was not merely a wrong-looking dropdown -- the empty
selection was then saved over the stored text, so opening the editor on
an archive whose reason was free text and pressing Save destroyed the
classification. An unrecognised value now keeps its own option and
survives a save.
A slicer preset is bound to a printer model: "Bambu PLA Basic @BBL X1C" is
not the same preset as "@BBL H2C", and Bambu names a nozzle size in it as
well. A spool carried exactly one, which was right until the same spool was
used on a second machine -- the AMS slot on the other one was then
configured with a preset that machine has no profile for. K profiles had
the matching gap from the other side: the tables have always been keyed per
hotend, but the picker could not express it.
spool_filament_preset and its Spoolman twin store the exceptions, keyed
(spool, printer_model, nozzle_diameter). Model rather than printer because
the preset is a property of the model -- "@BBL X1C" is the same preset on
every X1C, and asking per machine would mean picking the identical value
twice. K profiles stay on printer_id, because a K value is measured on one
physical hotend and two machines of the same model legitimately differ.
Resolution is exact (model, diameter) -> (model, "") -> the spool's own
preset, so a spool nobody has configured behaves exactly as it did before.
The form writes one row per nozzle size and never the "" row; that level is
kept for API clients wanting one value to cover a model.
Both halves cover every standard nozzle size rather than the size currently
fitted, because a spool is configured once and nozzles get swapped. The PA
Profile tab becomes a Printers tab: a model list beside a detail pane
holding a preset row per size and a K-profile grid of size by hotend. Each
model is offered only the presets that name it, through the same matcher
the Configure AMS Slot modal filters with, which moves out of that
component into utils/slicerPrinterMatch. Presets whose name identifies no
model -- most user-authored and OrcaSlicer ones -- stay offered everywhere,
as does whatever is already selected, so a saved override cannot vanish
from the control that shows it. Every preset carries an origin badge in the
wording and colours that modal already uses.
Every path that configures a slot now respects both: manual assign in
either inventory mode, RFID auto-assign, the Spoolman tag link, the re-fire
when a slot goes empty to loaded, the re-apply after a calibration-table
refresh, and the re-selection when a Filament Track Switch moves an AMS to
the other nozzle. Which nozzle a slot feeds, and how wide it is, was worked
out independently in seven of those places, each reading nozzles[0] for
every slot on the machine -- correct on a single-nozzle printer and on a
dual-nozzle printer with matching nozzles, wrong the moment two sizes are
fitted. That resolution is now services/slot_nozzle.
Which array entry belongs to which hotend is no longer inferred. Measured
on an H2D fitted with a 0.4 high flow on the left and a 0.6 on the right,
nozzles[0] reads the right hotend, so the array is indexed by extruder id
and the H2/X2 parser's convention is the one that holds. The legacy
parser's opposite convention never governs a real dual-nozzle machine:
every model in DUAL_NOZZLE_MODELS reports device.nozzle.info, and
left_nozzle_diameter appears in no log or wire capture. Two comments that
said otherwise were wrong and are fixed; amsHelpers' code was right all
along and only its comment lied.
Four defects surfaced while wiring it, all pre-existing except the last.
The picker identified a chosen calibration by cali_idx alone, and the
printer numbers its calibration table per nozzle -- on a dual-nozzle
machine the same index exists on both hotends meaning different things, so
saving could persist the other hotend's K value and diameter; SpoolBuddy's
write-tag page carried a verbatim copy and gets the same fix. RFID
auto-assign chose a K profile with no extruder test at all, so a spool
calibrated on both hotends had a coin toss decide which pressure-advance
value the slot got, on the path that runs unattended every time a Bambu
spool is loaded. The Spoolman tag-link path resolved no preset whatsoever,
configuring every linked slot with a generic material id and discarding a
preset set in inventory -- the same defect #1713 fixed on the assign path,
one function over. And an FTS inlet move re-selected K for nozzle 0 rather
than for the nozzle the AMS had just been moved to.
The last one is new here: a per-model override can be a cloud USER preset,
whose PFUS-prefixed id the slicer rejects, and passing it straight into
extrusion_cali_sel would silently lose the K-profile link. Reached the
printer only where such an override exists, which is why nothing in the
suite caught it. printer_safe_filament_id falls through to the spool's own
preset and then the tray's RFID value instead.
Reading a printer's calibration table asks for one nozzle size at a time.
H2-series firmware answers only the first one or two of a concurrent burst
of extrusion_cali_get and silently drops the rest, each dropped request
costing a five-second timeout before its retry: measured at 11 and 23
seconds on an H2C and an H2D for four parallel requests, against roughly
one second in series. An X1C answers all four at once, which is why this
only ever surfaced on dual-diameter printers. Printers themselves are read
in parallel -- separate machines are separate connections.
The Configure AMS Slot dialog opens on the spool's own configured values,
falling back to the slot's last manual configuration and then the tray's
RFID data. The spool form is wider for the two-pane layout, colour, weight,
cost and location move to their own tab in two columns, and a printer card
in expanded view lists every fitted nozzle size rather than the first entry
alone.
Adding on_stock_reorder_alert and on_stock_break_alert to the provider
schema made them required on the way out as well as in: the response model
inherits the write model. Every on_* column on notification_providers is
nullable with no server default, and where the table was created from
Base.metadata before run_migrations, the ALTER ... DEFAULT false that
introduced those columns was swallowed as a duplicate and never backfilled
existing rows. Those NULLs were harmless until the flags were read, at
which point the row failed validation -- and a list is validated as a
whole, so one row took every provider with it. The route returned 500 and
the UI rendered an empty list, so configured providers looked deleted.
Backfill them to off, which is what the sender already assumed: it selects
providers with IS TRUE, so a NULL flag never sent anything. A NULL flag now
also reads as off rather than failing the response, across all of them, so
the next flag added to this schema cannot repeat it. Writes are unchanged.
PostgreSQL renders its messages in the server's lc_messages locale. _safe_execute
recognised an already-applied statement by searching the error text for
"already exists", so a Russian-locale server -- which says "уже существует" --
re-raised it and aborted startup.
The column already existing is the expected outcome: create_all() builds the
tables from the models before the migration list runs, so on a fresh database
essentially every ADD COLUMN in that list is a duplicate by design, and all 382
of them relied on that recognition. No PostgreSQL server outside an English
locale could start Bambuddy at all, fresh install or upgrade.
Classify on SQLSTATE instead -- 42701, 42P07, 42710, 23505 -- which PostgreSQL
never translates. The existing narrowing is kept and now rests on a code rather
than a phrase: a missing column counts as already-applied only for RENAME COLUMN,
so a missing column during ADD COLUMN or CREATE INDEX still aborts rather than
hiding a corrupt schema. SQLite keeps the text match; its driver publishes no
SQLSTATE and it does not localise. The OIDC auto-link constraint read message
text the same way and gets the same treatment.
Verified against PostgreSQL 15 under ru_RU, en_US and C: init_db() completes on
a fresh database and on a re-run in all three, and the schema the Russian server
ends up with is byte-identical to the English one.
The Storage Location dropdown listed entries like "H2D-1 - AMS A1" next to
real locations, and they could not be got rid of.
They were never locations. Bambuddy used to record which slot a spool was
loaded into by writing that string into Spoolman's location field, and the
writer went away when Storage Location became something the user picks --
but the strings stayed on people's Spoolman spools, and the location sync
imports every distinct one it finds, so they have been coming back in
through the front door ever since. A printer slot is where a spool is
loaded, not where it is put away, and slot assignments already track the
first.
Deleting one by hand did not work either, which is what made this a dead
end rather than an annoyance: the delete route refuses a location that has
spools, and in Spoolman mode it counts them by matching that same string,
so every marker still sitting on a loaded spool answered 409 -- and the two
that were empty were back on the next sync a minute later.
The import now skips them and a one-shot migration clears the ones already
in the catalogue. The shape is defined once and used by both: an optional
printer-name prefix followed by AMS A1, AMS-HT A1 or External Spool, which
is exactly what convert_ams_slot_to_location produced. It stays narrow on
purpose -- "AMS Drybox" and "Spare AMS trays" are somebody's shelf, and
anything the filter swallowed would be a place they could no longer file a
spool under -- so both directions are pinned by tests.
A row is only removed when no spool in this database points at it, by id or
by legacy free-text name, so an internal-mode user who has deliberately
filed spools under such a name keeps it. Spools in Spoolman are neither
consulted nor touched: their location strings are the user's data on the
user's server, and one that still reads "H2D-1 - AMS A1" in the inventory
list is telling the truth about what Spoolman holds. It simply stops being
offered as a destination.
Verified on a live Postgres instance carrying the reported symptom: 13
locations down to 3, all ten markers removed, the two real shelves and one
hand-typed Spoolman name left alone.
Two faults in the same auto-add path, both found while tracing why an H2C
slot named a wood roll as plain PLA.
A spool's swatch is composed from effect_type and extra_colors, and the
RFID auto-add set neither. It reads the colour catalogue to name the
colour and took the name alone, even though the row it had in hand also
carries those two columns -- the spool form's own colour picker hands both
to a spool a user adds by hand, so the same roll rendered one way when you
typed it in and another when the printer identified it for you.
Both columns now travel with the name. That alone changes nothing on a
stock install, because the shipped catalogue carries an effect on none of
its 600-odd rows, so the subtype is read where the catalogue has none: it
is already derived from what the printer reports, and the two vocabularies
line up -- Wood, Silk, Sparkle, Marble, Glow, Galaxy, Metal, Rainbow,
Translucent, Matte, and the Gradient, Dual Color and Tri Color that the
M*/T* colour codes upgrade a subtype to. "Silk+" reads as Silk, since the
plus is on the product name rather than the finish. A subtype that names
no effect -- Basic, Tough, CF -- leaves the column empty rather than
inventing an overlay, and a value already set is never overwritten, so the
column stays what it is documented to be: a rendering hint the user can
override without touching Bambu's categorical label.
------
Correct the spool tare an RFID roll was added with (#2909)
The lookup that gave an auto-added spool its core_weight asked for the
first catalogue row whose name starts "Bambu Lab" and took whatever came
back. There are three, and which is first is the database's business:
SQLite returns insertion order in practice, Postgres promises nothing once
a table has seen an update. The same roll was therefore recorded with the
216 g High Temp tare on one install and correctly with the 250 g Low Temp
one on another. @ojimpo's forward fix picks the row by name; this repairs
the rows already written, which the forward fix cannot reach -- 22 of 26
RFID-added spools on the instance this was traced on.
The tare is not cosmetic. A spool weighed on SpoolBuddy has its remaining
filament worked out as the scale reading minus the tare, so a 34 g low
tare credits the roll with 34 g that is not there and writes a used weight
34 g short. That error is a constant -- every later print adds to the used
weight on top of it -- so adding the difference back is exact however much
has been printed since. It is applied only to spools that have been on the
scale; one that never was has a used weight derived from the AMS remaining
percentage, which the tare never entered into.
Rows are identified by the signature of the broken lookup: added by RFID,
carrying the weight of one of the other Bambu catalogue rows, with the
weights read out of the catalogue rather than hardcoded so an install
whose rows have been re-measured is repaired to its own numbers. Keying on
whether a catalogue row had been recorded would not have worked -- the
weight picker auto-selects the only row matching the weight and writes its
id on the next save, so that column says only whether the form was ever
opened. The one case that cannot be told apart is stated rather than
hidden: someone who moved an RFID roll onto a genuine High Temp spool and
set 216 g by hand is normalised with the rest. Runs exactly once, so a
tare set afterwards is kept.
On a UTC+3 install every AMS humidity reading and every archive was
stamped three hours ahead of when it happened. Bambuddy stores naive
timestamps that hold UTC and the frontend's parseUTCDate reads an
offsetless timestamp as UTC, so the display added the offset to a value
that was already local.
The Python side has honoured that contract since #504. The reporter's
timestamps were not written by Python. Around ninety-six columns take
their value from server_default=func.now() and the migration DDL carries
another forty-nine on DEFAULT CURRENT_TIMESTAMP -- the database fills
those, and on PostgreSQL now() is a timestamptz, so storing it into a
timestamp without time zone casts it through the session TimeZone. A
Postgres container started with TZ=Europe/Istanbul bakes that zone into
postgresql.conf at initdb, and every defaulted column then receives local
wall-clock. recorded_at is the clearest case: nothing in the codebase
ever assigns it, so its value is entirely whatever the database decided.
Connections now carry timezone=UTC, which makes the cast a no-op whatever
the server is set to. Measured through the real engine factory against a
live PostgreSQL, a session on the reporter's configuration stored +10800s
and the fixed one +0s. Pinning the session was preferred over a hundred
and forty-five individual edits partly for its size but mostly because
half of those sites are raw DDL that no model-level change can reach.
SQLite needed nothing and gets nothing: its CURRENT_TIMESTAMP is UTC by
definition and it has no session timezone to get wrong, which is why this
survived two years of timezone fixes without showing itself. That also
makes it the reference -- the change moves Postgres onto SQLite's
behaviour rather than introducing a third convention -- so the SQLite
behaviour is now pinned by a test instead of being assumed. asyncpg is
the documented driver and takes the setting in its startup packet; any
other Postgres driver gets the same setting the libpq way, so a psycopg
URL does not fail at connect on a keyword asyncpg alone accepts.
Rows already written are deliberately left alone. The inverse cast is
computable and DST-correct, but it cannot be applied safely: created_at
is assigned explicitly on some paths and defaulted on others, an install
that began on SQLite holds correct and shifted rows side by side, and
nothing distinguishes them after the fact. Timestamps are right from the
upgrade forward and history keeps the times it was given.
One related mismatch goes with it, because fixing the database side alone
would have made it start lying on exactly the installs this repairs. The
support package's oldest_pending_age_seconds subtracted a naive local
clock from a naive UTC column, with a comment claiming it was UTC; on the
reporter's install the two errors cancelled. It reported a job queued
five minutes ago as three hours old east of Greenwich and a negative age
west of it. The two AMS and printer-sensor retention cutoffs move to the
same utcnow_naive helper -- correct in value already, but deprecated in
3.12 and emitting warnings on every sweep.
* feat: add ScheduledDrying model for delayed drying runs (#2638)
* Release the printer when a scheduled dry ends (#2638)
_check_scheduled_dryings marks a printer as drying in _drying_in_progress,
which is shared with auto-drying. Auto-drying prunes that map in
_sync_drying_state(), but that call sits behind its enabled check, and this is
the first writer that runs whether auto-drying is on or not. With it off --
the default -- nothing dropped the entry short of a print being dispatched to
the same printer, so the next scheduled run parked on "already_drying"
forever and queue_drying_block held that printer's prints too. A nightly
off-peak dry with no printing in between is exactly the workflow this feature
is for: night one worked, every night after it silently did not. The check now
releases what it acquired, covering both a run that ends mid-pass and one
cancelled through the route between passes.
The retention prune ran on every pass. Issuing the DELETE is what opens a
write transaction, this method is called every 3s while the queue dispatches,
and rows only become prunable a week after they finish, so it is now gated to
hourly on a monotonic stamp -- with the first pass after a restart still
reaping whatever the dead process left behind.
Both drying paths now pick the blocking dry_sf_reason through one rule.
The immediate endpoint quoted whichever code the firmware listed first while
the scheduler prioritised power over retract, so one blocked AMS read two ways
depending on which button you pressed. drying_preflight.primary_reason_code
holds the order and both call it, including the flame button's tooltip, which
had no wording for filament at the outlet at all and sent those users to the
generic "can't start drying right now".
scheduled_drying joins the model list in init_db. The table was already
created -- importing the package registers it -- but it was the only model
relying on that indirection.
Tests: the release (completion and route-cancel), the prune throttle, the
shared reason rule, the tooltip priority, and four driving the real check_queue,
which nothing covered before -- a pass with no rows still dispatching prints, a
due row dispatching, and a failed row not stalling the queue behind it. Each
one fails against the code it guards.
---------
Co-authored-by: MartinNYHC <martin@bambuddy.cool>
Co-authored-by: maziggy <mz@v8w.de>
Every pipeline endpoint answered 403 for API keys whatever scopes the key
carried. PR A parked all three permissions on the admin denylist until the
run dispatch existed to decide about; it landed in PR C and the parking was
never revisited.
PIPELINES_READ now rides can_read_status. PIPELINES_RUN requires
can_queue AND can_manage_library together, so the allowlist gained tuple
values: a run slices into the library and then queues prints, and mapping
it to either flag alone would hand that flag the other one's authority.
The 403 names every flag the key is short of. PIPELINES_WRITE stays
admin-only -- a key can run the recipe, not rewrite it or clear the log.
Opening the run route also needed the cloud-owner fallback the direct
slice route makes: a pipeline can carry Bambu/Orca Cloud presets, and
resolving those reads a token off a user record that an API-keyed request
does not have. retry_failed forwards the new dependency explicitly,
since a direct call receives the Depends marker rather than None.
Two faults behind the same kind of print: one that arrives without a
retrievable 3MF, which on an H2S is any job started from the printer's
own internal library.
Such an archive has no file_path, and Path("").parent is Path("."), so
every site that derived the archive's folder from it landed on the data
directory itself. The finish-photo capture spotted that and wrote to
<archive_dir>/<id>/photos instead. Nothing else did. The photo was
written in one place and looked for in another: reads 404'd, deletes
dropped the name and left the file, and the notification attachment
never found the image. Hand-uploaded photos worked only because upload
and read agreed with each other rather than with the capture. Give the
question one owner in utils/archive_paths and have all four sites ask
it. Lookups check the old shared location too, so photos already
uploaded there stay reachable; uploads now go where captures go.
Separately, the remain%-delta fallback that stands in for a missing 3MF
can charge nothing for several reasons, and did so without a word. The
AMS reading is coarse and, on the reporter's printer, noisy: it rises
mid-print, swings five points over a job, sits at 100% through a
36-minute print on a fresh spool, and goes negative on a nearly empty
one -- which the start-of-print gate rejects, dropping the only slot
that was printing. Two of their prints went uncounted for two different
reasons and both read as "no spools updated", which is also what a print
with nothing to charge prints. Name the slot and the two readings in
each case, on the Spoolman path and on the internal-inventory path,
which has carried the same gates since #1119.
The Spoolman path also had no notion of which slots the print used, so a
spool swapped into an idle slot mid-print reads as consumption and is
billed to whoever that slot is assigned to -- the fault #1269 fixed for
the internal tracker, still open here, and likeliest on exactly the
prints this fallback serves, where nothing else narrows the field. Use
the same three pieces of evidence it does: the print's mapping, its
mid-print tray changes, and the tray it started on. The last needs
storing, because the internal tracker's row is deleted before this runs
and a screen-started print has no mapping to fall back on -- hence a new
nullable column, and no backfill, since a row from before it existed has
nothing to say. Where no evidence exists at all, every slot is still
considered.
Both paths also treated tray_now == 255 as naming a slot. It does not:
it is the field's initial value, the fallback for an unparseable
reading, and what it reports with nothing loaded. Mapped as a tray id it
becomes (255, 1), so as the only evidence it excluded every real slot
and charged nothing at all -- this issue's own bug, arriving by a new
route. On the internal path that is live today; on the Spoolman path it
would have shipped with the guard above. The external holder reports 254
when it is genuinely in use.
The arithmetic is untouched: at one percent per step this cannot resolve
a small print, and pretending otherwise would be worse than saying so.
The Vortek rack holds six hotends, and a multi-colour plate is sliced to
use a different one per colour so it can skip the purge. Which of the six
each colour takes is not in the 3MF. The same plate, sliced and sent twice
from Bambu Studio with a different choice each time, produces two files
that differ only in rounding in the last digit of a few extrusion figures
-- the filament grouping, the toolchange stream, the 120 nozzle-change
markers and project_settings.config are all identical. The choice travels
only in the dispatched nozzle_mapping.
Bambuddy had no way to state it, so those plates went out with no nozzle
assignment at all and the printer chose for itself. That is what levelled
on one hotend and printed with another, millimetres above the plate.
Every rack-bound filament now carries a position picker beside its AMS
slot dropdown, listing all six with the nozzle each holds. An empty
position, or one holding the wrong diameter or flow type, is shown greyed
out with the reason rather than hidden, so someone looking for position 4
finds it. The choice is per filament *group* rather than per slot, because
a group is one hotend: two filaments the slicer grouped together share it
and cannot point at different positions.
Nothing has to be picked. Positions are assigned automatically, preferring
one already loaded with that colour, which on the plate this was built
against reproduces Bambu Studio's own pick exactly.
A nozzle currently picked up onto the carriage is offered too. The
firmware drops its rack position from the report entirely rather than
sending a placeholder (#943), and refusing it would rule out the position
most likely to be wanted -- the one the last print left mounted. Only
recoverable when exactly one position is missing; two gaps are genuinely
ambiguous and stay unavailable.
Positions are re-checked at dispatch, not just when queued, because the
rack can be re-loaded in between. The two failure modes differ on purpose:
an explicitly chosen position that no longer fits stops the print, names
what the position now holds, and deletes the uploaded file from the SD
card so it cannot be started by hand either -- an operator who named a
hotend must not silently get a different one. An automatic assignment that
cannot be made instead falls back to letting the firmware choose, which is
what happened before any of this existed.
The pick is stored as {group: position} rather than as the expanded
nozzle_mapping, though that is what goes on the wire. That column means
"Bambu Studio decided, forward verbatim", and only the group-and-position
form can be re-checked against what is actually mounted at dispatch.
The existing multi-rack refusal in extract_nozzle_mapping_from_3mf stays.
It still guards the #2800 fallback, which can only ever name one rack id.
Measured on the maintainer's H2C: rack position n is physical nozzle id
15 + n, confirmed by cross-referencing two captured dispatches against
Bambu Studio's own dialog. extruder_max_nozzle_count names which carriage
is the rack straight from the file, and is read rather than assumed -- a
fourth independent confirmation of the carriage indices fixed in 45dc139.
The print dialog is also wider, on every printer. Its filament rows carry
the most horizontal content in it and adding a picker truncated names to
"Bamb...". The column widths themselves only change on a rack machine.
Tests: 44 unit covering the plan, the resolver, the mounted-nozzle
recovery and every refusal; 9 dispatch integration asserting the two real
captures end to end; 7 API round-trip; 33 frontend. The API ones exist
because two integration bugs got through a green suite that tested the
pieces and not the seams -- the group data reached only one of the three
filament-requirements routes, and the field was declared on every schema
except the create one, where Pydantic dropped it in silence.
Everything the completion path needs to split a print's filament across
the trays it fed from lived only in memory: the dispatched plate and
slot-to-tray mapping, the spool-assignment snapshot, and the tray-change
log. A print that outlived a restart lost all of it and fell back to
what the printer reports at completion -- which, with AMS Filament
Backup on, is the substitute tray. The whole print was charged to the
spool that only finished it while the spool that ran dry was charged
nothing.
Persist that context in a new active_print_sessions row, append tray
changes as they happen, and restore both the session and the printer's
tray-change log at restart recovery. Seed the log from the current tray
when there is nothing to restore, since last_loaded_tray advances even
when no change is logged.
Rank the queue item's stored ams_mapping above the printer's live
mapping field, which is what backup rewrites. Recover plate_id from the
archive or queue item, and give extract_layer_filament_usage_from_3mf a
plate_id instead of taking the first .gcode member -- a Bambu Studio
export stores plate 2 first, so per-layer figures were measured against
the wrong plate for both inventory backends.
Stop auto-unlinking a spool assignment when its slot reports empty
during a running print. At a runout the spool is still in the AMS, and
dropping the link leaves the completion path nothing to charge.
Capture the print-start context for both inventory backends. Spoolman's
own durable row (#1820) carries its plate-scoped figures and dispatched
mapping but not the tray-change log, and its slot assignments -- the
way. Registration in _active_sessions stays gated, since on_ams_change
reads it to decide whether to skip the remain%-based weight sync (#880).
The line carried "# noqa: S608", which is ruff's flake8-bandit code -- but S is
not in ruff's select list in pyproject.toml, so ruff never ran that rule and the
marker suppressed nothing. Bandit itself only honours "# nosec", so the query
went on being reported as B608 while the line read as already handled.
The finding is a false positive. The only interpolated fragments are source_expr
and model_expr, assigned just above from a two-branch is_sqlite() check where
both branches are string literals; no caller value reaches the string. They are
JSON expressions rather than values, so a bind parameter cannot express them.
Replaces the inert marker with "# nosec B608", matching the convention already
used across the test suite, and moves the reasoning into a comment above the
statement. Bandit's medium+ count drops to 16, none of them B608.
Archives, the queue and statistics report ownership as a numeric
created_by_id, and statistics accept it as a filter, but nothing let an
API key discover whose id was whose -- the only user listing returns
emails, roles, group membership and full permission sets, so it is
administrative and rejects keys.
Add GET /users/slim returning id + username only, gated on a new
users:read_slim permission mapped to can_read_status. That grants no
data a key could not already reach: for API-keyed requests the
permission deps return None as current_user, so the stats:filter_by_user
guard short-circuits and ?created_by_id=N is already honoured for every
N. What was missing was the ability to address the filter, not
permission to use it. The full listing stays unmapped = admin-only.
Also fix /auth/me, which answered an API key with a synthetic
administrator: id 0, role admin, is_admin true and every permission in
the enum. A key cannot reach an administrative route at all, so clients
building their UI from that response rendered actions that 403 on use.
It now reports the key owner's identity, is_admin false, and the
permissions the key's scopes actually admit. Ownerless legacy keys keep
id 0 but no longer claim admin.
---
Source user names from the slim listing where only names are needed (#1894)
Stats filter-by-user, the Archives print log filter, the File Manager
username autocomplete, the camera-token owner column and the Finance
member picker all render nothing but a username, but all of them read
the full user listing, which is gated on the admin-level users:read.
An operator granted stats:filter_by_user but not users:read got an
empty filter with no indication why.
Point them at /users/slim under a separate react-query key, since the
full listing shares the 'users' key and the two shapes would clobber
each other in the cache.
An H2D armed five 12-hour drying cycles inside four hours, one of them six
seconds after the previous one ended, and none ran more than a couple of
hours.
Two things combine. The firmware ends a cycle when it decides the filament
is dry rather than when the clock runs out, and reports no fault doing it --
across this printer's history the run length tracks how wet the spools were,
from nearly the full 12 hours starting at 32% down to minutes once the unit
sat at 10-13%. That part is the AMS doing its job.
The loop is ours. An AMS reports higher relative humidity while it is warm
than once it has cooled: the same unit read 10-13% cold and 15-20% through
every cycle. With the threshold at 14% the reading at the moment a cycle
ended was always still above it, so the next 30-second pass armed another
12-hour cycle. Nothing counted, nothing waited, and it only stopped when the
box finally cooled enough to read 13%.
Auto-drying now waits 30 minutes after a cycle ends before arming another on
the same unit, and gives up on a unit after two consecutive cycles that
bring the reading no lower -- logging why and sending a new notification,
on by default because it reports that Bambuddy has stopped acting. Progress
is judged against the lowest reading any cycle on that unit has ended at,
not against the threshold, so a genuinely wet spool in a humid room coming
down 40-37-35 keeps drying however far it still is from the target;
comparing against the best so far rather than the previous end stops a
sensor wobbling by one point reading as progress every other cycle. The
suspension lifts by itself once the reading falls below the threshold.
Neither guard can stop a running cycle, and a cycle Bambuddy cut short for a
print, or that the user stopped by hand, is not counted against the unit --
so a farm that dries between queue jobs is unaffected. The threshold field
now warns below 20%, and every cycle end logs the unit's temperature and
humidity, which is what made this diagnosable.
The same bundle showed unrelated tasks failing with "database is locked",
each inside a 30.000-second Discord connect timeout. Alarms are raised from
inside the loop that records sensor history, at a point where the new rows
are added but not committed; the first read in the notification path flushed
them to satisfy itself, opening a write transaction, and the provider was
then contacted over the network with that transaction still open. SQLite
allows one writer and 30 seconds outlives the 15-second busy timeout, so
every other write in that window failed. The two reads that run before a
provider is contacted no longer flush the caller's pending work, and the
connect timeout is 5 seconds rather than 30 -- the body keeps the full 30,
so image uploads on a slow uplink are unaffected. SQLite only; Postgres has
no single-writer limit.
both dialects rather than one.
Postgres upgrades never got as far as the finance schema.
database.py added on_billing_charge_failed with BOOLEAN DEFAULT 1. The 1 is a
SQLite-ism; Postgres answers DatatypeMismatchError, and _safe_execute
deliberately re-raises anything that is not an idempotency error, so
run_migrations died there and rolled the whole transaction back. No finance
tables, no columns, and the app does not start. Six lines above, the same
change gets is_voided right with an is_sqlite() branch, so this was an
oversight rather than a decision. Now branched the same way.
This also explains the four test_security.py::TestBackupKeyFiles failures
reporting "column print_archives.cost_center_id does not exist". That column's
migration exists and works -- it simply never ran, because every startup
aborted before committing. Reproduced against Postgres 16 by building a
pre-billing schema from dev and upgrading over it: fails without this,
completes with it, and re-running the migrations or starting from an empty
database are both clean.
test_billing_run_id_migration.py failed on any Postgres-configured checkout.
It builds its own SQLite engine, but run_migrations branches on the global
dialect rather than the connection in hand, so on a box whose DATABASE_URL
points at Postgres it emitted md5(random()::text) and btrim() into SQLite.
Given the same fixture test_ldap_migration.py already carries for exactly this
reason. The suite now agrees across dialects -- 9190 passed either way, where
it used to be 9184 on one and 9183 on the other.
The kill switch could not tell a print Bambuddy started from one it merely
watched.
Authorization fell back to a print_archives row in status="printing" matched on
subtask_id. But on_print_start archives every print it observes, including ones
started from Bambu Studio or Handy, and stamps them with the same status and
subtask_id -- the code says as much where it notes "a print Bambuddy didn't
dispatch". So a foreign print became authorized the moment its 3MF finished
downloading, and _active_prints was rehydrated from it, making that permanent.
The switch fired only inside the download race, and never afterwards. Neither
test caught it: one stubs the authorization call to False, the other stubs the
query to return an archive, so the real lookup was never exercised against a
foreign print.
Authorization now requires a marker Bambuddy writes itself: billing_run_id,
minted per dispatch in the scheduler, or created_by_id carried over from the
queue item. Failing that, it looks for a queue row in status="printing" on that
printer -- committed before the MQTT send, and the only durable trace a
library-file dispatch leaves, since those have no archive at send time and the
row created for them moments later carries neither marker. That row cannot be
tied to a subtask_id, so it defers rather than authorizes.
Deferring also closes a false positive the previous version shared: a restart
in the window between the send and the download left no archive at all, and a
Bambuddy print was stopped as unauthorized. Stopping a print is irreversible
and declining to act costs a log line, so ambiguity resolves that way.
Tests cover an unmarked archive not being authorization and not entering
_active_prints, either marker alone authorizing and rehydrating the fast path
without touching the queue, an unmarked archive with a live dispatch deferring,
a dispatch not yet archived deferring, and nothing at all being unauthorized.
Binds binary_sensor and reading-carrying sensor entities to a printer and
renders their state on its card, worded by Home Assistant's device_class.
Optional per-sensor alert condition drives a notification on the transition
into the alert state and an opt-in interlock that holds queued prints while
alerting -- a hold with a readable waiting_reason, never a failure, and only
ever on a sensor that was read successfully.
Sensors get their own table rather than a wider entity pattern on SmartPlug:
get_smart_plug_by_printer would otherwise hand the card's power button a door
contact to switch.
The hold is passed to the model matcher directly rather than merged into
busy_printers: _check_auto_drying reads that set as "is currently printing"
and would put an idle-but-held printer down the mid-print drying path.
The notification_providers migration spells its default FALSE, not 0 --
Postgres rejects an integer default for a boolean and _safe_execute swallows
the error.
Printing a multi-plate file in different quantities per plate meant
queueing each plate separately and tracking the counts by hand: one
shared Quantity field cannot say "plate 1 once, plate 2 twice, plate 3
three times". Each selected plate now carries its own quantity, and the
submission becomes an order on a new Batches tab.
The point is the distinction the old flat batch could not express.
print_batch_plates stores how many runs of each plate were wanted,
separately from what was queued, so a run that fails, is cancelled or is
skipped does not satisfy a target -- the order goes on saying it owes a
print instead of quietly under-delivering. Queue remaining re-queues
exactly what is missing, for the whole order or one plate, by cloning
the most recent item for that plate: that inherits the printer or model
target, AMS mapping, filament overrides and print options along with the
validation they already passed, rather than re-serialising twenty fields
through a template that would drift from the model the first time
someone adds a column. Clones append to the end of the relevant
printer's queue and take the same advisory lock the add-to-queue route
does; positions are per-printer sequences, not global.
Cost is measured, not estimated. print_log_entries gains queue_item_id,
set where the queue item is already in scope, so each run's material and
energy are attributed through the item that produced them -- an
unrelated reprint of the same archive never lands in an order's total,
and a multi-plate order gets each plate's own cost rather than the whole
file's via the plate-scoped estimate from #2614. Before any run has
completed there is no honest figure, so cost reads as unknown instead of
a fabricated 0.00.
The Batches tab wires up GET /queue/batches, which has been unreferenced
since the batch MVP shipped, along with six locale keys that were
translated and never used. It is a separate tab because an order
outlives the queue that produced it: once its runs finish they leave the
active queue, so Queue and History each hold half the picture.
completed was not a reachable status before now, so every batch created
since April is still marked active however long ago its last print
finished -- 73 of them on the development install. A startup pass closes
out the finished ones: those whose runs all completed become completed,
and groupings whose items were all cancelled become cancelled, which is
what they are. Not applied to orders, which state their intent
independently of their runs and still owe the work. Only batches with
nothing queued or printing are considered, and repeating the pass also
catches an order whose last run landed while the process was down.
Batches with neither items nor targets are no longer listed at all --
empty shells left when a grouping's items went with their source
archive.
Dispatch applies the same source-file gates as POST /queue/. It creates
queue items, so without them it would be a weaker door to the same
outcome; the archive and library-file checks move into shared helpers
so a third route cannot drift from them.
Adds /library/variant-groups for declaring that several sliced files are
the same job for different printers, and a variants payload on queue
creation that turns such a set into one queue item with a candidate per
file.
The candidate set is validated as a set: one file per printer model, each
file sliced for the model it is offered as, and at least one model with
an active printer. A cross-model item deliberately holds no file of its
own, because print_queue.library_file_id is ON DELETE CASCADE and would
destroy the whole job when a single alternative is deleted.
Fixes internal printer-model codes never being resolved on queue create
and update: normalize_printer_model returns unknown input unchanged, so
the or-chain never reached the code map and a "C13" target matched no
printer and waited forever.
Skips candidates whose file is trashed or missing. Library deletes are
soft, and SQLite runs with PRAGMA foreign_keys off, so neither case is
covered by the schema; the hard-delete paths now also drop the rows.
Adds library_files.variant_target_model so a user can say which printer
a file without slicer metadata is for, kept out of file_metadata so the
assertion is never mistaken for parsed data.
Adds file_variant_groups plus variant_group_id / variant_position on
library_files, so a set of files that are the same job sliced for
different printers can be resolved to whichever printer frees up first.
Backfills groups from the sliced_from_library_file_id provenance that
slice_and_persist and the pipeline runner have been writing into
file_metadata since they shipped, and which nothing has ever read.
Only sources with two or more children carrying distinct
sliced_for_model values are grouped: a single candidate is not a
choice, and two slices for the same printer give the resolver no basis
to prefer one.
unbounded repetition the match was quadratic in the subject length: on a run
of scheme-legal characters the engine restarted at every offset and consumed
to the end before failing to find "://". ffmpeg echoes the configured camera
URL into its stderr and the whole blob reaches the pattern before any
truncation, so the subject length is attacker-influenced.
A Kubernetes Secret written as a block scalar carries a trailing newline, and
the schema bounds the four required variables by max_length only, so an
unstripped issuer_url was stored and enabled and then raised httpx.InvalidURL
on the first click of the SSO button -- the authorize-time failure the
all-or-nothing rule exists to prevent. Whitespace-only values got through the
same way, contradicting the reader's own "an empty required var counts as
unset". The optional variables have always treated blank as unset; the
required ones now do too.
Also registers BAMBUDDY_LOCAL_LOGIN (#1589) in the typo guard, which logged
"possible typo" for it on every boot while listing every BAMBUDDY_OIDC_*
variable as legitimate.
Promoting env_bool to strict rejection made BAMBUDDY_LOCAL_LOGIN=on raise
EnvOIDCConfigError uncaught on the login/forgot-password path -- a 500 on
the exact recovery endpoint the bypass exists to keep open. env_bool gains
a strict flag (default True for the startup OIDC reader); the local-login
caller opts out so an unrecognized value falls back to "off" instead.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016q8EAf9Rj7ZHL92sPnXYxy
_env_bool returned the default for anything outside {true,1,yes}, so
BAMBUDDY_OIDC_REQUIRE_EMAIL_VERIFIED=on silently read as OFF and
BAMBUDDY_OIDC_ENABLED=on silently disabled the provider -- the exact
opposite of what .env.example claimed. Unrecognized values now raise
EnvOIDCConfigError, caught in _apply_env_oidc_provider the same way a
bad DEFAULT_GROUP or a ValidationError already is: logged and left
running, never released on a typo.
Also promotes _env_bool to env_bool now that it has a call site in
auth.py, and corrects the boolean-parsing sentence in .env.example.
BAMBUDDY_OIDC_SCOPES, _EMAIL_CLAIM and _ICON_URL fell back to their
default only when the key was absent, so `BAMBUDDY_OIDC_ICON_URL=` in
a compose file (as .env.example ships it, commented) reached the
schema validator as an empty string and got the whole provider
refused. default_group already treated blank as unset; these three
now follow the same rule.
The recovery rollback after a failed commit was itself unguarded, so a
rollback that raises on a wedged connection would still take the boot
down -- the exact failure the never-raise contract exists to prevent.
Suppressed; the caller's `async with` discards the session regardless.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016q8EAf9Rj7ZHL92sPnXYxy
A name collision with a provider that was NOT already env-managed
overwrites its issuer, client id and secret in place and locks it
behind the env-managed 409 -- but it logged the same routine "applied"
line as an ordinary re-apply, giving no signal a UI provider was just
taken over. Adoption is now a WARNING with its own wording; a routine
re-apply of an already env-managed provider keeps the INFO line.
The db.execute/db.commit calls in the upsert and release paths sat
outside the try/except that only wrapped OIDCProviderCreate, so a
commit failure at startup (connection blip, WAL lock) propagated out
of the lifespan and took the instance down -- the exact outcome this
module exists to avoid. The body now runs inside a private
_apply_env_oidc_provider(), with the public entry point catching,
logging and rolling back on any exception.
Lets a reprint reuse the AMS slot the slicer itself picked, instead of
re-deriving one from the file's static type/color.
When a Print Queue VP has "Save AMS mapping" on, the slicer's own
live-resolved ams_mapping (from the project_file MQTT command) is
persisted onto the archive as extra_data.slicer_ams_mapping. A later
reprint can reuse it via a new "Mapping" button in the filament-mapping
panel — one click snaps every slot to the saved pick, click again
reverts to auto-match. Archive cards and queue rows get an "AMS mapping
saved" badge so it's visible beforehand. add_to_queue also falls back
to the saved mapping automatically when the caller sends no explicit
ams_mapping (e.g. a plain reprint with no per-slot edits).
The queue item's own ams_mapping (used for that dispatch) is still
captured unconditionally whenever the slicer provides it — that part is
a correctness fix, not gated behind the toggle. Only the archive
persistence for future reprints is opt-in.
Split out from the original combined PR per review: this half is
genuinely opt-in and low-risk (#2684). The dispatch-time validation
gate that keeps a stored mapping honest (#1308) changes behaviour for
every existing user and will land as its own PR.
Review fixes applied:
- _extract_slicer_ams_mapping_json: dropped the unreachable `v is None`
arm and rejected bool explicitly (isinstance(v, int) accepts bool).
- Translated the Russian docstring text to English.
- save_ams_mapping's model comment moved to a trailing comment on the
column line, matching the file's convention.
- usingArchiveMapping now resets when the plate or archive changes, so
the Mapping button can't read ON against a mapping it never applied.
- Translated "Click to change slot assignment" and "Re-read".
- add_to_queue's fallback is now called out explicitly in code comments
and covered by three new integration tests (fallback fires, explicit
mapping wins, unrelated extra_data doesn't false-trigger).
Closes#2684
The two-flagged-row state was never released, so no installation can carry it
in -- the reason to release every flagged row is not repair, it is that the
sweep's invariant is not enforced anywhere and scalar_one_or_none() turns a
broken one into a failed boot. Comment, test docstring and test name say that
instead.
Without it every account auto-created through the env provider fell back to
Viewers (routes/mfa.py), and because the provider is locked the UI could not
correct it either -- a real limitation for a declarative deployment running
BAMBUDDY_OIDC_AUTO_CREATE_USERS=true.
BAMBUDDY_OIDC_DEFAULT_GROUP names a group rather than an id: ids are handed out
per installation, so the same compose file would point at a different group on
the next deployment. The name is matched exactly, resolved against the database
before anything is written, and default_group_id joins _APPLIED_FIELDS so
dropping the variable clears the group again -- the environment is the whole
truth for this row.
A name that matches no group is refused rather than defaulted: silently landing
users in Viewers is the failure this variable exists to remove, and the API
already answers 422 for a default_group_id that does not exist. The refusal is
logged and survivable, and it says which of the two cases happened, because
they differ sharply -- an existing provider keeps running on its last good
config, while on a first boot nothing is created and no SSO button appears.
Raised by maziggy in review of #2625 as a scope decision; documented in
.env.example and in the companion wiki PR.
apply_env_oidc_provider matches the provider by BAMBUDDY_OIDC_NAME but never
cleared is_env_managed from the row it managed previously. Renaming the
variable therefore left two flagged rows, and both consequences are reachable
by ordinary config edits: the old row stayed enabled with a stale issuer and
secret on the login page while _refuse_if_env_managed answered 409 to every
attempt to edit, disable or delete it -- the dead end reachable only through
the database that the release path exists to prevent -- and unsetting the
variables later hit scalar_one_or_none() on two rows, so MultipleResultsFound
propagated out of the lifespan and the app stopped booting.
The upsert now sweeps the flag off every other row, the same shape the
autologin sweep one block down already uses: disable and release rather than
delete, for the same cascade reason as everywhere else in this branch. The
release path releases every flagged row it finds instead of exactly one -- the
sweep should keep that at one, but a release path that dies with
MultipleResultsFound the moment that invariant breaks is a second way to lose
the boot, and the query costs the same either way.
Releasing now clears is_autologin as well. Without it a released row keeps a
latent autologin claim: update_oidc_provider only re-runs the exclusivity sweep
when a request sets is_autologin=True, so merely re-enabling the row in the UI
would silently make it the autologin target again.
Reported by maziggy in review of #2625, with the rename reproduction.
feat(db): warn when the connection pool can outgrow the PostgreSQL server
fix(mqtt): an unusable layer_num must not drop the printer connection
test: patch settings.base_dir via monkeypatch so it unwinds on error
test: restore the config module after reloading it
Subprocess output and user-supplied URLs are scrubbed of credentials
before they reach the application log. Adds a shared redaction helper in
core/logging_filters and routes the existing support-bundle sanitizer
through the same pattern.
Timelapse was on, the video never reached the archive, and Scan for Timelapse
found nothing afterwards. Across 247 support bundles this was the norm, not an
edge case: 457 automatic scans scheduled, 262 attached.
The scan looked four times over ~65s. The attempt that found the video was #1
272 times, then 17 / 13 / 13 — flat against the cutoff, not decaying, i.e.
files were still arriving when we stopped. What ran afterwards searched for the
print name inside the filename; Bambu only writes "video_<timestamp>", so it
fired 159 times and matched zero.
The manual Scan had no baseline at all and matched on filename timestamp, FTP
mtime, or "there is only one video" — all reading a clock a LAN-only printer
cannot sync. The reporter's P1S was six and a half days out.
- Poll for minutes instead of ~65s; drop the name-match fallback.
- Persist the print-start baseline on the archive, so the diff survives a
restart mid-print and the manual Scan runs the same comparison. With a
baseline present the clock-based strategies are skipped entirely — they can
only turn an honest "pick one" into a confident wrong answer.
- When several files are new (a previous print's video landing late), exclude
the ones already attached to another archive instead of ordering the
candidates. Ordering could only be done on the printer's clock.
- Delete the video from the printer once archived. Keeps /timelapse to
unclaimed files, which is what makes the diff unambiguous, and stops P1S
cards filling with AVIs.
- Gate that delete on a verified transfer: download_file now compares against
the size from the listing. An FTPS connection closing early does not always
raise, so a partial buffer was being attached as a complete video — which
would also have been the one case where deleting the source lost data.
Bounded twice on purpose: wall-clock deadline plus a derived round cap, since
the deadline stops bounding the loop as soon as the sleeps are shortened.
Per-round logging only speaks when the listing changed — 31 rounds of full
listings would bury the interesting line in the support bundle.
Migration adds print_archives.timelapse_baseline as JSON, spelled the same on
both dialects so a migrated database matches a fresh one.
-----------
fix(finish-photo): add the timelapse frame to the archive after the notification (#2704)
When a print records a timelapse, its last frame is the better finish photo:
the firmware stops recording with the toolhead parked and before the end
G-code drops the bed, where a live grab at that moment catches a lowered
plate. Bambuddy waited 60s for the video and then gave up, because the
print-complete notification blocks on that photo and holding a notification
for minutes is worse than sending it with the live grab.
P1-series printers write MJPEG AVI rather than H.264 MP4 and serve it slowly.
Measured over 261 attaches in the support bundles: P1S median 33s, p90 167s,
worst 546s, while every other model finished inside 26s. So the printers that
most needed the better framing were the ones that never got it.
Keep the notification on the same bound, and keep waiting off to the side.
_capture_finish_photo_from_timelapse now reports whether it ran out of time or
concluded — a video that landed and failed extraction is not worth retrying,
one that never arrived is. On the first, schedule a background task that waits
up to 15 minutes and inserts the extracted frame at the front of the archive's
photo list, where the gallery opens.
The live grab stays on disk: the notification already links to that exact
file, so removing it would leave a broken image in Discord or Telegram.
The length check proves we received what the listing said, not that the file
was finished. The first look happens ~5s after the print ends, while the
printer may still be writing, so a growing file can be listed short, served
short, and pass. Re-list after the download and only accept the video once its
size has stopped changing — a failed re-list counts as not settled, since
"could not check" must not mean "safe to delete".