Backend Tests (shard 1/4) timed out after 10 minutes on e8b901f54
("Updated BACKERS"), a docs-only commit. The step was not hung: it
reached 98%, every test passing, and was killed about 12 seconds short
of finishing.
pytest-split balances by duration only when it has a durations file,
and there was none. Without one it splits by test COUNT -- 2926 / 2926
/ 2926 / 2925, exactly 25% each, which is what the comment here claimed
was good enough. Count is not time. Measured over the full suite
(11703 tests, 1045.6s):
shard 1 2926 tests 658.2s
shard 2 2926 tests 173.1s
shard 3 2926 tests 113.3s
shard 4 2925 tests 101.1s
Shard 1 was carrying 63% of the suite's runtime -- a 6.5x spread -- and
had been walking toward the cap for weeks: 475s, 404s, 445s, then 587s
on 08-30, thirteen seconds under it, and 613s here. CI printed the
reason in its own log on every run: "[pytest-split] No test durations
found."
Commit tests/.test_durations and pass --durations-path explicitly:
shard 1 1273 tests 261.6s
shard 2 1070 tests 261.5s
shard 3 1522 tests 262.0s
shard 4 7838 tests 260.6s
A 1.01x spread. The lopsided test counts are the point: shard 4 takes
thousands of fast unit tests, shard 1 keeps the slow integration ones.
The four groups still partition the suite exactly -- union is 11703
with nothing dropped or duplicated, and all four pass.
--durations-path has to be explicit because pytest-split defaults it to
$CWD/.test_durations, and the two matrices run from different
directories: the native job from backend/, the Docker job from the
image root. Left implicit, the Docker one finds nothing and silently
falls back to the count split. Node IDs match across both because
backend/tests/pytest.ini pins rootdir to backend/tests either way, so a
single file serves them both; verified the file clears .dockerignore
and lands at /app/backend/tests/.test_durations at full size.
timeout-minutes 10 -> 15 for headroom, since the file goes stale as
tests are added. Staleness degrades slowly rather than breaking --
unknown tests are treated as average.
The SignPath Foundation production certificate does not sign on demand
the way the self-signed test certificate does. Every request has to be
approved by hand in the SignPath UI, because the Foundation verifies
what is being signed and which build produced it. The submitting action
waits for that approval with a default timeout of 600 seconds, which is
ample while the test policy approves in seconds and far too short once
the wait is a person noticing a tag went out. A tag pushed at night
would have failed the run ten minutes later with the installer already
compiled and thrown away.
The compile now ends in its own job, which uploads the unsigned artifact
and exposes its id. A second job downloads it, signs it, and does the
release-facing work, with the wait raised to an hour and the job timeout
sized to sit outside it. Separating them is what buys the recovery: the
artifact is uploaded before the wait begins and is addressed by id, so a
missed approval window costs a re-run of the second job alone rather
than a rebuild. Raising the timeout in place would not have given that.
The second job runs for unsigned builds too. Daily prereleases are
deliberately left unsigned to preserve the signing quota, and gating the
whole job on the signing decision would have meant a second copy of the
alias, artifact and release steps for them to run through.
The decision itself moves into a named step that echoes it, so a tag
that came out unsigned can be explained from the run log rather than by
re-reading the expression. It is one source of truth feeding both jobs,
which a job-level env could not be.
Every step body is otherwise unchanged. The property worth keeping is
that none of the alias, upload and release-attach steps carry always(),
so GitHub skips all three when signing fails or times out and an
unsigned .exe cannot reach a release; that is now written next to them,
because it is easy to break by adding a condition without noticing.
The policy slug stays at test-signing and the signature check stays
lenient -- the test certificate is self-signed and reports UnknownError,
so requiring Valid would fail every run until the production certificate
is imported. Both are the cutover. The restructure behaves identically
under the test policy, the request simply completing at once instead of
waiting, so it can be proven green beforehand.
The scan pinned Trivy v0.69.1, which aquasecurity have since deleted --
retained releases now run v0.74.0 down to v0.69.2 and then jump back to
v0.26.0. The tag survives, so setup-trivy resolves it, reports "found
version: 0.69.1" and then exits 1 with no asset to fetch.
This repository did not notice because the binary was coming back from
the Actions cache on every run, which skips the download. Forks have no
such cache, which is where it was reported from -- and the same failure
was due here the first time that entry went cold.
Both scans move to trivy-action v0.36.0 and Trivy v0.74.0; every input
they pass is still declared in the new action. The comment records that
this pin has to be bumped rather than left, and that a green run is not
evidence it still resolves.
The config scan is clean on v0.74.0, so the bump adds no new
misconfiguration alerts.
Frontend:
- react-router/-dom 7.18.1 -> 7.18.2. The RSC-mode CSRF advisory was carried
as a documented exception in the audit gate because its only fix was the
8.3.0 major; upstream backported it, so the exemption lapsed on its own --
an entry only holds while fixAvailable.isSemVerMajor is true. The allowlist
is now empty; the machinery stays for the next one.
- dompurify 3.4.12 -> 3.4.13. Ships in the app, but the path is unreachable:
no hooks registered, IN_PLACE never used.
- js-yaml override ^4.3.0 -> ^5.2.3 (fix not backported below 5.x, so a
major) and nanoid override ^3.3.18. Both dev-only, via eslint and postcss.
eslintrc calls only load(), on the legacy .eslintrc.yml path this repo does
not use; eslint, vite build and 2861 frontend tests pass on it.
Backend:
- cryptography >=48.0.1 -> >=50.0.0, aiohttp >=3.14.0 -> >=3.14.3, pyopenssl
>=26.3.0 -> >=26.4.0. CI resolves from scratch and was already installing
the fixed releases; the floors cover the case CI does not, an existing venv
where >= is satisfied and `pip install -r` upgrades nothing. pyOpenSSL has
to move with cryptography -- each release caps it to a narrow window, so a
stale pyOpenSSL pins cryptography below its own fix line.
- postcss 8.5.15 -> 8.5.23 (GHSA-r28c-9q8g-f849, source-map path traversal)
- brace-expansion override ^5.0.7 -> ^5.0.8 (GHSA-mh99-v99m-4gvg, DoS)
react-router: pin react-router-dom to exact 7.18.1 (direct dep) and react-router
to 7.18.1 via overrides (transitive). 7.18.1 is the most-patched 7.x -- it clears
14 advisories that older 7.x releases carry, several reachable from a SPA (open-
redirect XSS in Link/useNavigate, route-matching DoS). The one remaining advisory,
GHSA-qwww-vcr4-c8h2, is RSC-mode-only; Bambuddy is a Vite SPA using BrowserRouter
with no RSC runtime (@react-router/server not installed), so the path is
unreachable. The only version that fully clears npm audit is the 8.3.0 major
(no react-router-dom 8.x exists; it needs migrating 50 import sites plus a React
peer bump), deferred as its own change.
Because a version pin can't stop npm from reporting the theoretical 7.11.0
downgrade as fixAvailable, the ci.yml (hard) and security.yml (nightly issue)
audit gates gain a narrow, documented allowlist keyed on the GHSA id. It resolves
the react-router-dom -> react-router advisory chain and stays fail-closed: a
different advisory on react-router still fails the gate, and an isSemVerMajor
guard drops the exemption the moment a non-major fix ships, forcing us to take it.
The GitHub runner's Python toolcache ships setuptools 79.0.1, which
pip-audit flags for PYSEC-2026-3447 (fixed in 83.0.0), failing the
blocking Backend Security job. A fix version exists, so upgrade
setuptools in the install step rather than --ignore-vuln. Applied to
both ci.yml (blocking) and security.yml (scheduled scan).
GHCR exposes total + 30-day daily-pull counts only in the package page
HTML (no REST or GraphQL endpoint). jgehrcke/github-repo-stats has no
notion of container metrics, so post-process the report after it runs.
New: .github/scripts/ghcr_inject.py
- Scrapes Total downloads (exact integer from title="N", not the K-rounded
display) and the 30-day sparkline (rect data-merge-count, data-date).
- Merges per-day rows into maziggy/bambuddy/ghcr-pulls.csv on gh-pages.
Fresh window overwrites overlapping dates, so GitHub's late revisions
to the last 30 days self-correct; days older than 30 stay frozen at
whatever was captured while still in-window.
- Patches latest-report/report.html: adds a TOC entry, a Container
pulls (ghcr.io) section at the top, and a Vega-Lite line+point chart
whose theme/config is cloned from the existing Total clones chart so
it inherits the report's look-and-feel.
- Bracketed by HTML-comment markers so re-runs replace rather than
stack (jgehrcke regenerates report.html every tick; we re-inject).
- Hard-fails if either scrape pattern stops matching — silent fallbacks
would let the chart freeze without notice.
Workflow: after run-ghrs, checkout source + gh-pages, run the injector,
commit only if the diff is non-empty. Uses the existing contents: write
permission; no new secrets.
The opaque "exit code 1" from actions/checkout was actually masking a
"Remote branch gh-pages not found in upstream origin" — the bambuddy
repo doesn't have a gh-pages branch. jgehrcke/github-repo-stats writes
to a branch called `github-repo-stats` by default, and Pages is wired
to serve from that branch. The path under it (maziggy/bambuddy/latest-
report/report.html) is unchanged.
Switching the clone branch and renaming the local path from `gh-pages/`
to `data/` so the workflow reads correctly. The PAT-in-URL pattern and
the direct git clone (vs actions/checkout) stay — both still wanted so
the real git stderr reaches the log if anything goes wrong.
Verified by cloning `github-repo-stats` locally and running the
injector against its live report.html: clean patch, 30 days extracted,
all anchor markers (TOC, section, script) present.
The PAT-token swap didn't fix the gh-pages fetch — same opaque "exit
code 1" from actions/checkout@v4 on both attempts of the retry loop,
with no underlying git stderr surfaced. Likely the wildcard-refspec
+ shallow fetch pattern v4 uses combined with something at the runner
side, but the action's swallowed errors make it untriagable from logs.
Replacing the gh-pages checkout with `git clone --branch gh-pages
--depth 1` using the same PAT, embedded in the URL. Direct, explicit,
and if anything goes wrong the real git error reaches the log instead
of "exit code 1". The recorded remote keeps the PAT, so the later
`git push` reuses it — no separate auth setup needed.
Identity config (user.name / user.email) moved into the clone step so
it lives on the freshly cloned repo; dropped the now-duplicate config
calls from commit-and-push.
Source checkout still uses actions/checkout@v4 since it never had a
problem (the default ref is the workflow's own commit, no wildcard
refspec required).
The default GITHUB_TOKEN failed at `git fetch` for the gh-pages
checkout step with opaque "exit code 1" and no surfaced git stderr,
even with permissions: contents: write set at workflow level.
actions/checkout's retry loop didn't recover.
Switching the gh-pages checkout to secrets.GHRS_GITHUB_API_TOKEN —
the same PAT jgehrcke/github-repo-stats already writes the branch
with one step earlier — keeps the auth chain uniform and avoids the
mismatch. persist-credentials defaults to true, so the subsequent
commit-and-push step in the same gh-pages directory picks up the
PAT automatically; no separate change to the push step needed.
fetch-depth: 1 left explicit because checkout@v4 defaults to it but
the value's load-bearing for this workflow (we only need HEAD of
gh-pages, not history).
GHCR exposes total + 30-day daily-pull counts only in the package page
HTML (no REST or GraphQL endpoint). jgehrcke/github-repo-stats has no
notion of container metrics, so post-process the report after it runs.
New: .github/scripts/ghcr_inject.py
- Scrapes Total downloads (exact integer from title="N", not the K-rounded
display) and the 30-day sparkline (rect data-merge-count, data-date).
- Merges per-day rows into maziggy/bambuddy/ghcr-pulls.csv on gh-pages.
Fresh window overwrites overlapping dates, so GitHub's late revisions
to the last 30 days self-correct; days older than 30 stay frozen at
whatever was captured while still in-window.
- Patches latest-report/report.html: adds a TOC entry, a Container
pulls (ghcr.io) section at the top, and a Vega-Lite line+point chart
whose theme/config is cloned from the existing Total clones chart so
it inherits the report's look-and-feel.
- Bracketed by HTML-comment markers so re-runs replace rather than
stack (jgehrcke regenerates report.html every tick; we re-inject).
- Hard-fails if either scrape pattern stops matching — silent fallbacks
would let the chart freeze without notice.
Workflow: after run-ghrs, checkout source + gh-pages, run the injector,
commit only if the diff is non-empty. Uses the existing contents: write
permission; no new secrets.
GHCR exposes total + 30-day daily-pull counts only in the package page
HTML (no REST or GraphQL endpoint). jgehrcke/github-repo-stats has no
notion of container metrics, so post-process the report after it runs.
New: .github/scripts/ghcr_inject.py
- Scrapes Total downloads (exact integer from title="N", not the K-rounded
display) and the 30-day sparkline (rect data-merge-count, data-date).
- Merges per-day rows into maziggy/bambuddy/ghcr-pulls.csv on gh-pages.
Fresh window overwrites overlapping dates, so GitHub's late revisions
to the last 30 days self-correct; days older than 30 stay frozen at
whatever was captured while still in-window.
- Patches latest-report/report.html: adds a TOC entry, a Container
pulls (ghcr.io) section at the top, and a Vega-Lite line+point chart
whose theme/config is cloned from the existing Total clones chart so
it inherits the report's look-and-feel.
- Bracketed by HTML-comment markers so re-runs replace rather than
stack (jgehrcke regenerates report.html every tick; we re-inject).
- Hard-fails if either scrape pattern stops matching — silent fallbacks
would let the chart freeze without notice.
Workflow: after run-ghrs, checkout source + gh-pages, run the injector,
commit only if the diff is non-empty. Uses the existing contents: write
permission; no new secrets.
Brings the Windows installer work from dev to main without merging
the rest of the 0.2.5b1 release content. Squashes 12 commits from
dev (8711c54e..7bb11df2) into a single net-effect commit on main.
Includes:
- installers/windows/ — Inno Setup .iss script, build.py, vendored
NSSM 2.24, bambuddy.ico (multi-resolution app icon), service
install/uninstall .bat files, build pipeline README
- backend/app/services/network_utils.py — Windows psutil branch so
the VP bind-IP dropdown enumerates interfaces; Linux/macOS path
unchanged
- .github/workflows/windows-installer.yml — reconciles main's
kludge-pushed copy with dev's accumulated changes (NSSM
vendoring, version-from-tag, unversioned alias step, etc.)
CHANGELOG and README entries for the Windows installer stay on
dev — they reference unreleased 0.2.5b1 release notes that aren't
on main yet.
Adds bambuddy-windows-x64-setup.exe (unversioned) alongside the
date-stamped bambuddy-<version>-windows-x64-setup.exe for stable
and beta tag releases. Lets external surfaces (website, wiki,
newsletters) link to a stable URL that survives version bumps:
https://github.com/maziggy/bambuddy/releases/latest/download/
bambuddy-windows-x64-setup.exe
Daily prereleases are excluded — GitHub's `latest` redirect skips
prereleases so the alias would add no value there, and an
unversioned name next to a date-stamped versioned one on a daily
release page is semantically confusing.
Address CodeQL actions/missing-workflow-permissions finding. Least-
privilege at workflow level: contents: write is required by
softprops/action-gh-release to attach the installer .exe to a tag
release; all other steps are read-only.
Address CodeQL actions/missing-workflow-permissions finding. Least-
privilege at workflow level: contents: write is required by
softprops/action-gh-release to attach the installer .exe to a tag
release; all other steps are read-only.
windows-latest runners ship Inno Setup 6.7.1 pre-installed under the
same path we already hardcode for ISCC.exe; the choco install was trying
to downgrade to 6.2.2 and failing on the version mismatch.
windows-latest runners ship Inno Setup 6.7.1 pre-installed under the
same path we already hardcode for ISCC.exe; the choco install was trying
to downgrade to 6.2.2 and failing on the version mismatch.
Lays down the Inno Setup + embedded Python pipeline for producing a
self-contained Bambuddy Windows installer .exe. The installer ships
an embedded Python 3.13, the pre-built React bundle, NSSM (service
supervisor) and ffmpeg — no host Python or Node required on the
target machine.
Architecture:
- Install: C:\Program Files\Bambuddy (admin install, one-time UAC)
- Data: C:\ProgramData\Bambuddy\data (preserved on uninstall)
- Service: registered via NSSM, runs as LocalSystem, autostart on boot
- UI: browser at http://localhost:8000 (Start Menu shortcut)
Files:
- installers/windows/build.py stages embedded Python + deps,
frontend bundle, NSSM, ffmpeg
- installers/windows/bambuddy.iss Inno Setup compiler script
- installers/windows/service/*.bat NSSM register/deregister
- .github/workflows/windows-installer.yml CI build on tag push + manual
dispatch, uploads .exe artifact
build.py hard-fails on non-Windows hosts; Wine cross-build is an
unsupported escape hatch behind --allow-non-windows. v1 ships unsigned
(SmartScreen warns on first run) — production signing will be wired up
via SignPath OSS once the application is approved.
See installers/windows/README.md for build prerequisites and the
embedded-Python ._pth gotchas.
GitHub forces Node-20 actions to run on Node 24 starting 2026-06-02 and
removes Node 20 from the runner on 2026-09-16. Bumping each action to
its first Node-24 major now gets us ahead of both deadlines and silences
the deprecation warnings already firing in every CI run.
Bumps (across ci.yml, security.yml, codeql.yml, auto-label-area.yml,
issue-closed.yml, stale.yml):
- actions/checkout v4 -> v6
- actions/setup-python v5 -> v6
- actions/setup-node v4 -> v6
- actions/cache v4 -> v5
- actions/upload-artifact v4 -> v7
- actions/github-script v7 -> v9
- actions/stale v9 -> v10
- docker/setup-buildx v3 -> v4
- docker/build-push v5 -> v7
Verified each major's breaking-change notes against our usage:
- setup-node v6 limits auto-cache to npm only; we already pass
cache: 'npm' explicitly, so nothing changes.
- github-script v9 drops require('@actions/github'); none of our
scripts use it (only require('fs') and the injected github/context
globals).
- setup-buildx v4 removes deprecated inputs; we call it with no
inputs.
- build-push v6 enables build summaries by default; informational,
can disable via DOCKER_BUILD_SUMMARY=false env if it gets noisy.
codeql-action stays on v4 (already runs on Node 24). Trivy and
github-repo-stats are Docker actions and aren't affected by the
Node-20 deprecation.
GitHub forces Node-20 actions to run on Node 24 starting 2026-06-02 and
removes Node 20 from the runner on 2026-09-16. Bumping each action to
its first Node-24 major now gets us ahead of both deadlines and silences
the deprecation warnings already firing in every CI run.
Bumps (across ci.yml, security.yml, codeql.yml, auto-label-area.yml,
issue-closed.yml, stale.yml):
- actions/checkout v4 -> v6
- actions/setup-python v5 -> v6
- actions/setup-node v4 -> v6
- actions/cache v4 -> v5
- actions/upload-artifact v4 -> v7
- actions/github-script v7 -> v9
- actions/stale v9 -> v10
- docker/setup-buildx v3 -> v4
- docker/build-push v5 -> v7
Verified each major's breaking-change notes against our usage:
- setup-node v6 limits auto-cache to npm only; we already pass
cache: 'npm' explicitly, so nothing changes.
- github-script v9 drops require('@actions/github'); none of our
scripts use it (only require('fs') and the injected github/context
globals).
- setup-buildx v4 removes deprecated inputs; we call it with no
inputs.
- build-push v6 enables build summaries by default; informational,
can disable via DOCKER_BUILD_SUMMARY=false env if it gets noisy.
codeql-action stays on v4 (already runs on Node 24). Trivy and
github-repo-stats are Docker actions and aren't affected by the
Node-20 deprecation.
170 issues have been closed with the `invalid` label (61 of them in
the last 30 days alone — ~1 in 5 of all closed issues), almost always
because the reporter hadn't run the in-app Connection Diagnostic or
checked the documented troubleshooting page. The Connection Diagnostic
shipped weeks ago but the bug-report form let people skip it: the
"I ran it" checkbox was `required: false` and the Support Package
field was optional. Tighten both.
Form changes (.github/ISSUE_TEMPLATE/bug_report.yml):
- Connection Diagnostic checkbox: required: false → true
- Support Package field: required: false → true ("drag the .zip
or explain why you cannot attach one")
- New required textarea "Troubleshooting steps already taken" —
forces the reporter to type WHAT they tried and WHICH wiki pages
they checked before submitting. Empty answers can't submit.
- Pre-form intro spells out the search → wiki → diagnostic →
support package sequence and cites the 1-in-5 stat
- Final-checks list grew from one to three required confirmations
(searched issues + checked troubleshooting wiki + ran Connection
Diagnostic for any connection/printing/camera issue)
Bug categorization (the gap that motivated this):
- Old `Component` dropdown was Bambuddy / SpoolBuddy / Both — no
area triage signal
- Replaced with two required dropdowns:
- Product: Bambuddy / SpoolBuddy
- Area: 15 options covering the actual feature surface +
Other / not sure
- Auto-label workflow (.github/workflows/auto-label-area.yml)
reads the Area dropdown from the rendered issue body on
open/edit and applies the matching area:* label. Tolerant of
CRLF and the _No response_ placeholder, won't re-add on edit
re-fires, warns on unknown Area values
Maintainer hand-off — labels must exist BEFORE the workflow runs,
since github-script's addLabels throws on missing labels. Create the
16 labels (15 area:* + 1 area:unsorted) once via the `gh label create`
commands captured in CHANGELOG / commit context.
OS dropdown left untouched (Docker stays — per Martin).
Printer Model dropdown verified against backend/app/utils/printer_models.py
PRINTER_MODEL_MAP: all 13 current models present (X1 Carbon, X1, X1E,
X2D, P1S, P1P, P2S, A1, A1 Mini, H2D, H2D Pro, H2C, H2S).
170 issues have been closed with the `invalid` label (61 of them in
the last 30 days alone — ~1 in 5 of all closed issues), almost always
because the reporter hadn't run the in-app Connection Diagnostic or
checked the documented troubleshooting page. The Connection Diagnostic
shipped weeks ago but the bug-report form let people skip it: the
"I ran it" checkbox was `required: false` and the Support Package
field was optional. Tighten both.
Form changes (.github/ISSUE_TEMPLATE/bug_report.yml):
- Connection Diagnostic checkbox: required: false → true
- Support Package field: required: false → true ("drag the .zip
or explain why you cannot attach one")
- New required textarea "Troubleshooting steps already taken" —
forces the reporter to type WHAT they tried and WHICH wiki pages
they checked before submitting. Empty answers can't submit.
- Pre-form intro spells out the search → wiki → diagnostic →
support package sequence and cites the 1-in-5 stat
- Final-checks list grew from one to three required confirmations
(searched issues + checked troubleshooting wiki + ran Connection
Diagnostic for any connection/printing/camera issue)
Bug categorization (the gap that motivated this):
- Old `Component` dropdown was Bambuddy / SpoolBuddy / Both — no
area triage signal
- Replaced with two required dropdowns:
- Product: Bambuddy / SpoolBuddy
- Area: 15 options covering the actual feature surface +
Other / not sure
- Auto-label workflow (.github/workflows/auto-label-area.yml)
reads the Area dropdown from the rendered issue body on
open/edit and applies the matching area:* label. Tolerant of
CRLF and the _No response_ placeholder, won't re-add on edit
re-fires, warns on unknown Area values
Maintainer hand-off — labels must exist BEFORE the workflow runs,
since github-script's addLabels throws on missing labels. Create the
16 labels (15 area:* + 1 area:unsorted) once via the `gh label create`
commands captured in CHANGELOG / commit context.
OS dropdown left untouched (Docker stays — per Martin).
Printer Model dropdown verified against backend/app/utils/printer_models.py
PRINTER_MODEL_MAP: all 13 current models present (X1 Carbon, X1, X1E,
X2D, P1S, P1P, P2S, A1, A1 Mini, H2D, H2D Pro, H2C, H2S).
Earlier patch trimmed the duplicate unit-test re-run from docker-test
to drop a 5-10 min job that wasn't adding coverage. But "wasn't adding
coverage" only holds for pure-logic tests — system-touching tests
(ffmpeg version probes, ftp clients, subprocess shell-outs, locale/
timezone-sensitive assertions, paths) genuinely can pass on the GHA
host and fail in python:3.13-slim. Curation via a `docker_env`
marker is fragile (new tests get forgotten); gating on `main` only
defers the cost without removing it.
Instead, run the full backend suite IN Docker on every PR but make
it fast:
- New docker-backend-tests job runs the same 4-way pytest-split
matrix as the host backend-tests, just inside the test image.
- docker/setup-buildx-action + docker/build-push-action@v5 with
cache-from/cache-to: type=gha,scope=backend-test persist the
BuildKit cache (pip-install layer included) across CI runs and
across the 4 sibling shards. Cold build is ~150s/shard; warm
build drops to ~10s/shard.
- fail-fast: false so a single failing shard surfaces the rest's
output too.
Total CI wall-clock for a PR push is now gated by docker-test (the
image-build + integration HTTP smoke + integration test suite job)
at ~3 min, not by the unit-test re-run anymore.
The earlier ci.yml step that ran `docker compose run --rm
backend-test` synchronously in the docker-test job stays removed —
the new docker-backend-tests matrix covers the same ground and is
much faster.
Earlier patch trimmed the duplicate unit-test re-run from docker-test
to drop a 5-10 min job that wasn't adding coverage. But "wasn't adding
coverage" only holds for pure-logic tests — system-touching tests
(ffmpeg version probes, ftp clients, subprocess shell-outs, locale/
timezone-sensitive assertions, paths) genuinely can pass on the GHA
host and fail in python:3.13-slim. Curation via a `docker_env`
marker is fragile (new tests get forgotten); gating on `main` only
defers the cost without removing it.
Instead, run the full backend suite IN Docker on every PR but make
it fast:
- New docker-backend-tests job runs the same 4-way pytest-split
matrix as the host backend-tests, just inside the test image.
- docker/setup-buildx-action + docker/build-push-action@v5 with
cache-from/cache-to: type=gha,scope=backend-test persist the
BuildKit cache (pip-install layer included) across CI runs and
across the 4 sibling shards. Cold build is ~150s/shard; warm
build drops to ~10s/shard.
- fail-fast: false so a single failing shard surfaces the rest's
output too.
Total CI wall-clock for a PR push is now gated by docker-test (the
image-build + integration HTTP smoke + integration test suite job)
at ~3 min, not by the unit-test re-run anymore.
The earlier ci.yml step that ran `docker compose run --rm
backend-test` synchronously in the docker-test job stays removed —
the new docker-backend-tests matrix covers the same ground and is
much faster.
The "Docker Build" job in ci.yml was running the same 5287 backend
tests + 2022 frontend tests inside the bambuddy-backend-test /
bambuddy-frontend-test images that the host-side backend-tests and
frontend-tests jobs had already run. Same test code, same Python
version (env.PYTHON_VERSION), same requirements.txt the test image
installs. On 2-vCPU GHA runners that re-run added 5-10 min of
wall-clock for zero new coverage — and "frontend tests in Docker"
added another 2-3 min for the same reason.
Drop both steps from the CI job. Keep everything that validates the
Docker IMAGE specifically: production image build, backend module
import verification, static-files-copied check, integration
container bring-up + health/API/static HTTP smoke checks, and the
integration test suite (which IS genuinely Docker-specific — it
runs against the live container via BAMBUDDY_TEST_URL).
test_docker.sh keeps the unit-test reruns because devs running it
locally don't have a separate host-side pytest job to compare
against.
Combined with the earlier 4-way pytest-split shard on the host
backend-tests job, expected PR-push wall-clock drops from
~10-12 min to ~3 min, gated on max(backend-tests shard, frontend
tests, docker-image-build+integration).
The "Docker Build" job in ci.yml was running the same 5287 backend
tests + 2022 frontend tests inside the bambuddy-backend-test /
bambuddy-frontend-test images that the host-side backend-tests and
frontend-tests jobs had already run. Same test code, same Python
version (env.PYTHON_VERSION), same requirements.txt the test image
installs. On 2-vCPU GHA runners that re-run added 5-10 min of
wall-clock for zero new coverage — and "frontend tests in Docker"
added another 2-3 min for the same reason.
Drop both steps from the CI job. Keep everything that validates the
Docker IMAGE specifically: production image build, backend module
import verification, static-files-copied check, integration
container bring-up + health/API/static HTTP smoke checks, and the
integration test suite (which IS genuinely Docker-specific — it
runs against the live container via BAMBUDDY_TEST_URL).
test_docker.sh keeps the unit-test reruns because devs running it
locally don't have a separate host-side pytest job to compare
against.
Combined with the earlier 4-way pytest-split shard on the host
backend-tests job, expected PR-push wall-clock drops from
~10-12 min to ~3 min, gated on max(backend-tests shard, frontend
tests, docker-image-build+integration).
+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup
Root cause of the 4 CI failures on PR #1514 (all in
test_print_start_assigns_printer_id_to_vp_archive.py +
test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
in test_code_quality.py was deleting backend.app.main from sys.modules
and re-importing it via importlib.import_module. That created NEW
module-level dicts (_timelapse_baselines, _expected_prints,
_active_prints, …) and re-ran root_logger.addHandler — hence the
duplicate log lines at the same microsecond in captured stderr.
Any sibling test that bound those names via "from backend.app.main
import _timelapse_baselines" before the reimport now held a reference
to the OLD dict; production code (reached via "from backend.app.main
import on_print_start") resolved the symbol through the NEW module
instance. Production mutated the new dict, the test read the old one,
the assertion saw None / un-mutated mock_archive.
Locally with -n 30, xdist load-balanced test_code_quality.py to a
different worker process so the collision never happened (which is
why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
made the collision deterministic.
Fix: drop the "del sys.modules[name]" step. importlib.import_module
already returns the cached module if cached, or runs the import
machinery if not — either way, any import-time error surfaces. The
"fresh import" framing was theatre; in practice every module in the
list is already imported by other tests/fixtures before this test
runs, so we were never actually getting a fresh import anyway — just
destruction.
CI workflow tightening (separate concern, same PR since both touch
the test infrastructure):
- Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
runners. --tb=short is sufficient for failure context.
- Sharded backend-tests into a 4-way matrix via pytest-split (new
dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
all 4 run in parallel so wall-clock drops from 362s -> ~100s.
- fail-fast: false on the matrix so a single failing shard doesn't
hide failures in the other three — PRs see the complete failure
picture in one push.
+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup
Root cause of the 4 CI failures on PR #1514 (all in
test_print_start_assigns_printer_id_to_vp_archive.py +
test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
in test_code_quality.py was deleting backend.app.main from sys.modules
and re-importing it via importlib.import_module. That created NEW
module-level dicts (_timelapse_baselines, _expected_prints,
_active_prints, …) and re-ran root_logger.addHandler — hence the
duplicate log lines at the same microsecond in captured stderr.
Any sibling test that bound those names via "from backend.app.main
import _timelapse_baselines" before the reimport now held a reference
to the OLD dict; production code (reached via "from backend.app.main
import on_print_start") resolved the symbol through the NEW module
instance. Production mutated the new dict, the test read the old one,
the assertion saw None / un-mutated mock_archive.
Locally with -n 30, xdist load-balanced test_code_quality.py to a
different worker process so the collision never happened (which is
why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
made the collision deterministic.
Fix: drop the "del sys.modules[name]" step. importlib.import_module
already returns the cached module if cached, or runs the import
machinery if not — either way, any import-time error surfaces. The
"fresh import" framing was theatre; in practice every module in the
list is already imported by other tests/fixtures before this test
runs, so we were never actually getting a fresh import anyway — just
destruction.
CI workflow tightening (separate concern, same PR since both touch
the test infrastructure):
- Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
runners. --tb=short is sufficient for failure context.
- Sharded backend-tests into a 4-way matrix via pytest-split (new
dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
all 4 run in parallel so wall-clock drops from 362s -> ~100s.
- fail-fast: false on the matrix so a single failing shard doesn't
hide failures in the other three — PRs see the complete failure
picture in one push.
security.yml had this ignore added in 9d440beb but ci.yml runs its
own pip-audit step with a separate ignore list. CI was still failing
on main + dev. Reasoning identical to the security.yml comment —
disputed by PyJWT maintainers, no fix exists, Bambuddy uses
secrets.token_urlsafe(64) and rejects short secrets.
security.yml had this ignore added in 9d440beb but ci.yml runs its
own pip-audit step with a separate ignore list. CI was still failing
on main + dev. Reasoning identical to the security.yml comment —
disputed by PyJWT maintainers, no fix exists, Bambuddy uses
secrets.token_urlsafe(64) and rejects short secrets.
- requirements.txt: pin idna>=3.15 to clear ReDoS in idna.encode() on
crafted Unicode payloads. Transitive via anyio/httpx/requests/yarl,
so the explicit floor stops a future downstream loosening from
silently downgrading us.
- security.yml: permanently --ignore-vuln CVE-2025-45768 (PyJWT). The
advisory is disputed by the maintainers — "key length is chosen by
the application" — and no fix version exists. Bambuddy is safe:
auto-generates secrets via secrets.token_urlsafe(64) and rejects
file-loaded secrets shorter than 32 chars (auth.py:177, :184).
- security.yml: drop the stale Pygments --ignore-vuln CVE-2026-4539.
Pygments has been patched upstream; the ignore no longer matches
anything.
- requirements.txt: pin idna>=3.15 to clear ReDoS in idna.encode() on
crafted Unicode payloads. Transitive via anyio/httpx/requests/yarl,
so the explicit floor stops a future downstream loosening from
silently downgrading us.
- security.yml: permanently --ignore-vuln CVE-2025-45768 (PyJWT). The
advisory is disputed by the maintainers — "key length is chosen by
the application" — and no fix version exists. Bambuddy is safe:
auto-generates secrets via secrets.token_urlsafe(64) and rejects
file-loaded secrets shorter than 32 chars (auth.py:177, :184).
- security.yml: drop the stale Pygments --ignore-vuln CVE-2026-4539.
Pygments has been patched upstream; the ignore no longer matches
anything.