Backend Tests (shard 1/4) timed out after 10 minutes on e8b901f54
("Updated BACKERS"), a docs-only commit. The step was not hung: it
reached 98%, every test passing, and was killed about 12 seconds short
of finishing.
pytest-split balances by duration only when it has a durations file,
and there was none. Without one it splits by test COUNT -- 2926 / 2926
/ 2926 / 2925, exactly 25% each, which is what the comment here claimed
was good enough. Count is not time. Measured over the full suite
(11703 tests, 1045.6s):
shard 1 2926 tests 658.2s
shard 2 2926 tests 173.1s
shard 3 2926 tests 113.3s
shard 4 2925 tests 101.1s
Shard 1 was carrying 63% of the suite's runtime -- a 6.5x spread -- and
had been walking toward the cap for weeks: 475s, 404s, 445s, then 587s
on 08-30, thirteen seconds under it, and 613s here. CI printed the
reason in its own log on every run: "[pytest-split] No test durations
found."
Commit tests/.test_durations and pass --durations-path explicitly:
shard 1 1273 tests 261.6s
shard 2 1070 tests 261.5s
shard 3 1522 tests 262.0s
shard 4 7838 tests 260.6s
A 1.01x spread. The lopsided test counts are the point: shard 4 takes
thousands of fast unit tests, shard 1 keeps the slow integration ones.
The four groups still partition the suite exactly -- union is 11703
with nothing dropped or duplicated, and all four pass.
--durations-path has to be explicit because pytest-split defaults it to
$CWD/.test_durations, and the two matrices run from different
directories: the native job from backend/, the Docker job from the
image root. Left implicit, the Docker one finds nothing and silently
falls back to the count split. Node IDs match across both because
backend/tests/pytest.ini pins rootdir to backend/tests either way, so a
single file serves them both; verified the file clears .dockerignore
and lands at /app/backend/tests/.test_durations at full size.
timeout-minutes 10 -> 15 for headroom, since the file goes stale as
tests are added. Staleness degrades slowly rather than breaking --
unknown tests are treated as average.
Frontend:
- react-router/-dom 7.18.1 -> 7.18.2. The RSC-mode CSRF advisory was carried
as a documented exception in the audit gate because its only fix was the
8.3.0 major; upstream backported it, so the exemption lapsed on its own --
an entry only holds while fixAvailable.isSemVerMajor is true. The allowlist
is now empty; the machinery stays for the next one.
- dompurify 3.4.12 -> 3.4.13. Ships in the app, but the path is unreachable:
no hooks registered, IN_PLACE never used.
- js-yaml override ^4.3.0 -> ^5.2.3 (fix not backported below 5.x, so a
major) and nanoid override ^3.3.18. Both dev-only, via eslint and postcss.
eslintrc calls only load(), on the legacy .eslintrc.yml path this repo does
not use; eslint, vite build and 2861 frontend tests pass on it.
Backend:
- cryptography >=48.0.1 -> >=50.0.0, aiohttp >=3.14.0 -> >=3.14.3, pyopenssl
>=26.3.0 -> >=26.4.0. CI resolves from scratch and was already installing
the fixed releases; the floors cover the case CI does not, an existing venv
where >= is satisfied and `pip install -r` upgrades nothing. pyOpenSSL has
to move with cryptography -- each release caps it to a narrow window, so a
stale pyOpenSSL pins cryptography below its own fix line.
- postcss 8.5.15 -> 8.5.23 (GHSA-r28c-9q8g-f849, source-map path traversal)
- brace-expansion override ^5.0.7 -> ^5.0.8 (GHSA-mh99-v99m-4gvg, DoS)
react-router: pin react-router-dom to exact 7.18.1 (direct dep) and react-router
to 7.18.1 via overrides (transitive). 7.18.1 is the most-patched 7.x -- it clears
14 advisories that older 7.x releases carry, several reachable from a SPA (open-
redirect XSS in Link/useNavigate, route-matching DoS). The one remaining advisory,
GHSA-qwww-vcr4-c8h2, is RSC-mode-only; Bambuddy is a Vite SPA using BrowserRouter
with no RSC runtime (@react-router/server not installed), so the path is
unreachable. The only version that fully clears npm audit is the 8.3.0 major
(no react-router-dom 8.x exists; it needs migrating 50 import sites plus a React
peer bump), deferred as its own change.
Because a version pin can't stop npm from reporting the theoretical 7.11.0
downgrade as fixAvailable, the ci.yml (hard) and security.yml (nightly issue)
audit gates gain a narrow, documented allowlist keyed on the GHSA id. It resolves
the react-router-dom -> react-router advisory chain and stays fail-closed: a
different advisory on react-router still fails the gate, and an isSemVerMajor
guard drops the exemption the moment a non-major fix ships, forcing us to take it.
The GitHub runner's Python toolcache ships setuptools 79.0.1, which
pip-audit flags for PYSEC-2026-3447 (fixed in 83.0.0), failing the
blocking Backend Security job. A fix version exists, so upgrade
setuptools in the install step rather than --ignore-vuln. Applied to
both ci.yml (blocking) and security.yml (scheduled scan).
GitHub forces Node-20 actions to run on Node 24 starting 2026-06-02 and
removes Node 20 from the runner on 2026-09-16. Bumping each action to
its first Node-24 major now gets us ahead of both deadlines and silences
the deprecation warnings already firing in every CI run.
Bumps (across ci.yml, security.yml, codeql.yml, auto-label-area.yml,
issue-closed.yml, stale.yml):
- actions/checkout v4 -> v6
- actions/setup-python v5 -> v6
- actions/setup-node v4 -> v6
- actions/cache v4 -> v5
- actions/upload-artifact v4 -> v7
- actions/github-script v7 -> v9
- actions/stale v9 -> v10
- docker/setup-buildx v3 -> v4
- docker/build-push v5 -> v7
Verified each major's breaking-change notes against our usage:
- setup-node v6 limits auto-cache to npm only; we already pass
cache: 'npm' explicitly, so nothing changes.
- github-script v9 drops require('@actions/github'); none of our
scripts use it (only require('fs') and the injected github/context
globals).
- setup-buildx v4 removes deprecated inputs; we call it with no
inputs.
- build-push v6 enables build summaries by default; informational,
can disable via DOCKER_BUILD_SUMMARY=false env if it gets noisy.
codeql-action stays on v4 (already runs on Node 24). Trivy and
github-repo-stats are Docker actions and aren't affected by the
Node-20 deprecation.
Earlier patch trimmed the duplicate unit-test re-run from docker-test
to drop a 5-10 min job that wasn't adding coverage. But "wasn't adding
coverage" only holds for pure-logic tests — system-touching tests
(ffmpeg version probes, ftp clients, subprocess shell-outs, locale/
timezone-sensitive assertions, paths) genuinely can pass on the GHA
host and fail in python:3.13-slim. Curation via a `docker_env`
marker is fragile (new tests get forgotten); gating on `main` only
defers the cost without removing it.
Instead, run the full backend suite IN Docker on every PR but make
it fast:
- New docker-backend-tests job runs the same 4-way pytest-split
matrix as the host backend-tests, just inside the test image.
- docker/setup-buildx-action + docker/build-push-action@v5 with
cache-from/cache-to: type=gha,scope=backend-test persist the
BuildKit cache (pip-install layer included) across CI runs and
across the 4 sibling shards. Cold build is ~150s/shard; warm
build drops to ~10s/shard.
- fail-fast: false so a single failing shard surfaces the rest's
output too.
Total CI wall-clock for a PR push is now gated by docker-test (the
image-build + integration HTTP smoke + integration test suite job)
at ~3 min, not by the unit-test re-run anymore.
The earlier ci.yml step that ran `docker compose run --rm
backend-test` synchronously in the docker-test job stays removed —
the new docker-backend-tests matrix covers the same ground and is
much faster.
The "Docker Build" job in ci.yml was running the same 5287 backend
tests + 2022 frontend tests inside the bambuddy-backend-test /
bambuddy-frontend-test images that the host-side backend-tests and
frontend-tests jobs had already run. Same test code, same Python
version (env.PYTHON_VERSION), same requirements.txt the test image
installs. On 2-vCPU GHA runners that re-run added 5-10 min of
wall-clock for zero new coverage — and "frontend tests in Docker"
added another 2-3 min for the same reason.
Drop both steps from the CI job. Keep everything that validates the
Docker IMAGE specifically: production image build, backend module
import verification, static-files-copied check, integration
container bring-up + health/API/static HTTP smoke checks, and the
integration test suite (which IS genuinely Docker-specific — it
runs against the live container via BAMBUDDY_TEST_URL).
test_docker.sh keeps the unit-test reruns because devs running it
locally don't have a separate host-side pytest job to compare
against.
Combined with the earlier 4-way pytest-split shard on the host
backend-tests job, expected PR-push wall-clock drops from
~10-12 min to ~3 min, gated on max(backend-tests shard, frontend
tests, docker-image-build+integration).
+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup
Root cause of the 4 CI failures on PR #1514 (all in
test_print_start_assigns_printer_id_to_vp_archive.py +
test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
in test_code_quality.py was deleting backend.app.main from sys.modules
and re-importing it via importlib.import_module. That created NEW
module-level dicts (_timelapse_baselines, _expected_prints,
_active_prints, …) and re-ran root_logger.addHandler — hence the
duplicate log lines at the same microsecond in captured stderr.
Any sibling test that bound those names via "from backend.app.main
import _timelapse_baselines" before the reimport now held a reference
to the OLD dict; production code (reached via "from backend.app.main
import on_print_start") resolved the symbol through the NEW module
instance. Production mutated the new dict, the test read the old one,
the assertion saw None / un-mutated mock_archive.
Locally with -n 30, xdist load-balanced test_code_quality.py to a
different worker process so the collision never happened (which is
why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
made the collision deterministic.
Fix: drop the "del sys.modules[name]" step. importlib.import_module
already returns the cached module if cached, or runs the import
machinery if not — either way, any import-time error surfaces. The
"fresh import" framing was theatre; in practice every module in the
list is already imported by other tests/fixtures before this test
runs, so we were never actually getting a fresh import anyway — just
destruction.
CI workflow tightening (separate concern, same PR since both touch
the test infrastructure):
- Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
runners. --tb=short is sufficient for failure context.
- Sharded backend-tests into a 4-way matrix via pytest-split (new
dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
all 4 run in parallel so wall-clock drops from 362s -> ~100s.
- fail-fast: false on the matrix so a single failing shard doesn't
hide failures in the other three — PRs see the complete failure
picture in one push.
security.yml had this ignore added in 9d440beb but ci.yml runs its
own pip-audit step with a separate ignore list. CI was still failing
on main + dev. Reasoning identical to the security.yml comment —
disputed by PyJWT maintainers, no fix exists, Bambuddy uses
secrets.token_urlsafe(64) and rejects short secrets.
CI only installed requirements.txt, missing pyOpenSSL from
requirements-dev.txt. This caused an ImportError on
TLS_FTPHandler during test collection, blocking all
unit/services tests. Also adds pytest-timeout to dev deps
instead of ad-hoc pip install in CI.
- Test 1: Build production image, verify imports & static files
- Test 2: Backend unit tests in Docker container
- Test 3: Frontend unit tests in Docker container
- Test 4: Integration tests (health, API, static files)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Match test_docker.sh workflow:
- Build Docker image
- Verify backend imports
- Verify static files exist
- Test health endpoint
- Test API endpoint
- Test static files served
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Previously only ran tests/unit/, now runs tests/ to match
test_backend.sh workflow (513 tests instead of 239).
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Backend:
- pytest configuration with async support and coverage
- Unit tests for notification service (23 tests)
- Unit tests for smart plug manager (12 tests)
- Unit tests for archive service (16 tests)
- Integration tests for API endpoints
- Fix: notifications now send immediately (digest is summary only)
Frontend:
- Vitest configuration with jsdom and coverage
- MSW for API mocking
- Component tests for Toggle, Button, Card, ConfirmModal (77 tests)
- Test utilities with custom render wrapper
CI/CD:
- GitHub Actions workflow for automated testing
- Backend lint, unit tests, integration tests
- Frontend lint, type-check, unit tests, build