+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup
Root cause of the 4 CI failures on PR #1514 (all in
test_print_start_assigns_printer_id_to_vp_archive.py +
test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
in test_code_quality.py was deleting backend.app.main from sys.modules
and re-importing it via importlib.import_module. That created NEW
module-level dicts (_timelapse_baselines, _expected_prints,
_active_prints, …) and re-ran root_logger.addHandler — hence the
duplicate log lines at the same microsecond in captured stderr.
Any sibling test that bound those names via "from backend.app.main
import _timelapse_baselines" before the reimport now held a reference
to the OLD dict; production code (reached via "from backend.app.main
import on_print_start") resolved the symbol through the NEW module
instance. Production mutated the new dict, the test read the old one,
the assertion saw None / un-mutated mock_archive.
Locally with -n 30, xdist load-balanced test_code_quality.py to a
different worker process so the collision never happened (which is
why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
made the collision deterministic.
Fix: drop the "del sys.modules[name]" step. importlib.import_module
already returns the cached module if cached, or runs the import
machinery if not — either way, any import-time error surfaces. The
"fresh import" framing was theatre; in practice every module in the
list is already imported by other tests/fixtures before this test
runs, so we were never actually getting a fresh import anyway — just
destruction.
CI workflow tightening (separate concern, same PR since both touch
the test infrastructure):
- Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
runners. --tb=short is sufficient for failure context.
- Sharded backend-tests into a 4-way matrix via pytest-split (new
dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
all 4 run in parallel so wall-clock drops from 362s -> ~100s.
- fail-fast: false on the matrix so a single failing shard doesn't
hide failures in the other three — PRs see the complete failure
picture in one push.
security.yml had this ignore added in 9d440beb but ci.yml runs its
own pip-audit step with a separate ignore list. CI was still failing
on main + dev. Reasoning identical to the security.yml comment —
disputed by PyJWT maintainers, no fix exists, Bambuddy uses
secrets.token_urlsafe(64) and rejects short secrets.
- requirements.txt: pin idna>=3.15 to clear ReDoS in idna.encode() on
crafted Unicode payloads. Transitive via anyio/httpx/requests/yarl,
so the explicit floor stops a future downstream loosening from
silently downgrading us.
- security.yml: permanently --ignore-vuln CVE-2025-45768 (PyJWT). The
advisory is disputed by the maintainers — "key length is chosen by
the application" — and no fix version exists. Bambuddy is safe:
auto-generates secrets via secrets.token_urlsafe(64) and rejects
file-loaded secrets shorter than 32 chars (auth.py:177, :184).
- security.yml: drop the stale Pygments --ignore-vuln CVE-2026-4539.
Pygments has been patched upstream; the ignore no longer matches
anything.
CI only installed requirements.txt, missing pyOpenSSL from
requirements-dev.txt. This caused an ImportError on
TLS_FTPHandler during test collection, blocking all
unit/services tests. Also adds pytest-timeout to dev deps
instead of ad-hoc pip install in CI.
- Fix defusedxml import style in print_queue.py to be recognized by Bandit
(use `import defusedxml.ElementTree as ET` not `from defusedxml import`)
- Update Trivy scanner version from 0.65.0 to 0.69.1