fix(test): stop sys.modules-deleting backend.app.main in test_code_quality

+ ci: shard backend tests 4-way + drop -v for ~3.5x wall-clock speedup

  Root cause of the 4 CI failures on PR #1514 (all in
  test_print_start_assigns_printer_id_to_vp_archive.py +
  test_timelapse_baseline_restart_recovery.py): test_all_modules_importable
  in test_code_quality.py was deleting backend.app.main from sys.modules
  and re-importing it via importlib.import_module. That created NEW
  module-level dicts (_timelapse_baselines, _expected_prints,
  _active_prints, …) and re-ran root_logger.addHandler — hence the
  duplicate log lines at the same microsecond in captured stderr.

  Any sibling test that bound those names via "from backend.app.main
  import _timelapse_baselines" before the reimport now held a reference
  to the OLD dict; production code (reached via "from backend.app.main
  import on_print_start") resolved the symbol through the NEW module
  instance. Production mutated the new dict, the test read the old one,
  the assertion saw None / un-mutated mock_archive.

  Locally with -n 30, xdist load-balanced test_code_quality.py to a
  different worker process so the collision never happened (which is
  why the suite was green for me). CI's -n auto = -n 2 on ubuntu-latest
  made the collision deterministic.

  Fix: drop the "del sys.modules[name]" step. importlib.import_module
  already returns the cached module if cached, or runs the import
  machinery if not — either way, any import-time error surfaces. The
  "fresh import" framing was theatre; in practice every module in the
  list is already imported by other tests/fixtures before this test
  runs, so we were never actually getting a fresh import anyway — just
  destruction.

  CI workflow tightening (separate concern, same PR since both touch
  the test infrastructure):

  - Dropped -v from the pytest invocation. 5300+ "PASSED foo::bar"
    lines per worker were eating ~30-60s of stdout I/O on 2-vCPU
    runners. --tb=short is sufficient for failure context.
  - Sharded backend-tests into a 4-way matrix via pytest-split (new
    dev dep). Each shard runs ~1326 tests in ~95s on a 2-vCPU runner;
    all 4 run in parallel so wall-clock drops from 362s -> ~100s.
  - fail-fast: false on the matrix so a single failing shard doesn't
    hide failures in the other three — PRs see the complete failure
    picture in one push.
This commit is contained in:
maziggy
2026-05-24 12:50:35 +02:00
parent af03a6f384
commit 4fac9ff12c
4 changed files with 44 additions and 7 deletions
+21 -3
View File
@@ -84,10 +84,17 @@ jobs:
--ignore-vuln CVE-2025-45768
backend-tests:
name: Backend Tests
name: Backend Tests (shard ${{ matrix.shard }}/4)
runs-on: ubuntu-latest
if: github.event_name == 'push' || github.event.pull_request.user.login != github.repository_owner
needs: backend-lint
strategy:
# Don't cancel sibling shards if one fails — we want every shard's
# failure list, not just the first one, so a single PR push shows
# all broken tests in one go.
fail-fast: false
matrix:
shard: [1, 2, 3, 4]
steps:
- uses: actions/checkout@v4
@@ -110,11 +117,22 @@ jobs:
pip install -r requirements.txt
pip install -r requirements-dev.txt
- name: Run tests
- name: Run tests (shard ${{ matrix.shard }}/4)
timeout-minutes: 10
run: |
cd backend
python -m pytest tests/ -v --tb=short --timeout=60 --timeout-method=thread -n auto
# -v dropped: 5300+ "PASSED foo::bar" lines per worker eat 30-60s
# of stdout I/O time on 2-vCPU runners. --tb=short is enough.
# --splits 4 --group N uses pytest-split to slice the collected
# test set roughly evenly across the 4 matrix shards; first run
# is name-hash-based, subsequent runs improve via .test_durations
# if you ever commit one (we don't — even the naive hash split
# gets us ≈25% per shard given the test mix here).
python -m pytest tests/ \
--tb=short \
--timeout=60 --timeout-method=thread \
-n auto \
--splits 4 --group ${{ matrix.shard }}
# ============================================================================
# Frontend Checks