Files
bambuddy/backend/tests/unit/test_archive_run_aggregation.py
T
maziggy 1e08c25a9f fix(stats): #1593 multi-plate parser + per-run project rollup + carry-over system totals + accuracy band
Two stacked causes under-reported multi-plate prints in the project
  rollup and the archive card.

  Root cause 1 - parser only read plate 1.

  ThreeMFParser._parse_slice_info used root.find(".//plate") and pulled
  prediction / weight from that one element. Any multi-plate file's
  archive-level print_time_seconds / filament_used_grams reflected
  plate 1 alone. The /plates endpoint already looped findall and was
  correct, which is why the plate carousel showed the right numbers
  while the archive card was wrong.

  Fix: loop every <plate> and sum prediction + weight. Per-plate
  concepts (plate_number, _plate_index, printable_objects) only set
  when there's exactly one plate - for multi-plate exports the
  archive represents all plates and a single index doesn't apply at
  the file level. bed_type keeps the first plate's value as a
  best-effort default. Malformed prediction / weight on individual
  plates skip cleanly rather than poison the sum.

  Root cause 2 - project rollup aggregated PrintArchive, not the
  per-run log.

  compute_project_stats and list_projects quick-stats summed
  PrintArchive.print_time_seconds / filament_used_grams / cost /
  energy_* WHERE project_id. A reprint reuses the source archive row
  and writes a new PrintLogEntry, so 3 sequential runs collapsed to 1
  archive - and that archive's numbers were already plate-1-only from
  cause 1. The Archive Print Log path was already correct because it
  drove off print_log_entries (archives.py:420 comment).

  Fix: both compute_project_stats and the list_projects quick-stats
  block inner-join print_log_entries -> print_archives WHERE
  archives.project_id. total_archives becomes COUNT(PrintLogEntry.id),
  failed_prints counts runs in failed/aborted/cancelled/stopped,
  completed_items is SUM(PrintArchive.quantity) for runs where
  status='completed', time/filament/cost/energy from PrintLogEntry.
  Orphan log rows (archive_id IS NULL post archive deletion) are
  excluded by the inner join.

  Same-shape fixes carried forward (no follow-ups per project rule):

  system.py system-info totals: total_print_time / total_filament
  had the same bug shape - summed PrintArchive directly so reprints
  collapsed to one row. Now sums PrintLogEntry.duration_seconds /
  filament_used_grams. The semantic shift is also a correctness
  improvement: the field now reflects time the printer actually spent
  printing, not slicer-estimated time.

  archives.py time-accuracy metric: estimate / actual per run where
  estimate = PrintArchive.print_time_seconds. Post-parser-fix
  multi-plate archives have file-level estimate but per-run actual =
  one plate, so ratio = N x 100% for an N-plate file. The calc now
  clamps each row to the [50%, 200%] plausibility band before
  contributing to the printer-level average; single-plate accuracy
  (the case the metric is designed for) stays fully included.

  Backfill: users with AMS spool tracking - the reporter's case - have
  per-run filament_used_grams from the tracked spool delta, so stats
  become correct immediately. Users without tracking fall back to the
  archive estimate and undercount until they reprint. Archive card
  still reads PrintArchive.filament_used_grams directly so old
  multi-plate archives keep plate-1-only numbers until reslice -
  forward-only as the reporter accepted.
2026-06-02 13:35:06 +02:00

315 lines
11 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Tests for the PrintRun-based stats aggregation (#1378).
Statistics and per-archive aggregates now come from PrintLogEntry rows rather
than PrintArchive's runtime fields, so a reprint contributes new totals
instead of overwriting the source archive's first-run data.
"""
from datetime import datetime, timezone
import pytest
from httpx import AsyncClient
from backend.app.models.print_log import PrintLogEntry
@pytest.mark.asyncio
@pytest.mark.integration
async def test_stats_count_reprints_independently(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""A reprint adds to stats instead of overwriting the source archive."""
printer = await printer_factory()
archive = await archive_factory(
printer.id,
status="completed",
filament_used_grams=100.0,
cost=2.5,
print_time_seconds=3600,
with_run=False,
)
# First run — completed, 100g.
db_session.add(
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="completed",
started_at=datetime(2026, 5, 1, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 5, 1, 11, 0, tzinfo=timezone.utc),
duration_seconds=3600,
filament_used_grams=100.0,
cost=2.5,
created_at=datetime(2026, 5, 1, 11, 0, tzinfo=timezone.utc),
)
)
# Reprint — failed at 10g.
db_session.add(
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="failed",
started_at=datetime(2026, 5, 5, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 5, 5, 10, 5, tzinfo=timezone.utc),
duration_seconds=300,
filament_used_grams=10.0,
cost=0.25,
failure_reason="Cancelled by user",
created_at=datetime(2026, 5, 5, 10, 5, tzinfo=timezone.utc),
)
)
await db_session.commit()
response = await async_client.get("/api/v1/archives/stats")
assert response.status_code == 200
body = response.json()
# Both runs counted, not the single archive row.
assert body["total_prints"] == 2
assert body["successful_prints"] == 1
assert body["failed_prints"] == 1
# 100g + 10g — NOT 10g (which is what archives.filament_used_grams alone
# would give if the archive's runtime fields were the source of truth).
assert body["total_filament_grams"] == pytest.approx(110.0)
assert body["total_cost"] == pytest.approx(2.75)
@pytest.mark.asyncio
@pytest.mark.integration
async def test_archive_list_includes_run_aggregates(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""List response carries run_count, last_run_at, total_filament_actual_grams."""
printer = await printer_factory()
archive = await archive_factory(
printer.id,
status="completed",
filament_used_grams=100.0,
with_run=False,
)
db_session.add_all(
[
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="completed",
started_at=datetime(2026, 5, 1, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 5, 1, 11, 0, tzinfo=timezone.utc),
filament_used_grams=100.0,
created_at=datetime(2026, 5, 1, 11, 0, tzinfo=timezone.utc),
),
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="failed",
started_at=datetime(2026, 5, 10, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 5, 10, 10, 5, tzinfo=timezone.utc),
filament_used_grams=10.0,
created_at=datetime(2026, 5, 10, 10, 5, tzinfo=timezone.utc),
),
]
)
await db_session.commit()
response = await async_client.get("/api/v1/archives/")
assert response.status_code == 200
rows = response.json()
row = next(r for r in rows if r["id"] == archive.id)
assert row["run_count"] == 2
assert row["successful_run_count"] == 1
assert row["failed_run_count"] == 1
assert row["total_filament_actual_grams"] == pytest.approx(110.0)
assert row["last_run_at"] is not None # max(started_at) populated
@pytest.mark.asyncio
@pytest.mark.integration
async def test_runs_endpoint_returns_runs_newest_first(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""GET /archives/{id}/runs returns each PrintLogEntry for the archive."""
printer = await printer_factory()
archive = await archive_factory(
printer.id,
status="completed",
with_run=False,
)
db_session.add_all(
[
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="completed",
started_at=datetime(2026, 4, 1, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 4, 1, 11, 0, tzinfo=timezone.utc),
filament_used_grams=50.0,
),
PrintLogEntry(
archive_id=archive.id,
printer_id=archive.printer_id,
status="failed",
started_at=datetime(2026, 5, 1, 10, 0, tzinfo=timezone.utc),
completed_at=datetime(2026, 5, 1, 10, 5, tzinfo=timezone.utc),
filament_used_grams=5.0,
failure_reason="Cancelled by user",
),
]
)
await db_session.commit()
response = await async_client.get(f"/api/v1/archives/{archive.id}/runs")
assert response.status_code == 200
body = response.json()
assert body["total"] == 2
# Newest first
assert body["items"][0]["status"] == "failed"
assert body["items"][0]["failure_reason"] == "Cancelled by user"
assert body["items"][1]["status"] == "completed"
assert body["items"][1]["filament_used_grams"] == pytest.approx(50.0)
@pytest.mark.asyncio
@pytest.mark.integration
async def test_purge_stats_also_deletes_linked_runs(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""``DELETE /archives/{id}?purge_stats=true`` hard-deletes linked PrintLogEntry
rows so their filament / cost / count contributions truly leave Quick Stats.
Without this, ON DELETE SET NULL on the FK would orphan the runs and they'd
keep showing up in the new aggregate-from-PrintLogEntry totals (#1378)."""
from sqlalchemy import func, select
printer = await printer_factory()
keep = await archive_factory(printer.id, status="completed", filament_used_grams=50.0)
purge = await archive_factory(printer.id, status="completed", filament_used_grams=100.0)
# Extra runs on the archive about to be purged, to prove they all go.
db_session.add_all(
[
PrintLogEntry(
archive_id=purge.id,
printer_id=purge.printer_id,
status="failed",
filament_used_grams=10.0,
),
PrintLogEntry(
archive_id=purge.id,
printer_id=purge.printer_id,
status="completed",
filament_used_grams=100.0,
),
]
)
await db_session.commit()
resp = await async_client.delete(f"/api/v1/archives/{purge.id}?purge_stats=true")
assert resp.status_code == 200
assert resp.json()["purged_from_stats"] is True
remaining = await db_session.execute(
select(func.count(PrintLogEntry.id)).where(PrintLogEntry.archive_id == purge.id)
)
assert remaining.scalar() == 0
# The OTHER archive's auto-synthesized run is still there.
keep_remaining = await db_session.execute(
select(func.count(PrintLogEntry.id)).where(PrintLogEntry.archive_id == keep.id)
)
assert keep_remaining.scalar() == 1
@pytest.mark.asyncio
@pytest.mark.integration
async def test_soft_delete_keeps_runs_for_stats(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""Default soft-delete (without ``purge_stats=true``) keeps the archive's
PrintLogEntry rows so the #1343 stats-preservation contract still holds —
the archive disappears from listings, but its filament / time / cost stay
in Quick Stats."""
from sqlalchemy import func, select
printer = await printer_factory()
archive = await archive_factory(printer.id, status="completed", filament_used_grams=75.0)
resp = await async_client.delete(f"/api/v1/archives/{archive.id}")
assert resp.status_code == 200
assert resp.json()["purged_from_stats"] is False
# The run row is still there for stats aggregation.
runs = await db_session.execute(select(func.count(PrintLogEntry.id)).where(PrintLogEntry.archive_id == archive.id))
assert runs.scalar() == 1
stats = (await async_client.get("/api/v1/archives/stats")).json()
assert stats["total_prints"] >= 1
assert stats["total_filament_grams"] >= 75.0
@pytest.mark.asyncio
@pytest.mark.integration
async def test_time_accuracy_excludes_multi_plate_plate_by_plate_outliers(
async_client: AsyncClient, archive_factory, printer_factory, db_session
):
"""Per-run accuracy clamps to a plausible 50%-200% band so multi-plate
archives printed plate-by-plate don't poison the printer-level average.
Pre-#1593 the parser stored plate-1-only time in
``PrintArchive.print_time_seconds``, so a plate-by-plate run produced a
near-100% ratio by accident. Post-#1593 the field is the sum across
plates, so each plate-by-plate run produces estimate/actual = N×100%
for an N-plate file. Without the band filter a single 3-plate file
printed plate-by-plate would drag the printer's accuracy reading to
~300%, which is pure noise. The metric is designed for the
single-plate-file case and should reflect real slicer drift there.
"""
printer = await printer_factory()
# Archive 1: single-plate file. Estimate 3600s, actual 3700s
# → ratio 97.3% (well within band).
single = await archive_factory(
printer.id,
print_time_seconds=3600,
with_run=False,
)
db_session.add(
PrintLogEntry(
archive_id=single.id,
printer_id=printer.id,
status="completed",
duration_seconds=3700,
)
)
# Archive 2: multi-plate file (3 plates totaling 18000s). Two runs
# printed plate-by-plate at ~6000s each — ratio 18000/6000 = 300%.
# Both must be filtered out so the printer average stays at the
# single-plate file's 97.3% reading.
multi = await archive_factory(
printer.id,
print_time_seconds=18000,
with_run=False,
)
db_session.add(
PrintLogEntry(
archive_id=multi.id,
printer_id=printer.id,
status="completed",
duration_seconds=6000,
)
)
db_session.add(
PrintLogEntry(
archive_id=multi.id,
printer_id=printer.id,
status="completed",
duration_seconds=6100,
)
)
await db_session.commit()
body = (await async_client.get("/api/v1/archives/stats")).json()
assert body["average_time_accuracy"] == pytest.approx(97.3, abs=0.1)
assert body["time_accuracy_by_printer"][str(printer.id)] == pytest.approx(97.3, abs=0.1)