mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-10-08 23:21:58 +02:00
@Carter3DP's support package showed bambuddy.log filling with two
distinct cascades on long uploads:
ERROR sqlalchemy.pool Exception terminating connection ...
CancelledError: Cancelled via cancel scope
... by starlette.middleware.base
.BaseHTTPMiddleware.__call__.call_next
ERROR sqlalchemy.pool The garbage collector is trying to clean up
non-checked-in connection ... will be
terminated.
WARN backend.app.main Runtime tracking commit failed:
(sqlite3.OperationalError) database is locked
Single root cause. Starlette's BaseHTTPMiddleware (used under the hood
by every @app.middleware("http") decorator) cancels the inner task
scope when a client disconnects mid-request — common on long
multipart uploads where the client times out before the server's
response. Pre-fix get_db only caught Exception, but CancelledError
is BaseException, so cancellation skipped the rollback path entirely.
The SQLite write lock stayed held until GC reclaimed the connection
ages later, blocking every other writer in the meantime. On Postgres
the leak shape is identical; the symptom would be "QueuePool limit
... overflow" instead of "database is locked".
(1) get_db now catches BaseException so CancelledError triggers
rollback. Both rollback() and close() are wrapped in
asyncio.shield so the cleanup completes even when the await
itself is being cancelled by the same cancel scope. SQLite write
lock is released promptly; connection returns to the pool instead
of leaking until GC.
(2) CancelledPoolNoiseFilter (new filter on sqlalchemy.pool) drops
the residual records that pre-existing pools still emit during
their own cleanup. Two patterns suppressed:
- "Exception terminating connection ..." with a CancelledError
anywhere in the exc_info chain (walks __cause__/__context__
with a seen-set guard against pathological cycles)
- "The garbage collector is trying to clean up non-checked-in
connection ..." (always symptomatic of cancellation; never
independently actionable)
Real pool problems — broken connections, OSError on terminate,
pool exhaustion — keep flowing because they carry a different
exception chain or a different message prefix.
13 regression tests across test_get_db_cancel_safety.py (commit on
clean exit, rollback on regular Exception, rollback on CancelledError,
close runs even if rollback raises, close failure on clean exit
doesn't propagate, rollback + close both go through asyncio.shield)
and test_cancelled_pool_filter.py (drops cancellation-driven
terminate, drops GC-cleanup, keeps real OSError terminate, keeps
terminate without exc_info, keeps unrelated pool messages, drops
chained-cause CancelledError, defensive guard against self-referential
cause chains).
Applies to SQLite and PostgreSQL — get_db is dialect-agnostic and
the filtered messages come from base sqlalchemy.pool not from any
specific dialect.