Files
bambuddy/backend
maziggy 9884018497 fix: cancel-safe get_db + drop sqlalchemy.pool cancellation noise
@Carter3DP's support package showed bambuddy.log filling with two
  distinct cascades on long uploads:

    ERROR sqlalchemy.pool   Exception terminating connection ...
                            CancelledError: Cancelled via cancel scope
                            ... by starlette.middleware.base
                            .BaseHTTPMiddleware.__call__.call_next
    ERROR sqlalchemy.pool   The garbage collector is trying to clean up
                            non-checked-in connection ... will be
                            terminated.
    WARN  backend.app.main  Runtime tracking commit failed:
                            (sqlite3.OperationalError) database is locked

  Single root cause. Starlette's BaseHTTPMiddleware (used under the hood
  by every @app.middleware("http") decorator) cancels the inner task
  scope when a client disconnects mid-request — common on long
  multipart uploads where the client times out before the server's
  response. Pre-fix get_db only caught Exception, but CancelledError
  is BaseException, so cancellation skipped the rollback path entirely.
  The SQLite write lock stayed held until GC reclaimed the connection
  ages later, blocking every other writer in the meantime. On Postgres
  the leak shape is identical; the symptom would be "QueuePool limit
  ... overflow" instead of "database is locked".

  (1) get_db now catches BaseException so CancelledError triggers
      rollback. Both rollback() and close() are wrapped in
      asyncio.shield so the cleanup completes even when the await
      itself is being cancelled by the same cancel scope. SQLite write
      lock is released promptly; connection returns to the pool instead
      of leaking until GC.

  (2) CancelledPoolNoiseFilter (new filter on sqlalchemy.pool) drops
      the residual records that pre-existing pools still emit during
      their own cleanup. Two patterns suppressed:
        - "Exception terminating connection ..." with a CancelledError
          anywhere in the exc_info chain (walks __cause__/__context__
          with a seen-set guard against pathological cycles)
        - "The garbage collector is trying to clean up non-checked-in
          connection ..." (always symptomatic of cancellation; never
          independently actionable)
      Real pool problems — broken connections, OSError on terminate,
      pool exhaustion — keep flowing because they carry a different
      exception chain or a different message prefix.

  13 regression tests across test_get_db_cancel_safety.py (commit on
  clean exit, rollback on regular Exception, rollback on CancelledError,
  close runs even if rollback raises, close failure on clean exit
  doesn't propagate, rollback + close both go through asyncio.shield)
  and test_cancelled_pool_filter.py (drops cancellation-driven
  terminate, drops GC-cleanup, keeps real OSError terminate, keeps
  terminate without exc_info, keeps unrelated pool messages, drops
  chained-cause CancelledError, defensive guard against self-referential
  cause chains).

  Applies to SQLite and PostgreSQL — get_db is dialect-agnostic and
  the filtered messages come from base sqlalchemy.pool not from any
  specific dialect.
2026-04-27 16:32:10 +02:00
..
2025-11-28 10:23:59 +01:00