fix(shutdown): exec uvicorn as PID 1 in Docker, and bound the graceful-shutdown wait

Two defects, both invisible until you ask the app to stop.

Docker never shut down gracefully at all. CMD ["sh","-c","uvicorn ..."] left
the shell as PID 1 with uvicorn as its child, and dash does not forward
signals, so docker stop SIGTERMed the shell and uvicorn never heard about it.
Measured on the shipped image: the full 10s grace period, exit 137, and no
"Shutting down" line in the log. Every stop, restart and image update was a
hard kill -- no WAL checkpoint, no MQTT disconnect, no virtual-printer
teardown. `exec` makes uvicorn PID 1; the rebuilt image now stops in 1s with
exit 0 and checkpoints the WAL.

Separately, uvicorn's timeout_graceful_shutdown defaults to None -- wait
forever for in-flight requests. An MJPEG camera stream is a response that
never completes (httptools' connection shutdown() only flips keep_alive on an
in-flight cycle, it never closes the transport), so one open camera tile
pinned the process until systemd SIGKILLed at 90s. The ordering makes it
unfixable from inside the app: uvicorn fires the lifespan shutdown -- the code
that tears the streams down -- only after connections drain.

All six launchers now pass --timeout-graceful-shutdown 5: Dockerfile,
deploy/bambuddy.service, the systemd unit and launchd plist from
install/install.sh, the SpoolBuddy installer's unit, and the Windows NSSM
registration. On timeout uvicorn cancels the request tasks; the camera
generators already unwind cleanly on CancelledError.

TimeoutStopSec raised to 30s on the units and stop_grace_period: 30s added to
compose, as backstops rather than the mechanism. On Windows NSSM's default
1500ms AppStopMethodConsole was force-killing uvicorn mid-teardown; raised to
15s, with the WM_CLOSE and thread-message stages skipped (uvicorn is a console
app with neither a window nor a message loop).
This commit is contained in:
maziggy
2026-07-11 14:44:45 +02:00
parent aba00598bb
commit ba1394db3e
8 changed files with 202 additions and 7 deletions
+15 -1
View File
@@ -150,5 +150,19 @@ HEALTHCHECK --interval=30s --timeout=10s --start-period=10s --retries=3 \
# Port is configurable via PORT (default 8000); bind address via HOST (default
# 0.0.0.0). Set HOST=127.0.0.1 to bind loopback only, e.g. when a reverse proxy
# on the same host fronts the app.
#
# `exec` is load-bearing, not style. Without it the shell stays as PID 1 and
# uvicorn runs as its child; dash does not forward signals, so `docker stop`
# SIGTERMs the shell and uvicorn never hears about it. Every stop then ran to
# the end of the grace period and died on SIGKILL (exit 137) — no WAL
# checkpoint, no MQTT disconnect, no virtual-printer teardown, on every restart
# and every image update. With `exec`, uvicorn *is* PID 1 and gets the signal.
#
# --timeout-graceful-shutdown caps the wait on in-flight requests. Uvicorn's
# default is to wait forever, and an MJPEG camera stream is a response that
# never completes, so a single open camera tile would otherwise pin the process
# past Docker's 10s grace and back into SIGKILL. On timeout uvicorn cancels the
# request tasks; the camera generators already unwind cleanly on CancelledError.
ENV UVICORN_TIMEOUT_GRACEFUL_SHUTDOWN=5
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
CMD ["sh", "-c", "uvicorn backend.app.main:app --host ${HOST:-0.0.0.0} --port ${PORT:-8000} --loop asyncio"]
CMD ["sh", "-c", "exec uvicorn backend.app.main:app --host ${HOST:-0.0.0.0} --port ${PORT:-8000} --loop asyncio --timeout-graceful-shutdown ${UVICORN_TIMEOUT_GRACEFUL_SHUTDOWN}"]