mirror of
https://github.com/maziggy/bambuddy.git
synced 2026-09-29 18:51:43 +02:00
fix(shutdown): exec uvicorn as PID 1 in Docker, and bound the graceful-shutdown wait
Two defects, both invisible until you ask the app to stop. Docker never shut down gracefully at all. CMD ["sh","-c","uvicorn ..."] left the shell as PID 1 with uvicorn as its child, and dash does not forward signals, so docker stop SIGTERMed the shell and uvicorn never heard about it. Measured on the shipped image: the full 10s grace period, exit 137, and no "Shutting down" line in the log. Every stop, restart and image update was a hard kill -- no WAL checkpoint, no MQTT disconnect, no virtual-printer teardown. `exec` makes uvicorn PID 1; the rebuilt image now stops in 1s with exit 0 and checkpoints the WAL. Separately, uvicorn's timeout_graceful_shutdown defaults to None -- wait forever for in-flight requests. An MJPEG camera stream is a response that never completes (httptools' connection shutdown() only flips keep_alive on an in-flight cycle, it never closes the transport), so one open camera tile pinned the process until systemd SIGKILLed at 90s. The ordering makes it unfixable from inside the app: uvicorn fires the lifespan shutdown -- the code that tears the streams down -- only after connections drain. All six launchers now pass --timeout-graceful-shutdown 5: Dockerfile, deploy/bambuddy.service, the systemd unit and launchd plist from install/install.sh, the SpoolBuddy installer's unit, and the Windows NSSM registration. On timeout uvicorn cancels the request tasks; the camera generators already unwind cleanly on CancelledError. TimeoutStopSec raised to 30s on the units and stop_grace_period: 30s added to compose, as backstops rather than the mechanism. On Windows NSSM's default 1500ms AppStopMethodConsole was force-killing uvicorn mid-teardown; raised to 15s, with the WM_CLOSE and thread-message stages skipped (uvicorn is a console app with neither a window nor a message loop).
This commit is contained in:
+15
-1
@@ -150,5 +150,19 @@ HEALTHCHECK --interval=30s --timeout=10s --start-period=10s --retries=3 \
|
||||
# Port is configurable via PORT (default 8000); bind address via HOST (default
|
||||
# 0.0.0.0). Set HOST=127.0.0.1 to bind loopback only, e.g. when a reverse proxy
|
||||
# on the same host fronts the app.
|
||||
#
|
||||
# `exec` is load-bearing, not style. Without it the shell stays as PID 1 and
|
||||
# uvicorn runs as its child; dash does not forward signals, so `docker stop`
|
||||
# SIGTERMs the shell and uvicorn never hears about it. Every stop then ran to
|
||||
# the end of the grace period and died on SIGKILL (exit 137) — no WAL
|
||||
# checkpoint, no MQTT disconnect, no virtual-printer teardown, on every restart
|
||||
# and every image update. With `exec`, uvicorn *is* PID 1 and gets the signal.
|
||||
#
|
||||
# --timeout-graceful-shutdown caps the wait on in-flight requests. Uvicorn's
|
||||
# default is to wait forever, and an MJPEG camera stream is a response that
|
||||
# never completes, so a single open camera tile would otherwise pin the process
|
||||
# past Docker's 10s grace and back into SIGKILL. On timeout uvicorn cancels the
|
||||
# request tasks; the camera generators already unwind cleanly on CancelledError.
|
||||
ENV UVICORN_TIMEOUT_GRACEFUL_SHUTDOWN=5
|
||||
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
|
||||
CMD ["sh", "-c", "uvicorn backend.app.main:app --host ${HOST:-0.0.0.0} --port ${PORT:-8000} --loop asyncio"]
|
||||
CMD ["sh", "-c", "exec uvicorn backend.app.main:app --host ${HOST:-0.0.0.0} --port ${PORT:-8000} --loop asyncio --timeout-graceful-shutdown ${UVICORN_TIMEOUT_GRACEFUL_SHUTDOWN}"]
|
||||
|
||||
Reference in New Issue
Block a user