Gracefully Stopping a Python Daemon with the signal Module
A long-running daemon rarely lives in isolation. It usually sits behind a systemd unit, a container, or a small supervisor script, and somewhere in that lifecycle a process manager will ask it to stop. If the program responds by dying on the spot, the queue of work it was processing gets truncated, in-memory caches evaporate, and the next restart has to replay everything from scratch.
For systems developers in places like Sydney or Melbourne, where SaaS platforms often serve both Australian and overseas customers, the cost of an unclean exit is doubled. Customers on the east coast notice the blip, and the offshore side sees the warm-up penalty a few hours later. A predictable shutdown is part of the operational contract, not a luxury.
Python ships with a built-in module for exactly this job. The signal module exposes the POSIX signal interface and lets a program register callbacks that run whenever the kernel delivers SIGTERM, SIGINT, or SIGHUP. Wiring those callbacks into a daemon turns a hard kill into a clean handoff.
Why a polite exit matters for background services
Most daemons carry state that does not survive a process restart. A connection pool might hold open sockets, a worker queue might keep unsent messages, and a metrics buffer could contain the last few seconds of telemetry. None of that is safe to throw away mid-flight.
In Australian deployments this concern is amplified by the Privacy Act 1988 and the Notifiable Data Breaches scheme. If a consumer-facing service is collecting personal information, losing an in-flight batch because the process was killed at 02:00 AEDT can turn into an incident reportable to the Office of the Australian Information Commissioner. A graceful shutdown routine that flushes pending writes first is a quiet but effective control.
There is also a practical side. Process managers grade services on shutdown latency. Systemd's TimeoutStopSec defaults to ninety seconds, and an orchestrator like Kubernetes sends SIGTERM, waits thirty seconds, then escalates to SIGKILL. A daemon that ignores SIGTERM is one that always gets killed.
How the signal module exposes POSIX signals
The signal module is small and surprisingly approachable. signal.signal(signum, handler) installs a Python callable that runs when the operating system delivers a particular signal. SIGTERM is the polite request used by systemd, Docker, and most supervisors. SIGINT is what a terminal sends when the user presses Ctrl-C. SIGHUP is traditionally the cue to reload configuration.
A typical handler looks like a thin shim that flips a flag:
import signal
import threading
shutdown_event = threading.Event()
def request_stop(signum, frame):
shutdown_event.set()
signal.signal(signal.SIGTERM, request_stop)
signal.signal(signal.SIGINT, request_stop)
The handler itself runs on the main thread inside the interpreter's signal trampoline, so it should stay short. The right pattern is to record the intent and let the work loop notice the change between iterations. That approach avoids the classic bug of doing real work in a signal handler and tripping race conditions inside CPython.
Building a daemon that listens for shutdown requests
Putting the flag pattern into a real daemon means structuring the main loop so that it checks the event frequently. A simple example uses a worker thread that processes tasks from a queue:
import queue, threading
tasks = queue.Queue()
shutdown_event = threading.Event()
def worker():
while not shutdown_event.is_set():
try:
item = tasks.get(timeout=1)
except queue.Empty:
continue
process(item)
The worker drains the queue, sleeps briefly between checks, and exits as soon as the flag is raised. The main thread joins it with a timeout so that a misbehaving task cannot block the shutdown forever. This shape mirrors the background watchers that accompany utilities such as DenyHosts, where the daemon needs to react promptly to operator requests without dropping in-flight log lines.
Coordinating state and flushing buffers on the way out
Once the flag is set, the daemon still has housekeeping to do. SQLite WAL files, log buffers, configuration reloads, and AMQP or Redis connections all need a chance to close cleanly. A second, slower drain loop can iterate while there is still work, while a hard deadline forces the process to move on.
In Australian SOC environments following the ACSC Essential Eight, audit logs must be preserved across restarts. The shutdown routine should call logging.shutdown() after flushing any internal buffers, and it should not return until the final log line has been written. A common trick is to keep a small buffer of pending entries inside the worker and drain it before allowing the main thread to exit.
Shared mutable state also needs care. If multiple workers update a counter or a counter file, the shutdown sequence should stop accepting new work, let the in-flight tasks finish, then write the consolidated state. Using a threading.Event plus a context manager around the shared resource is usually enough.
When a worker refuses to cooperate
Sometimes a task hangs because the network is slow or a third-party library has its own blocking call. The graceful path needs an escape hatch. After the flag is set, the main thread waits on the worker join with a generous but bounded timeout. If the timeout expires, the worker is marked as a zombie and the daemon proceeds to close sockets, flush logs, and exit.
For services deployed in regions such as Perth or Brisbane, where round-trip latency to US-based APIs can stretch past three hundred milliseconds, that timeout should be tuned to the slowest expected dependency rather than the default one-second value. Logging the abandoned task with enough context for an operator to replay it manually is far more useful than letting the kernel decide with SIGKILL.
A more advanced pattern uses SIGUSR1 as a soft "please checkpoint" trigger and SIGUSR2 as a "stop accepting work but keep running" trigger. Production daemons written in Python rarely need those, but they are worth knowing when porting C-based tooling into Python.
Observability and the operator's view of shutdown
A graceful exit should leave a trail. Emitting a structured log event with the signal received, the drain duration, the number of items processed during shutdown, and the final exit code gives an SRE team something to alert on. Tools such as the ELK stack, Loki, or simply journalctl on a modern Linux host can pick up these events.
For Australian organisations operating under the Security of Critical Infrastructure Act, that trail also feeds into mandatory reporting pipelines. A daemon that logs its own lifecycle makes compliance reporting a matter of grep rather than archaeology. The signal module itself does not enforce any of this; it merely gives the daemon a chance to perform the bookkeeping before it leaves.
There is a temptation to skip this step because the daemon is "just internal". The same daemon that shuts down cleanly today is the one that, six months from now, will be running on a production host in Canberra or Adelaide with a paying customer depending on it. Treating shutdown as a first-class feature, rather than an afterthought, is what separates a script from a service.
