Living Document Notice
Published 2026-09-11. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Zero-Downtime Daemon Handoffs via Systemd Socket Activation
Summary
Across Fleet Operations, process restarts during software upgrades risk dropping in-flight TCP connections unless file descriptors are preserved across execution boundaries. Systemd socket activation isolates listener sockets in PID 1, allowing application binaries to restart or hot-swap with zero dropped packets and deterministic descriptor inheritance.
The File Descriptor Inheritance Problem
Traditional daemon deployment models terminate the running binary before starting the replacement process. Even when using graceful shutdown signals such as SIGTERM, an unavoidable race condition exists between the moment the old process calls close() on its listening socket and the moment the new process completes bind() and listen(). In-flight client SYN packets arriving during this inter-process window receive RST responses, causing connection resets across the fleet.
+-------------------------------------------------------------+
| Systemd Socket Activation Pipeline |
| |
| Client TCP Connection |
| | |
| v |
| [PID 1: systemd] === creates & binds socket ===> (FD 3) |
| | | |
| | (Spawns worker or hot-swaps binary) | |
| v v |
| [Worker Daemon PID 1420] <================= inherits FD 3 |
| | |
| +---> accept4() and services traffic |
+-------------------------------------------------------------+
By moving socket creation, port binding, and backlog buffering to systemd (PID 1), the listening file descriptor remains permanently open in the kernel. The daemon binary becomes an ephemeral execution target that can restart, crash, or reload without touching the transport endpoint.
Socket and Service Unit Definitions
Socket activation requires decoupling service configuration into paired .socket and .service units. The .socket unit instructs PID 1 to bind the interface and buffer incoming connection attempts in the kernel listen backlog:
# /etc/systemd/system/harbormaster-edge.socket
[Unit]
Description=Harbormaster Edge Ingress Socket Activation Listener
PartOf=harbormaster-edge.service
[Socket]
ListenStream=0.0.0.0:8443
ListenStream=[::]:8443
Backlog=2048
BindIPv6Only=both
FileDescriptorName=edge-ingress
NoDelay=true
[Install]
WantedBy=sockets.targetThe corresponding .service unit omits execution dependencies on external networking setup and inherits the allocated descriptors directly:
# /etc/systemd/system/harbormaster-edge.service
[Unit]
Description=Harbormaster Edge Ingress Worker
Requires=harbormaster-edge.socket
After=harbormaster-edge.socket
[Service]
Type=notify
ExecStart=/usr/local/bin/harbormaster-edge --supervised
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=2s
KillMode=mixed
TimeoutStopSec=15s
NonBlocking=true
LimitNOFILE=65536
[Install]
WantedBy=multi-user.targetSocket Descriptor Handoff Implementation
In standard Linux POSIX environments, systemd passes inherited listening sockets starting at file descriptor 3 (SD_LISTEN_FDS_START), accompanied by the environment variables LISTEN_FDS (indicating total descriptor count) and LISTEN_PID (enforcing process boundary isolation).
The receiving application validates ownership before polling or accepting connections:
import os
import socket
import sys
SD_LISTEN_FDS_START = 3
def acquire_systemd_socket() -> socket.socket:
listen_pid = os.environ.get("LISTEN_PID")
listen_fds = os.environ.get("LISTEN_FDS")
if not listen_pid or not listen_fds:
raise RuntimeError("Process was not launched via systemd socket activation")
if int(listen_pid) != os.getpid():
raise RuntimeError(f"LISTEN_PID ({listen_pid}) does not match current PID ({os.getpid()})")
num_fds = int(listen_fds)
if num_fds < 1:
raise ValueError("No file descriptors received from supervisor")
# Wrap the pre-bound file descriptor directly
fd = SD_LISTEN_FDS_START
server_sock = socket.fromfd(fd, socket.AF_INET6, socket.SOCK_STREAM)
server_sock.setblocking(False)
return server_sockProduction Verification and Connection Soak Test
To verify zero-downtime upgrades during active traffic, run a high-concurrency client loop while triggering a rolling daemon restart via systemd:
# In Terminal A: Run continuous TCP health probe
$ while true; do curl -s -o /dev/null -w "%{http_code} %{time_total}s
" http://127.0.0.1:8443/healthz; sleep 0.05; done
# In Terminal B: Execute service restart under load
$ systemctl restart harbormaster-edge.service
# In Terminal C: Inspect socket file descriptor state
$ ss -tlpn | grep 8443
LISTEN 0 2048 0.0.0.0:8443 0.0.0.0:* users:(("systemd",pid=1,fd=45))
LISTEN 0 2048 [::]:8443 [::]:* users:(("systemd",pid=1,fd=46))During service transition, in-flight connection requests queue within the 2048-entry kernel backlog buffer. When the new worker PID initializes and calls accept4(), all queued connections resolve immediately with zero dropped packets or TCP RST flags.
- Directus Target: on-the-line
- Garden Source Reference: zero-downtime-daemon-handoffs-via-systemd-socket-activation, systemd-socket-activation, harbormaster-ops, process-lifecycle, MOC - Fleet Operations, MOC - Bosun PKM Tools