Living Document Notice
Published 2026-09-15. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Kernel Out-Of-Memory Guardrails and Process Priority Tuning
Summary
When memory exhaustion occurs on single-node edge servers, default Linux OOM heuristics often terminate critical infrastructure daemons like SSH or primary database engines. Explicitly configuring /proc/[pid]/oom_score_adj and process scheduling priorities shields essential telemetry and access daemons during severe memory pressure.
The Unpredictability of Default OOM Scoring
When available physical RAM and swap space are exhausted, the Linux kernel invokes the Out-Of-Memory (OOM) killer to reclaim pages. By default, the kernel calculates an oom_score for every running task based on resident memory consumption (RSS), memory usage percentage, and the task’s root user status.
In edge environments running intensive indexing or large-scale document parsing tasks, this heuristic frequently causes catastrophic misfires. A background worker processing a large archive can push memory to the brink, leading the OOM killer to terminate the SSH daemon or API server simply because they hold substantial resident page allocations.
+-------------------------------------------------------------+
| Linux Kernel OOM Score Range |
| |
| -1000 (Immune) <--------------------> +1000 (First Victim)|
| ^ ^ |
| | | |
| [sshd / systemd-journald / db-engine] [batch workers] |
| OOMScoreAdjust=-1000 OOMScoreAdjust=+900
+-------------------------------------------------------------+
Controlling OOM Heuristics via Systemd
The oom_score_adj knob accepts integer values ranging from -1000 (completely disables OOM killing for the process) to +1000 (marks the process as the first target for termination).
Configure systemd unit drop-ins to protect remote management and storage daemons while directing the OOM killer toward batch background processes:
# /etc/systemd/system/sshd.service.d/oom-protect.conf
[Service]
# Shield remote recovery shell from OOM termination
OOMScoreAdjust=-1000# /etc/systemd/system/quartermaster-ingest.service.d/oom-sacrifice.conf
[Service]
# Designate batch conversion workers as preferred sacrifice targets
OOMScoreAdjust=800
Nice=15
CPUSchedulingPolicy=idleSystem-Wide Kernel Memory Policy
In addition to per-process scoring adjustments, configure the host memory overcommit and cache pressure settings via sysctl:
# /etc/sysctl.d/99-memory-protection.conf
# Prohibit unbounded overcommit allocations (Mode 2)
vm.overcommit_memory = 2
# Ratio of RAM considered for overcommit allocation calculations
vm.overcommit_ratio = 80
# Panic and reboot rather than hanging indefinitely if kernel memory locks up
vm.panic_on_oom = 0
# Retain inode and dentry caches during memory pressure
vm.vfs_cache_pressure = 50Under vm.overcommit_memory = 2, the kernel refuses memory allocations exceeding (Swap + (RAM * vm.overcommit_ratio / 100)), returning ENOMEM to callers instead of permitting runaway memory consumption.
OOM Score Verification Script
Operators can verify the active OOM priority hierarchy across all running processes using this shell audit:
#!/usr/bin/env bash
# Audit process memory consumption and OOM sacrifice order
printf "%-8s %-6s %-12s %-10s %s
" "PID" "SCORE" "SCORE_ADJ" "RSS(KiB)" "COMMAND"
printf "---------------------------------------------------------
"
for pid in /proc/[0-9]*; do
[ -d "$pid" ] || continue
PID_NUM=$(basename "$pid")
if [ -r "$pid/oom_score" ] && [ -r "$pid/oom_score_adj" ]; then
SCORE=$(cat "$pid/oom_score" 2>/dev/null || echo "0")
SCORE_ADJ=$(cat "$pid/oom_score_adj" 2>/dev/null || echo "0")
RSS=$(grep VmRSS "$pid/status" 2>/dev/null | awk '{print $2}' || echo "0")
COMM=$(cat "$pid/comm" 2>/dev/null || echo "unknown")
printf "%-8s %-6s %-12s %-10s %s
" "$PID_NUM" "$SCORE" "$SCORE_ADJ" "$RSS" "$COMM"
fi
done | sort -k2 -n -r | head -n 15Running this audit verifies that critical supervisor daemons maintain negative adjustment scores, ensuring remote access remains available during high memory pressure incidents.
- Directus Target: on-the-line
- Garden Source Reference: kernel-out-of-memory-guardrails-and-process-priority-tuning, linux-oom-heuristics, systemd-resource-control, edge-reliability, MOC - Fleet Operations, MOC - Bosun PKM Tools