Living Document Notice
Published 2026-09-16. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Informing Harbormaster of Node Ceilings (Exposing memory and bus topology to the scheduler)
Summary
Workload schedulers make suboptimal placement decisions when physical constraints remain hidden. Overcommitting memory or assigning compute threads across NUMA boundaries degrades performance on constrained nodes. Quartermaster extracts node hardware boundaries into a normalized capacity document that Harbormaster consumes before dispatching jobs.
The Problem of Topology-Blind Scheduling
Workload orchestrators often treat compute nodes as homogeneous pools of floating-point units and addressable RAM. In modest server fleets composed of diverse hardware, this abstraction breaks down quickly. Placing a high-throughput static garden generation task on a node with constrained memory bandwidth causes process starvation. Similarly, scheduling threads across Non-Uniform Memory Access (NUMA) sockets incurs heavy inter-socket bus penalties.
The Harbormaster KPP protocol governs workload distribution across the Bosun infrastructure. To make informed placement decisions without running dynamic resource agents, Harbormaster relies on static capacity definitions exported by Quartermaster.
By publishing explicit memory ceilings, PCIe bus bandwidths, and core topologies, Quartermaster gives the scheduler complete visibility into physical hardware boundaries.
Hardware Capacity Document Schema
Quartermaster compiles physical hardware facts into a standardized JSON capacity document stored at /var/run/quartermaster/capacity.json. This document serves as the local contract read by Harbormaster agents.
{
"schema_version": "1.0",
"node_id": "edge-lon-02",
"timestamp_utc": "2026-09-20T00:00:00Z",
"resources": {
"memory": {
"physical_total_bytes": 34359738368,
"system_reserved_bytes": 2147483648,
"allocatable_bytes": 32212254720,
"swap_total_bytes": 0,
"numa_nodes": [
{
"node_id": 0,
"allocatable_bytes": 32212254720,
"cpu_cores": [0, 1, 2, 3, 4, 5, 6, 7]
}
]
},
"compute": {
"logical_cores": 16,
"physical_cores": 8,
"base_frequency_mhz": 2300,
"max_frequency_mhz": 3000,
"isolated_cores": []
},
"storage_pools": [
{
"pool_name": "fast-root",
"mount_point": "/",
"filesystem": "ext4",
"available_bytes": 858993459200,
"io_read_bandwidth_mbps": 3200,
"io_write_bandwidth_mbps": 2800
}
],
"networking": {
"egress_bandwidth_limit_mbps": 1000,
"max_connections_concurrent": 65535
}
}
}System Reservation and Memory Ceilings
To prevent runaway build processes or database memory spikes from causing node lockups, Quartermaster partitions system memory into strict operational envelopes. Schedulers must calculate placement based strictly on allocatable capacity rather than physical total RAM.
| Operating Layer | Allocation Category | Minimum Quota | Failure Consequence |
|---|---|---|---|
| Linux Kernel Slab | dentry and inode_cache | 512 MB | Page allocation stalls during batch directory scans |
| Directus Node Instance | V8 Heap Ceiling | 384 MB | Process termination via SIGABRT upon memory breach |
| Caddy Edge Proxy | Worker Socket Buffer | 64 MB | Connection reset (ECONNRESET) for incoming HTTP clients |
| Ephemeral Run Buffer | tmpfs scratch space | 256 MB | File descriptor write errors (ENOSPC) on large builds |
Scheduler Allocation and Constraint Matching
Harbormaster uses the capacity document to evaluate job requirements against node hardware constraints before dispatch.
| Workload Class | Critical Parameter | Hardware Floor | Scheduler Action on Breach |
|---|---|---|---|
| Directus CMS API | memory.allocatable_bytes | 1024 MB free | Cordon node from receiving new application tenants |
| Quartz Static Build | compute.physical_cores | 2 dedicated cores | Queue build or route to larger compute node |
| SQLite Write Operations | storage_pools[].io_write_bandwidth | 500 MB/s minimum | Reject placement on SATA SSD; require NVMe pool |
| Edge Proxy (Caddy) | networking.max_connections | 10000 open sockets | Restrict concurrent worker spawn limits |
Capacity Generation Pipeline
Quartermaster generates this document deterministically through a single shell command without third-party runtime daemons.
#!/bin/sh
set -eu
OUTPUT="/var/run/quartermaster/capacity.json"
mkdir -p "$(dirname "$OUTPUT")"
MEM_TOTAL_KB=$(grep MemTotal /proc/meminfo | awk '{print $2}')
MEM_TOTAL_BYTES=$((MEM_TOTAL_KB * 1024))
# Reserve 2 GB for kernel and base system daemons
RESERVED_BYTES=$((2 * 1024 * 1024 * 1024))
ALLOCATABLE_BYTES=$((MEM_TOTAL_BYTES - RESERVED_BYTES))
CORES_LOGICAL=$(nproc)
CORES_PHYSICAL=$(lscpu -p | grep -v '^#' | cut -d, -f2 | sort -u | wc -l)
cat <<EOF > "$OUTPUT"
{
"schema_version": "1.0",
"node_id": "$(cat /etc/quartermaster/node_id)",
"resources": {
"memory": {
"physical_total_bytes": $MEM_TOTAL_BYTES,
"system_reserved_bytes": $RESERVED_BYTES,
"allocatable_bytes": $ALLOCATABLE_BYTES
},
"compute": {
"logical_cores": $CORES_LOGICAL,
"physical_cores": $CORES_PHYSICAL
}
}
}
EOF
chmod 0644 "$OUTPUT"Harbormaster Verification Invocations
The scheduler queries node capacity over local domain sockets or secure HTTPS endpoints:
# Query capacity metrics for edge node
curl -s --unix-socket /var/run/quartermaster.sock http://localhost/capacity
# Validate that node memory is sufficient for incoming task allocation
quartermaster-cli capacity check --required-memory-mb 1536 --file /var/run/quartermaster/capacity.json- Directus Target: quartermaster
- Garden Source Reference: MOC - Fleet Operations
- Garden Source Reference: MOC - Bosun PKM Tools
- Garden Source Reference: [QTM-1007 - Informing Harbormaster of Node Ceilings (Exposing memory and bus topology to the scheduler)](QTM-1007 - Informing Harbormaster of Node Ceilings (Exposing memory and bus topology to the scheduler))