Living Document Notice
Published 2026-09-19. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

The Real Cost of High-Frequency Sampling

The Real Cost of High-Frequency Sampling: Viridian green P31 radar grid with dense high-frequency warm polar amber P20 sampling bursts and persistence bloom trail

Summary

Telemetry systems frequently encourage sub-second metric collection under the assumption that higher resolution always produces superior visibility. In resource-constrained environments, aggressive sampling frequencies trigger severe hardware trade-offs that compromise host stability.

Crow’s Nest balances metric precision against physical hardware costs. This dispatch analyzes disk write amplification, CPU wakeups, flash memory wear, cache invalidation, and clock drift caused by aggressive sampling intervals.

The Fallacy of Sub-Second Metric Collection

Collecting metrics every 100 milliseconds sounds attractive in architectural whitepapers. In production systems, querying kernel structures at 10Hz generates significant operational friction:

  • CPU cores cannot enter deep C-states, consuming excess power and inducing thermal throttling on small edge devices.
  • Write amplification on cheap flash media (SD cards, budget VPS SSDs) degrades storage lifespan.
  • Ingest pipelines accumulate millions of identical data points for static metrics (e.g., total RAM, root filesystem size).

Observability requires engineering discipline: measuring at the frequency required to take corrective action, and no faster.

Quantifying Write Amplification and Disk Churn

Consider a monitoring agent writing a 256-byte JSON metric record to disk every second:

While 21MB appears modest, file storage does not write bare bytes. Filesystems allocate storage in 4096-byte blocks. If an agent calls fsync() after every write to ensure durability:

On consumer flash storage or shared cloud hypervisors with strict IOPS limits, this write amplification degrades neighboring workloads and consumes write endurance:

Sampling IntervalDaily Raw DataDaily Disk Writes (fsync)CPU Wakeups / DayEstimated Flash Lifespan
100ms211MB3.37GB864,00011 months
1s21.1MB337.5MB86,400~3.8 years
10s (Crow’s Nest)2.1MB33.7MB8,640> 10 years
60s0.35MB5.6MB1,440> 10 years

Crow’s Nest defaults to a 10-second polling cadence and buffers log writes in memory, flushing to disk only on block boundaries.

CPU Cache Invalidation and TLB Shootdown

Every timer interrupt wakes the telemetry daemon, forcing the CPU out of low-power sleep states (such as C6/C7 idle states) back to high-power C0 execution. This transition expels cached application data from L1 and L2 CPU caches:

// Crow's Nest avoids high-frequency wakeups via timerfd
struct itimerspec new_value;
new_value.it_interval.tv_sec = 10;
new_value.it_interval.tv_nsec = 0;
new_value.it_value.tv_sec = 10;
new_value.it_value.tv_nsec = 0;
timerfd_settime(tfd, 0, &new_value, NULL);

Using Linux timerfd synchronized with epoll eliminates busy-wait polling loops and ensures the CPU remains idle between sampling epochs, avoiding context switch thrashing.

Clock Skew and Monotonic Timers

When sampling high-frequency events, comparing timestamps derived from gettimeofday() or CLOCK_REALTIME introduces distortion when NTP synchronizes machine clocks. NTP steps can cause time to jump backward, corrupting rate calculations:

Crow’s Nest calculates all delta metrics, sliding windows, and timeouts using CLOCK_MONOTONIC_RAW, which advances at a steady rate regardless of system time adjustments:

uint64_t get_monotonic_ns(void) {
    struct timespec ts;
    clock_gettime(CLOCK_MONOTONIC_RAW, &ts);
    return ((uint64_t)ts.tv_sec * 1000000000ULL) + (uint64_t)ts.tv_nsec;
}

Using monotonic timestamps prevents zero-division errors in rate formulas when system clocks step backwards during leap second or NTP slew events:

# Verify system timer wakeups caused by monitoring processes
perf stat -e 'sched:sched_switch' -p $(pgrep crows-nest) sleep 30

  • Directus Target: crows-nest
  • Garden Source Reference: MOC - Ingestion & Capture
  • Garden Source Reference: MOC - Fleet Operations
  • Garden Source Reference: MOC - Bosun PKM Tools
  • Garden Source Reference: [CRW-1010 - The Real Cost of High-Frequency Sampling](CRW-1010 - The Real Cost of High-Frequency Sampling)