Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Killing Alert Fatigue with Sliding Windows

Killing Alert Fatigue with Sliding Windows: Viridian green P31 concentric threshold tolerance bands with warm polar amber P20 dynamic sliding window damping brackets

Summary

Point-in-time alerting systems fire notifications immediately when a single probe fails or a CPU metric crosses a threshold. Transient network route flapping, momentary GC pauses, and cron-induced I/O spikes generate false alarms that erode engineer trust in on-call systems.

Crow’s Nest eliminates false alarms by evaluating alert criteria across sliding observation windows. This dispatch details the mathematical structure of M-of-N threshold evaluation, circular ring buffer implementation, sample loss edge cases, and state transition hysteresis.

The Pathology of Point-in-Time Alerts

When an alerting system evaluates a condition strictly as , any transient spike triggers an operational notification. A single lost ICMP packet or a 504 Gateway Timeout during an application restart wakes on-call operators without indicating a persistent degradation.

Engineers respond to noisy systems by disabling notifications or setting excessively loose thresholds. Both outcomes undermine reliability.

Alert StrategyFalse Positive RateMean Time to Detect (MTTD)Memory FootprintOperator Trust
Point-in-Time ThresholdHigh (24–38 / day)0 seconds (Immediate)4 bytesDegraded
Exponential Moving Average (EWMA)Moderate (4–9 / day)45–90 seconds16 bytesModerate
M-of-N Sliding WindowLow (0–1 / day)Deterministic ()64 bytesHigh

Crow’s Nest rejects point-in-time triggers. A condition must persist across a specified window of observation samples before the daemon dispatches an incident ticket.

Discrete M-of-N Sliding Window Evaluation

The sliding window algorithm maintains a fixed-capacity ring buffer containing the boolean outcome of the last probe cycles. An alert state transition occurs only when at least of the historical samples violate the defined threshold.

Where . For standard API surveillance against the Harbor edge runtime:

  • Probe interval:
  • Window capacity: (30 seconds total evaluation horizon)
  • Breach threshold:

A single dropped packet or single 502 response sets 1 failure in the ring buffer (), keeping the monitor in state OK.

Ring Buffer Implementation in Memory

The probe window is implemented as a bitmask or a static array of unsigned integers, avoiding heap allocations:

#define WINDOW_CAPACITY 8
 
typedef struct {
    uint8_t buffer[WINDOW_CAPACITY];
    uint8_t head;
    uint8_t count;
    uint8_t threshold_m;
    uint8_t current_state; // 0 = OK, 1 = BREACH
} SlidingWindowFilter;
 
int record_sample(SlidingWindowFilter *filter, int is_failure) {
    filter->buffer[filter->head] = is_failure ? 1 : 0;
    filter->head = (filter->head + 1) % WINDOW_CAPACITY;
    if (filter->count < WINDOW_CAPACITY) {
        filter->count++;
    }
 
    // Evaluate M-of-N condition
    uint8_t failure_count = 0;
    for (uint8_t i = 0; i < filter->count; i++) {
        failure_count += filter->buffer[i];
    }
 
    uint8_t new_state = (failure_count >= filter->threshold_m) ? 1 : 0;
    int state_changed = (new_state != filter->current_state);
    filter->current_state = new_state;
 
    return state_changed ? (new_state ? 1 : -1) : 0;
}

The function returns 1 when entering a breach state and -1 when recovering, suppressing repetitive duplicate notices during sustained downtime.

State Hysteresis and Recovery Windows

To prevent alert bouncing at threshold boundaries (such as disk usage hovering between 89.9% and 90.1%), Crow’s Nest enforces asymmetric recovery thresholds.

# /etc/crows-nest/alerts.toml
[alert.disk_root]
metric = "disk_used_pct"
target = "/"
trigger_threshold = 90.0
trigger_window_m = 5
trigger_window_n = 5
 
recovery_threshold = 85.0
recovery_window_m = 6
recovery_window_n = 6

Clearing the alert requires disk usage to fall below 85.0% for six consecutive observation periods. If a probe cycle times out due to packet loss, it is counted as a failure, preventing network stalls from masquerading as healthy recoveries.

Operators verify active alert window states via local daemon inspection:

crows-nest alerts --status --window-dump

  • Directus Target: crows-nest
  • Garden Source Reference: MOC - Ingestion & Capture
  • Garden Source Reference: MOC - Fleet Operations
  • Garden Source Reference: MOC - Bosun PKM Tools
  • Garden Source Reference: [CRW-1005 - Killing Alert Fatigue with Sliding Windows](CRW-1005 - Killing Alert Fatigue with Sliding Windows)