Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers)

Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers): Emerald green P31 logarithmic wear endurance curve with warm golden-amber P20 replacement threshold alarm line and wear gauge bars

Summary

Solid-state drive degradation represents a primary hardware failure mode in continuous hosting environments. Unchecked write amplification and high sustained write volumes can exhaust drive endurance without operator notification. Quartermaster establishes automated health checks that monitor NVMe SMART logs, evaluate Terabytes Written against manufacturer ceilings, and trigger replacement procedures.

The Physical Reality of Flash Cell Degradation

NAND flash memory cells withstand a finite number of program-erase cycles before oxide layer degradation prevents reliable charge retention. In multi-tenant environments hosting dynamic logs, embedded SQLite write-ahead logs, and continuous deployment builds, write amplification can deplete consumer or entry-level enterprise SSD endurance rapidly.

Failing to track wear leaves infrastructure vulnerable to abrupt read-only drive lockups or silent sector corruptions. While traditional rotating disks often signaled failure through audible motor friction or progressive bad sector reallocations, NVMe drives often transition into emergency write-lock states instantaneously when spare flash blocks exhaust.

Quartermaster continuously reads the NVMe management log interface to track physical wear metrics, calculating consumed endurance long before write thresholds breach.

Critical NVMe Telemetry Attributes

The NVMe specification standardizes health reporting through the SMART / Health Information log page. Quartermaster interrogates this log using nvme-cli or low-level ioctl calls to collect critical wear indicators.

The key telemetry fields monitored include:

  • percentage_used: An estimate of vendor-rated life consumed, where 100 indicates the drive has reached its rated endurance.
  • available_spare: Normalized percentage of remaining spare flash capacity available to replace degraded blocks.
  • data_units_written: Number of 512-byte data units written, represented as a 128-bit integer in thousands (each unit equals 512,000 bytes).
  • critical_warning: Bitmask indicating spare capacity failure, temperature thresholds, or subsystem reliability issues.
  • media_errors: Unrecovered data integrity errors encountered by the internal drive controller.

Health Threshold and Action Matrix

Quartermaster classifies drive health into four explicit operational states based on physical telemetry thresholds.

Telemetry MetricRegister NameNormal BaselineWarning ThresholdCritical Replacement Action
Spare Block Reserveavailable_spare100%Under 20%Immediate node cordon and drive swap
Consumed Lifepercentage_used0% to 80%85% to 95%Schedule replacement maintenance window
Hardware Integritycritical_warning0x00Any non-zero bitEvacuate active databases immediately
Media Unrecoverablemedia_errors0 errorsGreater than 0Mark storage subsystem as degraded
Operating Temperaturetemperature30°C to 55°C60°C to 69°CIncrease chassis fan duty cycle

Automated NVMe Extraction Script

The following POSIX shell script extracts NVMe SMART metrics without requiring heavy monitoring daemons. It calculates cumulative Terabytes Written (TBW) directly from raw controller counters.

#!/bin/sh
set -eu
 
DEVICE="${1:-/dev/nvme0}"
 
if [ ! -e "$DEVICE" ]; then
  echo "ERROR: Block device $DEVICE does not exist" >&2
  exit 1
fi
 
# Extract SMART log in JSON format using nvme-cli
LOG_JSON=$(nvme smart-log "$DEVICE" -o json)
 
PERCENT_USED=$(echo "$LOG_JSON" | grep -o '"percent_used":[0-9]*' | cut -d: -f2)
SPARE_RESERVE=$(echo "$LOG_JSON" | grep -o '"avail_spare":[0-9]*' | cut -d: -f2)
CRIT_WARN=$(echo "$LOG_JSON" | grep -o '"critical_warning":[0-9]*' | cut -d: -f2)
DATA_UNITS=$(echo "$LOG_JSON" | grep -o '"data_units_written":[0-9]*' | cut -d: -f2)
 
# Calculate Terabytes Written (TBW)
# Each unit = 512,000 bytes. TBW = (DATA_UNITS * 512000) / 10^12
TBW=$(awk -v units="$DATA_UNITS" 'BEGIN { printf "%.2f", (units * 512000) / 1000000000000 }')
 
echo "Storage Report for $DEVICE:"
echo "  Percentage Used:       ${PERCENT_USED}%"
echo "  Available Spare:       ${SPARE_RESERVE}%"
echo "  Critical Warning Mask: ${CRIT_WARN}"
echo "  Total Terabytes Written: ${TBW} TBW"
 
# Evaluate critical triggers
if [ "$CRIT_WARN" -ne 0 ] || [ "$SPARE_RESERVE" -lt 10 ]; then
  echo "STATE: CRITICAL - Drive replacement required immediately" >&2
  exit 2
elif [ "$PERCENT_USED" -ge 90 ] || [ "$SPARE_RESERVE" -lt 20 ]; then
  echo "STATE: WARNING - Drive endurance nearing operational limit" >&2
  exit 1
fi
 
echo "STATE: HEALTHY"

Storage Degradation Drainage Invocations

When a drive enters the warning state, operators drain local services before physical replacement.

# Flush pending filesystem writes to disk
sync && fsfreeze -f /data && fsfreeze -u /data
 
# Unmount persistent volume cleanly
umount /data
 
# Put NVMe device into offline status prior to hot-swap
echo 1 > /sys/block/nvme0n1/device/device/remove

  • Directus Target: quartermaster
  • Garden Source Reference: MOC - Fleet Operations
  • Garden Source Reference: MOC - Bosun PKM Tools
  • Garden Source Reference: [QTM-1005 - Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers)](QTM-1005 - Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers))