Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers)
Summary
Solid-state drive degradation represents a primary hardware failure mode in continuous hosting environments. Unchecked write amplification and high sustained write volumes can exhaust drive endurance without operator notification. Quartermaster establishes automated health checks that monitor NVMe SMART logs, evaluate Terabytes Written against manufacturer ceilings, and trigger replacement procedures.
The Physical Reality of Flash Cell Degradation
NAND flash memory cells withstand a finite number of program-erase cycles before oxide layer degradation prevents reliable charge retention. In multi-tenant environments hosting dynamic logs, embedded SQLite write-ahead logs, and continuous deployment builds, write amplification can deplete consumer or entry-level enterprise SSD endurance rapidly.
Failing to track wear leaves infrastructure vulnerable to abrupt read-only drive lockups or silent sector corruptions. While traditional rotating disks often signaled failure through audible motor friction or progressive bad sector reallocations, NVMe drives often transition into emergency write-lock states instantaneously when spare flash blocks exhaust.
Quartermaster continuously reads the NVMe management log interface to track physical wear metrics, calculating consumed endurance long before write thresholds breach.
Critical NVMe Telemetry Attributes
The NVMe specification standardizes health reporting through the SMART / Health Information log page. Quartermaster interrogates this log using nvme-cli or low-level ioctl calls to collect critical wear indicators.
The key telemetry fields monitored include:
percentage_used: An estimate of vendor-rated life consumed, where 100 indicates the drive has reached its rated endurance.available_spare: Normalized percentage of remaining spare flash capacity available to replace degraded blocks.data_units_written: Number of 512-byte data units written, represented as a 128-bit integer in thousands (each unit equals 512,000 bytes).critical_warning: Bitmask indicating spare capacity failure, temperature thresholds, or subsystem reliability issues.media_errors: Unrecovered data integrity errors encountered by the internal drive controller.
Health Threshold and Action Matrix
Quartermaster classifies drive health into four explicit operational states based on physical telemetry thresholds.
| Telemetry Metric | Register Name | Normal Baseline | Warning Threshold | Critical Replacement Action |
|---|---|---|---|---|
| Spare Block Reserve | available_spare | 100% | Under 20% | Immediate node cordon and drive swap |
| Consumed Life | percentage_used | 0% to 80% | 85% to 95% | Schedule replacement maintenance window |
| Hardware Integrity | critical_warning | 0x00 | Any non-zero bit | Evacuate active databases immediately |
| Media Unrecoverable | media_errors | 0 errors | Greater than 0 | Mark storage subsystem as degraded |
| Operating Temperature | temperature | 30°C to 55°C | 60°C to 69°C | Increase chassis fan duty cycle |
Automated NVMe Extraction Script
The following POSIX shell script extracts NVMe SMART metrics without requiring heavy monitoring daemons. It calculates cumulative Terabytes Written (TBW) directly from raw controller counters.
#!/bin/sh
set -eu
DEVICE="${1:-/dev/nvme0}"
if [ ! -e "$DEVICE" ]; then
echo "ERROR: Block device $DEVICE does not exist" >&2
exit 1
fi
# Extract SMART log in JSON format using nvme-cli
LOG_JSON=$(nvme smart-log "$DEVICE" -o json)
PERCENT_USED=$(echo "$LOG_JSON" | grep -o '"percent_used":[0-9]*' | cut -d: -f2)
SPARE_RESERVE=$(echo "$LOG_JSON" | grep -o '"avail_spare":[0-9]*' | cut -d: -f2)
CRIT_WARN=$(echo "$LOG_JSON" | grep -o '"critical_warning":[0-9]*' | cut -d: -f2)
DATA_UNITS=$(echo "$LOG_JSON" | grep -o '"data_units_written":[0-9]*' | cut -d: -f2)
# Calculate Terabytes Written (TBW)
# Each unit = 512,000 bytes. TBW = (DATA_UNITS * 512000) / 10^12
TBW=$(awk -v units="$DATA_UNITS" 'BEGIN { printf "%.2f", (units * 512000) / 1000000000000 }')
echo "Storage Report for $DEVICE:"
echo " Percentage Used: ${PERCENT_USED}%"
echo " Available Spare: ${SPARE_RESERVE}%"
echo " Critical Warning Mask: ${CRIT_WARN}"
echo " Total Terabytes Written: ${TBW} TBW"
# Evaluate critical triggers
if [ "$CRIT_WARN" -ne 0 ] || [ "$SPARE_RESERVE" -lt 10 ]; then
echo "STATE: CRITICAL - Drive replacement required immediately" >&2
exit 2
elif [ "$PERCENT_USED" -ge 90 ] || [ "$SPARE_RESERVE" -lt 20 ]; then
echo "STATE: WARNING - Drive endurance nearing operational limit" >&2
exit 1
fi
echo "STATE: HEALTHY"Storage Degradation Drainage Invocations
When a drive enters the warning state, operators drain local services before physical replacement.
# Flush pending filesystem writes to disk
sync && fsfreeze -f /data && fsfreeze -u /data
# Unmount persistent volume cleanly
umount /data
# Put NVMe device into offline status prior to hot-swap
echo 1 > /sys/block/nvme0n1/device/device/remove- Directus Target: quartermaster
- Garden Source Reference: MOC - Fleet Operations
- Garden Source Reference: MOC - Bosun PKM Tools
- Garden Source Reference: [QTM-1005 - Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers)](QTM-1005 - Tracking Storage Wear and NVMe Health (Parsing SMART TBW, percentage used, replacement triggers))