Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Structured Incident Postmortems Without Blame or Retrospective Fiction
Summary
Post-incident analysis fails when teams reconstruct timelines from human memory rather than verifiable machine telemetry and shell logs. A grounded postmortem standardizes timeline ingestion around timestamped kernel logs, exact shell history, and causal dependency graphs rather than narrative rationalizations.
The Hazard of Narrative Reconstruction
Human recollections of live outages are shaped by cognitive bias, fatigue, and hindsight rationalization. When operators construct incident reports days after an event without machine-verifiable timestamps, critical details are smoothed over or reinterpreted to fit expected explanations.
Operational integrity requires anchoring every incident postmortem in verifiable machine records: POSIX shell logs, systemd journal timestamps, network interface counters, and packet traces. Replacing narrative storytelling with structured causality graphs turns post-incident analysis into reproducible systems engineering.
+-------------------------------------------------------------+
| Telemetry-Anchored Timeline Ingest |
| |
| 1. Systemd Journal Ingest (journalctl -o json) |
| | |
| 2. Reverse Proxy Access Logs (ISO-8601 UTC) |
| | |
| 3. Shell Command Audit Ledger (pam_tty_audit) |
| v |
| [Chronological Normalization & Merging Engine] |
| v |
| [Deterministic Four-Phase State Machine] |
| - Phase 1: Detection |
| - Phase 2: Containment |
| - Phase 3: Eradication |
| - Phase 4: Full Recovery |
+-------------------------------------------------------------+
The Four-Phase Incident Taxonomy
Every postmortem document organizes events into four non-overlapping chronological phases:
| Phase | Phase Boundary Definition | Verifiable Telemetry Artifact |
|---|---|---|
| 1. Detection | From first anomalous telemetry metric to active human acknowledgement | Watch log entry or Prometheus alert fire event |
| 2. Containment | From responder triage to isolation of blast radius | BGP withdrawal, NGINX upstream drain, or firewall rule |
| 3. Eradication | From isolation to root issue elimination | Service patch deployment or corrupt table restore |
| 4. Recovery | From root fix to complete return to normal baseline | Telemetry metrics return within SLA thresholds for 60m |
Standard Postmortem Note Schema
Postmortems are maintained in plain-text markdown, version-controlled directly in the operations repository alongside infrastructure code:
---
incident_id: "INC-20260913-01"
severity: "SEV-1"
start_time: "2026-09-13T04:12:08Z"
mitigated_time: "2026-09-13T04:38:15Z"
resolved_time: "2026-09-13T05:22:40Z"
impacted_services:
- "harbormaster-sync"
- "quartermaster-cdn"
lead_investigator: "ops-duty-node-4"
---
# Incident Analysis: INC-20260913-01
## Impact Summary
Between 04:12 UTC and 04:38 UTC, 14.2% of outbound vault synchronizations failed with HTTP 504 gateway timeouts. Inbound note reads remained unaffected. No data loss occurred.
## Causal Sequence
- **Proximate Trigger**: A scheduled cron job initiated a full-table vacuum on the sync catalog during peak sync traffic.
- **Latent Defect**: Connection pool sizing in Harbormaster lacked query timeout ceilings, allowing long-running table locks to starve incoming read requests.
- **Tooling Gap**: Grafana database connection alerts were set to evaluate over a 15-minute window rather than detecting sudden pool saturation.Timeline Ingest and Verification Script
To prevent timeline errors, this Python utility pulls journal events and validates that documented timestamps match recorded kernel entries:
import json
import subprocess
import sys
from datetime import datetime
def verify_journal_timeline(unit_name: str, start_iso: str, end_iso: str):
cmd = [
"journalctl",
f"-u", unit_name,
f"--since={start_iso}",
f"--until={end_iso}",
"-o", "json"
]
result = subprocess.run(cmd, capture_output=True, text=True, check=True)
entries = []
for line in result.stdout.strip().split("
"):
if not line:
continue
record = json.loads(line)
usec = int(record["__REALTIME_TIMESTAMP"])
dt = datetime.utcfromtimestamp(usec / 1e6)
msg = record.get("MESSAGE", "")
entries.append((dt.isoformat() + "Z", msg))
print(f"Verified {len(entries)} matching journal log frames.")
for ts, msg in entries[:5]:
print(f"[{ts}] {msg[:80]}")
if __name__ == "__main__":
verify_journal_timeline("harbormaster-sync", "2026-09-13 04:12:00", "2026-09-13 04:40:00")Systematic log-based verification prevents speculative narratives and keeps operational documentation grounded in observable facts.
- Directus Target: on-the-line
- Garden Source Reference: structured-incident-postmortems-without-blame-or-retrospective-fiction, incident-management, telemetry-correlation, root-cause-analysis, MOC - Fleet Operations, MOC - Bosun PKM Tools