Living Document Notice
Published 2026-09-11. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Fail-Closed Outbox Logging Surviving Process Termination Mid-Save

Summary

Abrupt operating system process termination during file write cycles poses a persistent risk of database corruption and index desynchronization. By establishing an append-only JSON Lines outbox journal, Bosun PKM guarantees zero knowledge loss during crashes. Filesystem mutations advance through deterministic lifecycle states, allowing the runtime to reconcile incomplete writes, handle editor swap files, and restore database projections upon restart.

Outbox Operation Lifecycle

Operating systems terminate background processes under memory pressure (OOM killer), battery depletion, or abrupt user termination. Updating database projections directly inside filesystem watcher callbacks exposes the system to torn writes.

To decouple observation from projection updates, the storage engine serializes all filesystem mutations into an append-only transaction journal (outbox.jsonl) prior to updating derived SQLite projections:

[Filesystem Event]
       │
       ▼
 [Pending State]     ──> Written to outbox.jsonl with operation UUID and timestamp
       │
       ▼
[Processing State]   ──> File payload read, SHA-256 computed, frontmatter validated
       │
       ├───────────────────────────────┐
       ▼                               ▼
[Committed State]              [Failed State]
- SQLite projection updated    - Recorded with typed error code
- Dedup ledger updated         - Projection remains untouched

Every operation record in outbox.jsonl contains immutable state attributes:

{
  "op_id": "9b1deb4d-3b7d-4bad-9bdd-2b0d7b3dcb6d",
  "locker_id": "lkr_01j7m8x4k2",
  "path": "engineering/storage-invariants.md",
  "event_type": "modify",
  "status": "committed",
  "content_hash": "a4f89d6e8b2c4f1a...",
  "created_at": "2026-09-11T14:22:05.120Z",
  "updated_at": "2026-09-11T14:22:05.145Z",
  "recovery_reason": null
}

Atomic Save-via-Rename Handling

Modern text editors and integrated development environments rarely overwrite files in place. Instead, editors follow an atomic save pattern: write modified contents to an ephemeral hidden file (e.g. .note.md.tmp or 4913), flush buffers to disk, and issue a POSIX rename() over the target path.

Naively processing file watcher events on hidden temporary files introduces state churn and race conditions:

# Ephemeral editor file filter
TEMPORARY_FILE_PATTERNS = (
    re.compile(r"^\..*\.tmp$"),
    re.compile(r".*~$"),
    re.compile(r"^\.#.*$"),
    re.compile(r"^[0-9]{4,}$"),
)
 
def is_transient_editor_buffer(relative_path: str) -> bool:
    filename = Path(relative_path).name
    return any(pattern.match(filename) for pattern in TEMPORARY_FILE_PATTERNS)

The storage engine ignores create events on transient editor buffers. When the editor issues the final rename() onto the destination path, the engine captures the target event, checks content hashes, and updates projections cleanly.

Crash Recovery and Replay Algorithm

On engine initialization, the recovery subsystem scans outbox.jsonl for operations left in the processing state:

def recover_crashed_operations(self) -> int:
    unresolved_ops = self.outbox.get_unresolved_operations()
    recovered_count = 0
 
    for op in unresolved_ops:
        target_path = self.locker_root / op.path
 
        if not target_path.exists():
            op.transition_to_failed(reason="crash_recovery_file_missing")
            self.outbox.update_operation(op)
            continue
 
        disk_hash = self.compute_sha256(target_path)
        if disk_hash == op.content_hash:
            # File write completed successfully prior to crash
            self.projection.upsert_note(target_path, disk_hash)
            op.transition_to_committed(reason="crash_recovery_hash_match")
            recovered_count += 1
        else:
            # File was partially written or modified by external tool
            op.transition_to_failed(reason="crash_recovery_hash_mismatch")
 
        self.outbox.update_operation(op)
 
    return recovered_count

The table below outlines failure modes and corresponding crash recovery outcomes:

Crash ScenarioState at InterruptionStartup Reconciliation Action
Kill during disk writestatus = "processing"SHA-256 mismatch detected; operation marked failed; file re-read on next touch
Kill after disk write, before SQLite updatestatus = "processing"SHA-256 matches; SQLite projection transactionally updated; marked committed
Corrupt JSONL entryIncomplete tail lineTail record discarded via EOF validation; previous committed entries preserved
Database file deletionMissing projection.dbOutbox journal replayed in full to reconstruct SQLite projection

The recovery process executes in sub-10 milliseconds for journals containing up to 10,000 recorded mutations, guaranteeing that vault state remains consistent without manual repair.


  • Directus Target: bosun-pkm
  • Garden Source Reference: MOC - Bosun PKM Tools, MOC - Bosun PKM Engine, MOC - The Plain-Text Longevity Standard, [BSN-1002 - Deterministic Round-Trip Serialization](BSN-1002 - Deterministic Round-Trip Serialization)