Living Document Notice
Published 2026-09-13. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Frontmatter as a Strongly Typed Schema
Summary
Treating YAML frontmatter as unstructured key-value dictionaries propagates schema decay across large vaults. As note collections expand, variations in date formats, tag types, and custom status fields cause silent data corruption in publication scripts and querying engines.
Bosun enforces strongly typed frontmatter models at parse time using Rust compile-time validation rules and custom byte-span tracking. By decoupling schema validation from full YAML AST allocation, the parser detects type violations and missing properties while preserving original source offsets for inline editor diagnostics.
Zero-Copy Frontmatter Slicing
Frontmatter extraction begins by locating the opening and closing YAML delimiters without loading external parsers. The scanner requires opening --- tokens at byte offset 0, followed by a terminating --- or ... marker on an isolated line.
pub struct FrontmatterSpan<'a> {
pub raw_yaml: &'a str,
pub start_byte: usize,
pub end_byte: usize,
pub body_start_byte: usize,
}This slicing operation evaluates only the leading bytes of the document buffer. If no frontmatter delimiters exist within the initial 4,096 bytes, the scanner exits early, skipping YAML initialization completely for raw markdown files.
The extracted slice references original document memory without copying bytes into temporary string allocations. This zero-copy design reduces initial file inspection overhead during startup scans.
Strongly Typed Models with Serde Validation
Extracted YAML slices deserialize directly into strongly typed structs defined with Serde attributes. Fields support strict typing, optional defaults, and fallback maps for unrecognized metadata keys.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct DocumentMetadata {
pub title: String,
pub date: chrono::NaiveDate,
#[serde(default)]
pub draft: bool,
#[serde(default)]
pub tags: Vec<String>,
#[serde(default)]
pub blog: Option<String>,
#[serde(flatten)]
pub extra: BTreeMap<String, serde_yaml::Value>,
}When a document specifies an invalid date format such as date: 2026/09/13 instead of ISO-8601 2026-09-13, deserialization halts with a structured error indicating the exact field and type expectation.
Schema definitions are compiled directly into binary format, eliminating runtime JSON schema parsing penalties during vault validation sweeps.
Precise Diagnostic Reporting with Source Spans
To display errors directly beneath editor text, deserialization errors must map to line and column numbers in the source file. Bosun pairs Serde deserialization with a custom span tracking parser.
pub struct Diagnostic {
pub line: u32,
pub column: u32,
pub byte_offset: u32,
pub message: String,
pub severity: DiagnosticSeverity,
}The error recorder computes line offsets using a monotonic byte-to-line index. This guarantees that diagnostics match editor cursor positions without rescanning file lines from disk.
Inline diagnostic messages provide exact line coordinates and suggest schema-compliant replacements, such as correcting malformed boolean values or unrecognized status codes.
Validation Throughput and Memory Allocations
The benchmark table below evaluates validation speeds between dynamic serde_yaml::Value maps and Bosun strongly typed schema extraction.
| Benchmark Configuration | Documents Parsed | Total Throughput (docs/sec) | Average Latency | Heap Allocations per Note |
|---|---|---|---|---|
| Untyped serde_yaml::Value | 10,000 | 18,400 docs/s | 54.3 μs | 42 allocations |
| Standard serde_yaml Typed | 10,000 | 28,100 docs/s | 35.5 μs | 18 allocations |
| Bosun Zero-Copy Scanner | 10,000 | 78,500 docs/s | 12.7 μs | 4 allocations |
| Bosun Cached Span Model | 10,000 | 114,200 docs/s | 8.7 μs | 2 allocations |
- Directus Target: bosunpkm-blog
- Garden Source Reference: MOC - Bosun PKM Engine, MOC - Bosun PKM Tools, MOC - The Plain-Text Longevity Standard