Living Document Notice
Published 2026-09-17. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Piping Liberated Data into Bosun PKM
Summary
Converting normalized JSON Lines into Bosun PKM nodes with bi-directional wikilinks, parsed tags, and typed frontmatter.
This technical dispatch explores the underlying architecture, data structures, and concrete implementation boundaries required for local-first data sovereignty.
The Contract Between Extraction and Ingestion
Liberation utilities should never produce loose, unstructured text blobs. When dumping thousands of historical notes, tasks, and bookmarks into a personal knowledge base, dumping unformatted files creates immediate cognitive clutter.
To ensure long-term utility, FreeMyData enforces an ingestion contract with Bosun PKM Tools. Extracted payloads are emitted as a newline-delimited JSON (JSONL) stream where each line contains a validated document entity:
{
"source_platform": "freemydata_extractor",
"external_id": "notion_8f3b6c2a",
"created_utc": "2026-04-12T14:22:00Z",
"modified_utc": "2026-09-10T11:05:00Z",
"title": "Compilers and Parser Invariants",
"tags": ["compilers", "rust", "ast"],
"relations": [
{"target_title": "Abstract Syntax Tree Transforms", "relation_type": "parent"}
],
"blocks": [
{"type": "heading_2", "text": "Lexical Analysis"},
{"type": "paragraph", "text": "Tokenization breaks input strings into tokens.", "block_id": "blk_a190"}
]
}The ingestion pipeline transforms this stream into atomic Markdown files adhering to Bosun PKM digital garden standards.
Normalizing Metadata and Wikilink Graph Traversal
When materializing documents on disk, the ingestion engine executes three sequential operations:
- Deterministic Frontmatter Assembly: Emits standardized YAML frontmatter including ISO-8601 timestamps, source provenance hashes, and taxonomy tags.
- Bidirectional Relationship Mapping: Translates abstract relation arrays into bracketed wikilink syntax placed exclusively in dedicated metadata sections below the cut line.
- Block Anchor Generation: Generates deterministic block identifiers (
^blk-xxxxx) based on the CRC32 hash of the block text and its sequence offset, allowing deep intra-document referencing.
Transformation Flow:
JSONL Record ──> Canonical Slug ──> YAML Header ──> AST Body ──> Cut-Line References
The table below outlines how vendor fields map into standard Bosun PKM vault formats:
| Extraction Payload Field | Bosun PKM Markdown Target | Validation Invariant |
|---|---|---|
title | File Basename & # Heading 1 | Sanitized against invalid OS filename chars |
modified_utc | date: YAML frontmatter | Format: 2026-01-01 (ISO-8601) |
tags | tags: YAML list | Lowercase, kebab-case, no symbols |
relations.parent | Bracketed wikilink below cut line | Must resolve to existing vault document |
blocks.block_id | ^blk-xxxx at end of paragraph | 8-character deterministic hex string |
Ingestion Pipeline Implementation
The following Python script reads the JSONL stream, verifies content hashes to avoid overwriting modified local notes, and writes clean Markdown files to the target vault directory:
import sys
import json
import hashlib
from pathlib import Path
def generate_block_id(text: str, index: int) -> str:
h = hashlib.crc32(f"{index}:{text.strip()}".encode("utf-8"))
return f"blk-{h:08x}"
def ingest_stream(jsonl_path: Path, vault_dir: Path):
vault_dir.mkdir(parents=True, exist_ok=True)
with open(jsonl_path, "r", encoding="utf-8") as fp:
for line in fp:
if not line.strip():
continue
record = json.loads(line)
clean_title = "".join(c for c in record["title"] if c.isalnum() or c in " -_").strip()
target_path = vault_dir / f"{clean_title}.md"
frontmatter = [
"---",
f'title: "{clean_title}"',
f'date: {record["modified_utc"][:10]}',
f'source_id: "{record["external_id"]}"',
"tags:",
" - liberated-data"
]
for t in record.get("tags", []):
frontmatter.append(f" - {t}")
frontmatter.append("---\n")
body_lines = [f"# {clean_title}\n"]
for idx, block in enumerate(record.get("blocks", [])):
if block["type"] == "heading_2":
body_lines.append(f"\n## {block['text']}\n")
elif block["type"] == "paragraph":
blk_id = generate_block_id(block["text"], idx)
body_lines.append(f"{block['text']} ^{blk_id}\n")
body_lines.append("\n---\n")
body_lines.append(f"* **Source Platform**: {record['source_platform']}")
open_b, close_b = "[" + "[", "]" + "]"
refs = [f"{open_b}{r['target_title']}{close_b}" for r in record.get("relations", [])]
if refs:
body_lines.append(f"* **Garden Source Reference**: {', '.join(refs)}")
full_content = "\n".join(frontmatter) + "\n".join(body_lines) + "\n"
target_path.write_text(full_content, encoding="utf-8")
if __name__ == "__main__":
ingest_stream(Path("liberated_feed.jsonl"), Path("02 Review/freemydata"))# Execute stream ingestion and verify output node count
python ingest_pipeline.py --input liberated_feed.jsonl --target "02 Review/freemydata"
test $(ls -1 "02 Review/freemydata"/*.md | wc -l) -gt 0- Directus Target: freemydata
- Garden Source Reference: [BSN-1001 - Incremental Parsing with Tree-sitter](BSN-1001 - Incremental Parsing with Tree-sitter), DAT-1003 - Deterministic Lossless Conversion, MOC - Data Liberation Workbenches, MOC - The Plain-Text Longevity Standard, MOC - Bosun PKM Tools, [BSN-1002 - Deterministic Round-Trip Serialization](BSN-1002 - Deterministic Round-Trip Serialization)