Living Document Notice
Published 2026-09-17. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Piping Liberated Data into Bosun PKM

Piping Liberated Data into Bosun PKM: Electric lime P1 vector CRT macro showing parallel incoming data pipelines funneling into an interconnected knowledge graph network

Summary

Converting normalized JSON Lines into Bosun PKM nodes with bi-directional wikilinks, parsed tags, and typed frontmatter.

This technical dispatch explores the underlying architecture, data structures, and concrete implementation boundaries required for local-first data sovereignty.

The Contract Between Extraction and Ingestion

Liberation utilities should never produce loose, unstructured text blobs. When dumping thousands of historical notes, tasks, and bookmarks into a personal knowledge base, dumping unformatted files creates immediate cognitive clutter.

To ensure long-term utility, FreeMyData enforces an ingestion contract with Bosun PKM Tools. Extracted payloads are emitted as a newline-delimited JSON (JSONL) stream where each line contains a validated document entity:

{
  "source_platform": "freemydata_extractor",
  "external_id": "notion_8f3b6c2a",
  "created_utc": "2026-04-12T14:22:00Z",
  "modified_utc": "2026-09-10T11:05:00Z",
  "title": "Compilers and Parser Invariants",
  "tags": ["compilers", "rust", "ast"],
  "relations": [
    {"target_title": "Abstract Syntax Tree Transforms", "relation_type": "parent"}
  ],
  "blocks": [
    {"type": "heading_2", "text": "Lexical Analysis"},
    {"type": "paragraph", "text": "Tokenization breaks input strings into tokens.", "block_id": "blk_a190"}
  ]
}

The ingestion pipeline transforms this stream into atomic Markdown files adhering to Bosun PKM digital garden standards.

When materializing documents on disk, the ingestion engine executes three sequential operations:

  1. Deterministic Frontmatter Assembly: Emits standardized YAML frontmatter including ISO-8601 timestamps, source provenance hashes, and taxonomy tags.
  2. Bidirectional Relationship Mapping: Translates abstract relation arrays into bracketed wikilink syntax placed exclusively in dedicated metadata sections below the cut line.
  3. Block Anchor Generation: Generates deterministic block identifiers (^blk-xxxxx) based on the CRC32 hash of the block text and its sequence offset, allowing deep intra-document referencing.
Transformation Flow:
JSONL Record ──> Canonical Slug ──> YAML Header ──> AST Body ──> Cut-Line References

The table below outlines how vendor fields map into standard Bosun PKM vault formats:

Extraction Payload FieldBosun PKM Markdown TargetValidation Invariant
titleFile Basename & # Heading 1Sanitized against invalid OS filename chars
modified_utcdate: YAML frontmatterFormat: 2026-01-01 (ISO-8601)
tagstags: YAML listLowercase, kebab-case, no symbols
relations.parentBracketed wikilink below cut lineMust resolve to existing vault document
blocks.block_id^blk-xxxx at end of paragraph8-character deterministic hex string

Ingestion Pipeline Implementation

The following Python script reads the JSONL stream, verifies content hashes to avoid overwriting modified local notes, and writes clean Markdown files to the target vault directory:

import sys
import json
import hashlib
from pathlib import Path
 
def generate_block_id(text: str, index: int) -> str:
    h = hashlib.crc32(f"{index}:{text.strip()}".encode("utf-8"))
    return f"blk-{h:08x}"
 
def ingest_stream(jsonl_path: Path, vault_dir: Path):
    vault_dir.mkdir(parents=True, exist_ok=True)
    
    with open(jsonl_path, "r", encoding="utf-8") as fp:
        for line in fp:
            if not line.strip():
                continue
            record = json.loads(line)
            clean_title = "".join(c for c in record["title"] if c.isalnum() or c in " -_").strip()
            target_path = vault_dir / f"{clean_title}.md"
            
            frontmatter = [
                "---",
                f'title: "{clean_title}"',
                f'date: {record["modified_utc"][:10]}',
                f'source_id: "{record["external_id"]}"',
                "tags:",
                "  - liberated-data"
            ]
            for t in record.get("tags", []):
                frontmatter.append(f"  - {t}")
            frontmatter.append("---\n")
            
            body_lines = [f"# {clean_title}\n"]
            for idx, block in enumerate(record.get("blocks", [])):
                if block["type"] == "heading_2":
                    body_lines.append(f"\n## {block['text']}\n")
                elif block["type"] == "paragraph":
                    blk_id = generate_block_id(block["text"], idx)
                    body_lines.append(f"{block['text']} ^{blk_id}\n")
            
            body_lines.append("\n---\n")
            body_lines.append(f"* **Source Platform**: {record['source_platform']}")
            open_b, close_b = "[" + "[", "]" + "]"
            refs = [f"{open_b}{r['target_title']}{close_b}" for r in record.get("relations", [])]
            if refs:
                body_lines.append(f"* **Garden Source Reference**: {', '.join(refs)}")
            
            full_content = "\n".join(frontmatter) + "\n".join(body_lines) + "\n"
            target_path.write_text(full_content, encoding="utf-8")
 
if __name__ == "__main__":
    ingest_stream(Path("liberated_feed.jsonl"), Path("02 Review/freemydata"))
# Execute stream ingestion and verify output node count
python ingest_pipeline.py --input liberated_feed.jsonl --target "02 Review/freemydata"
test $(ls -1 "02 Review/freemydata"/*.md | wc -l) -gt 0

  • Directus Target: freemydata
  • Garden Source Reference: [BSN-1001 - Incremental Parsing with Tree-sitter](BSN-1001 - Incremental Parsing with Tree-sitter), DAT-1003 - Deterministic Lossless Conversion, MOC - Data Liberation Workbenches, MOC - The Plain-Text Longevity Standard, MOC - Bosun PKM Tools, [BSN-1002 - Deterministic Round-Trip Serialization](BSN-1002 - Deterministic Round-Trip Serialization)