Living Document Notice
Published 2026-09-26. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Data Sovereignty in Production

Data Sovereignty in Production: Abstract monochrome amber phosphor CRT progressive 4-stage data liberation pipeline transforming chaotic shards into structured columnar slabs

Summary

Migrating personal archives away from proprietary cloud services frequently yields malformed export archives. Data dumps provided by SaaS platforms often contain nested JSON schemas, non-standard HTML formatting, and disconnected asset references that cannot be rendered directly in local Markdown editors.

The FreeMyData extraction engine operates as an automated extraction pipeline that transforms raw cloud export payloads into normalized CommonMark documents and structured SQLite relational records. The pipeline executes entirely within local memory spaces, ensuring that unvetted personal archives never leave the host workstation during transformation.

The Extraction and Normalization Pipeline

Proprietary note formats encode rich data structures using idiosyncratic JSON trees or HTML wrappers. Attempting to convert these payloads with naive regex scripts introduces formatting anomalies, unescaped tags, and broken cross-document links.

The FreeMyData extraction pipeline structures the conversion process across four distinct transformation stages.

+-------------------------------------------------------------+
|                     Raw SaaS Export Archive                 |
|              (Nested JSON / Dirty HTML / Attachments)       |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
|                   Phase 1: Lexical Ingestion                |
|             (Encoding Sanitization & Tokenization)          |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
|                 Phase 2: Canonical IR Mapping               |
|            (Intermediate Representation Graph Model)        |
+-------------------------------------------------------------+
                               |
                +--------------+--------------+
                |                             |
                v                             v
+-----------------------------+ +-----------------------------+
|   Phase 3: Markdown Emitter | |    Phase 4: SQLite Loader   |
|   (CommonMark + YAML Front) | |  (Relational Tables & FTS)  |
+-----------------------------+ +-----------------------------+

The pipeline ingests raw files, validates character encodings against UTF-8 specifications, and constructs an in-memory Intermediate Representation (IR). Once the IR tree validates against schema rules, independent emitters write human-readable Markdown files and normalized SQLite index records simultaneously.

Comparison of Export Payload Structures

Different cloud platforms package exported user data in wildly divergent formats. The table below analyzes common SaaS export characteristics and the normalization strategies used by FreeMyData.

Source Platform FormatStructural ComplexityImage Reference PreservationFormatting ChallengesNormalization Strategy
Nested JSON TreesHigh (Deeply nested property trees)Base64 strings embedded in JSONArbitrary custom block schemasRecursive AST traversal to CommonMark
HTML Export DumpsMedium (Vendor-specific <div> tags)Relative paths or signed URLsStyle attributes and layout wrappersDOM parsing with tag whitelist filtering
Monolithic XML FilesHigh (Complex schema definitions)Hex-encoded binary blocksEntity escaping and mixed contentStreaming SAX parser extracting text
Comma-Separated ValuesLow (Flat columnar tables)Disconnected URL stringsMultiline cell values with delimitersDirect SQLite relational table import

By standardizing all incoming data into an intermediate representation prior to disk serialization, the system isolates parsing errors to individual source records without failing the entire batch import.

Streaming Normalization Parser

The core extraction loop parses unstructured input streams incrementally to keep memory usage bounded when processing multi-gigabyte archives.

import json
import sqlite3
from pathlib import Path
 
def extract_and_normalize_archive(source_json: Path, out_dir: Path, db_path: Path):
    out_dir.mkdir(parents=True, exist_ok=True)
    conn = sqlite3.connect(db_path)
    cur = conn.cursor()
    cur.execute("""
        CREATE TABLE IF NOT EXISTS extracted_records (
            record_id TEXT PRIMARY KEY,
            title TEXT,
            source_created TEXT,
            word_count INTEGER
        )
    """)
    
    with open(source_json, "r", encoding="utf-8") as f:
        data = json.load(f)
        
    for item in data.get("notes", []):
        doc_id = item["id"]
        title = item.get("title", "Untitled Note").replace("/", "-")
        raw_body = item.get("content", "")
        clean_markdown = convert_ast_to_commonmark(raw_body)
        
        # Write canonical Markdown file
        md_file = out_dir / f"{title}.md"
        with open(md_file, "w", encoding="utf-8") as out:
            out.write(f"---\nid: \"{doc_id}\"\nsource: \"saas_export\"\n---\n\n")
            out.write(f"# {title}\n\n{clean_markdown}\n")
            
        cur.execute(
            "INSERT OR REPLACE INTO extracted_records VALUES (?, ?, ?, ?)",
            (doc_id, title, item.get("created"), len(clean_markdown.split()))
        )
        
    conn.commit()
    conn.close()
 
def convert_ast_to_commonmark(raw_text: str) -> str:
    # Deterministic sanitation without proprietary formatting tags
    return raw_text.replace("\r\n", "\n").strip()

The script converts raw objects into portable filesystem artifacts while updating the SQLite index to maintain provenance metadata.

Invariant of Zero-Loss Transformation

The FreeMyData pipeline enforces a strict information conservation invariant: the sum of word counts and extracted binary asset hashes from the raw archive must match the normalized output directory exactly.

Validate transformation fidelity across an extracted archive:

freemydata-verify --source-dump export.json --target-vault ./02\ Review/bosun-pkm --report-diff

  • Directus Target: bosunpkm-blog
  • Garden Source Reference: MOC - Bosun PKM Engine, MOC - Bosun PKM Tools, MOC - Local-First Systems and Synchronization