Living Document Notice
Published 2026-09-26. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Data Sovereignty in Production
Summary
Migrating personal archives away from proprietary cloud services frequently yields malformed export archives. Data dumps provided by SaaS platforms often contain nested JSON schemas, non-standard HTML formatting, and disconnected asset references that cannot be rendered directly in local Markdown editors.
The FreeMyData extraction engine operates as an automated extraction pipeline that transforms raw cloud export payloads into normalized CommonMark documents and structured SQLite relational records. The pipeline executes entirely within local memory spaces, ensuring that unvetted personal archives never leave the host workstation during transformation.
The Extraction and Normalization Pipeline
Proprietary note formats encode rich data structures using idiosyncratic JSON trees or HTML wrappers. Attempting to convert these payloads with naive regex scripts introduces formatting anomalies, unescaped tags, and broken cross-document links.
The FreeMyData extraction pipeline structures the conversion process across four distinct transformation stages.
+-------------------------------------------------------------+
| Raw SaaS Export Archive |
| (Nested JSON / Dirty HTML / Attachments) |
+-------------------------------------------------------------+
|
v
+-------------------------------------------------------------+
| Phase 1: Lexical Ingestion |
| (Encoding Sanitization & Tokenization) |
+-------------------------------------------------------------+
|
v
+-------------------------------------------------------------+
| Phase 2: Canonical IR Mapping |
| (Intermediate Representation Graph Model) |
+-------------------------------------------------------------+
|
+--------------+--------------+
| |
v v
+-----------------------------+ +-----------------------------+
| Phase 3: Markdown Emitter | | Phase 4: SQLite Loader |
| (CommonMark + YAML Front) | | (Relational Tables & FTS) |
+-----------------------------+ +-----------------------------+
The pipeline ingests raw files, validates character encodings against UTF-8 specifications, and constructs an in-memory Intermediate Representation (IR). Once the IR tree validates against schema rules, independent emitters write human-readable Markdown files and normalized SQLite index records simultaneously.
Comparison of Export Payload Structures
Different cloud platforms package exported user data in wildly divergent formats. The table below analyzes common SaaS export characteristics and the normalization strategies used by FreeMyData.
| Source Platform Format | Structural Complexity | Image Reference Preservation | Formatting Challenges | Normalization Strategy |
|---|---|---|---|---|
| Nested JSON Trees | High (Deeply nested property trees) | Base64 strings embedded in JSON | Arbitrary custom block schemas | Recursive AST traversal to CommonMark |
| HTML Export Dumps | Medium (Vendor-specific <div> tags) | Relative paths or signed URLs | Style attributes and layout wrappers | DOM parsing with tag whitelist filtering |
| Monolithic XML Files | High (Complex schema definitions) | Hex-encoded binary blocks | Entity escaping and mixed content | Streaming SAX parser extracting text |
| Comma-Separated Values | Low (Flat columnar tables) | Disconnected URL strings | Multiline cell values with delimiters | Direct SQLite relational table import |
By standardizing all incoming data into an intermediate representation prior to disk serialization, the system isolates parsing errors to individual source records without failing the entire batch import.
Streaming Normalization Parser
The core extraction loop parses unstructured input streams incrementally to keep memory usage bounded when processing multi-gigabyte archives.
import json
import sqlite3
from pathlib import Path
def extract_and_normalize_archive(source_json: Path, out_dir: Path, db_path: Path):
out_dir.mkdir(parents=True, exist_ok=True)
conn = sqlite3.connect(db_path)
cur = conn.cursor()
cur.execute("""
CREATE TABLE IF NOT EXISTS extracted_records (
record_id TEXT PRIMARY KEY,
title TEXT,
source_created TEXT,
word_count INTEGER
)
""")
with open(source_json, "r", encoding="utf-8") as f:
data = json.load(f)
for item in data.get("notes", []):
doc_id = item["id"]
title = item.get("title", "Untitled Note").replace("/", "-")
raw_body = item.get("content", "")
clean_markdown = convert_ast_to_commonmark(raw_body)
# Write canonical Markdown file
md_file = out_dir / f"{title}.md"
with open(md_file, "w", encoding="utf-8") as out:
out.write(f"---\nid: \"{doc_id}\"\nsource: \"saas_export\"\n---\n\n")
out.write(f"# {title}\n\n{clean_markdown}\n")
cur.execute(
"INSERT OR REPLACE INTO extracted_records VALUES (?, ?, ?, ?)",
(doc_id, title, item.get("created"), len(clean_markdown.split()))
)
conn.commit()
conn.close()
def convert_ast_to_commonmark(raw_text: str) -> str:
# Deterministic sanitation without proprietary formatting tags
return raw_text.replace("\r\n", "\n").strip()The script converts raw objects into portable filesystem artifacts while updating the SQLite index to maintain provenance metadata.
Invariant of Zero-Loss Transformation
The FreeMyData pipeline enforces a strict information conservation invariant: the sum of word counts and extracted binary asset hashes from the raw archive must match the normalized output directory exactly.
Validate transformation fidelity across an extracted archive:
freemydata-verify --source-dump export.json --target-vault ./02\ Review/bosun-pkm --report-diff- Directus Target: bosunpkm-blog
- Garden Source Reference: MOC - Bosun PKM Engine, MOC - Bosun PKM Tools, MOC - Local-First Systems and Synchronization