Living Document Notice
Published 2026-09-10. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Escaping Proprietary Vaults – The Client Side Parsing Approach
Summary
Personal knowledge management systems frequently trap personal data behind proprietary binary databases or XML schema variants. When users decide to migrate their notes to local Markdown repositories, existing migration tools often require transmitting complete archives to remote SaaS conversion servers. This architecture presents severe privacy risks, exposing personal journals and research to third-party databases.
The FreeNext project adopts an alternate approach: executing note parsing and asset extraction entirely inside the user’s browser runtime via Web Workers. This model guarantees that user notes never leave local device memory during conversion.
Processing Large Archives in Browser Memory
Desktop note archives often span multiple gigabytes and contain tens of thousands of attached media assets. Loading an entire two-gigabyte Evernote .enex export file into the JavaScript main thread triggers browser tab crashes and UI freezes.
To maintain interface responsiveness and manage constrained browser memory, the FreeNext pipeline distributes work across dedicated background Web Workers using streaming XML chunking.
// worker.ts: Streaming ENEX node extraction
import { XMLParser } from 'fast-xml-parser';
self.onmessage = async (event: MessageEvent<File>) => {
const file = event.data;
const stream = file.stream();
const reader = stream.getReader();
let buffer = '';
let noteIndex = 0;
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += new TextDecoder().decode(value, { stream: true });
// Extract complete <note>...</note> segments
let startIndex = buffer.indexOf('<note>');
let endIndex = buffer.indexOf('</note>');
while (startIndex !== -1 && endIndex !== -1 && endIndex > startIndex) {
const noteXml = buffer.slice(startIndex, endIndex + 7);
buffer = buffer.slice(endIndex + 7);
const parsedNote = parseSingleNote(noteXml);
self.postMessage({ type: 'NOTE_PARSED', note: parsedNote, index: ++noteIndex });
startIndex = buffer.indexOf('<note>');
endIndex = buffer.indexOf('</note>');
}
}
self.postMessage({ type: 'COMPLETE', totalNotes: noteIndex });
};This streaming pattern keeps heap consumption predictable. Instead of parsing the complete archive at once, the worker processes notes sequentially, serializing extracted Markdown and queuing binary attachments into local browser IndexedDB storage.
Normalizing Proprietary XML into CommonMark
Evernote formatting relies on ENML, a customized XHTML subset containing proprietary tags such as en-note, en-media, and en-todo. Transforming this markup into standard CommonMark requires careful DOM sanitation.
function transformEnmlToMarkdown(enmlContent: string): string {
// Replace Evernote media references with clean Markdown asset links
let md = enmlContent.replace(
new RegExp('<en-media[^>]*hash="([a-f0-9]+)"[^>]*type="image/([^"]+)"[^>]*/>', 'gi'),
'!alt text](assets/$1.$2)'
);
// Convert custom checkbox elements to GitHub Flavored Markdown task items
md = md.replace(/<en-todo\s+checked="true"\s*\/>/gi, '- [x] ');
md = md.replace(/<en-todo\s*\/?>/gi, '- [ ] ');
// Strip enclosing note wrappers and excess paragraph tags
md = md.replace(/<\/?en-note[^>]*>/gi, '');
md = md.replace(/<div><br\s*\/?><\/div>/gi, '
');
md = md.replace(/<div>/gi, '
').replace(/<\/div>/gi, '');
return md.trim();
}Handling embedded tables requires converting nested HTML table markup into standard Markdown grid tables. When encountering complex spans or invalid tag nestings common in legacy exports, the parser falls back to clean indented lists to preserve data integrity.
Architecture Comparison: Client Execution versus SaaS Converters
The table below contrasts the architectural differences between browser-local note conversion and traditional hosted migration services.
| Architectural Dimension | FreeNext Client-Side Model | Hosted Migration SaaS |
|---|---|---|
| Data Boundary | 100% within local browser tab | Transmitted over public networks |
| Server Compute Cost | Zero infrastructure spend | Proportional to archive gigabytes |
| Privacy Exposure | No persistence outside local disk | Potential logging and database backups |
| Archive Size Limits | Constrained by device RAM | Constrained by upload timeouts |
| Offline Availability | Functions without network connection | Requires constant high-bandwidth uplink |
Executing data conversion directly within the browser tab demonstrates that personal knowledge migration does not require sacrificing personal privacy.
- Directus Target: freenext
- Garden Source Reference: Client-Side Migration Architectures, Evernote ENEX Schema Normalization, MOC - Data Liberation Workbenches, MOC - The Plain-Text Longevity Standard, MOC - Local-First Systems and Synchronization, MOC - Bosun PKM Tools, [BSN-1001 - Incremental Parsing with Tree-sitter](BSN-1001 - Incremental Parsing with Tree-sitter), [BSN-1002 - Deterministic Round-Trip Serialization](BSN-1002 - Deterministic Round-Trip Serialization)