Living Document Notice
Published 2026-09-11. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Extraction is Assistance Not Authority
Summary
Why digital recipe collection fails when scrapers treat parsed text as canonical truth. We explain why Galley treats automated extraction as an assistive drafting layer and why human verification remains mandatory before data is committed to local storage.
Most automated scrapers optimize for collecting raw web content at high volume. When an ingestion engine visits a webpage, it often operates under the flawed assumption that heuristic parsing equals canonical truth. In the culinary domain, this assumption produces degraded archives. Misparsed fractions turn a tablespoon into an eighth of a cup, detailed prep instructions are dropped, and structural notes such as divided ingredients vanish during string tokenization.
Galley operates on a different engineering foundation: extraction is assistance, not authority. A parser exists to save manual keystrokes, not to make unilateral editorial decisions on behalf of the cook.
The Failure Modes of Autonomous Parsing
Modern recipe markup across the web is notoriously inconsistent. While many publishing platforms emit Schema.org Recipe microdata or embedded JSON-LD blocks, those structures are frequently populated by marketing plugins or non-standard CMS themes. Typical failure modes observed across thousands of culinary domains include:
- Fraction Loss and Coercion: Fractions such as 1 1/2 parsed by naive regular expressions frequently collapse into single integers, altering recipe chemistry and moisture balance.
- Context Stripping: Necessary preparation qualifiers, such as softened butter or chilled heavy cream, get stripped when values are coerced into rigid database columns.
- Ghost Ingredients: Auto-generated lists frequently include optional garnishes or serving suggestions as mandatory structural components without distinction.
When an application silently writes these parsed outputs directly to its primary database, data corruption accumulates invisibly over time. The cook discovers the missing ingredient or incorrect ratio only when the sauce breaks or the dough fails to rise.
The Human-in-the-Loop Review Boundary
To prevent silent archive decay, Galley establishes a boundary between automated extraction and persistent storage. Ingestion stages documents into an isolated holding inbox where an operator inspects and confirms each parsed field before anything is written to SQLite.
[ Web URL / Raw Text ]
│
▼
[ Ingestion Pipeline ]
├── Fetch & Sanitize DOM
├── Extract JSON-LD / Microdata
└── Heuristic Tokenization
│
▼
[ Staged Review Inbox ] <── Operator Verification
│
▼
[ Canonical Local SQLite Vault ]
By ensuring that the software acts purely as an assistive drafting tool, Galley preserves high data fidelity while reducing manual data entry overhead. The system presents its best extraction guess with visual confidence cues, allowing the cook to approve correct data with a single keystroke or amend subtle errors immediately.
Deterministic Parsing vs Speculative Inference
A critical distinction in Galley’s design is refusing to guess when incoming data is ambiguous. If an author writes “salt to taste” or “three glugs of olive oil,” the parser does not invent an arbitrary milliliter equivalent. It retains the literal string, flags the unit as indeterminate, and presents the raw token directly to the reviewer.
def tokenize_ingredient_line(raw_text: str) -> dict:
# Retain literal source string alongside heuristic breakdown
return {
"literal": raw_text.strip(),
"parsed_quantity": extract_quantity(raw_text),
"parsed_unit": extract_unit(raw_text),
"requires_confirmation": has_ambiguity(raw_text),
}This deterministic approach guarantees that machine speculation never overwrites human culinary judgment.
Assistive Architecture vs Autonomous Scraping
The table below contrasts standard web scrapers against Galley’s assistive ingestion pipeline:
| Dimension | Autonomous Scraper Model | Galley Assistive Model |
|---|---|---|
| Write Target | Direct write to primary tables | Isolated holding queue |
| Ambiguity Handling | Silent fallback or arbitrary guessing | Explicit amber review flags |
| Original Payload | Discarded after extraction | Preserved with cryptographic hashes |
| Operator Role | Passive consumer of scraped errors | Active verifier with keyboard ergonomics |
| Data Integrity | Degrades as web markup changes | Guaranteed clean via explicit sign-off |
Treating extraction as assistance ensures your personal recipe collection remains accurate and structurally dependable over years of practical kitchen use. The machine handles tedious text extraction, while the cook retains final editorial authority.
- Directus Target: galley
- Garden Source Reference: Galley Architecture, Provenance Protocols, Ingestion Pipeline, MOC - Culinary & Domain Workspaces, MOC - The Kitchen Chaos Factor, MOC - Bosun PKM Tools