Living Document Notice
Published 2026-09-11. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Extraction is Assistance Not Authority

Extraction is Assistance Not Authority: High-contrast P4 paper white and amber dual-trace vector CRT macro showing candidate extraction cards calibrated against precision reticle

Summary

Why digital recipe collection fails when scrapers treat parsed text as canonical truth. We explain why Galley treats automated extraction as an assistive drafting layer and why human verification remains mandatory before data is committed to local storage.

Most automated scrapers optimize for collecting raw web content at high volume. When an ingestion engine visits a webpage, it often operates under the flawed assumption that heuristic parsing equals canonical truth. In the culinary domain, this assumption produces degraded archives. Misparsed fractions turn a tablespoon into an eighth of a cup, detailed prep instructions are dropped, and structural notes such as divided ingredients vanish during string tokenization.

Galley operates on a different engineering foundation: extraction is assistance, not authority. A parser exists to save manual keystrokes, not to make unilateral editorial decisions on behalf of the cook.

The Failure Modes of Autonomous Parsing

Modern recipe markup across the web is notoriously inconsistent. While many publishing platforms emit Schema.org Recipe microdata or embedded JSON-LD blocks, those structures are frequently populated by marketing plugins or non-standard CMS themes. Typical failure modes observed across thousands of culinary domains include:

  1. Fraction Loss and Coercion: Fractions such as 1 1/2 parsed by naive regular expressions frequently collapse into single integers, altering recipe chemistry and moisture balance.
  2. Context Stripping: Necessary preparation qualifiers, such as softened butter or chilled heavy cream, get stripped when values are coerced into rigid database columns.
  3. Ghost Ingredients: Auto-generated lists frequently include optional garnishes or serving suggestions as mandatory structural components without distinction.

When an application silently writes these parsed outputs directly to its primary database, data corruption accumulates invisibly over time. The cook discovers the missing ingredient or incorrect ratio only when the sauce breaks or the dough fails to rise.

The Human-in-the-Loop Review Boundary

To prevent silent archive decay, Galley establishes a boundary between automated extraction and persistent storage. Ingestion stages documents into an isolated holding inbox where an operator inspects and confirms each parsed field before anything is written to SQLite.

[ Web URL / Raw Text ]
         │
         ▼
[ Ingestion Pipeline ]
  ├── Fetch & Sanitize DOM
  ├── Extract JSON-LD / Microdata
  └── Heuristic Tokenization
         │
         ▼
[ Staged Review Inbox ] <── Operator Verification
         │
         ▼
[ Canonical Local SQLite Vault ]

By ensuring that the software acts purely as an assistive drafting tool, Galley preserves high data fidelity while reducing manual data entry overhead. The system presents its best extraction guess with visual confidence cues, allowing the cook to approve correct data with a single keystroke or amend subtle errors immediately.

Deterministic Parsing vs Speculative Inference

A critical distinction in Galley’s design is refusing to guess when incoming data is ambiguous. If an author writes “salt to taste” or “three glugs of olive oil,” the parser does not invent an arbitrary milliliter equivalent. It retains the literal string, flags the unit as indeterminate, and presents the raw token directly to the reviewer.

def tokenize_ingredient_line(raw_text: str) -> dict:
    # Retain literal source string alongside heuristic breakdown
    return {
        "literal": raw_text.strip(),
        "parsed_quantity": extract_quantity(raw_text),
        "parsed_unit": extract_unit(raw_text),
        "requires_confirmation": has_ambiguity(raw_text),
    }

This deterministic approach guarantees that machine speculation never overwrites human culinary judgment.

Assistive Architecture vs Autonomous Scraping

The table below contrasts standard web scrapers against Galley’s assistive ingestion pipeline:

DimensionAutonomous Scraper ModelGalley Assistive Model
Write TargetDirect write to primary tablesIsolated holding queue
Ambiguity HandlingSilent fallback or arbitrary guessingExplicit amber review flags
Original PayloadDiscarded after extractionPreserved with cryptographic hashes
Operator RolePassive consumer of scraped errorsActive verifier with keyboard ergonomics
Data IntegrityDegrades as web markup changesGuaranteed clean via explicit sign-off

Treating extraction as assistance ensures your personal recipe collection remains accurate and structurally dependable over years of practical kitchen use. The machine handles tedious text extraction, while the cook retains final editorial authority.


  • Directus Target: galley
  • Garden Source Reference: Galley Architecture, Provenance Protocols, Ingestion Pipeline, MOC - Culinary & Domain Workspaces, MOC - The Kitchen Chaos Factor, MOC - Bosun PKM Tools