Living Document Notice
Published 2026-09-13. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Field Trust and Confidence Scoring in Recipe Ingestion

Field Trust and Confidence Scoring in Recipe Ingestion: High-contrast P4 paper white and amber dual-trace vector CRT macro showing statistical confidence scoring histogram and Gaussian threshold curve

Summary

How Galley assesses ingredient quantities, prep notes, and missing cooking units. Instead of silent truncation, ambiguous fields trigger explicit amber review flags.

Automated scrapers frequently drop prep notes like divided or room temperature, or coerce ambiguous quantities into incorrect decimal representations. In a kitchen, cooking requires precision in some areas and flexibility in others. When software silently guesses what an author meant, the cook pays the price at the stove.

Galley implements a field trust engine that evaluates extracted tokens and assigns confidence scores to individual components of a recipe. Rather than silently coercing ambiguous text, the engine raises explicit visual review flags.

The Multi-Token Ingredient AST

Parsing an ingredient line requires segmenting a single string into semantic tokens:

"2 1/2 tablespoons unsalted butter, softened and divided"
 ──┬── ────┬────── ───────┬───────  ────────────┬─────────
   │       │              │                     │
 Quantity Unit        Food Item             Prep Notes

Each token is evaluated against known vocabularies and grammatical rules:

  • Quantity: Recognizes integers, unicode fractions, standard fractions, and quantity ranges.
  • Unit: Matches metric units, US customary measures, and count indicators such as bunch or sprig.
  • Food Item: Strips leading articles and matches core culinary nouns.
  • Prep Modifiers: Detects state descriptors and procedural instructions.

When an ingredient line lacks an explicit unit (such as “3 large eggs”), the parser recognizes “eggs” as a discrete count noun, assigning a high confidence score without generating false unit warnings.

Confidence Scoring Thresholds

Each parsed token receives a normalized trust score between 0.0 and 1.0:

Confidence RangeUI Status IndicatorSystem BehaviorTypical Triggers
0.90 – 1.00Green (High Confidence)Field populated cleanly; cursor skips during review.Exact unit match, unambiguous single integer or fraction.
0.50 – 0.89Amber (Review Flagged)Field highlighted; requires single keystroke confirmation.Range quantities, compound units, prep cues.
0.00 – 0.49Red (Ambiguous / Unknown)Field focused first; displays raw string for quick edit.Unparsed measurement strings, foreign language units.

Fields rated with amber or red confidence scores are placed directly into the operator’s keyboard navigation path, ensuring that potential parsing inaccuracies are corrected before storage.

A Concrete Python Heuristic Example

Here is a simplified view of how Galley evaluates ingredient confidence in Python:

from dataclasses import dataclass
from typing import Optional
 
@dataclass
class IngredientToken:
    raw: str
    quantity: Optional[float]
    unit: Optional[str]
    item: str
    prep: Optional[str]
    confidence: float
 
def score_ingredient_line(raw_line: str) -> IngredientToken:
    confidence = 1.0
    known_units = ["cup", "tbsp", "tsp", "g", "ml", "oz", "lb"]
    has_known_unit = any(u in raw_line.lower() for u in known_units)
    has_prep_note = any(p in raw_line.lower() for p in ["divided", "chopped", "diced", "melted"])
    
    if not has_known_unit:
        confidence -= 0.35
    if has_prep_note:
        confidence -= 0.15
        
    return IngredientToken(
        raw=raw_line,
        quantity=1.0,
        unit="cup" if has_known_unit else None,
        item=raw_line.strip(),
        prep="checked" if has_prep_note else None,
        confidence=max(0.1, round(confidence, 2))
    )

Handling Compound Quantities and Packaging References

A frequent stumbling block in automated ingestion is packaging-based quantities, such as “two 14-ounce cans crushed tomatoes.” A naive parser extracts either “2” or “14,” losing the relationship between package count and unit weight.

Galley’s tokenizer identifies packaging nouns (“can”, “tin”, “package”, “bottle”) and preserves both the item count and the package volume as distinct attributes. This structured representation allows cooks to understand the original packaging requirements while still scaling recipes accurately.

By prioritizing transparency over false confidence, Galley keeps the user informed and in complete command of their recipe archive.


  • Directus Target: galley
  • Garden Source Reference: Parser Heuristics, Ingredient AST, Confidence Flagging, MOC - Culinary & Domain Workspaces, MOC - The Kitchen Chaos Factor, MOC - Bosun PKM Tools