Living Document Notice
Published 2026-09-16. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

The Fraction Parser

The Fraction Parser: Warm golden amber P20 vector CRT macro showing harmonic fractional intervals and caliper division arcs along graticule axis

Summary

Culinary recipes represent one of the most inconsistent domains for quantity and measurement parsing. Ingredient lists combine ASCII fractions, Unicode vulgar fraction code points, informal mixed numbers, range intervals, and colloquial kitchen measurements. Converting these strings into machine-readable numeric floats is required for automatic scaling and nutritional computation.

FreeMyRecipes incorporates a deterministic tokenizer that normalizes arbitrary ingredient quantity strings into structured float values paired with normalized unit types. This dispatch covers the tokenization grammar, Unicode normalization mapping, and parsing edge cases.

The Problem of Culinary Numeric Notation

Standard numerical parsers fail when processing kitchen notation. The string 1 1/2 cups contains an ASCII mixed fraction separated by whitespace, while 1½ cups contains a whole integer followed immediately by Unicode code point U+00BD.

Recipe authors also regularly use approximations such as 1/4 to 1/2 tsp or 2-3 cloves. Treating quantity strings as simple floats discards critical variance boundaries.

Raw Ingredient String
        │
        ▼
┌──────────────────────────────────────────────┐
│ Unicode Normalization (NFKD Decomposition)   │
│ Converts '½' (U+00BD) ──► '1/2'              │
└──────────────────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────────────────┐
│ Regex Tokenizer Pipeline                     │
│ Match: Range ("1-2"), Mixed ("1 1/2"), Unit  │
└──────────────────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────────────────┐
│ Quantity Struct:                             │
│   min_qty: 1.0, max_qty: 2.0, unit: "cup"    │
└──────────────────────────────────────────────┘

Deterministic parsing requires standardizing Unicode representations before running lexical analysis.

Unicode Vulgar Fraction Normalization Matrix

The table below catalogs common Unicode fraction characters, their decimal equivalents, and target normalized float values.

Unicode CharacterCode PointString RepresentationDecimal Float ValuePrimary Kitchen Use
¼U+00BC1/40.25Teaspoons, cups
½U+00BD1/20.50Teaspoons, tablespoons, cups
¾U+00BE3/40.75Cups, pounds
⅓U+21531/30.333Cups
⅔U+21542/30.666Cups
⅛U+215B1/80.125Teaspoons (spices)

Deterministic Tokenizer Implementation

The parsing engine uses regular expressions to capture whole integers, fractional components, and measurement units in a single pass.

import re
import unicodedata
from typing import Optional, Tuple
 
VULGAR_FRACTIONS = {
    '¼': 0.25,
    '½': 0.5,
    '¾': 0.75,
    '⅓': 1.0 / 3.0,
    '⅔': 2.0 / 3.0,
    '⅛': 0.125,
    '⅜': 0.375,
    '⅝': 0.625,
    '⅞': 0.875,
}
 
QUANTITY_REGEX = re.compile(
    r'^(?P<whole>\d+)?\s*'
    r'(?:(?P<num>\d+)/(?P<den>\d+)|(?P<vulgar>[¼½¾⅓⅔⅛⅜⅝⅞]))?'
    r'(?:\s*-\s*(?P<range_max>[\d\./¼½¾⅓⅔⅛⅜⅝⅞]+))?'
    r'\s*(?P<unit>[a-zA-Z]+)?\s*'
    r'(?P<ingredient>.*)$'
)
 
def parse_ingredient_line(raw_text: str) -> dict:
    text = raw_text.strip()
    match = QUANTITY_REGEX.match(text)
    if not match:
        return {"raw": text, "quantity": None, "unit": None, "ingredient": text}
        
    groups = match.groupdict()
    total_qty = 0.0
    
    if groups.get('whole'):
        total_qty += float(groups['whole'])
        
    if groups.get('num') and groups.get('den'):
        total_qty += float(groups['num']) / float(groups['den'])
    elif groups.get('vulgar'):
        total_qty += VULGAR_FRACTIONS.get(groups['vulgar'], 0.0)
        
    return {
        "quantity": total_qty if total_qty > 0 else None,
        "unit": groups.get('unit'),
        "ingredient": groups.get('ingredient', '').strip(),
        "raw": raw_text
    }

Unit Canonicalization and Range Handling

Once numeric floats are established, the engine normalizes unit strings against standard lookup dictionaries. Variations such as tbsp, Tbs, tablespoon, and tablespoons resolve to the canonical identifier tbsp.

# Test the tokenizer against edge-case input strings
freemyrecipes parse-qty "1 1/2 cups unbleached flour"
# Output: {"quantity": 1.5, "unit": "cup", "ingredient": "unbleached flour"}
 
freemyrecipes parse-qty "3/4 tsp kosher salt"
# Output: {"quantity": 0.75, "unit": "tsp", "ingredient": "kosher salt"}
 
freemyrecipes parse-qty "1-2 tbsp olive oil"
# Output: {"quantity_min": 1.0, "quantity_max": 2.0, "unit": "tbsp", "ingredient": "olive oil"}

  • Directus Target: freemyrecipes
  • Garden Source Reference: freemyrecipes-index, ast-normalization, MOC - Data Liberation Workbenches, MOC - Culinary & Domain Workspaces