Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

Handling Broken and Ill-Formed Schemas

Handling Broken and Ill-Formed Schemas: Warm golden amber P20 vector CRT macro showing fractured schema segments realigned by bridging arcs into balanced hierarchical tree

Summary

While Schema.org establishes formal definitions for culinary metadata, production implementations across WordPress plugins and custom web publisher CMS backends diverge significantly from the standard. Scrapers frequently encounter malformed JSON-LD graphs, string-wrapped instruction blocks, mixed array types, and unescaped control characters.

FreeMyRecipes applies an AST normalization pipeline that accepts non-conforming JSON structures and transforms them into standardized internal representations. This dispatch outlines the structural variations found in the wild and the normalizer transformations that resolve them.

Common Schema Degeneracies in Production CMS Engines

Publishing software packages generate Schema.org outputs through disparate templating filters. These generators often wrap fields in nested containers or flatten complex types into primitive strings.

Incoming Unsanitized JSON-LD
          │
          ▼
┌───────────────────────────────────────────────┐
│ Schema Normalization Pipeline                 │
│                                               │
│ 1. Graph Unwrapping: Flatten @graph arrays    │
│ 2. Instruction AST: Normalize HowToStep nodes │
│ 3. Ingredient Sanitization: Strip HTML markup │
│ 4. Type Coercion: Fix stringified durations   │
└───────────────────────────────────────────────┘
          │
          ▼
Standardized Recipe Entity

The most frequent defect occurs in instruction arrays. Standard schema specifies an array of HowToStep or HowToSection objects, but many sites emit an array of raw strings, HTML paragraph blocks, or single newline-delimited text blobs.

Structural Variances Across Publishing Engines

The table below catalogs recurrent schema defects identified across popular commercial food publishing stacks and their normalized targets.

CMS / Plugin EngineObserved Schema DefectTarget FieldNormalizer Resolution Strategy
WP Recipe Maker (Legacy)Nested @graph containing mixed WebPage and Recipe itemsRoot payloadTraverse @graph list; extract node matching @type: Recipe
Tasty RecipesRaw HTML tags inside recipeIngredient stringsrecipeIngredientStrip tag elements; decode HTML entities (½ → ½)
Mediavine CreateInstructions emitted as single string delimited by \n\nrecipeInstructionsSplit on regex newline patterns; wrap lines into HowToStep records
Custom Ghost ThemesDuration fields formatted as raw integers ("prepTime": "25")prepTimeCoerce plain integer to ISO 8601 duration format (PT25M)
SquareSpace Custom BlocksStringified JSON embedded inside script tag attributesRoot payloadRun secondary JSON decoder over raw string attribute contents

Recursive AST Normalizer

The FreeMyRecipes normalizer ingests raw JSON dictionaries and applies deterministic transformations across all keys.

import re
from typing import Any, Dict, List, Union
 
def unwrap_graph(payload: Union[Dict[str, Any], List[Any]]) -> Dict[str, Any]:
    if isinstance(payload, list):
        for item in payload:
            if isinstance(item, dict) and item.get("@type") == "Recipe":
                return item
    if isinstance(payload, dict):
        if payload.get("@type") == "Recipe":
            return payload
        if "@graph" in payload and isinstance(payload["@graph"], list):
            for item in payload["@graph"]:
                if isinstance(item, dict) and item.get("@type") == "Recipe":
                    return item
    raise ValueError("No Recipe entity discovered in payload")
 
def normalize_instructions(raw_steps: Any) -> List[Dict[str, str]]:
    normalized: List[Dict[str, str]] = []
    
    # Handle single delimited string
    if isinstance(raw_steps, str):
        lines = [line.strip() for line in re.split(r'[\r\n]+', raw_steps) if line.strip()]
        return [{"@type": "HowToStep", "text": line} for line in lines]
        
    # Handle mixed lists of strings, sections, and step objects
    if isinstance(raw_steps, list):
        for entry in raw_steps:
            if isinstance(entry, str):
                normalized.append({"@type": "HowToStep", "text": entry.strip()})
            elif isinstance(entry, dict):
                if entry.get("@type") == "HowToSection" and "itemListElement" in entry:
                    sub_steps = normalize_instructions(entry["itemListElement"])
                    normalized.extend(sub_steps)
                elif "text" in entry:
                    # Strip any leftover HTML tags inside text property
                    clean_text = re.sub(r'<[^>]+>', '', entry["text"]).strip()
                    normalized.append({"@type": "HowToStep", "text": clean_text})
                    
    return normalized

Normalizing Non-Standard ISO Durations

A recurring parsing error involves ISO 8601 time strings. Standard schemas define PT1H30M, but publishers occasionally emit 1 hr 30 mins, PT90M, or plain integer minute values.

def normalize_duration(duration_val: Union[str, int, None]) -> str:
    if duration_val is None:
        return "PT0M"
    if isinstance(duration_val, int):
        return f"PT{duration_val}M"
        
    val_str = str(duration_val).strip()
    if val_str.startswith("P"):
        return val_str
        
    match = re.search(r'(?:(\d+)\s*(?:hours|hour|hrs|hr|h))?\s*(?:(\d+)\s*(?:minutes|minute|mins|min|m))?', val_str)
    if match:
        hrs, mins = match.groups()
        hrs_part = f"{hrs}H" if hrs else ""
        mins_part = f"{mins}M" if mins else ""
        return f"P{hrs_part}{mins_part}" if (hrs_part or mins_part) else "PT0M"
        
    return "PT0M"

The pipeline executes these sanitizers before writing normalized JSON records to disk, avoiding downstream parsing exceptions during document generation.


  • Directus Target: freemyrecipes
  • Garden Source Reference: freemyrecipes-index, schema-edge-cases, MOC - Data Liberation Workbenches, MOC - Culinary & Domain Workspaces