Living Document Notice
Published 2026-09-14. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Handling Broken and Ill-Formed Schemas
Summary
While Schema.org establishes formal definitions for culinary metadata, production implementations across WordPress plugins and custom web publisher CMS backends diverge significantly from the standard. Scrapers frequently encounter malformed JSON-LD graphs, string-wrapped instruction blocks, mixed array types, and unescaped control characters.
FreeMyRecipes applies an AST normalization pipeline that accepts non-conforming JSON structures and transforms them into standardized internal representations. This dispatch outlines the structural variations found in the wild and the normalizer transformations that resolve them.
Common Schema Degeneracies in Production CMS Engines
Publishing software packages generate Schema.org outputs through disparate templating filters. These generators often wrap fields in nested containers or flatten complex types into primitive strings.
Incoming Unsanitized JSON-LD
│
▼
┌───────────────────────────────────────────────┐
│ Schema Normalization Pipeline │
│ │
│ 1. Graph Unwrapping: Flatten @graph arrays │
│ 2. Instruction AST: Normalize HowToStep nodes │
│ 3. Ingredient Sanitization: Strip HTML markup │
│ 4. Type Coercion: Fix stringified durations │
└───────────────────────────────────────────────┘
│
▼
Standardized Recipe Entity
The most frequent defect occurs in instruction arrays. Standard schema specifies an array of HowToStep or HowToSection objects, but many sites emit an array of raw strings, HTML paragraph blocks, or single newline-delimited text blobs.
Structural Variances Across Publishing Engines
The table below catalogs recurrent schema defects identified across popular commercial food publishing stacks and their normalized targets.
| CMS / Plugin Engine | Observed Schema Defect | Target Field | Normalizer Resolution Strategy |
|---|---|---|---|
| WP Recipe Maker (Legacy) | Nested @graph containing mixed WebPage and Recipe items | Root payload | Traverse @graph list; extract node matching @type: Recipe |
| Tasty Recipes | Raw HTML tags inside recipeIngredient strings | recipeIngredient | Strip tag elements; decode HTML entities (½ → ½) |
| Mediavine Create | Instructions emitted as single string delimited by \n\n | recipeInstructions | Split on regex newline patterns; wrap lines into HowToStep records |
| Custom Ghost Themes | Duration fields formatted as raw integers ("prepTime": "25") | prepTime | Coerce plain integer to ISO 8601 duration format (PT25M) |
| SquareSpace Custom Blocks | Stringified JSON embedded inside script tag attributes | Root payload | Run secondary JSON decoder over raw string attribute contents |
Recursive AST Normalizer
The FreeMyRecipes normalizer ingests raw JSON dictionaries and applies deterministic transformations across all keys.
import re
from typing import Any, Dict, List, Union
def unwrap_graph(payload: Union[Dict[str, Any], List[Any]]) -> Dict[str, Any]:
if isinstance(payload, list):
for item in payload:
if isinstance(item, dict) and item.get("@type") == "Recipe":
return item
if isinstance(payload, dict):
if payload.get("@type") == "Recipe":
return payload
if "@graph" in payload and isinstance(payload["@graph"], list):
for item in payload["@graph"]:
if isinstance(item, dict) and item.get("@type") == "Recipe":
return item
raise ValueError("No Recipe entity discovered in payload")
def normalize_instructions(raw_steps: Any) -> List[Dict[str, str]]:
normalized: List[Dict[str, str]] = []
# Handle single delimited string
if isinstance(raw_steps, str):
lines = [line.strip() for line in re.split(r'[\r\n]+', raw_steps) if line.strip()]
return [{"@type": "HowToStep", "text": line} for line in lines]
# Handle mixed lists of strings, sections, and step objects
if isinstance(raw_steps, list):
for entry in raw_steps:
if isinstance(entry, str):
normalized.append({"@type": "HowToStep", "text": entry.strip()})
elif isinstance(entry, dict):
if entry.get("@type") == "HowToSection" and "itemListElement" in entry:
sub_steps = normalize_instructions(entry["itemListElement"])
normalized.extend(sub_steps)
elif "text" in entry:
# Strip any leftover HTML tags inside text property
clean_text = re.sub(r'<[^>]+>', '', entry["text"]).strip()
normalized.append({"@type": "HowToStep", "text": clean_text})
return normalizedNormalizing Non-Standard ISO Durations
A recurring parsing error involves ISO 8601 time strings. Standard schemas define PT1H30M, but publishers occasionally emit 1 hr 30 mins, PT90M, or plain integer minute values.
def normalize_duration(duration_val: Union[str, int, None]) -> str:
if duration_val is None:
return "PT0M"
if isinstance(duration_val, int):
return f"PT{duration_val}M"
val_str = str(duration_val).strip()
if val_str.startswith("P"):
return val_str
match = re.search(r'(?:(\d+)\s*(?:hours|hour|hrs|hr|h))?\s*(?:(\d+)\s*(?:minutes|minute|mins|min|m))?', val_str)
if match:
hrs, mins = match.groups()
hrs_part = f"{hrs}H" if hrs else ""
mins_part = f"{mins}M" if mins else ""
return f"P{hrs_part}{mins_part}" if (hrs_part or mins_part) else "PT0M"
return "PT0M"The pipeline executes these sanitizers before writing normalized JSON records to disk, avoiding downstream parsing exceptions during document generation.
- Directus Target: freemyrecipes
- Garden Source Reference: freemyrecipes-index, schema-edge-cases, MOC - Data Liberation Workbenches, MOC - Culinary & Domain Workspaces