Living Document Notice
Published 2026-09-10. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.

From Bloated Food Blogs to Portable Markdown Recipe Cards

From Bloated Food Blogs to Portable Markdown Recipe Cards: Warm golden amber P20 vector CRT macro showing noise waveform filtering through sieve into binary ratio tree

Summary

Browsing modern recipe websites has become an exercise in digital endurance. A reader seeking simple cooking instructions must navigate multi-megabyte advertising scripts, video autoplay overlays, and lengthy autobiographical prose designed solely to optimize search engine ranking algorithms. The actual culinary formula often hides behind dozens of DOM layers, buried beneath promotional banners and dynamic tracking pixels.

FreeMyRecipes resolves this frustration by extracting structured culinary data directly from page metadata, converting bloated web documents into lean, durable Markdown recipe cards suitable for local knowledge archives.

Harvesting Structured Culinary Metadata via Schema.org

Fortunately, search engine indexing requirements force most food publishers to embed structured machine-readable metadata within their HTML documents. Publishers inject this data using Schema.org/Recipe microdata or JSON-LD script tags.

Rather than scraping arbitrary paragraph tags, the FreeMyRecipes parser queries the HTML document specifically for embedded application/ld+json script blocks.

import json
from bs4 import BeautifulSoup
 
def extract_recipe_schema(html_doc: str) -> dict:
    soup = BeautifulSoup(html_doc, 'html.parser')
    scripts = soup.find_all('script', type='application/ld+json')
    
    for script in scripts:
        if not script.string:
            continue
        try:
            data = json.loads(script.string)
        except json.JSONDecodeError:
            continue
            
        # Inspect direct object or graph array
        if isinstance(data, dict):
            if data.get('@type') == 'Recipe':
                return data
            if '@graph' in data:
                for item in data['@graph']:
                    if item.get('@type') == 'Recipe':
                        return item
        elif isinstance(data, list):
            for item in data:
                if isinstance(item, dict) and item.get('@type') == 'Recipe':
                    return item
                    
    raise ValueError("No valid Schema.org/Recipe metadata found in document")

Targeting the structured schema bypasses all page clutter, promotional copy, and advertising DOM structures. The extractor retrieves exact ingredient quantities, prep times, cooking steps, and nutritional facts directly from source data.

Normalizing Instructions into Portable Markdown Cards

Once extracted, the raw JSON metadata is transformed into clean CommonMark documents formatted with consistent YAML frontmatter. This format integrates directly into personal digital gardens, offline mobile readers, or terminal viewers.

---
title: "Skillet Pan-Seared Halibut with Herb Butter"
prep_time: "PT15M"
cook_time: "PT12M"
servings: 4
source_url: "https://example.com/halibut-recipe"
tags:
  - recipe
  - seafood
  - dinner
---
 
# Skillet Pan-Seared Halibut with Herb Butter
 
Fresh halibut fillets seared in a cast-iron skillet with clarifying butter and fresh thyme.
 
## Ingredients
- 4 halibut fillets (6 oz each, patted dry)
- 2 tbsp clarified butter
- 1 tsp coarse sea salt
- 0.5 tsp cracked black pepper
- 2 sprigs fresh thyme
- 1 lemon, sliced into wedges
 
## Instructions
1. Season halibut fillets generously on both sides with sea salt and cracked pepper.
2. Heat clarified butter in a heavy cast-iron skillet over medium-high heat until shimmering.
3. Place fillets in the skillet and sear undisturbed for 4 minutes until a golden crust forms.
4. Carefully flip fillets, add thyme sprigs, and baste with hot butter for an additional 4 minutes.
5. Transfer to warm plates and serve immediately with lemon wedges.

This clean representation strips away all tracking pixels and ephemeral styling. The recipe exists as a plain text artifact that remains readable across decades on any operating system.

Data Efficiency Metrics

The efficiency gain achieved by extracting pure recipe data from modern food blogs is substantial. The table below shows real-world measurements comparing original webpage transfer payloads against resulting Markdown cards.

Recipe SourceOriginal HTML PayloadTotal Network RequestsExtracted Markdown SizeByte Reduction Ratio
Commercial Food Blog A4.8 MB142 requests1.8 KB99.96% reduction
Lifestyle Publisher B6.2 MB218 requests2.4 KB99.96% reduction
Independent Cooking Site C2.1 MB68 requests1.4 KB99.93% reduction
Culinary Magazine D5.5 MB189 requests2.9 KB99.95% reduction

Transforming web clutter into structured Markdown cards preserves culinary knowledge in a format that honors user attention and local data autonomy.


  • Directus Target: freemyrecipes
  • Garden Source Reference: Schema.org Structured Recipe Extraction, Offline Culinary Formatting Standards, MOC - Data Liberation Workbenches, MOC - Culinary & Domain Workspaces