Living Document Notice
Published 2026-09-13. The evolving architecture, revisions, and connected notes for this dispatch live in the Stax Digital Garden.
Harvesting the Hidden JSON-LD
Summary
Commercial recipe sites rely heavily on search engine indexing to acquire organic visitors. To satisfy search crawler requirements, publishers embed machine-readable Schema.org/Recipe data structures directly into their HTML markup using JSON-LD script elements. While human visitors receive heavy client-side tracking frameworks, search indexing bots read clean structured JSON.
FreeMyRecipes exploits this dual-representation pattern. Rather than instantiating a headless browser runtime, our pipeline monitors incoming HTTP packet chunks, identifies JSON-LD script delimiters, and extracts structured recipe objects directly from raw network buffers.
Stream Parsing vs Headless Browser Execution
Automated scrapers often rely on headless Chromium instances managed by Playwright or Puppeteer. While full browser execution handles client-rendered SPAs, it imposes severe performance and operational penalties when extracting static recipes.
A headless browser downloads all referenced subresources, instantiates V8 contexts, and processes background tracking scripts. Stream parsing processes the incoming TCP stream directly, dropping irrelevant HTML nodes before memory allocation.
Incoming TCP Byte Stream
│
▼
┌──────────────────┐ Non-Script Tokens
│ Tokenizer Window │ ─────────────────────────► [ Discarded Buffer ]
└──────────────────┘
│
│ Match: <script type="application/ld+json">
▼
┌──────────────────┐
│ JSON-LD Buffer │
└──────────────────┘
│
│ Match: </script>
▼
┌──────────────────┐
│ SIMD JSON Parser │ ─────────────────────────► [ Structured Recipe Object ]
└──────────────────┘
The streaming parser terminates the HTTP connection the moment the target recipe schema closes, eliminating downstream asset downloads.
Benchmark: Ingestion Pipeline Overhead
The table below illustrates performance metrics across four extraction methods executed against a sample corpus of two hundred commercial food blogs.
| Pipeline Mechanism | Average Memory Footprint | Median Time to Recipe Object | Network Egress per Page | Process Isolation Required |
|---|---|---|---|---|
| Headless Chromium (Playwright) | 184 MB | 2,840 ms | 9.4 MB | Yes (Sandbox container) |
| Headless WebKit | 142 MB | 2,120 ms | 8.1 MB | Yes (Sandbox container) |
| DOM Parser (JSDOM / Cheerio) | 38 MB | 310 ms | 1.8 MB | No |
| FreeMyRecipes Stream Lexer | 4.2 MB | 18 ms | 48 KB | No |
Stream-level extraction cuts memory usage by over ninety-five percent compared to DOM emulation, while reducing latency to less than twenty milliseconds.
Token Scanner Implementation
The core scanner uses a sliding-window token matcher. It scans raw byte slices for the opening delimiter <script type="application/ld+json"> without allocating intermediate strings.
import io
import json
from typing import Optional, Iterator
OPEN_TAG = b'<script type="application/ld+json">'
CLOSE_TAG = b'</script>'
def stream_extract_jsonld(stream: io.RawIOBase) -> Iterator[dict]:
buffer = bytearray()
inside_script = False
while chunk := stream.read(4096):
buffer.extend(chunk)
while True:
if not inside_script:
idx = buffer.find(OPEN_TAG)
if idx == -1:
# Retain potential split prefix across chunk boundaries
buffer = buffer[-(len(OPEN_TAG) - 1):]
break
buffer = buffer[idx + len(OPEN_TAG):]
inside_script = True
if inside_script:
end_idx = buffer.find(CLOSE_TAG)
if end_idx == -1:
break
payload = buffer[:end_idx].strip()
buffer = buffer[end_idx + len(CLOSE_TAG):]
inside_script = False
try:
parsed = json.loads(payload.decode('utf-8'))
yield parsed
except (json.JSONDecodeError, UnicodeDecodeError):
continueWhen an upstream server returns chunked Transfer-Encoding, this generator yields valid JSON objects as packets arrive over the network socket.
Handling Early Stream Termination
To minimize egress consumption, FreeMyRecipes monitors parsed objects for target types. Once a valid Recipe entity resolves, the HTTP client aborts the socket connection immediately.
# Verify stream abort on recipe match via CLI
freemyrecipes harvest --abort-on-match "https://example.com/artisan-bread" \
--target-type "Recipe" \
--debug-socket-lifecycle[SOCKET] CONNECT 198.51.100.24:443 (TLSv1.3)
[STREAM] Read chunk 0 (4096 bytes) - Scanning tokens
[STREAM] Read chunk 1 (4096 bytes) - Found <script type="application/ld+json">
[PARSER] Matched @type: "Recipe" (Identifier: "Artisan Sourdough Boule")
[SOCKET] TCP RST injected; stream aborted at byte offset 8192
[OUTPUT] Schema extracted in 14.8ms. Saved to cache/artisan-bread.json
- Directus Target: freemyrecipes
- Garden Source Reference: freemyrecipes-index, schema-org-standards, REC-1001 - The Hostile Recipe Web, MOC - Data Liberation Workbenches, MOC - Culinary & Domain Workspaces