Why it is not optional bookkeeping
JSON-LD is the one output nobody looks at. A page renders identically whether its structured data is correct or ruined, so a regression here is invisible in review, invisible in the dev server, and invisible to the person whose site is publishing it.
The evidence is in SPEC-130's own history. Every defect it found was found by running extractSeo, never by reading a transform: MusicPlaylist on podcasts (BUG-013), the BreadcrumbList that renders and is never harvested, NonProfit — a type schema.org does not have — sitting inside a validated enum, seven entities asserting a type and describing nothing. And four Lumina rules pointing at attributes that moved years ago (BUG-015) survived every review in between.
This milestone rewrites that channel across 30 runes in 10 packages. The baseline is what turns each step from a leap into a diff.
Scope
The existing coverage is not this. packages/runes/test/seo.test.ts and the per-plugin seo.test.ts files assert a handful of runes — accordion, breadcrumb, playlist — with hand-written expectations, and several runes this milestone changes have no JSON-LD test at all.
What is wanted is per-rune fixtures run through extractSeo and committed as snapshots, covering every rune that emits a typeof today: all of Group A (7), Group B (14) and Group C (9).
Acceptance Criteria
- A committed JSON-LD snapshot covers all 30 emitting runes, generated by running
extractSeo over a per-rune fixture - The snapshot is regenerable by a documented command, so a deliberate change is a reviewable diff rather than a hand edit
- Comparison normalises object key order and preserves array order —
itemListElement, step, track and recipeInstructions are ordered sequences and a reordering is a real defect - Runes with no JSON-LD test today are covered — the gap is named in the item, not discovered during migration
- The rendered RDFa is captured alongside the JSON-LD for each fixture, so WORK-563's two-point invariant has something to compare against
- Nothing in the snapshot is edited to look correct — BUG-013's mistyped podcasts and
testimonial's jobTitle: ", CTO at Acme" comma defect are recorded as they are - The existing hand-written
seo.test.ts expectations are either folded in or left in place with a note saying which is authoritative
Approach
Snapshot the wrong output too. The instinct is to fix the obvious defects while writing fixtures for them. Don't — a baseline that has been tidied cannot prove that a later change was intentional. playlist type="podcast" should record MusicPlaylist today and change in WORK-569; that diff is the evidence the fix landed.
Fixtures should be minimal but real: enough content to populate every declared property, since SPEC-130's D6 turns on single- versus multi-item collections behaving differently (appendToProperty stores the first value as a scalar and only promotes on the second). Include at least one single-item and one multi-item case for every rune that emits a list, or D6's change will land as an unreviewable diff.
Rune-level fixtures are the cheap half. WORK-563 adds the page-level assertion through runPipeline, which is what catches the postProcess class — they are the same assertion at two granularities and the page-level one belongs with the harvest move that makes it meaningful.
References
- SPEC-130 — "Before anything moves: the baseline", and D6
- BUG-013 — recorded as-is, fixed in WORK-569
packages/runes/src/seo.ts:188 — extractSeopackages/runes/test/seo.test.ts — the coverage this replaces
Resolution
Completed: 2026-09-15
Branch: claude/v0.35-parallel-feasibility-eia5le
What was done
contracts/seo-baseline/fixtures/ — 42 fixtures covering all 30 emitting runes (Group A 7, B 14, C 9), in SPEC-102 fixture format.contracts/seo-baseline/baseline.json — the committed artifact. Per fixture: jsonLd (the pre-engine harvest site.ts publishes) and rendered (jsonLd + annotations after the identity transform). Object keys sorted, array order preserved.scripts/generate-seo-baseline.mjs — generator with --check. Reads the plugin list (and runes.prefer) from refrakt.config.json rather than hard-coding it, and mirrors transformContent / assembleSiteContext, both of which are module-private.scripts/generate-seo-baseline.test.mjs — 31 tests: unit coverage for normalize / rdfaOutline / coverageGaps, plus live assertions over the corpus.npm run seo:baseline / seo:baseline:check; documented in CLAUDE.md.biome.jsonc — artifact added to the generated-artifact exclusion list, matching contracts/structures.json.- A note in all seven existing
seo.test.ts files saying the baseline is authoritative for the catalog-wide question and they remain statements of intent.
What the baseline records (as-is, not tidied)
- Group A confirmed exactly:
gallery, datatable, budget, itinerary, map, symbol, blog emit a bare @type and nothing else — and no other rune does. A test pins that set both ways. - BUG-013:
playlist type="podcast" → MusicPlaylist with MusicRecording episodes; track type="episode" → MusicRecording. - D2:
organization type="NonProfit" publishes a type schema.org lacks. - The
jobTitle comma defect reproduced: **Alex Rivera**, CTO at Acme yields jobTitle: ", CTO at Acme". The em-dash form parses cleanly, so the defect is specific to the comma separator (testimonial.comma fixture). - D6 before-state pinned for all seven list emitters: a one-item collection is a scalar, a multi-item one an array.
Findings worth carrying forward
- The two harvest points already agree at rune level — 41/41 fixtures identical with keys normalised (SPEC-130 measured 15 runes, 11 identical, 4 key-order-only). Confirms WORK-563's invariant is a regression net, not a migration; the page-level
breadcrumb auto failure is the real one. - New defect, not previously recorded: the explicit
{% accordion-item %} form builds a Question with no name, and concatenates the heading into acceptedAnswer.text without a separator ("What is refrakt?A content framework."). The heading-based form is correct. Relevant to WORK-570. - Pre-existing engine diagnostic on
itinerary: "itinerary-stop requires parent itinerary — found nested directly in itinerary-day". Reproduces on the unabridged doc example, so it is the rune's own structure, not the fixture. Out of scope here; noted in the fixture. - 12 of the 30 runes had no JSON-LD assertion anywhere before this:
gallery, budget, itinerary, symbol, blog, cast-member, pricing, tier, timeline-entry, track, breadcrumb-item, accordion-item. breadcrumb-item is not an author-writable tag (absent from the tag map); breadcrumb builds those nodes itself, so it is baselined via its parent.
Notes
- Fixtures live beside the artifact rather than in
plugins/*/fixtures/. That directory is a real convention (discoverPluginFixtures, refrakt inspect) but no plugin ships one, so populating it would change inspect output and the examples generator — a deliberate change of its own, not this item's. - Five fixtures initially used the wrong content model and emitted thin entities while the generator passed. Each was caught by diffing the fixture's property set against the same rune's doc-page example; all canonical fixtures now match exactly. That comparison is deliberately not a test (doc pages change for their own reasons) — the committed artifact plus the "exactly Group A is bare" assertion are what guard the corpus.
- No changeset: nothing here ships in a published package.
pr attribute not set — no PR has been opened for this branch yet.
Verification
npm test — 363 files, 4458 tests, all passing. npm run format:check clean. npm run seo:baseline:check reports up to date, and the artifact is byte-stable across runs.