WORK-557
ID:WORK-557Status:done

Measure the validation blast radius across the dogfooded sites

SPEC-132 phase 2's gate. Before the attribute error ids are switched on, count what they would report across site/ and plan-site/ — and record the number, so the decision to proceed is made on evidence rather than optimism.

Priority:highComplexity:simpleMilestone:v0.34.0Source:SPEC-132

Criteria completion

Criteria completion: 6 of 6 (100%) checked; history from Sep 14 to Sep 150%25%50%75%100%Sep 14Sep 15
Branches 4
History 3
  1. e7405c6
    • ☑ A count per error id across `site/` and `plan-site/`, at minimum: `attribute-value-invalid`, `attribute-missing-required`, `attribute-type-invalid`
    • ☑ Findings grouped by rune, so a single misbehaving rune is distinguishable from a broad problem
    • ☑ Findings from the custom validators (`SeparatedString`, `SpaceSeparatedNumberList`, media's) are counted separately — they have never run, so they are the least predictable
    • ☑ `critical`-level findings counted separately from `error`, since D11 makes them non-suppressible
    • ☑ The numbers are written into this item's Resolution, not just reported in a PR comment — {% ref "WORK-558" /%} reads them
    • ☑ A recommendation: proceed inside {% ref "WORK-558" /%}, or split phase 2 out of the milestone
    by bjornolofandersson
  2. c22995b
    Created (ready)by bjornolofandersson
  3. d0e9b80
    Content editedby Claude
    docs(plan): SPEC-132 drops phase 3; break phases 1-2 into work items

Why it is its own item

Phase 2 enables attribute-value-invalid, attribute-missing-required and attribute-type-invalid, which also wake the custom attribute validators that have never executed in a build. Nobody knows what that reports. Two outcomes need different plans:

  • A handful of findings — fix them inside WORK-558 and proceed.
  • Many findings, or findings concentrated in one rune — phase 2 becomes its own piece of work and slips out of v0.34.0, which is a better outcome than discovering it halfway through.

Measuring costs an afternoon. Guessing costs the milestone's credibility.

Acceptance Criteria

  • A count per error id across site/ and plan-site/, at minimum: attribute-value-invalid, attribute-missing-required, attribute-type-invalid
  • Findings grouped by rune, so a single misbehaving rune is distinguishable from a broad problem
  • Findings from the custom validators (SeparatedString, SpaceSeparatedNumberList, media's) are counted separately — they have never run, so they are the least predictable
  • critical-level findings counted separately from error, since D11 makes them non-suppressible
  • The numbers are written into this item's Resolution, not just reported in a PR comment — WORK-558 reads them
  • A recommendation: proceed inside WORK-558, or split phase 2 out of the milestone

Approach

Measurement only — do not fix anything here, and do not enable the ids in the shipped path. A throwaway script or a temporarily widened allow-list on WORK-556's filter is enough; the deliverable is the numbers.

Run it against both dogfooded sites rather than just site/. plan-site/ uses a different rune mix, and a finding count from one says little about the other.

References

  • SPEC-132 — the phase table and D11
  • WORK-556 — provides the call and the id filter this widens
  • WORK-558 — consumes this measurement

Resolution

Completed: 2026-09-15

Branch: claude/milestone-v0-34-0-5zoab0

Recommendation

Proceed inside WORK-558. Phase 2 does not need to split out of the milestone. Eight genuine findings across both dogfooded sites, each a one-line fix, plus one upstream artifact that needs a decision rather than work.

How it was measured

scripts/validation-blast-radius.mjs — added by this item. It loads every site in refrakt.config.json through loadContent exactly as the adapter does (including each plugin's configure hook, without which the plan site reports 6 pages instead of 786), forces the full measured id set on regardless of the shipped allow-list, and groups the findings.

Measurement only — it fixes nothing and does not change the shipped default. Re-runnable: node scripts/validation-blast-radius.mjs.

variable-undefined is deliberately excluded from the measurement as well as from the product. SPEC-132 D4 already measured it as unusable, and including it would have buried every other number.

The numbers

sitepagesfindings
site/ (main)2353,310
plan-site/ (plan)7860

Per error id, site/ — plan-site/ reported nothing under any id:

idcount
attribute-type-invalid3,309
attribute-value-invalid1
attribute-missing-required0
tag-undefined0
attribute-undefined0

The headline number is one message repeated 3,302 times, which is why the raw total is misleading:

countmessage
3,302attribute-type-invalid: Attribute 'primary' must be type of 'Object'
2attribute-type-invalid: Attribute 'limit' must be type of 'String'
1attribute-value-invalid: Attribute 'frame-displace' … Got 'both' instead.
1attribute-type-invalid: Attribute 'showContrast' must be type of 'Boolean'
1attribute-type-invalid: Attribute 'showA11y' must be type of 'Boolean'
1attribute-type-invalid: Attribute 'showCharset' must be type of 'Boolean'
1attribute-type-invalid: Attribute 'route' must be type of 'Boolean'
1attribute-type-invalid: Attribute 'limit' must be type of 'Number'

By rune

Not a useful axis here, and the reason is worth recording. The item asked for it to distinguish "a single misbehaving rune" from "a broad problem", and the answer is the first — but the grouping that shows it is by message, not by rune. Markdoc's attribute messages name the attribute, not the owning tag, so a rune histogram comes out empty for this id class. The 3,302 all originate from one construct ({% if %}), and the remaining 8 are spread one apiece across collection, showcase, palette, typography, map and plan-activity.

Custom validators: zero

SeparatedString, SpaceSeparatedNumberList and the two in plugins/media/src/attributes.ts produce 0 findings across both sites.

They are told apart from Markdoc's own type check by message — all four share the attribute-type-invalid id, but only they emit is not a string or contains non-numeric value. Neither phrase appears anywhere in the output.

So the least predictable part of phase 2 turns out to be the quietest. That is a real result and not a null one: it means WORK-558's criterion that SpaceSeparatedNumberList demonstrably rejects non-numeric input in a build has to be proved by a test with deliberately bad input, because our own content will never exercise it.

Critical vs error

Zero critical findings on either site under the measured ids, after WORK-556's fixes. Before them there were 6,613 — 6,609 tag-placement-invalid from the bindRow inline defect, and 4 from a malformed {% ref %} — all resolved in that item.

One correction to how this criterion was framed. It asked for critical counted separately from error "since D11 makes them non-suppressible", on the spec's assumption that the phase 1 and 2 ids are all error. They are not: tag-undefined is critical in Markdoc 0.4.0, and attribute-type-invalid / attribute-value-invalid take their level from the attribute schema's own errorLevel — so hint.type, which sets errorLevel: 'critical', would have reported its finding as critical too. WORK-556's resolution has the full table.

The 3,302: an upstream artifact, not a defect in our content

{% if %} is Markdoc's own tag, and its schema declares primary: { type: Object, render: false } (@markdoc/markdoc@0.4.0, src/tags/conditional.ts:37). That type holds for what an author writes — {% if $var %} and {% if equals(a, b) %} both pass an expression node, which is an object.

It stops holding after substitution. bindRow replaces $row.required with the row's actual value before the page is validated, so the schema sees {% if true %} and rejects a boolean against Object. Verified in isolation:

{% if $row.required %}  → clean
{% if true %}           → attribute-type-invalid