feat(design-systems): the model, the parent chain, and the guard on it
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 7s
CI & Build / integration (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 20s
CI & Build / Python tests (push) Successful in 42s
CI & Build / Build & push image (push) Successful in 27s

Milestone #254 step 1 (#2286). A design system becomes a record Scribe holds
rather than prose in a rulebook: a named set of tokens with an OPTIONAL parent,
so a family system carries the house style and an app system carries only what
it changes. Answering "what does this app alter?" is then `list its tokens` —
nothing to compute.

`parent_id` is the whole model. It replaces both an `always_on` flag (a family
system is one with no parent) and a subscription join table (a project points at
ONE system; the chain supplies the rest) — less schema than the rulebook shape
it mirrors.

Two decisions the task left open, settled here:

- **Token values are JSONB keyed by mode**, not `value_light`/`value_dark`
  columns. The deciding argument was not flexibility, it was ambiguity: in a
  child system an unset mode means "inherit", in a root it means "not
  mode-dependent", and as columns both are NULL and the resolver cannot tell
  them apart. As a map, resolution is `{**parent, **child}` at every level with
  no special case for roots. Against it: queryability — but nothing filters
  tokens by value in SQL, so that buys a query no caller makes.
- **`group_name` is free text, no CHECK enum.** Groupings are each design
  system's own vocabulary; a whitelist would bake one install's kit into the
  schema. No CHECK is introduced anywhere, so rule #36 does not fire.

The cascade lives in `services/design_cascade.py` as pure functions over a
`{id: parent_id}` map, importing nothing — which is what lets both the service
and `access.py` use it without a cycle, and lets a test state a whole hierarchy
in one literal. Cycles are refused on WRITE by walking up from the proposed
parent (the cheap direction), and survived on READ by a visited-set, because a
loop from a direct DB edit must truncate rather than hang.

ACL (rule #78) is deliberately asymmetric: owning a system grants write,
reaching one through a project you can see grants READ ONLY. An editor on a
shared project must not be able to rewrite the family system every other project
in that family resolves through.

Also renames `services/design_system.py` -> `design_rulebook_import.py`. It is
the #251 prose extractor, whose role is already scheduled to become a one-shot
importer (#2288), and leaving it one character away from the new
`design_systems.py` was a trap for every later session.

Rule #115 throughout: nothing seeds a system or implies a default. An install
with zero design systems is ordinary, not degraded.
This commit is contained in:
2026-07-30 16:54:41 -04:00
parent 4ca3ab02c4
commit 03b3998585
13 changed files with 1022 additions and 3 deletions
+161
View File
@@ -0,0 +1,161 @@
"""Rulebook prose → checkable claims (milestone #251 step 2).
This is the piece of the design explorer that most needed to be testable, which
is why it lives in Python at all: the frontend has no test runner, so the fiddly
extraction happens server-side and the browser only does set arithmetic over it.
Rule text below is representative of a real design rulebook rather than copied
from this operator's — rule #115: the product must work for an install that has
none of their data, and a test that only passes against their exact wording would
be testing the instance, not the parser.
"""
from types import SimpleNamespace
from scribe.services.design_rulebook_import import (
expand_token_shorthand,
extract_expectations,
normalize_hex,
)
def _rule(rule_id, title, statement, how_to_apply=None):
return SimpleNamespace(
id=rule_id, title=title, statement=statement, how_to_apply=how_to_apply
)
# --- hex normalisation -------------------------------------------------------
def test_normalize_hex_makes_shorthand_and_case_comparable():
"""LOAD-BEARING. The rulebook writes `#FFFFFF` and components write `#fff`.
If those don't compare equal, the single largest drift finding — 67 hardcoded
white text colours (#2275) — reads as zero findings."""
assert normalize_hex("#fff") == normalize_hex("#FFFFFF") == "#ffffff"
assert normalize_hex("#E8E4D8") == "#e8e4d8"
assert normalize_hex("#14171a") == "#14171a"
def test_normalize_hex_keeps_alpha_rather_than_inventing_equality():
"""`#fff` and `#ffff` are different colours. Dropping the alpha to make them
match would manufacture agreement that isn't there."""
assert normalize_hex("#ffff") == "#ffffffff"
assert normalize_hex("#fff") != normalize_hex("#ffff")
def test_normalize_hex_rejects_non_colours():
for junk in ("", " ", "not-a-colour", "#", "#gg", "#12345"):
assert normalize_hex(junk) is None
# --- the slash shorthand -----------------------------------------------------
def test_expand_token_shorthand_handles_every_form_a_rulebook_uses():
"""One rule expands all three shapes: take everything up to and including the
LAST hyphen of the first segment as the prefix."""
assert expand_token_shorthand("--fs-radius-sm/md/lg/xl") == [
"--fs-radius-sm", "--fs-radius-md", "--fs-radius-lg", "--fs-radius-xl",
]
# Prefix is just `--fs-` here, and the same rule finds it.
assert expand_token_shorthand("--fs-obsidian/iron/slate/pewter") == [
"--fs-obsidian", "--fs-iron", "--fs-slate", "--fs-pewter",
]
assert expand_token_shorthand("--fs-dur-fast/base/slow") == [
"--fs-dur-fast", "--fs-dur-base", "--fs-dur-slow",
]
def test_expand_token_shorthand_passes_plain_names_through():
assert expand_token_shorthand("--fs-ease") == ["--fs-ease"]
# --- extraction --------------------------------------------------------------
def test_negation_is_scoped_to_the_sentence_not_the_rule():
"""THE trick that makes prohibition detection usable.
A single rule routinely states what the palette REQUIRES and what it FORBIDS
in consecutive sentences. Detecting negation across the whole statement would
mark the required colours as forbidden too — inverting the finding rather
than missing it, which is worse.
"""
rule = _rule(
52, "Text palette",
"Text tokens: Parchment #E8E4D8 (primary), Vellum #C2BFB4 (secondary), "
"Ash #9C9A92 (tertiary). Pure white #FFFFFF is NEVER used as text color.",
)
found = extract_expectations([rule])
required = {e.value for e in found if e.kind == "color"}
forbidden = {e.value for e in found if e.kind == "prohibited_color"}
assert required == {"#e8e4d8", "#c2bfb4", "#9c9a92"}
assert forbidden == {"#ffffff"}
assert not (required & forbidden)
def test_token_names_are_extracted_and_expanded():
rule = _rule(
72, "CSS custom properties",
"Expose the system as custom properties on :root — surfaces "
"(--fs-obsidian/iron/slate/pewter), radius (--fs-radius-sm/md/lg/xl), "
"and motion (--fs-ease).",
)
names = {e.value for e in extract_expectations([rule]) if e.kind == "token"}
assert "--fs-obsidian" in names and "--fs-pewter" in names
assert "--fs-radius-xl" in names
assert "--fs-ease" in names
assert len(names) == 9
def test_how_to_apply_is_read_as_well_as_the_statement():
"""Rulebooks routinely put the concrete values in how_to_apply and keep the
statement declarative, so ignoring it would miss the checkable half."""
rule = _rule(
56, "Per-app accent", "Each app owns exactly one accent.",
how_to_apply='[data-app="scribe"] #5B4A8A, [data-app="minstrel"] #4A6B5C.',
)
colours = {e.value for e in extract_expectations([rule]) if e.kind == "color"}
assert colours == {"#5b4a8a", "#4a6b5c"}
def test_claims_are_deduped_across_rules_keeping_the_first_source():
"""A colour named by several rules is one expectation, attributed to the rule
that introduced it — usually the most specific place to send a reader."""
rules = [
_rule(51, "Surfaces", "Obsidian #14171A is the page background."),
_rule(99, "Elsewhere", "Obsidian #14171A again, mentioned in passing."),
]
found = [e for e in extract_expectations(rules) if e.kind == "color"]
assert len(found) == 1
assert found[0].rule_id == 51
def test_prose_with_nothing_checkable_yields_nothing():
"""Most rules are judgement, not specification. They must contribute no
findings rather than a shrug — a panel that reports unparseable rules as
problems would be unusable."""
rule = _rule(
68, "Voice and tone",
"Voice is understated: plain language for anything functional, flavour "
"only where the user is waiting or failing. Be brief.",
)
assert extract_expectations([rule]) == []
def test_every_expectation_carries_the_sentence_it_came_from():
"""The panel has to show its working — "the rulebook says X" is only
actionable if you can see where, and in what context.
Asserts the context is the SENTENCE, not the whole statement: a rule that
states a requirement and a prohibition in consecutive sentences would
otherwise attribute both to the same undifferentiated blob of prose.
"""
rule = _rule(63, "Radius", "Radius: Small 4px. Pure white #FFFFFF is never used.")
found = extract_expectations([rule])
assert len(found) == 1
only = found[0]
assert only.kind == "prohibited_color"
assert only.rule_id == 63
assert only.rule_title == "Radius"
assert only.context == "Pure white #FFFFFF is never used."
assert "Radius: Small 4px" not in only.context