feat(snippets): add notes.data JSONB — the indexed mirror of snippet fields
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 12s
CI & Build / integration (push) Successful in 25s
CI & Build / Python tests (push) Failing after 29s
CI & Build / Build & push image (push) Has been skipped

Milestone #232 step 1 (task #2081). Takes the enabler first rather than the
write-path trigger: reverse lookup, drift checks and the duplicate finder all
need to QUERY structured fields, and building them on body-regex first means
writing them twice.

#227 deferred this bag "unless body-convention ergonomics prove insufficient" —
answering "which snippets live in this file?" by scanning every snippet and
regexing its body is that condition being met.

Migration 0070 adds `notes.data` (nullable JSONB) + a GIN index. The body is
UNCHANGED and still what gets embedded and read by humans; `data` mirrors the
same facts in a shape Postgres can index. Code is deliberately not copied into
it — the body holds it, and duplicating a blob into the column we index around
would be waste.

- compose_data() builds the mirror, omitting empties so the column stays sparse
- snippet_fields() prefers `data`, falling back to parsing the body. Rows written
  before 0070 have no `data` and are never backfilled, so a hand-edited body
  stays authoritative for them with no conversion deadline
- create / update / merge all write body and mirror from the same merged field
  set, so the two can't drift; merge in particular has to grow the mirror with
  the survivor's location set or a merged snippet would be unfindable at the very
  call sites the merge just recorded

Named `data`, not `metadata`, because that collides with SQLAlchemy's declarative
Base.metadata — which is why the pre-0069 model had to map an awkward
`entity_metadata` attribute. Not a revival of the column 0069 dropped: different
name, different purpose, nothing reads the old shape.

Two test fakes needed an explicit `data = None`: snippet_fields prefers `data`
when truthy and an auto-MagicMock attribute is truthy, so every parsed field
would have come back a MagicMock. Checked every fake reaching snippet code this
time rather than waiting for CI (note 2109).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RLwAaV4DQEmVyn496HnEvt
This commit is contained in:
2026-07-26 18:25:55 -04:00
parent 4b5d9005fd
commit cca40affe4
7 changed files with 264 additions and 9 deletions
+60
View File
@@ -0,0 +1,60 @@
"""add notes.data JSONB — queryable structured fields for typed records
Revision ID: 0070
Revises: 0069
Create Date: 2026-07-26
Snippets (note_type='snippet') carry structured fields — name, language,
signature, and a list of canonical locations (repo · path · symbol). Those were
stored as a markdown body-convention, which reads well and feeds the embedding
but cannot be QUERIED: answering "which snippets live in this file?" meant
scanning every snippet and regexing its body.
This adds a general `data` JSONB column plus a GIN index, so those fields become
indexable. The body stays exactly as it was — it is still the human-readable
form and still what gets embedded. `data` is a queryable mirror of the same
facts, not a replacement, and the code itself is deliberately NOT copied into it
(the body already holds it; duplicating a blob to index fields around it would
be waste).
Relationship to 0069: that migration DROPPED `notes.metadata`, a JSONB column
which only ever held person/place/list entity fields, when those surfaces were
removed. This is not a revival of it — different name, different purpose, and
nothing reads the old shape. The column is named `data` rather than `metadata`
because `metadata` collides with SQLAlchemy's declarative `Base.metadata`, which
is why the old model had to map an awkward `entity_metadata` attribute onto it.
Nullable with no backfill, deliberately: rows written before this migration keep
working because the service falls back to parsing the body when `data` is
absent. That means no migration deadline and no risk of a backfill mangling a
hand-edited body.
Downgrade drops the index and the column. Any structured fields it held remain
recoverable from the body convention, which is the same source they mirror.
"""
from alembic import op
import sqlalchemy as sa
from sqlalchemy.dialects.postgresql import JSONB
revision = "0070"
down_revision = "0069"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column("notes", sa.Column("data", JSONB, nullable=True))
# GIN supports containment (`data @> '{"locations":[{"repo":"x"}]}'`), which
# covers exact repo/path/symbol and language lookups. Path PREFIX matching
# ("everything under frontend/src/") is not an index-served operation here
# and still filters after the fact — acceptable while snippet counts are
# small, and a generated column is the escape hatch if that changes.
op.create_index(
"ix_notes_data_gin", "notes", ["data"], postgresql_using="gin",
)
def downgrade() -> None:
op.drop_index("ix_notes_data_gin", table_name="notes")
op.drop_column("notes", "data")