feat(embeddings): every vector records the model whose space it lives in (#4132)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Canceled after 27s

The four embedding tables stamped chunker_version but not the model, and
vector(384) is a width, not an identity: a same-width model swap would
write a second geometry beside the first with no error.

- embedding_model on note/rule/milestone/system embeddings (0109; existing
  rows stamped with the only model any install has ever run).
- Every write stamps EMBEDDING_MODEL; every backfill's "current" test is
  is_current_stamp(), both halves of calibration_stamp().
- migrate_floor refuses while any row its surface searches is off the live
  model, before sampling: re-embed, then migrate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-23 19:02:30 -04:00
co-authored by Claude Opus 5.5
parent 11b286d786
commit f4e9cd429b
10 changed files with 178 additions and 10 deletions
+2 -1
View File
@@ -18,7 +18,7 @@ from scribe.models.embedding import EMBEDDING_DIM, RuleEmbedding
from scribe.models.project import Project
from scribe.models.share import ProjectShare
from scribe.services import rulebooks as rulebooks_svc
from scribe.services.embeddings import CHUNKER_VERSION, semantic_search_rules
from scribe.services.embeddings import CHUNKER_VERSION, EMBEDDING_MODEL, semantic_search_rules
from tests.helpers import ensure_user
pytestmark = [pytest.mark.integration, pytest.mark.usefixtures("_dispose_engine")]
@@ -66,6 +66,7 @@ async def homes():
s.add(RuleEmbedding(
rule_id=rule.id, chunk_index=0, embedding=QUERY_VEC,
chunk_text=rule.title, chunker_version=CHUNKER_VERSION,
embedding_model=EMBEDDING_MODEL,
))
await s.commit()
ids.update(glob=glob.id, on_a=on_a.id, on_b=on_b.id)