feat(embeddings): every vector records the model whose space it lives in (#4132)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Canceled after 27s

The four embedding tables stamped chunker_version but not the model, and
vector(384) is a width, not an identity: a same-width model swap would
write a second geometry beside the first with no error.

- embedding_model on note/rule/milestone/system embeddings (0109; existing
  rows stamped with the only model any install has ever run).
- Every write stamps EMBEDDING_MODEL; every backfill's "current" test is
  is_current_stamp(), both halves of calibration_stamp().
- migrate_floor refuses while any row its surface searches is off the live
  model, before sampling: re-embed, then migrate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-23 19:02:30 -04:00
co-authored by Claude Opus 5.5
parent 11b286d786
commit f4e9cd429b
10 changed files with 178 additions and 10 deletions
+3 -2
View File
@@ -17,7 +17,7 @@ from scribe.models.embedding import EMBEDDING_DIM, MilestoneEmbedding
from scribe.models.milestone import Milestone
from scribe.models.project import Project
from scribe.services import dedup as dedup_svc
from scribe.services.embeddings import CHUNKER_VERSION, semantic_search_milestones
from scribe.services.embeddings import CHUNKER_VERSION, EMBEDDING_MODEL, semantic_search_milestones
from tests.helpers import ensure_user
pytestmark = [pytest.mark.integration, pytest.mark.usefixtures("_dispose_engine", "_no_embedding")]
@@ -48,7 +48,8 @@ async def roadmap():
await s.flush()
for ms, vec in ((m3, NEAR), (done, NEAR), (unrelated, FAR), (foreign, NEAR)):
s.add(MilestoneEmbedding(milestone_id=ms.id, chunk_index=0, embedding=vec,
chunk_text=ms.title, chunker_version=CHUNKER_VERSION))
chunk_text=ms.title, chunker_version=CHUNKER_VERSION,
embedding_model=EMBEDDING_MODEL))
ids = {"owner": owner.id, "stranger": stranger.id, "mine": mine.id,
"m3": m3.id, "done": done.id, "unrelated": unrelated.id, "foreign": foreign.id}
await s.commit()