Files
FabledScribe/tests/test_task_document_shape.py
T
bvandeusenandClaude Opus 5 aa95c109ea
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 27s
feat(retrieval): a task's work logs join the document it is embedded as (#4251)
Step 1 of #4251. A work log is the richest prose Scribe holds about WHY
something is the way it is — written during the work, recording what was tried
and ruled out. #4241 made it readable from the agent's door. It was still not
findable, so "has anyone tried this approach?" — precisely the question a log
answers — could not reach one. The cost is not hypothetical: #4208 was rebuilt
in this session because its logs were unreachable.

THE DESIGN QUESTION the issue left open was whether logs embed as part of their
task's document or as rows of their own. As their own rows, a hit has to be
resolved back to a task to be worth anything, and it needs a fourth search, a
fourth result shape and a fourth arm. As part of the task, the objection is
that a long log drowns a short title.

That objection was true before #280 and is not true now. Chunking made one
record into one vector per section, so each log becomes its own title-anchored
chunk, scored separately, and the task's own prose keeps the chunk it always
had — a task is as findable as its best-matching log rather than as the average
of everything in it. The other half is this session's other build: a search
hands back the chunk that won (#4243), so a hit earned by a log shows that
log's passage under the task's title. Without that a reader would have got
body[:240] of the task — the opening of a record whose relevance lives three
hundred lines further down.

So `task_document(title, body, logs)` sits beside `rule_document`: a
synthesised embed-time shape, because the stored record is the task row and the
logs live in their own table, so the document that should be searchable exists
nowhere until it is built. A task with no logs is returned untouched — most
notes are not tasks and most tasks carry no log, and their vectors are the
corpus every tuned number here was measured against.

CHUNKER_VERSION 1 → 2, and its comment now says what the version actually
means. It used to read "whenever chunk_document's output can change for the
same input", which this change would slip past: `chunk_document` is untouched
and every task with a log now embeds differently while its title, body and the
chunker all stand still. The invariant is the document a record is embedded as.
The startup backfill re-embeds on that.

Create, edit and delete of a log all refresh the task through `embed_note`, the
one path every writer shares — an edited log whose vectors still carry its old
wording keeps matching what it no longer says.

No new kind enters the auto-inject menu: tasks were always in it, and this
makes recall on them better rather than changing what the menu spans. The
calibration stamp will now report shape_version 2 against numbers measured at
1, which is exactly the report #4104 built it to make.

Also promoted `session_returning` into tests/helpers beside `make_mock_session`
(#2834) — two files had spelled it out identically and a third was about to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-21 11:20:27 -04:00

201 lines
8.5 KiB
Python

"""task_document — the document a TASK is embedded as, work logs included (#4251).
A work log is the richest record Scribe holds of *why* something is the way it
is: prose written during the work, saying what was tried and ruled out. #4241
made it readable. It was still not findable, so "has anyone tried this
approach?" — the question a log answers — could not reach one, and #4208 was
rebuilt in this very session because its logs were unreachable.
These tests pin the shape and the two properties the decision rested on: a task
without logs embeds EXACTLY as it always did, and a task with them stays as
discriminative as it was, because each log is its own title-anchored chunk
rather than prose averaged into one vector.
"""
import datetime
from scribe.services.embeddings import (
WORK_LOG_HEADING,
chunk_document,
embedding_text,
task_document,
)
D1 = datetime.datetime(2026, 9, 20, 10, 0)
D2 = datetime.datetime(2026, 9, 21, 11, 0)
# --- the identity half: a task with no logs is untouched ---------------------
def test_a_task_with_no_logs_embeds_exactly_as_it_did_before():
"""Most notes are not tasks and most tasks carry no log. Their vectors are
the corpus this instance's thresholds were measured against, and changing
them for nothing would move every tuned number underneath itself (#4225)."""
assert task_document("A title", "A body") == ("A title", "A body")
assert task_document("A title", "A body", []) == ("A title", "A body")
assert task_document("T", None) == ("T", None)
def test_an_entry_with_no_content_is_skipped_rather_than_emitted_empty():
"""A bare heading is a vector containing nothing but the task's title — it
would compete with the task's real chunk and say nothing."""
_t, body = task_document("T", "prose", [(D1, " "), (D2, None), (D1, "real")])
assert body.count(WORK_LOG_HEADING) == 1
assert "real" in body
def test_a_task_that_is_only_logs_still_yields_a_document():
"""A task opened with a title and filled in entirely through its log — the
shape every step of a milestone starts as."""
title, body = task_document("T", "", [(D1, "the only content")])
assert title == "T"
assert body.startswith(WORK_LOG_HEADING)
assert "the only content" in body
# --- the shape: the logs are sections, oldest first --------------------------
def test_logs_are_appended_as_dated_sections_after_the_tasks_own_prose():
title, body = task_document(
"Fix the thing", "The task prose.",
[(D1, "Tried X, ruled out."), (D2, "Y worked.")],
)
assert title == "Fix the thing"
assert body.index("The task prose.") < body.index("Tried X") < body.index("Y worked.")
assert f"{WORK_LOG_HEADING} — 2026-09-20" in body
assert f"{WORK_LOG_HEADING} — 2026-09-21" in body
def test_an_entry_with_no_timestamp_still_gets_its_own_section():
"""Degrades to an undated heading rather than dropping the entry or
crashing — a restored row or a hand-built one is not a reason to lose the
richest prose in the record."""
_t, body = task_document("T", "p", [(None, "content from nowhere")])
assert WORK_LOG_HEADING in body
assert "content from nowhere" in body
# --- why folding them in is safe: chunking, not averaging --------------------
def test_each_log_becomes_its_own_title_anchored_chunk():
"""THE decision this build turns on. The objection to putting logs in the
task's document is that a long log drowns a short title — true before #280,
when one vector per record meant a 2,000-word log averaged the task's
subject away and everything past ~400 words was truncated unread. Chunking
answers both: separate vectors, each carrying the title as its anchor."""
def _entry(subject: str) -> str:
para = " ".join([f"This log entry discusses the {subject} in detail."] * 8)
return "\n\n".join(f"{para} (p{i})" for i in range(6))
title, body = task_document(
"Short task title", "Short body.",
[(D1, _entry("approach")), (D2, _entry("alternative"))],
)
chunks = chunk_document(title, body)
assert len(chunks) > 1, "a long log must split rather than truncate"
for chunk in chunks:
assert chunk.startswith("Short task title\n")
# The task's own prose still has a chunk of its own — it is not merged into
# a log section and scored as part of it.
assert any("Short body." in c and "alternative" not in c for c in chunks)
# And nothing is lost: this is what #280 exists for. Paragraph-shaped,
# because that is what a log is and because a single unbroken 3,000-char
# line is hard-split mid-line by design — `test_chunking` owns that case.
joined = "\n".join(chunks)
for line in body.splitlines():
if line.strip():
assert line.strip() in joined, f"content dropped: {line[:60]!r}"
def test_a_short_task_with_a_short_log_is_still_one_sharp_chunk():
"""Over-splitting is not the goal either. A task and one short log stay a
single document, the shape note #2485 measured as the sharpest."""
title, body = task_document("T", "body", [(D1, "a short note about it")])
assert chunk_document(title, body) == [embedding_text(title, body)]
# --- keeping it current: a log write refreshes the task's vectors ------------
#
# The shaper above is pure. These pin that the writers actually call it — a
# task whose logs are in the document but never re-indexed after one is written
# is #4241's half-surface one layer down: the entry is readable, and the search
# still answers as though it were never written.
from unittest.mock import MagicMock, patch # noqa: E402
import pytest # noqa: E402
from scribe.services import task_logs as svc # noqa: E402
from tests.helpers import ( # noqa: E402
fake_note, make_mock_session, session_returning,
)
@pytest.mark.asyncio
async def test_writing_a_work_log_refreshes_the_tasks_embedding():
note = fake_note(id=7, title="A task", body="prose", is_task=True)
session = session_returning(note)
with (
patch.object(svc, "async_session", return_value=session),
patch("scribe.services.notes.embed_note") as embed,
):
await svc.create_log(42, 7, "what I tried")
embed.assert_called_once()
assert embed.call_args.args[0] is note
@pytest.mark.asyncio
async def test_editing_a_work_log_refreshes_it_too():
"""An edited log whose vectors still carry the old wording keeps matching
what it no longer says — the stale-vector case `upsert_note_embedding`
already refuses to leave behind for a body."""
note = fake_note(id=7, title="A task", body="prose", is_task=True)
log = MagicMock(id=3, task_id=7)
session = make_mock_session()
found_log, found_note = MagicMock(), MagicMock()
found_log.scalars.return_value.first.return_value = log
found_note.scalars.return_value.first.return_value = note
session.execute.side_effect = [found_log, found_note]
with (
patch.object(svc, "async_session", return_value=session),
patch("scribe.services.notes.embed_note") as embed,
):
await svc.update_log(42, 3, content="reworded")
embed.assert_called_once_with(note)
@pytest.mark.asyncio
async def test_deleting_a_work_log_refreshes_it_too():
"""And reads the task id BEFORE the row goes — after the delete there is
nothing left to ask which task it belonged to."""
note = fake_note(id=7, title="A task", body="prose", is_task=True)
log = MagicMock(id=3, task_id=7)
session = make_mock_session()
found_log, found_note = MagicMock(), MagicMock()
found_log.scalars.return_value.first.return_value = log
found_note.scalars.return_value.first.return_value = note
session.execute.side_effect = [found_log, found_note]
with (
patch.object(svc, "async_session", return_value=session),
patch("scribe.services.notes.embed_note") as embed,
):
assert await svc.delete_log(42, 3) is True
embed.assert_called_once_with(note)
@pytest.mark.asyncio
async def test_a_failed_refresh_does_not_fail_the_log_that_saved():
"""Indexing never breaks a write. The startup backfill is the backstop."""
session = session_returning(fake_note(id=7, title="T", body="b", is_task=True))
with (
patch.object(svc, "async_session", return_value=session),
patch("scribe.services.notes.embed_note", side_effect=RuntimeError("boom")),
):
log = await svc.create_log(42, 7, "what I tried")
assert log is not None
session.commit.assert_awaited()