Logs could be written and never read back from the tool that opens the task, so a session rebuilt work a previous one had already ruled out. Adds logs_for_task / count_logs_for_task / log_counts_for_tasks, scoped through services/access (rule 78) — by who may read the TASK, not by who wrote each entry, so a collaborator on a shared task doesn't get an empty list that reads as "no work has been done".
#4243 / #4252 — search showed a span it had already ranked lower
The search collapses to best-chunk-per-record and then previewed body[:240], unmarked. All three semantic searches now publish report["best_chunk"], and every surface that shows a span says which span it is (matched_passage vs body_opening). services/text.matched_excerpt is the one place that decides; elision cuts from the middle, keeping both ends.
#4250 — the agent's search could filter to 2 kinds, browse to 9
note_type and task_kind were engine parameters with no way to say them. Both non-web doors now derive from _FACETS via one content_type_filters. A third hand-kept copy turned up in routes/search.py whose unknown-value fallback was no filter — ?content_type=snippets returned the whole corpus while looking narrowed. Unknown kinds are now refused (ValueError / 400) rather than answered with an empty set.
#4251 — work logs and System charters become findable
Work logs join their task's embedded document (task_document, beside rule_document); each log is its own title-anchored chunk, so a task is as findable as its best log. Systems get system_embeddings (migration 0107) and search(content_type="system") — "where does this belong?" had no tool.
Step 3 (design guidance, project goals) was settled without building: both are one-per-project, so there is nothing for a search to select among. Residual filed as #4256.
CHUNKER_VERSION 1 → 2. The startup backfill will re-embed the note corpus and embed every System charter for the first time. Expect a longer first boot.
The calibration stamp will begin reporting shape_version: 2 against numbers measured at 1 — that is #4104 working, not a regression.
Nine commits. Four issues, all closed on `dev` with CI green on the tip (runs 7172, 7175, 7176).
## #4241 — `add_task_log` was write-only for an agent
Logs could be written and never read back from the tool that opens the task, so a session rebuilt work a previous one had already ruled out. Adds `logs_for_task` / `count_logs_for_task` / `log_counts_for_tasks`, scoped through `services/access` (rule 78) — by who may read the TASK, not by who wrote each entry, so a collaborator on a shared task doesn't get an empty list that reads as "no work has been done".
## #4243 / #4252 — search showed a span it had already ranked lower
The search collapses to best-chunk-per-record and then previewed `body[:240]`, unmarked. All three semantic searches now publish `report["best_chunk"]`, and every surface that shows a span says which span it is (`matched_passage` vs `body_opening`). `services/text.matched_excerpt` is the one place that decides; elision cuts from the middle, keeping both ends.
## #4250 — the agent's search could filter to 2 kinds, browse to 9
`note_type` and `task_kind` were engine parameters with no way to say them. Both non-web doors now derive from `_FACETS` via one `content_type_filters`. A third hand-kept copy turned up in `routes/search.py` whose unknown-value fallback was *no filter* — `?content_type=snippets` returned the whole corpus while looking narrowed. Unknown kinds are now refused (ValueError / 400) rather than answered with an empty set.
## #4251 — work logs and System charters become findable
Work logs join their task's embedded document (`task_document`, beside `rule_document`); each log is its own title-anchored chunk, so a task is as findable as its best log. Systems get `system_embeddings` (migration 0107) and `search(content_type="system")` — "where does this belong?" had no tool.
Step 3 (design guidance, project goals) was settled without building: both are one-per-project, so there is nothing for a search to select among. Residual filed as #4256.
## Deploy notes
- **Migration 0107** creates `system_embeddings` (HNSW index, derived — declared in backup's `_NOT_INCLUDED`).
- **CHUNKER_VERSION 1 → 2.** The startup backfill will re-embed the note corpus and embed every System charter for the first time. Expect a longer first boot.
- The calibration stamp will begin reporting `shape_version: 2` against numbers measured at 1 — that is #4104 working, not a regression.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Found by reading a live `retrieval_telemetry` readout after milestone 419
deployed, not by inspection. The readout said:
cannot_decline / report_preference — "45 calls, 0 of them returned
nothing. An arm that fires unasked has to be able to say nothing; this one
never has. Check that it applies its floor at all."
And printed, beside it, that arm's band: p10 = p50 = p90 = min = max = 0.791.
FIVE IDENTICAL PERCENTILES IS THE TELL. That is not a ranking, it is one
record at one score on every call — because `report_preference` searches a
fixed string (`reply_preferences.COMPLETION_QUERY`, a module constant, and
deliberately so).
For a fixed query against a stable corpus the top score is a CONSTANT, so the
arm's decline rate is 0% or 100% and never in between; which of the two it is
depends only on where the bar sits relative to that one number. "Never
returned nothing" is therefore arithmetic, not evidence, and the warning's own
remedy — check whether it applies a floor — cannot be answered from it.
The arm already knew this about itself; the warning did not:
"a fixed query makes this arm's score a constant and a floor a hair above
it produces a dead arm no amount of traffic will ever reveal"
— services/reply_preferences.py
THIS CLASS OF BUG ALREADY HAS A GUARD, which is the argument for the shape of
the fix. `Point.logs_unconditionally` exists because of #3497: both rule arms
once logged only their hits, so their zero count was structurally 0 and this
same warning would have fired on a LOGGING property while sending the reader
to move a threshold that was never involved. This is that one step over — a
QUERY-SHAPE property — and gets the same treatment: a declared field on
`Point`, and exclusion rather than trust.
AND THE WARNING THAT WOULD BE INFORMATIVE HERE DID NOT EXIST. For a fixed-query
arm the dangerous state is the mirror image: every call empty, meaning the bar
is above the constant and no further traffic will ever move it. The arm is off
rather than quiet, and nothing in the readout said so — `expects_traffic`
covers an arm with NO calls, not one with calls and a 100% decline rate. That
state is real and reached: `report_preference` once logged 69 consecutive
declines at 0.0006 under its bar.
So `fixed_query_never_clears` sends the reader to `near_miss_samples` and not
to the dial — because that incident is also the one where the statistic and
the correct action pointed opposite ways. Every percentile said lower the
floor; opening the refused record showed it was rule 77 arriving as a false
positive, and lowering it would have delivered that rule on every completion
report ever written.
Guards in tests/test_retrieval_warnings.py, including the falsifier that
matters most here: `cannot_decline` must still fire on an arm whose query
varies, or this change is a disabled check wearing a narrowed one's clothes.
Both boundaries tested from both sides, per that module's own standard.
The new code is documented in the `retrieval_telemetry` tool docstring beside
the others (rule 33) — an undocumented code in a readout is a reader meeting a
verdict with no way to disagree with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
The work log reached the web UI through routes/task_logs.py and nothing
else. get_task returned only the body — a claim written once, before the
work — with the record written during it invisible beside it. So a stale
body arrived with nothing to contradict it, and this session rebuilt work
that had already shipped, with the evidence sitting in the task's own logs.
Read side, scoped through the access layer (rule 78):
- logs_for_task / count_logs_for_task / log_counts_for_tasks in
services/task_logs.py. Scoped by who may read the TASK rather than by
who wrote the entry: list_logs filters TaskLog.user_id == user_id,
which hands a shared collaborator an empty list reading as "no work
has been done". The page query folds readable_notes_clause into the
same statement so the permission does not become an N+1.
- get_task returns work_log; list_tasks and get_milestone steps carry
log_count, zero-filled so "none" is a count and not a missing key.
Elision keeps both ends. The newest entry arrives whole to 4000 chars
because it answers "where does this stand"; older ones are shortened from
the MIDDLE, never the head. A head cut selects what a reader sees by
character position, which is uncorrelated with what matters — an entry
closing with "so this shipped in 04775c3" loses the one sentence that
answers the question, and a truncated flag says something went, never
whether it mattered. The gap states how many characters it covers.
conftest gains an autouse stub for the new read arm, same reasoning as
_no_rule_arm: three widely-called tools grew a database read, and the
existing call sites should not each have to learn about it.
Raised while reviewing this: search() has the same shape and worse —
body[:240] with no marker at all, while the chunk that actually matched
sits unused in the row that won. Filed as #4243, not fixed here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Raised by the operator: are we limiting what comes back by character count,
and how do we verify the pertinent part is the part displayed?
We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut,
with no marker that anything had been removed, so a 240-character preview of
a 4000-character record was indistinguishable from a complete short one.
The opening is the wrong span. The match is semantic and per chunk, and
semantic_search_notes collapses to best-chunk-per-note — its own comment at
the collapse says "the first appearance of a note is its best chunk". So the
system identified the passage that earned the hit and then discarded it:
select(Note, distance) kept no chunk column. A record could rank first on its
sixth paragraph, be previewed by its first, and be judged irrelevant on a
span the search had already scored lower. That biases against long records,
and it is self-concealing — the caller who does not open it never learns the
preview was misleading.
- embeddings: chunk_index/chunk_text ride along in the select, and the
collapse records the winner in report["best_chunk"]. Carried in `report`,
NOT by widening the return tuple: ten callers unpack (score, note) at
~18 sites and nothing would catch the misses (lesson #4207). `report` is
the side-channel this function already uses for best_available_score.
- search(): excerpt / excerpt_is / body_length, and read_full when there is
more. A caller that cannot tell a matched passage from a document opening
cannot judge whether to look deeper, which is the only decision the field
supports.
elide() moves to services/text.py so both callers share one copy, and it
keeps BOTH ends with a stated gap — it is the fallback for when nothing
identifies a better span than "all of it", not the goal.
Also fixes a guard that produced a false failure on the previous commit:
test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`,
which sees only the first line, so wrapping get_task's signature over four
lines made it report a function that does take the project as one that does
not. Parsed with ast now, and proven to still reject an absent or
wrongly-typed parameter rather than being appeased by reflowing the code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
test_chunking::test_search_collapses_chunk_rows_to_best_chunk_per_note feeds
rows straight into the real semantic_search_notes, so widening the select to
carry chunk_index/chunk_text broke its 2-tuple fakes. I had claimed no unit
test did this after grepping test_embeddings.py, which was the wrong file.
Rows are 4-tuples now, and the test additionally asserts the winning chunk is
REPORTED and not merely used for scoring — the property #4243 exists for, and
this is where it belongs, beside the collapse it comes from.
test_task_work_log_surface's ACL test asserted "task_logs.user_id" was absent
from the compiled statement. It is present in every SELECT as a projected
column; the claim was about the WHERE clause. Asserts on stmt.whereclause now,
which is what was actually meant and still fails if an owner filter returns.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
#4243 fixed one door. Scribe has three semantic searches over three chunk
tables, and all three collapsed chunk rows to the best one per record — each
of them KNEW which passage earned the hit, and each dropped it. Every surface
downstream then previewed the head of the document instead: a span the search
had already scored lower, with nothing saying so.
Mechanism, one place:
- embeddings.record_best_chunk publishes {id: {index, text}} into `report`.
Carried in `report`, NOT the return value: all three return
list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207).
- semantic_search_rules and semantic_search_milestones now select
chunk_index/chunk_text and publish the winner, as notes already did.
semantic_search_milestones gains `report`, which it had no way to take.
- services/text.matched_excerpt is the one choice of span, and
excerpt_fields the one result block. Doors keep their own field names —
the web renders `snippet`, MCP returns `excerpt` — because renaming a
field a frontend reads is a different change from fixing what goes in it.
Surfaces:
- knowledge.query_knowledge, whose own comment calls it "the human's MAIN
search surface", was `(note.body or "")[:200]` on every row alike. Now the
matched passage on a search, the opening on a browse, and `snippet_is`
saying which. KnowledgeView renders that snippet, so this was live.
- search(content_type='milestone') gains `matched` — the plan body stays
out, but the passage that matched comes along, because recognising a plan
means recognising the part you asked about and a description written at
the start need not mention it.
- The auto-inject menu and the write-path prior-art menu put the passage
under their line. Both were title-only, which answers "does this apply?"
for a lesson or snippet (the trigger is IN the title) and not at all for
an issue or dev-log. No fallback to the body's opening: on a menu that is
preamble dressed as a reason, and once indented it cannot be told apart.
Left alone deliberately: the rule arms. A rule hint already renders the rule's
TRIGGER, which is written to answer exactly "does this apply to me" and beats
a matched chunk at it; and that line's budget was measured at #3851. Adding a
passage there would duplicate the trigger and spend the budget twice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
test_milestone_search_is_its_own_shape asserts the result dict exactly, so
adding `matched` / `matched_is` / `body_length` broke it — which is the test
doing its job.
Its fake was a bare MagicMock, so `m.body` autovivified: `matched` came out as
a mock object and `body_length` as 0, and nothing in the old assertion touched
either. Lesson #2833 is about exactly this, and the failure output is what
surfaced it. The fake now carries a real body string and a report with the
chunk that matched, so the assertion is about the product rather than about
MagicMock's attribute behaviour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
The engine took `note_type` and `task_kind` all along. What was missing was a
way to say them: the MCP tool's `content_type` knew `note`, `task` and `all`,
and `/api/search` knew the same two — so an agent could not ask "has this
snippet already been recorded" or "what lessons apply here" without searching
everything and reading past the rest. Browse offered nine kinds from the same
data.
The cause is that each door kept its own map. `_FACETS` in services/knowledge
is where a kind is declared, and #3161 made adding one a single edit by
generating the SQL filter, the Python predicate and the door's validation from
it — but the two search doors were written before that and never joined. So
this adds the third dialect, `search_filters_for`, and one composition over it,
`content_type_filters`, and both doors now derive instead of listing.
Two names keep a meaning of their own, and the docstrings say so: `all` is no
filter, and `note` is BROAD — any non-task, snippets and lessons included —
where the browse facet of the same name is narrow (`note_type == 'note'`).
They are left different deliberately; narrowing this one would stop returning
snippets to every caller that already asks this way.
An unrecognised kind is now refused rather than answered. Both doors used to
fall through: the MCP tool into a filter matching no row, the route into no
filter at all, so `?content_type=snippets` returned the whole corpus while
looking like a narrowed search. An empty result set is a claim — "the corpus
holds nothing like this" — and an agent acts on that claim by building the
thing it could not find, so a typo must not be able to make it.
The docstring is the agent-facing contract (#2846), and a test now holds it to
the table: every kind `_FACETS` declares has to appear in it, because a filter
an agent has not been told about is unreachable however well it is wired.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Step 1 of #4251. A work log is the richest prose Scribe holds about WHY
something is the way it is — written during the work, recording what was tried
and ruled out. #4241 made it readable from the agent's door. It was still not
findable, so "has anyone tried this approach?" — precisely the question a log
answers — could not reach one. The cost is not hypothetical: #4208 was rebuilt
in this session because its logs were unreachable.
THE DESIGN QUESTION the issue left open was whether logs embed as part of their
task's document or as rows of their own. As their own rows, a hit has to be
resolved back to a task to be worth anything, and it needs a fourth search, a
fourth result shape and a fourth arm. As part of the task, the objection is
that a long log drowns a short title.
That objection was true before #280 and is not true now. Chunking made one
record into one vector per section, so each log becomes its own title-anchored
chunk, scored separately, and the task's own prose keeps the chunk it always
had — a task is as findable as its best-matching log rather than as the average
of everything in it. The other half is this session's other build: a search
hands back the chunk that won (#4243), so a hit earned by a log shows that
log's passage under the task's title. Without that a reader would have got
body[:240] of the task — the opening of a record whose relevance lives three
hundred lines further down.
So `task_document(title, body, logs)` sits beside `rule_document`: a
synthesised embed-time shape, because the stored record is the task row and the
logs live in their own table, so the document that should be searchable exists
nowhere until it is built. A task with no logs is returned untouched — most
notes are not tasks and most tasks carry no log, and their vectors are the
corpus every tuned number here was measured against.
CHUNKER_VERSION 1 → 2, and its comment now says what the version actually
means. It used to read "whenever chunk_document's output can change for the
same input", which this change would slip past: `chunk_document` is untouched
and every task with a log now embeds differently while its title, body and the
chunker all stand still. The invariant is the document a record is embedded as.
The startup backfill re-embeds on that.
Create, edit and delete of a log all refresh the task through `embed_note`, the
one path every writer shares — an edited log whose vectors still carry its old
wording keeps matching what it no longer says.
No new kind enters the auto-inject menu: tasks were always in it, and this
makes recall on them better rather than changing what the menu spans. The
calibration stamp will now report shape_version 2 against numbers measured at
1, which is exactly the report #4104 built it to make.
Also promoted `session_returning` into tests/helpers beside `make_mock_session`
(#2834) — two files had spelled it out identically and a third was about to.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Step 2 of #4251. A System's `description` is a charter — several hundred words
saying what belongs in that area and what does not — and it is the answer to
"which part of this codebase does X live in". There was no semantic path to
one: `list_systems` enumerates, and `search(system_id=…)` uses a System as a
FILTER over notes. So a System could narrow a search and could never be the
answer to one, and an agent asking where a record belonged had to read every
charter or guess.
ITS OWN SEARCH, not a `content_type` over notes, for the reason note 3163
gives about milestones: the row could be shared, the search cannot. A charter
competing with the whole note corpus for one top-k is outranked by the records
filed under it — the right answer crowded out by its own contents — and "where
does this belong?" is a different question from "what prior art is there?",
which a caller asking one should not have to read past answers to.
So `system_embeddings` (0107) joins note_, rule_ and milestone_embeddings as
the fourth sibling, with `system_document`, `upsert_system_embedding`,
`semantic_search_systems`, a startup backfill and `search(content_type=
"system")`. Scoped like milestones: with a project_id, that project's Systems
if the caller can read the project (rule 78); without one, the caller's own.
Archived Systems are excluded — an archived area is one the operator has said
is no longer where things go, which is exactly the question being asked.
`system_document` is the plainest of the four shapes on purpose. A charter is
already written as the thing this search has to match, in the words someone
asking would use — so there is no trigger to synthesise as `rule_document`
must, and no second record to gather as `task_document` must. The stored
charter IS the sharp document, the way a snippet's is. `color`, `status` and
`order_index` stay out: presentation and bookkeeping, and a vector carrying
them would be answering a question nobody asks of a charter.
The search publishes `report["best_chunk"]` from the start rather than being
retrofitted, which is what #4251 asked of any fourth search. It matters more
here than anywhere: a charter runs long and a result shows its NAME, so a match
on the paragraph that actually decides where a record belongs would otherwise
be previewed by two words that cannot say.
The id that comes back is the one `system_id`, `system_ids` and
`list_system_records` already take, so the answer to "where does this belong?"
is directly usable as "show me what is there" and as "file it here".
`embed_system` sits beside `notes.embed_note` at the service for #2056's
reason — every door gets it by construction. Not called on delete: that is a
soft delete and the search joins through `System`, so the vectors are already
unreachable, and leaving them means a restore is findable again immediately.
`system_embeddings` is declared in backup's `_NOT_INCLUDED` as derived, beside
its three siblings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Nine commits. Four issues, all closed on
devwith CI green on the tip (runs 7172, 7175, 7176).#4241 —
add_task_logwas write-only for an agentLogs could be written and never read back from the tool that opens the task, so a session rebuilt work a previous one had already ruled out. Adds
logs_for_task/count_logs_for_task/log_counts_for_tasks, scoped throughservices/access(rule 78) — by who may read the TASK, not by who wrote each entry, so a collaborator on a shared task doesn't get an empty list that reads as "no work has been done".#4243 / #4252 — search showed a span it had already ranked lower
The search collapses to best-chunk-per-record and then previewed
body[:240], unmarked. All three semantic searches now publishreport["best_chunk"], and every surface that shows a span says which span it is (matched_passagevsbody_opening).services/text.matched_excerptis the one place that decides; elision cuts from the middle, keeping both ends.#4250 — the agent's search could filter to 2 kinds, browse to 9
note_typeandtask_kindwere engine parameters with no way to say them. Both non-web doors now derive from_FACETSvia onecontent_type_filters. A third hand-kept copy turned up inroutes/search.pywhose unknown-value fallback was no filter —?content_type=snippetsreturned the whole corpus while looking narrowed. Unknown kinds are now refused (ValueError / 400) rather than answered with an empty set.#4251 — work logs and System charters become findable
Work logs join their task's embedded document (
task_document, besiderule_document); each log is its own title-anchored chunk, so a task is as findable as its best log. Systems getsystem_embeddings(migration 0107) andsearch(content_type="system")— "where does this belong?" had no tool.Step 3 (design guidance, project goals) was settled without building: both are one-per-project, so there is nothing for a search to select among. Residual filed as #4256.
Deploy notes
system_embeddings(HNSW index, derived — declared in backup's_NOT_INCLUDED).shape_version: 2against numbers measured at 1 — that is #4104 working, not a regression.🤖 Generated with Claude Code
https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
Found by reading a live `retrieval_telemetry` readout after milestone 419 deployed, not by inspection. The readout said: cannot_decline / report_preference — "45 calls, 0 of them returned nothing. An arm that fires unasked has to be able to say nothing; this one never has. Check that it applies its floor at all." And printed, beside it, that arm's band: p10 = p50 = p90 = min = max = 0.791. FIVE IDENTICAL PERCENTILES IS THE TELL. That is not a ranking, it is one record at one score on every call — because `report_preference` searches a fixed string (`reply_preferences.COMPLETION_QUERY`, a module constant, and deliberately so). For a fixed query against a stable corpus the top score is a CONSTANT, so the arm's decline rate is 0% or 100% and never in between; which of the two it is depends only on where the bar sits relative to that one number. "Never returned nothing" is therefore arithmetic, not evidence, and the warning's own remedy — check whether it applies a floor — cannot be answered from it. The arm already knew this about itself; the warning did not: "a fixed query makes this arm's score a constant and a floor a hair above it produces a dead arm no amount of traffic will ever reveal" — services/reply_preferences.py THIS CLASS OF BUG ALREADY HAS A GUARD, which is the argument for the shape of the fix. `Point.logs_unconditionally` exists because of #3497: both rule arms once logged only their hits, so their zero count was structurally 0 and this same warning would have fired on a LOGGING property while sending the reader to move a threshold that was never involved. This is that one step over — a QUERY-SHAPE property — and gets the same treatment: a declared field on `Point`, and exclusion rather than trust. AND THE WARNING THAT WOULD BE INFORMATIVE HERE DID NOT EXIST. For a fixed-query arm the dangerous state is the mirror image: every call empty, meaning the bar is above the constant and no further traffic will ever move it. The arm is off rather than quiet, and nothing in the readout said so — `expects_traffic` covers an arm with NO calls, not one with calls and a 100% decline rate. That state is real and reached: `report_preference` once logged 69 consecutive declines at 0.0006 under its bar. So `fixed_query_never_clears` sends the reader to `near_miss_samples` and not to the dial — because that incident is also the one where the statistic and the correct action pointed opposite ways. Every percentile said lower the floor; opening the refused record showed it was rule 77 arriving as a false positive, and lowering it would have delivered that rule on every completion report ever written. Guards in tests/test_retrieval_warnings.py, including the falsifier that matters most here: `cannot_decline` must still fire on an arm whose query varies, or this change is a disabled check wearing a narrowed one's clothes. Both boundaries tested from both sides, per that module's own standard. The new code is documented in the `retrieval_telemetry` tool docstring beside the others (rule 33) — an undocumented code in a readout is a reader meeting a verdict with no way to disagree with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92KuyThe work log reached the web UI through routes/task_logs.py and nothing else. get_task returned only the body — a claim written once, before the work — with the record written during it invisible beside it. So a stale body arrived with nothing to contradict it, and this session rebuilt work that had already shipped, with the evidence sitting in the task's own logs. Read side, scoped through the access layer (rule 78): - logs_for_task / count_logs_for_task / log_counts_for_tasks in services/task_logs.py. Scoped by who may read the TASK rather than by who wrote the entry: list_logs filters TaskLog.user_id == user_id, which hands a shared collaborator an empty list reading as "no work has been done". The page query folds readable_notes_clause into the same statement so the permission does not become an N+1. - get_task returns work_log; list_tasks and get_milestone steps carry log_count, zero-filled so "none" is a count and not a missing key. Elision keeps both ends. The newest entry arrives whole to 4000 chars because it answers "where does this stand"; older ones are shortened from the MIDDLE, never the head. A head cut selects what a reader sees by character position, which is uncorrelated with what matters — an entry closing with "so this shipped in 04775c3" loses the one sentence that answers the question, and a truncated flag says something went, never whether it mattered. The gap states how many characters it covers. conftest gains an autouse stub for the new read arm, same reasoning as _no_rule_arm: three widely-called tools grew a database read, and the existing call sites should not each have to learn about it. Raised while reviewing this: search() has the same shape and worse — body[:240] with no marker at all, while the chunk that actually matched sits unused in the row that won. Filed as #4243, not fixed here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92KuyRaised by the operator: are we limiting what comes back by character count, and how do we verify the pertinent part is the part displayed? We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut, with no marker that anything had been removed, so a 240-character preview of a 4000-character record was indistinguishable from a complete short one. The opening is the wrong span. The match is semantic and per chunk, and semantic_search_notes collapses to best-chunk-per-note — its own comment at the collapse says "the first appearance of a note is its best chunk". So the system identified the passage that earned the hit and then discarded it: select(Note, distance) kept no chunk column. A record could rank first on its sixth paragraph, be previewed by its first, and be judged irrelevant on a span the search had already scored lower. That biases against long records, and it is self-concealing — the caller who does not open it never learns the preview was misleading. - embeddings: chunk_index/chunk_text ride along in the select, and the collapse records the winner in report["best_chunk"]. Carried in `report`, NOT by widening the return tuple: ten callers unpack (score, note) at ~18 sites and nothing would catch the misses (lesson #4207). `report` is the side-channel this function already uses for best_available_score. - search(): excerpt / excerpt_is / body_length, and read_full when there is more. A caller that cannot tell a matched passage from a document opening cannot judge whether to look deeper, which is the only decision the field supports. elide() moves to services/text.py so both callers share one copy, and it keeps BOTH ends with a stated gap — it is the fallback for when nothing identifies a better span than "all of it", not the goal. Also fixes a guard that produced a false failure on the previous commit: test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`, which sees only the first line, so wrapping get_task's signature over four lines made it report a function that does take the project as one that does not. Parsed with ast now, and proven to still reject an absent or wrongly-typed parameter rather than being appeased by reflowing the code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy