feat(telemetry): pull-through per surface, not just per corpus (#3311)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 30s
CI & Build / TypeScript typecheck (push) Successful in 33s
CI & Build / Python tests (push) Successful in 1m5s
CI & Build / Build & push image (push) Successful in 25s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 30s
CI & Build / TypeScript typecheck (push) Successful in 33s
CI & Build / Python tests (push) Successful in 1m5s
CI & Build / Build & push image (push) Successful in 25s
The readout already grouped usage by source — `group_by(event, source)` — and the loop directly below it threw the source away, collapsing every surface into one corpus-wide ratio. So the question a threshold is actually tuned against, "is THIS surface worth its noise", could not be asked of any surface, while the data to answer it sat in the table. `usage.by_source` reports notes_surfaced / notes_pulled / pull_through per surface. The grain is the note, not the call: a pull records the door it came through, not the surface that led there, so grouping the pulled rows by source would answer a different question. Joining surfaced rows to pulled rows on note_id answers this one without the session identity #2085 declined to invent — at the cost of being an upper bound per surface, which the docstring says where it is read. Ambient surfaces report counts and a null ratio: nothing chose those records, so "surfaced often, opened never" is not a judgment about them. A surface that genuinely produced nothing reports 0.0, which must not look like the null. The join is guarded separately from the two reads above it. #2663 was a novel SQL shape the database rejected inside a broad except; this is the novel shape here, and it must not take down two readouts that work. Tests are integration for that same reason — a mock passes on a query Postgres refuses. They pin the distinct-first property (three surfacings of one note are one note), the ambient null, and the LIKE escape, since an unescaped `mcp_%` also matches `mcpXget_note` and nothing else in the payload would show the difference.
This commit is contained in:
@@ -182,6 +182,23 @@ async def retrieval_telemetry(days: int = 30) -> dict:
|
||||
tuned against — only by a pull the agent made. Aggregating across the
|
||||
mcp_/rest_ prefix would silently answer the wrong one.
|
||||
|
||||
`usage["by_source"]` — THE number to tune a threshold against, because the
|
||||
top-level `pull_through` is a corpus average and averages the surfaces
|
||||
together. Per surface: `notes_surfaced`, `notes_pulled`, `pull_through`,
|
||||
and `ambient: true` on surfaces whose surfacings were not scored choices
|
||||
(their ratio is null — "surfaced often, opened never" is not a judgment
|
||||
about a record nothing chose). Read it as: of the distinct notes THIS
|
||||
surface put in front of the agent, how many did the agent then open?
|
||||
|
||||
Two limits on it, both deliberate. It is an UPPER BOUND per surface: a pull
|
||||
records the door it came through, not the surface that led there, so a note
|
||||
surfaced by two surfaces and opened once counts for both — attribution
|
||||
would need the session identity #2085 declined to invent. And RULE
|
||||
surfacings are absent: `write_path_rule` appears in `sources` with its
|
||||
scores but has no usage counter at all, so it has no row here (#3311).
|
||||
`by_source_failed: true` means that one query failed while the rest of the
|
||||
readout stood.
|
||||
|
||||
Scoped to your own telemetry — a retrieval log records what your agent
|
||||
asked for, query text included, and is not a shared record kind.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user