feat(mcp): enter_project becomes a small primer: goal, recent work, open work, vocabulary (#4045)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 46s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 23s

The handshake carried the whole project record, every milestone's plan, full
rule text, the notes most recently edited and ~9k of design guidance. For
project 2 that was ~222k characters, past what an MCP client accepts as a tool
result. Each category was walked through with the operator and sized to what a
session needs on arrival; each names the call that has the rest.

- project: id, title, status and the full goal (session start's "full goal"
  pointer still lands here). get_project keeps the whole record.
- milestone_summary: the 5 most recently touched milestones, any status, most
  recent first, without plans. Summaries gain last_touched_at: the later of
  the milestone's own edit and its newest step update, from the query that
  already counts steps. milestone_summary_omitted counts the rest and points
  to list_milestones. get_project and list_milestones list every milestone,
  also without plans.
- open_tasks: the 10 most recently touched open tasks, with or without a
  milestone, each naming its milestone. list_notes gains sort="touched"
  (the later of updated_at and the newest work-log), because a log doesn't
  bump updated_at.
- recent_notes: dropped. Retrieval surfaces notes by relevance, and
  get_recent covers recency.
- systems: id and name.
- design_system: summary plus guidance_call. get_design_system gains
  resolved_guidance, the chain-merged prose; its own guidance field is only
  the departures, so session start's old pointer to it led to a fragment.
  The session start pointer and using-scribe's "Building UI" section now
  name resolved_guidance.
- rules: rules_payload(brief=True) gives project_rules as id and title plus
  subscribed_rulebooks, and records only what it shows. Retrieval delivers
  rules in full and ignores subscriptions (#4052). Other callers unchanged.
- pattern_coverage, inception and systems_bootstrap: unchanged.

Clients: the plugin's using-scribe skill, the compaction notice and session
start are updated here; the REST project summary only gains last_touched_at.
Plugin version minted.

Tests: a size ceiling on the handshake for a large project; milestone and
task selection and naming; brief rules; resolved_guidance; the session
start pointer; and a real-Postgres test that a work-log touches its task and
a step update touches its milestone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-14 22:04:38 -04:00
co-authored by Claude Opus 5
parent 9b2de3552f
commit 7f974d9749
17 changed files with 445 additions and 182 deletions
+14 -3
View File
@@ -14,6 +14,7 @@ So these tests assert the number of sessions opened, not only the values
returned. A version that produced identical output while opening a session per
project would pass a correctness test and reproduce the outage.
"""
from datetime import datetime, timezone
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
@@ -99,17 +100,27 @@ async def test_milestone_summaries_for_many_projects_open_ONE_session():
a session per MILESTONE, which is what turned 25 into ~250."""
from scribe.services import milestones as svc
m1 = MagicMock(id=10, project_id=1)
written = datetime(2026, 9, 1, tzinfo=timezone.utc)
m1 = MagicMock(id=10, project_id=1, updated_at=written)
m1.to_dict = MagicMock(return_value={"id": 10, "title": "A"})
m2 = MagicMock(id=11, project_id=2)
m2 = MagicMock(id=11, project_id=2, updated_at=written)
m2.to_dict = MagicMock(return_value={"id": 11, "title": "B"})
step_closed = datetime(2026, 9, 14, tzinfo=timezone.utc)
opened = [0]
rows = [[m1, m2], [(10, "done", 2), (10, "todo", 1), (11, "cancelled", 1)]]
rows = [[m1, m2], [
(10, "done", 2, step_closed),
(10, "todo", 1, datetime(2026, 9, 2, tzinfo=timezone.utc)),
(11, "cancelled", 1, datetime(2026, 8, 1, tzinfo=timezone.utc)),
]]
with patch.object(svc, "async_session", _session_factory(opened, rows)):
out = await svc.get_project_milestone_summaries(1, [1, 2])
assert opened[0] == 1
# Touched is the later of the milestone's own edit and its newest step
# update (#4045): steps closing today count, an old step doesn't pull it back.
assert out[1][0]["last_touched_at"] == step_closed.isoformat()
assert out[2][0]["last_touched_at"] == written.isoformat()
assert out[1][0]["completed"] == 2 and out[1][0]["total"] == 3
# Cancelled is excluded from the denominator, so a milestone whose only
# task was cancelled reads as complete rather than stalled at 0%.