fix(projects): batch the summary queries — the fan-out was exhausting the pool
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 43s
CI & Build / integration (push) Successful in 2m33s
CI & Build / Python tests (push) Successful in 3m1s
CI & Build / Build & push image (push) Successful in 44s

Reported live: Projects and Snippets showed skeletons that never resolved,
/knowledge worked intermittently. The logs named it exactly:

    QueuePool limit of size 5 overflow 10 reached, connection timed out, 30.00
    GET /api/settings  500  30584.0ms
    GET /api/projects  200  30882.9ms

/api/projects was not hanging — it was waiting out the 30-second checkout
timeout and then returning 200 with summaries silently missing, because
_attach swallowed the TimeoutError. Nobody waits 31 seconds, so it read as a
hang.

THE SHAPE: routes/projects.py ran asyncio.gather over every project. Each
_attach called get_project_summary, which opened its own session for three
queries and then called get_project_milestone_summary — which opened one more
session PER MILESTONE. So 25 projects asked for roughly 250 concurrent
checkouts against a pool of 15 (SQLAlchemy's default 5 + 10 overflow).

That is why unrelated routes failed too. Snippets and /knowledge were never
broken; they queued behind the burst and inherited its timeout. /api/settings
returning 500 while /api/projects returned 200 is the same cause wearing two
faces.

The comment above the gather said "one backend pass instead of N+1 frontend
calls". It did remove the N+1 from the network — and recreated it against the
connection pool, where it is worse, because the browser had at least been
serialising those calls.

Now: get_project_summaries() does all projects in four queries and one session,
and get_project_milestone_summaries() does all milestones in two. Two sessions
total for the whole page, independent of how many projects exist.

The progress calculation is extracted to _progress_from_counts and shared by
both the batch and single paths, so the cancelled-exclusion rule cannot drift
into two versions that disagree about whether a milestone is finished.

Tests assert the SESSION COUNT, not just the values. An implementation that
returned identical output while opening a session per project would pass a
correctness test and reproduce the outage.

Deliberately NOT done: raising pool_size. It would move the cliff rather than
remove it, and this endpoint now needs two connections regardless of scale.

Closes #2384.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
2026-08-02 19:09:36 -04:00
co-authored by Claude Opus 5
parent 5f8b824523
commit be3a0ffaf9
4 changed files with 304 additions and 28 deletions
+80
View File
@@ -127,6 +127,86 @@ async def delete_project(user_id: int, project_id: int) -> bool:
return True
async def get_project_summaries(
user_id: int, project_ids: list[int]
) -> dict[int, dict]:
"""Summaries for MANY projects — four queries and one session, total.
Replaces an `asyncio.gather` over the per-project version below, which was
a nested fan-out: each project opened its own session for three queries,
then called the milestone summary, which opened one more per milestone. For
25 projects that asked for roughly 250 pooled connections at once against a
pool of 15 (SQLAlchemy's default 5 + 10 overflow), so most of them sat out
the 30-second checkout timeout and everything else on the instance queued
behind them — including unrelated routes, which is why /api/settings
returned 500 while /api/projects took 30.9s (#2384).
The comment it replaced said "one backend pass instead of N+1 frontend
calls". It did remove the N+1 from the network — and recreated it against
the connection pool, where it is worse: the browser had at least been
serialising those calls.
"""
if not project_ids:
return {}
async with async_session() as session:
task_rows = await session.execute(
select(Note.project_id, Note.status, func.count(Note.id))
.where(
Note.user_id == user_id,
Note.project_id.in_(project_ids),
Note.status.isnot(None),
Note.deleted_at.is_(None),
)
.group_by(Note.project_id, Note.status)
)
task_counts: dict[int, dict[str, int]] = {}
for project_id, status, count in task_rows.fetchall():
task_counts.setdefault(project_id, {})[status] = count
note_rows = await session.execute(
select(Note.project_id, func.count(Note.id))
.where(
Note.user_id == user_id,
Note.project_id.in_(project_ids),
Note.status.is_(None),
Note.deleted_at.is_(None),
)
.group_by(Note.project_id)
)
note_counts = {pid: count for pid, count in note_rows.fetchall()}
# Deliberately NOT filtered by deleted_at, matching the per-project
# version: "last activity" includes trashing something.
activity_rows = await session.execute(
select(Note.project_id, func.max(Note.updated_at))
.where(Note.user_id == user_id, Note.project_id.in_(project_ids))
.group_by(Note.project_id)
)
last_activity = {pid: ts for pid, ts in activity_rows.fetchall()}
from scribe.services.milestones import get_project_milestone_summaries
milestones = await get_project_milestone_summaries(user_id, project_ids)
return {
pid: {
# All three lifecycle keys present so consumers can sum without
# `?? 0` guards — the frontend declares them required, and
# `undefined + N` renders as NaN.
"task_counts": {
"todo": 0, "in_progress": 0, "done": 0,
**task_counts.get(pid, {}),
},
"note_count": note_counts.get(pid, 0),
"last_activity": (
last_activity[pid].isoformat() if last_activity.get(pid) else None
),
"milestone_summary": milestones.get(pid, []),
}
for pid in project_ids
}
async def get_project_summary(user_id: int, project_id: int) -> dict:
"""Return task counts by status, note count, and last activity."""
async with async_session() as session: