fix(projects): batch the summary queries — the fan-out was exhausting the pool
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 43s
CI & Build / integration (push) Successful in 2m33s
CI & Build / Python tests (push) Successful in 3m1s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 43s
CI & Build / integration (push) Successful in 2m33s
CI & Build / Python tests (push) Successful in 3m1s
CI & Build / Build & push image (push) Successful in 44s
Reported live: Projects and Snippets showed skeletons that never resolved,
/knowledge worked intermittently. The logs named it exactly:
QueuePool limit of size 5 overflow 10 reached, connection timed out, 30.00
GET /api/settings 500 30584.0ms
GET /api/projects 200 30882.9ms
/api/projects was not hanging — it was waiting out the 30-second checkout
timeout and then returning 200 with summaries silently missing, because
_attach swallowed the TimeoutError. Nobody waits 31 seconds, so it read as a
hang.
THE SHAPE: routes/projects.py ran asyncio.gather over every project. Each
_attach called get_project_summary, which opened its own session for three
queries and then called get_project_milestone_summary — which opened one more
session PER MILESTONE. So 25 projects asked for roughly 250 concurrent
checkouts against a pool of 15 (SQLAlchemy's default 5 + 10 overflow).
That is why unrelated routes failed too. Snippets and /knowledge were never
broken; they queued behind the burst and inherited its timeout. /api/settings
returning 500 while /api/projects returned 200 is the same cause wearing two
faces.
The comment above the gather said "one backend pass instead of N+1 frontend
calls". It did remove the N+1 from the network — and recreated it against the
connection pool, where it is worse, because the browser had at least been
serialising those calls.
Now: get_project_summaries() does all projects in four queries and one session,
and get_project_milestone_summaries() does all milestones in two. Two sessions
total for the whole page, independent of how many projects exist.
The progress calculation is extracted to _progress_from_counts and shared by
both the batch and single paths, so the cancelled-exclusion rule cannot drift
into two versions that disagree about whether a milestone is finished.
Tests assert the SESSION COUNT, not just the values. An implementation that
returned identical output while opening a session per project would pass a
correctness test and reproduce the outage.
Deliberately NOT done: raising pool_size. It would move the cliff rather than
remove it, and this endpoint now needs two connections regardless of scale.
Closes #2384.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
@@ -127,6 +127,86 @@ async def delete_project(user_id: int, project_id: int) -> bool:
|
||||
return True
|
||||
|
||||
|
||||
async def get_project_summaries(
|
||||
user_id: int, project_ids: list[int]
|
||||
) -> dict[int, dict]:
|
||||
"""Summaries for MANY projects — four queries and one session, total.
|
||||
|
||||
Replaces an `asyncio.gather` over the per-project version below, which was
|
||||
a nested fan-out: each project opened its own session for three queries,
|
||||
then called the milestone summary, which opened one more per milestone. For
|
||||
25 projects that asked for roughly 250 pooled connections at once against a
|
||||
pool of 15 (SQLAlchemy's default 5 + 10 overflow), so most of them sat out
|
||||
the 30-second checkout timeout and everything else on the instance queued
|
||||
behind them — including unrelated routes, which is why /api/settings
|
||||
returned 500 while /api/projects took 30.9s (#2384).
|
||||
|
||||
The comment it replaced said "one backend pass instead of N+1 frontend
|
||||
calls". It did remove the N+1 from the network — and recreated it against
|
||||
the connection pool, where it is worse: the browser had at least been
|
||||
serialising those calls.
|
||||
"""
|
||||
if not project_ids:
|
||||
return {}
|
||||
|
||||
async with async_session() as session:
|
||||
task_rows = await session.execute(
|
||||
select(Note.project_id, Note.status, func.count(Note.id))
|
||||
.where(
|
||||
Note.user_id == user_id,
|
||||
Note.project_id.in_(project_ids),
|
||||
Note.status.isnot(None),
|
||||
Note.deleted_at.is_(None),
|
||||
)
|
||||
.group_by(Note.project_id, Note.status)
|
||||
)
|
||||
task_counts: dict[int, dict[str, int]] = {}
|
||||
for project_id, status, count in task_rows.fetchall():
|
||||
task_counts.setdefault(project_id, {})[status] = count
|
||||
|
||||
note_rows = await session.execute(
|
||||
select(Note.project_id, func.count(Note.id))
|
||||
.where(
|
||||
Note.user_id == user_id,
|
||||
Note.project_id.in_(project_ids),
|
||||
Note.status.is_(None),
|
||||
Note.deleted_at.is_(None),
|
||||
)
|
||||
.group_by(Note.project_id)
|
||||
)
|
||||
note_counts = {pid: count for pid, count in note_rows.fetchall()}
|
||||
|
||||
# Deliberately NOT filtered by deleted_at, matching the per-project
|
||||
# version: "last activity" includes trashing something.
|
||||
activity_rows = await session.execute(
|
||||
select(Note.project_id, func.max(Note.updated_at))
|
||||
.where(Note.user_id == user_id, Note.project_id.in_(project_ids))
|
||||
.group_by(Note.project_id)
|
||||
)
|
||||
last_activity = {pid: ts for pid, ts in activity_rows.fetchall()}
|
||||
|
||||
from scribe.services.milestones import get_project_milestone_summaries
|
||||
milestones = await get_project_milestone_summaries(user_id, project_ids)
|
||||
|
||||
return {
|
||||
pid: {
|
||||
# All three lifecycle keys present so consumers can sum without
|
||||
# `?? 0` guards — the frontend declares them required, and
|
||||
# `undefined + N` renders as NaN.
|
||||
"task_counts": {
|
||||
"todo": 0, "in_progress": 0, "done": 0,
|
||||
**task_counts.get(pid, {}),
|
||||
},
|
||||
"note_count": note_counts.get(pid, 0),
|
||||
"last_activity": (
|
||||
last_activity[pid].isoformat() if last_activity.get(pid) else None
|
||||
),
|
||||
"milestone_summary": milestones.get(pid, []),
|
||||
}
|
||||
for pid in project_ids
|
||||
}
|
||||
|
||||
|
||||
async def get_project_summary(user_id: int, project_id: int) -> dict:
|
||||
"""Return task counts by status, note count, and last activity."""
|
||||
async with async_session() as session:
|
||||
|
||||
Reference in New Issue
Block a user