Bug fix:
- ChatView.vue onMounted now skips fetchConversation when the conversation
is already loaded in the store (same guard that the convId watcher uses).
This prevents duplicate assistant messages when navigating from the
dashboard inline chat to /chat/:id after streaming completes.
Generation timing:
- logging.py: add log_generation() — persists per-generation timing
breakdown to app_logs (category=usage, action=generation) including
model, total_ms, intent_ms, ttft_ms, generation_ms, and per-tool timings.
Queryable via existing admin log viewer.
- generation_task.py: collect wall-clock timestamps at every pipeline stage:
intent classification, per-tool execution (both intent-routed and native),
time-to-first-token (measured from generation start to first content chunk),
LLM streaming round duration. Logs via log_generation() and includes timing
in the SSE 'done' event payload.
- types/chat.ts: add GenerationTiming interface; add optional timing field
to Message.
- chat.ts: capture timing from done event and attach to assistant message.
- ChatMessage.vue: show timing footer on assistant messages with breakdown:
"⏱ 4.2s total · first token 0.8s · analyzed 0.3s · created event 0.4s
· generated 3.5s". Visible this session; persisted to app_logs for
cross-session benchmarking.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>