- build.yml: clarify that runner needs Docker socket; SSH deploy uses
DEPLOY_PATH secret so app host path is configurable; --no-deps so
db/ollama containers are left untouched during deploy
- infra/runner-swarm.yml: act_runner as a Docker service on the Forgejo
host — mounts host socket for buildx, writes act_runner config via
one-shot init container, connects to fabled_backend network
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- .forgejo/workflows/ci.yml: fast checks on every push (TS typecheck,
Python lint via ruff, pytest) — no Docker build, ~30s with cache
- .forgejo/workflows/build.yml: Docker build with registry-side layer
caching + SSH deploy on main branch pushes only
- pyproject.toml: add ruff to dev deps, configure pytest and ruff rules
- tests/: conftest.py + first unit tests for _safe_filename (no DB needed)
- Makefile: add lint, fmt, typecheck, test, check targets for local use
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Panel state is saved to localStorage as workspace_panels_{projectId} on
every toggle and restored on mount. Each project remembers its own layout
independently. Fallback guard ensures at least one panel is always open
after restore.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Conversations are bucketed into: Today, Yesterday, This week, This month,
and older entries sub-grouped by calendar month (e.g. "February 2026").
Group labels are sticky so they stay visible while scrolling. Older
buckets use month+year sub-grouping rather than a single "Older" catch-all.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Completed (success) tool calls render collapsed: label + inline summary + ▶ chevron
- Clicking the header expands the detail section (result list, links, tags)
- Error cards start expanded so failures are immediately visible
- Running (in-progress) cards start expanded for live progress visibility
- Auto-collapses when a running call completes during streaming
- Suggested-tag pills remain accessible in both collapsed (via separate row) and expanded states
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
TiptapEditor.vue (central — applies to all three editors):
- Escape inside editor blurs it and emits 'escape' event to parent
TagInput.vue (central — applies everywhere tags are used):
- ArrowUp/ArrowDown navigate autocomplete suggestion list with visual highlight
- Enter confirms the keyboard-selected suggestion instead of typed text
NoteEditorView, TaskEditorView, WorkspaceNoteEditor (per-editor wiring):
- @escape on TiptapEditor returns focus to the title input (ref="titleRef")
- Ctrl+E from title input or editor main column jumps focus into editor body
- Ctrl+S on title input saves (was already on editor area; now consistent)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add MarkdownToolbar to the workspace note editor (was missing entirely)
- Add [[ button to MarkdownToolbar so wikilinks are discoverable in all editors;
clicking inserts [[ which immediately triggers the WikilinkSuggestion dropdown
- Add link suggestions strip: polls /api/notes/link-suggestions 2.5s after edit,
shows unlinked note-title mentions as clickable chips to wrap in [[...]], plus
"All" button to apply everything at once — directly addresses the heading→wikilink
workflow (note title appearing as heading text gets detected and offered for linking)
- Add Ctrl+S keyboard shortcut on title input and editor area to save
- Replace ✕ delete icon on note list items with trash can SVG (consistency with task panel)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- WorkspaceTaskPanel: add milestone <select> in task detail (PATCH /api/notes/:id),
replace delete ✕ with trash can SVG icon
- WorkspaceView: scroll to bottom when streaming ends so final message is visible
without a page refresh
- ToolCallCard: fix search_notes result count (was reading data.total; tool returns
data.count), so results no longer show "0 found"
- push.py: switch from deprecated WebPusher().send(vapid_private_key=...) to
webpush() function (pywebpush 2.x API compatibility)
- app.py: downgrade /api/health, /api/chat/status, and static asset requests from
INFO to DEBUG in after_request logger to reduce log noise at default log level
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- New /graph route with D3 force simulation (GraphView.vue)
- Tag nodes as first-class graph nodes (string IDs "tag:name") — clicking
navigates to /notes?tags=name; tags shown by default
- Invisible project hub nodes attract project members into clusters
- Physics panel with live sliders: repulsion, link distance, link strength,
project pull, gravity (forceX/forceY, not forceCenter)
- Wikilink edges retain directed arrowheads; tag edges are thin, no arrowhead
- Graph nav link in AppHeader; `g` shortcut in App.vue
- Backend: build_note_graph() emits tag nodes + note→tag edges instead of
O(n²) note→note shared-tag mesh
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add session_version to users table. Sessions now carry session_version
alongside user_id; @login_required rejects any session where the version
doesn't match the DB value.
- migration 0024: session_version INTEGER NOT NULL DEFAULT 1
- models/user.py: session_version Mapped[int] column
- auth.py: version mismatch → session.clear() + 401
- services/auth.py: change_password and reset_password_with_token both
increment session_version, evicting all other live sessions
- routes/auth.py: login/register/oauth_callback store session_version;
update_password rehydrates current session with new version so the
user who changed their password stays logged in
Existing sessions will require one re-login after the migration runs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
NoteEditorView: two-column sidebar layout (project/milestone/tags/assist
always visible), removed assist toggle button, InlineAssistPanel removed.
Writing assist: whole_doc mode rewrites entire document; DiffView.vue
replaces editor during review showing full-document diff. Scope dropdown
in sidebar switches between whole-document and section modes.
Persistent drafts: migration 0022 adds note_drafts (UNIQUE per note+user)
and note_versions (max 20, auto-pruned) tables. Draft saved after generation
completes, restored on editor mount, cleared on accept/reject. Version
snapshot created automatically whenever note body changes on save.
HistoryPanel.vue: version list + DiffView modal, restore button writes
body back to editor.
Config: OLLAMA_NUM_CTX default raised to 65536; assist num_predict now
tracks Config.OLLAMA_NUM_CTX instead of a hardcoded 4096.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stream_chat: add think=False parameter passed through to Ollama payload.
qwen3 models have thinking enabled by default; without this flag the model
spends minutes generating internal thinking tokens that stream_chat silently
discards, leaving the frontend spinner blank until the SSE connection times
out and the widget disappears.
- create_assist_buffer: orphan (overwrite) a still-running buffer instead of
raising. The old asyncio task holds a direct reference and completes
harmlessly against the stale buffer. New requests always win.
- assist_route: remove the 409 guard that blocked new requests when a previous
generation got stuck. create_assist_buffer now handles this transparently.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
stores/chat.ts:
- Replace 7 global stream refs with convStreams: Record<number, ConvStreamState>
- Each conversation tracks its own stream state; 7 names become computed getters
- Add isStreamingConv(id) public helper
- Increment message_count +2 immediately after POST 202 (not at done) so the
cleanup watcher never deletes a mid-generation conversation
- SSE reconnect retries now run for background conversations regardless of
which conv is currently viewed
- error toasts only shown when the erroring conv is currently viewed
- deleteConversation cleans up orphaned stream state
ChatView.vue:
- Remove invalid write to store.lastContextMeta (now a read-only computed)
- Add !store.isStreamingConv(newId) guard to watch(convId) fetch so navigating
back to a generating conversation shows accumulated content without stale refetch
- Add startNewConversation() helper; opening /chat with no ID auto-creates a
new conversation and redirects to it (on mount and on navigation)
- Deleting a conversation also auto-creates a new one
Also committing previous-session changes (keyboard nav + task routing):
- App.vue: e shortcut only on note-view (not task-view, which is removed)
- TaskCard.vue: route directly to /tasks/:id (viewer), remove View buttons
- router/index.ts: remove task-view route; /tasks/:id now maps to TaskEditorView
- TasksListView.vue: Enter key pushes to /tasks/:id
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- App.vue: e → edit on viewer pages; / → focus search (custom event); c → focus
chat on home or navigate to /chat; update shortcuts panel with Lists + Chat sections
- theme.css: add a:focus-visible to focus-ring rules (router-link cards now show
visible keyboard focus outline)
- SearchBar.vue: expose focus() via defineExpose for cross-component focus dispatch
- NotesListView / TasksListView: j/k vim-style navigation with kb-active-item
highlight; listen for shortcut:focus-search; cleanup on unmount
- HomeView.vue: listen for shortcut:focus-chat, call chatInputRef.focus()
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
First Escape blurs the active input (reaching neutral state where
shortcuts work), second Escape navigates home. Allows full keyboard
navigation without touching the mouse to leave a focused field.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Prevents content from escaping max-width boundaries across notes,
tasks, projects list, and project detail views.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Clicking a task card now goes directly to the edit view since that's
the primary use case. The hover button (compact) and button (full card)
are now labeled "View" and link to the detail/viewer view instead.
Sub-task links in TaskViewerView also route to edit.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Projects now appear at the top of the tasks column directly above
the task sections, laid out in a fixed 3-column grid. The previous
full-width widget below the main grid has been removed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Keyboard shortcuts (App.vue):
- g+h/n/t/p/c: navigate to home/notes/tasks/projects/chat
- n / t: new note / new task (when not in an input field)
- Escape: go home (when not typing)
- Shortcuts panel updated to document all working shortcuts
NoteViewerView:
- Context breadcrumb: parent note link (↑), project badge, milestone badge
- Fetches project and parent note titles in parallel on load
- Back button label improved to "← Notes"
TaskViewerView:
- Context breadcrumb: parent task link (uses parent_title from API), project badge, milestone badge
- Sub-tasks section with inline progress bar and status dots (clickable to advance)
- Sub-tasks loaded from /api/notes?parent_id=X&type=task
- Note type extended with optional parent_title field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add min-width: 0 to .col-tasks and .col-notes so CSS grid items
shrink to their track size. Change notes-mini-grid from repeat(2, 1fr)
to auto-fill minmax(220px, 1fr) so cards wrap instead of overflowing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
NoteCard: add compact prop — single-row title/tags/timestamp, no body preview
TaskCard: add compact prop — single-row with status dot, priority badge,
title, optional project breadcrumb, due date; edit button appears on hover.
Add projectTitle prop for cross-view context.
NotesListView: auto-fill CSS grid layout (2–4 columns); compact list toggle;
view mode persisted to localStorage.
TasksListView: compact rows throughout; group-by-project toggle with
collapsible sections, task counts, and "Open project" links; fetches
project list for group labels and task breadcrumbs; limit raised to 100
in grouped mode; view mode persisted to localStorage.
HomeView dashboard: task sections use compact rows (project breadcrumb shown);
notes column uses 2-per-row mini-grid; new "Active Projects" widget below
main grid shows milestone progress bars for up to 6 active projects.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
On first boot, ensure_vapid_keys() generates a fresh EC P-256 key pair
and saves it to /data/vapid_keys.json (inside the app_data volume) so
it survives container restarts. Subsequent boots load from that file.
VAPID_PRIVATE_KEY / VAPID_PUBLIC_KEY env vars still take precedence for
deployments that prefer to manage keys externally.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Startup no longer auto-warms Config.OLLAMA_MODEL. Instead it queries
all distinct default_model values from user settings, cross-references
with Ollama's installed models, and warms only the intersection.
Models that users have selected but not yet installed are skipped with
an info log — they are never auto-pulled. The embedding model pull
behaviour is unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Updates the startup warm-up and fallback default to match the
currently preferred model. Instances without a user setting will
also default to qwen3:14b.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
conv.model was stored at conversation creation time and short-circuited
the get_setting() call, so changing default_model in Settings had no
effect on existing conversations. Now the user's default_model setting
is always the source of truth for generation and summarization.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
tools.py:
- Default tag_mode changed from 'replace' to 'add' — existing tags are
preserved unless the user explicitly requests replacement
- create_task, create_note, update_note, create_milestone: project and
milestone lookup now uses get_by_title with fuzzy fallback; returns a
clear error (with a hint to use list_projects/list_milestones) instead
of silently creating duplicate projects or milestones via get_or_create
- create_task: project/milestone resolved before note creation so a bad
project name fails fast without leaving an orphaned task behind
- update_note return value now includes item_type, tags, project_id, and
an 'updated' summary string so the LLM can confirm what was modified
- research_topic stub now logs an error and returns failure instead of
silently returning success when reached via wrong code path
ProjectView.vue:
- Milestone headers now show edit (✎) and delete (✕) action buttons on hover
- Edit triggers inline rename input; Enter/blur commits, Escape cancels
- Delete opens a confirmation modal clarifying tasks are unlinked not deleted
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The separate intent model (OLLAMA_INTENT_MODEL / qwen2.5:7b) is removed
from every part of the system. All classification now uses the primary model.
Changes:
- config.py: remove OLLAMA_INTENT_MODEL
- intent.py: remove classify_intent() and all supporting infrastructure
(_SYSTEM_PROMPT_TEMPLATE, _RESEARCH_PREFIX, _PRIOR_WORK_REFS); file now
only contains the quick-capture classifier
- quick_capture.py: classify_capture_intent() now called with Config.OLLAMA_MODEL
- generation_task.py: remove intent_model_setting DB lookup and get_setting import;
history summarization and research pipeline use the primary model directly
- research.py: remove intent_model parameter from run_research_pipeline() and
_generate_sub_queries(); both use the model param throughout
- routes/settings.py: remove intent_model from model-key validation and response
- app.py: remove intent model pre-warming at startup
- SettingsView.vue: remove Intent Model selector and related refs/state
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The intent classifier (Phase 21) is removed from the main chat generation
path. The main model now handles all tool routing natively via Ollama's
structured tool-calling API, eliminating misidentification issues caused
by the small intent model.
Changes:
- generation_task.py: remove classify_intent call, intent_task, _WRITE_TOOLS,
_TOOL_ACTIONS, _INTENT_TRIGGER_WORDS, _should_skip_intent(), and the entire
round-0 intent-first + write-tool confirmation block (~315 lines removed)
- research_topic tool calls are now handled inline in the streaming loop:
runs run_research_pipeline, streams synthesis to buf, then breaks the round
loop (research is still the full response, no model follow-up)
- config.py: raise OLLAMA_NUM_CTX default from 8192 to 16384
The quick-capture dedicated classifier (classify_capture_intent) is unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add _process_note() — a second LLM pass using the main model that
transforms raw capture text into a well-formed note with a genuine
summary title and formatted body. Replaces the previous behaviour of
using the captured text verbatim as both title and body.
The processing prompt instructs the model to:
- Generate a 3-8 word summary title (never a verbatim copy)
- Format the body appropriately: bullet lists for items, clean prose
for stream-of-thought, organised paragraphs for raw notes/fragments
- Preserve all original information without inventing new facts
The enrichment pass runs for both the intent-classified create_note
path and the fallback path. On LLM/parse failure it degrades safely
to the old verbatim behaviour.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stream_chat and stream_chat_with_tools: remove read=300s per-chunk
timeout, replace with read=None. In httpx streaming mode, the read
timeout applies per-chunk — if Ollama pauses >300s while processing
a large input context before the first token, it raises ReadTimeout,
killing generation and leaving the assistant message as an empty stub.
With read=None the stream is unbounded; connect=30s still guards the
initial connection.
- chat_status_route: increase Ollama status check timeout 5s → 10s.
When Ollama is busy processing a large prompt it can be slow to
respond to /api/tags, causing the status indicator to briefly flip to
"offline" even though generation is running normally.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove all 6 CalDAV todo tools (create/list/update/complete/delete/search_todos)
from tools.py definitions, imports, execute_tool branches, intent routing rules,
generation_task labels/actions, and llm.py system prompt hints. CalDAV event
tools remain. Todo functions still exist in caldav.py but are no longer exposed.
- Quick-capture now uses a dedicated classify_capture_intent() with a focused
_CAPTURE_SYSTEM_PROMPT that always routes to a tool (never null). Tool set
expanded: create_note/task/event + update_note + research_topic.
- research_topic in quick-capture calls run_research_pipeline() directly (no SSE
buffer). run_research_pipeline() now accepts buf=None; all buf.append_event
calls are guarded so status events are skipped when no buffer is provided.
- Fallback note now always sets body=text (was empty for texts ≤80 chars).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Quart's send_file uses cache_timeout= not max_age=. The TypeError on
every /api/images/<id> request caused a 500, which the browser rendered
as alt text.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- tools.py: search_images result now includes 'embed' (ready-to-use
markdown image syntax) and 'citation' fields instead of raw 'local_url';
adds 'instructions' field so the model knows to render them verbatim
- llm.py: system prompt now explicitly tells the model to embed images
using the 'embed' field rather than describing or listing URLs
- markdown.ts: explicitly allow src/alt in PURIFY_OPTS_FULL so img tags
are never stripped by DOMPurify
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Single named volume app_data covering the entire /data directory so
future persistent storage (uploads, exports, etc.) doesn't need
additional volume entries.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Images found via SearXNG are fetched server-side, stored on disk, and
served from /api/images/<id> — the user's browser never contacts the
original image host. Original URLs are preserved for citation.
New files:
- alembic/versions/0016_add_image_cache.py — image_cache table
- src/fabledassistant/models/image_cache.py — SQLAlchemy model
- src/fabledassistant/services/images.py — fetch/store/serve logic
- src/fabledassistant/routes/images.py — GET /api/images/<id>
Modified:
- config.py: IMAGE_CACHE_DIR (/data/images), IMAGE_MAX_BYTES (5 MB)
- research.py: _search_searxng_images() — SearXNG categories=images
- tools.py: _IMAGE_TOOLS def + search_images branch in execute_tool
- intent.py: search_images routing rule (explicit visual language only)
- app.py: register images_bp
- docker-compose.yml: image_cache named volume mounted at /data/images
- ToolCallCard.vue: "image_search" label
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Raise similarity threshold 0.30 → 0.45: only genuinely relevant notes
shown; loosely-related notes no longer pad the sidebar
- Increase max suggested notes 3 → 8 (zero added compute — threshold is
the real gate; the embedding call is fixed regardless of limit)
- semantic_search_notes now returns list[tuple[float, Note]] instead of
list[Note] so scores propagate through context_meta to the frontend
- Keyword fallback notes carry score=null (no cosine similarity available)
- ChatView sidebar shows % badge on each suggested note:
green ≥75%, amber 60–74%, muted <60%
Hovering reveals the raw score in a tooltip
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Single POST that classifies natural-language text and creates the
appropriate item (note, task, event, or todo) in one synchronous
request — no SSE, no conversation context needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The same empty-string model guard added to chat.py was missing from the
notes assist route. If default_model was stored as "" in the DB, the
assist route would pass "" to Ollama which responds with 400, surfaced
to the user as a "400" error in the assist panel.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The model occasionally passes title as a list (e.g. when asked to create
multiple notes at once), causing asyncpg DataError on the INSERT. Return a
clear error result so the model sees the problem and retries with individual
calls instead of crashing with an unhandled exception.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously the paperclip button created a one-shot attachment that was
only included for the single message it was sent with, then discarded.
Now selecting a note via the picker calls includeNote() directly, so it
appears in the "In Context" sidebar and stays for the entire conversation
— consistent with clicking "+" on a suggested note.
Removed attachedNote ref, removeAttachedNote(), the pinned-note sidebar
block, and the contextNoteId/contextNoteTitle sendMessage arguments that
are no longer needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add a Python fast-path regex (_PRIOR_WORK_REFS) in classify_intent that
detects phrases like "research you did", "note you made", "using your
research", "based on the research" etc. and returns no-tool immediately —
saving the 19s intent LLM call and correctly letting the main model answer
using search_notes/context rather than firing off a web search.
Also tighten the intent prompt rules for search_web: explicitly prohibit
using it for creative/brainstorming requests or when the user references
existing notes, and add a rule that creative/ideation questions ("think of",
"come up with", "brainstorm") always route to null (chat).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When a user selected the 'Default' option in Settings, the dropdown
sent an empty string "" to the backend. The route saved it as a DB row,
which caused get_setting() to return "" instead of falling back to
Config defaults. The chat status endpoint then tried to match "" against
installed model names — always failing — resulting in model: "not_found"
and a permanently failing readiness indicator.
services/settings.py:
- Add delete_setting() helper: removes a setting row so get_setting()
correctly falls back to its hardcoded default argument
routes/settings.py:
- Import delete_setting
- When default_model or intent_model are saved as empty string, delete
the DB row instead of storing "" — cleanly restores Config fallback
routes/chat.py:
- chat_status_route: add explicit `or Config.OLLAMA_MODEL` guard for
any existing "" rows written before this fix (migration safety net)
- send_message and summarize routes: same guard on model resolution
so empty settings never cause silent generation failures
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
config.py:
- Default OLLAMA_INTENT_MODEL: qwen2.5:1.5b → qwen2.5:7b
- Startup will auto-pull and warm the new model on next container restart
intent.py:
- Replaced phrase-matching examples in search_web and research_topic rules
with semantic descriptions. The 7B model doesn't need example phrases to
understand intent — it can reason from the tool's purpose. Removes implied
usage patterns that caused misclassifications on conversational phrasing
(e.g. "I've been thinking about buying shirts, can you research this?").
- research_topic rule now explicitly covers any subject regardless of phrasing,
including shopping decisions, comparisons, how-things-work questions, etc.
- search_web rule clarified as "short summary, no note" vs research_topic's
"comprehensive written reference"
The 1.5B model required prescriptive phrase examples to route correctly; the
7B model has sufficient language understanding to classify from semantic intent.
Expected improvement: ~1-2s intent calls (vs 0.4-9s for the 1.5B model which
sometimes timed out or misclassified longer/conversational messages).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two fixes for the intent model failing to route 'Research: X' messages
to research_topic:
1. Fast-path in classify_intent: if the message matches ^Research:\s+.+
(the exact format the UI Research button always sends), skip the LLM
call entirely and return research_topic with high confidence. This is
100% reliable and saves an unnecessary model call for this pattern.
2. Expanded research_topic rule examples in the system prompt to include
"Research: X" prefix format, shopping-style queries ("research where
to buy X"), and clarification that the topic is everything after the
keyword — improves LLM routing for natural-language research requests
that don't match the previous narrow examples.
Root cause: qwen2.5:1.5b misclassified "Research: where to buy three-
quarter sleeve tee shirts" as general chat (shopping query phrasing
combined with the colon confused the small model).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When the intent model doesn't classify a research request (low confidence,
long message, etc.), the main model (qwen3) would correctly identify
research_topic itself and call it via the streaming tool loop. But
execute_tool("research_topic") only returns a dummy research_pending
placeholder, causing the model to see the result and retry — looping
up to MAX_TOOL_ROUNDS times.
Fix: filter research_topic out of stream_tools (the tool list given to
the main model via stream_chat_with_tools). research_topic is an
intent-only routing tool; the main model should never call it directly.
The full tools list (including research_topic) is still passed to
classify_intent so intent routing continues to work.
The _INTENT_ONLY_TOOLS frozenset makes this pattern explicit and
extensible for future intent-only tools.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
research.py:
- Parallelize all 5 SearXNG queries concurrently (200ms stagger via asyncio.gather)
- Parallelize all URL fetches in parallel (asyncio.gather) — up to 15 URLs at once
instead of sequential fetches; biggest performance win (was O(n) × 15s, now ~15s flat)
- _synthesize_note accepts buf: when provided uses stream_chat (num_ctx=16384,
num_predict=8192) to emit tokens into the chat buffer in real time so users see
the note being written; falls back to generate_completion when buf=None
- Added \n\n---\n\n separator before "Research complete!" to cleanly mark boundary
after streamed synthesis content
intent.py:
- classify_intent passes num_ctx=4096 to generate_completion — reduces VRAM pressure
and prefill time for the intent model call on every single request
generation_task.py:
- _INTENT_TRIGGER_WORDS frozenset (~50 action/object/date words) + _should_skip_intent()
skips intent classification for short messages (≤10 words) with no trigger words;
saves 400-800ms model call for conversational replies ("thanks", "okay", etc.)
- Added \n\n---\n\n separator before research "done" text in research_topic branch
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Ollama streams message.thinking tokens alongside message.content when
think=True — previously silently dropped. Now forwarded end-to-end.
Backend:
- llm.py: ChatChunk type gains "thinking" variant; stream_chat_with_tools
yields ChatChunk(type="thinking") for msg.thinking chunks before content
- generation_task.py: thinking chunks emit "thinking_chunk" SSE events
(not added to content_so_far — not persisted to DB)
Frontend:
- types/chat.ts: Message.thinking?: string (session-only, not from DB)
- stores/chat.ts: streamingThinking ref; thinking_chunk handler accumulates
chunks; on done, thinking carried into committed Message object then cleared
- ChatMessage.vue: collapsible <details class="thinking-block"> shown for
messages that have .thinking content (collapsed by default)
- ChatView.vue + ChatPanel.vue: live thinking block in streaming bubble —
open while only thinking is flowing, auto-collapses when content arrives;
typing indicator hidden while thinking is active
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Instead of relying solely on retry-on-500, poll /api/ps before starting
any LLM stream so the main model has time to fully load into VRAM.
- llm.py: add wait_for_model_loaded(model, timeout=90s) — polls /api/ps
every 2s, returns True when model appears in loaded list
- generation_task.py: launch model_load_task in parallel with build_context
and classify_intent (both use fast/small-model ops that don't need the
main model); after context is built, await the load task — shows
"Loading model..." status only if the user actually has to wait;
logs a warning and proceeds if 90s timeout elapses
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Tags are now a first-class field rather than being auto-extracted from
the note body. A new TagInput.vue chip component handles tag entry in
both editor views with autocomplete, Enter/comma/backspace UX, and
space-to-hyphen sanitization.
Backend:
- routes/notes.py: create reads tags from JSON; update accepts explicit
tags (omit = keep existing); append_tag writes to tags array with
dedup; suggest-tags accepts current_tags filter; remove extract_tags
- routes/tasks.py: same — explicit tags on create/update; remove extract_tags
- services/tag_suggestions.py: current_tags param replaces body extraction
- services/tools.py: create_note tool schema adds tags param; executor passes it
- services/llm.py: system prompt tells LLM to use tags param, not embed #tag in body
Frontend:
- components/TagInput.vue: new chip-based tag input (autocomplete, keyboard UX)
- NoteEditorView.vue / TaskEditorView.vue: tags ref loaded from note.tags;
TagInput placed between title and body; save/autosave include tags; suggest
now adds chips; fetchTagSuggestions passes current_tags; dirty tracks tags
- TiptapEditor.vue: remove fetchTags prop and TagSuggestion extension;
keep TagDecoration for legacy inline #tag highlighting
No DB migration needed — tags column already correct.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Tags with spaces (e.g. #science fiction) were breaking extraction because
TAG_RE only matched word characters — it would stop at the space and extract
#science instead of #science-fiction.
- TAG_RE (backend + frontend): add hyphens to character class so #science-fiction
is recognized as a single tag: [\w][\w-]* per segment
- System prompt: instruct LLM to use hyphens in multi-word tags, never spaces
- tag_suggestions.py: update prompt example + sanitize output by replacing
spaces with hyphens as a safety net regardless of LLM output
- append-tag route: sanitize incoming tag (spaces → hyphens) before appending
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds _stream_with_retry() async generator (wraps stream_chat_with_tools
with up to 2 retries on Ollama 500, 3s/6s delay). Previously only the
optimistic round 0 _fill_queue had retry logic. Two paths were still
bare: the declined-write-tool fresh stream, and the round 1+ stream.
Round 1 500s occur when tag suggestions (fire-and-forget inside
execute_tool) race the follow-up stream to the same model. The retry
waits for tag suggestions to complete before succeeding.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ensure_model only downloads a model if missing; it does not load it
into VRAM. The frontend warm call only covers the chat model (and only
after a user opens the dashboard). This left qwen2.5:1.5b (intent) cold,
causing simultaneous cold-load 500s when the first chat arrived.
Now both Config.OLLAMA_MODEL and Config.OLLAMA_INTENT_MODEL are warmed
at startup (after ensuring they're installed) via a fire-and-forget
/api/generate call with keep_alive=30m. The embedding model is still
pulled but not warmed (it's loaded on demand during backfill).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
With optimistic streaming, intent (qwen2.5:1.5b) and the main stream
(qwen3:latest) start concurrently. When both models are cold-loading,
Ollama returns 500 for both simultaneously. The intent 500 was already
handled silently in classify_intent; the stream 500 now retries up to
2 times (3s then 6s delay) before propagating as an error. 500s only
occur on the first cold-load pair — subsequent requests hit warm models.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Start the main LLM stream immediately after build_context finishes instead
of waiting for intent classification to complete. Race the two concurrently:
- Intent wins before first token → cancel stream, execute tool (tool path
unchanged: confirmation, acknowledgment, multi-round loop all preserved)
- First token wins → discard intent, user sees output immediately
For pure chat messages (no tool needed, the common case) this eliminates
the full intent classification RTT from TTFT. For tool calls, intent
typically wins the race since it finishes before the main model produces
its first token, so tool behaviour is unchanged in practice.
Also extracts _drain_queue() as a module-level async generator helper.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Default OLLAMA_INTENT_MODEL to qwen2.5:1.5b in code instead of empty
- Add GET /api/settings/models endpoint returning installed models and defaults
- Validate intent_model against installed models on save (same as default_model)
- Replace intent model text input with a dropdown of installed models
- Add chat model dropdown to Assistant settings section
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replaces the hardcoded num_ctx=32768 KV cache allocation with a
configurable env var defaulting to 8192. This significantly reduces
VRAM pressure when multiple services share the GPU.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace pure white backgrounds with off-white indigo-tinted values to
reduce glare, and switch the primary/accent color from Google blue to
the app's brand indigo (#6366f1) for consistency with the header and
email templates.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
favicon.svg:
- Light mode: replace near-black fill (#2d3748) with indigo brand color
(#6366f1 fill, #4f46e5 stroke, #a5b4fc page lines) — distinctive and
high-contrast without the dark/black appearance
- Dark mode unchanged
email.py:
- Add _EMAIL_LOGO_SVG: inline SVG with white palette for rendering on
the indigo header (white book, lavender lines, gold sparkle)
- Add _email_html(title, body): shared template wrapper — gray outer
background, white card with border-radius, indigo header with logo +
app name, content area, footer
notifications.py:
- Import and use _email_html for all six email functions: security alert,
password reset, password reset success, invitation, task reminder,
test email
- Clean up all inline HTML to match the new card layout and spacing
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Without pool_pre_ping, stale pooled connections (left over after a Postgres
restart or network blip) cause immediate query failures. SQLAlchemy then
propagates the error rather than transparently reconnecting, which crashes
the Quart request handler and triggers a Swarm restart loop.
- pool_pre_ping=True: issues a lightweight SELECT 1 before each checkout;
discards and replaces stale connections silently
- pool_recycle=1800: recycles connections every 30 minutes to prevent
long-idle connections from going stale at the TCP/firewall level
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove page title; move chat widget (quick actions + input) to top full-width
- Remove recent chats section beneath the widget
- Inline streaming response stays full-width between widget and grid
- Two-column grid below (3fr tasks / 2fr notes, collapses to 1 col on mobile)
- Left column: all active tasks categorized by urgency — Overdue, Due Today,
Due This Week, High Priority, In Progress, then Other (capped at 10)
- Other section: broad fetch of all non-done tasks deduped against shown
sections; sorted due-dated items first (asc), then undated by priority
- Right column: recent notes bumped from 5 to 8
- Max-width increased from 1200px to 1400px
- Section labels styled as small uppercase headings; overdue/high-priority
retain colored left-border treatment
- Single "See all" link per column header instead of per-section
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add IP column to the logs table (hidden on mobile)
- Fix expanded detail row condition: also expands when ip_address is set
even if there is no JSON details blob (login/logout events have ip_address
but null details, so they previously could not be expanded at all)
- Show ip_address as a plain line above the JSON blob in the detail row
- Update colspan from 6 to 7 for the new column
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- TRUST_PROXY_HEADERS=true: ensures rate limiting keys on the real client IP
(X-Forwarded-For / X-Real-IP) rather than the Traefik container IP
- SECURE_COOKIES=true: sets the Secure flag on session cookies since the app
is served behind TLS via Traefik
- Remove ports: "5000:5000" — Traefik routes traffic internally; publishing
the port directly would bypass Traefik middleware
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Backend security & correctness:
- Add rate_limit.py: sliding-window rate limiter (asyncio) applied to login,
register, forgot/reset password endpoints (10/60s or 5/300s per IP)
- app.py: add security headers in after_request (X-Frame-Options, CSP,
X-Content-Type-Options, Referrer-Policy) using setdefault to preserve SSE headers
- auth.py: refactor duplicate login_required/admin_required into shared _check_auth()
- config.py: add TRUST_PROXY_HEADERS for proxy-aware client IP resolution
- routes/auth.py: rate limiting, _client_ip() helper, cleaned-up reset_password route
- routes/chat.py, notes.py, tasks.py: int() DoS fix on last_event_id; limit capped
at 500; date.fromisoformat() wrapped in try/except → 400 on invalid dates
- services/auth.py: fix Setting.user_id update bug (filter on NULL not user.id);
reset_password_with_token returns int|None (user_id) instead of bool
- services/backup.py: add _security_notice to full backup JSON export
- services/assist.py: system prompt explicitly preserves markdown list structure
and nested indented sub-items
Infrastructure:
- docker-compose.yml: add healthcheck on app service (/api/health, 10s interval)
- .dockerignore: prevent secrets/node_modules/__pycache__/.env.* leaking into build
Frontend bug fixes:
- TaskCard.vue, TaskViewerView.vue: fix isOverdue() timezone bug (ISO string compare)
- useAssist.ts: accept() now resets state to idle when document changed since proposal
- stores/chat.ts: fix memory leak in _pollUntilLoaded() (try/catch around fetchStatus)
- TiptapEditor.vue: selection offset uses closest-match strategy (not first-match)
- utils/markdown.ts: explicit DOMPurify config with FORCE_BODY; remove as const
(DOMPurify expects mutable string[])
New features:
- Auto-save (5-minute interval) in NoteEditorView and TaskEditorView — only when
editing an existing dirty record; silent on error, shows "Auto-saved" toast
- sectionParser.ts: top-level bullet/numbered list items are now individual sections
in the AI Assist panel (previously treated as one undifferentiated block)
- editor-shared.css: extracted ~500 lines of CSS duplicated between both editors;
includes .inline-assist-btn at global scope (required for teleported elements)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously .btn-back and .btn-edit in editor/viewer toolbars were plain
primary-colored links while all other toolbar actions were proper buttons.
All four views now use the same bordered button style: subtle border,
secondary text color, hover highlights primary border/text — matching
the existing convert/assist-toggle button aesthetic.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Context sidebar + note title:
- ChatView: replace ephemeral context pills with a persistent right-panel sidebar;
auto-found notes accumulate across turns; attached note shows with pin icon;
× button excludes a note from future auto-search; hidden on mobile
- routes/chat.py: batch-fetch note titles via get_notes_by_ids() and inject
context_note_title into each message dict at conversation load time
- notes.py: add get_notes_by_ids() batch fetch helper
- types/chat.ts: add context_note_title field to Message interface
- stores/chat.ts: sendMessage accepts optional 5th arg contextNoteTitle,
included in optimistic user message
- ChatMessage.vue: context badge shows note title instead of 'Note #N'
Expanded LLM tool suite (all with intent router rules + ToolCallCard display):
- delete_note / delete_task: permanent delete with user confirmation (write tool),
type-safe (refuse to delete wrong type), clears note context cache on success
- get_note: fetch full note body by query (search_notes returns only 200-char preview)
- list_notes: browse notes by recency/keyword/tags with limit; notes only
- update_note: add tags + tag_mode (replace/add/remove) parameters
- search_notes: add optional type filter ("note" | "task")
- search_todos (CalDAV): keyword-filter todos, companion to list_todos
- caldav.py: add search_todos() built on top of list_todos()
- generation_task.py: register new tools in _WRITE_TOOLS, _TOOL_LABELS, _TOOL_ACTIONS
- llm.py: update available actions list and guidance in system prompt
- intent.py: routing rules for all new tools
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- New NoteEmbedding model + migration 0014 stores float embeddings (JSONB)
- services/embeddings.py: get_embedding, upsert_note_embedding,
semantic_search_notes (cosine similarity), backfill_note_embeddings
- build_context() now tries semantic search first, falls back to keyword search;
accepts cached_note_ids to reuse last-turn notes and stabilise the system
prompt prefix for Ollama's KV cache
- generation_buffer.py: per-conversation note ID cache (get/set/clear)
- generation_task.py: passes cached IDs into build_context, updates cache
after each turn, and invalidates it after create_note/update_note/create_task
- app.py: pulls nomic-embed-text at startup and launches a background backfill
to embed all existing notes (30 s delay so Ollama has time to load the model)
- routes/notes.py + services/tools.py: fire-and-forget embedding update on
every note create or update via the API or LLM tool calls
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When a conversation exceeds 20 messages (10 exchanges), the oldest
messages are summarized into a compact 3-5 sentence paragraph using the
intent model, and only the most recent 6 messages are passed verbatim.
The summary is injected into the system prompt so the model retains
context without the full token cost. For short conversations the check
is O(1) and returns immediately. The status indicator shows
"Summarizing conversation history..." when the LLM call is needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
stream_chat_with_tools now accepts a think parameter. run_generation
forwards it to Ollama. The message POST route reads think from the
request body. ChatView passes think=true so qwen3 uses chain-of-thought
reasoning for full conversations; the dashboard widget and ChatPanel
omit it, staying fast. Dashboard button updated to "Think it through
in Chat →" to signal the deeper capability.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Before executing any write tool (create/update/delete), the backend now
pauses with an asyncio.Future and emits a tool_pending SSE event. The
frontend displays a ToolConfirmCard with Accept and Decline buttons.
Clicking Accept resolves the Future and proceeds; Decline records a
declined tool_call chip and falls through to regular streaming. Typing
single-word yes/no responses (e.g. "yes", "cancel") also works as
confirmation. 120s timeout auto-declines if no response.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When the intent router detects a tool call, the acknowledgment sentence
and the tool now execute concurrently via asyncio.gather. The acknowledgment
uses the small intent model (already in VRAM) with max_tokens=40, so it
completes in ~200-400ms — the user sees text almost immediately instead of
staring at a status label for the full main-model TTFT (~22s).
The acknowledgment text is:
- Streamed to the client as a chunk event (clears the status spinner)
- Included in the assistant message for round 1 so the main LLM continues
coherently from where the acknowledgment left off
- Recorded in TTFT timing (acknowledgment counts as first token)
Varied phrasing is enforced in the system prompt so responses feel natural
rather than formulaic.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Pinned note: full body restored (truncation is wrong when the user is
explicitly asking about that note's content).
Auto-notes: restored to 2000 chars (800 was too restrictive for useful context).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Auto-notes (keyword-matched): 2000 → 800 chars each (×3 max = 6000 → 2400 chars).
Pinned note (explicit context): was unbounded → capped at 4000 chars with [truncated] marker.
The main post-GPU bottleneck is TTFT caused by the prefill phase — the model
processing the full input before generating any tokens. Shorter context =
faster prefill. Users can ask follow-up questions for more detail.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
If OLLAMA_INTENT_MODEL is configured and differs from OLLAMA_MODEL,
both are pulled concurrently as fire-and-forget tasks at startup.
Deduplicates via a set so pulling the same model twice is never attempted.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Docker Compose:
- Enable Ollama GPU passthrough (nvidia, count: all) in both dev and prod files
- Add OLLAMA_FLASH_ATTENTION=1 (faster attention on GPU in both files)
- Add OLLAMA_MAX_LOADED_MODELS=2 and OLLAMA_KEEP_ALIVE=30m to prod (was already in dev)
- Remove 8G memory limit from prod Ollama service (CPU-bound constraint, no longer valid)
llm.py:
- Increase num_ctx 16384 → 32768 in stream_chat and stream_chat_with_tools (GPU VRAM allows it)
- Increase num_predict cap 4096 → 8192 for tool-augmented responses
generation_task.py:
- Parallelize build_context, get_tools_for_user, and get_setting all from the start
- As soon as tools list is ready (fast DB call), launch classify_intent as an asyncio.Task
- Await build_context and classify_intent together via asyncio.gather
- Intent result is pre-computed before the generation loop; loop just reads pre_intent on round 0
- intent_ms timing now reflects wall-clock time from intent start to completion
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>