Trash had no end. A note sat in /trash until someone emptied it by hand, and its attachment BYTES sat on disk the whole time — the pile-up the operator asked about. Nothing purged; there was no scheduler at all. Retention is server-owned: `trash_retention_days` (default 30, 0 = keep forever) in the settings registry, so it lands in admin Settings with no migration and takes effect without a restart. A background sweep started in before_serving does the work. Clients learn about a purge the way they learn about any deletion — as a tombstone on the delta feed. An auto-purge nobody can see coming is data loss on a timer, so the window is now visible: /api/config publishes it, notes carry `deleted_at`, Trash leads with the policy, and each card counts down. The countdown rounds DOWN — saying "1 day left" for a note with ten minutes on the clock is the one error here that actually costs someone a note. Three things this turned up on the way: - `DELETE /api/notes/<id>` hard-deleted the row, leaving no tombstone at all. A permanent delete in the web UI never reached a linked device, which would keep its copy forever and push it back on the next edit. It now purges through the same path as everything else. - The purge left `note_revisions` and `note_link_previews` behind. A revision holds the full body, so the text of a "permanently deleted" note was still sitting in the database. - `deleted_at` now SURVIVES a purge instead of being cleared. It's still true, and it means every query that says "not trashed" excludes tombstones for free — without it a content-less row reads as a perfectly normal active note and shows up on the board as a blank card. Desktop keeps its own clock only when there's nobody else to keep one: the sweep runs at startup on an UNLINKED device and refuses otherwise. A linked client that expired notes on its own schedule could destroy something the server was deliberately keeping, then push that delete upstream. Local policy must never outrank the server's — so it also adopts the server's window for the countdown rather than showing its offline default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SreJkbxB4gx8pPsu8QbLPi
12 KiB
ThoughtSync sync protocol
The contract the local-first native clients (Tauri desktop, Android) implement
against. The server is the sync hub: each client keeps a full local store
(SQLite), works fully offline, and reconciles with the server when linked. The
web app does not use this API — it stays on the live REST API (/api/notes,
…) as an online-only stopgap.
All sync endpoints live under /api/sync. Everything is owner-scoped and
deterministic (no AI).
No Postgres CI lane. Trigger/migration/sync behavior is verified by the operator on deploy, not in CI. Pure logic (LWW comparator, paging cursor, token hashing) is unit-tested.
Protocol versioning — the compatibility handshake
Clients and servers update on their own schedules; a self-hosted server can sit on an older release than the desktop app for months. So the wire protocol is versioned separately from either program's release version, and each side declares two numbers: what it speaks, and the oldest counterpart it accepts.
server (src/thoughtsync/sync.py) |
client (desktop/src-tauri/src/sync/compat.rs) |
|
|---|---|---|
| speaks | SYNC_PROTOCOL_VERSION |
CLIENT_PROTOCOL_VERSION |
| accepts down to | MIN_CLIENT_PROTOCOL_VERSION |
MIN_SERVER_PROTOCOL_VERSION |
The server publishes its half on the public, unauthenticated GET /api/config
— a client must be able to ask "can I talk to you?" before it holds a device
token, or even has an account:
{ "site_name": "...", "version": "0.1.0",
"sync_protocol_version": 1,
"min_client_protocol_version": 1,
"sync_features": ["notes", "labels", "attachments", "tombstones", "revisions"],
"trash_retention_days": 30 }
The client identifies itself on every request with
X-ThoughtSync-Client: thoughtsync-desktop/<app version> and
X-ThoughtSync-Protocol: <n>.
sync_features — why versions alone aren't enough
A version number can only say "newer" or "older". sync_features names
capabilities, so a client tests for the one it needs instead of inferring it from
a number. That is what keeps an additive change from forcing a lockstep
upgrade: a newer client meeting an older server drops the missing feature and
syncs everything else.
The policy
- Any wire change → bump
SYNC_PROTOCOL_VERSION. - Additive change (a new field, a new capability) → add a
sync_featuresname. Do not raise a minimum. Old clients keep working. - Breaking change only → raise
MIN_CLIENT_PROTOCOL_VERSION(or the client'sMIN_SERVER_PROTOCOL_VERSION). This is the switch that hard-blocks the other side, so it is the one to be stingy with. - Never gate behavior on the release version (
version) — it's for display.
The three outcomes
The client evaluates the advertisement (compat::evaluate) and gets exactly one
of:
- ok — full parity; sync everything.
- degraded — safe to sync, but named capabilities are unavailable here; the UI says which.
- incompatible — do not sync. Carries
client_must_updateso the message can point at the side that can actually fix it, rather than just saying "incompatible".
A server that predates this handshake sends no protocol fields at all. That is treated as incompatible (update the server) — deliberately not as a parse error, which would look to the user like they mistyped the URL.
Authentication — device bearer tokens
Native clients authenticate with a long-lived device token, not a session cookie. Only the token's SHA-256 hash is stored server-side; the plaintext is shown once at creation.
- First link (no session):
POST /api/auth/device-loginwith{email, password, name}→{token, device, user}. Storetoken; send it asAuthorization: Bearer <token>on every subsequent call. - From the web app (already signed in): the user creates a token at
/account("Linked devices");POST /api/auth/devices{name}→{token, device}. They paste it into the native app. - Manage:
GET /api/auth/devices(list),DELETE /api/auth/devices/<id>(revoke). A revoked token stops authenticating immediately.
Every authenticated request (sync or otherwise) accepts the bearer token in place of the session cookie.
The revision cursor
Every syncable row carries a monotonic sync_revision (bigint), assigned by
a database trigger from a single shared sequence (sync_revision_seq) on every
insert/update. Because it's a shared sequence, one integer is a total-order
watermark across all of a user's notes and labels. It is not a timestamp —
it is immune to clock skew and is only ever compared with >.
The client persists the highest cursor it has fully consumed and passes it back
as ?since=. since=0 (or absent) is a full initial sync.
Entities and what syncs
- Note — the primary sync unit. It travels with its checklist items,
label memberships, and attachment metadata inline (same shape as the REST
serialization). A change to any child bumps the parent note's
sync_revision, so pulling the note re-syncs the whole thing. - Label — the label catalog (name + color) syncs as its own entity so a rename/recolor/delete propagates independently of notes.
- Derived, NOT synced:
[[wiki-links]]and#tagsare parsed from the note body. Clients recompute them locally; the server recomputes them on push. They never travel over the wire. - Attachment blobs sync by id over the existing upload/download routes (see Attachments below); only their metadata rides the delta feed.
Tombstones (deletes)
Two levels, both propagate:
- Trash —
deleted_atis a normal field, and it's on the wire: a trashed note still syncs with its content, the client shows it in its Trash, and the timestamp is what the client counts the retention window against. Restoring clears it. - Purge (permanent delete) — becomes a content-less tombstone:
purged_atis set, title/body/items/labels/attachments/previews/revisions are cleared or removed, and the row is kept. A client seeingpurged_at != nulldeletes the row from its local store.deleted_atdeliberately SURVIVES a purge, so ordinary server-side queries (deleted_at IS NULL) never see a tombstone as a live note. Tombstones themselves are retained indefinitely (cheap for a personal store); revisit if they ever grow large.
Retention — trash expires
A trashed note is purged automatically once it is older than the server's
trash_retention_days setting (default 30, 0 = keep forever), advertised on
/api/config so a client can show the countdown. A background sweep on the server
does the work; clients learn about it as ordinary tombstones and need no special
handling.
A linked client must not run its own expiry. The server owns the policy — one clock, one window. A client that purged on its own schedule could destroy a note the server was deliberately keeping and then push that delete upstream. An unlinked client (offline-only, no server to defer to) expires its own trash on its own default, which is the only case where nothing else can.
Pull — GET /api/sync/changes
Query: ?since=<cursor>&limit=<n> (limit default 500, max 1000).
Response:
{
"notes": [ { "...full note...", "trashed": false, "deleted_at": null,
"sync_revision": 42, "purged_at": null } ],
"labels": [ { "id": "...", "name": "...", "color": "...",
"sync_revision": 43, "purged_at": null, "created_at": "..." } ],
"cursor": 43,
"has_more": false
}
Returns all of the caller's notes + labels (any state — active, archived,
trash, purged) whose sync_revision > since, ascending by revision. Notes and
labels share the sequence, so paging merges the two streams: when either stream
fills a page, cursor advances only to the smaller of the two page
boundaries, so nothing between cursor and the next pull is skipped. Loop while
has_more is true, advancing since = cursor each time.
Note attachment metadata carries {id, url, mime, size, sha256}.
Push — POST /api/sync/push
Body: { "changes": [ ... ] } (max 1000 per batch). Each change:
{ "entity": "note", "id": "<uuid>", "op": "upsert", "edited_at": "<iso8601>",
"title": "...", "body": "...", "color": "blue", "kind": "text",
"pinned": false, "archived": false, "trashed": false, "remind_at": null,
"position": 0, "items": [ {"text": "...", "checked": false} ],
"label_ids": ["<uuid>", ...], "created_at": "<iso8601, on create>" }
- Client-generated ids. Notes/labels are UUIDs; the client mints the id when it creates the row offline and sends it here. Create-if-absent, else update.
- Whole-note semantics. A note upsert carries the client's full current
state (not a partial patch) — the server overwrites all scalar fields, replaces
items, and sets manual label memberships from
label_ids(tag-sourced labels are re-derived from the body).[[links]]/#tagsare recomputed server-side. op: "delete"purges (tombstones) the row. Trashing is just an upsert withtrashed: true.- Labels:
{entity: "label", op: "upsert"|"delete", id, edited_at, name, color}. A per-owner name clash on a different id is rejected (fix locally and retry).
Conflict resolution — last-write-wins + history
On a clash the most-recently-edited version wins, by edited_at
(client wall-clock) compared against the server row's last-edit time
(updated_at). The client applies iff client.edited_at >= server.updated_at.
When a winning upsert overwrites an existing title/body, the server first
snapshots the overwritten version into the note's version history
(note_revisions) — so a "lost" edit is never truly lost; it's one Restore away.
A stale delete loses to a newer server edit.
Response — per item, so the client can mark its local rows synced:
{ "results": [
{ "id": "...", "entity": "note", "status": "created", "sync_revision": 44 },
{ "id": "...", "entity": "note", "status": "kept", "sync_revision": 40 },
{ "id": "...", "entity": "label","status": "rejected","error": "name in use" }
] }
status ∈ created | applied | kept | noop | rejected. kept means the server
had a newer edit and the client should adopt the server version on its next pull.
Attachments (blobs)
Metadata rides the delta feed (id, url, mime, size, sha256); the bytes move
over the existing routes:
- Upload:
POST /api/notes/<note_id>/attachments(multipart, fieldfile; optional fieldidto keep a client-minted attachment id). Re-uploading an id the note already has is an idempotent no-op. Server stores + hashes the bytes. - Download:
GET /api/notes/<note_id>/attachments/<id>(owner/shared scoped). - The client uses
sha256to skip blobs it already holds and to verify integrity after download. (Currently image mimes only; broadening to any file is tracked separately.)
A sync cycle
- Push local changes since the last sync (batched). Apply the per-item
results (mark synced, adopt server version where
kept). - Pull from the stored
sincecursor untilhas_moreis false. Upsert notes/labels into the local store; delete rows whosepurged_atis set; download any attachment blobs referenced by a new/changedsha256. - Persist the new
cursor.
Initial sync is the same with since=0. Because everything is keyed by stable
ids and a monotonic cursor, the cycle is idempotent and resumable — a client
can crash mid-sync and simply resume from its last persisted cursor.