M10.7b: pull the change feed into the local store (task 2105)
Desktop (Tauri) / Tauri desktop (Linux) (push) Failing after 26s
Desktop (Tauri) / Windows installer (cross-compiled) (push) Successful in 1m53s

Server -> local. sync/wire.rs mirrors the delta-feed JSON exactly as
notes/serialize.py sends it; sync/pull.rs applies it.

ATOMICITY IS THE POINT. The cursor is written in the SAME transaction as the
page it describes. A cursor committed ahead of its data would skip those rows
forever while reporting a clean sync — the worst kind of failure, because
nothing looks wrong. A test forces a mid-page failure and asserts the cursor
stayed put.

Every degradation leans toward re-downloading rather than skipping: an
unparseable cursor means full sync, wire fields are all defaulted so a newer
server adding a field (or an older one omitting one) yields a partial note
instead of a rejected page, and a page that fails rolls back whole.

Labels are applied before notes so a membership never references a row that
doesn't exist. A note also carries enough of its labels to materialize them,
because notes and labels page from ONE shared sequence and a note can arrive
referencing a label whose own delta landed in an earlier page.

via_tag is applied verbatim rather than re-deriving #tags from the body. The
server already reconciled them on save, and re-deriving would go through the
local find-or-create path, which marks new labels dirty — pushing them
straight back. Sync churn manufactured out of nothing.

Duplicate-label merge, the subtle one: a label created offline can collide by
name with one the server already had under a different id. Both sides enforce
one label per name, so the server's row has to win — but simply deleting the
local duplicate would CASCADE its note_labels away, stripping the label off
notes this pull never mentions, with no later page to repair it. So we free
the name, insert the server's row, re-point the memberships, then drop the
husk. Tested.

Children (items/attachments/previews/labels) are replaced wholesale rather
than diffed: a delta carries the note's FULL state, so what arrived IS the
complete set, and diffing could strand a row the server no longer has.

The loop trusts the data over the flag — a server claiming has_more without
advancing its cursor stops with an error instead of spinning forever.

Pull can overwrite a row with unpushed local edits. The documented cycle is
push-then-pull (M10.7c), so that should never happen; when it does it's
counted as clobbered_dirty and logged rather than hidden.

17 tests, all against an in-memory database.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SreJkbxB4gx8pPsu8QbLPi
This commit is contained in:
2026-07-26 00:05:27 -04:00
co-authored by Claude Opus 5
parent 7d9a6509f3
commit dc8b2d360d
6 changed files with 938 additions and 5 deletions
+51 -5
View File
@@ -14,6 +14,7 @@ use reqwest::{RequestBuilder, StatusCode};
use serde::{Deserialize, Serialize};
use super::compat::{self, Compatibility, ServerInfo};
use super::wire;
/// Timeout for the short request/response calls in this module. Kept tight because a
/// user is watching a button while they run, and the most common mistake — a wrong
@@ -21,6 +22,15 @@ use super::compat::{self, Compatibility, ServerInfo};
/// just look frozen. The sync engine's bulk transfers will need their own, longer one.
const REQUEST_TIMEOUT: Duration = Duration::from_secs(10);
/// Bulk transfers get much longer: a first full sync can be thousands of notes, and
/// failing one at ten seconds would make a large store impossible to ever pull.
const SYNC_TIMEOUT: Duration = Duration::from_secs(120);
/// Shared by every call that presents a token, so a revoked one reads the same way
/// wherever it surfaces.
const TOKEN_REJECTED: &str = "This server rejected the device token — it may have been \
revoked. Unlink and link again to issue a new one.";
/// What the link UI needs after a handshake: where we ended up (the normalized URL,
/// which may differ from what was typed), who answered, and whether we can work
/// with them.
@@ -48,13 +58,17 @@ struct DeviceLoginResponse {
user: Identity,
}
fn http() -> Result<reqwest::Client, String> {
fn http_with(timeout: Duration) -> Result<reqwest::Client, String> {
reqwest::Client::builder()
.timeout(REQUEST_TIMEOUT)
.timeout(timeout)
.build()
.map_err(|e| format!("Could not start the network client: {e}"))
}
fn http() -> Result<reqwest::Client, String> {
http_with(REQUEST_TIMEOUT)
}
/// Attach the client-identity headers every request carries, plus a bearer token
/// when we hold one.
fn prepare(builder: RequestBuilder, token: Option<&str>) -> RequestBuilder {
@@ -179,6 +193,36 @@ pub async fn fetch_identity(base_url: &str, token: &str) -> Result<Identity, Str
.map_err(|_| format!("{base_url} accepted the token but sent an unexpected reply."))
}
/// Fetch one page of the change feed, starting after `since`.
///
/// The caller loops until `has_more` is false (see `pull::run`); paging lives there
/// rather than here so the transport stays a single request/response.
pub async fn fetch_changes(
base_url: &str,
token: &str,
since: i64,
) -> Result<wire::ChangesPage, String> {
let url = format!("{base_url}/api/sync/changes?since={since}");
let request = prepare(http_with(SYNC_TIMEOUT)?.get(url), Some(token));
let response = request
.send()
.await
.map_err(|e| describe_transport_error(base_url, &e))?;
let status = response.status();
if status == StatusCode::UNAUTHORIZED {
return Err(TOKEN_REJECTED.to_string());
}
if !status.is_success() {
return Err(unexpected_status(base_url, status));
}
response
.json()
.await
.map_err(|e| format!("Couldn't read the change feed from {base_url}: {e}"))
}
/// The public, unauthenticated endpoint carrying the handshake.
fn config_url(base_url: &str) -> String {
format!("{base_url}/api/config")
@@ -196,10 +240,12 @@ fn me_url(base_url: &str) -> String {
/// Display is accurate but reads like a stack trace.
fn describe_transport_error(base_url: &str, err: &reqwest::Error) -> String {
if err.is_timeout() {
// No specific duration here: these calls run under two different budgets
// (interactive vs bulk sync), and naming the wrong one is worse than naming
// none.
format!(
"{base_url} didn't respond within {} seconds. It may be offline, or \
unreachable from this network.",
REQUEST_TIMEOUT.as_secs()
"{base_url} didn't respond in time. It may be offline, or unreachable \
from this network."
)
} else if err.is_connect() {
format!(