grammar: a #tag starts after whitespace and with a letter, on every surface
CI & Build / Python lint (push) Successful in 2s
CI & Build / Build now, or wait for Android? (push) Successful in 2s
Android / Build, or is the channel already serving this? (push) Successful in 4s
Desktop (Tauri) / Build, or is the channel already serving this? (push) Successful in 2s
CI & Build / Web typecheck and unit tests (push) Successful in 8s
CI & Build / Python tests (push) Successful in 11s
CI & Build / integration (push) Successful in 38s
CI & Build / Build & push image (push) Skipped
Desktop (Tauri) / Web tests, clippy, Rust tests and rustfmt (push) Successful in 2m7s
Desktop (Tauri) / Windows installer (cross-compiled) (push) Successful in 3m26s
Desktop (Tauri) / Tauri desktop (Linux) (push) Successful in 4m38s
Desktop (Tauri) / Update manifest (push) Successful in 4s
Android / Kotlin + Rust (APK) (push) Successful in 10m35s

The shared fixture went red on the server (run 8468: 5 failed), because the three
tag rules disagreed:
- core and web: any non-tag character counts as a boundary, so `(#todo)`
  and `end.#tag` are tags, and so is the `/#section` of a pasted URL;
- server: only whitespace counts, but `#1st` and `#_x` are tags.

A note's labels could therefore change every time it synced.

All three now share the strict rule: start of line or whitespace, then a
letter, then letters, digits, `_` and `-`. Nothing becomes a tag that wasn't
already one everywhere, and URL anchors stop becoming labels on desktop and
Android. The server's existing `http://x/#nope` test already expected this.

derive.rs's boundary, markdown.ts's lookbehind and tags.py's regex change
together; `_is_tag` goes because the regex now requires the letter. The
fixture flips `(#todo)`, `end.#tag` and its lift case, and adds the URL case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-06 23:22:53 -04:00
co-authored by Claude Opus 5.5
parent 535331c5b2
commit 38da5160a0
4 changed files with 33 additions and 21 deletions
+5 -4
View File
@@ -46,15 +46,16 @@ export interface TaskMeta {
// chance to start inside `**bold #x**`.
//
// The grammar MIRRORS `line_tags` in core/src/local/derive.rs, which is the definition:
// a `#` at a word boundary (the preceding character is neither a tag character nor
// another `#`, so `a#b` and `##x` are not tags), a letter immediately after it, then
// a `#` at the start of the text or after whitespace (so `a#b`, `##x`, `(#x)` and the
// `/#section` of a URL are not tags), a letter immediately after it, then
// alphanumerics, `_` and `-`. Rust's `is_alphanumeric` is `Alphabetic | N`, hence the
// property escapes rather than `\w` — and hence the `u` flag.
// property escapes rather than `\w` — and hence the `u` flag. All three
// implementations run core/testdata/grammar.json (notes/grammar.test.ts here).
//
// A heading cannot collide with this: `parseMarkdown` requires a space after the `#`s,
// which `#tag` by definition does not have.
const INLINE_RE =
/(`[^`]+`)|(\*\*[^*]+\*\*)|(\*[^*]+\*)|(_[^_]+_)|((?<![\p{Alphabetic}\p{N}_#-])#\p{Alphabetic}[\p{Alphabetic}\p{N}_-]*)/gu;
/(`[^`]+`)|(\*\*[^*]+\*\*)|(\*[^*]+\*)|(_[^_]+_)|((?<!\S)#\p{Alphabetic}[\p{Alphabetic}\p{N}_-]*)/gu;
export function parseInline(text: string): InlineToken[] {
const tokens: InlineToken[] = [];