Files
minstrel/internal/db/queries/tracks.sql
T
bvandeusen f6d1cf24f0
test-go / test (push) Successful in 1m0s
test-go / integration (push) Successful in 5m10s
feat(library): detect missing files and stop offering them — #2523
Nothing in Minstrel ever noticed a deleted file. The walk only visits
paths that exist, so a row whose file was gone was never scanned, never
errored, never counted — permanently invisible. classifyEvent ignores
fsnotify removals by design, and the safety-net scan is the same walk, so
it covers additions only. Rows accumulated forever.

Found on the operator's library: a completed scan reported
skipped=24185 errored=0 while the MBID backfill (which opens files by DB
path rather than walking) logged ~40 "no such file or directory" across
three reorganised albums. Those rows also kept their pre-#2499 welded
genre, which is how this surfaced — the version-stamped tag re-read can
only reach files the walk visits.

The harm is not cosmetic. tracks is the candidate universe for
recommendation.sql / discover.sql / system_mixes.sql and nothing filtered
on file existence, so a mix could spend a slot on a track that cannot
stream.

Marks rather than deletes. A missing file is a claim about the filesystem
and the filesystem lies transiently — an unmounted volume, a network
blip, a container that started before its media mount attached. Every
sweep in internal/gc resolves a truth INSIDE the database and is safe to
run blind; this one is not, so no deletion happens here. Three guards
refuse to act on ambiguous evidence: every scan root must resolve to a
non-empty directory, the walk must have seen at least one file, and one
reconcile may newly mark at most 25% of the library. Clearing a mark is
never the dangerous direction, so it runs unconditionally — otherwise a
library that tripped the cap could never recover once the mount returned.

Only a full Scan reconciles. The walk's set of seen paths is the
evidence, and ScanFiles has no basis for concluding anything about files
it did not look at.

Excludes marked tracks from all 13 track-emitting queries (radio x2,
system mixes x5, discover x4, most-played x2), the 6 play-history seed
picks, and the genre browse axis. Deliberately NOT filtered: the shared
ListPlaylistTracks read path, because it also serves user-curated
playlists where hiding a track the user added would be wrong — system
playlists shed orphans on their next daily rebuild instead. History and
the taste profile also keep them: those record the past, and a track you
played 200 times still says something about your taste.

Reconcile tallies land in scan_runs so a disappearance is visible rather
than discovered when a mix comes up short.
2026-08-06 14:34:53 -04:00

171 lines
6.3 KiB
SQL

-- name: UpsertTrack :one
-- file_path is the canonical identity for library scan; mbid is secondary.
INSERT INTO tracks (
title, album_id, artist_id, track_number, disc_number,
duration_ms, file_path, file_size, file_format, bitrate, mbid, genre,
tag_read_version
) VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13)
ON CONFLICT (file_path) DO UPDATE SET
title = EXCLUDED.title,
album_id = EXCLUDED.album_id,
artist_id = EXCLUDED.artist_id,
track_number = EXCLUDED.track_number,
disc_number = EXCLUDED.disc_number,
duration_ms = EXCLUDED.duration_ms,
file_size = EXCLUDED.file_size,
file_format = EXCLUDED.file_format,
bitrate = EXCLUDED.bitrate,
mbid = EXCLUDED.mbid,
genre = EXCLUDED.genre,
-- Stamped on update too, so a tag-repair pass marks rows as done and the
-- next scan can short-circuit them again (#2499).
tag_read_version = EXCLUDED.tag_read_version,
updated_at = now()
RETURNING *;
-- name: ListTracksMissingMbidWithPath :many
-- Track recording-MBID backfill: tracks with NULL mbid that still have
-- a file to re-read. $1 caps the batch (mirrors the album backfill).
SELECT id, file_path
FROM tracks
WHERE mbid IS NULL
ORDER BY id
LIMIT $1;
-- name: SetTrackMbidIfNull :exec
-- Heal a track's recording MBID only while still NULL — idempotent, so
-- re-running the backfill is a no-op for already-healed rows.
UPDATE tracks
SET mbid = $2, updated_at = now()
WHERE id = $1 AND mbid IS NULL;
-- name: GetTrackByID :one
SELECT * FROM tracks WHERE id = $1;
-- name: GetTrackByPath :one
SELECT * FROM tracks WHERE file_path = $1;
-- name: ListTracksByAlbum :many
-- $1 = album_id, $2 = user_id. Pass pgtype.UUID{Valid: false} (NULL)
-- to skip the per-user quarantine filter; the NOT EXISTS clause on
-- a NULL user_id never matches a row, so every track passes through.
SELECT * FROM tracks
WHERE album_id = $1
AND NOT EXISTS (
SELECT 1 FROM lidarr_quarantine q
WHERE q.user_id = $2 AND q.track_id = tracks.id
)
ORDER BY disc_number NULLS LAST, track_number NULLS LAST;
-- name: CountTracksByAlbum :one
SELECT count(*) FROM tracks WHERE album_id = $1;
-- name: SearchTracks :many
-- $1 = title query, $2 = user_id (NULL to skip quarantine filter),
-- $3 = limit, $4 = offset.
SELECT * FROM tracks
WHERE title ILIKE '%' || $1::text || '%'
AND NOT EXISTS (
SELECT 1 FROM lidarr_quarantine q
WHERE q.user_id = $2 AND q.track_id = tracks.id
)
ORDER BY title
LIMIT $3 OFFSET $4;
-- name: CountTracksMatching :one
-- $1 = title query, $2 = user_id (NULL to skip quarantine filter).
SELECT COUNT(*) FROM tracks
WHERE title ILIKE '%' || $1::text || '%'
AND NOT EXISTS (
SELECT 1 FROM lidarr_quarantine q
WHERE q.user_id = $2 AND q.track_id = tracks.id
);
-- name: ListArtistTracksForUser :many
-- M6a: every track for the artist across their albums, with album_title
-- and artist_name joined. Honors per-user lidarr_quarantine. Used by
-- /api/artists/{id}/tracks for the artist-card play affordance, which
-- shuffles client-side. Ordering matches album/track natural order so
-- the shuffle has a deterministic input.
SELECT sqlc.embed(t),
albums.title AS album_title,
artists.name AS artist_name
FROM tracks t
JOIN albums ON albums.id = t.album_id
JOIN artists ON artists.id = t.artist_id
WHERE t.artist_id = $1
AND NOT EXISTS (
SELECT 1 FROM lidarr_quarantine q
WHERE q.user_id = $2 AND q.track_id = t.id
)
ORDER BY albums.release_date NULLS LAST, albums.sort_title,
t.disc_number NULLS FIRST, t.track_number NULLS FIRST, t.id;
-- name: ListRandomTracksForUser :many
-- #427 S4: backing query for GET /api/library/shuffle — the online
-- source for "Shuffle all". N random tracks across the whole
-- library, per-user-quarantine filtered. $1 user_id, $2 limit.
SELECT sqlc.embed(t),
albums.title AS album_title,
artists.name AS artist_name
FROM tracks t
JOIN albums ON albums.id = t.album_id
JOIN artists ON artists.id = t.artist_id
WHERE NOT EXISTS (
SELECT 1 FROM lidarr_quarantine q
WHERE q.user_id = $1 AND q.track_id = t.id
)
ORDER BY random()
LIMIT $2;
-- name: CountTracksByArtist :one
-- Used by request-progress reporting to count tracks ingested under a
-- matched artist (sum across all the artist's albums) while a request
-- is still in flight.
SELECT COUNT(*) FROM tracks WHERE artist_id = $1;
-- name: DeleteTrack :one
-- M7 #372: hard delete with FK cascade. The CASCADE on track_id from
-- play_events / general_likes_tracks / lidarr_quarantine /
-- lidarr_quarantine_actions handles their cleanup. RETURNING gives us
-- album_id + artist_id for the album-empty / artist-empty cascade
-- checks the service does next.
DELETE FROM tracks WHERE id = $1
RETURNING id, album_id, artist_id, file_path, mbid;
-- name: GetTracksByIDs :many
-- Batched lookup used by /api/library/sync to hydrate upsert payloads
-- (#357). Mirror of GetArtistsByIDs.
SELECT * FROM tracks WHERE id = ANY($1::uuid[]);
-- name: ListTrackPathsForReconcile :many
-- Every row's path + current missing mark, for the scanner's reconcile pass
-- (#2523). Deliberately unfiltered and unpaged: reconcile has to compare the
-- WHOLE table against what the walk saw, and a filtered subset would let rows
-- outside it drift forever. Three narrow columns keep it cheap even on a
-- library of a few hundred thousand tracks.
SELECT id, file_path, missing_since FROM tracks;
-- name: MarkTracksMissing :execrows
-- Marks rows whose file the walk did not see. `missing_since IS NULL` in the
-- predicate makes this idempotent: a row already marked keeps its ORIGINAL
-- timestamp, so "how long has it been gone" survives repeated scans. Losing
-- that would make any age-based cleanup policy meaningless.
--
-- updated_at is deliberately NOT touched. It tracks content changes and gates
-- the scanner's mtime skip; moving it here would make a returning file look
-- newer than its own mtime and stop its tags being re-read.
UPDATE tracks
SET missing_since = now()
WHERE id = ANY(sqlc.arg(ids)::uuid[])
AND missing_since IS NULL;
-- name: ClearTracksMissing :execrows
-- Clears the mark on rows whose file is back. Runs independently of the mtime
-- skip check, so a file that reappears unchanged is un-marked even though the
-- scanner skips re-reading its tags.
UPDATE tracks
SET missing_since = NULL
WHERE id = ANY(sqlc.arg(ids)::uuid[])
AND missing_since IS NOT NULL;