release / govulncheck (push) Successful in 45s
release / web (push) Successful in 1m27s
release / go (push) Successful in 1m51s
release / integration (push) Successful in 5m36s
release / android (push) Successful in 7m47s
release / Build signed APK (releases and dev) (push) Successful in 8m29s
release / Attach APK to the Release (tag releases only) (push) Skipped
release / Build + push container image (push) Successful in 1m47s
release / Verify release artifacts (tag releases only) (push) Skipped
The worker kept only the similar recordings already in the library, at most 20 of ListenBrainz's 50, and judged freshness by the edges it had written. Two failures followed, both measured on the operator's library (#3879): - A seed whose answer matched nothing wrote nothing, so it was never fresh. With the queue ordered by id, 25 such seeds held its head and were re-asked every hour; 17 of 2,466 played seeds had any edges. - A recording that reached the library after its seed was fetched (a Lidarr import, an MBID from the AcoustID lookup) was never linked until a refetch, which for the stuck seeds never came. Now every answer is cached whole in listenbrainz_similar_recordings and every answer, an empty one or a permanent 4xx included, is recorded in track_similarity_fetches. The queue reads the fetch record: never-fetched first, then the oldest, refreshed after 30 days. The listenbrainz edges are derived in SQL from the cache, one present track per recording and no cap, for the seed just fetched and for every seed once per tick, so new arrivals link within the hour without asking ListenBrainz again. Artists get the same queue fix via artist_similarity_fetches; their answer was already kept in artist_similarity and artist_similarity_unmatched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
145 lines
6.1 KiB
SQL
145 lines
6.1 KiB
SQL
-- name: ListPlayedTracksNeedingSimilarity :many
|
|
-- The worker's track queue: played, present tracks with a recording MBID that
|
|
-- ListenBrainz has never answered for, or not in 30 days. Never-fetched first,
|
|
-- then the oldest answer. Freshness is read from track_similarity_fetches, not
|
|
-- from the edges, so a seed whose answer matched nothing in the library still
|
|
-- counts as fetched (#5296); reading the edges is what let 25 such seeds hold
|
|
-- the head of the queue forever.
|
|
SELECT t.id, t.mbid
|
|
FROM tracks t
|
|
LEFT JOIN track_similarity_fetches f ON f.track_id = t.id
|
|
WHERE t.mbid IS NOT NULL
|
|
AND t.missing_since IS NULL
|
|
AND EXISTS (SELECT 1 FROM play_events pe WHERE pe.track_id = t.id)
|
|
AND (f.track_id IS NULL OR f.fetched_at < now() - interval '30 days')
|
|
ORDER BY f.fetched_at NULLS FIRST, t.id
|
|
LIMIT $1;
|
|
|
|
-- name: ListPlayedArtistsNeedingSimilarity :many
|
|
-- The artist queue, with the same shape and the same reason.
|
|
SELECT ar.id, ar.mbid
|
|
FROM artists ar
|
|
LEFT JOIN artist_similarity_fetches f ON f.artist_id = ar.id
|
|
WHERE ar.mbid IS NOT NULL
|
|
AND EXISTS (
|
|
SELECT 1 FROM tracks t
|
|
JOIN play_events pe ON pe.track_id = t.id
|
|
WHERE t.artist_id = ar.id
|
|
)
|
|
AND (f.artist_id IS NULL OR f.fetched_at < now() - interval '30 days')
|
|
ORDER BY f.fetched_at NULLS FIRST, ar.id
|
|
LIMIT $1;
|
|
|
|
-- name: GetArtistsByMBIDs :many
|
|
SELECT id, mbid FROM artists WHERE mbid = ANY($1::text[]);
|
|
|
|
-- name: DeleteSimilarRecordingsForSeed :exec
|
|
-- A fetch replaces the seed's cached answer whole.
|
|
DELETE FROM listenbrainz_similar_recordings WHERE seed_track_id = $1;
|
|
|
|
-- name: InsertSimilarRecordings :exec
|
|
-- The seed's whole answer, in or out of the library. An MBID LB lists twice
|
|
-- keeps its best score.
|
|
INSERT INTO listenbrainz_similar_recordings (seed_track_id, recording_mbid, score)
|
|
SELECT sqlc.arg(seed_track_id)::uuid, r.mbid, max(r.score)
|
|
FROM (SELECT unnest(sqlc.arg(mbids)::text[]) AS mbid,
|
|
unnest(sqlc.arg(scores)::float8[]) AS score) r
|
|
WHERE r.mbid <> ''
|
|
GROUP BY r.mbid;
|
|
|
|
-- name: RecordTrackSimilarityFetch :exec
|
|
INSERT INTO track_similarity_fetches (track_id, fetched_at, returned)
|
|
VALUES ($1, now(), $2)
|
|
ON CONFLICT (track_id) DO UPDATE SET fetched_at = now(), returned = EXCLUDED.returned;
|
|
|
|
-- name: RecordArtistSimilarityFetch :exec
|
|
INSERT INTO artist_similarity_fetches (artist_id, fetched_at, returned)
|
|
VALUES ($1, now(), $2)
|
|
ON CONFLICT (artist_id) DO UPDATE SET fetched_at = now(), returned = EXCLUDED.returned;
|
|
|
|
-- name: ResolveListenBrainzTrackEdges :exec
|
|
-- Derives the listenbrainz edges in track_similarity from the cached answers:
|
|
-- every cached recording that is a present track in the library becomes an
|
|
-- edge, whenever it got there. Scoped to the given seeds, or every seed when
|
|
-- seed_ids is NULL.
|
|
--
|
|
-- One track per recording MBID: two copies of a song (an mp3 beside its flac,
|
|
-- or one recording on two releases) would otherwise be two candidates, and the
|
|
-- pool would offer the same song twice. No cap per seed: LB's answer is already
|
|
-- at most 50.
|
|
--
|
|
-- Only edges of seeds the cache owns are removed, those with a fetch row.
|
|
-- Edges written before the cache existed, or copied onto a survivor by a fold
|
|
-- (merge.sql), stay until their seed is fetched.
|
|
WITH want AS (
|
|
SELECT DISTINCT ON (c.seed_track_id, c.recording_mbid)
|
|
c.seed_track_id AS track_a_id, t.id AS track_b_id, c.score
|
|
FROM listenbrainz_similar_recordings c
|
|
JOIN tracks t ON t.mbid = c.recording_mbid
|
|
AND t.missing_since IS NULL
|
|
AND t.id <> c.seed_track_id
|
|
WHERE sqlc.narg(seed_ids)::uuid[] IS NULL
|
|
OR c.seed_track_id = ANY(sqlc.narg(seed_ids)::uuid[])
|
|
ORDER BY c.seed_track_id, c.recording_mbid, t.id
|
|
),
|
|
upserted AS (
|
|
INSERT INTO track_similarity (track_a_id, track_b_id, score, source, fetched_at)
|
|
SELECT track_a_id, track_b_id, score, 'listenbrainz', now() FROM want
|
|
ON CONFLICT (track_a_id, track_b_id, source) DO UPDATE
|
|
SET score = EXCLUDED.score, fetched_at = EXCLUDED.fetched_at
|
|
WHERE track_similarity.score IS DISTINCT FROM EXCLUDED.score
|
|
RETURNING 1
|
|
)
|
|
DELETE FROM track_similarity ts
|
|
WHERE ts.source = 'listenbrainz'
|
|
AND (sqlc.narg(seed_ids)::uuid[] IS NULL OR ts.track_a_id = ANY(sqlc.narg(seed_ids)::uuid[]))
|
|
AND EXISTS (SELECT 1 FROM track_similarity_fetches f WHERE f.track_id = ts.track_a_id)
|
|
AND NOT EXISTS (
|
|
SELECT 1 FROM want w
|
|
WHERE w.track_a_id = ts.track_a_id AND w.track_b_id = ts.track_b_id
|
|
);
|
|
|
|
-- name: UpsertArtistSimilarity :exec
|
|
INSERT INTO artist_similarity (artist_a_id, artist_b_id, score, source, fetched_at)
|
|
VALUES ($1, $2, $3, 'listenbrainz', now())
|
|
ON CONFLICT (artist_a_id, artist_b_id, source)
|
|
DO UPDATE SET score = EXCLUDED.score, fetched_at = EXCLUDED.fetched_at;
|
|
|
|
-- name: UpsertArtistSimilarityUnmatched :exec
|
|
-- Persists an out-of-library similar-artist MBID. Idempotent on
|
|
-- (seed_artist_id, candidate_mbid, source) — re-fetches refresh the
|
|
-- name/score and bump fetched_at.
|
|
INSERT INTO artist_similarity_unmatched (
|
|
seed_artist_id, candidate_mbid, candidate_name, score, source
|
|
) VALUES ($1, $2, $3, $4, $5)
|
|
ON CONFLICT (seed_artist_id, candidate_mbid, source) DO UPDATE SET
|
|
candidate_name = EXCLUDED.candidate_name,
|
|
score = EXCLUDED.score,
|
|
fetched_at = now();
|
|
|
|
-- name: ListSimilarArtistsForArtist :many
|
|
-- In-library artists similar to a seed artist, ranked by best similarity
|
|
-- score across sources (deduped per candidate). cover_album_id + album_count
|
|
-- mirror ListArtistsAlphaWithCovers so the strip renders identical cards.
|
|
SELECT sqlc.embed(artists),
|
|
cov.id AS cover_album_id,
|
|
cnt.album_count::bigint AS album_count
|
|
FROM (
|
|
SELECT artist_b_id, max(score) AS sim_score
|
|
FROM artist_similarity
|
|
WHERE artist_a_id = sqlc.arg(seed_artist_id)
|
|
GROUP BY artist_b_id
|
|
) s
|
|
JOIN artists ON artists.id = s.artist_b_id
|
|
LEFT JOIN LATERAL (
|
|
SELECT id FROM albums
|
|
WHERE artist_id = artists.id AND cover_art_path IS NOT NULL
|
|
ORDER BY created_at DESC LIMIT 1
|
|
) cov ON true
|
|
LEFT JOIN LATERAL (
|
|
SELECT count(*) AS album_count
|
|
FROM albums WHERE artist_id = artists.id
|
|
) cnt ON true
|
|
ORDER BY s.sim_score DESC, artists.sort_name
|
|
LIMIT sqlc.arg(result_limit);
|