Files
minstrel/internal/db/migrations/0072_similarity_cache.up.sql
T
bvandeusenandClaude Opus 5.5 e5dac9ddf0
release / govulncheck (push) Successful in 45s
release / web (push) Successful in 1m27s
release / go (push) Successful in 1m51s
release / integration (push) Successful in 5m36s
release / android (push) Successful in 7m47s
release / Build signed APK (releases and dev) (push) Successful in 8m29s
release / Attach APK to the Release (tag releases only) (push) Skipped
release / Build + push container image (push) Successful in 1m47s
release / Verify release artifacts (tag releases only) (push) Skipped
fix(similarity): keep ListenBrainz's whole answer and resolve it locally (#5296)
The worker kept only the similar recordings already in the library, at most
20 of ListenBrainz's 50, and judged freshness by the edges it had written.
Two failures followed, both measured on the operator's library (#3879):

- A seed whose answer matched nothing wrote nothing, so it was never fresh.
  With the queue ordered by id, 25 such seeds held its head and were
  re-asked every hour; 17 of 2,466 played seeds had any edges.
- A recording that reached the library after its seed was fetched (a
  Lidarr import, an MBID from the AcoustID lookup) was never linked until
  a refetch, which for the stuck seeds never came.

Now every answer is cached whole in listenbrainz_similar_recordings and
every answer, an empty one or a permanent 4xx included, is recorded in
track_similarity_fetches. The queue reads the fetch record: never-fetched
first, then the oldest, refreshed after 30 days. The listenbrainz edges are
derived in SQL from the cache, one present track per recording and no cap,
for the seed just fetched and for every seed once per tick, so new arrivals
link within the hour without asking ListenBrainz again.

Artists get the same queue fix via artist_similarity_fetches; their answer
was already kept in artist_similarity and artist_similarity_unmatched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 19:56:22 -04:00

44 lines
2.0 KiB
SQL

-- #5296: keep ListenBrainz's whole similar-recordings answer, and record every
-- fetch, so the library match is made locally and can be made again.
--
-- Until now the worker kept only the recordings already in the library (at most
-- 20 of LB's 50) and threw the rest away. Two things followed:
--
-- * A recording that reached the library later — a Lidarr import, or an MBID
-- filled in by the AcoustID lookup — was not linked until its seed was
-- fetched again.
-- * A seed with no in-library match wrote nothing at all, and freshness was
-- read from track_similarity, so that seed was never fresh. With the queue
-- ordered by id, 25 such seeds held its head and were re-asked every hour
-- while 2,449 others waited (#3879).
--
-- The listenbrainz rows in track_similarity are now derived from this cache by
-- ResolveListenBrainzTrackEdges.
CREATE TABLE listenbrainz_similar_recordings (
seed_track_id uuid NOT NULL REFERENCES tracks(id) ON DELETE CASCADE,
recording_mbid text NOT NULL,
score DOUBLE PRECISION NOT NULL,
PRIMARY KEY (seed_track_id, recording_mbid)
);
-- The resolve joins the cache to tracks by recording MBID.
CREATE INDEX listenbrainz_similar_recordings_mbid_idx
ON listenbrainz_similar_recordings (recording_mbid);
-- One row per seed that ListenBrainz has answered for, an empty answer
-- included. The worker's queue reads this, not the edges.
CREATE TABLE track_similarity_fetches (
track_id uuid PRIMARY KEY REFERENCES tracks(id) ON DELETE CASCADE,
fetched_at timestamptz NOT NULL DEFAULT now(),
returned integer NOT NULL
);
-- The same for artists. Their answer is already kept, in artist_similarity and
-- artist_similarity_unmatched; this only stops an empty answer being re-asked.
CREATE TABLE artist_similarity_fetches (
artist_id uuid PRIMARY KEY REFERENCES artists(id) ON DELETE CASCADE,
fetched_at timestamptz NOT NULL DEFAULT now(),
returned integer NOT NULL
);