The worker kept only the similar recordings already in the library, at most
20 of ListenBrainz's 50, and judged freshness by the edges it had written.
Two failures followed, both measured on the operator's library (#3879):
- A seed whose answer matched nothing wrote nothing, so it was never fresh.
With the queue ordered by id, 25 such seeds held its head and were
re-asked every hour; 17 of 2,466 played seeds had any edges.
- A recording that reached the library after its seed was fetched (a
Lidarr import, an MBID from the AcoustID lookup) was never linked until
a refetch, which for the stuck seeds never came.
Now every answer is cached whole in listenbrainz_similar_recordings and
every answer, an empty one or a permanent 4xx included, is recorded in
track_similarity_fetches. The queue reads the fetch record: never-fetched
first, then the oldest, refreshed after 30 days. The listenbrainz edges are
derived in SQL from the cache, one present track per recording and no cap,
for the seed just fetched and for every seed once per tick, so new arrivals
link within the hour without asking ListenBrainz again.
Artists get the same queue fix via artist_similarity_fetches; their answer
was already kept in artist_similarity and artist_similarity_unmatched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The similar_artists and coplay_artists arms ordered by artist score, so the
closest related artist's catalogue filled the whole LIMIT. On the operator's
library the 30-row similar_artists arm held exactly one artist for all 17
seeds measured (#3879), though each seed had 7-33 similar artists in the
library. Both arms now rank tracks within each artist and take every
artist's first track, best artist first, before anyone's second.
Both also skip missing tracks: the outer select already dropped them, but
only after they had taken places in the arm's LIMIT.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>