Migration 0050 adds candidate_artist_tags + candidate_artist_tag_state:
folksonomy tags for artists NOT in the library, which track_tags cannot
hold because it's FK'd to tracks(id) and a Discover candidate has no local
row. Slice 6 ranks against these; this slice only fills the cache.
The reuse the task claimed is real and verified: MusicBrainz's
fetchEntityTags(ctx, "artist", mbid, scale) already existed for the #1519
recording→artist fallback, so FetchArtistTags is a thin wrapper. Two
subtleties it does NOT inherit:
- Weight scale is 1.0, not artistTagWeightFactor (0.6). That discount
exists because FetchTrackTags uses artist tags as a *proxy* for a
track's; here the artist IS the subject. Applying it would make these
weights incomparable with track_tags — exactly the comparison slice 6
depends on. Pinned by a test.
- fetchEntityTags reports existing-but-untagged as (empty, nil) so the
track path can fall through. There's no next level here, so empty
becomes the terminal ErrNotFound; otherwise the enricher would settle
a candidate as "enriched" with zero tags.
ArtistTagProvider is the split TrackTagProvider's own doc comment
anticipated ("e.g. artist-level tags"). Last.fm gains artist.getTopTags,
which returns the same toptags envelope, so the response type and
normalizer are reused unchanged.
Rather than write the merge-and-classify loop twice, extracted it from
EnrichTrack into runChain(). The ErrNotFound-vs-transient split is the
load-bearing part — those lead to opposite persistence decisions — so it
now has direct unit tests it never had while inlined.
Bookkeeping is a separate table, not columns, because the "providers had
nothing" outcome must be recordable for a candidate with zero tag rows,
and there is no per-candidate row to hang columns off (
artist_similarity_unmatched holds many rows per candidate). Absence of a
state row means "never processed", so a transient failure writes nothing
and stays eligible.
Two capacity realities are designed for, not papered over:
- The pool is O(library artists x neighbours) and MusicBrainz allows
~1 req/s, so it can never drain in one pass. The eligibility query
returns candidates in descending summed-similarity order, so the ones
that can actually reach a deck are enriched first.
- candidateBatch (50) is smaller than the track batch (200): tracks are
finite and drain to completion, candidates are effectively unbounded
and would otherwise starve the track arm forever.
GC sweeps both tables — the similarity feed churns, and a candidate that
joins the library has its tags in track_tags now. Tags swept before state
so a mid-sweep crash leaves a valid state, not a re-fetch loop.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>