Files
minstrel/internal/dbtest/reset.go
T
bvandeusenandClaude Opus 5 4f9b083eec
test-go / test (push) Failing after 32s
test-go / integration (push) Successful in 4m50s
feat(discover): artist-tag cache for out-of-library candidates — #2376
Migration 0050 adds candidate_artist_tags + candidate_artist_tag_state:
folksonomy tags for artists NOT in the library, which track_tags cannot
hold because it's FK'd to tracks(id) and a Discover candidate has no local
row. Slice 6 ranks against these; this slice only fills the cache.

The reuse the task claimed is real and verified: MusicBrainz's
fetchEntityTags(ctx, "artist", mbid, scale) already existed for the #1519
recording→artist fallback, so FetchArtistTags is a thin wrapper. Two
subtleties it does NOT inherit:

  - Weight scale is 1.0, not artistTagWeightFactor (0.6). That discount
    exists because FetchTrackTags uses artist tags as a *proxy* for a
    track's; here the artist IS the subject. Applying it would make these
    weights incomparable with track_tags — exactly the comparison slice 6
    depends on. Pinned by a test.
  - fetchEntityTags reports existing-but-untagged as (empty, nil) so the
    track path can fall through. There's no next level here, so empty
    becomes the terminal ErrNotFound; otherwise the enricher would settle
    a candidate as "enriched" with zero tags.

ArtistTagProvider is the split TrackTagProvider's own doc comment
anticipated ("e.g. artist-level tags"). Last.fm gains artist.getTopTags,
which returns the same toptags envelope, so the response type and
normalizer are reused unchanged.

Rather than write the merge-and-classify loop twice, extracted it from
EnrichTrack into runChain(). The ErrNotFound-vs-transient split is the
load-bearing part — those lead to opposite persistence decisions — so it
now has direct unit tests it never had while inlined.

Bookkeeping is a separate table, not columns, because the "providers had
nothing" outcome must be recordable for a candidate with zero tag rows,
and there is no per-candidate row to hang columns off (
artist_similarity_unmatched holds many rows per candidate). Absence of a
state row means "never processed", so a transient failure writes nothing
and stays eligible.

Two capacity realities are designed for, not papered over:
  - The pool is O(library artists x neighbours) and MusicBrainz allows
    ~1 req/s, so it can never drain in one pass. The eligibility query
    returns candidates in descending summed-similarity order, so the ones
    that can actually reach a deck are enriched first.
  - candidateBatch (50) is smaller than the track batch (200): tracks are
    finite and drain to completion, candidates are effectively unbounded
    and would otherwise starve the track arm forever.

GC sweeps both tables — the similarity feed churns, and a candidate that
joins the library has its tags in track_tags now. Tags swept before state
so a mid-sweep crash leaves a valid state, not a re-fetch loop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:04:15 -04:00

126 lines
5.1 KiB
Go

// Package dbtest provides shared helpers for integration tests that need
// a clean Postgres state without disturbing the operator's admin user.
//
// Background: when integration tests run against the same Postgres
// instance an operator uses for local dev (a common dev-laptop setup),
// a TRUNCATE that includes the users table wipes the operator's admin
// login between every test run. To avoid that, tests should:
//
// 1. Call ResetDB instead of issuing TRUNCATE statements directly. It
// truncates every data table EXCEPT users, then deletes only those
// users whose username starts with TestUserPrefix.
// 2. Always create test user rows with a username that begins with
// TestUserPrefix; otherwise those rows survive across runs and
// leak into other tests.
//
// One test legitimately needs to wipe the entire users table:
// internal/auth/bootstrap_test.go, which exercises first-time admin
// bootstrap. That file does its own TRUNCATE and does not use this
// helper.
package dbtest
import (
"context"
"strings"
"testing"
"github.com/jackc/pgx/v5/pgxpool"
)
// TestUserPrefix is the required username prefix for any user row
// created by an integration test. ResetDB removes every user matching
// this prefix and leaves all others intact.
const TestUserPrefix = "test-"
// dataTables is the union of every non-users table touched by any
// integration test in the repo. Truncating the union is harmless for
// callers that only care about a subset; truncating tables that don't
// exist would error, but every name here corresponds to a migration
// that has shipped.
var dataTables = []string{
"artist_similarity",
"track_similarity",
"artist_similarity_unmatched", // M5c
"scrobble_queue",
"contextual_likes",
"general_likes_albums",
"general_likes_artists",
"general_likes",
"play_events",
"skip_events",
"play_sessions",
"sessions",
"lidarr_quarantine_actions",
"lidarr_quarantine",
// #2374. Keyed by (user_id, candidate_mbid) with no FK to artists —
// candidates are out-of-library — so the CASCADE from artists/users
// does NOT reach it for a leftover row whose user survived. Truncate
// explicitly or a stale snooze silently hides a candidate from the
// next test's suggestion assertions.
"suggestion_snoozes",
// #2376. Same reasoning: keyed by candidate MBID with no FK anywhere,
// so nothing cascades to them. A leftover tag row would make a
// candidate look enriched to the next test, and a leftover state row
// would make it look already-settled and thus ineligible.
"candidate_artist_tags",
"candidate_artist_tag_state",
"playlist_tracks",
"playlists",
"library_changes", // M7 #357 — must reset to keep cursor isolated per test
// SettingsService.reconcile() idempotently re-UpsertProviderSettings
// for every registered provider at boot, so truncating this is the
// correct per-test reset (clears test-modified enabled/api_key rows).
// cover_art_sources_meta is NOT truncated — boot only READS it
// (never recreates the singleton, seeded once by 0018); ResetDB
// resets its counter via UPDATE below instead.
"cover_art_provider_settings",
// Same reasoning for the tag-enrichment settings (#1490): truncate the
// per-provider rows (reconcile re-seeds registered ones), and reset the
// tag_sources_meta counter via UPDATE below rather than truncating it.
"tag_provider_settings",
// recsettings.New reconciles shipped defaults on every construction
// (#1250), so truncating gives each test pristine tuning values.
"recommendation_weight_profiles",
"taste_tuning",
"recommendation_tuning_audit",
"tracks",
"albums",
"artists",
}
// ResetDB clears all data tables and removes any user whose username
// begins with TestUserPrefix. Users without that prefix (notably the
// operator's admin row) are left alone. Calls t.Fatalf on error.
func ResetDB(t *testing.T, pool *pgxpool.Pool) {
t.Helper()
ctx := context.Background()
stmt := "TRUNCATE " + strings.Join(dataTables, ", ") + " RESTART IDENTITY CASCADE"
if _, err := pool.Exec(ctx, stmt); err != nil {
t.Fatalf("dbtest.ResetDB truncate: %v", err)
}
if _, err := pool.Exec(ctx,
"DELETE FROM users WHERE username LIKE $1",
TestUserPrefix+"%",
); err != nil {
t.Fatalf("dbtest.ResetDB delete test users: %v", err)
}
// Reset the monotonic cover-art source-version counter to its
// post-migration seeded value. Truncating the row would break
// SettingsService boot, which reads (never recreates) this
// singleton; an UPDATE keeps the row while clearing cross-test
// version accumulation (CurrentVersion=4 want 1, key-only-bump).
if _, err := pool.Exec(ctx,
"UPDATE cover_art_sources_meta SET current_version = 1",
); err != nil {
t.Fatalf("dbtest.ResetDB reset cover-art version: %v", err)
}
// Same for the tag-sources counter (#1490) — clears cross-test version
// accumulation that would spuriously report version_bumped on a
// key-only change.
if _, err := pool.Exec(ctx,
"UPDATE tag_sources_meta SET current_version = 1, last_registered_providers_hash = ''",
); err != nil {
t.Fatalf("dbtest.ResetDB reset tag-sources version: %v", err)
}
}