Files
minstrel/internal/recommendation/candidates.go
T
bvandeusenandClaude Opus 5 f367eeaa9d
test-web / test (push) Successful in 1m7s
test-go / test (push) Successful in 1m31s
test-go / integration (push) Failing after 4m21s
release / Build signed APK (releases and dev) (push) Successful in 5m11s
release / Build + push container image (push) Successful in 1m52s
release / Verify release artifacts (tag releases only) (push) Skipped
fix(recommendation): Songs-like gets its own profile so it stops wandering
Operator, 2026-09-10: "when I play it I'm expecting to get a consistent
sound and style from the experience... I was getting a seeming wide variety
of music from each one when I was hoping to stay in a certain neighborhood."

Songs-like shared the `daily_mix` weight profile with For-You, and that
sharing WAS the bug. The two surfaces want opposite things: For-You answers
"what will they enjoy today" and is supposed to roam; Songs-like answers
"what sounds like THIS". Under one profile the broad answer wins.

The arithmetic, from the shared weights:

    unrelated track, liked, not played recently → 1.0 + 2.0 + 1.0 = 4.0
    PERFECT similarity match, not liked         → 1.0 + 1.5       = 2.5

Liking something outranked sounding like the seed, because LikeBoost (2.0)
exceeded SimilarityWeight's whole range (1.5) and TasteWeight (1.5, and
seed-INDEPENDENT) matched it outright. Under the new profile the same pair
scores 5.00 vs 2.00.

Two levers, because either alone leaves the other's failure intact:

POOL. Songs-like now takes its own CandidateSourceLimits. The default gave
~29% of candidates a sim_score of literally zero — `taste_overlap` and
`random_fill` are both `0.0::float8` in recommendation.sql, seed-independent
by construction. Same total pool size; composition shifts to arms that
measure distance from the seed, LBSimilar doubled.

WEIGHTS. A third profile beside radio and daily_mix, DB-backed and live per
rule 25, with the property that similarity's range exceeds the combined
range of every seed-independent differentiator — so a closer match cannot
be beaten on likes, freshness and taste alone, while tracks within ~0.39
similarity of each other still get ordered by what the user likes.

Rule 131 changed the pool design mid-way and for the better. Zeroing the
two seed-independent arms was the first instinct and is exactly the
vanish-or-nothing shape that rule forbids: a seed with thin ListenBrainz
coverage would yield a short mix or none. They are the tier-3 FLOOR — cut
hard, never removed — and the weights keep them at the bottom of the
ranking rather than out of the pool. "A few tracks further from the seed
than we'd like" beats "no playlist".

Caught while wiring it: switching only pickTopN's final Score would have
been nearly INERT. scoreAndSortCandidates does the selection sort, and the
caller caps and truncates in that order — so the playlist would still have
been chosen by daily_mix and merely relabelled with songs_like numbers. It
now takes the profile as a parameter, and each surface passes its own.

Also corrects the daily_mix card's blurb, which claimed Songs-like as one
of its surfaces and no longer is.

Guards pin behaviour rather than the numbers, since numbers get retuned:
that similarity beats an unrelated liked track, that daily_mix still
DOESN'T (or the split buys nothing), that the tier-3 floor is non-zero,
and that the UI card shows its own values rather than falling back. Each
falsified against its named regression first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SQ31KQpYbStyK5y58UmPLH
2026-09-10 20:51:15 -04:00

303 lines
10 KiB
Go
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
package recommendation
import (
"context"
"encoding/json"
"time"
"github.com/jackc/pgx/v5/pgtype"
"git.fabledsword.com/bvandeusen/minstrel/internal/db/dbq"
"git.fabledsword.com/bvandeusen/minstrel/internal/mood"
)
// LoadCandidates fetches the candidate pool for radio scoring. Combines
// the existing track+stats query with a one-shot bulk fetch of the user's
// active contextual_likes, mapping each candidate to its max similarity
// against currentVector. Pass currentVector with Seed=true to short-circuit
// the contextual term to 0 (cold-start path).
func LoadCandidates(
ctx context.Context,
q *dbq.Queries,
userID, seedID pgtype.UUID,
recentlyPlayedHours int,
currentVector SessionVector,
) ([]Candidate, error) {
rows, err := q.LoadRadioCandidates(ctx, dbq.LoadRadioCandidatesParams{
UserID: userID,
ID: seedID,
Column3: float64(recentlyPlayedHours),
})
if err != nil {
return nil, err
}
likes, err := loadContextualLikesByTrack(ctx, q, userID)
if err != nil {
return nil, err
}
profile, err := LoadTasteProfile(ctx, q, userID)
if err != nil {
return nil, err
}
affinity, err := LoadContextAffinity(ctx, q, userID, currentVector.DeviceClass)
if err != nil {
return nil, err
}
out := make([]Candidate, 0, len(rows))
for _, r := range rows {
var lpt *time.Time
if r.LastPlayedAt.Valid {
t := r.LastPlayedAt.Time
lpt = &t
}
ctxScore := ContextualMatchScore(currentVector, likes[r.Track.ID], DefaultSimilarityWeights)
out = append(out, Candidate{
Track: r.Track,
Inputs: ScoringInputs{
IsGeneralLiked: r.IsLiked,
LastPlayedAt: lpt,
PlayCount: int(r.PlayCount),
SkipCount: int(r.SkipCount),
ContextualMatchScore: ctxScore,
// Fallback path: mood is scored only in the primary
// (similarity) loader — loading per-candidate tags over this
// near-whole-library pool isn't worth it (nil moods → 0).
TasteMatchScore: profile.Match(r.Track.ArtistID, r.Track.Genre, r.ReleaseDate, nil),
ContextAffinityScore: affinity.Affinity(r.Track.ArtistID),
},
})
}
return out, nil
}
// CandidateSourceLimits controls per-source K values for the M4c
// similarity-driven pool. Defaults via DefaultCandidateSourceLimits().
type CandidateSourceLimits struct {
LBSimilar int
SimilarArtist int
TagOverlap int
LikesOverlap int
RandomFill int
// TasteOverlap (#796 phase 2b): tracks by the user's top positively-
// weighted taste-profile artists. 0 disables the arm (e.g. cold-start
// users have an empty profile, so it contributes nothing anyway).
TasteOverlap int
// UserCoplay (#1533): tracks by artists co-played across the instance
// with the seed's artist (source='user_cooccurrence'). Empty on
// single-user servers, so it contributes nothing there.
UserCoplay int
}
// DefaultCandidateSourceLimits returns the v1 hardcoded constants per spec.
func DefaultCandidateSourceLimits() CandidateSourceLimits {
return CandidateSourceLimits{
LBSimilar: 30,
SimilarArtist: 30,
TagOverlap: 20,
LikesOverlap: 20,
RandomFill: 30,
TasteOverlap: 20,
UserCoplay: 20,
}
}
// SongsLikeCandidateSourceLimits is the pool shape for "Songs like {X}"
// (#3881). Same total size as the default (~170) — the composition is what
// changes, shifted hard toward arms that actually measure distance from the
// seed.
//
// The surface answers "what sounds like THIS", and it shared the default
// pool with For-You, which answers the much broader "what will they enjoy
// today". Under the default, 50 of ~170 candidates carried sim_score = 0 by
// construction — `taste_overlap` (tracks by the user's top taste artists) and
// `random_fill` (literally any track not already in the pool), both of which
// are seed-INDEPENDENT. Nearly a third of the pool had no relationship to the
// seed at all, and the operator saw it: "I was getting a seeming wide variety
// of music from each one when I was hoping to stay in a certain neighborhood."
//
// TIERED, per rule 131 — a system mix degrades, it never vanishes:
//
// tier 1 lb_similar real track-level similarity. The exact promise.
// tier 2 similar_artist, tag_overlap, coplay, likes_overlap
// seed-RELATED but weaker signal.
// tier 3 taste_overlap, random_fill
// seed-independent. The floor, and nothing more.
//
// The tiering is enforced by SCORE rather than by a fallback ladder: tier-3
// arms carry sim_score 0, so under SongsLike weights (SimilarityWeight 4.0,
// everything seed-independent demoted) they rank below any real match and
// surface only when tiers 12 cannot fill the mix. That is the rule's
// "fill from tier 1 first, reach down only when a tier cannot fill".
//
// Which is exactly why tier 3 is REDUCED rather than removed. Zeroing those
// two arms was the first instinct and it is the vanish-or-nothing shape rule
// 131 exists to forbid: a seed whose artist has thin ListenBrainz coverage
// would produce a short mix or none at all, and "no playlist" is a worse
// answer than "a few tracks further from the seed than we would like".
//
// likes_overlap is cut hardest of the tier-2 arms for a specific reason: its
// SQL assigns a FLAT 0.6 sim_score (recommendation.sql:108) rather than
// measuring anything. It is a collaborative signal wearing similarity's
// clothes, and raising SimilarityWeight amplifies it — if real ListenBrainz
// scores commonly land below 0.6 it would outrank genuine matches. Halved
// here pending the fill-rate measurement in #3879; the honest fix is to stop
// it claiming a similarity score it did not compute.
func SongsLikeCandidateSourceLimits() CandidateSourceLimits {
return CandidateSourceLimits{
LBSimilar: 60, // tier 1 — doubled; the only arm that measures the seed
SimilarArtist: 40, // tier 2
TagOverlap: 20, // tier 2
UserCoplay: 20, // tier 2
LikesOverlap: 10, // tier 2, halved — flat 0.6, see above
TasteOverlap: 10, // tier 3 floor — halved, not removed
RandomFill: 10, // tier 3 floor — cut hard, never to zero
}
}
// LoadCandidatesFromSimilarity is M4c's primary candidate-pool loader.
// 5-way SQL UNION (LB-similar / similar-artist tracks / MB-tag overlap /
// likes-overlap / random fill) + dedup-by-max sim_score. Returns
// []Candidate (same shape as LoadCandidates) so Shuffle is unchanged.
//
// Caller (radio handler) falls back to LoadCandidates on error.
func LoadCandidatesFromSimilarity(
ctx context.Context,
q *dbq.Queries,
userID, seedID pgtype.UUID,
recentlyPlayedHours int,
currentVector SessionVector,
exclude []pgtype.UUID,
limits CandidateSourceLimits,
) ([]Candidate, error) {
if exclude == nil {
exclude = []pgtype.UUID{}
}
rows, err := q.LoadRadioCandidatesV2(ctx, dbq.LoadRadioCandidatesV2Params{
UserID: userID,
ID: seedID,
Column3: int64(recentlyPlayedHours),
Column4: exclude,
Limit: int32(limits.LBSimilar),
Limit_2: int32(limits.SimilarArtist),
Limit_3: int32(limits.TagOverlap),
Limit_4: int32(limits.LikesOverlap),
Limit_5: int32(limits.RandomFill),
Limit_6: int32(limits.TasteOverlap),
Limit_7: int32(limits.UserCoplay),
})
if err != nil {
return nil, err
}
likes, err := loadContextualLikesByTrack(ctx, q, userID)
if err != nil {
return nil, err
}
profile, err := LoadTasteProfile(ctx, q, userID)
if err != nil {
return nil, err
}
affinity, err := LoadContextAffinity(ctx, q, userID, currentVector.DeviceClass)
if err != nil {
return nil, err
}
trackIDs := make([]pgtype.UUID, 0, len(rows))
for _, r := range rows {
trackIDs = append(trackIDs, r.Track.ID)
}
moods, err := loadCandidateMoods(ctx, q, trackIDs)
if err != nil {
return nil, err
}
out := make([]Candidate, 0, len(rows))
for _, r := range rows {
var lpt *time.Time
if r.LastPlayedAt.Valid {
t := r.LastPlayedAt.Time
lpt = &t
}
// sqlc returns SimilarityScore as interface{} (couldn't infer the
// type through max(...) over a UNION). Type-assert; default to 0
// on the (impossible-but-defensive) case where it's nil/wrong type.
var simScore float64
if v, ok := r.SimilarityScore.(float64); ok {
simScore = v
}
ctxScore := ContextualMatchScore(currentVector, likes[r.Track.ID], DefaultSimilarityWeights)
out = append(out, Candidate{
Track: r.Track,
Inputs: ScoringInputs{
IsGeneralLiked: r.IsLiked,
LastPlayedAt: lpt,
PlayCount: int(r.PlayCount),
SkipCount: int(r.SkipCount),
ContextualMatchScore: ctxScore,
SimilarityScore: simScore,
TasteMatchScore: profile.Match(
r.Track.ArtistID, r.Track.Genre, r.ReleaseDate, moods[r.Track.ID]),
ContextAffinityScore: affinity.Affinity(r.Track.ArtistID),
},
})
}
return out, nil
}
// loadCandidateMoods fetches the enriched tags for the given candidate tracks
// and reduces each to its canonical mood buckets (internal/mood, #1534), so the
// scorer can apply the mood facet per candidate. Tracks with no mood-word tags
// are absent from the map (→ no mood signal). Empty input short-circuits.
func loadCandidateMoods(
ctx context.Context, q *dbq.Queries, trackIDs []pgtype.UUID,
) (map[pgtype.UUID][]string, error) {
if len(trackIDs) == 0 {
return map[pgtype.UUID][]string{}, nil
}
rows, err := q.ListTrackTagsForTracks(ctx, trackIDs)
if err != nil {
return nil, err
}
tagsByTrack := make(map[pgtype.UUID][]string)
for _, r := range rows {
tagsByTrack[r.TrackID] = append(tagsByTrack[r.TrackID], r.Tag)
}
out := make(map[pgtype.UUID][]string, len(tagsByTrack))
for id, tags := range tagsByTrack {
if m := mood.Of(tags); len(m) > 0 {
out[id] = m
}
}
return out, nil
}
// loadContextualLikesByTrack fetches the user's active contextual_likes in
// one query and groups them by track_id. Rows whose session_vector fails
// to unmarshal are skipped with no error (don't poison scoring over one
// bad row); the SQL query already filters NULL vectors.
func loadContextualLikesByTrack(
ctx context.Context,
q *dbq.Queries,
userID pgtype.UUID,
) (map[pgtype.UUID][]SessionVector, error) {
rows, err := q.ListActiveContextualLikesForUser(ctx, userID)
if err != nil {
return nil, err
}
out := make(map[pgtype.UUID][]SessionVector, len(rows))
for _, r := range rows {
var v SessionVector
if err := json.Unmarshal(r.SessionVector, &v); err != nil {
continue
}
out[r.TrackID] = append(out[r.TrackID], v)
}
return out, nil
}