799dab029a0f164c92dd3947384ee570c0efaf95
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
799dab029a |
feat(discover): rank suggestions by taste-tag overlap — #2377 (server)
The payoff slice. Until now a candidate's only claim on a slot was "some
artist you play is adjacent to it in a similarity graph" — a fact that says
nothing about whether the music sounds like anything you like. Now the
candidate's own folksonomy tags (cached by slice 5) are compared against
the user's taste-profile tags, so the deck ranks on taste and can say WHY.
The blend is MULTIPLICATIVE — score × (1 + weight × overlap) — and that
choice carries the whole safety argument:
- An untagged candidate has overlap 0, so its score is EXACTLY unchanged.
Tag coverage is permanently partial (#2376); it must cost a candidate
nothing, not sink it (rule #131).
- Nothing can leapfrog on tags alone. An additive term with a large
weight would let a near-zero-similarity artist outrank a strong match
for sharing one popular tag, which reads as noise.
- Weight 0 restores pure similarity order bit-for-bit, so the operator's
knob has a real off position.
overlap = Σ(shared) candWeight × normalizedTasteWeight ÷ Σ(all) candWeight.
Normalizing the taste side by the user's strongest tag makes the score
comparable across users (taste weights accumulate with listening, so a
heavy listener's raw numbers dwarf a new user's while meaning the same
thing). Dividing by the candidate's own mass makes it comparable across
candidates, so a densely-tagged artist can't win on tag count alone.
Applied to the whole over-fetched pool BEFORE selectSuggestions, so the
rotation and diversity rules operate on blended scores — boosting only the
twelve already chosen by similarity would leave the re-ranking undone.
A query failure is returned, NOT degraded past. Graceful degradation is
for expected absence (no taste profile, no cached tags) and both are
handled explicitly as empty inputs; swallowing a real error would hide a
broken DB behind a subtly worse ranking that nothing reports.
Migration 0051 adds a FOURTH tuning scope rather than columns on
taste_tuning, because snooze_days lives here too and a snooze must never
be read as taste signal (#2374) — filing it under 'taste' would put it one
careless join from the leak that design forbids. Expanding
recommendation_tuning_audit's CHECK is in the same migration per rule #36,
and a test asserts the audit row lands, which is what would catch its
absence.
snooze_days moves out of a Go constant onto the tuning card (rule #25),
closing the deferral from #2374.
Tag-overlap tests use deliberately SKEWED fixtures: an evenly-matching pool
cannot exercise a re-ranking, since every candidate gets the same
multiplier and the order is unchanged whether the blend works or not.
Admin UI + client attribution follow in this batch — rule #27.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b27029f674 |
feat(discover): rotate the suggestion deck daily + cap one seed's share — #2373
Second half of the reported symptom: suggestions "show the same artists until you request one". The ranking was `ORDER BY total_score DESC` with no randomization and no seen-state, so the only things that could ever change the deck were a candidate entering the library or the user filing a request. The tail of the ranking was unreachable — requesting was literally the only lever. No SQL change was needed. The query already takes a limit, so it over-fetches a pool (4x the slots, capped at 60) and the selection moves to Go, where it is a pure function of (pool, limit, day) — no DB, no clock — and therefore unit testable in the fast lane instead of behind the integration gate. Three rules. The best few by score always lead, so the strongest matches never rotate out of sight (For You's head/tail shape). The remaining slots are drawn by md5(mbid + day), the same daily-stable idiom the Home rows already use: stable within a day so pull-to-refresh doesn't reshuffle, different tomorrow, and no stored state. And a per-seed cap keeps roughly a quarter of the deck attributable to any one seed artist, so twelve neighbours of a single artist can't be the whole surface. The cap is a preference, not a quota. A user whose pool hangs off one or two seeds would otherwise get a three-card surface — worse than the monoculture being avoided, and exactly the vanish-or-nothing shape rule #131 exists to prevent — so a short deck tops up in score order from what the cap set aside. This is also what keeps the existing Top12Cap integration test honest: its 30 candidates share one seed, and without the top-up it would return 3. Eight unit tests, including one that had to be rewritten mid-change: the first version asserted the cap against an evenly-spread pool, where the top-N is already diverse and the assertion could not fail. It now uses a skewed pool where one seed owns the entire top of the ranking, which is the only shape that actually exercises a cap. Dropped two //nolint:gosec directives added in passing — gosec isn't in .golangci.yml, so they suppressed nothing and only implied a check that runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
14aa22198f |
feat(discover): seed request suggestions from the taste profile — #2372
The Discover request surface was the one recommendation surface still on its M5c implementation from early May. #796's taste profile, #1488's taste_unheard bucket and #1490's folksonomy enrichment all modernized in-library surfaces; this one was never in scope for any of them, so it still projected raw likes + plays through artist_similarity_unmatched. Two defects fall out of that signal, `5*liked + Σexp(-age/halflife)` summed over every play of the artist. It is unbounded, and contribution is signal × similarity — so a handful of heavily-played artists monopolize all twelve slots, and their share GROWS the more the user listens. The surface entrenched harder the better it knew you, which is exactly backwards and matches the reported "goes stale once it has a strong signal of your taste". It also counted every play_event with no was_skipped filter, so skipping an artist repeatedly INCREASED its signal and pushed more of its neighbours at the user. ListMostPlayedTracksForUser and the taste engine both filter skips; this query was the odd one out. Seeds now come from taste_profile_artists.weight, which the taste engine has already engagement-graded, time-decayed and signed — an artist the user drifted away from stops contributing instead of accumulating forever, and can even contribute negatively. Tiered per rule #131 rather than hard-switched: tier 1 is the profile, tier 2 is likes + completed plays for a user who has no profile rows yet (new account, or before the first daily recompute), so the surface never empties. The old unfiltered-play signal is gone, not kept behind a toggle. The signal is also log-damped, so one artist cannot take every slot even when its weight dwarfs the rest. $2 stays wired to the tier-2 decay: it is genuinely still used there, and dropping the parameter would have changed the generated signature. sqlc's image is not on this workstation and the change preserves the query signature exactly — same three params, same seven columns — so only the embedded SQL const moves. Both copies are edited and verified byte-identical rather than pulling a container onto the operator's machine; a malformed query fails the integration lane loudly, which is the real check either way. Four integration tests cover what changed: a taste weight alone seeds with no like or play; tier 2 does not run alongside tier 1; a non-positive weight never seeds (guarded by a second positive row, so an empty tier 1 can't make it pass for the wrong reason); and skip-only history seeds nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
277898a49a |
feat(recommendation): SuggestArtists service for M5c
Add per-user artist-suggestion service ranking out-of-library MBIDs by signal x similarity. Single-CTE SQL collects user likes (5x weight) and recency-decayed plays, joins against artist_similarity_unmatched, and filters in-library candidates plus non-terminal lidarr_requests. The service resolves top-3 attribution seeds to artist names in a batched GetArtistsByIDs call so the UI can render "because you liked X" reasons. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |