feat: duplicates resolve themselves where Lidarr says it is safe (M498)
release / govulncheck (push) Successful in 45s
release / web (push) Successful in 1m23s
release / go (push) Successful in 1m39s
release / integration (push) Successful in 4m25s
release / android (push) Successful in 6m17s
release / Build signed APK (releases and dev) (push) Successful in 5m57s
release / Attach APK to the Release (tag releases only) (push) Skipped
release / Build + push container image (push) Successful in 1m54s
release / Verify release artifacts (tag releases only) (push) Skipped
release / govulncheck (push) Successful in 45s
release / web (push) Successful in 1m23s
release / go (push) Successful in 1m39s
release / integration (push) Successful in 4m25s
release / android (push) Successful in 6m17s
release / Build signed APK (releases and dev) (push) Successful in 5m57s
release / Attach APK to the Release (tag releases only) (push) Skipped
release / Build + push container image (push) Successful in 1m54s
release / Verify release artifacts (tag releases only) (push) Skipped
The duplicate sweep proposed 4,197 groups and every one waited for the operator. Most are safe to settle, and Lidarr defines what safe means: it maps one file to each track of the release it monitors and downloads any mapped file that disappears. Deleting a mapped copy opens exactly the hole the operator saw Lidarr fill. Classify (#5435) - Migration 0075: duplicate_groups.class (same_release, cross_release, mismatch, review), resolve_note, resolved_automatically; duplicate_group_members.lidarr_state (tracked, unmapped); fingerprint_settings.auto_resolve; notification kind duplicates_resolved with both kind CHECKs swapped (rule 36). - library.ClassifyDuplicateGroup, with MatchTitleKey dropping featuring credits, remaster notes and video-rip markers, and keeping live, demo, remix and instrumental. The rip markers move from api to library. Choose the copy to keep (#5436) - ProposeSurvivor ranks the copy Lidarr maps first, then tag fit (a clash-free track number, no rip marker in the name, an MBID), then the quality rules. File size picked the wrong Humanz copy in 6 of 21 groups. Act (#5437) - An hourly resolver pass reads Lidarr's unmapped files, matched by the last three path components, and records each copy's state. - Same album, with at most one copy mapped: merged into the mapped copy. The merge is guarded, so a mapped copy can never be removed (MergeDuplicateGroupGuarded, ErrCopyTrackedByLidarr). - Same album, every copy mapped: the monitored release lists the song twice (Humanz's 14x12" box set). The pass moves Lidarr to the release that lists each song once and best covers what is on disk. It never picks one covering less, and is capped at 10 albums per pass. - Fixed point (lesson #4183): the chosen release no longer repeats. - The album is left alone for 24h while Lidarr rescans, so "every copy unmapped" mid-rescan is never read as licence to merge. - Both actions are audited with no actor and summarised to admins. The operator can switch them off in the Fingerprinting card (rule 25). - Manual merges use the same guard: 409 copy_tracked_by_lidarr, or 503 lidarr_unavailable when Lidarr cannot say. Web - Duplicates gets tabs: Needs review, Across releases, Resolved automatically. Each loads as you scroll (rule 172), replacing the pager. - Each copy says whether Lidarr uses it. - The resolver's note shows on each group. - The merge confirm blocks, before sending, a merge that would remove the copy Lidarr uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,143 @@
|
||||
package library
|
||||
|
||||
import (
|
||||
"regexp"
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
// Duplicate classification (M498 #5435).
|
||||
//
|
||||
// A duplicate group is one recording found more than once. What should happen
|
||||
// to it depends on where the copies sit, and the measured library (milestone
|
||||
// 498's body) splits into four shapes:
|
||||
//
|
||||
// - same_release: every copy is on one album. An album holds a song once, so
|
||||
// the extra copies are the ones to fold away — when Lidarr does not need
|
||||
// them (the resolver checks).
|
||||
// - cross_release: the copies are on different albums under one title. A
|
||||
// single and the album it came from, a deluxe edition and the standard one:
|
||||
// each copy fulfils its own release in Lidarr, so none is deleted.
|
||||
// - mismatch: byte-identical audio on different albums under different
|
||||
// titles. One file carries another song's tags — a wrong-file import.
|
||||
// - review: anything else. The operator decides.
|
||||
|
||||
// DuplicateClass is a group's shape. The values are duplicate_groups.class.
|
||||
type DuplicateClass string
|
||||
|
||||
const (
|
||||
ClassSameRelease DuplicateClass = "same_release"
|
||||
ClassCrossRelease DuplicateClass = "cross_release"
|
||||
ClassMismatch DuplicateClass = "mismatch"
|
||||
ClassReview DuplicateClass = "review"
|
||||
)
|
||||
|
||||
// ClassifyMember is what classifying needs to know about one copy.
|
||||
type ClassifyMember struct {
|
||||
AlbumID string
|
||||
Title string
|
||||
ArtistName string
|
||||
}
|
||||
|
||||
// ClassifyDuplicateGroup decides a group's shape from its tier ("exact" or
|
||||
// "acoustic") and its copies.
|
||||
//
|
||||
// Titles are compared by MatchTitleKey, which drops what a release or an upload
|
||||
// adds to a title (featuring credits, remaster notes, video markers) and keeps
|
||||
// what names a different recording (live, demo, remix, instrumental). So one
|
||||
// album holding "WWW" and "WWW (instrumental)" as an acoustic match is review,
|
||||
// not a copy to fold away: the #3885 case, where they were in fact different
|
||||
// recordings mislabelled. Byte-identical copies on one album are the same file
|
||||
// twice whatever their tags say.
|
||||
func ClassifyDuplicateGroup(tier string, members []ClassifyMember) DuplicateClass {
|
||||
if len(members) < 2 {
|
||||
return ClassReview
|
||||
}
|
||||
oneAlbum, oneTitle := true, true
|
||||
key := MatchTitleKey(members[0].Title, members[0].ArtistName)
|
||||
for _, m := range members[1:] {
|
||||
if m.AlbumID != members[0].AlbumID {
|
||||
oneAlbum = false
|
||||
}
|
||||
if MatchTitleKey(m.Title, m.ArtistName) != key {
|
||||
oneTitle = false
|
||||
}
|
||||
}
|
||||
exact := tier == string(tierExact)
|
||||
switch {
|
||||
case oneAlbum && (oneTitle || exact):
|
||||
return ClassSameRelease
|
||||
case oneAlbum:
|
||||
return ClassReview
|
||||
case oneTitle:
|
||||
return ClassCrossRelease
|
||||
case exact:
|
||||
return ClassMismatch
|
||||
default:
|
||||
return ClassReview
|
||||
}
|
||||
}
|
||||
|
||||
var (
|
||||
// A bracketed credit: "(feat. Panda Bear)", "[ft. X]", "(featuring Y)".
|
||||
featBracket = regexp.MustCompile(`(?i)[(\[][^)\]]*\b(feat\.?|ft\.?|featuring)\s[^)\]]*[)\]]`)
|
||||
// A credit trailing the title without brackets: "Doin' It Right ft. Panda Bear".
|
||||
featTrailing = regexp.MustCompile(`(?i)\s(feat\.?|ft\.?|featuring)\s.*$`)
|
||||
// "(2011 Remaster)", "[Remastered]", "- Remastered 2009".
|
||||
remasterBracket = regexp.MustCompile(`(?i)[(\[][^)\]]*remaster[^)\]]*[)\]]`)
|
||||
remasterSuffix = regexp.MustCompile(`(?i)\s-\s[^-]*remaster.*$`)
|
||||
// Any bracketed segment, to test each against the video-rip markers.
|
||||
bracketed = regexp.MustCompile(`[(\[][^)\]]*[)\]]`)
|
||||
)
|
||||
|
||||
// MatchTitleKey reduces a title to what identifies the song, for deciding
|
||||
// whether two copies carry the same one.
|
||||
//
|
||||
// Dropped: a leading "Artist - " (a video upload's title), bracketed featuring
|
||||
// credits and remaster notes, and bracketed segments carrying a video-rip marker
|
||||
// ("(Official Video)", "[HD]"). Kept: every other bracketed word, so "(live)",
|
||||
// "(demo)", "(acoustic)", "(remix)" and "(instrumental)" still tell recordings
|
||||
// apart. Then case, "&" against "and", and punctuation and spacing are ignored.
|
||||
func MatchTitleKey(title, artist string) string {
|
||||
t := strings.TrimSpace(title)
|
||||
if a := strings.TrimSpace(artist); a != "" {
|
||||
if prefix := a + " - "; len(t) > len(prefix) && strings.EqualFold(t[:len(prefix)], prefix) {
|
||||
t = t[len(prefix):]
|
||||
}
|
||||
}
|
||||
t = featBracket.ReplaceAllString(t, " ")
|
||||
t = remasterBracket.ReplaceAllString(t, " ")
|
||||
t = bracketed.ReplaceAllStringFunc(t, func(seg string) string {
|
||||
for _, re := range suspectSourceRegexps {
|
||||
if re.MatchString(seg) {
|
||||
return " "
|
||||
}
|
||||
}
|
||||
return seg
|
||||
})
|
||||
t = remasterSuffix.ReplaceAllString(t, "")
|
||||
t = featTrailing.ReplaceAllString(t, "")
|
||||
if key := alnumKey(t); key != "" {
|
||||
return key
|
||||
}
|
||||
return alnumKey(title)
|
||||
}
|
||||
|
||||
// ReleaseTitleKey is the lighter comparison for asking whether one release
|
||||
// lists a song twice: case, "&" and punctuation only. Every word stays, so a
|
||||
// box set's "Song" and "Song (demo)" are two songs, as the release means them.
|
||||
func ReleaseTitleKey(title string) string {
|
||||
return alnumKey(title)
|
||||
}
|
||||
|
||||
// alnumKey lowercases, reads "&" as "and", and keeps only letters and digits.
|
||||
func alnumKey(s string) string {
|
||||
s = strings.ReplaceAll(strings.ToLower(s), "&", " and ")
|
||||
var b strings.Builder
|
||||
for _, r := range s {
|
||||
if unicode.IsLetter(r) || unicode.IsDigit(r) {
|
||||
b.WriteRune(r)
|
||||
}
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
package library
|
||||
|
||||
import "testing"
|
||||
|
||||
// The shapes measured on the operator's library (milestone 498's body).
|
||||
func TestClassifyDuplicateGroup(t *testing.T) {
|
||||
m := func(album, title string) ClassifyMember {
|
||||
return ClassifyMember{AlbumID: album, Title: title, ArtistName: "Gorillaz"}
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
tier string
|
||||
members []ClassifyMember
|
||||
want DuplicateClass
|
||||
}{
|
||||
{"Humanz: one album, one song twice", "exact",
|
||||
[]ClassifyMember{m("humanz", "Saturnz Barz"), m("humanz", "Saturnz Barz (Official Video)")}, ClassSameRelease},
|
||||
{"one album, video upload title with artist prefix", "acoustic",
|
||||
[]ClassifyMember{m("humanz", "Andromeda"), m("humanz", "Gorillaz - Andromeda (Official Audio)")}, ClassSameRelease},
|
||||
{"one album, identical bytes under different titles", "exact",
|
||||
[]ClassifyMember{m("www", "WWW"), m("www", "WWW (instrumental)")}, ClassSameRelease},
|
||||
{"one album, acoustic match, different recordings named", "acoustic",
|
||||
[]ClassifyMember{m("www", "WWW"), m("www", "WWW (instrumental)")}, ClassReview},
|
||||
{"Big Me: the single and the album", "exact",
|
||||
[]ClassifyMember{m("single", "Big Me"), m("album", "Big Me")}, ClassCrossRelease},
|
||||
{"remaster and featuring credits are one song", "acoustic",
|
||||
[]ClassifyMember{m("a", "Feel Good Inc. (2017 Remaster)"), m("b", "Feel Good Inc feat. De La Soul")}, ClassCrossRelease},
|
||||
{"Nervosa / Anorexia: identical bytes, different songs", "exact",
|
||||
[]ClassifyMember{m("a", "Nervosa"), m("b", "Anorexia")}, ClassMismatch},
|
||||
{"studio and live across albums", "acoustic",
|
||||
[]ClassifyMember{m("studio", "Charger"), m("live", "Charger (live)")}, ClassReview},
|
||||
{"a lone copy", "exact", []ClassifyMember{m("a", "Song")}, ClassReview},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := ClassifyDuplicateGroup(tc.tier, tc.members); got != tc.want {
|
||||
t.Errorf("%s: got %s, want %s", tc.name, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestMatchTitleKey(t *testing.T) {
|
||||
cases := []struct{ title, artist, want string }{
|
||||
{"Saturnz Barz (Official Video)", "Gorillaz", "saturnzbarz"},
|
||||
{"Gorillaz - Saturnz Barz", "Gorillaz", "saturnzbarz"},
|
||||
{"Doin' It Right (Music Video) ft. Panda Bear", "Daft Punk", "doinitright"},
|
||||
{"Song [feat. X]", "", "song"},
|
||||
{"Song - Remastered 2009", "", "song"},
|
||||
{"Rock & Roll", "", "rockandroll"},
|
||||
// What names a different recording stays.
|
||||
{"Song (live)", "", "songlive"},
|
||||
{"Song (demo)", "", "songdemo"},
|
||||
{"Song (Remix)", "", "songremix"},
|
||||
// A title that is nothing but a marker keeps its words rather than vanishing.
|
||||
{"(Official Video)", "", "officialvideo"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := MatchTitleKey(tc.title, tc.artist); got != tc.want {
|
||||
t.Errorf("MatchTitleKey(%q, %q) = %q, want %q", tc.title, tc.artist, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The release check keeps every word: a box set's demo is its own song.
|
||||
func TestReleaseTitleKey(t *testing.T) {
|
||||
if ReleaseTitleKey("Song (demo)") == ReleaseTitleKey("Song") {
|
||||
t.Error("ReleaseTitleKey collapsed a demo into the song")
|
||||
}
|
||||
if ReleaseTitleKey("Hallelujah Money") != ReleaseTitleKey("hallelujah money!") {
|
||||
t.Error("ReleaseTitleKey kept case or punctuation")
|
||||
}
|
||||
}
|
||||
@@ -25,6 +25,16 @@ var ErrDuplicateGroupNotPending = errors.New("library: duplicate group is not pe
|
||||
// ErrSurvivorNotInGroup means the copy chosen to keep is not a member of the group.
|
||||
var ErrSurvivorNotInGroup = errors.New("library: survivor is not a member of the group")
|
||||
|
||||
// ErrCopyTrackedByLidarr means a copy the merge would remove is one Lidarr maps
|
||||
// to a track of the release it monitors (M498). Removing it would open a hole
|
||||
// that Lidarr fills by downloading the song again, so the merge refuses and
|
||||
// changes nothing.
|
||||
var ErrCopyTrackedByLidarr = errors.New("library: a copy to remove is tracked by Lidarr")
|
||||
|
||||
// RemovableFunc reports whether the file at filePath may be removed. A merge
|
||||
// asks it about every copy it would remove, before touching anything.
|
||||
type RemovableFunc func(filePath string) bool
|
||||
|
||||
// MergedCopy is one copy a merge kept or removed.
|
||||
type MergedCopy struct {
|
||||
TrackID pgtype.UUID
|
||||
@@ -75,6 +85,17 @@ type MergeResult struct {
|
||||
func MergeDuplicateGroup(
|
||||
ctx context.Context, pool *pgxpool.Pool, logger *slog.Logger, dataDir string,
|
||||
groupID, survivorID pgtype.UUID,
|
||||
) (MergeResult, error) {
|
||||
return MergeDuplicateGroupGuarded(ctx, pool, logger, dataDir, groupID, survivorID, nil)
|
||||
}
|
||||
|
||||
// MergeDuplicateGroupGuarded is MergeDuplicateGroup with a check on what it may
|
||||
// remove: when removable says no to any copy that would go, the merge returns
|
||||
// ErrCopyTrackedByLidarr (wrapped with the path) and changes nothing. A nil
|
||||
// removable allows every copy.
|
||||
func MergeDuplicateGroupGuarded(
|
||||
ctx context.Context, pool *pgxpool.Pool, logger *slog.Logger, dataDir string,
|
||||
groupID, survivorID pgtype.UUID, removable RemovableFunc,
|
||||
) (MergeResult, error) {
|
||||
if logger == nil {
|
||||
logger = slog.Default()
|
||||
@@ -109,6 +130,13 @@ func MergeDuplicateGroup(
|
||||
return MergeResult{}, err
|
||||
}
|
||||
|
||||
if removable != nil {
|
||||
for _, l := range losers {
|
||||
if !removable(l.FilePath) {
|
||||
return MergeResult{}, fmt.Errorf("%w: %s", ErrCopyTrackedByLidarr, l.FilePath)
|
||||
}
|
||||
}
|
||||
}
|
||||
for _, l := range losers {
|
||||
if err := removeTrackFileOnDisk(l.FilePath); err != nil {
|
||||
return MergeResult{}, err
|
||||
|
||||
@@ -0,0 +1,678 @@
|
||||
package library
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/jackc/pgx/v5"
|
||||
"github.com/jackc/pgx/v5/pgtype"
|
||||
"github.com/jackc/pgx/v5/pgxpool"
|
||||
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/audit"
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/db/dbq"
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/lidarr"
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/notifications"
|
||||
syncpkg "git.fabledsword.com/bvandeusen/minstrel/internal/sync"
|
||||
)
|
||||
|
||||
// Duplicate resolver (M498 #5435, #5437).
|
||||
//
|
||||
// The sweep proposes duplicate groups; the resolver decides what each one is
|
||||
// and settles the ones that are safe to settle without asking. Safety is
|
||||
// Lidarr's to define: Lidarr maps one file to each track of the release it
|
||||
// monitors, and re-downloads any mapped file that disappears. So:
|
||||
//
|
||||
// - A copy Lidarr maps is never removed.
|
||||
// - On one album, copies Lidarr does not map are merged into the copy it does
|
||||
// (or, when it maps none, into the best by ProposeSurvivor).
|
||||
// - When Lidarr maps every copy on one album, the release it monitors lists
|
||||
// the song twice: the Humanz box set, 14 vinyl sides then the digital
|
||||
// medium. The resolver moves Lidarr to a release of that album that lists
|
||||
// each song once. Lidarr then rescans the folder, maps one copy and leaves
|
||||
// the rest unmapped, and the next pass merges them.
|
||||
//
|
||||
// The release change has a fixed point (lesson #4183): it happens only while
|
||||
// the monitored release repeats a title, and the release it chooses repeats
|
||||
// none, so the next pass finds nothing to change.
|
||||
|
||||
// LidarrLibrary is the part of Lidarr's API the resolver uses. *lidarr.Client
|
||||
// satisfies it.
|
||||
type LidarrLibrary interface {
|
||||
ListUnmappedTrackFiles(ctx context.Context) ([]lidarr.TrackFile, error)
|
||||
LookupAlbumByMBID(ctx context.Context, mbid string) (lidarr.LidarrAlbum, error)
|
||||
ListAlbumTracks(ctx context.Context, albumID int) ([]lidarr.ReleaseTrack, error)
|
||||
GetAlbumReleases(ctx context.Context, albumID int) ([]lidarr.AlbumRelease, error)
|
||||
ListReleaseTracks(ctx context.Context, releaseID int) ([]lidarr.ReleaseTrack, error)
|
||||
SetMonitoredRelease(ctx context.Context, albumID, releaseID int) error
|
||||
}
|
||||
|
||||
const (
|
||||
// resolveMergeCap bounds the merges one pass makes. The first pass after
|
||||
// deploy meets the whole backlog (433 groups measured on 2026-10-09); a
|
||||
// cap spreads it over a few hours, each pass short.
|
||||
resolveMergeCap = 200
|
||||
// resolveReleaseCap bounds the Lidarr release changes one pass makes. Each
|
||||
// makes Lidarr rescan an artist folder.
|
||||
resolveReleaseCap = 10
|
||||
// releaseChangeSettle is how long after a release change the resolver
|
||||
// leaves that album's groups alone. Lidarr unmaps every file of the album
|
||||
// and rescans; until the rescan lands, every copy reads as unmapped, and a
|
||||
// merge then could remove the very file Lidarr is about to map.
|
||||
releaseChangeSettle = 24 * time.Hour
|
||||
// duplicateResolveTick is how often the worker runs a pass.
|
||||
duplicateResolveTick = time.Hour
|
||||
)
|
||||
|
||||
const (
|
||||
lidarrStateTracked = "tracked"
|
||||
lidarrStateUnmapped = "unmapped"
|
||||
)
|
||||
|
||||
// ReleaseChange is one Lidarr release the resolver switched.
|
||||
type ReleaseChange struct {
|
||||
AlbumTitle string
|
||||
ArtistName string
|
||||
From, To lidarr.AlbumRelease
|
||||
}
|
||||
|
||||
// DuplicateResolveResult tallies one pass.
|
||||
type DuplicateResolveResult struct {
|
||||
Groups int
|
||||
Classes map[DuplicateClass]int
|
||||
LidarrConsulted bool
|
||||
Merged int
|
||||
MergeFailed int
|
||||
ReleaseChanges []ReleaseChange
|
||||
}
|
||||
|
||||
// LidarrPathKey is the part of a path Minstrel and Lidarr agree on: the last
|
||||
// three components (artist folder, album folder, file). Both see the library
|
||||
// under their own mount; on the operator's deploy both happen to say /music,
|
||||
// and the key keeps the match from depending on that.
|
||||
func LidarrPathKey(p string) string {
|
||||
parts := strings.Split(strings.Trim(p, "/"), "/")
|
||||
if len(parts) > 3 {
|
||||
parts = parts[len(parts)-3:]
|
||||
}
|
||||
return strings.Join(parts, "/")
|
||||
}
|
||||
|
||||
// resolveGroup is one pending group and its present copies.
|
||||
type resolveGroup struct {
|
||||
id pgtype.UUID
|
||||
tier string
|
||||
members []dbq.ListDuplicateGroupsForResolveRow
|
||||
class DuplicateClass
|
||||
note string
|
||||
states []string // per member: tracked, unmapped, or "" when Lidarr was not asked
|
||||
}
|
||||
|
||||
func (g *resolveGroup) tracked() int {
|
||||
n := 0
|
||||
for _, s := range g.states {
|
||||
if s == lidarrStateTracked {
|
||||
n++
|
||||
}
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// ResolveDuplicates runs one pass. lid is nil when Lidarr is disabled; act is
|
||||
// the operator's auto-resolve setting. Without Lidarr, or with act off, the
|
||||
// pass only classifies: nothing is removed when Lidarr cannot say what is safe.
|
||||
func ResolveDuplicates(
|
||||
ctx context.Context, pool *pgxpool.Pool, logger *slog.Logger, dataDir string, lid LidarrLibrary, act bool,
|
||||
) (DuplicateResolveResult, error) {
|
||||
if logger == nil {
|
||||
logger = slog.Default()
|
||||
}
|
||||
q := dbq.New(pool)
|
||||
res := DuplicateResolveResult{Classes: map[DuplicateClass]int{}}
|
||||
|
||||
rows, err := q.ListDuplicateGroupsForResolve(ctx)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("list duplicate groups: %w", err)
|
||||
}
|
||||
groups := foldResolveGroups(rows)
|
||||
res.Groups = len(groups)
|
||||
|
||||
var unmapped map[string]bool
|
||||
if lid != nil {
|
||||
files, err := lid.ListUnmappedTrackFiles(ctx)
|
||||
if err != nil {
|
||||
logger.Warn("duplicate resolve: Lidarr's unmapped files unavailable; classifying only", "err", err)
|
||||
} else {
|
||||
unmapped = make(map[string]bool, len(files))
|
||||
for _, f := range files {
|
||||
unmapped[LidarrPathKey(f.Path)] = true
|
||||
}
|
||||
res.LidarrConsulted = true
|
||||
}
|
||||
}
|
||||
|
||||
for _, g := range groups {
|
||||
cm := make([]ClassifyMember, len(g.members))
|
||||
g.states = make([]string, len(g.members))
|
||||
for i, m := range g.members {
|
||||
cm[i] = ClassifyMember{AlbumID: syncpkg.FormatUUID(m.AlbumID), Title: m.Title, ArtistName: m.ArtistName}
|
||||
if res.LidarrConsulted {
|
||||
if unmapped[LidarrPathKey(m.FilePath)] {
|
||||
g.states[i] = lidarrStateUnmapped
|
||||
} else {
|
||||
g.states[i] = lidarrStateTracked
|
||||
}
|
||||
}
|
||||
}
|
||||
g.class = ClassifyDuplicateGroup(g.tier, cm)
|
||||
g.note = resolveNote(g)
|
||||
res.Classes[g.class]++
|
||||
}
|
||||
if err := writeResolveVerdicts(ctx, q, groups); err != nil {
|
||||
return res, err
|
||||
}
|
||||
|
||||
if !act || !res.LidarrConsulted {
|
||||
return res, nil
|
||||
}
|
||||
|
||||
settling, err := recentlyChangedAlbums(ctx, q, time.Now())
|
||||
if err != nil {
|
||||
return res, err
|
||||
}
|
||||
|
||||
// Merges first, on the state Lidarr reported at the start of the pass.
|
||||
removable := func(p string) bool { return unmapped[LidarrPathKey(p)] }
|
||||
for _, g := range groups {
|
||||
if res.Merged >= resolveMergeCap {
|
||||
break
|
||||
}
|
||||
if ctx.Err() != nil {
|
||||
return res, ctx.Err()
|
||||
}
|
||||
if g.class != ClassSameRelease || g.tracked() > 1 || settling[syncpkg.FormatUUID(g.members[0].AlbumID)] {
|
||||
continue
|
||||
}
|
||||
if err := autoMerge(ctx, pool, logger, dataDir, g, removable); err != nil {
|
||||
res.MergeFailed++
|
||||
logger.Warn("duplicate resolve: merge failed", "group_id", syncpkg.FormatUUID(g.id), "err", err)
|
||||
continue
|
||||
}
|
||||
res.Merged++
|
||||
}
|
||||
|
||||
// Then release changes, for albums where Lidarr maps every copy.
|
||||
changes, err := changeRepeatingReleases(ctx, q, pool, logger, lid, groups, settling)
|
||||
res.ReleaseChanges = changes
|
||||
if err != nil {
|
||||
return res, err
|
||||
}
|
||||
|
||||
if res.Merged > 0 || len(changes) > 0 {
|
||||
notifyAdmins(ctx, notifications.KindDuplicatesResolved, notifications.Payload{
|
||||
Count: int64(res.Merged + len(changes)),
|
||||
Detail: resolvedDetail(res.Merged, changes),
|
||||
})
|
||||
}
|
||||
return res, nil
|
||||
}
|
||||
|
||||
func foldResolveGroups(rows []dbq.ListDuplicateGroupsForResolveRow) []*resolveGroup {
|
||||
var groups []*resolveGroup
|
||||
for _, r := range rows {
|
||||
if n := len(groups); n == 0 || groups[n-1].id != r.GroupID {
|
||||
groups = append(groups, &resolveGroup{id: r.GroupID, tier: r.Tier})
|
||||
}
|
||||
g := groups[len(groups)-1]
|
||||
g.members = append(g.members, r)
|
||||
}
|
||||
// A group whose other copies went missing has nothing left to resolve.
|
||||
out := groups[:0]
|
||||
for _, g := range groups {
|
||||
if len(g.members) >= 2 {
|
||||
out = append(out, g)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// resolveNote is what the report says beside a group the resolver leaves for
|
||||
// the operator, when there is something to say.
|
||||
func resolveNote(g *resolveGroup) string {
|
||||
switch g.class {
|
||||
case ClassSameRelease:
|
||||
if g.tracked() > 1 {
|
||||
return "Lidarr tracks every copy as its own track, so removing one would make Lidarr download it again."
|
||||
}
|
||||
case ClassCrossRelease:
|
||||
return "Each copy belongs to its own release, and Lidarr keeps each one. Nothing is removed."
|
||||
case ClassMismatch:
|
||||
return "Identical audio filed under different titles: one file carries another song's tags."
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
func writeResolveVerdicts(ctx context.Context, q *dbq.Queries, groups []*resolveGroup) error {
|
||||
classes := dbq.SetDuplicateGroupClassesParams{}
|
||||
states := dbq.SetDuplicateMemberLidarrStatesParams{}
|
||||
for _, g := range groups {
|
||||
classes.GroupIds = append(classes.GroupIds, g.id)
|
||||
classes.Classes = append(classes.Classes, string(g.class))
|
||||
classes.Notes = append(classes.Notes, g.note)
|
||||
for i, m := range g.members {
|
||||
states.GroupIds = append(states.GroupIds, g.id)
|
||||
states.TrackIds = append(states.TrackIds, m.TrackID)
|
||||
states.States = append(states.States, g.states[i])
|
||||
}
|
||||
}
|
||||
if len(classes.GroupIds) == 0 {
|
||||
return nil
|
||||
}
|
||||
if err := q.SetDuplicateGroupClasses(ctx, classes); err != nil {
|
||||
return fmt.Errorf("write duplicate classes: %w", err)
|
||||
}
|
||||
if err := q.SetDuplicateMemberLidarrStates(ctx, states); err != nil {
|
||||
return fmt.Errorf("write lidarr states: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// survivorCandidates builds what ProposeSurvivor weighs for a group's copies.
|
||||
func survivorCandidates(g *resolveGroup) []SurvivorCandidate {
|
||||
cands := make([]SurvivorCandidate, len(g.members))
|
||||
for i, m := range g.members {
|
||||
cands[i] = SurvivorCandidate{
|
||||
TrackID: syncpkg.FormatUUID(m.TrackID),
|
||||
FileFormat: m.FileFormat,
|
||||
FileSize: m.FileSize,
|
||||
AddedAt: m.AddedAt.Time,
|
||||
LidarrTracked: g.states[i] == lidarrStateTracked,
|
||||
TagFit: TagFitScore(m.TrackNumber != nil, m.PositionClash, m.FilePath, m.HasMbid),
|
||||
}
|
||||
}
|
||||
return cands
|
||||
}
|
||||
|
||||
// autoMerge merges one group into the copy ProposeSurvivor picks, which is the
|
||||
// copy Lidarr tracks when there is one. removable is the guard: a copy Lidarr
|
||||
// maps can never be among those removed, whatever the survivor rule said.
|
||||
func autoMerge(
|
||||
ctx context.Context, pool *pgxpool.Pool, logger *slog.Logger, dataDir string, g *resolveGroup, removable RemovableFunc,
|
||||
) error {
|
||||
survivorKey, reason := ProposeSurvivor(survivorCandidates(g))
|
||||
var survivorID pgtype.UUID
|
||||
if err := survivorID.Scan(survivorKey); err != nil {
|
||||
return fmt.Errorf("parse survivor id: %w", err)
|
||||
}
|
||||
res, err := MergeDuplicateGroupGuarded(ctx, pool, logger, dataDir, g.id, survivorID, removable)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := dbq.New(pool).MarkDuplicateGroupResolvedAutomatically(ctx, g.id); err != nil {
|
||||
logger.Warn("duplicate resolve: marking the merge automatic failed", "group_id", syncpkg.FormatUUID(g.id), "err", err)
|
||||
}
|
||||
removed := make([]map[string]string, 0, len(res.Removed))
|
||||
for _, c := range res.Removed {
|
||||
removed = append(removed, map[string]string{"track_id": syncpkg.FormatUUID(c.TrackID), "file_path": c.FilePath})
|
||||
}
|
||||
audit.WriteOrLog(ctx, pool, logger, pgtype.UUID{}, pgtype.UUID{}, audit.ActionDuplicateMerge, map[string]any{
|
||||
"automatic": true,
|
||||
"group_id": syncpkg.FormatUUID(g.id),
|
||||
"tier": res.Tier,
|
||||
"class": string(g.class),
|
||||
"survivor_track_id": syncpkg.FormatUUID(res.Survivor.TrackID),
|
||||
"survivor_path": res.Survivor.FilePath,
|
||||
"survivor_reason": reason,
|
||||
"removed": removed,
|
||||
"moved": map[string]any{
|
||||
"play_events": res.PlayEvents, "skip_events": res.SkipEvents,
|
||||
"likes": res.Likes, "playlist_entries": res.PlaylistEntries,
|
||||
},
|
||||
})
|
||||
return nil
|
||||
}
|
||||
|
||||
// recentlyChangedAlbums is the Minstrel albums whose Lidarr release the
|
||||
// resolver changed within releaseChangeSettle, read from the audit log so a
|
||||
// restart does not forget them.
|
||||
func recentlyChangedAlbums(ctx context.Context, q *dbq.Queries, now time.Time) (map[string]bool, error) {
|
||||
rows, err := q.ListAuditLogByActions(ctx, dbq.ListAuditLogByActionsParams{
|
||||
Actions: []string{string(audit.ActionLidarrReleaseChange)}, SystemOnly: true,
|
||||
PageLimit: 200, PageOffset: 0,
|
||||
})
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read recent release changes: %w", err)
|
||||
}
|
||||
out := map[string]bool{}
|
||||
for _, r := range rows {
|
||||
if now.Sub(r.CreatedAt.Time) > releaseChangeSettle {
|
||||
break // newest first
|
||||
}
|
||||
var meta struct {
|
||||
AlbumID string `json:"album_id"`
|
||||
}
|
||||
if json.Unmarshal(r.Metadata, &meta) == nil && meta.AlbumID != "" {
|
||||
out[meta.AlbumID] = true
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// repeatingAlbum is a Minstrel album where Lidarr maps more than one copy of a
|
||||
// song, and the groups that showed it.
|
||||
type repeatingAlbum struct {
|
||||
albumID pgtype.UUID
|
||||
releaseGroupMbid string
|
||||
title, artist string
|
||||
groups []*resolveGroup
|
||||
}
|
||||
|
||||
func changeRepeatingReleases(
|
||||
ctx context.Context, q *dbq.Queries, pool *pgxpool.Pool, logger *slog.Logger,
|
||||
lid LidarrLibrary, groups []*resolveGroup, settling map[string]bool,
|
||||
) ([]ReleaseChange, error) {
|
||||
albums := map[string]*repeatingAlbum{}
|
||||
for _, g := range groups {
|
||||
if g.class != ClassSameRelease || g.tracked() < 2 {
|
||||
continue
|
||||
}
|
||||
m := g.members[0]
|
||||
key := syncpkg.FormatUUID(m.AlbumID)
|
||||
if settling[key] || m.ReleaseGroupMbid == nil || *m.ReleaseGroupMbid == "" {
|
||||
continue
|
||||
}
|
||||
a := albums[key]
|
||||
if a == nil {
|
||||
a = &repeatingAlbum{albumID: m.AlbumID, releaseGroupMbid: *m.ReleaseGroupMbid, title: m.AlbumTitle, artist: m.ArtistName}
|
||||
albums[key] = a
|
||||
}
|
||||
a.groups = append(a.groups, g)
|
||||
}
|
||||
keys := make([]string, 0, len(albums))
|
||||
for k := range albums {
|
||||
keys = append(keys, k)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
|
||||
var changes []ReleaseChange
|
||||
for _, k := range keys {
|
||||
if len(changes) >= resolveReleaseCap {
|
||||
break
|
||||
}
|
||||
if ctx.Err() != nil {
|
||||
return changes, ctx.Err()
|
||||
}
|
||||
a := albums[k]
|
||||
change, note, err := changeRelease(ctx, q, lid, a)
|
||||
if err != nil {
|
||||
logger.Warn("duplicate resolve: Lidarr release check failed", "album", a.title, "err", err)
|
||||
continue
|
||||
}
|
||||
if change != nil {
|
||||
changes = append(changes, *change)
|
||||
audit.WriteOrLog(ctx, pool, logger, pgtype.UUID{}, pgtype.UUID{}, audit.ActionLidarrReleaseChange, map[string]any{
|
||||
"automatic": true,
|
||||
"album_id": k,
|
||||
"album_title": a.title,
|
||||
"artist_name": a.artist,
|
||||
"from_release": releaseLabel(change.From),
|
||||
"to_release": releaseLabel(change.To),
|
||||
})
|
||||
}
|
||||
if note != "" {
|
||||
if err := setGroupNotes(ctx, q, a.groups, note); err != nil {
|
||||
return changes, err
|
||||
}
|
||||
}
|
||||
}
|
||||
return changes, nil
|
||||
}
|
||||
|
||||
// changeRelease checks one album's monitored release and, when it lists a
|
||||
// song twice, moves Lidarr to the release that lists each song once and best
|
||||
// covers what is on disk. It returns the change made (nil when none) and the
|
||||
// note for the album's groups.
|
||||
func changeRelease(
|
||||
ctx context.Context, q *dbq.Queries, lid LidarrLibrary, a *repeatingAlbum,
|
||||
) (*ReleaseChange, string, error) {
|
||||
la, err := lid.LookupAlbumByMBID(ctx, a.releaseGroupMbid)
|
||||
if errors.Is(err, lidarr.ErrNotFound) {
|
||||
return nil, "Lidarr does not hold this album, so its copies are left for you.", nil
|
||||
}
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
current, err := lid.ListAlbumTracks(ctx, la.ID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
if !releaseRepeats(current) {
|
||||
return nil, "Lidarr maps each copy to a different track of this album, so none can go without a download.", nil
|
||||
}
|
||||
releases, err := lid.GetAlbumReleases(ctx, la.ID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
titles, err := q.ListAlbumPresentTrackTitles(ctx, a.albumID)
|
||||
if err != nil {
|
||||
return nil, "", fmt.Errorf("list album titles: %w", err)
|
||||
}
|
||||
onDisk := map[string]bool{}
|
||||
for _, t := range titles {
|
||||
if k := ReleaseTitleKey(t); k != "" {
|
||||
onDisk[k] = true
|
||||
}
|
||||
}
|
||||
|
||||
var from lidarr.AlbumRelease
|
||||
var candidates []releaseCandidate
|
||||
for _, r := range releases {
|
||||
if r.Monitored {
|
||||
from = r
|
||||
continue
|
||||
}
|
||||
tracks, err := lid.ListReleaseTracks(ctx, r.ID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
if len(tracks) == 0 || releaseRepeats(tracks) {
|
||||
continue
|
||||
}
|
||||
candidates = append(candidates, releaseCandidate{release: r, coverage: coverage(tracks, onDisk)})
|
||||
}
|
||||
best, ok := pickRelease(candidates, coverage(current, onDisk))
|
||||
if !ok {
|
||||
return nil, "Lidarr's release of this album lists songs twice, and no other release lists each once while keeping what is on disk.", nil
|
||||
}
|
||||
if err := lid.SetMonitoredRelease(ctx, la.ID, best.ID); err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
note := fmt.Sprintf("Lidarr now monitors the %s release, which lists each song once. The extra copies are merged once Lidarr has rescanned.", releaseLabel(best))
|
||||
return &ReleaseChange{AlbumTitle: a.title, ArtistName: a.artist, From: from, To: best}, note, nil
|
||||
}
|
||||
|
||||
type releaseCandidate struct {
|
||||
release lidarr.AlbumRelease
|
||||
coverage int
|
||||
}
|
||||
|
||||
// pickRelease chooses among releases that list each song once: the one
|
||||
// covering most of the titles on disk, then the fewest tracks (the plainest
|
||||
// edition that holds them), then a digital one, then the lowest id so the
|
||||
// choice is stable. It refuses a release covering fewer titles on disk than
|
||||
// the current one does: the change must not drop a song Lidarr now keeps.
|
||||
func pickRelease(cands []releaseCandidate, currentCoverage int) (lidarr.AlbumRelease, bool) {
|
||||
if len(cands) == 0 {
|
||||
return lidarr.AlbumRelease{}, false
|
||||
}
|
||||
sort.SliceStable(cands, func(i, j int) bool {
|
||||
a, b := cands[i], cands[j]
|
||||
if a.coverage != b.coverage {
|
||||
return a.coverage > b.coverage
|
||||
}
|
||||
if a.release.TrackCount != b.release.TrackCount {
|
||||
return a.release.TrackCount < b.release.TrackCount
|
||||
}
|
||||
if da, db := isDigital(a.release), isDigital(b.release); da != db {
|
||||
return da
|
||||
}
|
||||
return a.release.ID < b.release.ID
|
||||
})
|
||||
if cands[0].coverage < currentCoverage {
|
||||
return lidarr.AlbumRelease{}, false
|
||||
}
|
||||
return cands[0].release, true
|
||||
}
|
||||
|
||||
func isDigital(r lidarr.AlbumRelease) bool {
|
||||
return strings.Contains(strings.ToLower(r.Format), "digital")
|
||||
}
|
||||
|
||||
// releaseRepeats reports whether a release lists one title more than once.
|
||||
func releaseRepeats(tracks []lidarr.ReleaseTrack) bool {
|
||||
seen := map[string]bool{}
|
||||
for _, t := range tracks {
|
||||
k := ReleaseTitleKey(t.Title)
|
||||
if k == "" {
|
||||
continue
|
||||
}
|
||||
if seen[k] {
|
||||
return true
|
||||
}
|
||||
seen[k] = true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// coverage counts the distinct titles on disk that the release lists.
|
||||
func coverage(tracks []lidarr.ReleaseTrack, onDisk map[string]bool) int {
|
||||
hit := map[string]bool{}
|
||||
for _, t := range tracks {
|
||||
if k := ReleaseTitleKey(t.Title); onDisk[k] {
|
||||
hit[k] = true
|
||||
}
|
||||
}
|
||||
return len(hit)
|
||||
}
|
||||
|
||||
func releaseLabel(r lidarr.AlbumRelease) string {
|
||||
label := r.Format
|
||||
if label == "" {
|
||||
label = r.Title
|
||||
}
|
||||
if r.Disambiguation != "" {
|
||||
label = r.Disambiguation + ", " + label
|
||||
}
|
||||
return fmt.Sprintf("%s (%d tracks)", label, r.TrackCount)
|
||||
}
|
||||
|
||||
func setGroupNotes(ctx context.Context, q *dbq.Queries, groups []*resolveGroup, note string) error {
|
||||
p := dbq.SetDuplicateGroupClassesParams{}
|
||||
for _, g := range groups {
|
||||
g.note = note
|
||||
p.GroupIds = append(p.GroupIds, g.id)
|
||||
p.Classes = append(p.Classes, string(g.class))
|
||||
p.Notes = append(p.Notes, note)
|
||||
}
|
||||
if err := q.SetDuplicateGroupClasses(ctx, p); err != nil {
|
||||
return fmt.Errorf("write duplicate notes: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func resolvedDetail(merged int, changes []ReleaseChange) string {
|
||||
var parts []string
|
||||
if merged > 0 {
|
||||
parts = append(parts, fmt.Sprintf("Merged %d %s Lidarr does not need.", merged, pluralWord(merged, "copy", "copies")))
|
||||
}
|
||||
switch len(changes) {
|
||||
case 0:
|
||||
case 1:
|
||||
c := changes[0]
|
||||
parts = append(parts, fmt.Sprintf("Lidarr now monitors the %s release of %s by %s, which lists each song once.",
|
||||
releaseLabel(c.To), c.AlbumTitle, c.ArtistName))
|
||||
default:
|
||||
parts = append(parts, fmt.Sprintf("Lidarr now monitors a release that lists each song once for %d albums.", len(changes)))
|
||||
}
|
||||
return strings.Join(parts, " ")
|
||||
}
|
||||
|
||||
func pluralWord(n int, one, many string) string {
|
||||
if n == 1 {
|
||||
return one
|
||||
}
|
||||
return many
|
||||
}
|
||||
|
||||
// DuplicateResolveWorker runs a resolver pass every hour.
|
||||
type DuplicateResolveWorker struct {
|
||||
pool *pgxpool.Pool
|
||||
logger *slog.Logger
|
||||
dataDir string
|
||||
settings *FingerprintSettingsService
|
||||
lidarr func() LidarrLibrary
|
||||
tick time.Duration
|
||||
}
|
||||
|
||||
// NewDuplicateResolveWorker builds a worker with the production cadence.
|
||||
// lidarrFn is asked each pass, so a Lidarr setting saved in admin takes effect
|
||||
// without a restart; it returns nil while Lidarr is disabled.
|
||||
func NewDuplicateResolveWorker(
|
||||
pool *pgxpool.Pool, logger *slog.Logger, dataDir string,
|
||||
settings *FingerprintSettingsService, lidarrFn func() LidarrLibrary,
|
||||
) *DuplicateResolveWorker {
|
||||
return &DuplicateResolveWorker{
|
||||
pool: pool, logger: logger, dataDir: dataDir, settings: settings, lidarr: lidarrFn, tick: duplicateResolveTick,
|
||||
}
|
||||
}
|
||||
|
||||
// Run blocks until ctx is cancelled, running a pass at start and then each tick.
|
||||
func (w *DuplicateResolveWorker) Run(ctx context.Context) {
|
||||
w.tickOnce(ctx)
|
||||
t := time.NewTicker(w.tick)
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
w.tickOnce(ctx)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// tickOnce contains one pass so nothing it does can stop the next tick (rule 157).
|
||||
func (w *DuplicateResolveWorker) tickOnce(ctx context.Context) {
|
||||
defer func() {
|
||||
if r := recover(); r != nil {
|
||||
w.logger.Error("duplicate resolve: tick panicked", "panic", r)
|
||||
}
|
||||
}()
|
||||
// A sweep in flight is rewriting the groups; the next tick sees its result.
|
||||
_, err := dbq.New(w.pool).GetInFlightDuplicateSweep(ctx)
|
||||
switch {
|
||||
case err == nil:
|
||||
return
|
||||
case !errors.Is(err, pgx.ErrNoRows):
|
||||
if ctx.Err() == nil {
|
||||
w.logger.Warn("duplicate resolve: sweep check failed", "err", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
var lid LidarrLibrary
|
||||
if w.lidarr != nil {
|
||||
lid = w.lidarr()
|
||||
}
|
||||
res, err := ResolveDuplicates(ctx, w.pool, w.logger, w.dataDir, lid, w.settings.Get().AutoResolve)
|
||||
if err != nil && ctx.Err() == nil {
|
||||
w.logger.Warn("duplicate resolve: pass failed", "err", err)
|
||||
}
|
||||
w.logger.Info("duplicate resolve complete",
|
||||
"groups", res.Groups, "lidarr", res.LidarrConsulted, "merged", res.Merged,
|
||||
"merge_failed", res.MergeFailed, "release_changes", len(res.ReleaseChanges))
|
||||
}
|
||||
@@ -0,0 +1,349 @@
|
||||
package library
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/jackc/pgx/v5/pgtype"
|
||||
"github.com/jackc/pgx/v5/pgxpool"
|
||||
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/db/dbq"
|
||||
"git.fabledsword.com/bvandeusen/minstrel/internal/lidarr"
|
||||
syncpkg "git.fabledsword.com/bvandeusen/minstrel/internal/sync"
|
||||
)
|
||||
|
||||
func TestLidarrPathKey(t *testing.T) {
|
||||
// Minstrel and Lidarr see the library under different mounts; the key is
|
||||
// what they agree on.
|
||||
a := LidarrPathKey("/music/Gorillaz/Humanz (2017)/Gorillaz - Humanz - 01 - Ascension.mp3")
|
||||
b := LidarrPathKey("/data/media/music/Gorillaz/Humanz (2017)/Gorillaz - Humanz - 01 - Ascension.mp3")
|
||||
if a != b || a != "Gorillaz/Humanz (2017)/Gorillaz - Humanz - 01 - Ascension.mp3" {
|
||||
t.Errorf("keys differ or are wrong: %q, %q", a, b)
|
||||
}
|
||||
if LidarrPathKey("/music/Gorillaz/Humanz (2017)/a.mp3") == LidarrPathKey("/music/Gorillaz/Demon Days (2005)/a.mp3") {
|
||||
t.Error("two albums' files share a key")
|
||||
}
|
||||
}
|
||||
|
||||
func rt(titles ...string) []lidarr.ReleaseTrack {
|
||||
out := make([]lidarr.ReleaseTrack, len(titles))
|
||||
for i, title := range titles {
|
||||
out[i] = lidarr.ReleaseTrack{ID: i + 1, Title: title}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func TestReleaseRepeats(t *testing.T) {
|
||||
if !releaseRepeats(rt("Ascension", "Strobelite", "Ascension")) {
|
||||
t.Error("the box set lists Ascension twice, but no repeat was found")
|
||||
}
|
||||
if releaseRepeats(rt("Charger", "Charger (demo)")) {
|
||||
t.Error("a demo was read as a repeat of the song")
|
||||
}
|
||||
}
|
||||
|
||||
// The Humanz releases as Lidarr lists them (2026-10-09): the box set is
|
||||
// monitored, and the 40-track digital release is the one that lists each song
|
||||
// once and covers what is on disk.
|
||||
func TestPickRelease(t *testing.T) {
|
||||
digital40 := lidarr.AlbumRelease{ID: 10149, Format: "Digital Media", TrackCount: 40}
|
||||
deluxe26 := lidarr.AlbumRelease{ID: 10150, Format: "CD", TrackCount: 26}
|
||||
standard20 := lidarr.AlbumRelease{ID: 10151, Format: "Digital Media", TrackCount: 20}
|
||||
|
||||
got, ok := pickRelease([]releaseCandidate{
|
||||
{release: standard20, coverage: 20}, {release: deluxe26, coverage: 26}, {release: digital40, coverage: 34},
|
||||
}, 34)
|
||||
if !ok || got.ID != digital40.ID {
|
||||
t.Errorf("picked %+v (ok=%v), want the 40-track digital release", got, ok)
|
||||
}
|
||||
|
||||
// Equal coverage: the plainer edition, then digital.
|
||||
cd := lidarr.AlbumRelease{ID: 1, Format: "CD", TrackCount: 26}
|
||||
web := lidarr.AlbumRelease{ID: 2, Format: "Digital Media", TrackCount: 26}
|
||||
if got, _ := pickRelease([]releaseCandidate{{release: cd, coverage: 26}, {release: web, coverage: 26}}, 26); got.ID != web.ID {
|
||||
t.Errorf("picked %+v, want the digital one of two equal releases", got)
|
||||
}
|
||||
|
||||
// Never a release that keeps fewer of the songs on disk than the current one.
|
||||
if got, ok := pickRelease([]releaseCandidate{{release: deluxe26, coverage: 26}}, 34); ok {
|
||||
t.Errorf("picked %+v, which covers fewer songs on disk than the current release", got)
|
||||
}
|
||||
}
|
||||
|
||||
// fakeLidarr is one Lidarr album with several releases. Switching the
|
||||
// monitored release changes what ListAlbumTracks answers, as in Lidarr.
|
||||
type fakeLidarr struct {
|
||||
unmapped []string
|
||||
unmappedErr error
|
||||
album lidarr.LidarrAlbum
|
||||
releases []lidarr.AlbumRelease
|
||||
releaseTracks map[int][]lidarr.ReleaseTrack
|
||||
setCalls [][2]int
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) ListUnmappedTrackFiles(context.Context) ([]lidarr.TrackFile, error) {
|
||||
out := make([]lidarr.TrackFile, len(f.unmapped))
|
||||
for i, p := range f.unmapped {
|
||||
out[i] = lidarr.TrackFile{ID: i + 1, Path: p}
|
||||
}
|
||||
return out, f.unmappedErr
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) LookupAlbumByMBID(_ context.Context, mbid string) (lidarr.LidarrAlbum, error) {
|
||||
if f.album.ForeignAlbumID != mbid {
|
||||
return lidarr.LidarrAlbum{}, lidarr.ErrNotFound
|
||||
}
|
||||
return f.album, nil
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) ListAlbumTracks(context.Context, int) ([]lidarr.ReleaseTrack, error) {
|
||||
for _, r := range f.releases {
|
||||
if r.Monitored {
|
||||
return f.releaseTracks[r.ID], nil
|
||||
}
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) GetAlbumReleases(context.Context, int) ([]lidarr.AlbumRelease, error) {
|
||||
return append([]lidarr.AlbumRelease(nil), f.releases...), nil
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) ListReleaseTracks(_ context.Context, releaseID int) ([]lidarr.ReleaseTrack, error) {
|
||||
return f.releaseTracks[releaseID], nil
|
||||
}
|
||||
|
||||
func (f *fakeLidarr) SetMonitoredRelease(_ context.Context, albumID, releaseID int) error {
|
||||
f.setCalls = append(f.setCalls, [2]int{albumID, releaseID})
|
||||
for i := range f.releases {
|
||||
f.releases[i].Monitored = f.releases[i].ID == releaseID
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// resolveFixture is the Humanz shape: one album, one song twice, both copies on
|
||||
// disk under an artist/album folder, in one pending exact-tier group.
|
||||
type resolveFixture struct {
|
||||
pool *pgxpool.Pool
|
||||
clean, rip dbq.Track
|
||||
cleanPath, ripPath string
|
||||
groupID pgtype.UUID
|
||||
}
|
||||
|
||||
func newResolveFixture(t *testing.T) resolveFixture {
|
||||
t.Helper()
|
||||
pool := newPool(t)
|
||||
ctx := context.Background()
|
||||
q := dbq.New(pool)
|
||||
dir := filepath.Join(t.TempDir(), "Gorillaz", "Humanz (2017)")
|
||||
if err := os.MkdirAll(dir, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f := resolveFixture{pool: pool}
|
||||
f.cleanPath = filepath.Join(dir, "Gorillaz - Humanz - 03 - Saturnz Barz.mp3")
|
||||
f.ripPath = filepath.Join(dir, "Gorillaz - Humanz - 03 - Saturnz Barz (Official Video).mp3")
|
||||
for _, p := range []string{f.cleanPath, f.ripPath} {
|
||||
if err := os.WriteFile(p, []byte("audio"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
var album dbq.Album
|
||||
var artist dbq.Artist
|
||||
f.clean, album, artist = seedTrack(t, pool, f.cleanPath)
|
||||
rip, err := q.UpsertTrack(ctx, dbq.UpsertTrackParams{
|
||||
Title: "Saturnz Barz (Official Video)", AlbumID: album.ID, ArtistID: artist.ID,
|
||||
DurationMs: 1000, FilePath: f.ripPath, FileSize: 100, FileFormat: "mp3",
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f.rip = rip
|
||||
exec := func(sql string, args ...any) {
|
||||
t.Helper()
|
||||
if _, err := pool.Exec(ctx, sql, args...); err != nil {
|
||||
t.Fatalf("exec %q: %v", sql, err)
|
||||
}
|
||||
}
|
||||
exec(`UPDATE tracks SET title = 'Saturnz Barz' WHERE id = $1`, f.clean.ID)
|
||||
exec(`UPDATE albums SET release_group_mbid = 'rg-humanz' WHERE id = $1`, album.ID)
|
||||
if err := pool.QueryRow(ctx,
|
||||
`INSERT INTO duplicate_groups (member_key, tier) VALUES ('resolve-fixture', 'exact') RETURNING id`,
|
||||
).Scan(&f.groupID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
exec(`INSERT INTO duplicate_group_members (group_id, track_id) VALUES ($1, $2), ($1, $3)`, f.groupID, f.clean.ID, f.rip.ID)
|
||||
return f
|
||||
}
|
||||
|
||||
// lidarrPath is where Lidarr, mounted elsewhere, sees a fixture file.
|
||||
func lidarrPath(p string) string { return "/music/" + LidarrPathKey(p) }
|
||||
|
||||
func (f resolveFixture) group(t *testing.T) (status string, class *string, note *string, auto bool) {
|
||||
t.Helper()
|
||||
if err := f.pool.QueryRow(context.Background(),
|
||||
`SELECT status, class, resolve_note, resolved_automatically FROM duplicate_groups WHERE id = $1`, f.groupID,
|
||||
).Scan(&status, &class, ¬e, &auto); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return status, class, note, auto
|
||||
}
|
||||
|
||||
func exists(p string) bool { _, err := os.Stat(p); return err == nil }
|
||||
|
||||
// Lidarr maps the clean copy and holds the rip unmapped: the rip goes, the
|
||||
// clean copy stays, and the merge is recorded as the resolver's.
|
||||
func TestResolveDuplicates_MergesTheUnmappedCopy_Integration(t *testing.T) {
|
||||
f := newResolveFixture(t)
|
||||
lid := &fakeLidarr{unmapped: []string{lidarrPath(f.ripPath)}}
|
||||
|
||||
res, err := ResolveDuplicates(context.Background(), f.pool, nil, "", lid, true)
|
||||
if err != nil {
|
||||
t.Fatalf("resolve: %v", err)
|
||||
}
|
||||
if res.Merged != 1 || !res.LidarrConsulted {
|
||||
t.Fatalf("result = %+v, want one merge with Lidarr consulted", res)
|
||||
}
|
||||
if exists(f.ripPath) || !exists(f.cleanPath) {
|
||||
t.Errorf("rip exists=%v clean exists=%v; want only the clean copy left", exists(f.ripPath), exists(f.cleanPath))
|
||||
}
|
||||
status, class, _, auto := f.group(t)
|
||||
if status != "merged" || class == nil || *class != string(ClassSameRelease) || !auto {
|
||||
t.Errorf("group = (%s, %v, auto=%v), want merged same_release by the resolver", status, class, auto)
|
||||
}
|
||||
var audits int
|
||||
if err := f.pool.QueryRow(context.Background(),
|
||||
`SELECT count(*) FROM audit_log WHERE action = 'duplicate_merge' AND actor_id IS NULL
|
||||
AND metadata->>'group_id' = $1 AND (metadata->>'automatic')::boolean`,
|
||||
syncpkg.FormatUUID(f.groupID)).Scan(&audits); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if audits != 1 {
|
||||
t.Errorf("automatic merge audit rows = %d, want 1", audits)
|
||||
}
|
||||
}
|
||||
|
||||
// Lidarr maps both copies, and its release lists the song twice: nothing is
|
||||
// deleted, the release moves to the one listing each song once, and the next
|
||||
// pass leaves it there (lesson #4183's fixed point).
|
||||
func TestResolveDuplicates_ChangesARepeatingRelease_Integration(t *testing.T) {
|
||||
f := newResolveFixture(t)
|
||||
lid := &fakeLidarr{
|
||||
album: lidarr.LidarrAlbum{ID: 4723, ForeignAlbumID: "rg-humanz", Title: "Humanz"},
|
||||
releases: []lidarr.AlbumRelease{
|
||||
{ID: 10148, Format: "14x12\" Vinyl, Digital Media", TrackCount: 4, Monitored: true},
|
||||
{ID: 10149, Format: "Digital Media", TrackCount: 2},
|
||||
},
|
||||
releaseTracks: map[int][]lidarr.ReleaseTrack{
|
||||
10148: rt("Saturnz Barz", "Ascension", "Saturnz Barz", "Ascension"),
|
||||
10149: rt("Saturnz Barz", "Ascension"),
|
||||
},
|
||||
}
|
||||
|
||||
res, err := ResolveDuplicates(context.Background(), f.pool, nil, "", lid, true)
|
||||
if err != nil {
|
||||
t.Fatalf("resolve: %v", err)
|
||||
}
|
||||
if res.Merged != 0 || !exists(f.ripPath) || !exists(f.cleanPath) {
|
||||
t.Fatalf("merged %d; a copy Lidarr maps was touched", res.Merged)
|
||||
}
|
||||
if len(lid.setCalls) != 1 || lid.setCalls[0] != [2]int{4723, 10149} {
|
||||
t.Fatalf("release changes = %v, want album 4723 to release 10149", lid.setCalls)
|
||||
}
|
||||
status, _, note, _ := f.group(t)
|
||||
if status != "pending" || note == nil || *note == "" {
|
||||
t.Errorf("group = (%s, note %v), want pending with a note saying what changed", status, note)
|
||||
}
|
||||
|
||||
// The next pass: Lidarr has not rescanned yet, so both copies still read
|
||||
// as tracked. The album is settling, so nothing changes.
|
||||
if _, err := ResolveDuplicates(context.Background(), f.pool, nil, "", lid, true); err != nil {
|
||||
t.Fatalf("second resolve: %v", err)
|
||||
}
|
||||
if len(lid.setCalls) != 1 {
|
||||
t.Errorf("release changed again: %v", lid.setCalls)
|
||||
}
|
||||
// And once it has settled, the check itself finds nothing to do: the
|
||||
// release it chose lists each song once.
|
||||
change, _, err := changeRelease(context.Background(), dbq.New(f.pool), lid,
|
||||
&repeatingAlbum{albumID: f.clean.AlbumID, releaseGroupMbid: "rg-humanz"})
|
||||
if err != nil || change != nil || len(lid.setCalls) != 1 {
|
||||
t.Errorf("after the change, changeRelease = (%+v, %v) with %d calls; want no change", change, err, len(lid.setCalls))
|
||||
}
|
||||
}
|
||||
|
||||
// After a release change Lidarr unmaps every file of the album until it has
|
||||
// rescanned. A pass in that window must not read "every copy unmapped" as
|
||||
// licence to merge: the copy it would remove may be the one Lidarr maps next.
|
||||
func TestResolveDuplicates_LeavesAnAlbumAloneWhileLidarrRescans_Integration(t *testing.T) {
|
||||
f := newResolveFixture(t)
|
||||
lid := &fakeLidarr{
|
||||
album: lidarr.LidarrAlbum{ID: 4723, ForeignAlbumID: "rg-humanz"},
|
||||
releases: []lidarr.AlbumRelease{
|
||||
{ID: 10148, TrackCount: 2, Monitored: true}, {ID: 10149, TrackCount: 1},
|
||||
},
|
||||
releaseTracks: map[int][]lidarr.ReleaseTrack{10148: rt("Saturnz Barz", "Saturnz Barz"), 10149: rt("Saturnz Barz")},
|
||||
}
|
||||
if _, err := ResolveDuplicates(context.Background(), f.pool, nil, "", lid, true); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
lid.unmapped = []string{lidarrPath(f.cleanPath), lidarrPath(f.ripPath)}
|
||||
res, err := ResolveDuplicates(context.Background(), f.pool, nil, "", lid, true)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if res.Merged != 0 || !exists(f.ripPath) || !exists(f.cleanPath) {
|
||||
t.Errorf("merged %d during the rescan window", res.Merged)
|
||||
}
|
||||
}
|
||||
|
||||
// Without Lidarr, or with auto-resolve off, the pass classifies and removes
|
||||
// nothing.
|
||||
func TestResolveDuplicates_ClassifiesOnlyWithoutLidarrOrPermission_Integration(t *testing.T) {
|
||||
for name, run := range map[string]func(resolveFixture) (DuplicateResolveResult, error){
|
||||
"lidarr disabled": func(f resolveFixture) (DuplicateResolveResult, error) {
|
||||
return ResolveDuplicates(context.Background(), f.pool, nil, "", nil, true)
|
||||
},
|
||||
"lidarr unreachable": func(f resolveFixture) (DuplicateResolveResult, error) {
|
||||
return ResolveDuplicates(context.Background(), f.pool, nil, "", &fakeLidarr{unmappedErr: errors.New("timeout")}, true)
|
||||
},
|
||||
"auto-resolve off": func(f resolveFixture) (DuplicateResolveResult, error) {
|
||||
lid := &fakeLidarr{unmapped: []string{lidarrPath(f.ripPath)}}
|
||||
return ResolveDuplicates(context.Background(), f.pool, nil, "", lid, false)
|
||||
},
|
||||
} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
f := newResolveFixture(t)
|
||||
res, err := run(f)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if res.Merged != 0 || !exists(f.ripPath) {
|
||||
t.Errorf("merged %d", res.Merged)
|
||||
}
|
||||
status, class, _, _ := f.group(t)
|
||||
if status != "pending" || class == nil || *class != string(ClassSameRelease) {
|
||||
t.Errorf("group = (%s, %v), want pending and classified", status, class)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// The guard itself: a merge asked to remove a copy the check refuses changes
|
||||
// nothing, file or row.
|
||||
func TestMergeDuplicateGroupGuarded_RefusesATrackedCopy_Integration(t *testing.T) {
|
||||
f := newResolveFixture(t)
|
||||
refuse := func(p string) bool { return p != f.ripPath }
|
||||
_, err := MergeDuplicateGroupGuarded(context.Background(), f.pool, nil, "", f.groupID, f.clean.ID, refuse)
|
||||
if !errors.Is(err, ErrCopyTrackedByLidarr) {
|
||||
t.Fatalf("err = %v, want ErrCopyTrackedByLidarr", err)
|
||||
}
|
||||
if !exists(f.ripPath) {
|
||||
t.Error("the refused copy's file was removed")
|
||||
}
|
||||
if status, _, _, _ := f.group(t); status != "pending" {
|
||||
t.Errorf("group status = %s, want pending", status)
|
||||
}
|
||||
}
|
||||
@@ -13,6 +13,36 @@ type SurvivorCandidate struct {
|
||||
FileFormat string
|
||||
FileSize int64
|
||||
AddedAt time.Time
|
||||
|
||||
// LidarrTracked: Lidarr maps this file to a track of the release it
|
||||
// monitors. Removing it opens a hole Lidarr downloads again (M498), so a
|
||||
// tracked copy outranks everything else. False when Lidarr was not asked.
|
||||
LidarrTracked bool
|
||||
// TagFit is TagFitScore: how well the copy's tags and name fit the album.
|
||||
TagFit int
|
||||
}
|
||||
|
||||
// TagFitScore counts the signs that a copy belongs where it is filed (M498
|
||||
// #5436), one point each:
|
||||
// - it has a track number that no other track on its album also claims
|
||||
// - its file name carries no video-rip marker (SourceMarkersFor)
|
||||
// - it has a MusicBrainz recording id
|
||||
//
|
||||
// Two copies of one song differ in these far more often than in anything a
|
||||
// listener hears: the Humanz rips differ by a few bytes, and file size picked
|
||||
// the wrong one in 6 of 21 groups.
|
||||
func TagFitScore(hasTrackNumber, positionClash bool, filePath string, hasMBID bool) int {
|
||||
score := 0
|
||||
if hasTrackNumber && !positionClash {
|
||||
score++
|
||||
}
|
||||
if len(SourceMarkersFor(filePath)) == 0 {
|
||||
score++
|
||||
}
|
||||
if hasMBID {
|
||||
score++
|
||||
}
|
||||
return score
|
||||
}
|
||||
|
||||
// losslessFormats are the scanned extensions that are lossless by definition.
|
||||
@@ -26,6 +56,9 @@ var losslessFormats = map[string]bool{"flac": true, "wav": true}
|
||||
// the report shows it and the merge (#3911) lets the operator choose another.
|
||||
//
|
||||
// In order:
|
||||
// 0. the copy Lidarr maps, then the copy whose tags fit the album best
|
||||
// (TagFitScore). Both are zero for every copy when the caller does not
|
||||
// know them, and the order falls through to the quality rules below
|
||||
// 1. lossless over lossy — the one difference no later step can recover
|
||||
// 2. the larger file — for one recording at one duration that is the higher
|
||||
// bitrate. The scanner does not record bitrate (tracks.bitrate is never
|
||||
@@ -48,6 +81,10 @@ func ProposeSurvivor(cands []SurvivorCandidate) (trackID, reason string) {
|
||||
// runner-up — the rule that actually decided, not every rule it passed.
|
||||
next := ranked[1]
|
||||
switch {
|
||||
case best.LidarrTracked != next.LidarrTracked:
|
||||
return best.TrackID, "the copy Lidarr tracks"
|
||||
case best.TagFit != next.TagFit:
|
||||
return best.TrackID, "tags fit the album"
|
||||
case isLossless(best) != isLossless(next):
|
||||
return best.TrackID, "lossless (" + strings.ToLower(best.FileFormat) + ")"
|
||||
case best.FileSize != next.FileSize:
|
||||
@@ -60,6 +97,12 @@ func ProposeSurvivor(cands []SurvivorCandidate) (trackID, reason string) {
|
||||
}
|
||||
|
||||
func survivorBefore(a, b SurvivorCandidate) bool {
|
||||
if a.LidarrTracked != b.LidarrTracked {
|
||||
return a.LidarrTracked
|
||||
}
|
||||
if a.TagFit != b.TagFit {
|
||||
return a.TagFit > b.TagFit
|
||||
}
|
||||
if isLossless(a) != isLossless(b) {
|
||||
return isLossless(a)
|
||||
}
|
||||
|
||||
@@ -89,3 +89,66 @@ func TestProposeSurvivor_ReasonIsTheDecidingRule(t *testing.T) {
|
||||
t.Fatalf("got (%q, %q), want (flac-big, largest file)", id, reason)
|
||||
}
|
||||
}
|
||||
|
||||
// M498 #5436: the copy Lidarr maps outranks everything, then the copy whose
|
||||
// tags fit the album; quality decides only between copies equal on both. Each
|
||||
// case is built so the older rules would choose the other copy, so a reorder
|
||||
// that drops the new rules fails.
|
||||
func TestProposeSurvivor_LidarrThenTagFit(t *testing.T) {
|
||||
at := time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC)
|
||||
cases := []struct {
|
||||
name string
|
||||
cands []SurvivorCandidate
|
||||
wantID string
|
||||
wantReason string
|
||||
}{
|
||||
{
|
||||
name: "tracked beats lossless and better tags",
|
||||
cands: []SurvivorCandidate{
|
||||
{TrackID: "flac", FileFormat: "flac", FileSize: 40_000_000, AddedAt: at, TagFit: 3},
|
||||
{TrackID: "tracked", FileFormat: "mp3", FileSize: 5_000_000, AddedAt: at, TagFit: 1, LidarrTracked: true},
|
||||
},
|
||||
wantID: "tracked", wantReason: "the copy Lidarr tracks",
|
||||
},
|
||||
{
|
||||
// The Humanz shape: two rips a few bytes apart, one carrying a
|
||||
// video marker and clashing with another track's position.
|
||||
name: "tag fit beats a few bytes",
|
||||
cands: []SurvivorCandidate{
|
||||
{TrackID: "rip", FileFormat: "mp3", FileSize: 5_104_498, AddedAt: at, TagFit: 1},
|
||||
{TrackID: "clean", FileFormat: "mp3", FileSize: 5_104_492, AddedAt: at, TagFit: 3},
|
||||
},
|
||||
wantID: "clean", wantReason: "tags fit the album",
|
||||
},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
id, reason := ProposeSurvivor(tc.cands)
|
||||
if id != tc.wantID || reason != tc.wantReason {
|
||||
t.Fatalf("ProposeSurvivor = (%q, %q), want (%q, %q)", id, reason, tc.wantID, tc.wantReason)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestTagFitScore(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
hasTrackNumber bool
|
||||
clash bool
|
||||
path string
|
||||
hasMBID bool
|
||||
want int
|
||||
}{
|
||||
{"clean, numbered, identified", true, false, "/music/A/B/A - B - 01 - Song.flac", true, 3},
|
||||
{"position taken by another track", true, true, "/music/A/B/A - B - 01 - Song.flac", true, 2},
|
||||
{"no track number", false, false, "/music/A/B/Song.flac", true, 2},
|
||||
{"video rip", true, false, "/music/A/B/A - B - 01 - Song (Official Video).mp3", true, 2},
|
||||
{"rip with nothing else", false, false, "/music/A/B/Song (Official Video).mp3", false, 0},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := TagFitScore(tc.hasTrackNumber, tc.clash, tc.path, tc.hasMBID); got != tc.want {
|
||||
t.Errorf("%s: TagFitScore = %d, want %d", tc.name, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -55,6 +55,10 @@ type FingerprintSettings struct {
|
||||
AcousticMaxBitErrorRate float64
|
||||
BackfillConcurrency int32
|
||||
SweepIntervalHours int32
|
||||
// AutoResolve lets the duplicate resolver (M498) act on its own: merge
|
||||
// copies Lidarr does not need and change a Lidarr release that lists songs
|
||||
// twice. Off, it only classifies, and every group waits for the operator.
|
||||
AutoResolve bool
|
||||
// UpdatedAt is when the settings were last saved. Set by the database;
|
||||
// ignored by Set.
|
||||
UpdatedAt time.Time
|
||||
@@ -68,6 +72,7 @@ var DefaultFingerprintSettings = FingerprintSettings{
|
||||
AcousticMaxBitErrorRate: defaultAcousticMaxBitErrorRate,
|
||||
BackfillConcurrency: fingerprintBackfillConcurrency,
|
||||
SweepIntervalHours: defaultSweepIntervalHours,
|
||||
AutoResolve: true,
|
||||
}
|
||||
|
||||
// ErrFingerprintSettingOutOfRange is returned by Set for a value migration
|
||||
@@ -121,6 +126,7 @@ func (s *FingerprintSettingsService) Set(ctx context.Context, in FingerprintSett
|
||||
AcousticMaxBitErrorRate: in.AcousticMaxBitErrorRate,
|
||||
BackfillConcurrency: in.BackfillConcurrency,
|
||||
SweepIntervalHours: in.SweepIntervalHours,
|
||||
AutoResolve: in.AutoResolve,
|
||||
})
|
||||
if err != nil {
|
||||
return FingerprintSettings{}, fmt.Errorf("fingerprint settings: save: %w", err)
|
||||
@@ -158,6 +164,7 @@ func fingerprintSettingsFromRow(row dbq.FingerprintSetting) FingerprintSettings
|
||||
AcousticMaxBitErrorRate: row.AcousticMaxBitErrorRate,
|
||||
BackfillConcurrency: row.BackfillConcurrency,
|
||||
SweepIntervalHours: row.SweepIntervalHours,
|
||||
AutoResolve: row.AutoResolve,
|
||||
UpdatedAt: row.UpdatedAt.Time,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -102,6 +102,7 @@ func TestFingerprintSettingsService_Integration(t *testing.T) {
|
||||
AcousticMaxBitErrorRate: minAcousticMaxBitErrorRate,
|
||||
BackfillConcurrency: minBackfillConcurrency,
|
||||
SweepIntervalHours: minSweepIntervalHours,
|
||||
AutoResolve: false,
|
||||
}
|
||||
highest := FingerprintSettings{
|
||||
Enabled: true,
|
||||
@@ -109,6 +110,7 @@ func TestFingerprintSettingsService_Integration(t *testing.T) {
|
||||
AcousticMaxBitErrorRate: maxAcousticMaxBitErrorRate,
|
||||
BackfillConcurrency: maxBackfillConcurrency,
|
||||
SweepIntervalHours: maxSweepIntervalHours,
|
||||
AutoResolve: true,
|
||||
}
|
||||
for _, step := range []struct {
|
||||
name string
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
package library
|
||||
|
||||
import (
|
||||
"path"
|
||||
"regexp"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// sourceMarker is one sign that a file was ripped from a video rather than
|
||||
// taken from a release: words a video title carries and an album track does
|
||||
// not (#5410).
|
||||
//
|
||||
// Each pattern is used twice, so it is written in the subset that Postgres's
|
||||
// regex flavour and Go's RE2 read the same way: no \b (Postgres spells word
|
||||
// boundaries \y), no lookaround, and brackets written as (\(|\[) rather than
|
||||
// as a bracket expression. The SQL side matches it case-insensitively with ~*,
|
||||
// the Go side with (?i); [^a-zA-Z] is spelled out so neither has to fold case
|
||||
// inside a negated class.
|
||||
type sourceMarker struct {
|
||||
label string
|
||||
pattern string
|
||||
}
|
||||
|
||||
// suspectSourceMarkers was calibrated against the operator's library on
|
||||
// 2026-10-08. Two candidates were left out on purpose:
|
||||
// - "live in"/"live at": 229 files, nearly all from real live albums.
|
||||
// - a bare "reaction": it caught Beck's "Chain Reaction". The marker below
|
||||
// wants the phrases a reaction video actually uses.
|
||||
var suspectSourceMarkers = []sourceMarker{
|
||||
{"music video", `official[^a-zA-Z]*(music[^a-zA-Z]*)?video|music[^a-zA-Z]*video`},
|
||||
{"official audio", `official[^a-zA-Z]*audio`},
|
||||
{"lyric video", `lyrics?[^a-zA-Z]*video`},
|
||||
{"visualiser", `visuali[sz]er`},
|
||||
{"reaction", `reaction[^a-zA-Z]*(video|mashup)|reacts?[^a-zA-Z]+to[^a-zA-Z]|first[^a-zA-Z]*time[^a-zA-Z]*(hearing|listening)`},
|
||||
{"MV", `(^|[^a-zA-Z])(mv|m/v)([^a-zA-Z]|$)`},
|
||||
{"[Audio]", `(\(|\[)audio(\)|\])`},
|
||||
{"[HD]", `(\(|\[)(hd|hq|4k)(\)|\])`},
|
||||
}
|
||||
|
||||
// SuspectSourcePattern is every marker as one alternation, for the SQL filter
|
||||
// behind the admin suspect-sources report.
|
||||
var SuspectSourcePattern = func() string {
|
||||
parts := make([]string, len(suspectSourceMarkers))
|
||||
for i, m := range suspectSourceMarkers {
|
||||
parts[i] = "(" + m.pattern + ")"
|
||||
}
|
||||
return strings.Join(parts, "|")
|
||||
}()
|
||||
|
||||
var suspectSourceRegexps = func() []*regexp.Regexp {
|
||||
out := make([]*regexp.Regexp, len(suspectSourceMarkers))
|
||||
for i, m := range suspectSourceMarkers {
|
||||
out[i] = regexp.MustCompile("(?i)" + m.pattern)
|
||||
}
|
||||
return out
|
||||
}()
|
||||
|
||||
// SourceMarkersFor returns the labels of every marker the file's basename
|
||||
// carries, in list order. Empty for a file the SQL filter would not return.
|
||||
func SourceMarkersFor(filePath string) []string {
|
||||
base := path.Base(filePath)
|
||||
labels := make([]string, 0, 2)
|
||||
for i, re := range suspectSourceRegexps {
|
||||
if re.MatchString(base) {
|
||||
labels = append(labels, suspectSourceMarkers[i].label)
|
||||
}
|
||||
}
|
||||
return labels
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
package library
|
||||
|
||||
import (
|
||||
"slices"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Real basenames from the operator's library (2026-10-08), each with the
|
||||
// labels it should carry. The negatives are the near misses that shaped the
|
||||
// list: a song called "Chain Reaction", a live album, an ordinary track.
|
||||
func TestSourceMarkersFor(t *testing.T) {
|
||||
cases := []struct {
|
||||
base string
|
||||
want []string
|
||||
}{
|
||||
{"Daft Punk - Random Access Memories - 06 - Daft Punk - Doin' It Right (Music Video) ft. Panda Bear.mp3", []string{"music video"}},
|
||||
{"Gorillaz - Humanz - 03 - Gorillaz - Saturnz Barz (Official Video).mp3", []string{"music video"}},
|
||||
{"Gorillaz - Humanz - 05 - Gorillaz - Andromeda (Official Audio).mp3", []string{"official audio"}},
|
||||
{"Gorillaz - Gorillaz - 06 - Gorillaz - P45 (Visualizer).mp3", []string{"visualiser"}},
|
||||
{"Gorillaz - Demon Days - 16 - Gorillaz - Don Quixote's Christmas Bonanza (Visualiser).mp3", []string{"visualiser"}},
|
||||
{"Artist - Album - 01 - Artist - Song (Lyric Video).mp3", []string{"lyric video"}},
|
||||
{"Gorillaz - Humanz - 02 - FIRST TIME HEARING Gorillaz - Ascension REACTION.mp3", []string{"reaction"}},
|
||||
{"Watsky - INTENTION - 05 - MANIAC Reacts to Watsky - AWW SHiT.mp3", []string{"reaction"}},
|
||||
{"米津玄師 - diorama - 09 - 【MV】米津玄師 - 恋と病熱.mp3", []string{"MV"}},
|
||||
{"Andora - Ego - 01 - Andora - Ego (feat. Will Stetson) MV.mp3", []string{"MV"}},
|
||||
{"Jimmy Eat World - Something(s) Loud - 05 - Jimmy Eat World - Call to Love (Audio).mp3", []string{"[Audio]"}},
|
||||
{"Record Heat - World War IV - 03 - Record Heat - Front Seat Feelin' [Audio].mp3", []string{"[Audio]"}},
|
||||
{"Aphex Twin - Come To Daddy - 03 - Aphex Twin - Bucephalus Bouncing Ball (HQ).mp3", []string{"[HD]"}},
|
||||
{"Artist - Album - 01 - Artist - Song (Official Lyric Video) [4K].mp3", []string{"lyric video", "[HD]"}},
|
||||
|
||||
{"Beck - Guero - 15 - Beck - Chain Reaction.mp3", nil},
|
||||
{"Nirvana - MTV Unplugged in New York - 01 - About a Girl (live in New York).flac", nil},
|
||||
{"Boards of Canada - Music Has the Right to Children - 05 - Roygbiv.flac", nil},
|
||||
{"Artist - Album - 01 - Mvula.flac", nil},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got := SourceMarkersFor("/music/x/" + c.base)
|
||||
if len(got) == 0 && len(c.want) == 0 {
|
||||
continue // nil and empty both mean "no markers"
|
||||
}
|
||||
if !slices.Equal(got, c.want) {
|
||||
t.Errorf("%s:\n got %v\n want %v", c.base, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The folder is not evidence: a marker word in a directory name must not
|
||||
// flag the files inside it.
|
||||
func TestSourceMarkersFor_ReadsOnlyTheBasename(t *testing.T) {
|
||||
if got := SourceMarkersFor("/music/Official Video Collection/01 - Song.flac"); len(got) != 0 {
|
||||
t.Errorf("directory name flagged the file: %v", got)
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user