fix(scanner): read multi-value genre frames correctly — #2499
dhowden/tag's readTFrame splits ID3v2 null-separated multi-value text frames and rejoins them with the EMPTY string, so a file tagged "Alternative Rock" + "Rock" was stored as "Alternative RockRock". It also leaves bare numeric ID3v1 references unresolved, which is why the library showed genres like "4017" and "526617". This corrupted more than the browse axis added in #367: taste_profile.sql reads tracks.genre directly, so the welded tokens were entering the taste profile's tag vocabulary, and recommendation.sql/discover.sql were comparing them as single opaque tags. Genre counts were wrong everywhere. ffprobe is not a fix — ffmpeg's read_ttag calls decode_str once with no loop, keeping only the first value. Truncating multi-genre tags would blunt the similarity signal genre mainly feeds. So the TCON frame is now parsed directly (ID3v2.2/2.3/2.4, all four text encodings, per-frame and tag-level unsynchronisation, numeric and parenthesised ID3v1 references); everything else still comes from dhowden/tag. Values are stored ";"-delimited, which the read side already splits on, so no query changes. Existing rows are repaired without an operator-run rebuild: migration 0054 adds tracks.tag_read_version DEFAULT 0, below the scanner's current tagReadVersion, so the next scan re-reads tags it would otherwise skip on mtime. Such a re-read reuses the stored duration instead of re-running ffprobe, keeping a repair pass tag-read-bound rather than one fork+exec per file. Bumping the constant is how a future extraction fix reaches an existing library. Only ID3v2 is in scope — dhowden welds nowhere else. The Vorbis/MP4 repeated-field question is #2500, unproven and deliberately not built.
This commit is contained in:
@@ -9,11 +9,16 @@ import (
|
||||
|
||||
// genreCount is one row of the genre browse index (#367).
|
||||
//
|
||||
// Genres are the raw ID3 strings, split on [;,] but otherwise untouched — no
|
||||
// Genres are the tag's own strings, split on [;,] but otherwise untouched — no
|
||||
// case folding and no synonym mapping. So "Rock" and "rock" can both appear,
|
||||
// as can "Rock/Pop" alongside "Rock" and "Pop". That's deliberate for v1: the
|
||||
// alternative is a normalisation table to invent and maintain, and the raw
|
||||
// spread has to be visible before anyone can judge whether it's a problem.
|
||||
//
|
||||
// The first look at that spread found it dominated by welded tokens like
|
||||
// "Alternative RockRock" — the scanner's own bug, not the operator's tagging
|
||||
// (#2499). Judge the "is a taxonomy needed" question (#2468) only against a
|
||||
// library re-scanned since that fix.
|
||||
type genreCount struct {
|
||||
Genre string `json:"genre"`
|
||||
TrackCount int `json:"track_count"`
|
||||
|
||||
Reference in New Issue
Block a user