CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m19s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Counting files punished a piece for the one thing a working name is
guaranteed to do — have several exports. Measured across the operator's four
dual-platform artists: knuxy carries `p217` on four files spanning exactly two
posts, the Patreon post and the Discord drop, and ~200 comic-page tokens have
that shape. Every one scored half strength. Counting posts scores them 1.00
and still catches the habits, which span many posts rather than many files:
tamadaheijun's `comic2` spans 8, conto's `seth2` 5, `maid` 4. Knuxy's
ambiguous band drops from 268 tokens to 55.
Three false positives found the same way — by running the module over the real
library rather than reading it:
* The screenshot guard matches from the start of the stem, so it never fired
on the legacy `<post id>_media_<media id>_` era. `Screenshot 2025-07-27
182450ab` sailed through and contributed `2025-07-27` — the same-day date
collision this module's docstring exists to refuse.
* tamadaheijun names screenshots in Japanese. A guard that knows only the
English word is a guard for one artist.
* `timeline 3-0002` spans three unrelated posts. Keeping hyphens inside
tokens for `0-k`'s sake let the number survive whole, so an identity token
now has to contain a letter — which is the property behind the old
bare-year and all-digits rules anyway, and leaves `0-k`, `680lc` and `p59`
untouched.
Also stopwords for `the`, `gif`, `main`, `patreon`, `capture`, `timeline`,
each measured carrying a false match inside the admitted band.
IDENTITY_FLOOR names what a shared name must reach to propose a link with no
corroboration. At 0.75, 13 of artist 8's 15 name-sharing pairs clear it,
including the operator's own example. The two that do not are real pairs this
signal will not carry alone — the stated cost of refusing the four-post band,
where conto's `illustration9` and `maid` also sit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
379 lines
17 KiB
Python
379 lines
17 KiB
Python
"""The creator's own working name for a piece, recovered from a filename.
|
|
|
|
Milestone 388. Pure functions, no DB and no ML — the whole point is that this
|
|
signal is free.
|
|
|
|
## What this is for
|
|
|
|
Two of the operator's artists post a censored or cropped teaser on Patreon and
|
|
the real release in their Discord. Matching those by IMAGE is the pair a
|
|
whole-image comparison handles worst: the teaser is a crop with a censor bar,
|
|
which is exactly the local edit that moves a perceptual hash and blurs a
|
|
semantic embedding.
|
|
|
|
But the creator names both exports after the same internal working title, and
|
|
that name survives into both platforms untouched. Measured on the live instance
|
|
2026-09-24, artist 8:
|
|
|
|
01_((0-k <-> 0-k_base (1.3h apart)
|
|
01_680LC <-> 680LC_Border (21.0h apart)
|
|
01_cnni18x <-> cnni18x (21.5h apart)
|
|
|
|
Three pairs, no false positives, and **two of them are 21 hours apart** — far
|
|
enough that time proximity scores them ~0.10 and could never propose them. The
|
|
naming signal is orthogonal to the timing one: each finds pairs the other
|
|
cannot, which is why both are kept rather than one being tuned to cover both.
|
|
|
|
## Why a filename and not a perceptual hash
|
|
|
|
A shared working-name token is IDENTITY evidence — `680lc` appearing on both
|
|
platforms is not a coincidence. Proximity is CIRCUMSTANTIAL: it says two things
|
|
happened near each other, never that they are the same thing. The distinction
|
|
drives the weighting in `post_association_service`, and it is why a rare enough
|
|
token is allowed to carry a proposal on its own while no amount of circumstance
|
|
is.
|
|
|
|
## The one false-positive class found, and why the fix is shaped this way
|
|
|
|
A first pass matched `01_Screenshot 2026-08-13 000004` to
|
|
`Screenshot_2026-08-13_032144` on the token `2026-08-13`, twice.
|
|
|
|
A screenshot filename is a camera artifact. It carries no working name, and the
|
|
date inside it collides across platforms on the same day BY CONSTRUCTION — the
|
|
teaser and the release are posted the same day, so their screenshot names
|
|
always share a date token. That is a signal that fires precisely when it is
|
|
least informative.
|
|
|
|
So a filename with no working name contributes NOTHING, rather than the
|
|
plausible-looking date match it could be squeezed for. Re-run with that rule:
|
|
the same three true pairs, zero false. Half of this creator's recent teasers
|
|
are screenshots, and those pairs are simply out of this signal's reach — which
|
|
is where crop-to-source matching earns its cost, and nowhere else.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
from collections import Counter
|
|
from collections.abc import Iterable
|
|
from pathlib import PurePosixPath
|
|
|
|
# A screenshot name, on either platform. Patreon's importer writes
|
|
# `01_Screenshot 2026-08-13 000004`; gallery-dl's Discord naming writes
|
|
# `Screenshot_2026-09-22_003651`. Matched after the index/message prefixes are
|
|
# stripped, so both shapes reach this as a bare `Screenshot ...`.
|
|
# NOT `\b` after "shot": `\b` needs a word/non-word transition and `_` is a
|
|
# WORD character, so `Screenshot_2026-08-13_032144` — gallery-dl's Discord
|
|
# spelling — sailed straight past the guard while the space-separated Patreon
|
|
# spelling was caught. Found by running this against the live library rather
|
|
# than by reading it. Assert the next character is not a letter instead.
|
|
_SCREENSHOT = re.compile(r"^(?:screen[ _-]?shot(?![a-z])|\u30b9\u30af\u30ea\u30fc\u30f3\u30b7\u30e7\u30c3\u30c8|\u30b9\u30af\u30b7\u30e7)", re.I)
|
|
|
|
# The importer's per-post media index: `01_`, `02_`. Not part of any name.
|
|
_MEDIA_INDEX = re.compile(r"^\d{1,3}_")
|
|
|
|
# gallery-dl's Discord filename pattern (#3999):
|
|
# `{date:%Y%m%d}_{message_id}_{num:>02}_{filename}`.
|
|
_DISCORD_PREFIX = re.compile(r"^\d{8}_\d{6,}_\d{1,3}_")
|
|
|
|
# The legacy era (#4002): images sit FLAT at the artist root as
|
|
# `<post id>_media_<media id>_<name>`. Stripping it is not cosmetic — the
|
|
# SCREENSHOT guard below matches from the start of the stem, so while this
|
|
# prefix was left on, a legacy screenshot never looked like one. Measured on
|
|
# tamadaheijun: `109078417_media_334848471_Screenshot 2025-07-27 182450ab`
|
|
# sailed through and contributed `2025-07-27`, which is precisely the
|
|
# same-day date collision this module was built to refuse.
|
|
_LEGACY_PREFIX = re.compile(r"^\d+_media_\d+_")
|
|
|
|
# The importer's content-hash suffix, `__<10 hex>`, sometimes doubled on files
|
|
# that went through an older import era.
|
|
_HASH_SUFFIX = re.compile(r"(?:__[0-9a-f]{10})+$")
|
|
|
|
# Generic export decorations. Stripped as SUFFIXES so the stem survives:
|
|
# `cnni18x_wip3` and `cnni18x` must yield the same token, or a work-in-progress
|
|
# would never match the piece it became.
|
|
_DECORATION = re.compile(
|
|
r"(?:[_-]?(?:wip|base|final|alt|alts|edit|edits|border|clean|raw|hd|full|"
|
|
r"censored|uncensored|nsfw|sfw|ver|v)\d*)+$",
|
|
re.I,
|
|
)
|
|
|
|
# Tokens that carry no identity even when they survive the rules above.
|
|
_STOPWORDS = frozenset({
|
|
"img", "image", "untitled", "new", "test", "page", "final", "copy",
|
|
"post", "media", "file", "avatar", "cover", "banner", "icon", "splash",
|
|
# Each of these was MEASURED carrying a false match in the 3..6 frequency
|
|
# band, where the rarity gate still admits a token: `capture` across six
|
|
# unrelated knuxy posts, `the` and `patreon` out of legacy title-derived
|
|
# names, `main` out of `Anya Main CST`, `timeline` and `gif` out of
|
|
# tamadaheijun's exports.
|
|
"the", "gif", "main", "patreon", "capture", "timeline", "screenshot",
|
|
# Literally the string "None": issue #3999's Discord naming rendered
|
|
# `{user[name]}` as it for ~1,600 files, so it is the single most common
|
|
# "name" in the library and identifies nothing.
|
|
"none",
|
|
})
|
|
|
|
# A bare year: still needed for the TEXT signal, where words and numbers are
|
|
# tokenised separately.
|
|
_YEAR = re.compile(r"^(?:19|20)\d{2}$")
|
|
|
|
# An identity token must contain a LETTER. This replaces separate "all digits"
|
|
# and "bare year" rules with the property behind both, and it is the rule that
|
|
# holds once hyphens are kept inside tokens for `0-k`'s sake: without it
|
|
# `2025-07-27` and `3-0002` survive as single tokens, and both were measured
|
|
# linking unrelated posts — the second across three of them, out of
|
|
# tamadaheijun's `timeline 3-0002` exports.
|
|
#
|
|
# `0-k`, `680lc`, `cnni18x` and `p59` all keep a letter and are unaffected.
|
|
_HAS_LETTER = re.compile(r"[A-Za-z]")
|
|
|
|
MIN_TOKEN_LEN = 3
|
|
|
|
# A token appearing in more than this many of ONE ARTIST's POSTS is a habit,
|
|
# not an identity — a character name, a series tag, a recurring export preset.
|
|
#
|
|
# POSTS, not files, and the difference is not bookkeeping. Counting files
|
|
# punishes a piece for having many exports, which is the one thing a working
|
|
# name is GUARANTEED to do. Measured across the operator's four dual-platform
|
|
# artists: knuxy's comic pages carry `p217` on four files spread over exactly
|
|
# two posts — the Patreon post and the Discord drop — and file-counting scored
|
|
# every one of ~200 such tokens at half strength for it. Counting posts scores
|
|
# them 1.00 while still catching the real habits, which span many posts:
|
|
# tamadaheijun's `comic2` spans 8, conto's `seth2` 5, `maid` 4.
|
|
#
|
|
# Six rather than two, although two posts IS the shape of a teaser and its
|
|
# drop, because a creator legitimately revisits one working name: a wip post,
|
|
# then an alt, then the release. Measured on artist 8, `cnni18x` and `680lc`
|
|
# each span four posts and both are genuine. IDENTITY_FLOOR below is what
|
|
# decides how much span a link may carry on its own.
|
|
MAX_TOKEN_POSTS = 6
|
|
|
|
# A shared name at or above this strength is enough to propose a link with NO
|
|
# corroboration — it is identity evidence, and the whole reason this module
|
|
# exists is that identity survives where circumstance does not.
|
|
#
|
|
# 0.75 is a token spanning three posts or fewer. Measured on artist 8: of the
|
|
# 15 same-artist pairs that share a name, 13 clear this bar, including the
|
|
# operator's own example (`0-k`, three posts, 1.3h apart). The two that do not
|
|
# — `680lc` and `cnni18x`, four posts each, 21h apart — are real pairs this
|
|
# signal will not carry alone; they are the measured cost of not admitting the
|
|
# four-post band, where conto's `illustration9` and `maid` also sit.
|
|
IDENTITY_FLOOR = 0.75
|
|
|
|
|
|
def _strip_prefixes(stem: str) -> str:
|
|
"""Remove the framing each platform's importer adds around the real name."""
|
|
stem = _DISCORD_PREFIX.sub("", stem)
|
|
stem = _LEGACY_PREFIX.sub("", stem)
|
|
stem = _MEDIA_INDEX.sub("", stem)
|
|
return _HASH_SUFFIX.sub("", stem)
|
|
|
|
|
|
def working_name_tokens(path: str) -> set[str]:
|
|
"""The identity-bearing tokens in one image's filename.
|
|
|
|
Returns an EMPTY set for a name that carries no working title — a
|
|
screenshot, a bare number, a stopword. Empty means "no evidence", which the
|
|
caller must treat as silence rather than as a weak match; see the module
|
|
docstring for the false positive that rule exists for.
|
|
"""
|
|
stem = _strip_prefixes(PurePosixPath(path).stem)
|
|
if _SCREENSHOT.match(stem.strip()):
|
|
return set()
|
|
|
|
out: set[str] = set()
|
|
# Hyphens are kept INSIDE tokens — `0-k` is a real working name on the live
|
|
# instance, and splitting on hyphen would reduce it to a single character
|
|
# and then discard it for being too short.
|
|
for raw in re.split(r"[^0-9A-Za-z-]+", stem.lower()):
|
|
tok = _DECORATION.sub("", raw).strip("-")
|
|
if len(tok) < MIN_TOKEN_LEN:
|
|
continue
|
|
if tok in _STOPWORDS or not _HAS_LETTER.search(tok):
|
|
continue
|
|
out.add(tok)
|
|
return out
|
|
|
|
|
|
def token_frequencies(posts: Iterable[Iterable[str]]) -> Counter[str]:
|
|
"""How many of ONE ARTIST's POSTS each working-name token appears in.
|
|
|
|
Takes posts — each an iterable of that post's image paths — rather than a
|
|
flat list of paths, because the unit of the count is the post. See
|
|
MAX_TOKEN_POSTS for what that buys; the short version is that a piece with
|
|
six exports in one post has used its name once.
|
|
|
|
Scoped to the ARTIST, not the library: a working name belongs to the person
|
|
who chose it, and the same string can be one creator's piece and another's
|
|
boilerplate. Built once per artist per sweep, not per candidate pair.
|
|
"""
|
|
counts: Counter[str] = Counter()
|
|
for paths in posts:
|
|
counts.update({t for path in paths for t in working_name_tokens(path)})
|
|
return counts
|
|
|
|
|
|
def _rarity(freq: int, max_frequency: int) -> float:
|
|
"""Rarity of one token within an artist's own corpus, in [0, 1].
|
|
|
|
Shared by BOTH signals deliberately. They carried one formula each
|
|
until 2026-09-24, and the copies drifted: the filename signal grew a
|
|
frequency gate and the marker signal never did, so a creator's habitual
|
|
emoji scored the same 1.00 as a marker they had used twice. One
|
|
definition cannot drift from itself.
|
|
|
|
Full strength at 2 rather than 1: a genuine match means the token is on
|
|
at least two things, so demanding uniqueness would reject every real
|
|
pair. Decays to zero AT the cap rather than falling off it, so nothing
|
|
sits on a cliff edge.
|
|
"""
|
|
if freq <= 2:
|
|
return 1.0
|
|
if freq >= max_frequency:
|
|
return 0.0
|
|
return (max_frequency - freq) / (max_frequency - 2)
|
|
|
|
|
|
def shared_identity(
|
|
left: Iterable[str],
|
|
right: Iterable[str],
|
|
frequencies: Counter[str],
|
|
*,
|
|
max_frequency: int = MAX_TOKEN_POSTS,
|
|
) -> tuple[float, str | None]:
|
|
"""Strength in [0, 1] that two sets of filenames name the SAME piece.
|
|
|
|
Returns `(strength, token)` — the token is carried back so the proposal can
|
|
say WHY it was made. A review queue that cannot explain itself is one the
|
|
operator learns to click through without reading.
|
|
|
|
Strength is a function of the winning token's rarity within the artist's
|
|
own library, not of how many tokens matched. One decisive token beats three
|
|
vague ones, and a token appearing across forty of this artist's posts is a
|
|
habit rather than an identity however exactly it matches.
|
|
|
|
`frequencies` must be the POST counts from `token_frequencies`.
|
|
"""
|
|
shared = {t for t in set(left) & set(right) if frequencies.get(t, 0) <= max_frequency}
|
|
if not shared:
|
|
return 0.0, None
|
|
|
|
# The rarest shared token decides — one decisive token beats three vague
|
|
# ones.
|
|
token = min(shared, key=lambda t: (frequencies.get(t, 0), -len(t), t))
|
|
strength = round(_rarity(max(frequencies.get(token, 1), 1), max_frequency), 4)
|
|
# A token sitting exactly ON the cap decays to zero, and naming it anyway
|
|
# would hand the review queue a reason that carries no weight — "matched on
|
|
# loislanetb2", with nothing behind it. Measured: that token is on 6 of this
|
|
# artist's images. Report a token only when it is doing work.
|
|
return (strength, token) if strength > 0 else (0.0, None)
|
|
|
|
|
|
# --- the body/title signal ---------------------------------------------------
|
|
#
|
|
# The same idea applied to TEXT. The operator's example pair carries `🍈🍈` in
|
|
# the Patreon title and `@everyone 🍈 🍈` in the Discord message — a marker the
|
|
# creator uses to tie the two together, which no vocabulary list would predict.
|
|
#
|
|
# Rarity-gated, exactly as the filename signal is, and the gate is here because
|
|
# the first pass did NOT have one. Measured on artist 8, 300 posts:
|
|
#
|
|
# 💦 11 posts (4%) 🫴 6 🌰 5 🍗 5 🫣 4
|
|
#
|
|
# 💦 is punctuation for this creator — about one post in twenty-five. Ungated it
|
|
# scored a full 1.00 and was the DECIDING term in a proposal that proximity
|
|
# alone (0.441) could not carry. A habitual marker riding along with proximity
|
|
# is just proximity wearing a hat, which is the exact failure the matcher's
|
|
# threshold sits above 0.55 to prevent. The operator's 🍈🍈 is the opposite
|
|
# case: two posts, and they are the pair itself.
|
|
|
|
_WORD = re.compile(r"[0-9A-Za-z]{3,}")
|
|
# Anything outside the Basic Multilingual Plane's text ranges: emoji, symbols,
|
|
# kaomoji parts. These are the tokens creators actually use as markers, and
|
|
# they are rare enough in prose to be evidence on their own.
|
|
# U+1F000-1FAFF is the emoji planes; U+2190-2BFF covers arrows, dingbats and
|
|
# the miscellaneous-symbol blocks, which already contains U+2600-27BF.
|
|
_SYMBOL = re.compile(r"[\U0001F000-\U0001FAFF\u2190-\u2BFF]")
|
|
|
|
_COMMON_TEXT = frozenset({
|
|
"the", "and", "for", "you", "new", "out", "now", "this", "that", "with",
|
|
"everyone", "here", "post", "all", "art", "one", "get", "has", "are",
|
|
})
|
|
|
|
# A marker in more than this many of ONE ARTIST's posts is a signature, not a
|
|
# tie-back. A marker tying an announcement to its drop lands on two posts —
|
|
# the two.
|
|
#
|
|
# Tighter than MAX_TOKEN_POSTS. Both count posts, so the numbers are directly
|
|
# comparable and the gap between them is the claim being made: a working name
|
|
# is the creator's private label for one piece and may honestly recur as they
|
|
# revisit it, while a marker is public decoration and stops being evidence the
|
|
# moment it is reused. Measured on artist 8: 💦 spans 13 posts, the word
|
|
# "like" 43, and 🌗 — a real tie-back — exactly 2.
|
|
MAX_MARKER_POSTS = 4
|
|
|
|
|
|
def text_markers(text: str | None) -> set[str]:
|
|
"""Distinctive tokens in a post body or title: symbols, and rare-ish words.
|
|
|
|
Symbols count individually rather than as a run, so `🍈🍈` and `🍈 🍈` —
|
|
which is how the same marker appears on the two platforms — reduce to the
|
|
same token. Spacing is a platform's rendering, not the creator's intent.
|
|
"""
|
|
if not text:
|
|
return set()
|
|
out = {m.group(0) for m in _SYMBOL.finditer(text)}
|
|
out |= {
|
|
w.lower() for w in _WORD.findall(text)
|
|
if w.lower() not in _COMMON_TEXT and not _YEAR.match(w)
|
|
}
|
|
return out
|
|
|
|
|
|
def marker_frequencies(texts: Iterable[str | None]) -> Counter[str]:
|
|
"""How many of ONE ARTIST's posts each marker appears in.
|
|
|
|
Per POST, not per occurrence: a creator who repeats an emoji six times in
|
|
one body has used it once as far as identity goes. Scoped to the artist for
|
|
the same reason `token_frequencies` is — a marker is a personal habit, and
|
|
one creator's signature is another's whole vocabulary.
|
|
"""
|
|
counts: Counter[str] = Counter()
|
|
for t in texts:
|
|
counts.update(text_markers(t))
|
|
return counts
|
|
|
|
|
|
def marker_overlap(
|
|
left: str | None,
|
|
right: str | None,
|
|
frequencies: Counter[str],
|
|
*,
|
|
max_frequency: int = MAX_MARKER_POSTS,
|
|
) -> float:
|
|
"""Strength in [0, 1] that two texts share a DELIBERATE marker.
|
|
|
|
`frequencies` is required rather than defaulted to "no gate". An ungated
|
|
call is the bug this signature exists to make impossible to write by
|
|
accident, and a default would have kept it one keyword away.
|
|
|
|
Symbols weigh full and words a quarter, because prose shares words by
|
|
accident: a creator who writes "commission" in both posts on a Tuesday has
|
|
told us nothing that the timestamps did not already say.
|
|
|
|
There is no divisor. An earlier pass halved the total so that a long body
|
|
could not out-vote a short one, which the rarity gate now does properly —
|
|
and halving meant the operator's own 🍈🍈 pair, a marker on exactly two
|
|
posts, could reach only 0.5. One marker the creator uses nowhere else is
|
|
the whole signal, not half of it.
|
|
"""
|
|
shared = text_markers(left) & text_markers(right)
|
|
if not shared:
|
|
return 0.0
|
|
score = sum(
|
|
(1.0 if _SYMBOL.match(t) else 0.25) * _rarity(frequencies.get(t, 1), max_frequency)
|
|
for t in shared
|
|
)
|
|
return round(min(1.0, score), 4)
|