CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 34s
CI and images / integration (push) Successful in 2m39s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 5s
CI and images / smoke-web (push) Canceled after 0s
CI and images / promote (push) Canceled after 0s
#4392's third cause, and the only one of the three about whether a pair is SCORED AT ALL rather than how well. A drop's post_date is backdated to its first message, but FC cannot author the drop until that message has an embedding and the hourly grouper has run. So a drop created this minute lands wherever its messages were — days or weeks back in the feed. The sweep only looked at announcements published within twice the window of NOW, so by the time the drop existed its neighbours were already outside the horizon, and nothing brought the sweep back to them. Measured on the live instance: a pair scoring 0.800 with `associations: []` and an empty queue. The two scoring fixes in the previous commit would not have helped it, because nothing scored it. So the sweep now also gathers announcements sitting beside any drop whose `last_grew_at`/`downloaded_at` is inside the horizon — the drop's own clock rather than its backdated position. Capped at MAX_RECENT_DROPS so a backfill authoring thousands at once does not quietly become the full-library rescan the manual button exists for, and the id set is sorted before the loop because a sweep visiting posts in a different order each run is one whose failures cannot be reproduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
483 lines
20 KiB
Python
483 lines
20 KiB
Python
"""The announcement matcher: which Patreon post announced which Discord drop.
|
|
|
|
Milestone 388, step E5.
|
|
|
|
Two of the operator's artists post a deliberately CROPPED fragment on Patreon
|
|
to signal that the real thing has landed in their Discord. This service
|
|
proposes those pairs, and proposes only — the operator accepts or dismisses,
|
|
following the FC-6.3 series matcher (task 737) rather than linking on its own.
|
|
|
|
## Why confirm-only is not caution for its own sake
|
|
|
|
A wrongly-asserted association tells the operator that two different pieces are
|
|
one. That is strictly worse than no link at all: no link leaves them exactly
|
|
where they already were, a wrong one actively misinforms and then propagates
|
|
into whatever reads the association. So the matcher's job is to make a SHORT
|
|
list worth reading, not a long list worth trusting.
|
|
|
|
## Two routes, because the evidence is of two different kinds
|
|
|
|
CIRCUMSTANTIAL evidence says two things happened near each other. It is
|
|
additive, weighted, and no single one of its signals may reach the threshold:
|
|
|
|
1. **Time proximity.** The Patreon post exists in order to announce the drop,
|
|
so the two are minutes-to-hours apart. Nearly free, and strong.
|
|
2. **The post says so.** These announcements routinely name Discord or carry
|
|
an invite link, which is close to a declaration.
|
|
3. **A shared marker.** The creator's own tie-back — `🍈🍈` in the Patreon
|
|
title and `@everyone 🍈 🍈` in the Discord message — gated on how rare that
|
|
marker is in THIS artist's posts, because a habitual emoji is punctuation.
|
|
|
|
IDENTITY evidence says two things are the same thing, and it gets its own
|
|
route (see `IDENTITY_FLOOR`):
|
|
|
|
4. **A shared working name.** The creator exports the teaser and the release
|
|
from one file, and the internal name survives into both platforms
|
|
untouched. Measured on the operator's artist: `ConnFront` ↔ `ConnFront`.
|
|
This is the only signal that reaches a pair 23.8 hours apart, which
|
|
proximity scores at 0.005.
|
|
|
|
## The one deliberately NOT built
|
|
|
|
**Crop-to-source matching is HELD, on the plan's own instruction** — it is
|
|
real work with real false-positive risk, and it is only worth building once
|
|
the cheap signals are shown to be insufficient against the operator's actual
|
|
artists. Half of this creator's recent teasers are screenshots carrying no
|
|
working name at all, and those pairs are out of reach here; that, measured, is
|
|
what would justify it.
|
|
|
|
Note also that a naive whole-image SigLIP similarity is NOT that signal. A
|
|
cropped teaser and its full version are exactly the pair a whole-image
|
|
comparison handles worst, so adding one as a "bonus" would mostly add noise
|
|
while looking like progress.
|
|
|
|
## Creator identity comes free, so E4 is not actually a prerequisite
|
|
|
|
The plan listed E4 (creator identity across the two channels) as a dependency.
|
|
It is not one for the pairs that matter today: a Patreon `Source` and a Discord
|
|
`Source` the operator has added under the same `Artist` already share
|
|
`Post.artist_id`, and the synthetic grouping inherits it (E2). E4 EXTENDS this
|
|
to creators whose association FC has to learn rather than being told; it is not
|
|
needed to represent an association FC already knows.
|
|
|
|
## A correction the plan carried, worth restating
|
|
|
|
`link_extract.py` does exist and does capture off-platform links — but only
|
|
file hosts (`SUPPORTED_HOSTS` is mega/gdrive/mediafire/dropbox/pixeldrain).
|
|
`host_for()` returns None for a Discord URL, so no `ExternalLink` row is ever
|
|
written for one. The declaration signal therefore reads the post body itself
|
|
rather than the extracted-links table the plan assumed it could use.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import logging
|
|
import re
|
|
from collections import Counter
|
|
from dataclasses import dataclass
|
|
from datetime import UTC, datetime, timedelta
|
|
|
|
from sqlalchemy import and_, func, or_, select
|
|
from sqlalchemy.ext.asyncio import AsyncSession
|
|
|
|
from ..models import ImageRecord, ImportSettings, Post, PostAssociation
|
|
from ..utils.text import html_to_plain
|
|
from .discord_grouping import DROP_GROUPER
|
|
from .post_naming import (
|
|
IDENTITY_FLOOR,
|
|
marker_frequencies,
|
|
marker_overlap,
|
|
shared_identity,
|
|
token_frequencies,
|
|
working_name_tokens,
|
|
)
|
|
|
|
log = logging.getLogger(__name__)
|
|
|
|
# Additive weights, summing to 1.0. Kept as constants rather than settings —
|
|
# the sensitivity knob that matters is the threshold, and per-signal weights
|
|
# are an over-tune (same call as series_match_service.WEIGHTS).
|
|
#
|
|
# THE RELATIONSHIP TO THE THRESHOLD IS THE DESIGN. No single weight may reach
|
|
# the default threshold, which is what makes "time proximity alone is never
|
|
# enough" arithmetic rather than aspirational: on a busy day an artist posts
|
|
# several times, and a matcher that could pair on proximity alone would turn
|
|
# every busy day into false pairs. A guard test pins this.
|
|
WEIGHTS = {"proximity": 0.45, "declared": 0.35, "marker": 0.20}
|
|
|
|
# A Discord INVITE in the body is close to a declaration; the bare word is
|
|
# weaker but still meaningful, because these posts are short and on-topic.
|
|
#
|
|
# "the server" and its possessives are here because the word `discord` is NOT
|
|
# how these creators actually write. Measured across 20,558 Patreon bodies:
|
|
# `discord` appears in 486 and `the server` in 37 — but the distribution is the
|
|
# point, not the totals. For the artist this step was built for, 21 of 42 posts
|
|
# say `discord` and 7 say `the server`, and it is the RECENT ones that say the
|
|
# latter: the phrasing drifted once the audience already knew where the server
|
|
# was. A vocabulary list written from old posts silently stops matching.
|
|
_INVITE = re.compile(r"discord\.(?:gg|com/invite)/", re.I)
|
|
_MENTION = re.compile(r"\b(?:discord|(?:the|our|my)\s+server)\b", re.I)
|
|
|
|
DECLARED_INVITE = 1.0
|
|
DECLARED_MENTION = 0.6
|
|
|
|
MAX_CANDIDATES = 25
|
|
|
|
# How many just-grouped drops one sweep will look around. A ceiling on the
|
|
# `or_` the sweep builds, not a policy — a backfill that authors thousands of
|
|
# drops at once should not turn one sweep into a full-library rescan, which is
|
|
# the manual button's job.
|
|
MAX_RECENT_DROPS = 200
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class _Corpus:
|
|
"""One artist's rare-token evidence, gathered once rather than per pair.
|
|
|
|
Both rare-token signals are scoped to a single artist — a working name and
|
|
a marker belong to the person who chose them — so the counts are useless
|
|
across artists and expensive to rebuild per candidate. A sweep touches an
|
|
artist's posts many times over; this is loaded on the first touch and kept
|
|
for the life of the service.
|
|
"""
|
|
|
|
tokens_by_post: dict[int, set[str]]
|
|
token_posts: Counter[str]
|
|
text_by_post: dict[int, str]
|
|
marker_posts: Counter[str]
|
|
|
|
|
|
def proximity_signal(gap: timedelta, window: timedelta) -> float:
|
|
"""1.0 when the two posts are simultaneous, decaying linearly to 0 at the
|
|
window's edge. Linear rather than a step, so a pair an hour outside a
|
|
hand-tuned window degrades instead of vanishing."""
|
|
if window <= timedelta(0):
|
|
return 0.0
|
|
seconds = abs(gap.total_seconds())
|
|
if seconds >= window.total_seconds():
|
|
return 0.0
|
|
return round(1.0 - (seconds / window.total_seconds()), 4)
|
|
|
|
|
|
def declared_signal(description: str | None) -> float:
|
|
"""Does the announcement say, in its own body, that this is about Discord?
|
|
|
|
The invite is matched against the RAW body and the bare mention against the
|
|
stripped text, which is not fussiness — post bodies are HTML, and these
|
|
creators put the invite in an anchor's `href`. `html_to_plain` discards
|
|
attributes, so stripping first would have thrown away the strongest form of
|
|
the signal and left only whatever the link text happened to say.
|
|
|
|
The mention still reads stripped text, so `\bdiscord\b` is matched against
|
|
prose rather than against markup and URLs, where it would fire on any link
|
|
that merely passes through a discord domain.
|
|
"""
|
|
if not description:
|
|
return 0.0
|
|
if _INVITE.search(description):
|
|
return DECLARED_INVITE
|
|
if _MENTION.search(html_to_plain(description) or ""):
|
|
return DECLARED_MENTION
|
|
return 0.0
|
|
|
|
|
|
def weighted_score(signals: dict) -> float:
|
|
return round(sum(WEIGHTS[k] * signals.get(k, 0.0) for k in WEIGHTS), 4)
|
|
|
|
|
|
def _post_time(post: Post) -> datetime:
|
|
return post.post_date or post.downloaded_at
|
|
|
|
|
|
class PostAssociationService:
|
|
def __init__(self, session: AsyncSession):
|
|
self.session = session
|
|
self._corpora: dict[int, _Corpus] = {}
|
|
|
|
async def _corpus(self, artist_id: int) -> _Corpus:
|
|
if artist_id in self._corpora:
|
|
return self._corpora[artist_id]
|
|
|
|
paths_by_post: dict[int, list[str]] = {}
|
|
rows = await self.session.execute(
|
|
select(ImageRecord.primary_post_id, ImageRecord.path).where(
|
|
ImageRecord.artist_id == artist_id,
|
|
ImageRecord.primary_post_id.is_not(None),
|
|
)
|
|
)
|
|
for post_id, path in rows:
|
|
paths_by_post.setdefault(post_id, []).append(path)
|
|
|
|
text_by_post: dict[int, str] = {}
|
|
rows = await self.session.execute(
|
|
select(Post.id, Post.post_title, Post.description).where(
|
|
Post.artist_id == artist_id
|
|
)
|
|
)
|
|
for post_id, title, description in rows:
|
|
text_by_post[post_id] = "\n".join(
|
|
part for part in (title, html_to_plain(description) or "") if part
|
|
)
|
|
|
|
corpus = _Corpus(
|
|
tokens_by_post={
|
|
pid: {t for path in paths for t in working_name_tokens(path)}
|
|
for pid, paths in paths_by_post.items()
|
|
},
|
|
# Both counts take POSTS, which is why they are built from these
|
|
# groupings rather than from flat lists — see post_naming.
|
|
token_posts=token_frequencies(paths_by_post.values()),
|
|
text_by_post=text_by_post,
|
|
marker_posts=marker_frequencies(text_by_post.values()),
|
|
)
|
|
self._corpora[artist_id] = corpus
|
|
return corpus
|
|
|
|
async def _decided(self, announcement_id: int) -> set[int]:
|
|
"""Payload posts already proposed for this announcement, in ANY status.
|
|
|
|
Dismissed pairs are included deliberately: re-proposing a pair the
|
|
operator has already rejected on every subsequent scan is the single
|
|
behaviour that makes a review queue get ignored.
|
|
"""
|
|
rows = (await self.session.execute(
|
|
select(PostAssociation.payload_post_id)
|
|
.where(PostAssociation.announcement_post_id == announcement_id)
|
|
)).scalars().all()
|
|
return set(rows)
|
|
|
|
async def _candidate_groups(
|
|
self, announcement: Post, *, window: timedelta,
|
|
) -> list[Post]:
|
|
"""Synthetic Discord groupings by the SAME artist, inside the window.
|
|
|
|
Same-artist is the identity signal and it is free (see the module
|
|
docstring on E4). It is also a hard filter rather than a scored one:
|
|
two different creators posting minutes apart is a coincidence, not
|
|
evidence, and letting it score at all would mean a busy hour across the
|
|
library could out-vote everything else.
|
|
"""
|
|
at = _post_time(announcement)
|
|
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
|
|
return (await self.session.execute(
|
|
select(Post)
|
|
.where(
|
|
Post.artist_id == announcement.artist_id,
|
|
Post.synthesized_by == DROP_GROUPER,
|
|
Post.id != announcement.id,
|
|
sort_key >= at - window,
|
|
sort_key <= at + window,
|
|
)
|
|
.order_by(sort_key)
|
|
.limit(MAX_CANDIDATES)
|
|
)).scalars().all()
|
|
|
|
async def match_post(
|
|
self, announcement_id: int, *, threshold: float, window_hours: float,
|
|
) -> int:
|
|
"""Score one announcement against nearby groupings. Returns proposals made."""
|
|
announcement = await self.session.get(Post, announcement_id)
|
|
if announcement is None or announcement.synthesized_by is not None:
|
|
# A synthetic post cannot announce anything — FC wrote it.
|
|
return 0
|
|
|
|
window = timedelta(hours=window_hours)
|
|
declared = declared_signal(announcement.description)
|
|
already = await self._decided(announcement_id)
|
|
corpus = await self._corpus(announcement.artist_id)
|
|
here = corpus.tokens_by_post.get(announcement.id, set())
|
|
here_text = corpus.text_by_post.get(announcement.id, "")
|
|
|
|
made = 0
|
|
for group in await self._candidate_groups(announcement, window=window):
|
|
if group.id in already:
|
|
continue
|
|
identity, token = shared_identity(
|
|
here,
|
|
corpus.tokens_by_post.get(group.id, set()),
|
|
corpus.token_posts,
|
|
)
|
|
circumstantial = {
|
|
"proximity": proximity_signal(
|
|
_post_time(group) - _post_time(announcement), window,
|
|
),
|
|
"declared": declared,
|
|
"marker": marker_overlap(
|
|
here_text,
|
|
corpus.text_by_post.get(group.id, ""),
|
|
corpus.marker_posts,
|
|
),
|
|
}
|
|
score = weighted_score(circumstantial)
|
|
# THE TWO ROUTES, and why identity is not simply a fourth weight.
|
|
#
|
|
# Circumstance and identity answer different questions. Proximity
|
|
# and a declaration say two things happened near each other and
|
|
# that one of them mentioned Discord; a working name the creator
|
|
# uses on these two posts and nowhere else says they are the same
|
|
# piece. Averaging those makes the threshold uninterpretable, and
|
|
# it costs both: adding identity as a weight dilutes the others
|
|
# enough that measured teaser/drop pairs an hour apart stop
|
|
# proposing, while capping identity's contribution at its weight
|
|
# means the strongest evidence available can never carry a pair on
|
|
# its own.
|
|
#
|
|
# So identity may override, never dilute. Below the floor it is
|
|
# recorded for the operator to read and moves nothing — which is
|
|
# the conservative direction, since a wrong link asserts that two
|
|
# different pieces are one.
|
|
if identity >= IDENTITY_FLOOR:
|
|
score = max(score, identity)
|
|
if score < threshold:
|
|
continue
|
|
signals = {**circumstantial, "identity": identity}
|
|
if token:
|
|
# Carried so the queue can say WHY. A review queue that cannot
|
|
# explain itself is one the operator learns to click through.
|
|
signals["identity_token"] = token
|
|
self.session.add(PostAssociation(
|
|
announcement_post_id=announcement.id,
|
|
payload_post_id=group.id,
|
|
score=score,
|
|
signals=signals,
|
|
status="pending",
|
|
))
|
|
made += 1
|
|
return made
|
|
|
|
async def list_pending(self) -> list[dict]:
|
|
rows = (await self.session.execute(
|
|
select(PostAssociation)
|
|
.where(PostAssociation.status == "pending")
|
|
.order_by(PostAssociation.score.desc(), PostAssociation.id.desc())
|
|
)).scalars().all()
|
|
return [
|
|
{
|
|
"id": a.id,
|
|
"announcement_post_id": a.announcement_post_id,
|
|
"payload_post_id": a.payload_post_id,
|
|
"score": a.score,
|
|
"signals": a.signals,
|
|
}
|
|
for a in rows
|
|
]
|
|
|
|
async def accept(self, association_id: int) -> dict | None:
|
|
a = await self.session.get(PostAssociation, association_id)
|
|
if a is None:
|
|
return None
|
|
a.status = "linked"
|
|
return {"id": a.id, "status": a.status}
|
|
|
|
async def dismiss(self, association_id: int) -> dict | None:
|
|
a = await self.session.get(PostAssociation, association_id)
|
|
if a is None:
|
|
return None
|
|
# Kept, not deleted — the row is what remembers the rejection.
|
|
a.status = "dismissed"
|
|
return {"id": a.id, "status": a.status}
|
|
|
|
async def linked_for(self, post_ids: list[int]) -> dict[int, list[dict]]:
|
|
"""Accepted links touching these posts, keyed by post id, BOTH ways.
|
|
|
|
A post is either end of the relationship, and each end wants the other
|
|
one: the teaser wants "the full set is over here", the grouping wants
|
|
"this is what announced me". One query, both directions.
|
|
"""
|
|
if not post_ids:
|
|
return {}
|
|
rows = (await self.session.execute(
|
|
select(PostAssociation).where(
|
|
PostAssociation.status == "linked",
|
|
or_(
|
|
PostAssociation.announcement_post_id.in_(post_ids),
|
|
PostAssociation.payload_post_id.in_(post_ids),
|
|
),
|
|
)
|
|
)).scalars().all()
|
|
out: dict[int, list[dict]] = {}
|
|
for a in rows:
|
|
if a.announcement_post_id in post_ids:
|
|
out.setdefault(a.announcement_post_id, []).append(
|
|
{"role": "announces", "post_id": a.payload_post_id, "id": a.id}
|
|
)
|
|
if a.payload_post_id in post_ids:
|
|
out.setdefault(a.payload_post_id, []).append(
|
|
{"role": "announced_by", "post_id": a.announcement_post_id, "id": a.id}
|
|
)
|
|
return out
|
|
|
|
|
|
async def rescan(session: AsyncSession, *, now: datetime | None = None) -> dict:
|
|
"""Score every recent non-synthetic post against nearby groupings."""
|
|
settings = await ImportSettings.load(session)
|
|
if not settings.discord_link_enabled:
|
|
return {"enabled": False, "scanned": 0, "proposed": 0}
|
|
|
|
now = now or datetime.now(UTC)
|
|
window_hours = float(settings.discord_link_window_hours)
|
|
window = timedelta(hours=window_hours)
|
|
# Only look at announcements that could still have a partner in range —
|
|
# a full-library rescan is the manual button's job, not the sweep's.
|
|
horizon = now - timedelta(hours=window_hours * 2)
|
|
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
|
|
ids = set((await session.execute(
|
|
select(Post.id).where(
|
|
Post.synthesized_by.is_(None),
|
|
Post.absorbed_by_post_id.is_(None),
|
|
sort_key >= horizon,
|
|
)
|
|
)).scalars().all())
|
|
|
|
# ...and announcements sitting next to a drop FC has only JUST authored.
|
|
#
|
|
# A drop's `post_date` is backdated to its first message, but FC cannot
|
|
# write the drop until the message has an embedding and the hourly grouper
|
|
# has run — so a drop created this minute can land weeks back in the feed.
|
|
# Its neighbours were last swept before it existed, and a sweep keyed only
|
|
# on how recent the ANNOUNCEMENT is will never look at them again.
|
|
#
|
|
# That is #4392's third cause, and it is the one that left a measured 0.800
|
|
# pair with an empty review queue on the live instance. The other two were
|
|
# about scoring; this one meant nothing was scored at all.
|
|
drop_times = (await session.execute(
|
|
select(sort_key).where(
|
|
Post.synthesized_by == DROP_GROUPER,
|
|
func.coalesce(Post.last_grew_at, Post.downloaded_at) >= horizon,
|
|
)
|
|
.order_by(func.coalesce(Post.last_grew_at, Post.downloaded_at).desc())
|
|
.limit(MAX_RECENT_DROPS)
|
|
)).scalars().all()
|
|
# The interval arithmetic is done in Python rather than SQL: a handful of
|
|
# literal ranges is portable, and `now - INTERVAL` is not.
|
|
ranges = [
|
|
and_(sort_key >= at - window, sort_key <= at + window)
|
|
for at in drop_times
|
|
if at is not None
|
|
]
|
|
if ranges:
|
|
ids |= set((await session.execute(
|
|
select(Post.id).where(
|
|
Post.synthesized_by.is_(None),
|
|
Post.absorbed_by_post_id.is_(None),
|
|
or_(*ranges),
|
|
)
|
|
)).scalars().all())
|
|
|
|
svc = PostAssociationService(session)
|
|
proposed = 0
|
|
# Sorted because `ids` is now a union of two queries: set iteration order
|
|
# is arbitrary, and a sweep that visits posts in a different order each
|
|
# run is one whose failures cannot be reproduced.
|
|
for pid in sorted(ids):
|
|
proposed += await svc.match_post(
|
|
pid,
|
|
threshold=float(settings.discord_link_threshold),
|
|
window_hours=window_hours,
|
|
)
|
|
log.info(
|
|
"discord announcement matcher: scanned %d post(s), proposed %d pair(s)",
|
|
len(ids), proposed,
|
|
)
|
|
return {"enabled": True, "scanned": len(ids), "proposed": proposed}
|