Files
FabledCurator/backend/app/services/post_association_service.py
T
bvandeusenandClaude Opus 5 31de54e326
CI and images / lint (push) Failing after 2s
CI and images / extension-version (push) Successful in 2s
CI and images / frontend-build (push) Successful in 20s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Failing after 2m15s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
feat: FC links a conclusive pair itself instead of asking (4392)
Operator, 2026-09-24: *"I don't want this to be manual that defeats the
convenience that I'm going for."*

Confirm-only was right while every signal was circumstantial. Time proximity
and a body that mentions Discord can never be more than suggestive, so asking
was the honest response to what FC actually knew. A shared working name is
different in kind: when the creator's own name for a piece appears in exactly
these two posts and nowhere else in their library, there is nothing left for
the operator to adjudicate, and asking is a chore FC invented for them.

AUTO_LINK_FLOOR is 1.0 and sits deliberately above IDENTITY_FLOOR's 0.75. The
gap between them IS the review queue — real evidence, not certain enough for
FC to act on alone. Measured on artist 8, of 15 name-sharing pairs: 11 are
conclusive, 2 more propose, 2 fall short of both.

Three refusals, because an auto-link is FC asserting something the operator
never saw:

  * Exactly one candidate may be conclusive. Two is not a tie to be broken by
    score — one post can tease two pieces dropped separately, and then each
    drop carries a different name from the same teaser, each individually
    conclusive. Two conclusive answers to "which drop is this" means the
    question was wrong, so both queue and FC says nothing.
  * Neither end may already be claimed by an accepted link. That link is the
    operator's decision and reassigning its other end would overrule them
    silently.
  * The whole thing is one setting, defaulting on, reversible in the UI —
    an accepted link is a row they can dismiss.

`match_post` now returns (proposed, linked) so a sweep can report what it did
on its own rather than only what it queued.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-24 08:14:04 -04:00

550 lines
23 KiB
Python

"""The announcement matcher: which Patreon post announced which Discord drop.
Milestone 388, step E5.
Two of the operator's artists post a deliberately CROPPED fragment on Patreon
to signal that the real thing has landed in their Discord. This service
proposes those pairs, and proposes only — the operator accepts or dismisses,
following the FC-6.3 series matcher (task 737) rather than linking on its own.
## Why confirm-only is not caution for its own sake
A wrongly-asserted association tells the operator that two different pieces are
one. That is strictly worse than no link at all: no link leaves them exactly
where they already were, a wrong one actively misinforms and then propagates
into whatever reads the association. So the matcher's job is to make a SHORT
list worth reading, not a long list worth trusting.
## Two routes, because the evidence is of two different kinds
CIRCUMSTANTIAL evidence says two things happened near each other. It is
additive, weighted, and no single one of its signals may reach the threshold:
1. **Time proximity.** The Patreon post exists in order to announce the drop,
so the two are minutes-to-hours apart. Nearly free, and strong.
2. **The post says so.** These announcements routinely name Discord or carry
an invite link, which is close to a declaration.
3. **A shared marker.** The creator's own tie-back — `🍈🍈` in the Patreon
title and `@everyone 🍈 🍈` in the Discord message — gated on how rare that
marker is in THIS artist's posts, because a habitual emoji is punctuation.
IDENTITY evidence says two things are the same thing, and it gets its own
route (see `IDENTITY_FLOOR`):
4. **A shared working name.** The creator exports the teaser and the release
from one file, and the internal name survives into both platforms
untouched. Measured on the operator's artist: `ConnFront` ↔ `ConnFront`.
This is the only signal that reaches a pair 23.8 hours apart, which
proximity scores at 0.005.
## The one deliberately NOT built
**Crop-to-source matching is HELD, on the plan's own instruction** — it is
real work with real false-positive risk, and it is only worth building once
the cheap signals are shown to be insufficient against the operator's actual
artists. Half of this creator's recent teasers are screenshots carrying no
working name at all, and those pairs are out of reach here; that, measured, is
what would justify it.
Note also that a naive whole-image SigLIP similarity is NOT that signal. A
cropped teaser and its full version are exactly the pair a whole-image
comparison handles worst, so adding one as a "bonus" would mostly add noise
while looking like progress.
## Creator identity comes free, so E4 is not actually a prerequisite
The plan listed E4 (creator identity across the two channels) as a dependency.
It is not one for the pairs that matter today: a Patreon `Source` and a Discord
`Source` the operator has added under the same `Artist` already share
`Post.artist_id`, and the synthetic grouping inherits it (E2). E4 EXTENDS this
to creators whose association FC has to learn rather than being told; it is not
needed to represent an association FC already knows.
## A correction the plan carried, worth restating
`link_extract.py` does exist and does capture off-platform links — but only
file hosts (`SUPPORTED_HOSTS` is mega/gdrive/mediafire/dropbox/pixeldrain).
`host_for()` returns None for a Discord URL, so no `ExternalLink` row is ever
written for one. The declaration signal therefore reads the post body itself
rather than the extracted-links table the plan assumed it could use.
"""
from __future__ import annotations
import logging
import re
from collections import Counter
from dataclasses import dataclass
from datetime import UTC, datetime, timedelta
from sqlalchemy import and_, func, or_, select
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ImageRecord, ImportSettings, Post, PostAssociation
from ..utils.text import html_to_plain
from .discord_grouping import DROP_GROUPER
from .post_naming import (
IDENTITY_FLOOR,
marker_frequencies,
marker_overlap,
shared_identity,
token_frequencies,
working_name_tokens,
)
log = logging.getLogger(__name__)
# Additive weights, summing to 1.0. Kept as constants rather than settings —
# the sensitivity knob that matters is the threshold, and per-signal weights
# are an over-tune (same call as series_match_service.WEIGHTS).
#
# THE RELATIONSHIP TO THE THRESHOLD IS THE DESIGN. No single weight may reach
# the default threshold, which is what makes "time proximity alone is never
# enough" arithmetic rather than aspirational: on a busy day an artist posts
# several times, and a matcher that could pair on proximity alone would turn
# every busy day into false pairs. A guard test pins this.
WEIGHTS = {"proximity": 0.45, "declared": 0.35, "marker": 0.20}
# A Discord INVITE in the body is close to a declaration; the bare word is
# weaker but still meaningful, because these posts are short and on-topic.
#
# "the server" and its possessives are here because the word `discord` is NOT
# how these creators actually write. Measured across 20,558 Patreon bodies:
# `discord` appears in 486 and `the server` in 37 — but the distribution is the
# point, not the totals. For the artist this step was built for, 21 of 42 posts
# say `discord` and 7 say `the server`, and it is the RECENT ones that say the
# latter: the phrasing drifted once the audience already knew where the server
# was. A vocabulary list written from old posts silently stops matching.
_INVITE = re.compile(r"discord\.(?:gg|com/invite)/", re.I)
_MENTION = re.compile(r"\b(?:discord|(?:the|our|my)\s+server)\b", re.I)
DECLARED_INVITE = 1.0
DECLARED_MENTION = 0.6
MAX_CANDIDATES = 25
# How many just-grouped drops one sweep will look around. A ceiling on the
# `or_` the sweep builds, not a policy — a backfill that authors thousands of
# drops at once should not turn one sweep into a full-library rescan, which is
# the manual button's job.
MAX_RECENT_DROPS = 200
# What a shared name must reach before FC links a pair WITHOUT asking.
#
# 1.0, which under post_naming's post-span counting means the name appears in
# exactly these two posts and nowhere else in the artist's library. That is not
# "strong evidence" — within the library it is conclusive, and the remaining
# ways to be wrong are a mis-parse or the creator reusing a name for a genuinely
# different piece on the same day.
#
# Deliberately above IDENTITY_FLOOR, which is what a name needs to PROPOSE.
# The gap between them is the review queue: real evidence, not certain enough
# for FC to act on by itself. Measured on artist 8, 15 name-sharing pairs: 11
# are conclusive, 2 more propose, 2 fall short of both.
AUTO_LINK_FLOOR = 1.0
@dataclass(frozen=True)
class _Corpus:
"""One artist's rare-token evidence, gathered once rather than per pair.
Both rare-token signals are scoped to a single artist — a working name and
a marker belong to the person who chose them — so the counts are useless
across artists and expensive to rebuild per candidate. A sweep touches an
artist's posts many times over; this is loaded on the first touch and kept
for the life of the service.
"""
tokens_by_post: dict[int, set[str]]
token_posts: Counter[str]
text_by_post: dict[int, str]
marker_posts: Counter[str]
def proximity_signal(gap: timedelta, window: timedelta) -> float:
"""1.0 when the two posts are simultaneous, decaying linearly to 0 at the
window's edge. Linear rather than a step, so a pair an hour outside a
hand-tuned window degrades instead of vanishing."""
if window <= timedelta(0):
return 0.0
seconds = abs(gap.total_seconds())
if seconds >= window.total_seconds():
return 0.0
return round(1.0 - (seconds / window.total_seconds()), 4)
def declared_signal(description: str | None) -> float:
"""Does the announcement say, in its own body, that this is about Discord?
The invite is matched against the RAW body and the bare mention against the
stripped text, which is not fussiness — post bodies are HTML, and these
creators put the invite in an anchor's `href`. `html_to_plain` discards
attributes, so stripping first would have thrown away the strongest form of
the signal and left only whatever the link text happened to say.
The mention still reads stripped text, so `\bdiscord\b` is matched against
prose rather than against markup and URLs, where it would fire on any link
that merely passes through a discord domain.
"""
if not description:
return 0.0
if _INVITE.search(description):
return DECLARED_INVITE
if _MENTION.search(html_to_plain(description) or ""):
return DECLARED_MENTION
return 0.0
def weighted_score(signals: dict) -> float:
return round(sum(WEIGHTS[k] * signals.get(k, 0.0) for k in WEIGHTS), 4)
def _post_time(post: Post) -> datetime:
return post.post_date or post.downloaded_at
class PostAssociationService:
def __init__(self, session: AsyncSession):
self.session = session
self._corpora: dict[int, _Corpus] = {}
async def _corpus(self, artist_id: int) -> _Corpus:
if artist_id in self._corpora:
return self._corpora[artist_id]
paths_by_post: dict[int, list[str]] = {}
rows = await self.session.execute(
select(ImageRecord.primary_post_id, ImageRecord.path).where(
ImageRecord.artist_id == artist_id,
ImageRecord.primary_post_id.is_not(None),
)
)
for post_id, path in rows:
paths_by_post.setdefault(post_id, []).append(path)
text_by_post: dict[int, str] = {}
rows = await self.session.execute(
select(Post.id, Post.post_title, Post.description).where(
Post.artist_id == artist_id
)
)
for post_id, title, description in rows:
text_by_post[post_id] = "\n".join(
part for part in (title, html_to_plain(description) or "") if part
)
corpus = _Corpus(
tokens_by_post={
pid: {t for path in paths for t in working_name_tokens(path)}
for pid, paths in paths_by_post.items()
},
# Both counts take POSTS, which is why they are built from these
# groupings rather than from flat lists — see post_naming.
token_posts=token_frequencies(paths_by_post.values()),
text_by_post=text_by_post,
marker_posts=marker_frequencies(text_by_post.values()),
)
self._corpora[artist_id] = corpus
return corpus
async def _decided(self, announcement_id: int) -> set[int]:
"""Payload posts already proposed for this announcement, in ANY status.
Dismissed pairs are included deliberately: re-proposing a pair the
operator has already rejected on every subsequent scan is the single
behaviour that makes a review queue get ignored.
"""
rows = (await self.session.execute(
select(PostAssociation.payload_post_id)
.where(PostAssociation.announcement_post_id == announcement_id)
)).scalars().all()
return set(rows)
async def _candidate_groups(
self, announcement: Post, *, window: timedelta,
) -> list[Post]:
"""Synthetic Discord groupings by the SAME artist, inside the window.
Same-artist is the identity signal and it is free (see the module
docstring on E4). It is also a hard filter rather than a scored one:
two different creators posting minutes apart is a coincidence, not
evidence, and letting it score at all would mean a busy hour across the
library could out-vote everything else.
"""
at = _post_time(announcement)
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
return (await self.session.execute(
select(Post)
.where(
Post.artist_id == announcement.artist_id,
Post.synthesized_by == DROP_GROUPER,
Post.id != announcement.id,
sort_key >= at - window,
sort_key <= at + window,
)
.order_by(sort_key)
.limit(MAX_CANDIDATES)
)).scalars().all()
async def _claimed(self, announcement_id: int, payload_id: int) -> bool:
"""Is either end of this pair already spoken for by an accepted link?
An auto-link is FC asserting something the operator never saw, so it
only happens where there is nothing to contradict. A drop already
linked to a different announcement is exactly such a contradiction, and
resolving it is a judgement about which one is right — which is the
operator's, not FC's.
"""
return (await self.session.execute(
select(PostAssociation.id).where(
PostAssociation.status == "linked",
or_(
PostAssociation.payload_post_id == payload_id,
PostAssociation.announcement_post_id == announcement_id,
),
).limit(1)
)).scalar() is not None
async def match_post(
self, announcement_id: int, *, threshold: float, window_hours: float,
auto_link: bool = False,
) -> tuple[int, int]:
"""Score one announcement against nearby groupings.
Returns `(proposed, linked)` — how many pairs were written, and how
many of those were linked outright rather than queued.
"""
announcement = await self.session.get(Post, announcement_id)
if announcement is None or announcement.synthesized_by is not None:
# A synthetic post cannot announce anything — FC wrote it.
return 0, 0
window = timedelta(hours=window_hours)
declared = declared_signal(announcement.description)
already = await self._decided(announcement_id)
corpus = await self._corpus(announcement.artist_id)
here = corpus.tokens_by_post.get(announcement.id, set())
here_text = corpus.text_by_post.get(announcement.id, "")
made = 0
scored: list[tuple[Post, float, dict, float]] = []
for group in await self._candidate_groups(announcement, window=window):
if group.id in already:
continue
identity, token = shared_identity(
here,
corpus.tokens_by_post.get(group.id, set()),
corpus.token_posts,
)
circumstantial = {
"proximity": proximity_signal(
_post_time(group) - _post_time(announcement), window,
),
"declared": declared,
"marker": marker_overlap(
here_text,
corpus.text_by_post.get(group.id, ""),
corpus.marker_posts,
),
}
score = weighted_score(circumstantial)
# THE TWO ROUTES, and why identity is not simply a fourth weight.
#
# Circumstance and identity answer different questions. Proximity
# and a declaration say two things happened near each other and
# that one of them mentioned Discord; a working name the creator
# uses on these two posts and nowhere else says they are the same
# piece. Averaging those makes the threshold uninterpretable, and
# it costs both: adding identity as a weight dilutes the others
# enough that measured teaser/drop pairs an hour apart stop
# proposing, while capping identity's contribution at its weight
# means the strongest evidence available can never carry a pair on
# its own.
#
# So identity may override, never dilute. Below the floor it is
# recorded for the operator to read and moves nothing — which is
# the conservative direction, since a wrong link asserts that two
# different pieces are one.
if identity >= IDENTITY_FLOOR:
score = max(score, identity)
if score < threshold:
continue
signals = {**circumstantial, "identity": identity}
if token:
# Carried so the queue can say WHY. A review queue that cannot
# explain itself is one the operator learns to click through.
signals["identity_token"] = token
scored.append((group, score, signals, identity))
# Who, if anyone, FC links without asking.
#
# EXACTLY ONE candidate may be conclusive. Two drops sharing a name
# with one teaser at full strength is not a tie to be broken by score —
# it means the name identifies something other than what FC thinks it
# does, and the right response is to queue both and say nothing.
auto_id = None
if auto_link:
conclusive = [c for c in scored if c[3] >= AUTO_LINK_FLOOR]
if len(conclusive) == 1 and not await self._claimed(
announcement.id, conclusive[0][0].id
):
auto_id = conclusive[0][0].id
linked = 0
for group, score, signals, identity in scored:
status = "linked" if group.id == auto_id else "pending"
linked += status == "linked"
self.session.add(PostAssociation(
announcement_post_id=announcement.id,
payload_post_id=group.id,
score=score,
signals=signals,
status=status,
))
made += 1
return made, linked
async def list_pending(self) -> list[dict]:
rows = (await self.session.execute(
select(PostAssociation)
.where(PostAssociation.status == "pending")
.order_by(PostAssociation.score.desc(), PostAssociation.id.desc())
)).scalars().all()
return [
{
"id": a.id,
"announcement_post_id": a.announcement_post_id,
"payload_post_id": a.payload_post_id,
"score": a.score,
"signals": a.signals,
}
for a in rows
]
async def accept(self, association_id: int) -> dict | None:
a = await self.session.get(PostAssociation, association_id)
if a is None:
return None
a.status = "linked"
return {"id": a.id, "status": a.status}
async def dismiss(self, association_id: int) -> dict | None:
a = await self.session.get(PostAssociation, association_id)
if a is None:
return None
# Kept, not deleted — the row is what remembers the rejection.
a.status = "dismissed"
return {"id": a.id, "status": a.status}
async def linked_for(self, post_ids: list[int]) -> dict[int, list[dict]]:
"""Accepted links touching these posts, keyed by post id, BOTH ways.
A post is either end of the relationship, and each end wants the other
one: the teaser wants "the full set is over here", the grouping wants
"this is what announced me". One query, both directions.
"""
if not post_ids:
return {}
rows = (await self.session.execute(
select(PostAssociation).where(
PostAssociation.status == "linked",
or_(
PostAssociation.announcement_post_id.in_(post_ids),
PostAssociation.payload_post_id.in_(post_ids),
),
)
)).scalars().all()
out: dict[int, list[dict]] = {}
for a in rows:
if a.announcement_post_id in post_ids:
out.setdefault(a.announcement_post_id, []).append(
{"role": "announces", "post_id": a.payload_post_id, "id": a.id}
)
if a.payload_post_id in post_ids:
out.setdefault(a.payload_post_id, []).append(
{"role": "announced_by", "post_id": a.announcement_post_id, "id": a.id}
)
return out
async def rescan(session: AsyncSession, *, now: datetime | None = None) -> dict:
"""Score every recent non-synthetic post against nearby groupings."""
settings = await ImportSettings.load(session)
if not settings.discord_link_enabled:
return {"enabled": False, "scanned": 0, "proposed": 0, "linked": 0}
now = now or datetime.now(UTC)
window_hours = float(settings.discord_link_window_hours)
window = timedelta(hours=window_hours)
# Only look at announcements that could still have a partner in range —
# a full-library rescan is the manual button's job, not the sweep's.
horizon = now - timedelta(hours=window_hours * 2)
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
ids = set((await session.execute(
select(Post.id).where(
Post.synthesized_by.is_(None),
Post.absorbed_by_post_id.is_(None),
sort_key >= horizon,
)
)).scalars().all())
# ...and announcements sitting next to a drop FC has only JUST authored.
#
# A drop's `post_date` is backdated to its first message, but FC cannot
# write the drop until the message has an embedding and the hourly grouper
# has run — so a drop created this minute can land weeks back in the feed.
# Its neighbours were last swept before it existed, and a sweep keyed only
# on how recent the ANNOUNCEMENT is will never look at them again.
#
# That is #4392's third cause, and it is the one that left a measured 0.800
# pair with an empty review queue on the live instance. The other two were
# about scoring; this one meant nothing was scored at all.
drop_times = (await session.execute(
select(sort_key).where(
Post.synthesized_by == DROP_GROUPER,
func.coalesce(Post.last_grew_at, Post.downloaded_at) >= horizon,
)
.order_by(func.coalesce(Post.last_grew_at, Post.downloaded_at).desc())
.limit(MAX_RECENT_DROPS)
)).scalars().all()
# The interval arithmetic is done in Python rather than SQL: a handful of
# literal ranges is portable, and `now - INTERVAL` is not.
ranges = [
and_(sort_key >= at - window, sort_key <= at + window)
for at in drop_times
if at is not None
]
if ranges:
ids |= set((await session.execute(
select(Post.id).where(
Post.synthesized_by.is_(None),
Post.absorbed_by_post_id.is_(None),
or_(*ranges),
)
)).scalars().all())
svc = PostAssociationService(session)
proposed = 0
linked = 0
# Sorted because `ids` is now a union of two queries: set iteration order
# is arbitrary, and a sweep that visits posts in a different order each
# run is one whose failures cannot be reproduced.
for pid in sorted(ids):
made, auto = await svc.match_post(
pid,
threshold=float(settings.discord_link_threshold),
window_hours=window_hours,
auto_link=bool(settings.discord_link_auto),
)
proposed += made
linked += auto
log.info(
"discord announcement matcher: scanned %d post(s), proposed %d pair(s), "
"linked %d outright",
len(ids), proposed, linked,
)
return {
"enabled": True, "scanned": len(ids), "proposed": proposed,
"linked": linked,
}