Files
FabledCurator/backend/app/models/post.py
T
bvandeusenandClaude Opus 5 1e45e2c56c
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m6s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m59s
Build images / promote (push) Skipped
CI / integration (push) Failing after 2m7s
feat: an open grouping — a later drop joins its post (milestone 388 step E3)
A synthetic post is no longer sealed at creation. A creator who adds two more
variants the next day extends the existing post, its body grows with the new
messages, and no rival post appears. That is what makes chat capture read as
content trickling in rather than as a stream of separate arrivals.

The sweep now runs two passes per source and the ORDER is load-bearing: offer
new messages to still-open groups BEFORE founding new ones, because whichever
runs first claims a message.

E3's three named problems, each answered rather than discovered later:

**Bridging.** A candidate near two groups joins NEITHER. Nearest-wins would
silently make an arbitrary choice between two posts the operator may already
have seen; merging them is worse still, because a merge rewrites history and
anything pointing at the absorbed post dangles. Leaving it to found its own
group is the recoverable failure. AMBIGUITY_MARGIN is a module constant and
deliberately not a setting — it is not a quality dial anyone would tune toward
a better feed, and exposing it would invite turning it to zero, which is
exactly the silent arbitrary choice it prevents.

**Re-surfacing without thrashing.** A grouping has two dates, and which one
orders the feed is a real decision, so the feed orders by neither directly.
Ordering by when the drop STARTED buries a group that grows a week later under
a week of other posts — defeating the point of keeping it open. Ordering by
every growth lets a group gaining one image a day live permanently at the top,
so chat out-competes authored posts for the front page — the opposite of "post
pacing stays front and centre". Instead `resurfaced_at` moves only when growth
clears BOTH a minimum-images bar and a cooldown, so a drip-feed updates in
place and a genuine second wave resurfaces exactly once. It is NULL on every
ordinary post, so the sort key COALESCEs through it without moving anything
that is not a grouping.

**Reopening forever.** Groups close after a quiet period — artists reuse
characters for years, and a group left open indefinitely will eventually
absorb something it shouldn't. Openness is DERIVED, not stored: a group is
open if it grew (or started) within the window. Lowering the setting closes
old groups and raising it reopens them, with nothing to repair either way; a
stored closed_at would have needed a sweep to set it and a repair path to ever
change the policy.

Rule 89 is satisfied structurally rather than by a parallel mechanism:
celery_signals writes a TaskRun for every task, which already supplies
duration, the 5-minute stalled-run recovery, and retention pruning. What this
step owed on top of that was a wall-clock limit (present) and idempotence —
re-running the joiner adds nothing, asserted directly rather than left to the
unique (image, post) constraint to catch.

Two bugs fixed in the writing, one of which my own test would have hit:

* `assign_to_group` sorted bare (distance, Post) tuples, which falls through
  to comparing Posts when two distances tie — and a perfectly symmetric
  bridge, the exact case the function exists for, would have raised TypeError
  instead of declining to choose. Now keyed on the distance alone.
* The cursor was still built from `post_date or downloaded_at` while the
  ORDER BY had gained `resurfaced_at`. Two expressions that disagree at a page
  boundary don't error, they silently skip or repeat rows; both sites now go
  through one `_post_sort_value`, and a test pages through one row at a time
  to prove the walk matches the whole list.

Image linking is now one shared helper rather than written twice, because
creation and joining would otherwise be free to drift on exactly the detail
(which post owns the image) that makes a grouping reversible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:30:17 -04:00

163 lines
8.1 KiB
Python

"""Post — provenance anchor for one creator post (may contain many images).
`source_id` is nullable since alembic 0030 — filesystem-imported posts
with no live subscription have NULL source_id. `artist_id` is the
denormalized always-present link to the creator (added in 0030 so
artist-filter queries don't depend on the Source detour).
"""
from datetime import datetime
from sqlalchemy import (
JSON,
CheckConstraint,
DateTime,
ForeignKey,
Index,
Integer,
String,
Text,
UniqueConstraint,
func,
text,
)
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class Post(Base):
__tablename__ = "post"
__table_args__ = (
# alembic 0030. The comment above described this index; nothing declared
# it, so autogenerate proposed dropping it (#3275).
Index("uq_post_artist_external_id_null_source", "artist_id", "external_post_id",
unique=True, postgresql_where=text("source_id IS NULL")),
# Source-bound dedup. Postgres treats NULL != NULL so rows
# with source_id IS NULL aren't deduped by this constraint;
# the partial unique index `uq_post_artist_external_id_null_source`
# (created in alembic 0030) covers that case via
# (artist_id, external_post_id).
UniqueConstraint("source_id", "external_post_id", name="uq_post_source_external_id"),
CheckConstraint(
"translation_override IN ('auto', 'force', 'original')",
# Bare name: Base.metadata's naming convention prepends
# ck_<table>_. Pre-prefixing it here doubles the prefix — see
# alembic 0088, which renames the four constraints that shipped
# that way (#3275).
name="translation_override",
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
source_id: Mapped[int | None] = mapped_column(
ForeignKey("source.id", ondelete="SET NULL"), nullable=True, index=True
)
# Denormalized; always equals source.artist_id when source_id is set
# (the importer is responsible for keeping them consistent on insert).
# Filter queries (artist detail, artist-scoped posts feed) use this
# directly instead of joining through Source.
artist_id: Mapped[int] = mapped_column(
ForeignKey("artist.id", ondelete="CASCADE"),
nullable=False, index=True,
)
external_post_id: Mapped[str] = mapped_column(String(128), nullable=False)
post_url: Mapped[str | None] = mapped_column(Text, nullable=True)
post_title: Mapped[str | None] = mapped_column(Text, nullable=True)
post_date: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
raw_metadata: Mapped[dict | None] = mapped_column(JSON, nullable=True)
description: Mapped[str | None] = mapped_column(Text, nullable=True)
attachment_count: Mapped[int | None] = mapped_column(Integer, nullable=True)
# -- Post-text translation (milestone 143). Filled by the translate_posts
# sweep via the Interpreter LAN service so viewing is instant.
# translated_source_lang is the DETECTED original language; "en" (or a
# passthrough) means nothing to translate and the *_translated columns stay
# NULL. engine_version keys the Interpreter cache — re-runs are ~1ms and a
# model upgrade re-translates instead of serving stale.
post_title_translated: Mapped[str | None] = mapped_column(Text, nullable=True)
description_translated: Mapped[str | None] = mapped_column(Text, nullable=True)
translated_source_lang: Mapped[str | None] = mapped_column(
String(8), nullable=True
)
translation_engine_version: Mapped[str | None] = mapped_column(
String(128), nullable=True
)
translated_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# Sticky per-post override of the translation decision (milestone 155):
# 'auto' = the acceptance gate decides; 'force' = always store Interpreter's
# translation even below the confidence floor (rescue a skipped legit-foreign
# title); 'original' = never translate, keep the original (kill a confidently
# mis-flagged one the floor can't catch). The sweep reads this on every run,
# and re-translate leaves 'original' posts alone, so the choice survives a
# Re-translate-all.
translation_override: Mapped[str] = mapped_column(
String(16), nullable=False, default="auto", server_default="auto",
)
downloaded_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
# -- Synthetic posts (milestone 388). ----------------------------------
# Discord is a delivery CHANNEL, not a publisher: one message is not one
# post. So FC authors the post itself, grouping a creator's variant drop
# into a single row (services/discord_grouping.py).
#
# NULL for every post a creator actually wrote — which is all of them until
# a grouper runs. Non-NULL names the grouper that authored this row, and is
# the ONE flag the UI keys off to say so. The honesty rule is the whole
# point: a synthetic post must never present itself as authored, and a
# column that is absent-or-a-name makes "was this us?" answerable from the
# row rather than inferred from its shape.
#
# Plain String, no CHECK (rule 36 considered and declined) — same reasoning
# as source.error_type and service_seen.kind. There is exactly one grouper
# today; a second would be a value, not an invariant.
synthesized_by: Mapped[str | None] = mapped_column(String(32), nullable=True)
# What it was built from, so the operator can audit a grouping FC invented:
# member post ids, message count, and the thresholds in force when the
# decision was made. That last part matters — the thresholds are operator-
# tunable, so "why did it group these" is unanswerable a month later
# without recording the values that produced it.
synthesis_details: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# Set on a MEMBER post, pointing at the synthetic post that absorbed it.
# The feed hides absorbed posts (they are the chat lines the synthetic post
# replaced); every other surface still reaches them by id, because they
# remain the image's true origin and the grouping has to be inspectable.
#
# Self-FK, ON DELETE SET NULL: deleting a synthetic post un-absorbs its
# members and they return to the feed on their own. That is the reversal
# path, and it is one DELETE — nothing to undo by hand.
absorbed_by_post_id: Mapped[int | None] = mapped_column(
ForeignKey("post.id", ondelete="SET NULL"), nullable=True, index=True
)
# -- An OPEN grouping (milestone 388 E3) -------------------------------
# A synthetic post is not sealed at creation: a creator who adds two more
# variants the next day extends the existing post rather than starting a
# new one. These two columns are what make that possible without the post
# either freezing or thrashing the feed.
#
# `last_grew_at` is when the group last absorbed something. It answers two
# questions: how long the group stays JOINABLE (a group closes after a
# quiet period — artists reuse characters for years, and a group left open
# forever will eventually absorb something it shouldn't), and what the card
# shows as "updated N ago". NULL means it has never grown since creation.
last_grew_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# The feed position, and ONLY set when the anti-thrash rule fires — see
# discord_grouping.should_resurface. A group that gains one image a day
# must not sit permanently at the top of the feed, so growth updates the
# post without necessarily moving it; a genuine second wave moves it once.
#
# NULL on every ordinary post, which is why the feed's sort key can
# COALESCE through it without changing where anything else lands.
resurfaced_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)