Files
FabledCurator/backend/app/services/wip_title.py
T
bvandeusenandClaude Opus 5.5 7b1570f2a5
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 20s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m35s
CI and images / smoke-web (push) Successful in 56s
CI and images / promote (push) Successful in 1s
feat: retire the sketch/doodle WIP title tier, and the review strip asks "Is this a WIP?" (milestone 430)
The soft tier (#1474) tagged `wip` on any post titled sketch, doodle or
scribble: 6,096 of the library's 8,876 wip tags. Its conflict audit flagged
most of them, because finished art scores >= 0.5 on some content head,
which filled the Gallery strip with 2,086 cards. A "sketch" is usually
finished work, so the operator retired the tier. The artist's own
"WIP" / "work in progress" title rule and human wip tags stay.

- Removed: the soft matcher, source and prefilter; the importer and backfill
  soft branches; soft_wip_conflict_audit with its task and beat; the
  wip_soft_title_tagging_enabled setting, API field and toggle; and
  wip_title_soft from _AUTO_SOURCES.
- Migration 0112: a soft tag the operator confirmed, or kept from the strip,
  becomes `manual`. Every other soft tag is deleted, along with the open
  review cards whose tag is gone. The column is dropped.
- Review strip (#4424): each card asks "Is this a WIP?" (or banner, and so on),
  and the buttons read "Is a WIP" / "Is not a WIP". The content tag it also
  scored on is shown as the reason it was flagged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 08:20:28 -04:00

102 lines
4.5 KiB
Python

"""Title-based WIP auto-tagging (task #1458).
Deterministic heuristic: when a post's TITLE explicitly declares work-in-progress
(the artist's own "WIP" / "work in progress" label), the ``wip`` system tag is
applied to that post's images — a cheap, high-precision complement to the
image-based ML ``wip`` head. WIP images are excluded from the Explore/gallery
browse (see gallery_service ``excluded_system_tags``), so honouring the artist's
own label keeps unfinished pieces out of the main browse right at import.
Precision over recall — a false WIP tag HIDES a finished post — so matching is
token-anchored: ``swipe`` / ``wiped`` / ``wiping`` never trip it (a letter on the
boundary blocks the match).
Sync-only: both consumers (the importer and the backfill Celery task) run on a
sync Session. Application is idempotent-additive (ON CONFLICT DO NOTHING) and
stamps a distinct ``image_tag.source`` so a later pass can tell where a wip tag
came from — the "manual" / "head_auto" / "ccip_auto" / "ml_accepted" provenance
family gains one member.
"""
import re
from sqlalchemy import select
from sqlalchemy.orm import Session
from ..models.tag import WIP_SYSTEM_TAG, Tag, image_tag
from .image_tag_apply import insert_image_tags
# image_tag.source stamped on title-heuristic WIP tags — distinct from the other
# apply sources so provenance stays legible and a future undo can target only these.
# Only the artist's own "WIP"/"work in progress" label counts — high-precision, so it
# trains the wip head. A sketch/doodle/scribble tier (#1474) was retired in milestone
# 430: a "sketch" is usually finished art, and its 6k tags flooded the review strip.
WIP_TITLE_SOURCE = "wip_title"
# A standalone "WIP" / "W.I.P" token, or the phrase "work in progress"
# (space/underscore/hyphen separated). The letter-boundary lookarounds are what
# make this precision-first: `s|wip|e`, `|wip|ed`, `|wip|ing` all have a letter
# abutting the token, so they're rejected. A trailing digit is allowed so
# "WIP2" (= WIP part 2) still matches.
_WIP_RE = re.compile(
r"(?<![A-Za-z])(?:w\.?i\.?p\.?|work[\s_-]+in[\s_-]+progress)(?![A-Za-z])",
re.IGNORECASE,
)
# Coarse SQL prefilter for the backfill sweep — narrows the post scan to rows that
# COULD match before the precise regex confirms. Case-insensitive ILIKE patterns.
# It MUST stay a SUPERSET of the regex or the sweep would silently miss posts.
WIP_TITLE_SQL_PREFILTER = ("%wip%", "%work%progress%")
# Chunk bulk inserts so a large sweep can't blow past psycopg's 65535-parameter
# ceiling (3 params/row → ~21k rows max; 5k stays comfortably under).
_INSERT_CHUNK = 5000
def matches_wip_title(title: str | None) -> bool:
"""True when a post title explicitly marks it work-in-progress."""
if not title:
return False
return _WIP_RE.search(title) is not None
def resolve_wip_tag_id(session: Session) -> int | None:
"""The seeded ``wip`` system tag's id (migration 0075), or None if absent."""
return session.execute(
select(Tag.id).where(Tag.name == WIP_SYSTEM_TAG, Tag.is_system.is_(True))
).scalar_one_or_none()
def apply_wip_image_tags(
session: Session, image_ids, tag_id: int, *, source: str = WIP_TITLE_SOURCE
) -> int:
"""Attach ``tag_id`` (stamped with ``source``) to each image id, idempotently —
never disturbs an existing tag or its source. Returns the number of image_tag
rows newly inserted. Does NOT commit.
The insert count is computed from a pre-SELECT of already-tagged ids rather
than the statement's ``rowcount``: psycopg reports -1 for a multi-row
ON CONFLICT DO NOTHING insert (it runs via an executemany path), so rowcount
is unusable here. The SELECT is accurate within this single transaction (no
concurrent writer touches these (image, wip) rows); ON CONFLICT DO NOTHING
stays as a race-safety belt so a rare concurrent insert can't error."""
ids = list({int(i) for i in image_ids})
if not ids:
return 0
inserted = 0
for start in range(0, len(ids), _INSERT_CHUNK):
chunk = ids[start:start + _INSERT_CHUNK]
already = set(session.execute(
select(image_tag.c.image_record_id)
.where(image_tag.c.tag_id == tag_id)
.where(image_tag.c.image_record_id.in_(chunk))
).scalars())
to_insert = [iid for iid in chunk if iid not in already]
if not to_insert:
continue
insert_image_tags(session, [
{"image_record_id": iid, "tag_id": tag_id, "source": source}
for iid in to_insert
])
inserted += len(to_insert)
return inserted