feat: a tick keeps looking back 30 days, so an EDITED post is reached (4386)
CI and images / extension-version (push) Successful in 3s
CI and images / lint (push) Successful in 3s
CI and images / frontend-build (push) Successful in 19s
CI and images / backend-lint-and-test (push) Successful in 40s
CI and images / integration (push) Failing after 2m17s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
CI and images / extension-version (push) Successful in 3s
CI and images / lint (push) Successful in 3s
CI and images / frontend-build (push) Successful in 19s
CI and images / backend-lint-and-test (push) Successful in 40s
CI and images / integration (push) Failing after 2m17s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Operator, 2026-09-23, on a Floppystack post: "this post has been updated as he implements hot fixes — any chance we have a way to scan for or see updated posts so we can update ours to match and pull the new attachments and pictures etc." The download half already worked: extract_media reads the media list off the LIVE feed response every walk, so a newly attached hotfix build is a ledger key we have never seen. Only REACHING the post was missing — a tick stopped after 20 contiguous already-have-it items, and a post edited three days after publication sits well below twenty. Not a bug in the early-out; a count cannot express "recent". The early-out now needs BOTH conditions: the run of seen items AND a post published before the horizon. Strictly a widening — window 0 is exactly the old behaviour, and no window can make a tick stop EARLIER than it used to, so a source paused for months still walks its whole unseen backlog. The horizon is a floor on how far to look, never a ceiling. Inside the window the post-record gate is bypassed too (write_post_record revisit=True): the body is re-read from the feed response already in hand, so a revisit costs zero requests, and a body that comes back empty writes NOTHING rather than blanking one a detail-fetch had filled. Revisits are kept out of the #862 body-drift canary's sample for the same reason — an empty revisit is healthy, and counting it would walk the alarm toward firing on good ticks. The run summary names what changed ("3 post(s) updated (5 new file(s))") with a line per post; the ask was to SEE updated posts, not only to end up with their bytes. download_revisit_days is a settings row, not a constant (rule 25) — how long a creator keeps editing is a property of the creator. Default 30, 0 turns it off. Migration 0108. Also corrects two stale docstrings: both clients described post_meta as feeding an Ingester.preview that no longer calls it. It had no consumer at all until this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
@@ -68,6 +68,16 @@ class ImportSettings(Base):
|
||||
Integer, nullable=False, default=90,
|
||||
server_default="90",
|
||||
)
|
||||
# How far back a routine tick keeps looking after it has run out of new
|
||||
# posts, so a creator who EDITS an older post to attach a hotfix build is
|
||||
# still reached (ingest_core.DEFAULT_REVISIT_DAYS carries the reasoning).
|
||||
# A knob rather than a constant because how long a creator keeps editing is
|
||||
# a property of the creator, not of FabledCurator: 0 turns the revisit off
|
||||
# and restores the pure count early-out.
|
||||
download_revisit_days: Mapped[int] = mapped_column(
|
||||
Integer, nullable=False, default=30,
|
||||
server_default="30",
|
||||
)
|
||||
download_failure_warning_threshold: Mapped[int] = mapped_column(
|
||||
Integer, nullable=False, default=5,
|
||||
server_default="5",
|
||||
|
||||
Reference in New Issue
Block a user