mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 12:55:00 +02:00
Tweet and tweet_thread entries store their payload under artifact_role `raw_tweet_json`, not `primary_media`. `build_summary_input` filtered strictly for `primary_media LIMIT 1`, so both cases silently failed with 'entry X has no primary_media artifact to summarize'. Threads compound the problem: the tweet scraper writes ONE json file per status, so even a fixed lookup that took the first row would summarize only the initial tweet and lose the rest of the conversation. Fixes: - New `load_summary_artifacts` helper returns every artifact for a role in insertion order. - For entity_kind `tweet` / `tweet_thread`, load all `raw_tweet_json` artifacts (falling back to `primary_media` for archives predating that role convention). - Iterate artifacts, extract text per file with the existing markdown/html/json branches, then join thread pieces with a `---` separator so the model sees a real paragraph break between statuses instead of one flowing document. Single-tweet entries produce one piece and the separator never renders. Non-tweet entries behave exactly as before. |
||
|---|---|---|
| .. | ||
| archivr-cli | ||
| archivr-core | ||
| archivr-server | ||