1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
Commit graph

3 commits

Author SHA1 Message Date
archivr-qa
189fe2d392
fix(core): summarize tweets + walk all tweets in a thread
Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.

Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.

Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
  role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
  `raw_tweet_json` artifacts (falling back to `primary_media` for
  archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
  markdown/html/json branches, then join thread pieces with a
  `---` separator so the model sees a real paragraph break between
  statuses instead of one flowing document.

Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.
2026-08-23 19:02:29 +02:00
archivr-qa
a4fb2ed795
fix(core): codex_cli — auto-discover binary + use --output-last-message
Two related fixes for the codex_cli summary provider:

1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but
   without it the code resolved to bare `codex` and relied on PATH.
   The ChatGPT desktop app installs codex at
   `/Applications/ChatGPT.app/Contents/Resources/codex` and does not
   put it on PATH, so users who only have the desktop app saw
   'No such file or directory' with no hint. `resolve_cli` now walks
   env override → a small set of well-known absolute paths → HOME
   /.local/bin/<bare> → bare fallback. Same treatment applied to
   claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/
   claude, HOME/.local/bin/claude).

2. Clean output. `codex exec -` writes a runtime header ("OpenAI
   Codex vX", session id, sandbox, model), the assistant reply, and a
   footer ("tokens used", replay of the reply) to stdout. The JSON
   extractor took the first '{' from the *user prompt echo* and the
   last '}' from the trailing replay, producing invalid text that
   fell through to the "raw text under summary" fallback path. Now
   uses `--output-last-message <tempfile>` and reads only the final
   assistant message. Fallback (positional prompt) uses the same
   flag. Tempfile is cleaned up on all paths, incl. spawn failure.
2026-08-23 18:41:57 +02:00
fd3f6903e2
feat(core): add entry_summaries schema + summarizer trait/providers
Per-entry LLM summaries as a regenerable child record, not a column on
archived_entries and not an on-disk artifact: an entry may carry several
summaries (one per provider/model/prompt version), any of which can be
discarded and recomputed. Generation is manual-only — nothing in capture.rs
calls into this module.

- database.rs: entry_summaries table + index, EntrySummaryRecord, and
  upsert/update/find/latest helpers mirroring the capture_jobs style.
  provider_model is stored as '' rather than NULL because SQLite treats
  NULLs as distinct inside a UNIQUE index, which would stop the CLI
  providers (no model) from ever deduping on the cache key.
- summarizer.rs: SummaryProvider trait with four implementations —
  Anthropic Messages API, OpenAI-compatible chat completions, `claude -p`
  and `codex exec -`. Configuration comes from env vars only (never TOML),
  matching how yt-dlp / single-file / tweet-scraper are resolved, which
  also keeps API keys out of anything the archive persists.
- archive.rs: EntryDetail gains latest_summary, populated by one extra
  LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row
  rather than duplicating it.

Implementation notes:
- No tokio in core. CLI timeouts are enforced structurally: stdout is
  drained on its own thread and handed back over a channel so the calling
  thread can recv_timeout and kill an overrunning child; stdin is written
  on a third thread so a 48 KB prompt cannot deadlock against a child
  waiting for us to read.
- HTML is reduced with regex rather than a parser: html5ever is not in the
  tree, and a model tolerates imperfect whitespace. Paired tags are spelled
  out per tag because Rust's regex engine has no backreferences by design.
- reqwest is declared with only the `blocking` feature here, so bodies are
  serialized via .body(value.to_string()) instead of widening the
  workspace dependency for .json().
- input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one
  hashing primitive; the content is truncated to 48 KB *before* hashing so
  the cache key describes exactly the bytes the model saw.

Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for
one JSON shape is a poor trade, so the two halves that can actually break
are tested directly — request-body builders and response parsers — leaving
only reqwest's own transport uncovered. Plus schema idempotency, cache-key
dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML
and tweet extraction, output normalization, and the CLI runner's stdin
round-trip, timeout kill, and nonzero-exit paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:41:46 +02:00