From af6e473bdeb0e49574cbffbd930d0264af118719 Mon Sep 17 00:00:00 2001 From: archivr-qa Date: Sun, 23 Aug 2026 21:57:04 +0200 Subject: [PATCH] docs: explain X Article, image summaries and summary search --- AGENTS.md | 12 +++++++++++- ARCHIVR-MENTAL-MODEL.md | 19 ++++++++++++++----- docs/README.md | 24 +++++++++++++++++++----- 3 files changed, 44 insertions(+), 11 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 2ac8c2b..098232c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -28,6 +28,16 @@ LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/sr table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the provider set and the status lifecycle. +`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and +the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row. +Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each, +and 12 MiB in aggregate. Anthropic HTTP, OpenAI-compatible HTTP, and Codex support images; Claude CLI does not. The core +remains synchronous: the server puts provider work in its blocking boundary rather than introducing async to +`archivr-core`. + +Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only. +Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary. + Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8). The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`. @@ -87,7 +97,7 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy *pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. -- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. +- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`. - **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. - **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract. - **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums. diff --git a/ARCHIVR-MENTAL-MODEL.md b/ARCHIVR-MENTAL-MODEL.md index d3ffd54..00ab7a1 100644 --- a/ARCHIVR-MENTAL-MODEL.md +++ b/ARCHIVR-MENTAL-MODEL.md @@ -185,6 +185,13 @@ summarised by two providers, two models, or after a prompt change yields distinc request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`. +The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and +cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is +present, `SummaryBuildOptions::include_images` admits only bounded `media` image candidates and the digest includes both +the flag and selected blob identity, MIME type, and size. The core is still synchronous; the server owns the blocking +boundary. Anthropic HTTP, OpenAI-compatible HTTP, and Codex can transport the selected image data; Claude CLI receives +text only. + ```mermaid flowchart LR UI["ContextRail Summary"] -->|POST .../summary| Server @@ -196,13 +203,14 @@ flowchart LR UI -->|GET .../summary poll| Row2 ``` -**Tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads +**X Articles and tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads the single `primary_media` artifact. A `tweet` or `tweet_thread` entry has no `primary_media` — it has N `raw_tweet_json` artifacts, one per status in the thread. So the summarizer selects on the `raw_tweet_json` role instead, loads **all** matching artifacts in order, and joins them with `\n\n---\n\n`; a `---` line reads as a hard paragraph break to every model, keeping individual -statuses from bleeding into one another. Any change to how thread artifacts are stored has to be -mirrored here. +statuses from bleeding into one another. For an X Article, the reducer prefers article text over the tweet's body: +`plain_text`, then flattened ordered blocks, then `preview_text`, then `summary_text`; only then does it fall back to +`full_text`/`text`/`content`/`body`. Any change to article reduction or thread status storage must be mirrored here. ## yt-dlp Lifecycle @@ -272,7 +280,8 @@ The server both reads and writes archive data. Capture jobs are asynchronous: `P **Auth model.** A separate `archivr-auth.sqlite` (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are `u32` flags (`GUEST`, `USER`, `ADMIN`, `OWNER`) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is `setup_guard` → `login_rate_limit` → `security_headers`; route families are classified `READ / ADMIN / WRITE / STATIC` in `routes.rs`. -**Search** is client-side filtering over entries the frontend has already fetched. +**Search** is server-side free-text filtering over entry fields and the latest completed summary. The summary JSON is +searched as text, so generated `tags` participate. Older completed summaries stay searchable while a newer request is +pending or failed; rows with no completed summary contribute no summary-derived match. **Admin view** covers mounted archives, users, sessions, and API tokens. - diff --git a/docs/README.md b/docs/README.md index 6a2e841..901776e 100644 --- a/docs/README.md +++ b/docs/README.md @@ -44,11 +44,11 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y - **Web pages** — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode - **Local files** — import any file from disk by `file://` path - **Deduplication** — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once -- **Tags and search** — hierarchical tag tree, full-text search, filterable entry list +- **Tags and search** — hierarchical tag tree, full-text search (including the latest completed summary and its generated JSON tags), filterable entry list - **Multiple archives** — the server mounts any number of separate archives from a single TOML config - **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords - **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download -- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture +- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option - **Text notes** — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser - **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block @@ -220,8 +220,22 @@ The Nix wrapper and Docker image set `ARCHIVR_STATIC_DIR`, `ARCHIVR_SINGLE_FILE` #### LLM providers -Summaries are opt-in and provider-agnostic. Only the variables for the provider you actually select are read; the two -HTTP providers refuse to start without their API key. +Summaries are manual and provider-agnostic. Only the variables for the provider you actually select are read; the two +HTTP providers refuse to start without their API key. They are text-only by default. Selecting `Include attached images` +explicitly sends eligible archived image data to the chosen provider; it is never attached automatically. + +| Provider | Attached images | +|---|---| +| Anthropic HTTP | Supported | +| OpenAI-compatible HTTP | Supported | +| Codex CLI | Supported | +| Claude CLI | Not supported | + +Image inclusion considers only `media` artifacts with `jpg`, `jpeg`, `png`, `webp`, `gif`, or `avif` files. At most four +images are sent, each no larger than 5 MiB and no more than 12 MiB in total. + +Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no +summary, or only a pending or failed summary, get no summary-derived match. | Variable | Default | Description | |---|---|---| @@ -404,4 +418,4 @@ nix build .#archivr-server ## License MIT — see [LICENSE](../LICENSE.md). -\n \ No newline at end of file +\n