mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 12:55:00 +02:00
docs: explain X Article, image summaries and summary search
This commit is contained in:
parent
ecc25af05a
commit
af6e473bde
3 changed files with 44 additions and 11 deletions
12
AGENTS.md
12
AGENTS.md
|
|
@ -28,6 +28,16 @@ LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/sr
|
|||
table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the
|
||||
provider set and the status lifecycle.
|
||||
|
||||
`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and
|
||||
the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row.
|
||||
Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each,
|
||||
and 12 MiB in aggregate. Anthropic HTTP, OpenAI-compatible HTTP, and Codex support images; Claude CLI does not. The core
|
||||
remains synchronous: the server puts provider work in its blocking boundary rather than introducing async to
|
||||
`archivr-core`.
|
||||
|
||||
Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only.
|
||||
Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary.
|
||||
|
||||
Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8).
|
||||
|
||||
The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`.
|
||||
|
|
@ -87,7 +97,7 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
|
|||
*pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a
|
||||
self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison
|
||||
entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver.
|
||||
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH.
|
||||
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`.
|
||||
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
|
||||
- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract.
|
||||
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
|
||||
|
|
|
|||
|
|
@ -185,6 +185,13 @@ summarised by two providers, two models, or after a prompt change yields distinc
|
|||
request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`,
|
||||
mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`.
|
||||
|
||||
The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and
|
||||
cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is
|
||||
present, `SummaryBuildOptions::include_images` admits only bounded `media` image candidates and the digest includes both
|
||||
the flag and selected blob identity, MIME type, and size. The core is still synchronous; the server owns the blocking
|
||||
boundary. Anthropic HTTP, OpenAI-compatible HTTP, and Codex can transport the selected image data; Claude CLI receives
|
||||
text only.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
UI["ContextRail Summary"] -->|POST .../summary| Server
|
||||
|
|
@ -196,13 +203,14 @@ flowchart LR
|
|||
UI -->|GET .../summary poll| Row2
|
||||
```
|
||||
|
||||
**Tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads
|
||||
**X Articles and tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads
|
||||
the single `primary_media` artifact. A `tweet` or `tweet_thread` entry has no `primary_media` — it has
|
||||
N `raw_tweet_json` artifacts, one per status in the thread. So the summarizer selects on the
|
||||
`raw_tweet_json` role instead, loads **all** matching artifacts in order, and joins them with
|
||||
`\n\n---\n\n`; a `---` line reads as a hard paragraph break to every model, keeping individual
|
||||
statuses from bleeding into one another. Any change to how thread artifacts are stored has to be
|
||||
mirrored here.
|
||||
statuses from bleeding into one another. For an X Article, the reducer prefers article text over the tweet's body:
|
||||
`plain_text`, then flattened ordered blocks, then `preview_text`, then `summary_text`; only then does it fall back to
|
||||
`full_text`/`text`/`content`/`body`. Any change to article reduction or thread status storage must be mirrored here.
|
||||
|
||||
## yt-dlp Lifecycle
|
||||
|
||||
|
|
@ -272,7 +280,8 @@ The server both reads and writes archive data. Capture jobs are asynchronous: `P
|
|||
|
||||
**Auth model.** A separate `archivr-auth.sqlite` (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are `u32` flags (`GUEST`, `USER`, `ADMIN`, `OWNER`) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is `setup_guard` → `login_rate_limit` → `security_headers`; route families are classified `READ / ADMIN / WRITE / STATIC` in `routes.rs`.
|
||||
|
||||
**Search** is client-side filtering over entries the frontend has already fetched.
|
||||
**Search** is server-side free-text filtering over entry fields and the latest completed summary. The summary JSON is
|
||||
searched as text, so generated `tags` participate. Older completed summaries stay searchable while a newer request is
|
||||
pending or failed; rows with no completed summary contribute no summary-derived match.
|
||||
|
||||
**Admin view** covers mounted archives, users, sessions, and API tokens.
|
||||
|
||||
|
|
|
|||
|
|
@ -44,11 +44,11 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y
|
|||
- **Web pages** — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
|
||||
- **Local files** — import any file from disk by `file://` path
|
||||
- **Deduplication** — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
|
||||
- **Tags and search** — hierarchical tag tree, full-text search, filterable entry list
|
||||
- **Tags and search** — hierarchical tag tree, full-text search (including the latest completed summary and its generated JSON tags), filterable entry list
|
||||
- **Multiple archives** — the server mounts any number of separate archives from a single TOML config
|
||||
- **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
|
||||
- **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
|
||||
- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture
|
||||
- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option
|
||||
- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser
|
||||
- **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
|
||||
|
||||
|
|
@ -220,8 +220,22 @@ The Nix wrapper and Docker image set `ARCHIVR_STATIC_DIR`, `ARCHIVR_SINGLE_FILE`
|
|||
|
||||
#### LLM providers
|
||||
|
||||
Summaries are opt-in and provider-agnostic. Only the variables for the provider you actually select are read; the two
|
||||
HTTP providers refuse to start without their API key.
|
||||
Summaries are manual and provider-agnostic. Only the variables for the provider you actually select are read; the two
|
||||
HTTP providers refuse to start without their API key. They are text-only by default. Selecting `Include attached images`
|
||||
explicitly sends eligible archived image data to the chosen provider; it is never attached automatically.
|
||||
|
||||
| Provider | Attached images |
|
||||
|---|---|
|
||||
| Anthropic HTTP | Supported |
|
||||
| OpenAI-compatible HTTP | Supported |
|
||||
| Codex CLI | Supported |
|
||||
| Claude CLI | Not supported |
|
||||
|
||||
Image inclusion considers only `media` artifacts with `jpg`, `jpeg`, `png`, `webp`, `gif`, or `avif` files. At most four
|
||||
images are sent, each no larger than 5 MiB and no more than 12 MiB in total.
|
||||
|
||||
Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no
|
||||
summary, or only a pending or failed summary, get no summary-derived match.
|
||||
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
|
|
@ -404,4 +418,4 @@ nix build .#archivr-server
|
|||
## License
|
||||
|
||||
MIT — see [LICENSE](../LICENSE.md).
|
||||
\n
|
||||
\n
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue