From 44137a8eb100579ec6835d1d15a23ce873358352 Mon Sep 17 00:00:00 2001 From: archivr-qa Date: Sun, 23 Aug 2026 19:47:48 +0200 Subject: [PATCH] docs(maintainer): document summarizer, text capture, and yt-dlp lifecycle Co-Authored-By: Claude Opus 5 --- AGENTS.md | 38 +++++++++++++++++--- ARCHIVR-MENTAL-MODEL.md | 78 ++++++++++++++++++++++++++++++++++++++++- 2 files changed, 111 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 5233785..2ac8c2b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -18,6 +18,16 @@ Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the playlist probe (fetching per-video quality metadata before archiving), the multi-video download loop, and sync mode (skipping already-archived videos when re-archiving a playlist or channel). +Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection +and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the +blob lands in `raw/` like any other artifact. Entrypoint is +`POST /api/archives/:archive_id/captures/text`. + +LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind +`GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries` +table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the +provider set and the status lifecycle. + Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8). The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`. @@ -32,6 +42,7 @@ The server mounts multiple archives from a TOML registry (`crates/archivr-server | `frontend/src/` | React app: `App.jsx` (root state + custom routing), `api.js` (fetch client), `components/`, `styles.css` | | `docs/` | User docs (`README.md`), `superpowers/plans/` and `superpowers/specs/` (dated design docs — write plans there before large features) | | `modules/nixos/` | NixOS module (`services.archivr-server`) | +| `.github/workflows/` | `update-ytdlp.yml` — weekly cron that PRs a yt-dlp version bump into `flake.nix` | | `vendor/twitter/` | Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually. | | `vendor/readability/` | Mozilla `Readability.js`, concatenated into the SingleFile reader-mode browser script by `downloader/singlefile.rs`. | | `testing/` | Legacy scraping scripts + sample data. `testing/creds.txt` holds real tokens — never read, commit, or print it. | @@ -72,8 +83,13 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy - **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2. - **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal. - **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`. Downloaders shell out to subprocesses; resolve binaries through these vars. -- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. -- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. +- **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the + *pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a + self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison + entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. +- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. +- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. +- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract. - **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums. ## Important Files @@ -81,9 +97,16 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy - `crates/archivr-server/src/main.rs` — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup). - `crates/archivr-server/src/routes.rs` — all HTTP handlers and the router; grep here first for API work. - `crates/archivr-core/src/capture.rs` — `perform_capture()`, `Source` enum, shorthand parsing. -- `crates/archivr-core/src/downloader/ytdlp.rs` — yt-dlp integration; YouTube playlist/channel probe and download, sync mode logic. +- `crates/archivr-core/src/downloader/ytdlp.rs` — every yt-dlp shell-out (playlist/channel probe and download, sync mode) **plus** the binary resolver: `resolve_yt_dlp()`, `state_dir()`, `probe_version()`. +- `crates/archivr-core/src/downloader/text.rs` — pasted-text staging + hashing (`save()` → `StagedText`); accepts only `text/plain` and `text/markdown`. +- `crates/archivr-core/src/summarizer.rs` — the `SummaryProvider` trait and its four implementations (Anthropic HTTP, OpenAI-compatible HTTP, `claude` CLI, `codex` CLI), `PROMPT_VERSION`, `resolve_cli()`, prompt assembly, and `build_summary_input()` (artifact selection + HTML/text/JSON reduction). +- `crates/archivr-cli/src/main.rs` — CLI entry point, including the `archivr yt-dlp update|status` subcommand (staged, atomic zipapp install into the state dir). +- `.github/workflows/update-ytdlp.yml` — weekly (`0 6 * * 1`) + manual auto-bump of the `flake.nix` yt-dlp pin. - `crates/archivr-core/src/database.rs` — single source of truth for all SQLite schema and queries (both archive and auth DBs). - `frontend/src/App.jsx` / `frontend/src/api.js` — frontend root state and API surface. +- `frontend/src/components/CaptureDialog.jsx` — `CaptureRow` (locator input, playlist quality selectors) and `CaptureTextRow` (the "Add text" flow), plus job polling. +- `frontend/src/components/ContextRail.jsx` — the Summary section: provider selector, generate/regenerate, and summary polling. +- `frontend/src/components/TextPreview.jsx` — preview renderer for text/markdown entries. - `docker/config.example.toml` — server config schema: `bind`, `auth_db_path`, repeated `[[archives]]` (`id`, `label`, `archive_path`). - `flake.nix`, `modules/nixos/archivr-server.nix`, `Dockerfile`, `docker-compose.yml` — deployment surfaces; config schema changes must be reflected in all of them plus `docs/README.md`. @@ -94,9 +117,16 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy - Runtime binaries the app expects on PATH or via env vars: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset. - `.gitignore` is **default-deny with an allowlist** — new top-level files/dirs are invisible to git until explicitly allowed there. - Frontend build output (`crates/archivr-server/static/`) is generated; never hand-edit it. +- **yt-dlp is pinned to a specific GitHub release** in the `ytDlp` derivation in `flake.nix` (zipapp + fetched from `github.com/yt-dlp/yt-dlp/releases`, wrapped with `python312` + `ffmpeg`) — not taken + from nixpkgs. Both the `archivr` and `archivr-server` wrappers set `ARCHIVR_YT_DLP` from it. Three + ways to bump: the weekly `Update yt-dlp` workflow (automatic PR), `archivr yt-dlp update` + (per-machine, into the state dir), or editing the three fields of the `ytDlp` block by hand. Which + binary actually runs is decided at runtime by `resolve_yt_dlp()` in `downloader/ytdlp.rs`; use + `archivr yt-dlp status` to see the candidates and the winner. ## Testing & QA -- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p `. +- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`, and newer: `summarizer.rs` — 23 tests over provider construction, CLI/env resolution, output extraction, HTML/text/JSON input reduction and tweet-thread joining; `downloader/ytdlp.rs` — 13 tests, several covering resolver priority and version tie-breaking). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p `. - **Frontend**: no test framework. Storybook (`bun run storybook`) is the component QA surface — stories are colocated `*.stories.jsx` files; add one when adding a nontrivial component. - Manual smoke test for server changes: build frontend, `cargo run -p archivr-server -- `, exercise `/api/*`. diff --git a/ARCHIVR-MENTAL-MODEL.md b/ARCHIVR-MENTAL-MODEL.md index 7f8c3ec..d3ffd54 100644 --- a/ARCHIVR-MENTAL-MODEL.md +++ b/ARCHIVR-MENTAL-MODEL.md @@ -81,7 +81,7 @@ There are two user-facing binaries: | Binary | Purpose | |---|---| -| `archivr` | CLI for initializing archives and capturing material into one archive | +| `archivr` | CLI for initializing archives and capturing material into one archive (also `yt-dlp status\|update`) | | `archivr-server` | Web server for browsing one or more existing archives | The CLI writes archive data: @@ -128,6 +128,15 @@ sequenceDiagram CLI->>User: terminal result ``` +**Pasted text short-circuits most of that.** `perform_text_capture` (`capture.rs`) is not a `Source` +route: there is no locator to classify, no URL probe, and no downloader subprocess. It validates the +title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME type (`text/plain` or +`text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp//` +and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry +and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI +drives it from `CaptureTextRow` in `CaptureDialog.jsx`; entries produced this way render through +`TextPreview.jsx`. + ## Web Capture Pipeline Web pages (`Source::WebPage`) take a longer path than yt-dlp or tweets: @@ -158,6 +167,66 @@ sequenceDiagram Server-->>Browser: JSON ``` +## LLM Summaries + +Summaries are a **post-capture, manually triggered** subsystem. Nothing in `capture.rs` calls the +summarizer; a summary exists only because someone pressed generate in the Summary section of +`ContextRail.jsx`. + +`crates/archivr-core/src/summarizer.rs` defines one `SummaryProvider` trait with four implementations: +Anthropic HTTP, OpenAI-compatible HTTP, the local `claude` CLI, and the local `codex` CLI. Each is +built purely from environment variables (`provider_from_env`), so no key or model name is ever written +into archive data. `PROMPT_VERSION` in the same file stamps every row, so changing the prompt +invalidates the cache instead of silently mixing generations. + +`entry_summaries` (schema in `database.rs`) is that cache, unique on +(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the same entry +summarised by two providers, two models, or after a prompt change yields distinct rows, while a repeat +request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`, +mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`. + +```mermaid +flowchart LR + UI["ContextRail Summary"] -->|POST .../summary| Server + Server --> Input["build_summary_input()"] + Input --> Artifacts["entry artifacts on disk"] + Server --> Row["entry_summaries: pending → running"] + Server --> Provider["SummaryProvider (HTTP or CLI)"] + Provider --> Row2["completed / failed"] + UI -->|GET .../summary poll| Row2 +``` + +**Tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads +the single `primary_media` artifact. A `tweet` or `tweet_thread` entry has no `primary_media` — it has +N `raw_tweet_json` artifacts, one per status in the thread. So the summarizer selects on the +`raw_tweet_json` role instead, loads **all** matching artifacts in order, and joins them with +`\n\n---\n\n`; a `---` line reads as a hard paragraph break to every model, keeping individual +statuses from bleeding into one another. Any change to how thread artifacts are stored has to be +mirrored here. + +## yt-dlp Lifecycle + +There is no single yt-dlp. Up to three can exist on one machine: + +1. **The flake pin** — the `ytDlp` derivation in `flake.nix` fetches an exact release zipapp from + `github.com/yt-dlp/yt-dlp/releases` and wraps it with `python312` + `ffmpeg`. Both the `archivr` and + `archivr-server` wrappers export it as `ARCHIVR_YT_DLP`. +2. **A state-dir install** — `archivr yt-dlp update` downloads the latest zipapp and installs it + atomically (staged file, then rename) at `/yt-dlp/yt-dlp` with a sibling `.version` + sentinel that lets repeat runs skip the download. +3. **Whatever is on PATH** — the historical behaviour, and the last-resort fallback. + +`resolve_yt_dlp()` in `downloader/ytdlp.rs` picks between them once per process (cached in a +`OnceLock`): `ARCHIVR_YT_DLP_FORCE` wins outright if it points at a real file; otherwise the pinned and +state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`, +so plain string ordering is chronological — with exact ties going to the state-dir copy the user +deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status` +prints every candidate, its version, and the winner. + +Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the +current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp` +block and opens a PR), `archivr yt-dlp update` for one machine, or editing `flake.nix` by hand. + ## Where To Edit | Feature kind | Edit here | @@ -167,6 +236,10 @@ sequenceDiagram | Archive opening, listing entries, entry detail, runs | `crates/archivr-core/src/archive.rs` | | Download/save behavior | `crates/archivr-core/src/downloader/` | | YouTube playlist/channel download, playlist probe, sync mode | `crates/archivr-core/src/downloader/ytdlp.rs` and `capture.rs` | +| Which yt-dlp binary runs (resolver, state dir, version probe) | `crates/archivr-core/src/downloader/ytdlp.rs` | +| Pasted-text capture (staging, hashing, MIME allowlist) | `crates/archivr-core/src/downloader/text.rs` and `capture.rs` | +| LLM summary providers, prompt, `PROMPT_VERSION`, input building | `crates/archivr-core/src/summarizer.rs` | +| `entry_summaries` schema and summary CRUD | `crates/archivr-core/src/database.rs` | | CLI commands, argument parsing, terminal output | `crates/archivr-cli/src/main.rs` | | Server API routes | `crates/archivr-server/src/routes.rs` | | Auth model (users, sessions, tokens, roles) | `crates/archivr-server/src/auth.rs` | @@ -174,6 +247,9 @@ sequenceDiagram | Frontend root state + routing | `frontend/src/App.jsx` | | Frontend API client | `frontend/src/api.js` | | Frontend components | `frontend/src/components/` | +| Summary UI (provider selector, generate, polling) | `frontend/src/components/ContextRail.jsx` | +| Text/Markdown entry preview | `frontend/src/components/TextPreview.jsx` | +| "Add text" capture row | `frontend/src/components/CaptureDialog.jsx` | | Frontend styling | `frontend/src/styles.css` | ## Practical Feature Rule