diff --git a/AGENTS.md b/AGENTS.md index 098232c..9dea540 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -21,13 +21,18 @@ YouTube playlists and channels produce a **parent container entry** with each vi Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the blob lands in `raw/` like any other artifact. Entrypoint is -`POST /api/archives/:archive_id/captures/text`. +`POST /api/archives/:archive_id/captures/text`. Its body is byte-preserving, its entry has no fabricated +`original_url`, and its normal text preview opens from the entry rail. LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind `GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries` table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the provider set and the status lifecycle. +The requested provider model is the cache identity; a provider-returned resolved model is display attribution. +At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed +summary visible until its replacement completes; public readers receive completed content only, never diagnostics. + `SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row. Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each, @@ -96,9 +101,10 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy - **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the *pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison - entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. + entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR` + relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. - **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`. -- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. +- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. - **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract. - **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums. diff --git a/ARCHIVR-MENTAL-MODEL.md b/ARCHIVR-MENTAL-MODEL.md index 00ab7a1..3b4053e 100644 --- a/ARCHIVR-MENTAL-MODEL.md +++ b/ARCHIVR-MENTAL-MODEL.md @@ -134,7 +134,8 @@ title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME t `text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp//` and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI -drives it from `CaptureTextRow` in `CaptureDialog.jsx`; entries produced this way render through +drives it from `CaptureTextRow` in `CaptureDialog.jsx`. The body is preserved byte-for-byte, the entry's +`original_url` remains empty (no fabricated `text:` URL), and its normal entry-rail preview renders through `TextPreview.jsx`. ## Web Capture Pipeline @@ -180,10 +181,17 @@ into archive data. `PROMPT_VERSION` in the same file stamps every row, so changi invalidates the cache instead of silently mixing generations. `entry_summaries` (schema in `database.rs`) is that cache, unique on -(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the same entry -summarised by two providers, two models, or after a prompt change yields distinct rows, while a repeat -request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`, -mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`. +(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the requested provider +model is the cache identity, so the same entry summarised by two providers, two requested models, or after a +prompt change yields distinct rows, while a repeat request with identical inputs reuses one. When a provider +returns its concrete resolved model, it is stored separately and displayed as attribution without changing that +identity. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked; +the frontend polls only for the currently selected entry, and its generate callbacks are scoped to that same +selection. Image-selection behavior remains unchanged. + +On startup the server marks interrupted `pending` or `running` attempts failed. Regeneration is non-destructive: +the prior completed summary stays visible until a replacement completes successfully. Public readers receive only +completed summary content, never pending/failed state or diagnostic error text. The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is @@ -229,7 +237,8 @@ There is no single yt-dlp. Up to three can exist on one machine: state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`, so plain string ordering is chronological — with exact ties going to the state-dir copy the user deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status` -prints every candidate, its version, and the winner. +prints every candidate, its version, and the winner; when the force variable applies, it includes that +forced candidate and selects it as the winner. Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp` diff --git a/docs/README.md b/docs/README.md index 901776e..8e4d3c9 100644 --- a/docs/README.md +++ b/docs/README.md @@ -49,7 +49,7 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y - **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords - **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download - **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option -- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser +- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview - **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block ## Quick Start @@ -177,6 +177,8 @@ Two body types are accepted: `text/markdown` (saved as `.md`) and `text/plain` ( rejected. The body lands in `store/raw/…` under its SHA3-256 content hash, exactly like every other capture, so an identical note captured twice is stored once. +Text notes have no synthetic source URL: the original-URL field stays empty rather than inventing a `text:` locator. + ## Configuration ### TOML config file @@ -237,6 +239,14 @@ images are sent, each no larger than 5 MiB and no more than 12 MiB in total. Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no summary, or only a pending or failed summary, get no summary-derived match. +Each request is cached under the provider and the **requested** model identifier. If a provider reports a more precise +resolved model (for example, an alias's concrete version), the UI displays that resolved name as attribution without +changing the cache identity. + +Summary attempts move from `pending` to `running` and then to `completed` or `failed`. On server startup, interrupted +pending or running attempts are marked failed. Regenerating does not replace an earlier completed summary until the +replacement succeeds, and public readers receive completed content only—never pending state or diagnostic errors. + | Variable | Default | Description | |---|---|---| | `ARCHIVR_ANTHROPIC_API_KEY` | *(required for `anthropic_http`)* | API key for the Anthropic Messages API | @@ -285,6 +295,8 @@ archivr yt-dlp update # download the latest zipapp into t archivr yt-dlp update --version 2026.09.15 # pin a specific release tag ``` +When `ARCHIVR_YT_DLP_FORCE` applies, `status` shows that forced candidate and selects it as the winner. + The released artifact is a Python zipapp, so this path needs `python3` on `PATH` at run time. **2. Automatic weekly bump.** `.github/workflows/update-ytdlp.yml` runs every Monday at 06:00 UTC, queries GitHub for