1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00

Merge branch 'docs-sol-review' into integration-all-three

This commit is contained in:
archivr-qa 2026-08-24 16:36:50 +02:00
commit f5123ee64d
No known key found for this signature in database
3 changed files with 37 additions and 10 deletions

View file

@ -21,13 +21,18 @@ YouTube playlists and channels produce a **parent container entry** with each vi
Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection
and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the
blob lands in `raw/` like any other artifact. Entrypoint is blob lands in `raw/` like any other artifact. Entrypoint is
`POST /api/archives/:archive_id/captures/text`. `POST /api/archives/:archive_id/captures/text`. Its body is byte-preserving, its entry has no fabricated
`original_url`, and its normal text preview opens from the entry rail.
LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind
`GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries` `GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries`
table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the
provider set and the status lifecycle. provider set and the status lifecycle.
The requested provider model is the cache identity; a provider-returned resolved model is display attribution.
At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed
summary visible until its replacement completes; public readers receive completed content only, never diagnostics.
`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and `SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and
the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row. the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row.
Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each, Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each,
@ -96,9 +101,10 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
- **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the - **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the
*pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a *pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a
self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison
entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR`
relocates the state dir. Never spawn bare `yt-dlp` — call the resolver.
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`. - **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`.
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. - **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract. - **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract.
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums. - **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.

View file

@ -134,7 +134,8 @@ title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME t
`text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp/<timestamp>/` `text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp/<timestamp>/`
and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry
and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI
drives it from `CaptureTextRow` in `CaptureDialog.jsx`; entries produced this way render through drives it from `CaptureTextRow` in `CaptureDialog.jsx`. The body is preserved byte-for-byte, the entry's
`original_url` remains empty (no fabricated `text:` URL), and its normal entry-rail preview renders through
`TextPreview.jsx`. `TextPreview.jsx`.
## Web Capture Pipeline ## Web Capture Pipeline
@ -180,10 +181,17 @@ into archive data. `PROMPT_VERSION` in the same file stamps every row, so changi
invalidates the cache instead of silently mixing generations. invalidates the cache instead of silently mixing generations.
`entry_summaries` (schema in `database.rs`) is that cache, unique on `entry_summaries` (schema in `database.rs`) is that cache, unique on
(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the same entry (`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the requested provider
summarised by two providers, two models, or after a prompt change yields distinct rows, while a repeat model is the cache identity, so the same entry summarised by two providers, two requested models, or after a
request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`, prompt change yields distinct rows, while a repeat request with identical inputs reuses one. When a provider
mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`. returns its concrete resolved model, it is stored separately and displayed as attribution without changing that
identity. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked;
the frontend polls only for the currently selected entry, and its generate callbacks are scoped to that same
selection. Image-selection behavior remains unchanged.
On startup the server marks interrupted `pending` or `running` attempts failed. Regeneration is non-destructive:
the prior completed summary stays visible until a replacement completes successfully. Public readers receive only
completed summary content, never pending/failed state or diagnostic error text.
The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and
cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is
@ -229,7 +237,8 @@ There is no single yt-dlp. Up to three can exist on one machine:
state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`, state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`,
so plain string ordering is chronological — with exact ties going to the state-dir copy the user so plain string ordering is chronological — with exact ties going to the state-dir copy the user
deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status` deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status`
prints every candidate, its version, and the winner. prints every candidate, its version, and the winner; when the force variable applies, it includes that
forced candidate and selects it as the winner.
Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the
current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp` current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp`

View file

@ -49,7 +49,7 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y
- **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords - **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
- **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download - **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option - **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option
- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser - **Text notes** — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview
- **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block - **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
## Quick Start ## Quick Start
@ -177,6 +177,8 @@ Two body types are accepted: `text/markdown` (saved as `.md`) and `text/plain` (
rejected. The body lands in `store/raw/…` under its SHA3-256 content hash, exactly like every other capture, so an rejected. The body lands in `store/raw/…` under its SHA3-256 content hash, exactly like every other capture, so an
identical note captured twice is stored once. identical note captured twice is stored once.
Text notes have no synthetic source URL: the original-URL field stays empty rather than inventing a `text:` locator.
## Configuration ## Configuration
### TOML config file ### TOML config file
@ -237,6 +239,14 @@ images are sent, each no larger than 5 MiB and no more than 12 MiB in total.
Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no
summary, or only a pending or failed summary, get no summary-derived match. summary, or only a pending or failed summary, get no summary-derived match.
Each request is cached under the provider and the **requested** model identifier. If a provider reports a more precise
resolved model (for example, an alias's concrete version), the UI displays that resolved name as attribution without
changing the cache identity.
Summary attempts move from `pending` to `running` and then to `completed` or `failed`. On server startup, interrupted
pending or running attempts are marked failed. Regenerating does not replace an earlier completed summary until the
replacement succeeds, and public readers receive completed content only—never pending state or diagnostic errors.
| Variable | Default | Description | | Variable | Default | Description |
|---|---|---| |---|---|---|
| `ARCHIVR_ANTHROPIC_API_KEY` | *(required for `anthropic_http`)* | API key for the Anthropic Messages API | | `ARCHIVR_ANTHROPIC_API_KEY` | *(required for `anthropic_http`)* | API key for the Anthropic Messages API |
@ -285,6 +295,8 @@ archivr yt-dlp update # download the latest zipapp into t
archivr yt-dlp update --version 2026.09.15 # pin a specific release tag archivr yt-dlp update --version 2026.09.15 # pin a specific release tag
``` ```
When `ARCHIVR_YT_DLP_FORCE` applies, `status` shows that forced candidate and selects it as the winner.
The released artifact is a Python zipapp, so this path needs `python3` on `PATH` at run time. The released artifact is a Python zipapp, so this path needs `python3` on `PATH` at run time.
**2. Automatic weekly bump.** `.github/workflows/update-ytdlp.yml` runs every Monday at 06:00 UTC, queries GitHub for **2. Automatic weekly bump.** `.github/workflows/update-ytdlp.yml` runs every Monday at 06:00 UTC, queries GitHub for