mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 12:55:00 +02:00
docs: record summary lifecycle and review hardening
This commit is contained in:
parent
80f6b84af2
commit
f8626e07d9
3 changed files with 37 additions and 10 deletions
12
AGENTS.md
12
AGENTS.md
|
|
@ -21,13 +21,18 @@ YouTube playlists and channels produce a **parent container entry** with each vi
|
|||
Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection
|
||||
and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the
|
||||
blob lands in `raw/` like any other artifact. Entrypoint is
|
||||
`POST /api/archives/:archive_id/captures/text`.
|
||||
`POST /api/archives/:archive_id/captures/text`. Its body is byte-preserving, its entry has no fabricated
|
||||
`original_url`, and its normal text preview opens from the entry rail.
|
||||
|
||||
LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind
|
||||
`GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries`
|
||||
table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the
|
||||
provider set and the status lifecycle.
|
||||
|
||||
The requested provider model is the cache identity; a provider-returned resolved model is display attribution.
|
||||
At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed
|
||||
summary visible until its replacement completes; public readers receive completed content only, never diagnostics.
|
||||
|
||||
`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and
|
||||
the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row.
|
||||
Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each,
|
||||
|
|
@ -96,9 +101,10 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
|
|||
- **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the
|
||||
*pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a
|
||||
self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison
|
||||
entirely; `ARCHIVR_STATE_DIR` relocates the state dir. Never spawn bare `yt-dlp` — call the resolver.
|
||||
entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR`
|
||||
relocates the state dir. Never spawn bare `yt-dlp` — call the resolver.
|
||||
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`.
|
||||
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
|
||||
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
|
||||
- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract.
|
||||
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
|
||||
|
||||
|
|
|
|||
|
|
@ -134,7 +134,8 @@ title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME t
|
|||
`text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp/<timestamp>/`
|
||||
and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry
|
||||
and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI
|
||||
drives it from `CaptureTextRow` in `CaptureDialog.jsx`; entries produced this way render through
|
||||
drives it from `CaptureTextRow` in `CaptureDialog.jsx`. The body is preserved byte-for-byte, the entry's
|
||||
`original_url` remains empty (no fabricated `text:` URL), and its normal entry-rail preview renders through
|
||||
`TextPreview.jsx`.
|
||||
|
||||
## Web Capture Pipeline
|
||||
|
|
@ -180,10 +181,17 @@ into archive data. `PROMPT_VERSION` in the same file stamps every row, so changi
|
|||
invalidates the cache instead of silently mixing generations.
|
||||
|
||||
`entry_summaries` (schema in `database.rs`) is that cache, unique on
|
||||
(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the same entry
|
||||
summarised by two providers, two models, or after a prompt change yields distinct rows, while a repeat
|
||||
request with identical inputs reuses one. Rows move `pending` → `running` → `completed` | `failed`,
|
||||
mirroring how capture jobs are tracked, and the frontend polls until the row leaves `running`.
|
||||
(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the requested provider
|
||||
model is the cache identity, so the same entry summarised by two providers, two requested models, or after a
|
||||
prompt change yields distinct rows, while a repeat request with identical inputs reuses one. When a provider
|
||||
returns its concrete resolved model, it is stored separately and displayed as attribution without changing that
|
||||
identity. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked;
|
||||
the frontend polls only for the currently selected entry, and its generate callbacks are scoped to that same
|
||||
selection. Image-selection behavior remains unchanged.
|
||||
|
||||
On startup the server marks interrupted `pending` or `running` attempts failed. Regeneration is non-destructive:
|
||||
the prior completed summary stays visible until a replacement completes successfully. Public readers receive only
|
||||
completed summary content, never pending/failed state or diagnostic error text.
|
||||
|
||||
The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and
|
||||
cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is
|
||||
|
|
@ -229,7 +237,8 @@ There is no single yt-dlp. Up to three can exist on one machine:
|
|||
state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`,
|
||||
so plain string ordering is chronological — with exact ties going to the state-dir copy the user
|
||||
deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status`
|
||||
prints every candidate, its version, and the winner.
|
||||
prints every candidate, its version, and the winner; when the force variable applies, it includes that
|
||||
forced candidate and selects it as the winner.
|
||||
|
||||
Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the
|
||||
current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp`
|
||||
|
|
|
|||
|
|
@ -49,7 +49,7 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y
|
|||
- **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
|
||||
- **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
|
||||
- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option
|
||||
- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser
|
||||
- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview
|
||||
- **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
|
||||
|
||||
## Quick Start
|
||||
|
|
@ -177,6 +177,8 @@ Two body types are accepted: `text/markdown` (saved as `.md`) and `text/plain` (
|
|||
rejected. The body lands in `store/raw/…` under its SHA3-256 content hash, exactly like every other capture, so an
|
||||
identical note captured twice is stored once.
|
||||
|
||||
Text notes have no synthetic source URL: the original-URL field stays empty rather than inventing a `text:` locator.
|
||||
|
||||
## Configuration
|
||||
|
||||
### TOML config file
|
||||
|
|
@ -237,6 +239,14 @@ images are sent, each no larger than 5 MiB and no more than 12 MiB in total.
|
|||
Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no
|
||||
summary, or only a pending or failed summary, get no summary-derived match.
|
||||
|
||||
Each request is cached under the provider and the **requested** model identifier. If a provider reports a more precise
|
||||
resolved model (for example, an alias's concrete version), the UI displays that resolved name as attribution without
|
||||
changing the cache identity.
|
||||
|
||||
Summary attempts move from `pending` to `running` and then to `completed` or `failed`. On server startup, interrupted
|
||||
pending or running attempts are marked failed. Regenerating does not replace an earlier completed summary until the
|
||||
replacement succeeds, and public readers receive completed content only—never pending state or diagnostic errors.
|
||||
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
| `ARCHIVR_ANTHROPIC_API_KEY` | *(required for `anthropic_http`)* | API key for the Anthropic Messages API |
|
||||
|
|
@ -285,6 +295,8 @@ archivr yt-dlp update # download the latest zipapp into t
|
|||
archivr yt-dlp update --version 2026.09.15 # pin a specific release tag
|
||||
```
|
||||
|
||||
When `ARCHIVR_YT_DLP_FORCE` applies, `status` shows that forced candidate and selects it as the winner.
|
||||
|
||||
The released artifact is a Python zipapp, so this path needs `python3` on `PATH` at run time.
|
||||
|
||||
**2. Automatic weekly bump.** `.github/workflows/update-ytdlp.yml` runs every Monday at 06:00 UTC, queries GitHub for
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue