1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
archivr/AGENTS.md
TheGeneralist 253f779216 Add YouTube subtitles, local transcription, self-updating yt-dlp/Deno, X Article and thread titles
- Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles)
- Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error
- Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only)
- Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s)
- Settings > Instance > yt-dlp: status and in-app update without restart
- X Article titles from article.title, with idempotent startup backfill
- Thread title generation (single and bulk) with per-provider cheap models
- Per-provider title model settings in Settings > Instance
- Docs, mental model, AGENTS.md and transcription spec updated
2026-10-05 19:39:51 +02:00

196 lines
26 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Repository Guidelines
## Project Overview
archivr is a self-hosted archival tool that captures and preserves digital content — YouTube/Twitter/Instagram/TikTok/Reddit posts, arbitrary URLs, full web pages (via SingleFile + Chromium, optionally through a Freedium mirror for paywalled articles), and local files — into self-contained, SQLite-backed archive directories with blob deduplication, hierarchical tags, collections, and role-based auth. Rust workspace + React frontend.
Read `ARCHIVR-MENTAL-MODEL.md` before making structural changes.
## Architecture & Data Flow
Three crates with a strict ownership split — **core owns truth; CLI and server are adapters**:
- `crates/archivr-core` — domain library: capture orchestration, SQLite schema/CRUD, downloaders, hashing. New archive features start here.
- `crates/archivr-server` — Axum HTTP API + auth + static frontend serving.
- `crates/archivr-cli` — clap-based CLI (`archivr` binary): `init`, `archive`, `yt-dlp status|update` subcommands.
Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture.rs`) routes by platform/shorthand (`yt:`, `x:`, `tweet:` …) → platform downloader (`downloader/ytdlp.rs`, `tweets.rs`, `singlefile.rs`, `http.rs`, `local.rs`) stages into `temp/` → SHA3-256 dedup (`hash.rs`, `downloader/store.rs`) moves blobs to `raw/A/B/HASH.EXT` → rows written to `archivr.sqlite` (runs, entries, artifacts, blobs) → served via `/api/archives/:id/...`. `CaptureConfig` carries per-request toggles (uBlock, reader mode, Freedium mirror, YouTube subtitles, etc.); when `via_freedium` is set, the fetch URL is rewritten through `freedium-mirror.cfd` while the canonical DB URL stays the original locator.
YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the flat-playlist probe (fetching per-video quality metadata before archiving); the multi-video download loop and sync mode (skipping already-archived videos when re-archiving a playlist or channel) live in `capture.rs`. Child order is persisted in `archived_entries.position` (append-only at insert in `database::create_archived_entry`; rewritten only by `database::reorder_child_entries` behind `PUT …/entries/:entry_uid/children/order` (allowed roles = `InstanceSettings::reorder_children_role_bits`, default ADMIN|OWNER, Owner-editable)).
YouTube videos (`Source::YouTubeVideo`: single videos and playlist/channel children; no other platform) also get up
to two subtitle tracks (English + original language, manual over auto, VTT/SRT) in the same yt-dlp call by default.
`CaptureConfig::download_subtitles` (manual `Default` = true) gates it; the API body field `download_subtitles` (absent
= true), the CaptureDialog "Download subtitles" toggle and CLI `archivr archive --no-subtitles` feed it. Files are
archived into `raw/` and registered as `subtitle` artifacts (`crates/archivr-core/src/subtitles.rs`). Subtitle
failures are `eprintln!` warnings, never capture failures.
Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection
and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the
blob lands in `raw/` like any other artifact. Entrypoint is
`POST /api/archives/:archive_id/captures/text`. Its body is byte-preserving, its entry has no fabricated
`original_url`, and its normal text preview opens from the entry rail.
LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind
`GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries`
table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the
provider set and the status lifecycle.
YouTube video entries are summarized from the best-ranked `subtitle` artifact reduced to a transcript, not the mp4.
With no usable track the POST still returns 202: the row is created `pending` with the placeholder `input_sha256`
`pending-subtitle-fetch`, and the background task fetches subtitles only (`build_summary_input_with_subtitle_fetch` →
`subtitles::fetch_subtitles_for_entry`), writes the real hash, then runs the provider. The order is fixed (user
requirement): archived subtitles → fetched subtitles → local transcription (`transcriber::transcribe_entry`, only when
the POST body names a `transcribe_engine` and both earlier steps gave nothing) → error. If there are still none, the
row fails with `NO_SUBTITLES_SUMMARY_MESSAGE` (or a transcription-specific copy when an engine ran) and no provider is
called. Transcripts are `subtitle` artifacts with `kind: "transcribed"`, `origin: "transcription"`, `engine`, `model`.
Spec + deviations: `docs/superpowers/specs/2026-10-05-local-transcription-fallback.md`.
X thread titles: `POST /api/archives/:archive_id/entries/:entry_uid/thread-title` (`ROLE_USER`, same gate as title
PATCH; body `{provider}`) runs `thread_title::generate_thread_title` synchronously in one blocking task and saves
`Thread about <topic> — @author` via `database::update_entry_title`. Never touches `entry_summaries`.
The requested provider model is the cache identity; a provider-returned resolved model is display attribution.
At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed
summary visible until its replacement completes; public readers receive completed content only, never diagnostics.
`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and
the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row.
Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each,
and 12 MiB in aggregate. Anthropic HTTP, OpenAI-compatible HTTP, and Codex support images; Claude CLI does not. The core
remains synchronous: the server puts provider work in its blocking boundary rather than introducing async to
`archivr-core`.
Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only.
Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary.
Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8; per-action role masks (e.g. `reorder_children_role_bits`) and per-provider thread-title models (`title_model_<kind>`) live on its `instance_settings` row).
The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`.
## Key Directories
| Path | Purpose |
|---|---|
| `crates/archivr-core/src/` | Domain logic: `capture.rs`, `archive.rs`, `database.rs` (~2700 lines, all schema/CRUD), `downloader/` |
| `crates/archivr-server/src/` | `main.rs` (bootstrap), `routes.rs` (~3000 lines: AppState, handlers, middleware), `auth.rs`, `registry.rs` |
| `crates/archivr-cli/src/` | CLI entry point |
| `frontend/src/` | React app: `App.jsx` (root state + custom routing), `api.js` (fetch client), `components/`, `styles.css` |
| `docs/` | User docs (`README.md`), `superpowers/plans/` and `superpowers/specs/` (dated design docs — write plans there before large features) |
| `modules/nixos/` | NixOS module (`services.archivr-server`) |
| `.github/workflows/` | `update-ytdlp.yml` — weekly cron that PRs a yt-dlp version bump into `flake.nix` |
| `vendor/twitter/` | Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually. |
| `vendor/readability/` | Mozilla `Readability.js`, concatenated into the SingleFile reader-mode browser script by `downloader/singlefile.rs`. |
| `testing/` | Legacy scraping scripts + sample data. `testing/creds.txt` holds real tokens — never read, commit, or print it. |
## Development Commands
```bash
# Rust
cargo build # whole workspace
cargo test # all unit tests
cargo test -p archivr-core # single crate
cargo build --release -p archivr-server
# Frontend (Bun, from frontend/)
bun install
bun run dev # Vite dev server
bun run build # → ../crates/archivr-server/static (gitignored; only needed for bare cargo run)
# Nix
nix develop # devshell: yt-dlp, deno, nushell, uv, twitter-api-client
nix build .#archivr-server # also .#archivr-cli, .#archivr-all; builds frontend automatically
# If frontendDeps hash is stale (after bun.lock/package.json change):
# nix build 2>&1 | grep "got:" → paste hash into flake.nix frontendDepsHash
# Docker
docker compose up -d # port 8080; config in ./config/archivr-server.toml (see docker/config.example.toml)
# Run server locally
cargo run -p archivr-server -- path/to/archivr-server.toml
```
No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy` settings apply.
## Code Conventions & Common Patterns
- **Errors**: `anyhow::Result<T>` everywhere in core; no custom error enums. The server converts to `ApiError { status, message }` (`routes.rs`) with helpers `not_found`/`bad_request`/`unauthorized`/`forbidden`; any `anyhow` error maps to 500.
- **Async**: Tokio only at the server boundary (`#[tokio::main]`, async handlers, `tokio::spawn` for the 24h session-cleanup task). **Core is synchronous** — blocking `reqwest`, subprocess downloaders, sync `rusqlite`. Keep it that way; don't introduce async into archivr-core.
- **State**: Axum `AppState { registry, auth_db_path, login_attempts }` via `State` extractor; middleware stack = `setup_guard` (503 until owner exists) → `login_rate_limit` (5/15min per IP) → `security_headers`.
- **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2.
- **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal.
- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_DENO`, `ARCHIVR_JS_RUNTIME`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`, `ARCHIVR_FFMPEG`. Local transcription: `ARCHIVR_TRANSCRIBE_ENGINES` (the gate; unset = off), `ARCHIVR_TRANSCRIBE_TIMEOUT` (default 3600s, whole job), `ARCHIVR_WHISPER_BACKEND` (`whisper_cpp`|`script`) / `ARCHIVR_WHISPER_CLI` / `ARCHIVR_WHISPER_MODEL` / `ARCHIVR_WHISPER_LANGUAGES`, `ARCHIVR_PARAKEET_CLI` / `ARCHIVR_PARAKEET_MODEL` / `ARCHIVR_PARAKEET_LANGUAGES`, `ARCHIVR_PHONON2_CLI` / `ARCHIVR_PHONON2_MODEL`. Downloaders shell out to subprocesses; resolve binaries through these vars. Env helpers (`required_env`, `env_or`, `optional_env`, `env_timeout`, `resolve_cli`) live in `env_config.rs`; bounded subprocesses go through `process::run_with_timeout` (yt-dlp calls keep `ytdlp.rs`'s private runner).
- **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the
*pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a
self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison
entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR`
relocates the state dir. The JS runtime works the same way: `resolve_js_runtime()` (`downloader/js_runtime.rs`)
picks the newest Deno ≥ 2.3.0 from `ARCHIVR_DENO` and `<state_dir>/deno/deno` (ties → state dir), then PATH;
`ARCHIVR_JS_RUNTIME=RUNTIME[:ABS_PATH]` forces it. Both caches are `RwLock`s refreshed by `refresh_yt_dlp()` /
`refresh_js_runtime()` after each successful component install (`ytdlp_tools::update_tools`), so a UI update needs
no restart; `resolve_js_runtime()` returns an owned `Option<JsRuntime>`. Never cache a resolved path across
operations. Never spawn bare yt-dlp — build every command with
`yt_dlp_command()` so the resolver-chosen binary *and* `--js-runtimes` args are applied.
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Thread titles use `ARCHIVR_ANTHROPIC_TITLE_MODEL` (default `claude-haiku-4-5`), `ARCHIVR_OPENAI_TITLE_MODEL` (`gpt-4o-mini`), `ARCHIVR_CLAUDE_TITLE_MODEL` (`haiku`), `ARCHIVR_CODEX_TITLE_MODEL` (`gpt-6-luna`, override if unavailable) — never the summary model; `thread_title.rs` reuses provider transports via `summarizer::complete_plain`. The one exception to env-only config: admins may override the title model per provider in the auth DB (`instance_settings.title_model_{anthropic_http,openai_compatible,claude_cli,codex_cli}`; PATCH `/api/admin/instance-settings` trims, blank clears, rejects overlong or whitespace/control chars; GET returns `title_models.<kind>` with `model`/`source`/`fallback`). Precedence instance > `ARCHIVR_*_TITLE_MODEL` > default; the server passes the instance value to core as an `Option<&str>` override — core never reads the auth DB. Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`.
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract.
- **Transcription engines follow the same rule**: whisper.cpp writes `-ovtt`; script engines (Whisper `script`, Parakeet) are run as `<script> --input <wav> --output <vtt> --model <m> [--language xx]` and must write WebVTT to `--output` (optional `<output>.lang`); the only stdout parsed is Phonon-2's documented `--json`, converted to VTT by `phonon_json_to_vtt`. Phonon-2 is English-only, hard-coded. One job at a time (process-wide slot).
- **Log prefixes**: `info:`/`warn:` only. The one exception is the yt-dlp installer's `warning: python3 was not found on PATH …` (`ytdlp_tools.rs`), kept verbatim for byte-identical CLI stderr.
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
## Important Files
- `crates/archivr-server/src/main.rs` — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup), X Article title backfill (`capture::backfill_x_article_titles`; idempotent, startup only — CLI-only installs never run it).
- `crates/archivr-server/src/routes.rs` — all HTTP handlers and the router; grep here first for API work.
- `crates/archivr-core/src/capture.rs` — `perform_capture()`, `Source` enum, shorthand parsing; tweet titles prefer an X Article's `article.title` (`<title> — @handle`).
- `crates/archivr-core/src/downloader/ytdlp.rs` — every yt-dlp shell-out (playlist/channel probe and download, sync mode, subtitle planning/args/staging, the combined media+subtitle call with its media-only retry, and the subtitles-only `download_subtitles`) **plus** the binary resolver: `resolve_yt_dlp()`, `state_dir()`, `probe_version()`.
- `crates/archivr-core/src/subtitles.rs` — `subtitle` artifacts: archive/register (deduped per entry+blob), `subtitle_to_transcript()` (VTT/SRT reduction, rolling-caption dedup), `subtitle_track_rank()`, and summary-time `fetch_subtitles_for_entry()`.
- `crates/archivr-core/src/downloader/text.rs` — pasted-text staging + hashing (`save()` → `StagedText`); accepts only `text/plain` and `text/markdown`.
- `crates/archivr-core/src/summarizer.rs` — the `SummaryProvider` trait and its four implementations (Anthropic HTTP, OpenAI-compatible HTTP, `claude` CLI, `codex` CLI), `PROMPT_VERSION`, `resolve_cli()`, prompt assembly, `build_summary_input()` (artifact selection + HTML/text/JSON reduction; YouTube videos via subtitle transcript), `build_summary_input_with_subtitle_fetch()`, and the no-subtitles error/copy.
- `crates/archivr-core/src/transcriber.rs` — local transcription: engine config from env (`transcriber_from_env`, `available_transcribers`), whisper.cpp/script/Phonon-2 adapters, audio acquisition + ffmpeg, the process-wide job slot, `transcribe_entry`, user-facing error copy.
- `crates/archivr-core/src/process.rs` — `run_with_timeout` (kill on deadline, stderr tail) and the `ProcessTimedOut` sentinel; also backs summarizer `run_cli`.
- `crates/archivr-core/src/env_config.rs` — shared `ARCHIVR_*` env helpers (`required_env`, `env_or`, `optional_env`, `env_timeout`, `resolve_cli`).
- `crates/archivr-core/src/thread_title.rs` — thread-title prompt, topic sanitization, `Thread about … — @author` format, per-provider title model via `resolve_title_model` (precedence: instance setting `instance_settings.title_model_<kind>` passed in by the server > `ARCHIVR_*_TITLE_MODEL` > cheap default; core never reads the auth DB).
- `crates/archivr-core/src/downloader/ytdlp_tools.rs` — yt-dlp + Deno update orchestration (`update_tools`, `install_yt_dlp`) and the status model (`tools_status`), shared by `archivr yt-dlp` and `/api/admin/yt-dlp[/update]`.
- `crates/archivr-core/src/downloader/js_runtime.rs` — JS runtime for yt-dlp: `ARCHIVR_JS_RUNTIME` parsing, `DenoVersion`, `deno_candidates()`, `JsRuntimeRole` + `resolve_js_runtime_with_role()` (uncached; reports the winning slot so `status` stars exactly one row), `resolve_js_runtime()` (cached `RwLock`, owned clone, warns per resolution), `refresh_js_runtime()` and `js_runtime_args()`.
- `crates/archivr-cli/src/main.rs` — CLI entry point; `archivr yt-dlp update|status` is a thin renderer over core `ytdlp_tools`.
- `crates/archivr-core/src/downloader/deno_install.rs` — Deno release lookup, per-platform asset, staged + `--version`-verified atomic install into `<state_dir>/deno/deno`.
- `.github/workflows/update-ytdlp.yml` — weekly (`0 6 * * 1`) + manual auto-bump of the `flake.nix` yt-dlp pin.
- `crates/archivr-core/src/database.rs` — single source of truth for all SQLite schema and queries (both archive and auth DBs).
- `frontend/src/App.jsx` / `frontend/src/api.js` — frontend root state and API surface.
- `frontend/src/components/CaptureDialog.jsx` — `CaptureRow` (locator input, playlist quality selectors) and `CaptureTextRow` (the "Add text" flow), plus job polling.
- `frontend/src/components/ContextRail.jsx` — the Summary section (provider selector, local-transcription engine selector, generate/regenerate, polling), the thread **Generate title** button, and the bulk panel's **Generate titles** (`handleBulkGenerateTitles`: ≥2 selected with ≥1 `tweet_thread`; skips non-threads; rail's Summary provider; per-entry `thread-title` calls 2 at a time; progress + `X updated, Y failed`; a selection change stops picking up new entries).
- `frontend/src/components/SettingsView.jsx` — Settings, incl. the admin `YtDlpSection` (Instance › yt-dlp status + update).
- `frontend/src/components/TextPreview.jsx` — preview renderer for text/markdown entries.
- `docker/config.example.toml` — server config schema: `bind`, `auth_db_path`, repeated `[[archives]]` (`id`, `label`, `archive_path`).
- `flake.nix`, `modules/nixos/archivr-server.nix`, `Dockerfile`, `docker-compose.yml` — deployment surfaces; config schema changes must be reflected in all of them plus `docs/README.md`.
## Runtime/Tooling Preferences
- **Rust edition 2024** (root `Cargo.toml`); shared deps live in `[workspace.dependencies]` — add new deps there and reference with `workspace = true`.
- **Bun** is the frontend package manager (`frontend/bun.lock`); use `bun`, not npm/yarn.
- Runtime binaries the app expects on PATH or via env vars: `yt-dlp`, Deno (≥ 2.3.0, YouTube challenge solving), Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
- `.gitignore` is **default-deny with an allowlist** — new top-level files/dirs are invisible to git until explicitly allowed there.
- Frontend build output (`crates/archivr-server/static/`) is generated by the `frontendStatic` Nix derivation; never hand-edit it, and do not commit it — it is excluded from git tracking. `nix build` is the standard workflow everywhere (local and NixOS) and builds the frontend automatically. `bun run build` only needed for bare `cargo run` one-off testing. When `bun.lock` or `package.json` changes, update `frontendDepsHash` in `flake.nix` for each system by running `nix build 2>&1 | grep "got:"` and pasting the reported hash.
- **yt-dlp is pinned to a specific GitHub release** in the `ytDlp` derivation in `flake.nix` (zipapp
fetched from `github.com/yt-dlp/yt-dlp/releases`, wrapped with `python312` + `ffmpeg`) — not taken
from nixpkgs. Both the `archivr` and `archivr-server` wrappers set `ARCHIVR_YT_DLP` from it. Three
ways to bump: the weekly `Update yt-dlp` workflow (automatic PR), `archivr yt-dlp update`
(per-machine, into the state dir), or editing the three fields of the `ytDlp` block by hand. Which
binary actually runs is decided at runtime by `resolve_yt_dlp()` in `downloader/ytdlp.rs`; use
`archivr yt-dlp status` to see the candidates and the winner.
- **Deno** comes from nixpkgs `pkgs.deno` in both wrappers (`ARCHIVR_DENO`, also on their PATH) and the devShell.
The `Dockerfile` pins Deno 2.9.7 with a sha256 per arch — no auto-bump; change version and both hashes together.
Docker installs yt-dlp via pip as `"yt-dlp[default]==<version>"`; the `[default]` extra pulls `yt-dlp-ejs`
(the challenge solver) — dropping it brings back YouTube 403s even with Deno. The weekly `Update yt-dlp`
workflow only bumps `flake.nix`, never the Dockerfile pin. Docker sets `ARCHIVR_STATE_DIR=/data/archivr-state`
(on the persistent `/data` volume) so in-container `archivr yt-dlp update` survives restarts. On NixOS
without `programs.nix-ld` the upstream Deno can't execute: `install_deno` (`archivr-core/src/downloader/deno_install.rs`)
detects the spawn `NotFound`, skips, and keeps the pinned `ARCHIVR_DENO` if it is usable (≥ 2.3.0) —
reported as ok, not a failure (exit 0 unless yt-dlp itself failed); with no usable pin it errors.
## Testing & QA
- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`, and newer: `summarizer.rs` — 53 tests over provider construction, CLI/env resolution, output extraction, HTML/text/JSON input reduction, tweet-thread joining, YouTube subtitle selection/digest/no-subtitles errors and the transcription fallback order; `downloader/ytdlp.rs` — tests covering resolver priority and refresh, version tie-breaking, `yt_dlp_command_with` args, subtitle planning, argument construction, staging, the media-only retry decision and the timeout runner; `downloader/js_runtime.rs` — runtime spec parsing, Deno version parsing, candidate priority/ties/minimum version and the winning role, PATH fallback via `resolve_js_runtime_with_path` (never mutate the process `PATH`), refresh and `--js-runtimes` args; `downloader/deno_install.rs` — 7 tests: asset selection, release parsing and `verify_staged` (version mismatch, the "cannot execute" case); `downloader/ytdlp_tools.rs` — 2; `subtitles.rs` — 13 tests over VTT/SRT reduction, rolling-caption dedup, track ranking and artifact dedup; `transcriber.rs` — 40 (env config, argument builders, Phonon JSON → VTT against a real sample, language gating, end-to-end with fake engines); `process.rs` — 5; `env_config.rs` — 2; `thread_title.rs` — 8). Test locks: tests that set resolver env vars take `downloader::RESOLVER_ENV_LOCK` and call `refresh_*()` on cleanup so the cache never points into a deleted tempdir; tests that set transcription env vars or run `transcribe_entry` (including `summarizer.rs`'s fallback tests) take `transcriber::TRANSCRIBE_TEST_LOCK`; `summarizer.rs` and `env_config.rs` provider-env tests use their module-local `ENV_LOCK`. Tests that exec a freshly written script write it to a fresh path and wait out ETXTBSY first (`fake_deno`/`deno_install::tests::write_script` in core, `write_script` in the CLI). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p <crate>`.
- **Frontend**: component render tests are colocated `*.test.jsx` files on `bun:test` (`bun test` from `frontend/`); there is no Storybook.
- Manual smoke test for server changes: build frontend, `cargo run -p archivr-server -- <config.toml>`, exercise `/api/*`.