1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
archivr/AGENTS.md
TheGeneralist 253f779216 Add YouTube subtitles, local transcription, self-updating yt-dlp/Deno, X Article and thread titles
- Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles)
- Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error
- Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only)
- Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s)
- Settings > Instance > yt-dlp: status and in-app update without restart
- X Article titles from article.title, with idempotent startup backfill
- Thread title generation (single and bulk) with per-provider cheap models
- Per-provider title model settings in Settings > Instance
- Docs, mental model, AGENTS.md and transcription spec updated
2026-10-05 19:39:51 +02:00

26 KiB
Raw Blame History

Repository Guidelines

Project Overview

archivr is a self-hosted archival tool that captures and preserves digital content — YouTube/Twitter/Instagram/TikTok/Reddit posts, arbitrary URLs, full web pages (via SingleFile + Chromium, optionally through a Freedium mirror for paywalled articles), and local files — into self-contained, SQLite-backed archive directories with blob deduplication, hierarchical tags, collections, and role-based auth. Rust workspace + React frontend.

Read ARCHIVR-MENTAL-MODEL.md before making structural changes.

Architecture & Data Flow

Three crates with a strict ownership split — core owns truth; CLI and server are adapters:

  • crates/archivr-core — domain library: capture orchestration, SQLite schema/CRUD, downloaders, hashing. New archive features start here.
  • crates/archivr-server — Axum HTTP API + auth + static frontend serving.
  • crates/archivr-cli — clap-based CLI (archivr binary): init, archive, yt-dlp status|update subcommands.

Capture flow: locator → determine_source() (crates/archivr-core/src/capture.rs) routes by platform/shorthand (yt:, x:, tweet: …) → platform downloader (downloader/ytdlp.rs, tweets.rs, singlefile.rs, http.rs, local.rs) stages into temp/ → SHA3-256 dedup (hash.rs, downloader/store.rs) moves blobs to raw/A/B/HASH.EXT → rows written to archivr.sqlite (runs, entries, artifacts, blobs) → served via /api/archives/:id/.... CaptureConfig carries per-request toggles (uBlock, reader mode, Freedium mirror, YouTube subtitles, etc.); when via_freedium is set, the fetch URL is rewritten through freedium-mirror.cfd while the canonical DB URL stays the original locator.

YouTube playlists and channels produce a parent container entry with each video captured as a child entry. downloader/ytdlp.rs handles the flat-playlist probe (fetching per-video quality metadata before archiving); the multi-video download loop and sync mode (skipping already-archived videos when re-archiving a playlist or channel) live in capture.rs. Child order is persisted in archived_entries.position (append-only at insert in database::create_archived_entry; rewritten only by database::reorder_child_entries behind PUT …/entries/:entry_uid/children/order (allowed roles = InstanceSettings::reorder_children_role_bits, default ADMIN|OWNER, Owner-editable)).

YouTube videos (Source::YouTubeVideo: single videos and playlist/channel children; no other platform) also get up to two subtitle tracks (English + original language, manual over auto, VTT/SRT) in the same yt-dlp call by default. CaptureConfig::download_subtitles (manual Default = true) gates it; the API body field download_subtitles (absent = true), the CaptureDialog "Download subtitles" toggle and CLI archivr archive --no-subtitles feed it. Files are archived into raw/ and registered as subtitle artifacts (crates/archivr-core/src/subtitles.rs). Subtitle failures are eprintln! warnings, never capture failures.

Pasted text takes a much shorter path: perform_text_capture() (capture.rs) skips source detection and every downloader shell-out — downloader/text.rs stages the body under temp/, hashes it, and the blob lands in raw/ like any other artifact. Entrypoint is POST /api/archives/:archive_id/captures/text. Its body is byte-preserving, its entry has no fabricated original_url, and its normal text preview opens from the entry rail.

LLM summaries are a post-capture, manual-only subsystem: crates/archivr-core/src/summarizer.rs behind GET/POST /api/archives/:archive_id/entries/:entry_uid/summary, cached in the entry_summaries table per (entry, provider, model, prompt version, input hash). See ARCHIVR-MENTAL-MODEL.md for the provider set and the status lifecycle.

YouTube video entries are summarized from the best-ranked subtitle artifact reduced to a transcript, not the mp4. With no usable track the POST still returns 202: the row is created pending with the placeholder input_sha256 pending-subtitle-fetch, and the background task fetches subtitles only (build_summary_input_with_subtitle_fetch → subtitles::fetch_subtitles_for_entry), writes the real hash, then runs the provider. The order is fixed (user requirement): archived subtitles → fetched subtitles → local transcription (transcriber::transcribe_entry, only when the POST body names a transcribe_engine and both earlier steps gave nothing) → error. If there are still none, the row fails with NO_SUBTITLES_SUMMARY_MESSAGE (or a transcription-specific copy when an engine ran) and no provider is called. Transcripts are subtitle artifacts with kind: "transcribed", origin: "transcription", engine, model. Spec + deviations: docs/superpowers/specs/2026-10-05-local-transcription-fallback.md.

X thread titles: POST /api/archives/:archive_id/entries/:entry_uid/thread-title (ROLE_USER, same gate as title PATCH; body {provider}) runs thread_title::generate_thread_title synchronously in one blocking task and saves Thread about <topic> — @author via database::update_entry_title. Never touches entry_summaries.

The requested provider model is the cache identity; a provider-returned resolved model is display attribution. At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed summary visible until its replacement completes; public readers receive completed content only, never diagnostics.

SummaryBuildOptions keeps summaries text-only unless include_images is set. The input digest includes that flag and the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row. Candidates are media artifacts only: jpg/jpeg, png, webp, gif, and avif, capped at four images, 5 MiB each, and 12 MiB in aggregate. Anthropic HTTP, OpenAI-compatible HTTP, and Codex support images; Claude CLI does not. The core remains synchronous: the server puts provider work in its blocking boundary rather than introducing async to archivr-core.

Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only. Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary.

Per-archive layout (created by archivr init): .archivr/ (name, store_path, archivr.sqlite) + sibling store/ (raw/, raw_tweets/, structured/, temp/). Server-level auth lives in a separate archivr-auth.sqlite (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8; per-action role masks (e.g. reorder_children_role_bits) and per-provider thread-title models (title_model_<kind>) live on its instance_settings row).

The server mounts multiple archives from a TOML registry (crates/archivr-server/src/registry.rs); routes are parameterized by :archive_id.

Key Directories

Path Purpose
crates/archivr-core/src/ Domain logic: capture.rs, archive.rs, database.rs (~2700 lines, all schema/CRUD), downloader/
crates/archivr-server/src/ main.rs (bootstrap), routes.rs (~3000 lines: AppState, handlers, middleware), auth.rs, registry.rs
crates/archivr-cli/src/ CLI entry point
frontend/src/ React app: App.jsx (root state + custom routing), api.js (fetch client), components/, styles.css
docs/ User docs (README.md), superpowers/plans/ and superpowers/specs/ (dated design docs — write plans there before large features)
modules/nixos/ NixOS module (services.archivr-server)
.github/workflows/ update-ytdlp.yml — weekly cron that PRs a yt-dlp version bump into flake.nix
vendor/twitter/ Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually.
vendor/readability/ Mozilla Readability.js, concatenated into the SingleFile reader-mode browser script by downloader/singlefile.rs.
testing/ Legacy scraping scripts + sample data. testing/creds.txt holds real tokens — never read, commit, or print it.

Development Commands

# Rust
cargo build                        # whole workspace
cargo test                         # all unit tests
cargo test -p archivr-core         # single crate
cargo build --release -p archivr-server

# Frontend (Bun, from frontend/)
bun install
bun run dev                        # Vite dev server
bun run build                      # → ../crates/archivr-server/static (gitignored; only needed for bare cargo run)

# Nix
nix develop                        # devshell: yt-dlp, deno, nushell, uv, twitter-api-client
nix build .#archivr-server         # also .#archivr-cli, .#archivr-all; builds frontend automatically
                                   # If frontendDeps hash is stale (after bun.lock/package.json change):
                                   #   nix build 2>&1 | grep "got:" → paste hash into flake.nix frontendDepsHash

# Docker
docker compose up -d               # port 8080; config in ./config/archivr-server.toml (see docker/config.example.toml)

# Run server locally
cargo run -p archivr-server -- path/to/archivr-server.toml

No CI is configured; no rustfmt.toml/clippy.toml — default cargo fmt/clippy settings apply.

Code Conventions & Common Patterns

  • Errors: anyhow::Result<T> everywhere in core; no custom error enums. The server converts to ApiError { status, message } (routes.rs) with helpers not_found/bad_request/unauthorized/forbidden; any anyhow error maps to 500.
  • Async: Tokio only at the server boundary (#[tokio::main], async handlers, tokio::spawn for the 24h session-cleanup task). Core is synchronous — blocking reqwest, subprocess downloaders, sync rusqlite. Keep it that way; don't introduce async into archivr-core.
  • State: Axum AppState { registry, auth_db_path, login_attempts } via State extractor; middleware stack = setup_guard (503 until owner exists) → login_rate_limit (5/15min per IP) → security_headers.
  • Auth extraction: AuthUser implements FromRequestParts — session cookie (session) or Authorization: Bearer token (stored SHA3-256-hashed). Passwords are Argon2.
  • Logging: eprintln! with info:/warn: prefixes. No tracing/log — don't add structured logging piecemeal.
  • External tools by env var: ARCHIVR_YT_DLP, ARCHIVR_DENO, ARCHIVR_JS_RUNTIME, ARCHIVR_CHROME, ARCHIVR_SINGLE_FILE, ARCHIVR_TWEET_PYTHON, ARCHIVR_TWEET_SCRAPER, ARCHIVR_STATIC_DIR, ARCHIVR_BIND, ARCHIVR_FFMPEG. Local transcription: ARCHIVR_TRANSCRIBE_ENGINES (the gate; unset = off), ARCHIVR_TRANSCRIBE_TIMEOUT (default 3600s, whole job), ARCHIVR_WHISPER_BACKEND (whisper_cpp|script) / ARCHIVR_WHISPER_CLI / ARCHIVR_WHISPER_MODEL / ARCHIVR_WHISPER_LANGUAGES, ARCHIVR_PARAKEET_CLI / ARCHIVR_PARAKEET_MODEL / ARCHIVR_PARAKEET_LANGUAGES, ARCHIVR_PHONON2_CLI / ARCHIVR_PHONON2_MODEL. Downloaders shell out to subprocesses; resolve binaries through these vars. Env helpers (required_env, env_or, optional_env, env_timeout, resolve_cli) live in env_config.rs; bounded subprocesses go through process::run_with_timeout (yt-dlp calls keep ytdlp.rs's private runner).
  • yt-dlp is resolved, not just read: ARCHIVR_YT_DLP (set by the flake wrappers) is only the pinned candidate handed to resolve_yt_dlp() (downloader/ytdlp.rs), which compares it against a self-updated copy in the state dir. ARCHIVR_YT_DLP_FORCE (absolute path) bypasses that comparison entirely; archivr yt-dlp status shows and chooses that forced candidate when it applies; ARCHIVR_STATE_DIR relocates the state dir. The JS runtime works the same way: resolve_js_runtime() (downloader/js_runtime.rs) picks the newest Deno ≥ 2.3.0 from ARCHIVR_DENO and <state_dir>/deno/deno (ties → state dir), then PATH; ARCHIVR_JS_RUNTIME=RUNTIME[:ABS_PATH] forces it. Both caches are RwLocks refreshed by refresh_yt_dlp() / refresh_js_runtime() after each successful component install (ytdlp_tools::update_tools), so a UI update needs no restart; resolve_js_runtime() returns an owned Option<JsRuntime>. Never cache a resolved path across operations. Never spawn bare yt-dlp — build every command with yt_dlp_command() so the resolver-chosen binary and --js-runtimes args are applied.
  • LLM summaries by env var: ARCHIVR_ANTHROPIC_API_KEY / ARCHIVR_ANTHROPIC_URL / ARCHIVR_ANTHROPIC_MODEL, ARCHIVR_OPENAI_API_KEY / ARCHIVR_OPENAI_URL / ARCHIVR_OPENAI_MODEL, ARCHIVR_CLAUDE_CLI / ARCHIVR_CLAUDE_MODEL, ARCHIVR_CODEX_CLI / ARCHIVR_CODEX_MODEL, plus ARCHIVR_SUMMARY_HTTP_TIMEOUT (default 120s) and ARCHIVR_SUMMARY_CLI_TIMEOUT (default 300s). Thread titles use ARCHIVR_ANTHROPIC_TITLE_MODEL (default claude-haiku-4-5), ARCHIVR_OPENAI_TITLE_MODEL (gpt-4o-mini), ARCHIVR_CLAUDE_TITLE_MODEL (haiku), ARCHIVR_CODEX_TITLE_MODEL (gpt-6-luna, override if unavailable) — never the summary model; thread_title.rs reuses provider transports via summarizer::complete_plain. The one exception to env-only config: admins may override the title model per provider in the auth DB (instance_settings.title_model_{anthropic_http,openai_compatible,claude_cli,codex_cli}; PATCH /api/admin/instance-settings trims, blank clears, rejects overlong or whitespace/control chars; GET returns title_models.<kind> with model/source/fallback). Precedence instance > ARCHIVR_*_TITLE_MODEL > default; the server passes the instance value to core as an Option<&str> override — core never reads the auth DB. Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in capture.rs triggers them. The two CLI vars are optional overrides: unset, resolve_cli() auto-discovers well-known absolute installs first (/opt/homebrew/bin/claude, /usr/local/bin/claude; /Applications/ChatGPT.app/Contents/Resources/codex, /opt/homebrew/bin/codex, /usr/local/bin/codex), then $HOME/.local/bin/<name>, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships codex off PATH. Frontend static output is generated; never hand-edit crates/archivr-server/static/.
  • Frontend: JSX (no TypeScript), PascalCase components in frontend/src/components/, kebab-case CSS classes, plain CSS with custom properties in styles.css (no Tailwind/CSS-in-JS). No router — App.jsx parses window.location.pathname + history.pushState. State = useState + one AuthContext; sessionStorage for refresh-resilient dialog state (see CaptureDialog.jsx job polling, 500ms). All API calls through frontend/src/api.js with relative /api/* URLs — add new endpoints there, not inline fetch. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through SkeletonEntryRow.jsx, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), not a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. .capture-text-row in styles.css), never from fallthrough on a generic row class; give a new row shape its own class.
  • CLI providers parse a file, not stdout: codex is invoked as codex exec --output-last-message <tempfile> - (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a tokens used footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject -. Keep any new CLI provider on the same "give me only the final message" contract.
  • Transcription engines follow the same rule: whisper.cpp writes -ovtt; script engines (Whisper script, Parakeet) are run as <script> --input <wav> --output <vtt> --model <m> [--language xx] and must write WebVTT to --output (optional <output>.lang); the only stdout parsed is Phonon-2's documented --json, converted to VTT by phonon_json_to_vtt. Phonon-2 is English-only, hard-coded. One job at a time (process-wide slot).
  • Log prefixes: info:/warn: only. The one exception is the yt-dlp installer's warning: python3 was not found on PATH … (ytdlp_tools.rs), kept verbatim for byte-identical CLI stderr.
  • Naming (Rust): standard snake_case/PascalCase; visibility and roles are bitflag u32s, not enums.

Important Files

  • crates/archivr-server/src/main.rs — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup), X Article title backfill (capture::backfill_x_article_titles; idempotent, startup only — CLI-only installs never run it).
  • crates/archivr-server/src/routes.rs — all HTTP handlers and the router; grep here first for API work.
  • crates/archivr-core/src/capture.rs — perform_capture(), Source enum, shorthand parsing; tweet titles prefer an X Article's article.title (<title> — @handle).
  • crates/archivr-core/src/downloader/ytdlp.rs — every yt-dlp shell-out (playlist/channel probe and download, sync mode, subtitle planning/args/staging, the combined media+subtitle call with its media-only retry, and the subtitles-only download_subtitles) plus the binary resolver: resolve_yt_dlp(), state_dir(), probe_version().
  • crates/archivr-core/src/subtitles.rs — subtitle artifacts: archive/register (deduped per entry+blob), subtitle_to_transcript() (VTT/SRT reduction, rolling-caption dedup), subtitle_track_rank(), and summary-time fetch_subtitles_for_entry().
  • crates/archivr-core/src/downloader/text.rs — pasted-text staging + hashing (save() → StagedText); accepts only text/plain and text/markdown.
  • crates/archivr-core/src/summarizer.rs — the SummaryProvider trait and its four implementations (Anthropic HTTP, OpenAI-compatible HTTP, claude CLI, codex CLI), PROMPT_VERSION, resolve_cli(), prompt assembly, build_summary_input() (artifact selection + HTML/text/JSON reduction; YouTube videos via subtitle transcript), build_summary_input_with_subtitle_fetch(), and the no-subtitles error/copy.
  • crates/archivr-core/src/transcriber.rs — local transcription: engine config from env (transcriber_from_env, available_transcribers), whisper.cpp/script/Phonon-2 adapters, audio acquisition + ffmpeg, the process-wide job slot, transcribe_entry, user-facing error copy.
  • crates/archivr-core/src/process.rs — run_with_timeout (kill on deadline, stderr tail) and the ProcessTimedOut sentinel; also backs summarizer run_cli.
  • crates/archivr-core/src/env_config.rs — shared ARCHIVR_* env helpers (required_env, env_or, optional_env, env_timeout, resolve_cli).
  • crates/archivr-core/src/thread_title.rs — thread-title prompt, topic sanitization, Thread about … — @author format, per-provider title model via resolve_title_model (precedence: instance setting instance_settings.title_model_<kind> passed in by the server > ARCHIVR_*_TITLE_MODEL > cheap default; core never reads the auth DB).
  • crates/archivr-core/src/downloader/ytdlp_tools.rs — yt-dlp + Deno update orchestration (update_tools, install_yt_dlp) and the status model (tools_status), shared by archivr yt-dlp and /api/admin/yt-dlp[/update].
  • crates/archivr-core/src/downloader/js_runtime.rs — JS runtime for yt-dlp: ARCHIVR_JS_RUNTIME parsing, DenoVersion, deno_candidates(), JsRuntimeRole + resolve_js_runtime_with_role() (uncached; reports the winning slot so status stars exactly one row), resolve_js_runtime() (cached RwLock, owned clone, warns per resolution), refresh_js_runtime() and js_runtime_args().
  • crates/archivr-cli/src/main.rs — CLI entry point; archivr yt-dlp update|status is a thin renderer over core ytdlp_tools.
  • crates/archivr-core/src/downloader/deno_install.rs — Deno release lookup, per-platform asset, staged + --version-verified atomic install into <state_dir>/deno/deno.
  • .github/workflows/update-ytdlp.yml — weekly (0 6 * * 1) + manual auto-bump of the flake.nix yt-dlp pin.
  • crates/archivr-core/src/database.rs — single source of truth for all SQLite schema and queries (both archive and auth DBs).
  • frontend/src/App.jsx / frontend/src/api.js — frontend root state and API surface.
  • frontend/src/components/CaptureDialog.jsx — CaptureRow (locator input, playlist quality selectors) and CaptureTextRow (the "Add text" flow), plus job polling.
  • frontend/src/components/ContextRail.jsx — the Summary section (provider selector, local-transcription engine selector, generate/regenerate, polling), the thread Generate title button, and the bulk panel's Generate titles (handleBulkGenerateTitles: ≥2 selected with ≥1 tweet_thread; skips non-threads; rail's Summary provider; per-entry thread-title calls 2 at a time; progress + X updated, Y failed; a selection change stops picking up new entries).
  • frontend/src/components/SettingsView.jsx — Settings, incl. the admin YtDlpSection (Instance › yt-dlp status + update).
  • frontend/src/components/TextPreview.jsx — preview renderer for text/markdown entries.
  • docker/config.example.toml — server config schema: bind, auth_db_path, repeated [[archives]] (id, label, archive_path).
  • flake.nix, modules/nixos/archivr-server.nix, Dockerfile, docker-compose.yml — deployment surfaces; config schema changes must be reflected in all of them plus docs/README.md.

Runtime/Tooling Preferences

  • Rust edition 2024 (root Cargo.toml); shared deps live in [workspace.dependencies] — add new deps there and reference with workspace = true.
  • Bun is the frontend package manager (frontend/bun.lock); use bun, not npm/yarn.
  • Runtime binaries the app expects on PATH or via env vars: yt-dlp, Deno (≥ 2.3.0, YouTube challenge solving), Chromium, single-file (Node), Python 3 with twitter-api-client, ffmpeg. nix develop provides the dev subset.
  • .gitignore is default-deny with an allowlist — new top-level files/dirs are invisible to git until explicitly allowed there.
  • Frontend build output (crates/archivr-server/static/) is generated by the frontendStatic Nix derivation; never hand-edit it, and do not commit it — it is excluded from git tracking. nix build is the standard workflow everywhere (local and NixOS) and builds the frontend automatically. bun run build only needed for bare cargo run one-off testing. When bun.lock or package.json changes, update frontendDepsHash in flake.nix for each system by running nix build 2>&1 | grep "got:" and pasting the reported hash.
  • yt-dlp is pinned to a specific GitHub release in the ytDlp derivation in flake.nix (zipapp fetched from github.com/yt-dlp/yt-dlp/releases, wrapped with python312 + ffmpeg) — not taken from nixpkgs. Both the archivr and archivr-server wrappers set ARCHIVR_YT_DLP from it. Three ways to bump: the weekly Update yt-dlp workflow (automatic PR), archivr yt-dlp update (per-machine, into the state dir), or editing the three fields of the ytDlp block by hand. Which binary actually runs is decided at runtime by resolve_yt_dlp() in downloader/ytdlp.rs; use archivr yt-dlp status to see the candidates and the winner.
  • Deno comes from nixpkgs pkgs.deno in both wrappers (ARCHIVR_DENO, also on their PATH) and the devShell. The Dockerfile pins Deno 2.9.7 with a sha256 per arch — no auto-bump; change version and both hashes together. Docker installs yt-dlp via pip as "yt-dlp[default]==<version>"; the [default] extra pulls yt-dlp-ejs (the challenge solver) — dropping it brings back YouTube 403s even with Deno. The weekly Update yt-dlp workflow only bumps flake.nix, never the Dockerfile pin. Docker sets ARCHIVR_STATE_DIR=/data/archivr-state (on the persistent /data volume) so in-container archivr yt-dlp update survives restarts. On NixOS without programs.nix-ld the upstream Deno can't execute: install_deno (archivr-core/src/downloader/deno_install.rs) detects the spawn NotFound, skips, and keeps the pinned ARCHIVR_DENO if it is usable (≥ 2.3.0) — reported as ok, not a failure (exit 0 unless yt-dlp itself failed); with no usable pin it errors.

Testing & QA

  • Rust: unit tests only, in #[cfg(test)] modules inside source files (e.g. capture.rs, database.rs, registry.rs, routes.rs, hash.rs, and newer: summarizer.rs — 53 tests over provider construction, CLI/env resolution, output extraction, HTML/text/JSON input reduction, tweet-thread joining, YouTube subtitle selection/digest/no-subtitles errors and the transcription fallback order; downloader/ytdlp.rs — tests covering resolver priority and refresh, version tie-breaking, yt_dlp_command_with args, subtitle planning, argument construction, staging, the media-only retry decision and the timeout runner; downloader/js_runtime.rs — runtime spec parsing, Deno version parsing, candidate priority/ties/minimum version and the winning role, PATH fallback via resolve_js_runtime_with_path (never mutate the process PATH), refresh and --js-runtimes args; downloader/deno_install.rs — 7 tests: asset selection, release parsing and verify_staged (version mismatch, the "cannot execute" case); downloader/ytdlp_tools.rs — 2; subtitles.rs — 13 tests over VTT/SRT reduction, rolling-caption dedup, track ranking and artifact dedup; transcriber.rs — 40 (env config, argument builders, Phonon JSON → VTT against a real sample, language gating, end-to-end with fake engines); process.rs — 5; env_config.rs — 2; thread_title.rs — 8). Test locks: tests that set resolver env vars take downloader::RESOLVER_ENV_LOCK and call refresh_*() on cleanup so the cache never points into a deleted tempdir; tests that set transcription env vars or run transcribe_entry (including summarizer.rs's fallback tests) take transcriber::TRANSCRIBE_TEST_LOCK; summarizer.rs and env_config.rs provider-env tests use their module-local ENV_LOCK. Tests that exec a freshly written script write it to a fresh path and wait out ETXTBSY first (fake_deno/deno_install::tests::write_script in core, write_script in the CLI). No tests/ integration dir. Patterns: tempfile for scratch archives, config round-trip assertions, regex/parser validation. Run cargo test or cargo test -p <crate>.
  • Frontend: component render tests are colocated *.test.jsx files on bun:test (bun test from frontend/); there is no Storybook.
  • Manual smoke test for server changes: build frontend, cargo run -p archivr-server -- <config.toml>, exercise /api/*.