1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 21:03:17 +02:00
archivr/ARCHIVR-MENTAL-MODEL.md
TheGeneralist 253f779216 Add YouTube subtitles, local transcription, self-updating yt-dlp/Deno, X Article and thread titles
- Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles)
- Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error
- Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only)
- Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s)
- Settings > Instance > yt-dlp: status and in-app update without restart
- X Article titles from article.title, with idempotent startup backfill
- Thread title generation (single and bulk) with per-provider cheap models
- Per-provider title model settings in Settings > Instance
- Docs, mental model, AGENTS.md and transcription spec updated
2026-10-05 19:39:51 +02:00

412 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Archivr Mental Model
This document explains the current project shape after the workspace refactor.
## Key Documents
| Document | Role |
|---|---|
| `ARCHIVR-MENTAL-MODEL.md` | **This file.** Current architecture, data flows, and where to edit. |
| `docs/README.md` | User-facing docs: how to run the tool, supported inputs, environment variables. |
## The Big Model
archivr is now a Rust workspace with three crates:
```mermaid
flowchart LR
CLI["archivr-cli"] --> Core["archivr-core"]
Server["archivr-server"] --> Core
UI["static web UI"] --> Server
ServerConfig["server TOML registry"] --> Server
Core --> DB["archive/.archivr/archivr.sqlite"]
Core --> Store["archive store: raw, raw_tweets, structured, temp"]
```
The key rule:
> `archivr-core` owns archive behavior. `archivr-cli` and `archivr-server` are adapters.
## Crates
| Crate | Responsibility |
|---|---|
| `archivr-core` | Archive/domain logic, database schema, queries, download/store helpers |
| `archivr-cli` | Command-line interface, argument parsing, terminal behavior |
| `archivr-server` | Web server, API routes, mounted archive registry, static UI |
## Archive Model
Each archive is still self-contained:
```text
some-archive/
.archivr/
archivr.sqlite
name
store_path
store/
raw/
raw_tweets/
structured/
temp/
```
The web server can mount many independent archives through its own TOML registry.
That registry is separate from the archives themselves.
Example:
```toml
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/path/to/archive/.archivr"
```
## Entry Nesting
Entries support a two-level parent/child hierarchy. A **container entry** (playlist or channel) holds zero or more child entries (individual videos). Container entries have no primary media artifact of their own; their `total_artifact_bytes` is the sum of their children's bytes.
Rules:
- Maximum nesting depth is 2 (root → child). Children cannot have children.
- The UI shows child entries collapsed under their parent, expandable with a chevron.
- Container entries are created by playlist/channel captures. Single-video and all other source types produce a standalone root entry with no children.
- Children have a persisted sibling order: `archived_entries.position` (0-based per parent, `NULL` for roots). `list_child_entries` orders by it (tie-break `archived_at, id`).
- New children are always appended (`MAX(position)+1` in `create_archived_entry`): an initial playlist/channel capture keeps playlist enumeration order; sync appends newly found videos after the existing (possibly user-reordered) children.
- Users whose roles are allowed by the instance setting `reorder_children_role_bits` (auth DB `instance_settings`; default ADMIN|OWNER = 12; editable only by the Owner in Settings → Instance → Permissions) reorder an expanded parent's children on the main page. Exactly one control is shown, selected purely in CSS: the ↑/↓ buttons are the default, and a single `(hover: hover) and (pointer: fine) and (min-width: 641px)` query swaps in the drag handle (HTML5 drag-and-drop). Touch, no-pointer, phone-width (≤ 640px) viewports, and browsers that can't evaluate the query keep the arrows, since HTML5 DnD is unreliable there and no feature test detects it. Alt+↑/↓ on a focused child row works on all inputs (advertised via `aria-keyshortcuts`). The UI sends the full child UID list to `PUT /api/archives/:archive_id/entries/:entry_uid/children/order` (401 guest, 403 role not in mask, 404 unknown parent or a parent the caller can't see under the `list_child_entries` rule (`database::caller_sees_all_children`), 400 unless the list is exactly the current children). `/api/auth/me` and login return `can_reorder_children`; the UI hides reorder controls when it is false.
If a feature touches how entries are parented or how the UI groups them, start in `archivr-core` (`database.rs` for schema, `archive.rs` for listing, `capture.rs` for creation). Child-order UI lives in `routes.rs` (`reorder_entry_children_handler`) and `frontend/src/components/EntryRow.jsx`.
## How To Run It
There are two user-facing binaries:
| Binary | Purpose |
|---|---|
| `archivr` | CLI for initializing archives and capturing material into one archive (also `yt-dlp status\|update`; the web UI equivalent is Settings › Instance › yt-dlp) |
| `archivr-server` | Web server for browsing one or more existing archives |
The CLI writes archive data:
```sh
nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
```
The server reads archive data:
```sh
nix run .#archivr-server -- ./archivr-server.toml
```
If no config path is passed, the server reads `./archivr-server.toml`.
The config is a server registry, not archive data:
```toml
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
```
The packaged Nix server wrapper sets `ARCHIVR_STATIC_DIR` so the server can find the installed web UI assets. Source-tree runs do not need that variable because they fall back to `crates/archivr-server/static`.
## Write Data Flow
When archiving something through the CLI:
```mermaid
sequenceDiagram
participant User
participant CLI
participant Core
participant Store
participant DB
User->>CLI: archivr archive path-or-url
CLI->>Core: classify source and call downloader/store helpers
Core->>Store: save raw/structured artifacts
Core->>DB: insert run, source identity, entry, artifacts
CLI->>User: terminal result
```
**YouTube subtitles are sidecar artifacts.** For `Source::YouTubeVideo` (single videos and YouTube playlist/channel
children) with `CaptureConfig::download_subtitles` set (the default), the yt-dlp media call also writes up to two
subtitle files (see [yt-dlp Lifecycle](#yt-dlp-lifecycle)). `subtitles::archive_staged_subtitles` moves them into
`raw/` through `store::archive_staged_file` (same SHA3 dedup) before the temp dir is removed, and
`subtitles::register_subtitle_artifacts` records each one as a `subtitle` artifact once the entry exists:
`storage_area = "raw"`, MIME `text/vtt` or `application/x-subrip`, and `metadata_json`
`{language, kind, format, original_language, origin}` with `kind` `manual`/`auto`/`unknown` and `origin` `capture` or
`summary_fetch`. Registration runs in one `BEGIN IMMEDIATE` transaction and skips an existing
(entry, `subtitle`, blob) row. Subtitle archive and register errors are warnings; the capture still succeeds. The CLI
(`--no-subtitles`), the capture API (`download_subtitles: false`) and the capture dialog toggle all turn it off.
**Pasted text short-circuits most of that.** `perform_text_capture` (`capture.rs`) is not a `Source`
route: there is no locator to classify, no URL probe, and no downloader subprocess. It validates the
title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME type (`text/plain` or
`text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp/<timestamp>/`
and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry
and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI
drives it from `CaptureTextRow` in `CaptureDialog.jsx`. The body is preserved byte-for-byte, the entry's
`original_url` remains empty (no fabricated `text:` URL), and its normal entry-rail preview renders through
`TextPreview.jsx`.
## Web Capture Pipeline
Web pages (`Source::WebPage`) take a longer path than yt-dlp or tweets:
1. **Fetch URL selection.** If `via_freedium` is set and the locator isn't already a Freedium URL, the downloader fetches the page through `freedium-mirror.cfd` with no forwarded cookies. The canonical DB URL stays the original locator.
2. **Browser capture.** `downloader/singlefile.rs` shells out to `single-file-cli` driving headless Chromium. Extensions (uBlock, cookie-consent) and injected browser scripts (modal closer, reader mode) attach per `CaptureConfig`.
3. **Reader mode** (optional) concatenates Mozilla's `Readability.js` from `vendor/readability/` into the SingleFile browser script and stamps absolute URLs on lazy images so the serialised DOM points to fetchable sources.
4. **Freedium cleanup** strips mirror UI (nav, footer, toaster, author header, download control) using multi-signal selectors so article-authored controls aren't hit.
5. **Rust post-processing.** After SingleFile writes the HTML, a Rust pass fetches any images the browser couldn't inline (bounded reads, same-origin cookie forwarding) and embeds them as data URIs. Title extraction runs after embedded font blocks are stripped so large fonts don't push `<title>` beyond the read window.
## Read Data Flow
When opening the web UI:
```mermaid
sequenceDiagram
participant Browser
participant Server
participant Core
participant DB
Browser->>Server: GET /api/archives
Browser->>Server: GET /api/archives/:id/entries
Server->>Core: list_root_entries(conn)
Core->>DB: query archive SQLite
DB-->>Core: rows
Core-->>Server: summaries
Server-->>Browser: JSON
```
## LLM Summaries
Summaries are a **post-capture, manually triggered** subsystem. Nothing in `capture.rs` calls the
summarizer; a summary exists only because someone pressed generate in the Summary section of
`ContextRail.jsx`.
`crates/archivr-core/src/summarizer.rs` defines one `SummaryProvider` trait with four implementations:
Anthropic HTTP, OpenAI-compatible HTTP, the local `claude` CLI, and the local `codex` CLI. Each is
built purely from environment variables (`provider_from_env`), so no key or model name is ever written
into archive data. `PROMPT_VERSION` in the same file stamps every row, so changing the prompt
invalidates the cache instead of silently mixing generations.
`entry_summaries` (schema in `database.rs`) is that cache, unique on
(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the requested provider
model is the cache identity, so the same entry summarised by two providers, two requested models, or after a
prompt change yields distinct rows, while a repeat request with identical inputs reuses one. When a provider
returns its concrete resolved model, it is stored separately and displayed as attribution without changing that
identity. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked;
the frontend polls only for the currently selected entry, and its generate callbacks are scoped to that same
selection. Image-selection behavior remains unchanged.
On startup the server marks interrupted `pending` or `running` attempts failed. Regeneration is non-destructive:
the prior completed summary stays visible until a replacement completes successfully. Public readers receive only
completed summary content, never pending/failed state or diagnostic error text.
The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and
cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is
present, `SummaryBuildOptions::include_images` admits only bounded `media` image candidates and the digest includes both
the flag and selected blob identity, MIME type, and size. The core is still synchronous; the server owns the blocking
boundary. Anthropic HTTP, OpenAI-compatible HTTP, and Codex can transport the selected image data; Claude CLI receives
text only.
```mermaid
flowchart LR
UI["ContextRail Summary"] -->|POST .../summary| Server
Server --> Input["build_summary_input()"]
Input --> Artifacts["entry artifacts on disk"]
Server -->|YouTube, no usable subtitles| Fetch["fetch_subtitles_for_entry (yt-dlp)"]
Fetch --> Artifacts
Fetch -->|still none + transcribe_engine| Transcribe["transcribe_entry (audio → ffmpeg → engine)"]
Transcribe --> Artifacts
Server --> Row["entry_summaries: pending → running"]
Server --> Provider["SummaryProvider (HTTP or CLI)"]
Provider --> Row2["completed / failed"]
UI -->|GET .../summary poll| Row2
```
**X Articles and tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads
the single `primary_media` artifact. A `tweet` or `tweet_thread` entry has no `primary_media` — it has
N `raw_tweet_json` artifacts, one per status in the thread. So the summarizer selects on the
`raw_tweet_json` role instead, loads **all** matching artifacts in order, and joins them with
`\n\n---\n\n`; a `---` line reads as a hard paragraph break to every model, keeping individual
statuses from bleeding into one another. For an X Article, the reducer prefers article text over the tweet's body:
`plain_text`, then flattened ordered blocks, then `preview_text`, then `summary_text`; only then does it fall back to
`full_text`/`text`/`content`/`body`. Any change to article reduction or thread status storage must be mirrored here.
**Tweet entry titles.** When the status is an X Article, the title is `<article.title> — @handle`; otherwise the
tweet-text excerpt (`caption_excerpt`). On server startup `capture::backfill_x_article_titles` retitles only rows whose title still
byte-equals the legacy bare-link title recomputed from their raw JSON and whose raw JSON has an Article title. It is
compare-and-set and idempotent. Titles carry no provenance flag, so exact equality is the guard: renamed titles are
never touched, but a title renamed back to the exact legacy string is retitled. It runs only in `archivr-server`;
CLI-only installs never backfill. Rearchive still does not change titles.
**Thread titles** (`thread_title.rs`) are manual and synchronous: the rail's **Generate title** button on
`tweet_thread` entries calls `POST .../entries/:entry_uid/thread-title` (user role, the same gate as rename). It reuses
the selected provider's transport (`summarizer::complete_plain`) with a cheap title model — admin instance setting
`instance_settings.title_model_<kind>`, else `ARCHIVR_*_TITLE_MODEL`, else a per-provider default, never the summary
model; the server passes the instance value to core as an `Option<&str>` override (core never reads the auth DB) — and
never touches `entry_summaries`. The model returns only the topic; the server sanitizes it and owns the
`Thread about <topic> — @author` format, saved via `update_entry_title`. The bulk panel's **Generate titles** (≥2
selected, ≥1 thread) loops the same endpoint client-side over the selected threads, 2 at a time, with the rail's
provider; a selection change stops it from starting further entries.
**YouTube videos are summarized from a `subtitle` artifact, never the mp4.** For `youtube`/`video` entries
`build_summary_input` skips `primary_media` and reads the entry's `subtitle` artifacts (VTT/SRT by extension or MIME).
`subtitles::subtitle_track_rank` orders them — 0 manual English, 1 manual original language, 2 other manual, 3
transcribed (any language), 4 auto/unknown original language, 5 auto/unknown English, 6 the rest; ties go to the lowest
artifact id — and the first
track that reduces to a non-empty transcript wins. `subtitles::subtitle_to_transcript` drops header/`NOTE`/`STYLE`
blocks, cue ids, timing lines, cue settings, inline tags and ASS overrides, decodes entities, and collapses rolling
auto-caption repeats. Content is `Transcript ({language}, {kind} subtitles):` plus the transcript, truncated at
48,000 chars like every input; the digest covers it, so adding or switching subtitles changes `input_sha256`.
No usable track yields `NoSubtitlesAvailable` (`is_no_subtitles_error`, distinct from unsupported content). The server
preflight then inserts a `pending` row whose `input_sha256` is the placeholder `SUBTITLE_FETCH_PENDING_INPUT_SHA256`
(`"pending-subtitle-fetch"`, never a real digest), skips the cache lookup, and returns 202. Its blocking task loads
cookie rules from the auth DB and calls `build_summary_input_with_subtitle_fetch`: `fetch_subtitles_for_entry` returns
early for non-YouTube entries, non-`http(s)` canonical URLs, or an entry that already has a usable track; otherwise it
runs `fetch_metadata`, `plan_subtitle_request` and a subtitles-only `download_subtitles`, and registers the results
with origin `summary_fetch`. An unreachable video or any yt-dlp error counts as zero subtitles. Then the input is
rebuilt. On success `update_entry_summary_input_sha256` writes the real hash and the provider runs. If there are still
no subtitles and the request named a `transcribe_engine`, `transcriber::transcribe_entry` runs (step 3 of the fixed
order archived → fetched → transcribed → error), then the input is rebuilt once more. Transcription never runs when
either earlier step yields a usable track. Otherwise the row fails with `NO_SUBTITLES_SUMMARY_MESSAGE` (or a
transcription-specific copy when an engine ran) and no provider is called. Entries that already have usable subtitles
keep the synchronous preflight.
**Local transcription** (`transcriber.rs`): engines `whisper` (whisper.cpp or a `script` wrapper), `parakeet` (script),
`phonon2` (English only, `--json` stdout → VTT), configured only by env and gated by `ARCHIVR_TRANSCRIBE_ENGINES`
(`GET /api/summary/transcription-engines` lists enabled ones). One job at a time per process; audio from the archived
media or a yt-dlp audio download, ffmpeg to 16 kHz mono WAV, all within one `ARCHIVR_TRANSCRIBE_TIMEOUT` budget
(ffmpeg and engines via `process::run_with_timeout`; the yt-dlp audio call via `ytdlp.rs`'s own runner). The result is a `subtitle` artifact with kind `transcribed`, origin `transcription`, plus
`engine` and `model` metadata. Spec and deviations:
`docs/superpowers/specs/2026-10-05-local-transcription-fallback.md`.
## yt-dlp Lifecycle
There is no single yt-dlp. Up to three can exist on one machine:
1. **The flake pin** — the `ytDlp` derivation in `flake.nix` fetches an exact release zipapp from
`github.com/yt-dlp/yt-dlp/releases` and wraps it with `python312` + `ffmpeg`. Both the `archivr` and
`archivr-server` wrappers export it as `ARCHIVR_YT_DLP`.
2. **A state-dir install** — `archivr yt-dlp update` downloads the latest zipapp and installs it
atomically (staged file, then rename) at `<state_dir>/yt-dlp/yt-dlp` with a sibling `.version`
sentinel that lets repeat runs skip the download.
3. **Whatever is on PATH** — the historical behaviour, and the last-resort fallback.
`resolve_yt_dlp()` in `downloader/ytdlp.rs` picks between them and caches the result in an `RwLock` until
`refresh_yt_dlp()`: `ARCHIVR_YT_DLP_FORCE` wins outright if it points at a real file; otherwise the pinned and
state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`,
so plain string ordering is chronological — with exact ties going to the state-dir copy the user
deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status`
prints every candidate, its version, and the winner; when the force variable applies, it includes that
forced candidate and selects it as the winner.
Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the
current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp`
block and opens a PR), `archivr yt-dlp update` / Settings › Instance › yt-dlp for one machine, or editing `flake.nix` by
hand.
**The JS runtime has the same shape.** YouTube's player challenges are solved by yt-dlp's EJS solver, which needs
Deno ≥ 2.3.0. Candidates: the Nix/Docker pin in `ARCHIVR_DENO`, `<state_dir>/deno/deno`, and `deno` on PATH.
`resolve_js_runtime()` in `downloader/js_runtime.rs` (cached in an `RwLock` until `refresh_js_runtime()`, returns an owned clone, warnings printed once per resolution) returns a valid
`ARCHIVR_JS_RUNTIME` force (`RUNTIME[:ABS_PATH]`, `deno|node|bun|quickjs`, invalid values warned and ignored)
outright; otherwise it probes the pinned and state-dir Deno, drops anything below 2.3.0, compares real semver
(`DenoVersion`, so 2.10.0 > 2.9.7) and keeps the newest, ties to the state dir; then PATH; else `None` plus a one-time
warning. Only Deno is chosen automatically. Every yt-dlp process is built by `yt_dlp_command()` in `ytdlp.rs`, which
appends `js_runtime_args()` (`--js-runtimes deno:<path>`; non-Deno forces get `--no-js-runtimes` first) — the
`download` closure (incl. the media-only retry), `download_subtitles`, `fetch_metadata_with_timeout`,
`fetch_playlist_info` and `probe_playlist_qualities`. The update also installs the latest Deno
(`crates/archivr-core/src/downloader/deno_install.rs`): download, extract to `deno.new`, require `--version` to equal the release,
then atomic rename. `status` adds a JS runtime table with rows `force (ARCHIVR_JS_RUNTIME)`, `env (ARCHIVR_DENO)`,
`state-dir`, `path (deno)`; the star goes to the winning `JsRuntimeRole` from `resolve_js_runtime_with_role()`, not to
every row whose path matches (the Nix wrappers' pinned Deno is also on PATH).
**One updater, two front ends.** `downloader/ytdlp_tools.rs` owns `install_yt_dlp`, `update_tools` and the
`tools_status()` model; the CLI renders it as text and `GET /api/admin/yt-dlp` / `POST /api/admin/yt-dlp/update`
(admin, 409 while an update runs) serve it to Settings › Instance › yt-dlp. `update_tools` calls `refresh_yt_dlp()` /
`refresh_js_runtime()` after each successful component, so a UI update takes effect without a restart; commands
already built keep their old binary. A CLI update runs in another process, so a running server still needs a restart.
**Subtitles ride on the media call.** When capture wants subtitles, `plan_subtitle_request` builds a bounded (≤ 2
tracks) request from the `--dump-json` metadata capture already fetched. `download` then appends `--write-subs` and/or
`--write-auto-subs` (only the kinds planned), `--sub-langs <codes>`, `--sub-format vtt/srt/best` and `--ignore-errors`.
There is no `--convert-subs`, so ffmpeg is never needed for subtitles. yt-dlp treats `--sub-langs` entries as regexes,
so planned codes are limited to `[A-Za-z0-9][A-Za-z0-9-]*`. Without metadata the request falls back to `en` and
`.*-orig`. If the combined call exits non-zero without staging media and its stderr mentions subtitles, it is
retried once with the exact legacy media-only arguments (`should_retry_media_only`), and subtitle files from the
first attempt are still collected; other failures (private, deleted, geo-blocked) fail without a retry.
`collect_staged_outputs` splits the staging dir into the
media file and `<stem>.<lang>.<vtt|srt>` sidecars; other subtitle formats are dropped with a warning. Summary-time
fetches use `download_subtitles` instead: `--skip-download --no-playlist --ignore-no-formats-error` plus the same
subtitle args and `--ignore-errors`, staged under `temp/subs-<uuid>/`. A non-zero exit is tolerated if any subtitle
file was written. The summary-time metadata probe and subtitle call are killed after `ARCHIVR_SUMMARY_CLI_TIMEOUT`
(a timeout counts as "no subtitles"); capture-time calls stay unbounded. Blob cleanup refuses to run while such a
fetch is in flight (`has_pending_subtitle_fetches`). Every one of these calls is built with `yt_dlp_command()`.
## Where To Edit
| Feature kind | Edit here |
|---|---|
| DB schema, inserts, archive runs, entries, tags | `crates/archivr-core/src/database.rs` |
| Capture orchestration, `Source` routing, `CaptureConfig` | `crates/archivr-core/src/capture.rs` |
| Archive opening, listing entries, entry detail, runs | `crates/archivr-core/src/archive.rs` |
| Download/save behavior | `crates/archivr-core/src/downloader/` |
| YouTube playlist/channel download, playlist probe, sync mode | `crates/archivr-core/src/downloader/ytdlp.rs` and `capture.rs` |
| Which yt-dlp binary runs (resolver, state dir, version probe) | `crates/archivr-core/src/downloader/ytdlp.rs` |
| Which JS runtime yt-dlp gets (Deno resolver, `ARCHIVR_JS_RUNTIME`, `--js-runtimes` args) | `crates/archivr-core/src/downloader/js_runtime.rs` |
| yt-dlp/Deno update orchestration and status model (CLI + admin API) | `crates/archivr-core/src/downloader/ytdlp_tools.rs` |
| Deno installer (CLI and UI update) | `crates/archivr-core/src/downloader/deno_install.rs` |
| Settings › Instance › yt-dlp section | `frontend/src/components/SettingsView.jsx` (`YtDlpSection`) |
| YouTube subtitles (track planning, yt-dlp args, staging) | `crates/archivr-core/src/downloader/ytdlp.rs` |
| Subtitle artifacts, transcript reduction, track ranking, summary-time fetch | `crates/archivr-core/src/subtitles.rs` |
| Pasted-text capture (staging, hashing, MIME allowlist) | `crates/archivr-core/src/downloader/text.rs` and `capture.rs` |
| LLM summary providers, prompt, `PROMPT_VERSION`, input building | `crates/archivr-core/src/summarizer.rs` |
| Local transcription engines, audio, job slot, `transcribe_entry` | `crates/archivr-core/src/transcriber.rs` |
| Subprocess timeout runner | `crates/archivr-core/src/process.rs` |
| Shared `ARCHIVR_*` env helpers | `crates/archivr-core/src/env_config.rs` |
| Thread-title generation, cheap title models | `crates/archivr-core/src/thread_title.rs` |
| X Article titles + startup backfill | `crates/archivr-core/src/capture.rs` (`backfill_x_article_titles`), called from `archivr-server/src/main.rs` |
| `entry_summaries` schema and summary CRUD | `crates/archivr-core/src/database.rs` |
| CLI commands, argument parsing, terminal output | `crates/archivr-cli/src/main.rs` |
| Server API routes | `crates/archivr-server/src/routes.rs` |
| Auth model (users, sessions, tokens, roles) | `crates/archivr-server/src/auth.rs` |
| Mounted archive config model | `crates/archivr-server/src/registry.rs` |
| Frontend root state + routing | `frontend/src/App.jsx` |
| Frontend API client | `frontend/src/api.js` |
| Frontend components | `frontend/src/components/` |
| Summary UI (provider selector, transcription engine, generate, polling, thread Generate title) | `frontend/src/components/ContextRail.jsx` |
| Text/Markdown entry preview | `frontend/src/components/TextPreview.jsx` |
| "Add text" capture row | `frontend/src/components/CaptureDialog.jsx` |
| Frontend styling | `frontend/src/styles.css` |
## Practical Feature Rule
If a feature affects archive truth, start in `archivr-core`.
If a feature is only how the terminal behaves, edit `archivr-cli`.
If a feature is only how the browser sees or calls things, edit `archivr-server` and the static UI.
If a browser feature needs new data, the usual order is:
1. Add or query the data in `archivr-core`.
2. Expose it in `archivr-server`.
3. Render it in the static UI.
## Server Capabilities
The server both reads and writes archive data. Capture jobs are asynchronous: `POST /api/archives/:id/captures` inserts a job row, spawns a blocking task, and returns immediately; the frontend polls until the job completes or fails. Heavy work stays synchronous inside `archivr-core`.
**Auth model.** A separate `archivr-auth.sqlite` (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are `u32` flags (`GUEST`, `USER`, `ADMIN`, `OWNER`) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is `setup_guard` → `login_rate_limit` → `security_headers`; route families are classified `READ / ADMIN / WRITE / STATIC` in `routes.rs`. Configurable per-action permissions are `u32` role masks on the auth `instance_settings` singleton (currently `reorder_children_role_bits`). Handlers read them per request and allow when `caller_bits & mask != 0`, so changes take effect without re-login. Only the Owner may change them, and masks may contain only existing non-guest role bits (`database::grantable_role_bits`).
**Search** is server-side free-text filtering over entry fields and the latest completed summary. The summary JSON is
searched as text, so generated `tags` participate. Older completed summaries stay searchable while a newer request is
pending or failed; rows with no completed summary contribute no summary-derived match.
**Admin view** covers mounted archives, users, sessions, and API tokens.