From 253f7792167d10bb6e66444e2ce302c9b969f274 Mon Sep 17 00:00:00 2001 From: TheGeneralist <180094941+thegeneralist01@users.noreply.github.com> Date: Mon, 5 Oct 2026 00:30:39 +0200 Subject: [PATCH] Add YouTube subtitles, local transcription, self-updating yt-dlp/Deno, X Article and thread titles - Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles) - Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error - Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only) - Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s) - Settings > Instance > yt-dlp: status and in-app update without restart - X Article titles from article.title, with idempotent startup backfill - Thread title generation (single and bulk) with per-provider cheap models - Per-provider title model settings in Settings > Instance - Docs, mental model, AGENTS.md and transcription spec updated --- .gitignore | 4 + AGENTS.md | 77 +- ARCHIVR-MENTAL-MODEL.md | 123 +- Cargo.lock | 47 +- Cargo.toml | 2 + Dockerfile | 37 +- crates/archivr-cli/Cargo.toml | 2 - crates/archivr-cli/src/main.rs | 325 ++- crates/archivr-core/Cargo.toml | 4 + crates/archivr-core/src/capture.rs | 377 +++- crates/archivr-core/src/database.rs | 396 +++- .../src/downloader/deno_install.rs | 328 ++++ .../archivr-core/src/downloader/js_runtime.rs | 684 +++++++ crates/archivr-core/src/downloader/mod.rs | 29 + crates/archivr-core/src/downloader/ytdlp.rs | 1191 ++++++++++- .../src/downloader/ytdlp_tools.rs | 559 ++++++ crates/archivr-core/src/env_config.rs | 109 ++ crates/archivr-core/src/lib.rs | 5 + crates/archivr-core/src/process.rs | 381 ++++ crates/archivr-core/src/subtitles.rs | 844 ++++++++ crates/archivr-core/src/summarizer.rs | 1085 +++++++--- crates/archivr-core/src/thread_title.rs | 650 ++++++ crates/archivr-core/src/transcriber.rs | 1742 +++++++++++++++++ crates/archivr-server/src/main.rs | 19 + crates/archivr-server/src/routes.rs | 1172 ++++++++++- docker-compose.yml | 8 + docs/README.md | 304 ++- ...2026-10-05-local-transcription-fallback.md | 874 +++++++++ flake.nix | 18 +- frontend/src/api.js | 59 +- frontend/src/components/CaptureDialog.jsx | 18 + frontend/src/components/ContextRail.jsx | 187 +- frontend/src/components/SettingsView.jsx | 245 ++- frontend/src/styles.css | 19 + modules/nixos/archivr-server.nix | 36 + 35 files changed, 11377 insertions(+), 583 deletions(-) create mode 100644 crates/archivr-core/src/downloader/deno_install.rs create mode 100644 crates/archivr-core/src/downloader/js_runtime.rs create mode 100644 crates/archivr-core/src/downloader/ytdlp_tools.rs create mode 100644 crates/archivr-core/src/env_config.rs create mode 100644 crates/archivr-core/src/process.rs create mode 100644 crates/archivr-core/src/subtitles.rs create mode 100644 crates/archivr-core/src/thread_title.rs create mode 100644 crates/archivr-core/src/transcriber.rs create mode 100644 docs/superpowers/specs/2026-10-05-local-transcription-fallback.md diff --git a/.gitignore b/.gitignore index cddda57..b951746 100644 --- a/.gitignore +++ b/.gitignore @@ -9,6 +9,10 @@ !docs/README* !docs/branding/ !docs/branding/** +# Dated design specs are tracked; docs/superpowers/plans/ stays ignored. +!docs/superpowers/ +!docs/superpowers/specs/ +!docs/superpowers/specs/** !crates !crates/** # Static assets are built by Nix (frontendStatic derivation in flake.nix). diff --git a/AGENTS.md b/AGENTS.md index 6566445..0b92f12 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -12,12 +12,19 @@ Three crates with a strict ownership split — **core owns truth; CLI and server - `crates/archivr-core` — domain library: capture orchestration, SQLite schema/CRUD, downloaders, hashing. New archive features start here. - `crates/archivr-server` — Axum HTTP API + auth + static frontend serving. -- `crates/archivr-cli` — clap-based CLI (`archivr` binary): `init`, `archive` subcommands. +- `crates/archivr-cli` — clap-based CLI (`archivr` binary): `init`, `archive`, `yt-dlp status|update` subcommands. -Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture.rs`) routes by platform/shorthand (`yt:`, `x:`, `tweet:` …) → platform downloader (`downloader/ytdlp.rs`, `tweets.rs`, `singlefile.rs`, `http.rs`, `local.rs`) stages into `temp/` → SHA3-256 dedup (`hash.rs`, `downloader/store.rs`) moves blobs to `raw/A/B/HASH.EXT` → rows written to `archivr.sqlite` (runs, entries, artifacts, blobs) → served via `/api/archives/:id/...`. `CaptureConfig` carries per-request toggles (uBlock, reader mode, Freedium mirror, etc.); when `via_freedium` is set, the fetch URL is rewritten through `freedium-mirror.cfd` while the canonical DB URL stays the original locator. +Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture.rs`) routes by platform/shorthand (`yt:`, `x:`, `tweet:` …) → platform downloader (`downloader/ytdlp.rs`, `tweets.rs`, `singlefile.rs`, `http.rs`, `local.rs`) stages into `temp/` → SHA3-256 dedup (`hash.rs`, `downloader/store.rs`) moves blobs to `raw/A/B/HASH.EXT` → rows written to `archivr.sqlite` (runs, entries, artifacts, blobs) → served via `/api/archives/:id/...`. `CaptureConfig` carries per-request toggles (uBlock, reader mode, Freedium mirror, YouTube subtitles, etc.); when `via_freedium` is set, the fetch URL is rewritten through `freedium-mirror.cfd` while the canonical DB URL stays the original locator. YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the flat-playlist probe (fetching per-video quality metadata before archiving); the multi-video download loop and sync mode (skipping already-archived videos when re-archiving a playlist or channel) live in `capture.rs`. Child order is persisted in `archived_entries.position` (append-only at insert in `database::create_archived_entry`; rewritten only by `database::reorder_child_entries` behind `PUT …/entries/:entry_uid/children/order` (allowed roles = `InstanceSettings::reorder_children_role_bits`, default ADMIN|OWNER, Owner-editable)). +YouTube videos (`Source::YouTubeVideo`: single videos and playlist/channel children; no other platform) also get up +to two subtitle tracks (English + original language, manual over auto, VTT/SRT) in the same yt-dlp call by default. +`CaptureConfig::download_subtitles` (manual `Default` = true) gates it; the API body field `download_subtitles` (absent += true), the CaptureDialog "Download subtitles" toggle and CLI `archivr archive --no-subtitles` feed it. Files are +archived into `raw/` and registered as `subtitle` artifacts (`crates/archivr-core/src/subtitles.rs`). Subtitle +failures are `eprintln!` warnings, never capture failures. + Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the blob lands in `raw/` like any other artifact. Entrypoint is @@ -29,6 +36,20 @@ LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/sr table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the provider set and the status lifecycle. +YouTube video entries are summarized from the best-ranked `subtitle` artifact reduced to a transcript, not the mp4. +With no usable track the POST still returns 202: the row is created `pending` with the placeholder `input_sha256` +`pending-subtitle-fetch`, and the background task fetches subtitles only (`build_summary_input_with_subtitle_fetch` → +`subtitles::fetch_subtitles_for_entry`), writes the real hash, then runs the provider. The order is fixed (user +requirement): archived subtitles → fetched subtitles → local transcription (`transcriber::transcribe_entry`, only when +the POST body names a `transcribe_engine` and both earlier steps gave nothing) → error. If there are still none, the +row fails with `NO_SUBTITLES_SUMMARY_MESSAGE` (or a transcription-specific copy when an engine ran) and no provider is +called. Transcripts are `subtitle` artifacts with `kind: "transcribed"`, `origin: "transcription"`, `engine`, `model`. +Spec + deviations: `docs/superpowers/specs/2026-10-05-local-transcription-fallback.md`. + +X thread titles: `POST /api/archives/:archive_id/entries/:entry_uid/thread-title` (`ROLE_USER`, same gate as title +PATCH; body `{provider}`) runs `thread_title::generate_thread_title` synchronously in one blocking task and saves +`Thread about — @author` via `database::update_entry_title`. Never touches `entry_summaries`. + The requested provider model is the cache identity; a provider-returned resolved model is display attribution. At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed summary visible until its replacement completes; public readers receive completed content only, never diagnostics. @@ -43,7 +64,7 @@ remains synchronous: the server puts provider work in its blocking boundary rath Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only. Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary. -Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8; per-action role masks (e.g. `reorder_children_role_bits`) live on its `instance_settings` row). +Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8; per-action role masks (e.g. `reorder_children_role_bits`) and per-provider thread-title models (`title_model_`) live on its `instance_settings` row). The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`. @@ -77,7 +98,7 @@ bun run dev # Vite dev server bun run build # → ../crates/archivr-server/static (gitignored; only needed for bare cargo run) # Nix -nix develop # devshell: yt-dlp, nushell, uv, twitter-api-client +nix develop # devshell: yt-dlp, deno, nushell, uv, twitter-api-client nix build .#archivr-server # also .#archivr-cli, .#archivr-all; builds frontend automatically # If frontendDeps hash is stale (after bun.lock/package.json change): # nix build 2>&1 | grep "got:" → paste hash into flake.nix frontendDepsHash @@ -98,31 +119,48 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy - **State**: Axum `AppState { registry, auth_db_path, login_attempts }` via `State` extractor; middleware stack = `setup_guard` (503 until owner exists) → `login_rate_limit` (5/15min per IP) → `security_headers`. - **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2. - **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal. -- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`. Downloaders shell out to subprocesses; resolve binaries through these vars. +- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_DENO`, `ARCHIVR_JS_RUNTIME`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`, `ARCHIVR_FFMPEG`. Local transcription: `ARCHIVR_TRANSCRIBE_ENGINES` (the gate; unset = off), `ARCHIVR_TRANSCRIBE_TIMEOUT` (default 3600s, whole job), `ARCHIVR_WHISPER_BACKEND` (`whisper_cpp`|`script`) / `ARCHIVR_WHISPER_CLI` / `ARCHIVR_WHISPER_MODEL` / `ARCHIVR_WHISPER_LANGUAGES`, `ARCHIVR_PARAKEET_CLI` / `ARCHIVR_PARAKEET_MODEL` / `ARCHIVR_PARAKEET_LANGUAGES`, `ARCHIVR_PHONON2_CLI` / `ARCHIVR_PHONON2_MODEL`. Downloaders shell out to subprocesses; resolve binaries through these vars. Env helpers (`required_env`, `env_or`, `optional_env`, `env_timeout`, `resolve_cli`) live in `env_config.rs`; bounded subprocesses go through `process::run_with_timeout` (yt-dlp calls keep `ytdlp.rs`'s private runner). - **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the *pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR` - relocates the state dir. Never spawn bare `yt-dlp` — call the resolver. -- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`. + relocates the state dir. The JS runtime works the same way: `resolve_js_runtime()` (`downloader/js_runtime.rs`) + picks the newest Deno ≥ 2.3.0 from `ARCHIVR_DENO` and `/deno/deno` (ties → state dir), then PATH; + `ARCHIVR_JS_RUNTIME=RUNTIME[:ABS_PATH]` forces it. Both caches are `RwLock`s refreshed by `refresh_yt_dlp()` / + `refresh_js_runtime()` after each successful component install (`ytdlp_tools::update_tools`), so a UI update needs + no restart; `resolve_js_runtime()` returns an owned `Option`. Never cache a resolved path across + operations. Never spawn bare yt-dlp — build every command with + `yt_dlp_command()` so the resolver-chosen binary *and* `--js-runtimes` args are applied. +- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Thread titles use `ARCHIVR_ANTHROPIC_TITLE_MODEL` (default `claude-haiku-4-5`), `ARCHIVR_OPENAI_TITLE_MODEL` (`gpt-4o-mini`), `ARCHIVR_CLAUDE_TITLE_MODEL` (`haiku`), `ARCHIVR_CODEX_TITLE_MODEL` (`gpt-6-luna`, override if unavailable) — never the summary model; `thread_title.rs` reuses provider transports via `summarizer::complete_plain`. The one exception to env-only config: admins may override the title model per provider in the auth DB (`instance_settings.title_model_{anthropic_http,openai_compatible,claude_cli,codex_cli}`; PATCH `/api/admin/instance-settings` trims, blank clears, rejects overlong or whitespace/control chars; GET returns `title_models.` with `model`/`source`/`fallback`). Precedence instance > `ARCHIVR_*_TITLE_MODEL` > default; the server passes the instance value to core as an `Option<&str>` override — core never reads the auth DB. Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`. - **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class. - **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract. +- **Transcription engines follow the same rule**: whisper.cpp writes `-ovtt`; script engines (Whisper `script`, Parakeet) are run as `