Archivr — Preserve what matters. Forever.

MIT License Rust 2024 Self-hosted

--- Archivr is a self-hosted tool for capturing and preserving digital content — YouTube videos and playlists, tweets and threads, Instagram, TikTok, web pages, and local files — into self-contained, locally-owned archives. Content is stored in SQLite with SHA3-256 blob deduplication, hierarchical tags, a browser-based UI, and role-based auth. ## Table of Contents - [Features](#features) - [Quick Start](#quick-start) - [With Nix](#with-nix) - [With Docker](#with-docker-1) - [Architecture](#architecture) - [Supported Inputs](#supported-inputs) - [YouTube playlists and channels](#youtube-playlists-and-channels) - [Video quality and audio-only](#video-quality-and-audio-only) - [YouTube subtitles](#youtube-subtitles) - [Text notes](#text-notes) - [Configuration](#configuration) - [TOML config file](#toml-config-file) - [Environment variables](#environment-variables) - [LLM providers](#llm-providers) - [Keeping yt-dlp and its JS runtime fresh](#keeping-yt-dlp-and-its-js-runtime-fresh) - [JavaScript runtime (Deno)](#javascript-runtime-deno) - [Troubleshooting](#troubleshooting) - [Deployment](#deployment) - [Security](#security) - [NixOS](#hosting-on-nixos) - [Docker](#hosting-with-docker) - [Development](#development) - [License](#license) ## Features - **Social media** — YouTube (videos, shorts, playlists, channels with sync mode; subtitles saved with each video by default), X/Twitter (tweet and thread JSON + media downloads), Instagram, TikTok, Facebook, Reddit, Snapchat via yt-dlp - **Web pages** — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode - **Local files** — import any file from disk by `file://` path - **Deduplication** — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once - **Tags and search** — hierarchical tag tree, full-text search (including the latest completed summary and its generated JSON tags), filterable entry list - **Multiple archives** — the server mounts any number of separate archives from a single TOML config - **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords; the Owner can choose which roles (including custom ones) may reorder child entries - **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download - **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option; YouTube videos are summarized from their subtitles - **Text notes** — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview - **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block ## Quick Start ### With Nix ```sh # Create an archive nix run github:thegeneralist/archivr#archivr -- init ./my-archive --name "My Archive" # Archive something nix run github:thegeneralist/archivr#archivr -- archive https://www.youtube.com/watch?v=dQw4w9WgXcQ # Start the web UI (reads ./archivr-server.toml) nix run github:thegeneralist/archivr#archivr-server ``` Create `archivr-server.toml` next to where you run the command: ```toml auth_db_path = "/absolute/path/to/archivr-auth.sqlite" [[archives]] id = "personal" label = "Personal" archive_path = "/absolute/path/to/my-archive/.archivr" ``` Then open `http://127.0.0.1:8080`. On the first visit you will be prompted to create the owner account. ### With Docker ```sh mkdir config cp docker/config.example.toml config/archivr-server.toml # Edit config/archivr-server.toml # Initialize the archive on the persistent volume (run once) docker compose run --rm archivr archivr init \ /data/archives/main /data/archives/main/.archivr/store \ --name "Main Archive" docker compose up -d ``` Open `http://localhost:8080`. See [Hosting with Docker](#hosting-with-docker) for volume layout and Twitter/X credential setup. ## Architecture Two binaries: | Binary | Purpose | |---|---| | `archivr` | CLI — create archives (`init`) and add content (`archive`) | | `archivr-server` | Web server — browse and search one or more archives via browser UI | Archive layout created by `archivr init`: ``` my-archive/ ├── .archivr/ # metadata: name, store_path, archivr.sqlite └── store/ ├── raw/ # deduplicated blobs: raw/A/B/.ext ├── raw_tweets/ # tweet and thread JSON ├── structured/ # structured metadata outputs └── temp/ # staging area during capture ``` A separate auth database (`archivr-auth.sqlite`, path set in TOML) holds users, sessions, API tokens, and role bits. It is independent of individual archives. ## Supported Inputs `archivr archive ` accepts URLs and platform shorthands: | Platform | Input examples | |---|---| | Local file | `file:///absolute/path/to/file.pdf` | | YouTube video / short | `https://youtube.com/watch?v=ID` · `yt:video/ID` · `yt:short/ID` | | YouTube playlist | `https://youtube.com/playlist?list=ID` | | YouTube channel | `https://youtube.com/@handle` | | X/Twitter tweet (JSON) | `tweet:ID` · `x:tweet:ID` · `twitter:tweet:ID` | | X/Twitter thread (JSON) | `x:thread:ID` · `twitter:thread:ID` | | X/Twitter media download | `tweet:media:ID` | | Instagram | Direct URL · `instagram:ID` | | TikTok | Direct URL · `tiktok:ID` | | Facebook | Direct URL · `facebook:ID` | | Reddit | Direct URL · `reddit:ID` | | Snapchat | Direct URL · `snapchat:ID` | | Arbitrary URL / web page | Any `https://` URL | X Articles are titled `
— @handle` from the article's own title instead of the tweet text (usually a bare `t.co` link). Article entries archived before this change are retitled on the next `archivr-server` start, but only while their title still equals the old auto-generated bare-link title; a renamed entry is left alone. Titles have no "edited" flag, so an entry you renamed back to exactly that old title is retitled too. CLI-only installs never run this pass — it runs at server startup only. X threads keep `Thread by @handle` at capture. **Generate title** in the entry rail (thread entries, user role and up) asks the selected Summary provider's cheap title model for a short topic and saves `Thread about — @handle` as the entry title (`POST /api/archives/:id/entries/:uid/thread-title`, body `{"provider": ""}`); rename it like any title. With ≥2 entries selected and at least one X thread, the bulk panel's **Generate titles** does the same for each selected thread (non-threads skipped; uses the Summary provider selected in the rail), two at a time, with progress and an `X updated, Y failed` summary; changing the selection stops it picking up further entries. Models are listed under [LLM providers](#llm-providers). ### YouTube playlists and channels Capturing a playlist or channel creates a **container entry** with each video archived as a child beneath it. Before downloading, the UI probes each video for available quality options — set quality per-video or apply one to the whole batch. Individual videos can be excluded with the remove button. **Sync mode:** when re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container is reused. **Reordering:** if your role is allowed (by default Admin and Owner), expand a container on the main page and drag a video by its handle to change its position. On touch screens and phone-sized windows, where dragging isn't available, ↑/↓ buttons appear instead. Alt+↑/↓ on a selected row works everywhere. The order is saved to the archive. Videos added later by sync mode appear at the end. The Owner chooses which roles, including custom roles, may reorder under **Settings → Instance → Permissions**. ### Video quality and audio-only When capturing a yt-dlp-backed source through the web UI, a metadata probe runs first and populates the quality selector with heights actually available in that video: | `qualities` | `has_audio` | UI shows | |---|---|---| | `["1080p", "720p", …]` | `true` | Best / heights / Audio only | | `["1080p", …]` | `false` | Best / heights | | `[]` | `true` | Audio only (pre-selected) | | `[]` | `false` | "No media detected" | | probe fails (502) | — | picker hidden; capture still submittable | The `POST /api/archives/:id/captures` endpoint accepts an optional `quality` field: ```json { "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" } { "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" } ``` `"audio"` selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting `quality` or passing `"best"` downloads at the highest available quality. ### YouTube subtitles YouTube video captures save subtitles next to the video by default: single videos, shorts, and every video archived from a YouTube playlist or channel, at any quality including audio-only. Subtitles apply to YouTube videos only — YouTube Music, Spotify, X, TikTok, and the other yt-dlp sources are downloaded without them. At most two tracks are saved, chosen from the metadata yt-dlp already fetches for the capture: - English, plus the video's original language when that isn't English. - Manual (uploader-provided) tracks are preferred. If the original language has no manual track, its auto-generated track is used; auto-generated English is used only when nothing else was found. - If the metadata probe fails, yt-dlp is asked for `en` and any `-orig` (original-language) track instead. Tracks are requested as VTT, with SRT accepted; nothing is converted, and other subtitle formats are dropped. Each file goes through the usual SHA3-256 dedup into `store/raw/` and is recorded as a `subtitle` artifact of the entry, along with its language, whether it was manual or auto-generated, and its format. Subtitle failures never fail a capture. Subtitles are requested in the same yt-dlp call as the media with `--ignore-errors`, so a missing track or a rate-limited caption request only logs a warning. If that call still fails, the media is retried once without subtitles. Subtitles are on unless you turn them off for a capture: - **Web UI:** the **Download subtitles** toggle in the capture dialog. - **API:** `"download_subtitles": false` in the `POST /api/archives/:id/captures` body. Omitting the field means `true`. - **CLI:** `archivr archive --no-subtitles `. Videos captured without subtitles can still be summarized; see [LLM providers](#llm-providers) and [Local transcription](#local-transcription-optional). #### Local transcription (optional) When a YouTube video has no subtitles at summary time, Archivr can transcribe its audio on the server. The order is fixed and transcription never runs if an earlier step yields usable subtitles: 1. archived subtitles; 2. subtitles fetched from the original video (subtitles only, no media); 3. local transcription with the engine chosen in the Summary panel; 4. the no-subtitles error (or a transcription-specific error if step 3 ran and failed). The feature is off until `ARCHIVR_TRANSCRIBE_ENGINES` lists at least one configured engine (env vars in [Local transcription env](#local-transcription)). Then the Summary panel shows a second selector on YouTube videos — **No local transcription** (default) or an enabled engine — remembered for the browser session. The API field is `"transcribe_engine": ""` in the summary POST body; an unknown or unconfigured engine is a 400. Enabled engines are listed by `GET /api/summary/transcription-engines`. | Engine (`kind`) | Languages | Runs on | Install | |---|---|---|---| | Whisper (`whisper`) | ~99 (`*.en` models English only) | CPU, Metal, CUDA/Vulkan | whisper.cpp `whisper-cli` + a ggml model (default backend), or a faster-whisper wrapper (`script` backend) | | NVIDIA Parakeet (`parakeet`) | v2 English; v3 25 European | NVIDIA GPU (NeMo), Apple silicon (parakeet-mlx), CPU (ONNX) | your own wrapper script; weights CC-BY-4.0 | | Fermion Phonon-2 (`phonon2`) | **English only** (hard-coded) | Apple silicon (MLX), x86-64/Arm CPU, CUDA | `pip install fermion-research` + platform runtime; weights CC-BY-4.0, CLI licence unknown | Setup examples: ```sh # whisper.cpp nix shell nixpkgs#whisper-cpp curl -LO https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo.bin export ARCHIVR_TRANSCRIBE_ENGINES=whisper ARCHIVR_WHISPER_MODEL=$PWD/ggml-large-v3-turbo.bin # Parakeet via parakeet-mlx (Apple silicon), wrapper from spec Appendix A.3 export ARCHIVR_TRANSCRIBE_ENGINES=parakeet ARCHIVR_PARAKEET_CLI=/opt/transcribe/parakeet.py # Phonon-2 (vendor package; Apple silicon shown) python3 -m venv /opt/transcribe && /opt/transcribe/bin/pip install fermion-research mlx mlx-audio mlx-lm soundfile scipy zstandard export ARCHIVR_TRANSCRIBE_ENGINES=phonon2 ARCHIVR_PHONON2_CLI=/opt/transcribe/bin/fermion ``` **Script contract** (Whisper `script` backend and Parakeet). Archivr runs `