18 KiB
Archivr is a self-hosted tool for capturing and preserving digital content — YouTube videos and playlists, tweets and threads, Instagram, TikTok, web pages, and local files — into self-contained, locally-owned archives. Content is stored in SQLite with SHA3-256 blob deduplication, hierarchical tags, a browser-based UI, and role-based auth.
Table of Contents
- Features
- Quick Start
- Architecture
- Supported Inputs
- Configuration
- Keeping yt-dlp fresh
- Deployment
- Development
- License
Features
- Social media — YouTube (videos, shorts, playlists, channels with sync mode), X/Twitter (tweet and thread JSON + media downloads), Instagram, TikTok, Facebook, Reddit, Snapchat via yt-dlp
- Web pages — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
- Local files — import any file from disk by
file://path - Deduplication — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
- Tags and search — hierarchical tag tree, full-text search, filterable entry list
- Multiple archives — the server mounts any number of separate archives from a single TOML config
- Role-based auth — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
- Quality selection — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
- LLM summaries — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local
claudeCLI, or a localcodexCLI; triggered manually from the entry rail, never automatically on capture - Text notes — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser
- In-progress capture indicator — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
Quick Start
With Nix
# Create an archive
nix run github:thegeneralist/archivr#archivr -- init ./my-archive --name "My Archive"
# Archive something
nix run github:thegeneralist/archivr#archivr -- archive https://www.youtube.com/watch?v=dQw4w9WgXcQ
# Start the web UI (reads ./archivr-server.toml)
nix run github:thegeneralist/archivr#archivr-server
Create archivr-server.toml next to where you run the command:
auth_db_path = "/absolute/path/to/archivr-auth.sqlite"
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
Then open http://127.0.0.1:8080. On the first visit you will be prompted to create the owner account.
With Docker
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit config/archivr-server.toml
# Initialize the archive on the persistent volume (run once)
docker compose run --rm archivr archivr init \
/data/archives/main /data/archives/main/.archivr/store \
--name "Main Archive"
docker compose up -d
Open http://localhost:8080. See Hosting with Docker for volume layout and Twitter/X credential setup.
Architecture
Two binaries:
| Binary | Purpose |
|---|---|
archivr |
CLI — create archives (init) and add content (archive) |
archivr-server |
Web server — browse and search one or more archives via browser UI |
Archive layout created by archivr init:
my-archive/
├── .archivr/ # metadata: name, store_path, archivr.sqlite
└── store/
├── raw/ # deduplicated blobs: raw/A/B/<sha3-256>.ext
├── raw_tweets/ # tweet and thread JSON
├── structured/ # structured metadata outputs
└── temp/ # staging area during capture
A separate auth database (archivr-auth.sqlite, path set in TOML) holds users, sessions, API tokens, and role bits. It is independent of individual archives.
Supported Inputs
archivr archive <locator> accepts URLs and platform shorthands:
| Platform | Input examples |
|---|---|
| Local file | file:///absolute/path/to/file.pdf |
| YouTube video / short | https://youtube.com/watch?v=ID · yt:video/ID · yt:short/ID |
| YouTube playlist | https://youtube.com/playlist?list=ID |
| YouTube channel | https://youtube.com/@handle |
| X/Twitter tweet (JSON) | tweet:ID · x:tweet:ID · twitter:tweet:ID |
| X/Twitter thread (JSON) | x:thread:ID · twitter:thread:ID |
| X/Twitter media download | tweet:media:ID |
Direct URL · instagram:ID |
|
| TikTok | Direct URL · tiktok:ID |
Direct URL · facebook:ID |
|
Direct URL · reddit:ID |
|
| Snapchat | Direct URL · snapchat:ID |
| Arbitrary URL / web page | Any https:// URL |
YouTube playlists and channels
Capturing a playlist or channel creates a container entry with each video archived as a child beneath it. Before downloading, the UI probes each video for available quality options — set quality per-video or apply one to the whole batch. Individual videos can be excluded with the remove button.
Sync mode: when re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container is reused.
Video quality and audio-only
When capturing a yt-dlp-backed source through the web UI, a metadata probe runs first and populates the quality selector with heights actually available in that video:
qualities |
has_audio |
UI shows |
|---|---|---|
["1080p", "720p", …] |
true |
Best / heights / Audio only |
["1080p", …] |
false |
Best / heights |
[] |
true |
Audio only (pre-selected) |
[] |
false |
"No media detected" |
| probe fails (502) | — | picker hidden; capture still submittable |
The POST /api/archives/:id/captures endpoint accepts an optional quality field:
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }
"audio" selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting quality or passing "best" downloads at the highest available quality.
Text notes
Not every capture has a URL. Add text in the capture dialog takes a title and a body and turns them into a self-contained entry — useful for a scrap of prose, a quote, or a note attached to the surrounding archive. The body is stored verbatim; no network fetch happens.
Two body types are accepted: text/markdown (saved as .md) and text/plain (saved as .txt). Anything else is
rejected. The body lands in store/raw/… under its SHA3-256 content hash, exactly like every other capture, so an
identical note captured twice is stored once.
Configuration
TOML config file
# Optional. Default: 127.0.0.1:8080
bind = "127.0.0.1:8080"
# Required. Persists across upgrades; must be on a writable path.
auth_db_path = "/var/lib/archivr/archivr-auth.sqlite"
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/srv/archivr/personal/.archivr"
[[archives]]
id = "work"
label = "Work"
archive_path = "/srv/archivr/work/.archivr"
See docker/config.example.toml for a complete annotated example.
Environment variables
| Variable | Default | Description |
|---|---|---|
ARCHIVR_BIND |
127.0.0.1:8080 |
Bind address; overrides bind in TOML |
ARCHIVR_STATIC_DIR |
crates/archivr-server/static |
Pre-built frontend asset directory |
ARCHIVR_YT_DLP |
yt-dlp |
yt-dlp binary used for video and social downloads; the Nix wrappers point this at the pinned release |
ARCHIVR_YT_DLP_FORCE |
— | Absolute path to a yt-dlp binary that MUST be used, bypassing the resolver. Prefer ARCHIVR_YT_DLP unless you are overriding for a specific run |
ARCHIVR_SINGLE_FILE |
single-file |
single-file-cli binary for web page archiving |
ARCHIVR_CHROME |
chromium |
Chromium executable passed to single-file |
ARCHIVR_CHROME_ARGS |
— | Extra space-separated Chromium flags (Docker sets --no-sandbox) |
ARCHIVR_TWITTER_CREDENTIALS_FILE |
— | Cookies file for tweet/thread scraping — required for tweet:ID and x:thread:ID inputs |
ARCHIVR_TWEET_SCRAPER |
vendor/twitter/scrape_user_tweet_contents.py |
Tweet scraper script path |
ARCHIVR_TWEET_PYTHON |
python3 |
Python executable for the tweet scraper |
The Nix wrapper and Docker image set ARCHIVR_STATIC_DIR, ARCHIVR_SINGLE_FILE, and ARCHIVR_CHROME automatically.
LLM providers
Summaries are opt-in and provider-agnostic. Only the variables for the provider you actually select are read; the two HTTP providers refuse to start without their API key.
| Variable | Default | Description |
|---|---|---|
ARCHIVR_ANTHROPIC_API_KEY |
(required for anthropic_http) |
API key for the Anthropic Messages API |
ARCHIVR_ANTHROPIC_URL |
https://api.anthropic.com/v1/messages |
Endpoint override, e.g. an internal proxy |
ARCHIVR_ANTHROPIC_MODEL |
claude-3-5-sonnet-latest |
Model id used for Anthropic summaries |
ARCHIVR_OPENAI_API_KEY |
(required for openai_compatible) |
API key for any OpenAI-compatible endpoint |
ARCHIVR_OPENAI_URL |
https://api.openai.com/v1/chat/completions |
Endpoint override; point this at a local server to run offline |
ARCHIVR_OPENAI_MODEL |
gpt-4o-mini |
Model id used for OpenAI-compatible summaries |
ARCHIVR_CLAUDE_CLI |
(auto-discovered) | Path to a local claude binary |
ARCHIVR_CLAUDE_MODEL |
(the CLI's own default) | Optional model override for the local Claude CLI |
ARCHIVR_CODEX_CLI |
(auto-discovered) | Path to a local codex binary |
ARCHIVR_CODEX_MODEL |
(the CLI's own default) | Optional model override for the local Codex CLI |
ARCHIVR_SUMMARY_HTTP_TIMEOUT |
120 |
Seconds before an HTTP-provider summary is killed |
ARCHIVR_SUMMARY_CLI_TIMEOUT |
300 |
Seconds before a CLI-provider summary is killed |
When ARCHIVR_CLAUDE_CLI / ARCHIVR_CODEX_CLI is unset the binary is auto-discovered, in this order: the well-known
absolute paths, then $HOME/.local/bin/<name>, then the bare name resolved through PATH. Note that PATH is
consulted last — if a stale binary sits at one of the well-known paths it wins over a newer one on PATH, so set
the variable explicitly when you have both. The well-known paths are /opt/homebrew/bin/claude and
/usr/local/bin/claude for Claude, and /Applications/ChatGPT.app/Contents/Resources/codex,
/opt/homebrew/bin/codex, and /usr/local/bin/codex for Codex.
Keeping yt-dlp fresh
yt-dlp is the download engine behind every video and social capture. YouTube rotates its player-signature and API surfaces on a days-to-weeks cadence, so a binary that worked last month starts returning HTTP 403 on downloads. Keeping it current is ordinary maintenance, not an emergency.
What ships. flake.nix pins a specific yt-dlp release fetched straight from github.com/yt-dlp/yt-dlp/releases,
not from nixpkgs — that channel usually lags months behind. The archivr-server and archivr wrappers set
ARCHIVR_YT_DLP to that pinned binary.
How the resolver picks. At runtime archivr probes --version on each candidate — the pinned binary from
ARCHIVR_YT_DLP and any user-installed binary at <state_dir>/yt-dlp/yt-dlp — and runs the newest. An exact version
tie resolves in favour of your own install. Setting ARCHIVR_YT_DLP_FORCE=/path/to/yt-dlp bypasses the comparison
entirely. The state dir is ~/Library/Application Support/archivr on macOS, and $XDG_STATE_HOME/archivr (default
~/.local/state/archivr) elsewhere.
There are three ways to get a fresh version, cheapest first.
1. Self-update — no rebuild required.
archivr yt-dlp status # every candidate, its version, and which one wins
archivr yt-dlp update # download the latest zipapp into the state dir
archivr yt-dlp update --version 2026.09.15 # pin a specific release tag
The released artifact is a Python zipapp, so this path needs python3 on PATH at run time.
2. Automatic weekly bump. .github/workflows/update-ytdlp.yml runs every Monday at 06:00 UTC, queries GitHub for
the latest release, and opens a PR bumping version and hash in flake.nix via peter-evans/create-pull-request.
It also accepts workflow_dispatch for an on-demand run.
3. Manual bump, when you need it now and do not want to wait for the weekly:
NEW=$(curl -s https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest | jq -r .tag_name)
HASH=$(nix hash file --sri --type sha256 <(curl -sL "https://github.com/yt-dlp/yt-dlp/releases/download/${NEW}/yt-dlp"))
# In flake.nix, inside the `ytDlp = pkgs.stdenv.mkDerivation { … }` block:
# version = "OLD"; → version = "$NEW";
# url = ".../download/OLD/yt-dlp"; → .../download/$NEW/yt-dlp
# hash = "sha256-OLD…"; → hash = "$HASH";
nix build .#archivr-server
./result/bin/archivr yt-dlp status # the env row should report the new version
git commit -am "chore(nix): yt-dlp OLD → $NEW"
Deployment
Security
archivr-server binds to 127.0.0.1:8080 by default. Do not expose it to a public network without understanding the risks. When started on a non-loopback address the server logs a warning to stderr.
Hosting on NixOS
The flake exposes nixosModules.default:
# flake.nix (your system flake)
{
inputs.archivr.url = "github:thegeneralist/archivr";
outputs = { nixpkgs, archivr, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
archivr.nixosModules.default
{
services.archivr-server = {
enable = true;
# listenAddress defaults to "127.0.0.1"
# port defaults to 8080
archives = [
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
];
};
}
];
};
};
}
The module creates an archivr system user and group, generates the TOML config from your options, stores the auth database at /var/lib/archivr-server/ (persists across upgrades), and runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp). Archive directories are whitelisted for read-write access.
Set openFirewall = true with a non-loopback listenAddress only when LAN or remote access is required.
Archive directories must be owned by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.
Hosting with Docker
# 1. Configure
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit archivr-server.toml — set id, label, archive_path, and auth_db_path
# 2. Initialize each archive (run once per archive)
docker compose run --rm archivr archivr init \
/data/archives/main /data/archives/main/.archivr/store \
--name "Main Archive"
# 3. Start
docker compose up -d
| Mount | Purpose |
|---|---|
./config (read-only) |
Directory containing archivr-server.toml |
archivr-data named volume |
Auth database (/data/archivr-auth.sqlite) and archive directories |
Important:
auth_db_pathmust point to a path on the writable data volume (e.g./data/archivr-auth.sqlite). The example config sets this correctly. A baremkdiris not enough to initialise an archive —archivr initwrites metadata files the server requires.
Twitter/X archiving: supply a cookies file inside the config volume and reference it in docker-compose.yml:
environment:
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
Building locally:
docker build -t archivr-server .
The image compiles the Rust binary in a separate build stage; only runtime dependencies (Chromium, Node.js, Python) land in the final layer.
Development
Runtime dependencies beyond Rust and Node: yt-dlp, Chromium, single-file (Node), Python 3 with twitter-api-client, ffmpeg. nix develop provides the dev subset.
Entry summaries are served by one of four interchangeable providers — anthropic_http, openai_compatible,
claude_cli, or codex_cli — each configured entirely through the environment; see
LLM providers for the full variable list. The archivr CLI itself exposes archive, init, and
yt-dlp status / yt-dlp update; summaries are triggered from the web UI rather than the command line.
# Rust (workspace root)
cargo build
cargo test
cargo test -p archivr-core
cargo run -p archivr-server -- ./archivr-server.toml
# Frontend (from frontend/)
bun install
bun run dev # Vite dev server
bun run build # → crates/archivr-server/static/
bun run storybook # Component QA on :6006
# Nix
nix develop # dev shell
nix build .#archivr-server
License
MIT — see LICENSE. \n