1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 21:03:17 +02:00
archivr/docs/README.md
archivr-qa ccb4d6f8a4
docs(readme): document LLM summaries, text notes, yt-dlp resolver + bump paths
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 20:02:29 +02:00

18 KiB

Archivr — Preserve what matters. Forever.

MIT License Rust 2024 Self-hosted


Archivr is a self-hosted tool for capturing and preserving digital content — YouTube videos and playlists, tweets and threads, Instagram, TikTok, web pages, and local files — into self-contained, locally-owned archives. Content is stored in SQLite with SHA3-256 blob deduplication, hierarchical tags, a browser-based UI, and role-based auth.

Table of Contents

Features

  • Social media — YouTube (videos, shorts, playlists, channels with sync mode), X/Twitter (tweet and thread JSON + media downloads), Instagram, TikTok, Facebook, Reddit, Snapchat via yt-dlp
  • Web pages — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
  • Local files — import any file from disk by file:// path
  • Deduplication — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
  • Tags and search — hierarchical tag tree, full-text search, filterable entry list
  • Multiple archives — the server mounts any number of separate archives from a single TOML config
  • Role-based auth — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
  • Quality selection — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
  • LLM summaries — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local claude CLI, or a local codex CLI; triggered manually from the entry rail, never automatically on capture
  • Text notes — capture a plain-text or Markdown note with a title and no URL; the note is stored as a normal deduplicated blob and previews in-browser
  • In-progress capture indicator — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block

Quick Start

With Nix

# Create an archive
nix run github:thegeneralist/archivr#archivr -- init ./my-archive --name "My Archive"

# Archive something
nix run github:thegeneralist/archivr#archivr -- archive https://www.youtube.com/watch?v=dQw4w9WgXcQ

# Start the web UI (reads ./archivr-server.toml)
nix run github:thegeneralist/archivr#archivr-server

Create archivr-server.toml next to where you run the command:

auth_db_path = "/absolute/path/to/archivr-auth.sqlite"

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"

Then open http://127.0.0.1:8080. On the first visit you will be prompted to create the owner account.

With Docker

mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit config/archivr-server.toml

# Initialize the archive on the persistent volume (run once)
docker compose run --rm archivr archivr init \
  /data/archives/main /data/archives/main/.archivr/store \
  --name "Main Archive"

docker compose up -d

Open http://localhost:8080. See Hosting with Docker for volume layout and Twitter/X credential setup.

Architecture

Two binaries:

Binary Purpose
archivr CLI — create archives (init) and add content (archive)
archivr-server Web server — browse and search one or more archives via browser UI

Archive layout created by archivr init:

my-archive/
├── .archivr/          # metadata: name, store_path, archivr.sqlite
└── store/
    ├── raw/           # deduplicated blobs: raw/A/B/<sha3-256>.ext
    ├── raw_tweets/    # tweet and thread JSON
    ├── structured/    # structured metadata outputs
    └── temp/          # staging area during capture

A separate auth database (archivr-auth.sqlite, path set in TOML) holds users, sessions, API tokens, and role bits. It is independent of individual archives.

Supported Inputs

archivr archive <locator> accepts URLs and platform shorthands:

Platform Input examples
Local file file:///absolute/path/to/file.pdf
YouTube video / short https://youtube.com/watch?v=ID · yt:video/ID · yt:short/ID
YouTube playlist https://youtube.com/playlist?list=ID
YouTube channel https://youtube.com/@handle
X/Twitter tweet (JSON) tweet:ID · x:tweet:ID · twitter:tweet:ID
X/Twitter thread (JSON) x:thread:ID · twitter:thread:ID
X/Twitter media download tweet:media:ID
Instagram Direct URL · instagram:ID
TikTok Direct URL · tiktok:ID
Facebook Direct URL · facebook:ID
Reddit Direct URL · reddit:ID
Snapchat Direct URL · snapchat:ID
Arbitrary URL / web page Any https:// URL

YouTube playlists and channels

Capturing a playlist or channel creates a container entry with each video archived as a child beneath it. Before downloading, the UI probes each video for available quality options — set quality per-video or apply one to the whole batch. Individual videos can be excluded with the remove button.

Sync mode: when re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container is reused.

Video quality and audio-only

When capturing a yt-dlp-backed source through the web UI, a metadata probe runs first and populates the quality selector with heights actually available in that video:

qualities has_audio UI shows
["1080p", "720p", …] true Best / heights / Audio only
["1080p", …] false Best / heights
[] true Audio only (pre-selected)
[] false "No media detected"
probe fails (502) — picker hidden; capture still submittable

The POST /api/archives/:id/captures endpoint accepts an optional quality field:

{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }

"audio" selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting quality or passing "best" downloads at the highest available quality.

Text notes

Not every capture has a URL. Add text in the capture dialog takes a title and a body and turns them into a self-contained entry — useful for a scrap of prose, a quote, or a note attached to the surrounding archive. The body is stored verbatim; no network fetch happens.

Two body types are accepted: text/markdown (saved as .md) and text/plain (saved as .txt). Anything else is rejected. The body lands in store/raw/… under its SHA3-256 content hash, exactly like every other capture, so an identical note captured twice is stored once.

Configuration

TOML config file

# Optional. Default: 127.0.0.1:8080
bind = "127.0.0.1:8080"

# Required. Persists across upgrades; must be on a writable path.
auth_db_path = "/var/lib/archivr/archivr-auth.sqlite"

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/srv/archivr/personal/.archivr"

[[archives]]
id = "work"
label = "Work"
archive_path = "/srv/archivr/work/.archivr"

See docker/config.example.toml for a complete annotated example.

Environment variables

Variable Default Description
ARCHIVR_BIND 127.0.0.1:8080 Bind address; overrides bind in TOML
ARCHIVR_STATIC_DIR crates/archivr-server/static Pre-built frontend asset directory
ARCHIVR_YT_DLP yt-dlp yt-dlp binary used for video and social downloads; the Nix wrappers point this at the pinned release
ARCHIVR_YT_DLP_FORCE — Absolute path to a yt-dlp binary that MUST be used, bypassing the resolver. Prefer ARCHIVR_YT_DLP unless you are overriding for a specific run
ARCHIVR_SINGLE_FILE single-file single-file-cli binary for web page archiving
ARCHIVR_CHROME chromium Chromium executable passed to single-file
ARCHIVR_CHROME_ARGS — Extra space-separated Chromium flags (Docker sets --no-sandbox)
ARCHIVR_TWITTER_CREDENTIALS_FILE — Cookies file for tweet/thread scraping — required for tweet:ID and x:thread:ID inputs
ARCHIVR_TWEET_SCRAPER vendor/twitter/scrape_user_tweet_contents.py Tweet scraper script path
ARCHIVR_TWEET_PYTHON python3 Python executable for the tweet scraper

The Nix wrapper and Docker image set ARCHIVR_STATIC_DIR, ARCHIVR_SINGLE_FILE, and ARCHIVR_CHROME automatically.

LLM providers

Summaries are opt-in and provider-agnostic. Only the variables for the provider you actually select are read; the two HTTP providers refuse to start without their API key.

Variable Default Description
ARCHIVR_ANTHROPIC_API_KEY (required for anthropic_http) API key for the Anthropic Messages API
ARCHIVR_ANTHROPIC_URL https://api.anthropic.com/v1/messages Endpoint override, e.g. an internal proxy
ARCHIVR_ANTHROPIC_MODEL claude-3-5-sonnet-latest Model id used for Anthropic summaries
ARCHIVR_OPENAI_API_KEY (required for openai_compatible) API key for any OpenAI-compatible endpoint
ARCHIVR_OPENAI_URL https://api.openai.com/v1/chat/completions Endpoint override; point this at a local server to run offline
ARCHIVR_OPENAI_MODEL gpt-4o-mini Model id used for OpenAI-compatible summaries
ARCHIVR_CLAUDE_CLI (auto-discovered) Path to a local claude binary
ARCHIVR_CLAUDE_MODEL (the CLI's own default) Optional model override for the local Claude CLI
ARCHIVR_CODEX_CLI (auto-discovered) Path to a local codex binary
ARCHIVR_CODEX_MODEL (the CLI's own default) Optional model override for the local Codex CLI
ARCHIVR_SUMMARY_HTTP_TIMEOUT 120 Seconds before an HTTP-provider summary is killed
ARCHIVR_SUMMARY_CLI_TIMEOUT 300 Seconds before a CLI-provider summary is killed

When ARCHIVR_CLAUDE_CLI / ARCHIVR_CODEX_CLI is unset the binary is auto-discovered, in this order: the well-known absolute paths, then $HOME/.local/bin/<name>, then the bare name resolved through PATH. Note that PATH is consulted last — if a stale binary sits at one of the well-known paths it wins over a newer one on PATH, so set the variable explicitly when you have both. The well-known paths are /opt/homebrew/bin/claude and /usr/local/bin/claude for Claude, and /Applications/ChatGPT.app/Contents/Resources/codex, /opt/homebrew/bin/codex, and /usr/local/bin/codex for Codex.

Keeping yt-dlp fresh

yt-dlp is the download engine behind every video and social capture. YouTube rotates its player-signature and API surfaces on a days-to-weeks cadence, so a binary that worked last month starts returning HTTP 403 on downloads. Keeping it current is ordinary maintenance, not an emergency.

What ships. flake.nix pins a specific yt-dlp release fetched straight from github.com/yt-dlp/yt-dlp/releases, not from nixpkgs — that channel usually lags months behind. The archivr-server and archivr wrappers set ARCHIVR_YT_DLP to that pinned binary.

How the resolver picks. At runtime archivr probes --version on each candidate — the pinned binary from ARCHIVR_YT_DLP and any user-installed binary at <state_dir>/yt-dlp/yt-dlp — and runs the newest. An exact version tie resolves in favour of your own install. Setting ARCHIVR_YT_DLP_FORCE=/path/to/yt-dlp bypasses the comparison entirely. The state dir is ~/Library/Application Support/archivr on macOS, and $XDG_STATE_HOME/archivr (default ~/.local/state/archivr) elsewhere.

There are three ways to get a fresh version, cheapest first.

1. Self-update — no rebuild required.

archivr yt-dlp status                        # every candidate, its version, and which one wins
archivr yt-dlp update                        # download the latest zipapp into the state dir
archivr yt-dlp update --version 2026.09.15   # pin a specific release tag

The released artifact is a Python zipapp, so this path needs python3 on PATH at run time.

2. Automatic weekly bump. .github/workflows/update-ytdlp.yml runs every Monday at 06:00 UTC, queries GitHub for the latest release, and opens a PR bumping version and hash in flake.nix via peter-evans/create-pull-request. It also accepts workflow_dispatch for an on-demand run.

3. Manual bump, when you need it now and do not want to wait for the weekly:

NEW=$(curl -s https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest | jq -r .tag_name)
HASH=$(nix hash file --sri --type sha256 <(curl -sL "https://github.com/yt-dlp/yt-dlp/releases/download/${NEW}/yt-dlp"))

# In flake.nix, inside the `ytDlp = pkgs.stdenv.mkDerivation { … }` block:
#   version = "OLD";                   → version = "$NEW";
#   url  = ".../download/OLD/yt-dlp";  → .../download/$NEW/yt-dlp
#   hash = "sha256-OLD…";              → hash = "$HASH";

nix build .#archivr-server
./result/bin/archivr yt-dlp status   # the env row should report the new version
git commit -am "chore(nix): yt-dlp OLD → $NEW"

Deployment

Security

archivr-server binds to 127.0.0.1:8080 by default. Do not expose it to a public network without understanding the risks. When started on a non-loopback address the server logs a warning to stderr.

Hosting on NixOS

The flake exposes nixosModules.default:

# flake.nix (your system flake)
{
  inputs.archivr.url = "github:thegeneralist/archivr";

  outputs = { nixpkgs, archivr, ... }: {
    nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
      modules = [
        archivr.nixosModules.default
        {
          services.archivr-server = {
            enable = true;
            # listenAddress defaults to "127.0.0.1"
            # port defaults to 8080
            archives = [
              { id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
              { id = "work";     label = "Work";     path = "/srv/archivr/work/.archivr"; }
            ];
          };
        }
      ];
    };
  };
}

The module creates an archivr system user and group, generates the TOML config from your options, stores the auth database at /var/lib/archivr-server/ (persists across upgrades), and runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp). Archive directories are whitelisted for read-write access.

Set openFirewall = true with a non-loopback listenAddress only when LAN or remote access is required.

Archive directories must be owned by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.

Hosting with Docker

# 1. Configure
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit archivr-server.toml — set id, label, archive_path, and auth_db_path

# 2. Initialize each archive (run once per archive)
docker compose run --rm archivr archivr init \
  /data/archives/main /data/archives/main/.archivr/store \
  --name "Main Archive"

# 3. Start
docker compose up -d
Mount Purpose
./config (read-only) Directory containing archivr-server.toml
archivr-data named volume Auth database (/data/archivr-auth.sqlite) and archive directories

Important: auth_db_path must point to a path on the writable data volume (e.g. /data/archivr-auth.sqlite). The example config sets this correctly. A bare mkdir is not enough to initialise an archive — archivr init writes metadata files the server requires.

Twitter/X archiving: supply a cookies file inside the config volume and reference it in docker-compose.yml:

environment:
  ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt

Building locally:

docker build -t archivr-server .

The image compiles the Rust binary in a separate build stage; only runtime dependencies (Chromium, Node.js, Python) land in the final layer.

Development

Runtime dependencies beyond Rust and Node: yt-dlp, Chromium, single-file (Node), Python 3 with twitter-api-client, ffmpeg. nix develop provides the dev subset.

Entry summaries are served by one of four interchangeable providers — anthropic_http, openai_compatible, claude_cli, or codex_cli — each configured entirely through the environment; see LLM providers for the full variable list. The archivr CLI itself exposes archive, init, and yt-dlp status / yt-dlp update; summaries are triggered from the web UI rather than the command line.

# Rust (workspace root)
cargo build
cargo test
cargo test -p archivr-core
cargo run -p archivr-server -- ./archivr-server.toml

# Frontend (from frontend/)
bun install
bun run dev        # Vite dev server
bun run build      # → crates/archivr-server/static/
bun run storybook  # Component QA on :6006

# Nix
nix develop        # dev shell
nix build .#archivr-server

License

MIT — see LICENSE. \n