* ui: show spinner for pending captures
* feat(core): add text capture path with title + Markdown/plain body
- Add downloader/text.rs module with save() function that stages and hashes text content
- Support text/markdown and text/plain MIME types with .md and .txt extensions
- Add perform_text_capture() function for capturing user-supplied text
- Validates title (non-empty, max 500 chars) and body (non-empty, max 2 MiB)
- Creates blob records and entries with source_kind='text', entity_kind='document'
- Includes comprehensive unit tests for markdown, plain text, and validation
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(server): add POST /api/archives/:archive_id/captures/text
- Add CaptureTextBody struct for title, body, and optional MIME type
- Implement capture_text_handler with validation for empty fields and MIME type
- Route text submissions to perform_text_capture() in background
- Reuse existing capture job tracking and polling infrastructure
- Default MIME type to text/markdown when not specified
- Include route tests covering happy path, validation, auth, and error cases
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(frontend): add text-capture form to CaptureDialog
- Add submitTextCapture API client function with same error handling as submitCapture
- Create makeTextItem() factory for text capture state
- Implement CaptureTextRow component with title, body textarea, and MIME selector
- Add 'Add text' button in capture dialog toolbar
- Update handleArchive to filter and route text submissions
- Modify submitBgJob to detect and submit text items via submitTextCapture
- Skip probe and conflict checks for text items
- Reuse job tracking and batch settlement for text captures
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(core): add entry_summaries schema + summarizer trait/providers
Per-entry LLM summaries as a regenerable child record, not a column on
archived_entries and not an on-disk artifact: an entry may carry several
summaries (one per provider/model/prompt version), any of which can be
discarded and recomputed. Generation is manual-only — nothing in capture.rs
calls into this module.
- database.rs: entry_summaries table + index, EntrySummaryRecord, and
upsert/update/find/latest helpers mirroring the capture_jobs style.
provider_model is stored as '' rather than NULL because SQLite treats
NULLs as distinct inside a UNIQUE index, which would stop the CLI
providers (no model) from ever deduping on the cache key.
- summarizer.rs: SummaryProvider trait with four implementations —
Anthropic Messages API, OpenAI-compatible chat completions, `claude -p`
and `codex exec -`. Configuration comes from env vars only (never TOML),
matching how yt-dlp / single-file / tweet-scraper are resolved, which
also keeps API keys out of anything the archive persists.
- archive.rs: EntryDetail gains latest_summary, populated by one extra
LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row
rather than duplicating it.
Implementation notes:
- No tokio in core. CLI timeouts are enforced structurally: stdout is
drained on its own thread and handed back over a channel so the calling
thread can recv_timeout and kill an overrunning child; stdin is written
on a third thread so a 48 KB prompt cannot deadlock against a child
waiting for us to read.
- HTML is reduced with regex rather than a parser: html5ever is not in the
tree, and a model tolerates imperfect whitespace. Paired tags are spelled
out per tag because Rust's regex engine has no backreferences by design.
- reqwest is declared with only the `blocking` feature here, so bodies are
serialized via .body(value.to_string()) instead of widening the
workspace dependency for .json().
- input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one
hashing primitive; the content is truncated to 48 KB *before* hashing so
the cache key describes exactly the bytes the model saw.
Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for
one JSON shape is a poor trade, so the two halves that can actually break
are tested directly — request-body builders and response parsers — leaving
only reqwest's own transport uncovered. Plus schema idempotency, cache-key
dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML
and tweet extraction, output normalization, and the CLI runner's stdin
round-trip, timeout kill, and nonzero-exit paths.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(server): add GET/POST /api/archives/:id/entries/:uid/summary
GET is read-only and gated exactly like entry detail, so a guest can read a
summary only for an entry whose content they could already read. POST
requires ROLE_USER, matching capture / tags / patch / rearchive; no auth
roles change.
Both the provider config and the content extraction resolve on the request
thread, before spawn_blocking. That is what lets a missing env var come back
as a synchronous 400 naming the exact variable, and an unsummarizable
artifact (video, audio) as a 400 saying so, rather than becoming a
background job the caller must poll only to learn about a config typo.
The pending row is claimed before spawning so the 202 can name a summary_uid
the client can poll immediately. summarize_entry owns the
pending → running → completed/failed transitions for that same row — the
cache key is identical, so both upserts resolve to one row — leaving the
handler to catch only the case where it fails before recording anything.
When !force and an identical cache key already completed, the existing row
comes back as a 200 with no new work.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(frontend): render Summary rail section + provider selector
New "Summary" rail section between the URL/Preview controls and .meta-list.
A completed summary renders as bold tl;dr, body paragraph, tag chips, and a
provider · model footer; missing or failed shows Generate; pending/running
shows an inline spinner and polls GET every 1500 ms until terminal.
- api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps
ApiError's { "error": ... } body so the missing-env-var message reaches
the user verbatim rather than as a bare status code.
- ContextRail.jsx: state seeds from detail.latest_summary so the section
renders immediately on selection. Polling is anchored on the summary
status rather than started inside the click handler, so a job still
running when the user navigates away and back is picked up again. A
transient poll failure is swallowed — the next tick retries, and a real
failure arrives as status === 'failed'.
- Regenerate passes force:true only when a completed summary is already
shown; otherwise the request can take the server's 200 cache-hit path.
- Provider choice persists in sessionStorage under archivr:summary:provider,
with try/catch around both accessors for private-mode browsers.
- Public sessions never see the selector or the Generate button, and the
section renders at all only when a completed summary made it through the
server's visibility gate.
- styles.css: .rail-summary-* only; spacing and the action button reuse
.rail-section and .rail-rearchive-btn. The spinner honours
prefers-reduced-motion — the text alone conveys the state.
- AGENTS.md: document the summary env vars alongside the existing
external-tool convention.
Smoke-tested end to end against a scratch archive with a seeded markdown
entry: claude_cli produced a real summary (pending → running → completed in
~11s); a local mock server exercised the openai_compatible transport and
confirmed the Bearer header, model, and system/user role split on the wire;
unconfigured providers return 400 naming the exact variable; a video entry
returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without
adding a row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): codex_cli — auto-discover binary + use --output-last-message
Two related fixes for the codex_cli summary provider:
1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but
without it the code resolved to bare `codex` and relied on PATH.
The ChatGPT desktop app installs codex at
`/Applications/ChatGPT.app/Contents/Resources/codex` and does not
put it on PATH, so users who only have the desktop app saw
'No such file or directory' with no hint. `resolve_cli` now walks
env override → a small set of well-known absolute paths → HOME
/.local/bin/<bare> → bare fallback. Same treatment applied to
claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/
claude, HOME/.local/bin/claude).
2. Clean output. `codex exec -` writes a runtime header ("OpenAI
Codex vX", session id, sandbox, model), the assistant reply, and a
footer ("tokens used", replay of the reply) to stdout. The JSON
extractor took the first '{' from the *user prompt echo* and the
last '}' from the trailing replay, producing invalid text that
fell through to the "raw text under summary" fallback path. Now
uses `--output-last-message <tempfile>` and reads only the final
assistant message. Fallback (positional prompt) uses the same
flag. Tempfile is cleaned up on all paths, incl. spawn failure.
* fix(frontend): give text-capture row real CSS
The text row shipped with semantic classnames (`capture-text-inputs`,
`capture-text-title`, `capture-text-body`, `capture-text-mime`,
`capture-text-icon`) but no CSS rules. Falling through to the parent
`.capture-row-main` flex-row (`display: flex; align-items: center`)
meant the title, textarea, and mime-select stacked as intrinsic-width
boxes centered on the tall body, producing a layout where the body
floated to the top-right, the title box appeared BELOW it, and the mime
selector rendered as an unstyled OS dropdown.
Fix:
- `.capture-text-row .capture-row-main` uses `align-items: flex-start`
so the leading icon and trailing × pin to the top of the block.
- `.capture-text-inputs` is now a full-width column-flex container with
proper gaps.
- `.capture-text-title` reuses the 44px input height and typography of
`.capture-input`; `.capture-text-body` gets a 140px min-height,
vertical resize, and matching border/focus treatment.
- `.capture-text-mime` is styled as a small chip with a custom caret
so it matches `.capture-quality` and stops looking like a raw
`<select>`. Sits in a right-aligned footer under the body.
- `.capture-text-icon` gets a 44px column so it aligns with the title
input; remove button gets a small top-margin for the same reason.
Rebuilt static bundle bumped as well (`index-BLxoi9rt.css`,
`index-CQcpPA_I.js`).
* chore(static): rebuild bundle after text-row CSS merge
* fix(core): summarize tweets + walk all tweets in a thread
Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.
Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.
Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
`raw_tweet_json` artifacts (falling back to `primary_media` for
archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
markdown/html/json branches, then join thread pieces with a
`---` separator so the model sees a real paragraph break between
statuses instead of one flowing document.
Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.
* feat(frontend): preview text-capture entries (.md / .txt)
Text captures land as `.md` (Markdown) or `.txt` (plain) blobs, but
PreviewPanel only dispatched on video/audio/image/pdf/html extensions,
so opening a text entry hit the 'No preview available' fallback with
the raw artifact path exposed.
- New `TextPreview` component fetches the primary artifact as text,
renders it in a monospace `<pre>` with word-wrap, and shows the
entry title on top and the MIME as a small trailing tag. Handles
loading/error states.
- `PreviewPanel` gains a `TEXT_EXTS` set + a branch that dispatches
to `TextPreview` for `md` / `markdown` / `txt`.
- CSS is padded and centered to ~780px so a text note reads like a
document rather than an edge-to-edge terminal dump.
v1 intentionally does NOT parse Markdown: keeping frontend deps at
react+react-dom only. Bump to a real Markdown renderer if we start
capturing Markdown-authored notes.
* fix(nix): pin yt-dlp from its own release + wire into server wrapper
Two independent problems, one commit:
1. Stale binary. nixpkgs-provided `pkgs.yt-dlp` on the pinned
nixos-unstable rev is 2026.03.17 (Mar 2026). yt-dlp itself
releases days-to-weeks, and YouTube frequently rotates the
player-signature / client surfaces the older builds request
(`android_vr` is the current casualty), which returns HTTP 403
mid-download for the format specs archivr passes (`-f
bestvideo+bestaudio/best`). Even bumping the nixpkgs input would
leave us dependent on that channel's yt-dlp cadence.
Fetch the upstream zipapp directly instead
(github.com/yt-dlp/yt-dlp/releases/download/<ver>/yt-dlp), wrap so
`python3` and `ffmpeg` are on PATH, and pin version+hash in one
place. Bumping is: change version, replace hash from
`nix hash file <url>`.
2. Missing pin in server wrapper. `archivr-cli` was already wrapped
with `--set ARCHIVR_YT_DLP` + a PATH prefix; `archivr-server`
was NOT — it only pinned single-file, chrome, and the tweet
scraper, silently falling back to whatever `yt-dlp` the user
happened to have on PATH. Server captures therefore inherited
the user's (often stale) system yt-dlp regardless of the flake
pin. Same wrapper flags now apply to both binaries.
devShell keeps `pkgs.yt-dlp` for now: the dev shell is a
convenience, not a release surface, and matching wouldn't fit in this
commit without duplicating the derivation across let-scopes.
* chore(static): rebuild bundle for round-3 fixes
* chore(nix): pin python 3.12 for yt-dlp zipapp (avoid py3.14 libffi crash on darwin/arm64)
* feat(core): resolve_yt_dlp picks the newer of pinned vs state-dir
The nix flake wrapper pins a yt-dlp via ARCHIVR_YT_DLP, but yt-dlp rots
fast — extractors break within weeks of a pin. Add a resolver that probes
`--version` on both the pinned binary and a user-installed copy under the
mutable state dir, and runs whichever is newer.
Version strings are YYYY.MM.DD, so plain string ordering is chronological.
Ties resolve toward the state dir: a user who installed it there did so
deliberately. ARCHIVR_YT_DLP_FORCE bypasses the comparison entirely, and
with no candidate at all we fall back to bare `yt-dlp` on PATH — exactly
the previous behaviour.
Resolution is cached in a OnceLock so `--version` costs one subprocess per
process, and all four inline env::var lookups now go through it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: bump yt-dlp from upstream releases, not nixpkgs
The flake no longer takes yt-dlp from nixpkgs; a dedicated `ytDlp`
derivation fetches the upstream release binary directly and pins both
`version` and an SRI `hash`. That makes the previous workflow inert: it
ran `nix flake update nixpkgs` and compared `nixpkgs#yt-dlp.version`
before and after, so it could churn the lockfile forever without ever
moving the version we actually ship.
The workflow now reads the pinned version straight out of the `ytDlp`
block in flake.nix, asks the GitHub API for yt-dlp's latest release tag,
short-circuits when they already match, downloads the new release to
recompute its SRI hash (required — the hash is part of the derivation's
identity, so the URL cannot be changed alone), and rewrites the three
pinned fields under a sed range address scoped to that block so sibling
pins like ublockLite and isdcac are untouched. It asserts only flake.nix
changed and that the new version appears exactly twice before opening
the PR.
* feat(cli): add `archivr yt-dlp update|status` subcommand
`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.
The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.
`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.
reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(maintainer): document summarizer, text capture, and yt-dlp lifecycle
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(readme): document LLM summaries, text notes, yt-dlp resolver + bump paths
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(plan): specify X Article, image summaries, summary search
* fix: summarize X Article text
* feat: search completed summary tags
* fix: preserve X Article block order
* feat: model opt-in summary images
* feat: attach opted-in images to summaries
* feat: accept image summary requests
* feat: add summary image consent control
* docs: explain X Article, image summaries and summary search
* chore(static): rebuild bundle for summary image consent
* fix: show civilized unsupported summary errors
* chore(static): rebuild bundle for civilized summary errors
* fix: keep text previews and summary state scoped to entry
* fix: show forced yt-dlp candidate in status
* fix: infer trusted MIME for tweet images
* fix: preserve text capture bytes and hide synthetic URL
* fix: recover and preserve summary attempts
* fix: bound codex fallback and record resolved model
* fix: protect public summary diagnostics
* fix: guard summary callbacks during entry render
* fix: scope summary callbacks to selected entry
* test: cover terminal newline in text capture
* fix: preserve missing summary entry status
* test: cover tweet image summary selection
* docs: record summary lifecycle and review hardening
* chore(static): rebuild bundle for Sol review fixes
* fix: retain completed summary during regeneration display
* fix: hide superseded summary attempts
* chore(static): rebuild bundle after regeneration display fix
* fix: allow full-size text capture requests
* fix: allow escaped full-size text captures
* fix: preserve text draft whitespace in capture UI
* chore(static): rebuild bundle for text whitespace fix
20 KiB
Archivr is a self-hosted tool for capturing and preserving digital content — YouTube videos and playlists, tweets and threads, Instagram, TikTok, web pages, and local files — into self-contained, locally-owned archives. Content is stored in SQLite with SHA3-256 blob deduplication, hierarchical tags, a browser-based UI, and role-based auth.
Table of Contents
- Features
- Quick Start
- Architecture
- Supported Inputs
- Configuration
- Keeping yt-dlp fresh
- Deployment
- Development
- License
Features
- Social media — YouTube (videos, shorts, playlists, channels with sync mode), X/Twitter (tweet and thread JSON + media downloads), Instagram, TikTok, Facebook, Reddit, Snapchat via yt-dlp
- Web pages — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
- Local files — import any file from disk by
file://path - Deduplication — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
- Tags and search — hierarchical tag tree, full-text search (including the latest completed summary and its generated JSON tags), filterable entry list
- Multiple archives — the server mounts any number of separate archives from a single TOML config
- Role-based auth — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
- Quality selection — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
- LLM summaries — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local
claudeCLI, or a localcodexCLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicitInclude attached imagesoption - Text notes — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview
- In-progress capture indicator — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
Quick Start
With Nix
# Create an archive
nix run github:thegeneralist/archivr#archivr -- init ./my-archive --name "My Archive"
# Archive something
nix run github:thegeneralist/archivr#archivr -- archive https://www.youtube.com/watch?v=dQw4w9WgXcQ
# Start the web UI (reads ./archivr-server.toml)
nix run github:thegeneralist/archivr#archivr-server
Create archivr-server.toml next to where you run the command:
auth_db_path = "/absolute/path/to/archivr-auth.sqlite"
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
Then open http://127.0.0.1:8080. On the first visit you will be prompted to create the owner account.
With Docker
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit config/archivr-server.toml
# Initialize the archive on the persistent volume (run once)
docker compose run --rm archivr archivr init \
/data/archives/main /data/archives/main/.archivr/store \
--name "Main Archive"
docker compose up -d
Open http://localhost:8080. See Hosting with Docker for volume layout and Twitter/X credential setup.
Architecture
Two binaries:
| Binary | Purpose |
|---|---|
archivr |
CLI — create archives (init) and add content (archive) |
archivr-server |
Web server — browse and search one or more archives via browser UI |
Archive layout created by archivr init:
my-archive/
├── .archivr/ # metadata: name, store_path, archivr.sqlite
└── store/
├── raw/ # deduplicated blobs: raw/A/B/<sha3-256>.ext
├── raw_tweets/ # tweet and thread JSON
├── structured/ # structured metadata outputs
└── temp/ # staging area during capture
A separate auth database (archivr-auth.sqlite, path set in TOML) holds users, sessions, API tokens, and role bits. It is independent of individual archives.
Supported Inputs
archivr archive <locator> accepts URLs and platform shorthands:
| Platform | Input examples |
|---|---|
| Local file | file:///absolute/path/to/file.pdf |
| YouTube video / short | https://youtube.com/watch?v=ID · yt:video/ID · yt:short/ID |
| YouTube playlist | https://youtube.com/playlist?list=ID |
| YouTube channel | https://youtube.com/@handle |
| X/Twitter tweet (JSON) | tweet:ID · x:tweet:ID · twitter:tweet:ID |
| X/Twitter thread (JSON) | x:thread:ID · twitter:thread:ID |
| X/Twitter media download | tweet:media:ID |
Direct URL · instagram:ID |
|
| TikTok | Direct URL · tiktok:ID |
Direct URL · facebook:ID |
|
Direct URL · reddit:ID |
|
| Snapchat | Direct URL · snapchat:ID |
| Arbitrary URL / web page | Any https:// URL |
YouTube playlists and channels
Capturing a playlist or channel creates a container entry with each video archived as a child beneath it. Before downloading, the UI probes each video for available quality options — set quality per-video or apply one to the whole batch. Individual videos can be excluded with the remove button.
Sync mode: when re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container is reused.
Video quality and audio-only
When capturing a yt-dlp-backed source through the web UI, a metadata probe runs first and populates the quality selector with heights actually available in that video:
qualities |
has_audio |
UI shows |
|---|---|---|
["1080p", "720p", …] |
true |
Best / heights / Audio only |
["1080p", …] |
false |
Best / heights |
[] |
true |
Audio only (pre-selected) |
[] |
false |
"No media detected" |
| probe fails (502) | — | picker hidden; capture still submittable |
The POST /api/archives/:id/captures endpoint accepts an optional quality field:
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }
"audio" selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting quality or passing "best" downloads at the highest available quality.
Text notes
Not every capture has a URL. Add text in the capture dialog takes a title and a body and turns them into a self-contained entry — useful for a scrap of prose, a quote, or a note attached to the surrounding archive. The body is stored verbatim; no network fetch happens.
Two body types are accepted: text/markdown (saved as .md) and text/plain (saved as .txt). Anything else is
rejected. The body lands in store/raw/… under its SHA3-256 content hash, exactly like every other capture, so an
identical note captured twice is stored once.
Text notes have no synthetic source URL: the original-URL field stays empty rather than inventing a text: locator.
Configuration
TOML config file
# Optional. Default: 127.0.0.1:8080
bind = "127.0.0.1:8080"
# Required. Persists across upgrades; must be on a writable path.
auth_db_path = "/var/lib/archivr/archivr-auth.sqlite"
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/srv/archivr/personal/.archivr"
[[archives]]
id = "work"
label = "Work"
archive_path = "/srv/archivr/work/.archivr"
See docker/config.example.toml for a complete annotated example.
Environment variables
| Variable | Default | Description |
|---|---|---|
ARCHIVR_BIND |
127.0.0.1:8080 |
Bind address; overrides bind in TOML |
ARCHIVR_STATIC_DIR |
crates/archivr-server/static |
Pre-built frontend asset directory |
ARCHIVR_YT_DLP |
yt-dlp |
yt-dlp binary used for video and social downloads; the Nix wrappers point this at the pinned release |
ARCHIVR_YT_DLP_FORCE |
— | Absolute path to a yt-dlp binary that MUST be used, bypassing the resolver. Prefer ARCHIVR_YT_DLP unless you are overriding for a specific run |
ARCHIVR_SINGLE_FILE |
single-file |
single-file-cli binary for web page archiving |
ARCHIVR_CHROME |
chromium |
Chromium executable passed to single-file |
ARCHIVR_CHROME_ARGS |
— | Extra space-separated Chromium flags (Docker sets --no-sandbox) |
ARCHIVR_TWITTER_CREDENTIALS_FILE |
— | Cookies file for tweet/thread scraping — required for tweet:ID and x:thread:ID inputs |
ARCHIVR_TWEET_SCRAPER |
vendor/twitter/scrape_user_tweet_contents.py |
Tweet scraper script path |
ARCHIVR_TWEET_PYTHON |
python3 |
Python executable for the tweet scraper |
The Nix wrapper and Docker image set ARCHIVR_STATIC_DIR, ARCHIVR_SINGLE_FILE, and ARCHIVR_CHROME automatically.
LLM providers
Summaries are manual and provider-agnostic. Only the variables for the provider you actually select are read; the two
HTTP providers refuse to start without their API key. They are text-only by default. Selecting Include attached images
explicitly sends eligible archived image data to the chosen provider; it is never attached automatically.
| Provider | Attached images |
|---|---|
| Anthropic HTTP | Supported |
| OpenAI-compatible HTTP | Supported |
| Codex CLI | Supported |
| Claude CLI | Not supported |
Image inclusion considers only media artifacts with jpg, jpeg, png, webp, gif, or avif files. At most four
images are sent, each no larger than 5 MiB and no more than 12 MiB in total.
Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no summary, or only a pending or failed summary, get no summary-derived match.
Each request is cached under the provider and the requested model identifier. If a provider reports a more precise resolved model (for example, an alias's concrete version), the UI displays that resolved name as attribution without changing the cache identity.
Summary attempts move from pending to running and then to completed or failed. On server startup, interrupted
pending or running attempts are marked failed. Regenerating does not replace an earlier completed summary until the
replacement succeeds, and public readers receive completed content only—never pending state or diagnostic errors.
| Variable | Default | Description |
|---|---|---|
ARCHIVR_ANTHROPIC_API_KEY |
(required for anthropic_http) |
API key for the Anthropic Messages API |
ARCHIVR_ANTHROPIC_URL |
https://api.anthropic.com/v1/messages |
Endpoint override, e.g. an internal proxy |
ARCHIVR_ANTHROPIC_MODEL |
claude-3-5-sonnet-latest |
Model id used for Anthropic summaries |
ARCHIVR_OPENAI_API_KEY |
(required for openai_compatible) |
API key for any OpenAI-compatible endpoint |
ARCHIVR_OPENAI_URL |
https://api.openai.com/v1/chat/completions |
Endpoint override; point this at a local server to run offline |
ARCHIVR_OPENAI_MODEL |
gpt-4o-mini |
Model id used for OpenAI-compatible summaries |
ARCHIVR_CLAUDE_CLI |
(auto-discovered) | Path to a local claude binary |
ARCHIVR_CLAUDE_MODEL |
(the CLI's own default) | Optional model override for the local Claude CLI |
ARCHIVR_CODEX_CLI |
(auto-discovered) | Path to a local codex binary |
ARCHIVR_CODEX_MODEL |
(the CLI's own default) | Optional model override for the local Codex CLI |
ARCHIVR_SUMMARY_HTTP_TIMEOUT |
120 |
Seconds before an HTTP-provider summary is killed |
ARCHIVR_SUMMARY_CLI_TIMEOUT |
300 |
Seconds before a CLI-provider summary is killed |
When ARCHIVR_CLAUDE_CLI / ARCHIVR_CODEX_CLI is unset the binary is auto-discovered, in this order: the well-known
absolute paths, then $HOME/.local/bin/<name>, then the bare name resolved through PATH. Note that PATH is
consulted last — if a stale binary sits at one of the well-known paths it wins over a newer one on PATH, so set
the variable explicitly when you have both. The well-known paths are /opt/homebrew/bin/claude and
/usr/local/bin/claude for Claude, and /Applications/ChatGPT.app/Contents/Resources/codex,
/opt/homebrew/bin/codex, and /usr/local/bin/codex for Codex.
Keeping yt-dlp fresh
yt-dlp is the download engine behind every video and social capture. YouTube rotates its player-signature and API surfaces on a days-to-weeks cadence, so a binary that worked last month starts returning HTTP 403 on downloads. Keeping it current is ordinary maintenance, not an emergency.
What ships. flake.nix pins a specific yt-dlp release fetched straight from github.com/yt-dlp/yt-dlp/releases,
not from nixpkgs — that channel usually lags months behind. The archivr-server and archivr wrappers set
ARCHIVR_YT_DLP to that pinned binary.
How the resolver picks. At runtime archivr probes --version on each candidate — the pinned binary from
ARCHIVR_YT_DLP and any user-installed binary at <state_dir>/yt-dlp/yt-dlp — and runs the newest. An exact version
tie resolves in favour of your own install. Setting ARCHIVR_YT_DLP_FORCE=/path/to/yt-dlp bypasses the comparison
entirely. The state dir is ~/Library/Application Support/archivr on macOS, and $XDG_STATE_HOME/archivr (default
~/.local/state/archivr) elsewhere.
There are three ways to get a fresh version, cheapest first.
1. Self-update — no rebuild required.
archivr yt-dlp status # every candidate, its version, and which one wins
archivr yt-dlp update # download the latest zipapp into the state dir
archivr yt-dlp update --version 2026.09.15 # pin a specific release tag
When ARCHIVR_YT_DLP_FORCE applies, status shows that forced candidate and selects it as the winner.
The released artifact is a Python zipapp, so this path needs python3 on PATH at run time.
2. Automatic weekly bump. .github/workflows/update-ytdlp.yml runs every Monday at 06:00 UTC, queries GitHub for
the latest release, and opens a PR bumping version and hash in flake.nix via peter-evans/create-pull-request.
It also accepts workflow_dispatch for an on-demand run.
3. Manual bump, when you need it now and do not want to wait for the weekly:
NEW=$(curl -s https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest | jq -r .tag_name)
HASH=$(nix hash file --sri --type sha256 <(curl -sL "https://github.com/yt-dlp/yt-dlp/releases/download/${NEW}/yt-dlp"))
# In flake.nix, inside the `ytDlp = pkgs.stdenv.mkDerivation { … }` block:
# version = "OLD"; → version = "$NEW";
# url = ".../download/OLD/yt-dlp"; → .../download/$NEW/yt-dlp
# hash = "sha256-OLD…"; → hash = "$HASH";
nix build .#archivr-server
./result/bin/archivr yt-dlp status # the env row should report the new version
git commit -am "chore(nix): yt-dlp OLD → $NEW"
Deployment
Security
archivr-server binds to 127.0.0.1:8080 by default. Do not expose it to a public network without understanding the risks. When started on a non-loopback address the server logs a warning to stderr.
Hosting on NixOS
The flake exposes nixosModules.default:
# flake.nix (your system flake)
{
inputs.archivr.url = "github:thegeneralist/archivr";
outputs = { nixpkgs, archivr, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
archivr.nixosModules.default
{
services.archivr-server = {
enable = true;
# listenAddress defaults to "127.0.0.1"
# port defaults to 8080
archives = [
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
];
};
}
];
};
};
}
The module creates an archivr system user and group, generates the TOML config from your options, stores the auth database at /var/lib/archivr-server/ (persists across upgrades), and runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp). Archive directories are whitelisted for read-write access.
Set openFirewall = true with a non-loopback listenAddress only when LAN or remote access is required.
Archive directories must be owned by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.
Hosting with Docker
# 1. Configure
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# Edit archivr-server.toml — set id, label, archive_path, and auth_db_path
# 2. Initialize each archive (run once per archive)
docker compose run --rm archivr archivr init \
/data/archives/main /data/archives/main/.archivr/store \
--name "Main Archive"
# 3. Start
docker compose up -d
| Mount | Purpose |
|---|---|
./config (read-only) |
Directory containing archivr-server.toml |
archivr-data named volume |
Auth database (/data/archivr-auth.sqlite) and archive directories |
Important:
auth_db_pathmust point to a path on the writable data volume (e.g./data/archivr-auth.sqlite). The example config sets this correctly. A baremkdiris not enough to initialise an archive —archivr initwrites metadata files the server requires.
Twitter/X archiving: supply a cookies file inside the config volume and reference it in docker-compose.yml:
environment:
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
Building locally:
docker build -t archivr-server .
The image compiles the Rust binary in a separate build stage; only runtime dependencies (Chromium, Node.js, Python) land in the final layer.
Development
Runtime dependencies beyond Rust and Node: yt-dlp, Chromium, single-file (Node), Python 3 with twitter-api-client, ffmpeg. nix develop provides the dev subset.
Entry summaries are served by one of four interchangeable providers — anthropic_http, openai_compatible,
claude_cli, or codex_cli — each configured entirely through the environment; see
LLM providers for the full variable list. The archivr CLI itself exposes archive, init, and
yt-dlp status / yt-dlp update; summaries are triggered from the web UI rather than the command line.
# Rust (workspace root)
cargo build
cargo test
cargo test -p archivr-core
cargo run -p archivr-server -- ./archivr-server.toml
# Frontend (from frontend/)
bun install
bun run dev # Vite dev server
bun run build # → crates/archivr-server/static/
bun run storybook # Component QA on :6006
# Nix
nix develop # dev shell
nix build .#archivr-server
License
MIT — see LICENSE. \n