1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-07-22 11:15:41 +02:00
archivr/docs/README.md

374 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# archivr
An open-source self-hosted archiving tool. Work in progress.
- [x] Archiving
- [x] Archiving media files from social media platforms
- [x] YouTube Videos
- [x] YouTube Playlists
- [x] YouTube Channels
- [x] Twitter Videos
- [x] Instagram
- [x] Facebook
- [x] TikTok
- [x] Reddit
- [x] Snapchat
- [ ] YouTube Posts (postponed)
- [x] Archiving local files
- [x] Archiving Twitter Tweets, Threads, and Articles
- [ ] Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
- [x] URLs
- [ ] Google Drive
- [ ] Dropbox
- [ ] OneDrive
- (Some of these could be postponed for later.)
- [x] Archive web pages (HTML, CSS, JS, images)
- [ ] Archiving emails (???)
- [ ] Gmail
- [ ] Outlook
- [ ] Yahoo Mail
- [x] Management
- [x] Deduplication
- [x] Tagging system
- [x] Search functionality
- [ ] Categorization
- [x] Metadata extraction and storage
- [x] User Interface
- [x] Web-based UI
- [x] Authentication and login
- [x] Archive setup
- [x] Browse and view entries
- [x] Tag management and filtering
- [x] Search entries
- [x] View archive runs
- [x] Capture dialog
- [x] User settings and API tokens
- [x] Admin panel
- [ ] Backup and Sync
- [ ] Cloud backup (AWS S3, Google Cloud Storage)
- [ ] Local backup
## Motivation
There are two driving factors behind this project:
- In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is _very important_ for preserving personal memories and digital history.
- I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.
This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.
## Archive Inputs
`archivr archive <path>` currently accepts three kinds of inputs:
- Local files via `file://...`
- Direct platform URLs
- Platform shorthand inputs such as `tweet:...`, `yt:...`, or `instagram:...`
## Running Archivr
Archivr currently ships as two binaries:
- `archivr`
- The CLI for creating and writing to one archive.
- Use this for `init` and `archive`.
- `archivr-server`
- The web server for reading one or more existing archives through the browser UI.
- Use this after archives already exist.
With Nix, run the CLI with:
```sh
nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
```
Run the web server with:
```sh
nix run .#archivr-server -- ./archivr-server.toml
```
The server expects a TOML registry file. If no path is passed, it reads `./archivr-server.toml`.
Example:
```toml
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
```
Then open:
```text
http://127.0.0.1:8080
```
When installed through Nix, `archivr-server` is wrapped so it can find the static web UI assets automatically. The wrapper sets `ARCHIVR_STATIC_DIR` to the installed static asset directory. Running from source with `cargo run -p archivr-server` falls back to `crates/archivr-server/static`.
### Security and Deployment
`archivr-server` is a **local-only tool by default**. It binds to `127.0.0.1:8080` and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.
**Changing the bind address**
You can set the bind address in your TOML config:
```toml
# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"
```
Or override it with the `ARCHIVR_BIND` environment variable:
```sh
ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml
```
If the server is started with a non-loopback address (e.g. `0.0.0.0`), it prints a warning to stderr:
```text
warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.
```
**When will auth be added?**
Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See `crates/archivr-server/src/routes.rs` for the route classification that will guide where middleware is applied.
### Supported Platforms
- Local files: `file:///absolute/path/to/file.ext`
- YouTube media: individual videos/shorts, playlists, and channels; standard URLs or [shorthand video inputs](#supported-shorthand-inputs). Playlists and channels archive as a container entry with each video stored as a child entry beneath it.
- X/Twitter media from Tweets: normal Tweet URLs or the `tweet:media:ID` shorthand
- X/Twitter Tweet content scrape: [Tweet and Thread shorthands](#supported-shorthand-inputs). (These are saved as JSON files in `raw_tweets/`)
- Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to `yt-dlp`
#### Video quality and audio-only downloads
When capturing via the web UI, entering a URL for a yt-dlp-backed source (YouTube, Instagram, TikTok, Facebook, Reddit, Snapchat, X media) triggers a metadata probe via `GET /api/archives/:id/captures/probe`. The quality selector then shows only the heights actually available in that video plus **Best quality** (default). An **Audio only** option is appended whenever the probe confirms an audio track exists. UI behaviour by probe outcome:
| `qualities` | `has_audio` | UI shows |
|---|---|---|
| `["1080p", "720p", …]` | `true` | Best / heights / Audio only |
| `["1080p", …]` | `false` | Best / heights |
| `[]` | `true` | Audio only (pre-selected, no Best option) |
| `[]` | `false` | "No media detected" |
| probe fails (502) | — | picker hidden, capture still submittable |
The `POST /api/archives/:id/captures` endpoint accepts an optional `quality` field: `"best"`, `"audio"`, or any `"NNNp"` height string:
```json
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }
```
`"audio"` selects the most efficient native audio track without transcoding: Opus/WebM is preferred (smallest at equivalent quality), then AAC/M4A, then whatever yt-dlp considers best. The saved file's extension matches the native format (`.webm` for Opus, `.m4a` for AAC, etc.) — no ffmpeg re-encode, no size inflation. Any `"NNNp"` height is accepted; the server builds the yt-dlp format selector with an unconditional `/best` fallback so the download succeeds even if the exact height is unavailable. Omitting `quality` or passing `"best"` downloads at the highest available quality. Anything else is rejected with HTTP 400.
The probe endpoint (`GET /api/archives/:id/captures/probe?locator=…`) requires auth and returns 200 with:
```json
{ "has_video": true, "has_audio": true, "qualities": ["1080p", "720p", "480p"] }
{ "has_video": false, "has_audio": true, "qualities": [] }
{ "has_video": false, "has_audio": false, "qualities": [] }
```
`has_video: false, has_audio: false` means yt-dlp found no downloadable tracks (e.g. a tweet with no media). A 502 means yt-dlp itself failed (transient network error, rate-limit, unsupported extractor) — treat as inconclusive, not "no media."
#### YouTube playlists and channels
Capturing a YouTube playlist or channel URL creates a **container entry** for the playlist or channel, with each video archived as a child entry beneath it. Before downloading, the capture UI probes each video to fetch available quality options, letting you set quality per-video or apply a single quality to the whole playlist.
**Incremental sync:** When re-archiving a playlist or channel, enable **sync mode** in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container entry is reused.
**Excluding individual videos:** In the expanded per-video list, each video has a remove button (×) to exclude it from the current capture. Removed videos are not downloaded; the rest proceed normally.
### Hosting on NixOS
The flake exposes a `nixosModules.default` output. Add it to your system flake and
enable the service:
```nix
# flake.nix (your system flake)
{
inputs.archivr.url = "github:thegeneralist/archivr";
outputs = { nixpkgs, archivr, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
archivr.nixosModules.default
{
services.archivr-server = {
enable = true;
# listenAddress defaults to "127.0.0.1" (loopback only)
# port defaults to 8080
archives = [
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
];
};
}
];
};
};
}
```
The module:
- Creates an `archivr` system user and group.
- Generates the TOML config from your options and stores the auth database under
`/var/lib/archivr-server/` (persists across upgrades).
- Runs under a hardened systemd unit (`ProtectSystem = strict`, `NoNewPrivileges`,
`PrivateTmp`, etc.). Archive directories are whitelisted for read-write access.
- Restarts automatically on failure.
**`openFirewall`** — set to `true` to open the TCP port derived from `bind`.
Only needed when binding to a non-loopback address:
```nix
services.archivr-server = {
listenAddress = "0.0.0.0";
port = 8080; # explicit, though 8080 is the default
openFirewall = true;
};
```
**Archive directories** must be readable and writable by the `archivr` user.
Initialise them with `archivr init` first, then `chown -R archivr:archivr /srv/archivr`.
### Hosting with Docker
A `Dockerfile` and `docker-compose.yml` are provided for self-hosting without Nix.
**Quickstart**
1. Copy the example config and edit it:
```sh
mkdir config
cp docker/config.example.toml config/archivr-server.toml
# edit config/archivr-server.toml — set archive id, label, and archive_path
```
2. Initialize each archive on the persistent data volume before the first start.
The image includes the `archivr` CLI for this purpose:
```sh
docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"
```
This creates `/data/archives/main/.archivr/` with the metadata the server requires.
A bare `mkdir` is not enough — the server reads `name` and `store_path` files that
only `archivr init` writes.
3. Start the server:
```sh
docker compose up -d
```
Then open `http://localhost:8080`.
**Volumes**
| Mount | Purpose |
|-------|---------|
| `./config` (read-only) | Directory containing `archivr-server.toml` |
| `archivr-data` named volume | Auth database (`/data/archivr-auth.sqlite`) and archive directories |
> **Important:** `auth_db_path` must be set explicitly in `archivr-server.toml` to a
> path on the writable data volume (e.g. `/data/archivr-auth.sqlite`). If left unset,
> the server defaults to writing the auth database next to the config file — which is
> on the read-only `/config` mount and will fail. The example config sets this correctly.
**Twitter/X archiving**
Supply a cookies file inside the config volume and set `ARCHIVR_TWITTER_CREDENTIALS_FILE` in `docker-compose.yml`:
```yaml
environment:
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
```
**Building the image locally**
```sh
docker build -t archivr-server .
```
The image compiles the Rust binary in a separate build stage so only the runtime
dependencies (Chromium, Node.js, Python) land in the final layer.
### Supported Shorthand Inputs
- YouTube video/short media:
- `yt:video/ID`
- `youtube:video/ID`
- `yt:short/ID`
- `yt:shorts/ID`
- `youtube:shorts/ID`
- X/Twitter tweet JSON content:
- `tweet:ID`
- `x:tweet:ID`
- `x:x:ID`
- `twitter:x:ID`
- `twitter:tweet:ID`
- X/Twitter media/video download:
- `tweet:media:ID`
- X/Twitter thread JSON content:
- `x:thread:ID`
- `twitter:thread:ID`
- Other platform shorthands:
- `instagram:ID`
- `facebook:ID`
- `tiktok:ID`
- `reddit:ID`
- `snapchat:ID`
### Environment Variables
- `ARCHIVR_BIND`
- Optional.
- Overrides the bind address from the TOML config. Useful in Docker where you need
`0.0.0.0:8080` without editing the config file. Default: `127.0.0.1:8080`.
- `ARCHIVR_STATIC_DIR`
- Optional.
- Path to the directory of pre-built frontend assets served by the web UI.
Set automatically by the Nix wrapper and the Docker image. When running from
source with `cargo run`, falls back to `crates/archivr-server/static`.
- `ARCHIVR_YT_DLP`
- Optional.
- Overrides the `yt-dlp` binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
- `ARCHIVR_SINGLE_FILE`
- Optional.
- Overrides the `single-file` binary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
- `ARCHIVR_CHROME`
- Optional.
- Overrides the Chromium/Chrome executable passed to `single-file` via `--browser-executable-path`. Set automatically by the Nix wrapper and the Docker image. Default: `chromium`.
- `ARCHIVR_CHROME_ARGS`
- Optional.
- Space-separated extra flags appended to Chromium's `--browser-args`. The Docker
image sets this to `--no-sandbox` because Chromium refuses to run as root without
it. Leave unset when running natively (Nix, Linux desktop).
A `--window-size=1920,1080` is always passed to provide a realistic desktop
viewport (so responsive @media rules and styles are evaluated and preserved
correctly). Supply your own `--window-size=...` here to override.
- `ARCHIVR_TWITTER_CREDENTIALS_FILE`
- Required for tweet/thread scraping inputs such as `tweet:ID` and `x:thread:ID`.
- Must point to a cookies file for the vendored scraper.
- `ARCHIVR_TWEET_SCRAPER`
- Optional.
- Overrides the tweet scraper script path. Default: `vendor/twitter/scrape_user_tweet_contents.py`.
- `ARCHIVR_TWEET_PYTHON`
- Optional.
- Overrides the Python executable used to run the tweet scraper. Default: `python3`.
### Current Limitations
- Arbitrary `http://` or `https://` URLs that return HTML are archived as self-contained single-file HTML snapshots via `single-file-cli` (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires `single-file` and a Chromium binary on PATH, or the `ARCHIVR_SINGLE_FILE` / `ARCHIVR_CHROME` env vars set.
- Local files currently need to be passed as `file://...` paths.
## License
This project is licensed under the MIT License. See the [LICENSE](LICENSE.md) file for details.