mirror of
https://github.com/thegeneralist01/archivr
synced 2026-07-21 18:55:36 +02:00
* chore: add Dockerfile, docker-compose, and Docker docs - Multi-stage Dockerfile: Rust builder stage + debian:bookworm-slim runtime with Chromium, Node/single-file-cli, Python venv (yt-dlp + twitter-api-client) - docker-compose.yml: wires ARCHIVR_BIND, config volume, and persistent data volume - docker/config.example.toml: annotated TOML template for Docker deployments - docs/README.md: add Hosting with Docker section; add ARCHIVR_BIND and ARCHIVR_STATIC_DIR to the Environment Variables reference * fix: address code review issues with Docker setup - .gitignore: whitelist Dockerfile, docker-compose.yml, docker/ so they are actually tracked (the * catch-all was silently dropping them) - Dockerfile: build and ship the archivr CLI alongside archivr-server so users can run `archivr init` inside the container on first setup - docker/config.example.toml: fix archive_path to point at the .archivr subdirectory that archivr init creates (not the parent directory), which is what read_archive_paths expects - docs/README.md: replace the bare mkdir quickstart step with `archivr init`, explain why mkdir is insufficient; add a callout that auth_db_path must be set explicitly to a writable path when the config mount is read-only * fix: address second round of Docker review issues Chromium sandbox (P2): - singlefile.rs: add ARCHIVR_CHROME_ARGS env var (space-separated flags appended to Chromium's --browser-args JSON array); Dockerfile sets it to --no-sandbox because Chromium refuses to start as root without it Store-path outside volume (P1): - README: pass explicit absolute store-path as the second positional arg to `archivr init` so the blob store lands on /data instead of the container layer (CLI default is ./.archivr/store, resolved from cwd, which is / with no WORKDIR set) ENTRYPOINT vs CMD (P2): - Dockerfile: switch from ENTRYPOINT to CMD so `docker compose run archivr archivr init …` overrides the full command instead of being appended to the server invocation ffmpeg missing (P2): - Dockerfile: add ffmpeg to the apt-get install block (required by yt-dlp --merge-output-format mp4 for bestvideo+bestaudio streams) Node version (P2): - Dockerfile: replace Debian bookworm's nodejs (18.x) with Node 20 via the NodeSource setup script (single-file-cli declares engines.node >=20) Build context secrets (P2): - Add .dockerignore excluding config/ and docker/ from the build context so runtime secrets (e.g. twitter-cookies.txt) are never sent to the builder - Whitelist .dockerignore in .gitignore docs: - README: document ARCHIVR_CHROME_ARGS in the Environment Variables section * fix: third round of Docker review issues Rust toolchain (P1): - Dockerfile: bump builder from rust:1.87 to rust:1.88; time@0.3.51, time-core@0.1.9, and time-macros@0.2.30 (present in Cargo.lock) all require MSRV 1.88, so the real cargo build --release step was failing single-file-cli wait mode (P2): - singlefile.rs: replace --browser-wait-until=networkidle2 with networkAlmostIdle; the single-file-cli option only accepts InteractiveTime/networkIdle/networkAlmostIdle/load/domContentLoaded (verified in options.js); networkidle2 is a Puppeteer concept that the CLI does not recognise, causing silent fallback to the earliest state and incomplete captures. networkAlmostIdle is the closest equivalent (<=2 open connections, matching Puppeteer's networkidle2 semantics) Build context size (P3): - .dockerignore: add target/, frontend/node_modules/, frontend/dist/; these can reach 1.4G+ after a local dev build and are never read by the Dockerfile, so sending them to the builder wastes time and memory
325 lines
11 KiB
Markdown
325 lines
11 KiB
Markdown
# archivr
|
|
|
|
An open-source self-hosted archiving tool. Work in progress.
|
|
|
|
## Milestones
|
|
|
|
- [ ] Archiving
|
|
- [x] Archiving media files from social media platforms
|
|
- [x] YouTube Videos
|
|
- [x] Twitter Videos
|
|
- [x] Instagram
|
|
- [x] Facebook
|
|
- [x] TikTok
|
|
- [x] Reddit
|
|
- [x] Snapchat
|
|
- [ ] YouTube Posts (postponed)
|
|
- [x] Archiving local files
|
|
- [x] Archiving Twitter Tweets, Threads, and Articles
|
|
- [ ] Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
|
|
- [x] URLs
|
|
- [ ] Google Drive
|
|
- [ ] Dropbox
|
|
- [ ] OneDrive
|
|
- (Some of these could be postponed for later.)
|
|
- [x] Archive web pages (HTML, CSS, JS, images)
|
|
- [ ] Archiving emails (???)
|
|
- [ ] Gmail
|
|
- [ ] Outlook
|
|
- [ ] Yahoo Mail
|
|
- [ ] Management
|
|
- [ ] Deduplication
|
|
- [ ] Tagging system
|
|
- [ ] Search functionality
|
|
- [ ] Categorization
|
|
- [ ] Metadata extraction and storage
|
|
- [ ] User Interface
|
|
- [ ] Web-based UI
|
|
- [ ] Backup and Sync
|
|
- [ ] Cloud backup (AWS S3, Google Cloud Storage)
|
|
- [ ] Local backup
|
|
|
|
## Motivation
|
|
|
|
There are two driving factors behind this project:
|
|
|
|
- In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is _very important_ for preserving personal memories and digital history.
|
|
- I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.
|
|
|
|
This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.
|
|
|
|
## Archive Inputs
|
|
|
|
`archivr archive <path>` currently accepts three kinds of inputs:
|
|
|
|
- Local files via `file://...`
|
|
- Direct platform URLs
|
|
- Platform shorthand inputs such as `tweet:...`, `yt:...`, or `instagram:...`
|
|
|
|
## Running Archivr
|
|
|
|
Archivr currently ships as two binaries:
|
|
|
|
- `archivr`
|
|
- The CLI for creating and writing to one archive.
|
|
- Use this for `init` and `archive`.
|
|
- `archivr-server`
|
|
- The web server for reading one or more existing archives through the browser UI.
|
|
- Use this after archives already exist.
|
|
|
|
With Nix, run the CLI with:
|
|
|
|
```sh
|
|
nix run .#archivr -- init ./my-archive --name "My Archive"
|
|
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
|
|
```
|
|
|
|
Run the web server with:
|
|
|
|
```sh
|
|
nix run .#archivr-server -- ./archivr-server.toml
|
|
```
|
|
|
|
The server expects a TOML registry file. If no path is passed, it reads `./archivr-server.toml`.
|
|
|
|
Example:
|
|
|
|
```toml
|
|
[[archives]]
|
|
id = "personal"
|
|
label = "Personal"
|
|
archive_path = "/absolute/path/to/my-archive/.archivr"
|
|
```
|
|
|
|
Then open:
|
|
|
|
```text
|
|
http://127.0.0.1:8080
|
|
```
|
|
|
|
When installed through Nix, `archivr-server` is wrapped so it can find the static web UI assets automatically. The wrapper sets `ARCHIVR_STATIC_DIR` to the installed static asset directory. Running from source with `cargo run -p archivr-server` falls back to `crates/archivr-server/static`.
|
|
|
|
### Security and Deployment
|
|
|
|
`archivr-server` is a **local-only tool by default**. It binds to `127.0.0.1:8080` and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.
|
|
|
|
**Changing the bind address**
|
|
|
|
You can set the bind address in your TOML config:
|
|
|
|
```toml
|
|
# Optional. Default: 127.0.0.1:8080
|
|
# Only change this if you know what you are doing — the server has no authentication.
|
|
bind = "127.0.0.1:9090"
|
|
```
|
|
|
|
Or override it with the `ARCHIVR_BIND` environment variable:
|
|
|
|
```sh
|
|
ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml
|
|
```
|
|
|
|
If the server is started with a non-loopback address (e.g. `0.0.0.0`), it prints a warning to stderr:
|
|
|
|
```text
|
|
warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.
|
|
```
|
|
|
|
**When will auth be added?**
|
|
|
|
Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See `crates/archivr-server/src/routes.rs` for the route classification that will guide where middleware is applied.
|
|
|
|
### Supported Platforms
|
|
|
|
- Local files: `file:///absolute/path/to/file.ext`
|
|
- YouTube media: standard video/short URLs, plus [shorthand video inputs](#supported-shorthand-inputs)
|
|
- X/Twitter media from Tweets: normal Tweet URLs or the `tweet:media:ID` shorthand
|
|
- X/Twitter Tweet content scrape: [Tweet and Thread shorthands](#supported-shorthand-inputs). (These are saved as JSON files in `raw_tweets/`)
|
|
- Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to `yt-dlp`
|
|
|
|
### Hosting on NixOS
|
|
|
|
The flake exposes a `nixosModules.default` output. Add it to your system flake and
|
|
enable the service:
|
|
|
|
```nix
|
|
# flake.nix (your system flake)
|
|
{
|
|
inputs.archivr.url = "github:thegeneralist/archivr";
|
|
|
|
outputs = { nixpkgs, archivr, ... }: {
|
|
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
|
|
modules = [
|
|
archivr.nixosModules.default
|
|
{
|
|
services.archivr-server = {
|
|
enable = true;
|
|
# listenAddress defaults to "127.0.0.1" (loopback only)
|
|
# port defaults to 8080
|
|
archives = [
|
|
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
|
|
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
|
|
];
|
|
};
|
|
}
|
|
];
|
|
};
|
|
};
|
|
}
|
|
```
|
|
|
|
The module:
|
|
- Creates an `archivr` system user and group.
|
|
- Generates the TOML config from your options and stores the auth database under
|
|
`/var/lib/archivr-server/` (persists across upgrades).
|
|
- Runs under a hardened systemd unit (`ProtectSystem = strict`, `NoNewPrivileges`,
|
|
`PrivateTmp`, etc.). Archive directories are whitelisted for read-write access.
|
|
- Restarts automatically on failure.
|
|
|
|
**`openFirewall`** — set to `true` to open the TCP port derived from `bind`.
|
|
Only needed when binding to a non-loopback address:
|
|
|
|
```nix
|
|
services.archivr-server = {
|
|
listenAddress = "0.0.0.0";
|
|
port = 8080; # explicit, though 8080 is the default
|
|
openFirewall = true;
|
|
};
|
|
```
|
|
|
|
**Archive directories** must be readable and writable by the `archivr` user.
|
|
Initialise them with `archivr init` first, then `chown -R archivr:archivr /srv/archivr`.
|
|
|
|
|
|
### Hosting with Docker
|
|
|
|
A `Dockerfile` and `docker-compose.yml` are provided for self-hosting without Nix.
|
|
|
|
**Quickstart**
|
|
|
|
1. Copy the example config and edit it:
|
|
|
|
```sh
|
|
mkdir config
|
|
cp docker/config.example.toml config/archivr-server.toml
|
|
# edit config/archivr-server.toml — set archive id, label, and archive_path
|
|
```
|
|
|
|
2. Initialize each archive on the persistent data volume before the first start.
|
|
The image includes the `archivr` CLI for this purpose:
|
|
|
|
```sh
|
|
docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"
|
|
```
|
|
|
|
This creates `/data/archives/main/.archivr/` with the metadata the server requires.
|
|
A bare `mkdir` is not enough — the server reads `name` and `store_path` files that
|
|
only `archivr init` writes.
|
|
|
|
3. Start the server:
|
|
|
|
```sh
|
|
docker compose up -d
|
|
```
|
|
|
|
Then open `http://localhost:8080`.
|
|
|
|
**Volumes**
|
|
|
|
| Mount | Purpose |
|
|
|-------|---------|
|
|
| `./config` (read-only) | Directory containing `archivr-server.toml` |
|
|
| `archivr-data` named volume | Auth database (`/data/archivr-auth.sqlite`) and archive directories |
|
|
|
|
> **Important:** `auth_db_path` must be set explicitly in `archivr-server.toml` to a
|
|
> path on the writable data volume (e.g. `/data/archivr-auth.sqlite`). If left unset,
|
|
> the server defaults to writing the auth database next to the config file — which is
|
|
> on the read-only `/config` mount and will fail. The example config sets this correctly.
|
|
|
|
**Twitter/X archiving**
|
|
|
|
Supply a cookies file inside the config volume and set `ARCHIVR_TWITTER_CREDENTIALS_FILE` in `docker-compose.yml`:
|
|
|
|
```yaml
|
|
environment:
|
|
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
|
|
```
|
|
|
|
**Building the image locally**
|
|
|
|
```sh
|
|
docker build -t archivr-server .
|
|
```
|
|
|
|
The image compiles the Rust binary in a separate build stage so only the runtime
|
|
dependencies (Chromium, Node.js, Python) land in the final layer.
|
|
|
|
### Supported Shorthand Inputs
|
|
|
|
- YouTube video/short media:
|
|
- `yt:video/ID`
|
|
- `youtube:video/ID`
|
|
- `yt:short/ID`
|
|
- `yt:shorts/ID`
|
|
- `youtube:shorts/ID`
|
|
- X/Twitter tweet JSON content:
|
|
- `tweet:ID`
|
|
- `x:tweet:ID`
|
|
- `x:x:ID`
|
|
- `twitter:x:ID`
|
|
- `twitter:tweet:ID`
|
|
- X/Twitter media/video download:
|
|
- `tweet:media:ID`
|
|
- X/Twitter thread JSON content:
|
|
- `x:thread:ID`
|
|
- `twitter:thread:ID`
|
|
- Other platform shorthands:
|
|
- `instagram:ID`
|
|
- `facebook:ID`
|
|
- `tiktok:ID`
|
|
- `reddit:ID`
|
|
- `snapchat:ID`
|
|
|
|
### Environment Variables
|
|
|
|
- `ARCHIVR_BIND`
|
|
- Optional.
|
|
- Overrides the bind address from the TOML config. Useful in Docker where you need
|
|
`0.0.0.0:8080` without editing the config file. Default: `127.0.0.1:8080`.
|
|
- `ARCHIVR_STATIC_DIR`
|
|
- Optional.
|
|
- Path to the directory of pre-built frontend assets served by the web UI.
|
|
Set automatically by the Nix wrapper and the Docker image. When running from
|
|
source with `cargo run`, falls back to `crates/archivr-server/static`.
|
|
- `ARCHIVR_YT_DLP`
|
|
- Optional.
|
|
- Overrides the `yt-dlp` binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
|
|
- `ARCHIVR_SINGLE_FILE`
|
|
- Optional.
|
|
- Overrides the `single-file` binary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
|
|
- `ARCHIVR_CHROME`
|
|
- Optional.
|
|
- Overrides the Chromium/Chrome executable passed to `single-file` via `--browser-executable-path`. Set automatically by the Nix wrapper and the Docker image. Default: `chromium`.
|
|
- `ARCHIVR_CHROME_ARGS`
|
|
- Optional.
|
|
- Space-separated extra flags appended to Chromium's `--browser-args`. The Docker
|
|
image sets this to `--no-sandbox` because Chromium refuses to run as root without
|
|
it. Leave unset when running natively (Nix, Linux desktop).
|
|
- `ARCHIVR_TWITTER_CREDENTIALS_FILE`
|
|
- Required for tweet/thread scraping inputs such as `tweet:ID` and `x:thread:ID`.
|
|
- Must point to a cookies file for the vendored scraper.
|
|
- `ARCHIVR_TWEET_SCRAPER`
|
|
- Optional.
|
|
- Overrides the tweet scraper script path. Default: `vendor/twitter/scrape_user_tweet_contents.py`.
|
|
- `ARCHIVR_TWEET_PYTHON`
|
|
- Optional.
|
|
- Overrides the Python executable used to run the tweet scraper. Default: `python3`.
|
|
|
|
### Current Limitations
|
|
|
|
- Arbitrary `http://` or `https://` URLs that return HTML are archived as self-contained single-file HTML snapshots via `single-file-cli` (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires `single-file` and a Chromium binary on PATH, or the `ARCHIVR_SINGLE_FILE` / `ARCHIVR_CHROME` env vars set.
|
|
- Local files currently need to be passed as `file://...` paths.
|
|
|
|
## License
|
|
|
|
This project is licensed under the MIT License. See the [LICENSE](LICENSE.md) file for details.
|