Add --window-size=1920,1080 to the Chromium flags passed via --browser-args. This makes the existing --remove-unused-styles=false and --remove-alternative-medias=false effective for real @media rules and responsive styles (headless default is small). Also document in ARCHIVR_CHROME_ARGS that users can override by supplying their own --window-size in the env var. |
||
|---|---|---|
| .. | ||
| superpowers | ||
| LICENSE.md | ||
| README.md | ||
archivr
An open-source self-hosted archiving tool. Work in progress.
- Archiving
- Archiving media files from social media platforms
- YouTube Videos
- YouTube Playlists
- YouTube Channels
- Twitter Videos
- TikTok
- Snapchat
- YouTube Posts (postponed)
- Archiving local files
- Archiving Twitter Tweets, Threads, and Articles
- Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
- URLs
- Google Drive
- Dropbox
- OneDrive
- (Some of these could be postponed for later.)
- Archive web pages (HTML, CSS, JS, images)
- Archiving emails (???)
- Gmail
- Outlook
- Yahoo Mail
- Archiving media files from social media platforms
- Management
- Deduplication
- Tagging system
- Search functionality
- Categorization
- Metadata extraction and storage
- User Interface
- Web-based UI
- Authentication and login
- Archive setup
- Browse and view entries
- Tag management and filtering
- Search entries
- View archive runs
- Capture dialog
- User settings and API tokens
- Admin panel
- Web-based UI
- Backup and Sync
- Cloud backup (AWS S3, Google Cloud Storage)
- Local backup
Motivation
There are two driving factors behind this project:
- In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
- I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.
This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.
Archive Inputs
archivr archive <path> currently accepts three kinds of inputs:
- Local files via
file://... - Direct platform URLs
- Platform shorthand inputs such as
tweet:...,yt:..., orinstagram:...
Running Archivr
Archivr currently ships as two binaries:
archivr- The CLI for creating and writing to one archive.
- Use this for
initandarchive.
archivr-server- The web server for reading one or more existing archives through the browser UI.
- Use this after archives already exist.
With Nix, run the CLI with:
nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
Run the web server with:
nix run .#archivr-server -- ./archivr-server.toml
The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.
Example:
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
Then open:
http://127.0.0.1:8080
When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.
Security and Deployment
archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.
Changing the bind address
You can set the bind address in your TOML config:
# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"
Or override it with the ARCHIVR_BIND environment variable:
ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml
If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:
warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.
When will auth be added?
Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.
Supported Platforms
- Local files:
file:///absolute/path/to/file.ext - YouTube media: standard video/short URLs, plus shorthand video inputs
- X/Twitter media from Tweets: normal Tweet URLs or the
tweet:media:IDshorthand - X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in
raw_tweets/) - Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to
yt-dlp
Hosting on NixOS
The flake exposes a nixosModules.default output. Add it to your system flake and
enable the service:
# flake.nix (your system flake)
{
inputs.archivr.url = "github:thegeneralist/archivr";
outputs = { nixpkgs, archivr, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
archivr.nixosModules.default
{
services.archivr-server = {
enable = true;
# listenAddress defaults to "127.0.0.1" (loopback only)
# port defaults to 8080
archives = [
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
];
};
}
];
};
};
}
The module:
- Creates an
archivrsystem user and group. - Generates the TOML config from your options and stores the auth database under
/var/lib/archivr-server/(persists across upgrades). - Runs under a hardened systemd unit (
ProtectSystem = strict,NoNewPrivileges,PrivateTmp, etc.). Archive directories are whitelisted for read-write access. - Restarts automatically on failure.
openFirewall — set to true to open the TCP port derived from bind.
Only needed when binding to a non-loopback address:
services.archivr-server = {
listenAddress = "0.0.0.0";
port = 8080; # explicit, though 8080 is the default
openFirewall = true;
};
Archive directories must be readable and writable by the archivr user.
Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.
Hosting with Docker
A Dockerfile and docker-compose.yml are provided for self-hosting without Nix.
Quickstart
-
Copy the example config and edit it:
mkdir config cp docker/config.example.toml config/archivr-server.toml # edit config/archivr-server.toml — set archive id, label, and archive_path -
Initialize each archive on the persistent data volume before the first start. The image includes the
archivrCLI for this purpose:docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"This creates
/data/archives/main/.archivr/with the metadata the server requires. A baremkdiris not enough — the server readsnameandstore_pathfiles that onlyarchivr initwrites. -
Start the server:
docker compose up -dThen open
http://localhost:8080.
Volumes
| Mount | Purpose |
|---|---|
./config (read-only) |
Directory containing archivr-server.toml |
archivr-data named volume |
Auth database (/data/archivr-auth.sqlite) and archive directories |
Important:
auth_db_pathmust be set explicitly inarchivr-server.tomlto a path on the writable data volume (e.g./data/archivr-auth.sqlite). If left unset, the server defaults to writing the auth database next to the config file — which is on the read-only/configmount and will fail. The example config sets this correctly.
Twitter/X archiving
Supply a cookies file inside the config volume and set ARCHIVR_TWITTER_CREDENTIALS_FILE in docker-compose.yml:
environment:
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
Building the image locally
docker build -t archivr-server .
The image compiles the Rust binary in a separate build stage so only the runtime dependencies (Chromium, Node.js, Python) land in the final layer.
Supported Shorthand Inputs
- YouTube video/short media:
yt:video/IDyoutube:video/IDyt:short/IDyt:shorts/IDyoutube:shorts/ID
- X/Twitter tweet JSON content:
tweet:IDx:tweet:IDx:x:IDtwitter:x:IDtwitter:tweet:ID
- X/Twitter media/video download:
tweet:media:ID
- X/Twitter thread JSON content:
x:thread:IDtwitter:thread:ID
- Other platform shorthands:
instagram:IDfacebook:IDtiktok:IDreddit:IDsnapchat:ID
Environment Variables
ARCHIVR_BIND- Optional.
- Overrides the bind address from the TOML config. Useful in Docker where you need
0.0.0.0:8080without editing the config file. Default:127.0.0.1:8080.
ARCHIVR_STATIC_DIR- Optional.
- Path to the directory of pre-built frontend assets served by the web UI.
Set automatically by the Nix wrapper and the Docker image. When running from
source with
cargo run, falls back tocrates/archivr-server/static.
ARCHIVR_YT_DLP- Optional.
- Overrides the
yt-dlpbinary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
ARCHIVR_SINGLE_FILE- Optional.
- Overrides the
single-filebinary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
ARCHIVR_CHROME- Optional.
- Overrides the Chromium/Chrome executable passed to
single-filevia--browser-executable-path. Set automatically by the Nix wrapper and the Docker image. Default:chromium.
ARCHIVR_CHROME_ARGS- Optional.
- Space-separated extra flags appended to Chromium's
--browser-args. The Docker image sets this to--no-sandboxbecause Chromium refuses to run as root without it. Leave unset when running natively (Nix, Linux desktop). A--window-size=1920,1080is always passed to provide a realistic desktop viewport (so responsive @media rules and styles are evaluated and preserved correctly). Supply your own--window-size=...here to override.
ARCHIVR_TWITTER_CREDENTIALS_FILE- Required for tweet/thread scraping inputs such as
tweet:IDandx:thread:ID. - Must point to a cookies file for the vendored scraper.
- Required for tweet/thread scraping inputs such as
ARCHIVR_TWEET_SCRAPER- Optional.
- Overrides the tweet scraper script path. Default:
vendor/twitter/scrape_user_tweet_contents.py.
ARCHIVR_TWEET_PYTHON- Optional.
- Overrides the Python executable used to run the tweet scraper. Default:
python3.
Current Limitations
- Arbitrary
http://orhttps://URLs that return HTML are archived as self-contained single-file HTML snapshots viasingle-file-cli(requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requiressingle-fileand a Chromium binary on PATH, or theARCHIVR_SINGLE_FILE/ARCHIVR_CHROMEenv vars set. - Local files currently need to be passed as
file://...paths.
License
This project is licensed under the MIT License. See the LICENSE file for details.