1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-07-21 18:55:36 +02:00
No description
Find a file
2026-06-26 11:19:40 +02:00
crates fix(capture): remove debug println! from hash_exists 2026-06-25 16:19:33 +02:00
docs docs: add Track 4 auth foundation implementation plan 2026-06-25 17:38:30 +02:00
frontend feat(ui): show website favicon in entry list for webpage entries 2026-06-24 19:16:21 +02:00
vendor/twitter fix: extract full tweet text from note_tweet field when available 2026-04-06 11:05:32 +02:00
.gitignore chore: remove docs/superpowers from gitignore (not for VCS) 2026-06-26 11:19:40 +02:00
ARCHIVR-MENTAL-MODEL.md docs: remove completed planning docs and stale superpowers plan 2026-06-23 17:24:24 +02:00
Cargo.lock feat(singlefile): add SaveResult, favicon extraction, wait for networkidle2 2026-06-24 19:16:21 +02:00
Cargo.toml feat(singlefile): add SaveResult, favicon extraction, wait for networkidle2 2026-06-24 19:16:21 +02:00
flake.lock feat: add archiving of platform media files (#1) 2026-03-31 12:39:35 +02:00
flake.nix fix(nix): guard pkgs.chromium behind isLinux for aarch64-darwin 2026-06-24 18:46:11 +02:00
NEXT.md docs: add Track 4 auth foundation design spec + whitelist docs/superpowers in gitignore 2026-06-25 17:20:17 +02:00

archivr

An open-source self-hosted archiving tool. Work in progress.

Milestones

  • Archiving
    • Archiving media files from social media platforms
      • YouTube Videos
      • Twitter Videos
      • Instagram
      • Facebook
      • TikTok
      • Reddit
      • Snapchat
      • YouTube Posts (postponed)
    • Archiving local files
    • Archiving Twitter Tweets, Threads, and Articles
    • Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
      • URLs
      • Google Drive
      • Dropbox
      • OneDrive
      • (Some of these could be postponed for later.)
    • Archive web pages (HTML, CSS, JS, images)
    • Archiving emails (???)
      • Gmail
      • Outlook
      • Yahoo Mail
  • Management
    • Deduplication
    • Tagging system
    • Search functionality
    • Categorization
    • Metadata extraction and storage
  • User Interface
    • Web-based UI
  • Backup and Sync
    • Cloud backup (AWS S3, Google Cloud Storage)
    • Local backup

Motivation

There are two driving factors behind this project:

  • In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
  • I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.

This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.

Archive Inputs

archivr archive <path> currently accepts three kinds of inputs:

  • Local files via file://...
  • Direct platform URLs
  • Platform shorthand inputs such as tweet:..., yt:..., or instagram:...

Running Archivr

Archivr currently ships as two binaries:

  • archivr
    • The CLI for creating and writing to one archive.
    • Use this for init and archive.
  • archivr-server
    • The web server for reading one or more existing archives through the browser UI.
    • Use this after archives already exist.

With Nix, run the CLI with:

nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf

Run the web server with:

nix run .#archivr-server -- ./archivr-server.toml

The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.

Example:

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"

Then open:

http://127.0.0.1:8080

When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.

Security and Deployment

archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.

Changing the bind address

You can set the bind address in your TOML config:

# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"

Or override it with the ARCHIVR_BIND environment variable:

ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml

If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:

warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.

When will auth be added?

Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.

Supported Platforms

  • Local files: file:///absolute/path/to/file.ext
  • YouTube media: standard video/short URLs, plus shorthand video inputs
  • X/Twitter media from Tweets: normal Tweet URLs or the tweet:media:ID shorthand
  • X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in raw_tweets/)
  • Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to yt-dlp

Supported Shorthand Inputs

  • YouTube video/short media:
    • yt:video/ID
    • youtube:video/ID
    • yt:short/ID
    • yt:shorts/ID
    • youtube:shorts/ID
  • X/Twitter tweet JSON content:
    • tweet:ID
    • x:tweet:ID
    • x:x:ID
    • twitter:x:ID
    • twitter:tweet:ID
  • X/Twitter media/video download:
    • tweet:media:ID
  • X/Twitter thread JSON content:
    • x:thread:ID
    • twitter:thread:ID
  • Other platform shorthands:
    • instagram:ID
    • facebook:ID
    • tiktok:ID
    • reddit:ID
    • snapchat:ID

Environment Variables

  • ARCHIVR_YT_DLP
    • Optional.
    • Overrides the yt-dlp binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
  • ARCHIVR_SINGLE_FILE
    • Optional.
    • Overrides the single-file binary used for web page archiving. When installed through Nix, this is set automatically to the Nixpkgs single-file-cli binary.
  • ARCHIVR_CHROME
    • Optional.
    • Overrides the Chromium/Chrome executable passed to single-file via --browser-executable-path. When installed through Nix, this is set automatically to the Nixpkgs chromium binary. Default: chromium.
  • ARCHIVR_TWITTER_CREDENTIALS_FILE
    • Required for tweet/thread scraping inputs such as tweet:ID and x:thread:ID.
    • Must point to a cookies file for the vendored scraper.
  • ARCHIVR_TWEET_SCRAPER
    • Optional.
    • Overrides the tweet scraper script path. Default: vendor/twitter/scrape_user_tweet_contents.py.
  • ARCHIVR_TWEET_PYTHON
    • Optional.
    • Overrides the Python executable used to run the tweet scraper. Default: python3.

Current Limitations

  • Arbitrary http:// or https:// URLs that return HTML are archived as self-contained single-file HTML snapshots via single-file-cli (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires single-file and a Chromium binary on PATH, or the ARCHIVR_SINGLE_FILE / ARCHIVR_CHROME env vars set.
  • Local files currently need to be passed as file://... paths.

License

This project is licensed under the MIT License. See the LICENSE file for details.