1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
No description
Find a file
TheGeneralist 1af920eb63
feat: share videos to TV (Chromecast + AirPlay) (#33)
* server: add scoped media-token endpoint for Cast/AirPlay auth bypass

Chromecast and Apple TV fetch media URLs as independent HTTP clients
with no session cookie. The existing serve_artifact handler requires
auth_user.require_auth(), so those devices always received 401.

Changes:
- MediaToken struct stored in AppState (Arc<Mutex<HashMap>>), scoped to
  a single (archive_id, entry_uid, artifact_index) tuple with a 2-hour TTL
- POST /api/archives/:id/entries/:uid/artifacts/:idx/media-token
  requires an authenticated session, verifies the artifact exists,
  prunes expired tokens, mints a 43-char URL-safe token, and returns
  { url, expires_in_secs }
- serve_artifact now accepts an optional ?token= query param; a valid
  scoped token bypasses require_auth() while a missing/invalid/expired
  token falls through to the normal 401 path
- CSP script-src extended to include https://www.gstatic.com so the
  Cast sender SDK script (injected lazily by VideoPreview) is not blocked
- 4 new tests: bare-URL still 401, tokenized fetch succeeds without
  session cookie, bogus token 401, wrong-artifact-index 401

* frontend: Cast/AirPlay overlay in VideoPreview

When the user opens a video archive entry, VideoPreview now:

1. Issues a signed media token (POST .../artifacts/:idx/media-token) and
   uses the returned signed URL as <video src>. This ensures the video
   element's src is one that Cast devices and Apple TV can fetch without a
   session cookie.

2. Lazily injects the Google Cast SDK script (cast_sender.js from
   gstatic.com, now allowed by the updated CSP). Once the SDK reports
   available, a <google-cast-launcher> web component appears as an overlay
   button in the top-right corner of the video. Selecting a Cast device
   triggers loadMedia() with the signed URL and the artifact's MIME type.

3. Detects AirPlay support (webkitShowPlaybackTargetPicker on
   HTMLVideoElement) and shows an AirPlay icon button alongside Cast.
   The <video> element carries x-webkit-airplay='allow', so Safari's native
   controls also surface the AirPlay option. The explicit overlay button
   calls webkitShowPlaybackTargetPicker() for consistent placement.

Both buttons are hidden when the respective APIs are unavailable (HTTP
pages, non-Safari for AirPlay, no Cast extension/devices), so there is no
UI regression for users who don't cast.

PreviewPanel now passes contentType (derived from artifact extension) to
VideoPreview so Cast receives a correct MIME type.

New CSS: .video-tv-controls (absolute overlay), .video-tv-btn (frosted
glass icon button), .video-tv-loading (placeholder during token fetch).

* server: fix serve_artifact auth OR logic — bogus token falls back to session

Previously a request carrying ?token=<expired> was immediately rejected
with 401, even if the user held a valid session cookie. This broke
logged-in browser playback after the 2-hour signed-URL window expired,
because VideoPreview uses the signed URL as <video src>.

Fix: compute token_valid first; if the token is absent or invalid, fall
through to auth_user.require_auth() instead of returning early.
Effect: valid token skips session check, invalid/missing token checks
session, both invalid → 401 as before.

Updated the bogus-token-no-session test docstring to clarify it tests
the no-auth path specifically. Added new test:
  media_token_bogus_token_with_session_returns_200 — verifies a logged-in
  user can still fetch the artifact via a URL carrying a stale token.

* frontend: guard token-fetch effect against stale async resolution

A slow issueMediaToken() response for video A could resolve after the
user selected video B and call setSignedSrc(urlA), making the
preview/Cast play the wrong file.

Add a cancelled flag set in the effect cleanup; both .then and .catch
check it before touching state, so only the most recent src wins.

* frontend: load Cast media immediately if session already exists

Previously the effect only sent video to the TV on SESSION_STARTED /
SESSION_RESUMED events. Two gaps:

1. If a Cast session was already active when signedSrc became ready
   (e.g. the SDK resumed a session before the token fetch finished, or
   the user switches videos while already casting), nothing was sent.

2. Same gap if castReady fired after an already-established session.

Fix: extract loadMedia(session) and call it against
ctx.getCurrentSession() immediately when castReady + signedSrc are both
truthy, in addition to keeping the event listener for future connects.

* server: staged file-upload endpoint

POST /api/archives/:id/uploads streams a multipart body to a temp file
under the archive's store/temp/ directory and returns a staged_path the
capture pipeline can move into place.

- Routes: /api/archives/:id/uploads (POST, requires auth)
- Body cap: 10 GiB; chunk-streamed to disk, never buffered in memory
- Path-traversal sanitised on the filename field
- Temp files are cleaned up on error paths (disk-leak fix)
- main.rs wires the new route into the server startup
- Cargo: adds the multipart dependency

* frontend: file upload in Capture dialog

Drag-and-drop or 'Upload file' button stages files for archiving:

- File items sit alongside URL rows in the same list; each shows the
  original filename, a live progress bar during upload, and a check badge
  when ready.  The locator input is replaced entirely — no editable field.
- Archive button is disabled until all uploads finish; each file item
  contributes to the Archive N count once its upload is done.
- File items are excluded from sessionStorage persistence (they are
  transient — the staged server path would be invalid after a reload).

Staged-file cleanup is handled at every exit path so temp/uploads/ does
not accumulate:
  • removeRow on an in-progress item aborts the XHR; removeRow on a done
    item calls DELETE /archives/:id/uploads.
  • Dialog cancel (Escape / Cancel button) aborts all in-flight XHRs and
    DELETEs all completed staged files via the close-event handler.
  • handleArchive sets isSubmittingRef=true before dialog.close() so the
    close handler skips cleanup — the background capture job handles
    staged-file removal on success instead.
  • uploadFile() returns { promise, abort } so the component can cancel
    the XHR without any visible fetch.

api.js additions: uploadFile (XHR with progress + abort), deleteUpload.

(Static assets rebuilt from combined source to include screensharing
changes from this branch.)

* fix: collection enrollment with default_visibility_bits

Two related fixes from feat-file-uploading:

core: fix collection enrollment using default_visibility_bits instead of
entry.visibility — entries were being enrolled with the entry-level
visibility rather than the collection's configured default.

server: allow changing default_visibility_bits on the default collection
— the PATCH handler was incorrectly blocking updates to the default
collection's visibility configuration.

* server: fix unbounded staged-upload disk growth

Two review findings:

P2 — delete staged file on capture failure (routes.rs)
When perform_capture returns Err, the job was marked failed but
staged_upload_path was never removed. With a 10 GiB body cap a few
failed imports could exhaust archive storage before the next restart.
Mirror the success-path cleanup into the Err arm so the file is removed
immediately regardless of outcome.

P1 — periodic staged-upload pruning (main.rs)
The startup prune of temp/uploads/ only ran once, so uploads abandoned
mid-session (browser crash, navigation away) accumulated forever on a
long-running server. Folded the pruning logic into the existing 24 h
maintenance task alongside session cleanup, so stale dirs are swept
continuously without requiring a restart.

* server+frontend: fix staged-upload disk-growth and prune safety

Server (main.rs + routes.rs):

- Extract prune_stale_upload_dirs() helper called by both startup and
  the periodic 24h task, eliminating the duplicated loop.

- Sentinel (.uploading) created in the UUID dir before streaming begins;
  removed on successful completion; error path uses remove_dir_all so
  the partial file and sentinel are cleaned up together.
  The periodic prune skips any dir containing .uploading (active XHR).

- Startup prune passes cleanup_stale_sentinels=true: the server has not
  started accepting connections yet so any sentinel is a crash remnant —
  it is removed and the dir proceeds to the age check, preventing leaked
  dirs from a previous crash accumulating forever.

- Staleness measured from the newest non-sentinel child file mtime so a
  just-finished slow upload (dir mtime stale, file mtime fresh) is not
  pruned before the user can submit it for capture. Empty dirs fall back
  to dir mtime.

- Failed captures (Err branch in spawn_blocking) now also delete the
  staged file and UUID dir immediately, matching the success path.

Frontend (api.js + CaptureDialog.jsx):

- submitCapture attaches err.status = res.status on non-2xx responses
  so callers can distinguish a definite HTTP rejection from a network
  error where the response may have been lost.

- submitBgJob catch deletes the staged file only when e.status is set
  (server definitively rejected the POST /captures request). A network
  error leaves the file in place because the server may have accepted
  the job and the response was lost — deleting would race the capture.
2026-07-23 20:12:22 +02:00
.github/workflows ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
crates feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
docker chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
docs feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
frontend feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
modules/nixos fix: address 6 code-review findings 2026-06-29 21:30:45 +02:00
vendor feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
.dockerignore chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
.gitignore ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
AGENTS.md feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
ARCHIVR-MENTAL-MODEL.md feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
Cargo.lock feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
Cargo.toml feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
docker-compose.yml chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
Dockerfile docker: add uBlock Lite + ISDCAC extensions, mirror flake.nix parity 2026-07-12 17:09:17 +02:00
flake.lock chore: yt-dlp 2026.06.09 → 2026.07.04 (#25) 2026-07-13 14:58:29 +02:00
flake.nix feat: uBlock Origin Lite + cookie consent extension + reader mode + ad placeholder cleanup (#21) 2026-07-08 23:26:48 +02:00

archivr

An open-source self-hosted archiving tool. Work in progress.

  • Archiving
    • Archiving media files from social media platforms
      • YouTube Videos
      • YouTube Playlists
      • YouTube Channels
      • Twitter Videos
      • Instagram
      • Facebook
      • TikTok
      • Reddit
      • Snapchat
      • YouTube Posts (postponed)
    • Archiving local files
    • Archiving Twitter Tweets, Threads, and Articles
    • Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
      • URLs
      • Google Drive
      • Dropbox
      • OneDrive
      • (Some of these could be postponed for later.)
    • Archive web pages (HTML, CSS, JS, images)
    • Archiving emails (???)
      • Gmail
      • Outlook
      • Yahoo Mail
  • Management
    • Deduplication
    • Tagging system
    • Search functionality
    • Categorization
    • Metadata extraction and storage
  • User Interface
    • Web-based UI
      • Authentication and login
      • Archive setup
      • Browse and view entries
      • Tag management and filtering
      • Search entries
      • View archive runs
      • Capture dialog
      • User settings and API tokens
      • Admin panel
  • Backup and Sync
    • Cloud backup (AWS S3, Google Cloud Storage)
    • Local backup

Motivation

There are two driving factors behind this project:

  • In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
  • I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.

This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.

Archive Inputs

archivr archive <path> currently accepts three kinds of inputs:

  • Local files via file://...
  • Direct platform URLs
  • Platform shorthand inputs such as tweet:..., yt:..., or instagram:...

Running Archivr

Archivr currently ships as two binaries:

  • archivr
    • The CLI for creating and writing to one archive.
    • Use this for init and archive.
  • archivr-server
    • The web server for reading one or more existing archives through the browser UI.
    • Use this after archives already exist.

With Nix, run the CLI with:

nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf

Run the web server with:

nix run .#archivr-server -- ./archivr-server.toml

The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.

Example:

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"

Then open:

http://127.0.0.1:8080

When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.

Security and Deployment

archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.

Changing the bind address

You can set the bind address in your TOML config:

# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"

Or override it with the ARCHIVR_BIND environment variable:

ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml

If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:

warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.

When will auth be added?

Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.

Supported Platforms

  • Local files: file:///absolute/path/to/file.ext
  • YouTube media: individual videos/shorts, playlists, and channels; standard URLs or shorthand video inputs. Playlists and channels archive as a container entry with each video stored as a child entry beneath it.
  • X/Twitter media from Tweets: normal Tweet URLs or the tweet:media:ID shorthand
  • X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in raw_tweets/)
  • Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to yt-dlp

Video quality and audio-only downloads

When capturing via the web UI, entering a URL for a yt-dlp-backed source (YouTube, Instagram, TikTok, Facebook, Reddit, Snapchat, X media) triggers a metadata probe via GET /api/archives/:id/captures/probe. The quality selector then shows only the heights actually available in that video plus Best quality (default). An Audio only option is appended whenever the probe confirms an audio track exists. UI behaviour by probe outcome:

qualities has_audio UI shows
["1080p", "720p", …] true Best / heights / Audio only
["1080p", …] false Best / heights
[] true Audio only (pre-selected, no Best option)
[] false "No media detected"
probe fails (502) — picker hidden, capture still submittable

The POST /api/archives/:id/captures endpoint accepts an optional quality field: "best", "audio", or any "NNNp" height string:

{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }

"audio" selects the most efficient native audio track without transcoding: Opus/WebM is preferred (smallest at equivalent quality), then AAC/M4A, then whatever yt-dlp considers best. The saved file's extension matches the native format (.webm for Opus, .m4a for AAC, etc.) — no ffmpeg re-encode, no size inflation. Any "NNNp" height is accepted; the server builds the yt-dlp format selector with an unconditional /best fallback so the download succeeds even if the exact height is unavailable. Omitting quality or passing "best" downloads at the highest available quality. Anything else is rejected with HTTP 400.

The probe endpoint (GET /api/archives/:id/captures/probe?locator=…) requires auth and returns 200 with:

{ "has_video": true,  "has_audio": true,  "qualities": ["1080p", "720p", "480p"] }
{ "has_video": false, "has_audio": true,  "qualities": [] }
{ "has_video": false, "has_audio": false, "qualities": [] }

has_video: false, has_audio: false means yt-dlp found no downloadable tracks (e.g. a tweet with no media). A 502 means yt-dlp itself failed (transient network error, rate-limit, unsupported extractor) — treat as inconclusive, not "no media."

YouTube playlists and channels

Capturing a YouTube playlist or channel URL creates a container entry for the playlist or channel, with each video archived as a child entry beneath it. Before downloading, the capture UI probes each video to fetch available quality options, letting you set quality per-video or apply a single quality to the whole playlist.

Incremental sync: When re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container entry is reused.

Excluding individual videos: In the expanded per-video list, each video has a remove button (×) to exclude it from the current capture. Removed videos are not downloaded; the rest proceed normally.

Hosting on NixOS

The flake exposes a nixosModules.default output. Add it to your system flake and enable the service:

# flake.nix (your system flake)
{
  inputs.archivr.url = "github:thegeneralist/archivr";

  outputs = { nixpkgs, archivr, ... }: {
    nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
      modules = [
        archivr.nixosModules.default
        {
          services.archivr-server = {
            enable = true;
            # listenAddress defaults to "127.0.0.1" (loopback only)
            # port defaults to 8080
            archives = [
              { id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
              { id = "work";     label = "Work";     path = "/srv/archivr/work/.archivr"; }
            ];
          };
        }
      ];
    };
  };
}

The module:

  • Creates an archivr system user and group.
  • Generates the TOML config from your options and stores the auth database under /var/lib/archivr-server/ (persists across upgrades).
  • Runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp, etc.). Archive directories are whitelisted for read-write access.
  • Restarts automatically on failure.

openFirewall — set to true to open the TCP port derived from bind. Only needed when binding to a non-loopback address:

services.archivr-server = {
  listenAddress = "0.0.0.0";
  port = 8080;            # explicit, though 8080 is the default
  openFirewall = true;
};

Archive directories must be readable and writable by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.

Hosting with Docker

A Dockerfile and docker-compose.yml are provided for self-hosting without Nix.

Quickstart

  1. Copy the example config and edit it:

    mkdir config
    cp docker/config.example.toml config/archivr-server.toml
    # edit config/archivr-server.toml — set archive id, label, and archive_path
    
  2. Initialize each archive on the persistent data volume before the first start. The image includes the archivr CLI for this purpose:

    docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"
    

    This creates /data/archives/main/.archivr/ with the metadata the server requires. A bare mkdir is not enough — the server reads name and store_path files that only archivr init writes.

  3. Start the server:

    docker compose up -d
    

    Then open http://localhost:8080.

Volumes

Mount Purpose
./config (read-only) Directory containing archivr-server.toml
archivr-data named volume Auth database (/data/archivr-auth.sqlite) and archive directories

Important: auth_db_path must be set explicitly in archivr-server.toml to a path on the writable data volume (e.g. /data/archivr-auth.sqlite). If left unset, the server defaults to writing the auth database next to the config file — which is on the read-only /config mount and will fail. The example config sets this correctly.

Twitter/X archiving

Supply a cookies file inside the config volume and set ARCHIVR_TWITTER_CREDENTIALS_FILE in docker-compose.yml:

environment:
  ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt

Building the image locally

docker build -t archivr-server .

The image compiles the Rust binary in a separate build stage so only the runtime dependencies (Chromium, Node.js, Python) land in the final layer.

Supported Shorthand Inputs

  • YouTube video/short media:
    • yt:video/ID
    • youtube:video/ID
    • yt:short/ID
    • yt:shorts/ID
    • youtube:shorts/ID
  • X/Twitter tweet JSON content:
    • tweet:ID
    • x:tweet:ID
    • x:x:ID
    • twitter:x:ID
    • twitter:tweet:ID
  • X/Twitter media/video download:
    • tweet:media:ID
  • X/Twitter thread JSON content:
    • x:thread:ID
    • twitter:thread:ID
  • Other platform shorthands:
    • instagram:ID
    • facebook:ID
    • tiktok:ID
    • reddit:ID
    • snapchat:ID

Environment Variables

  • ARCHIVR_BIND
    • Optional.
    • Overrides the bind address from the TOML config. Useful in Docker where you need 0.0.0.0:8080 without editing the config file. Default: 127.0.0.1:8080.
  • ARCHIVR_STATIC_DIR
    • Optional.
    • Path to the directory of pre-built frontend assets served by the web UI. Set automatically by the Nix wrapper and the Docker image. When running from source with cargo run, falls back to crates/archivr-server/static.
  • ARCHIVR_YT_DLP
    • Optional.
    • Overrides the yt-dlp binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
  • ARCHIVR_SINGLE_FILE
    • Optional.
    • Overrides the single-file binary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
  • ARCHIVR_CHROME
    • Optional.
    • Overrides the Chromium/Chrome executable passed to single-file via --browser-executable-path. Set automatically by the Nix wrapper and the Docker image. Default: chromium.
  • ARCHIVR_CHROME_ARGS
    • Optional.
    • Space-separated extra flags appended to Chromium's --browser-args. The Docker image sets this to --no-sandbox because Chromium refuses to run as root without it. Leave unset when running natively (Nix, Linux desktop). A --window-size=1920,1080 is always passed to provide a realistic desktop viewport (so responsive @media rules and styles are evaluated and preserved correctly). Supply your own --window-size=... here to override.
  • ARCHIVR_TWITTER_CREDENTIALS_FILE
    • Required for tweet/thread scraping inputs such as tweet:ID and x:thread:ID.
    • Must point to a cookies file for the vendored scraper.
  • ARCHIVR_TWEET_SCRAPER
    • Optional.
    • Overrides the tweet scraper script path. Default: vendor/twitter/scrape_user_tweet_contents.py.
  • ARCHIVR_TWEET_PYTHON
    • Optional.
    • Overrides the Python executable used to run the tweet scraper. Default: python3.

Current Limitations

  • Arbitrary http:// or https:// URLs that return HTML are archived as self-contained single-file HTML snapshots via single-file-cli (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires single-file and a Chromium binary on PATH, or the ARCHIVR_SINGLE_FILE / ARCHIVR_CHROME env vars set.
  • Local files currently need to be passed as file://... paths.

License

This project is licensed under the MIT License. See the LICENSE file for details.