1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
No description
Find a file
TheGeneralist e1ee05bd41
feat(collections): public collections, per-collection auth, UX improvements (#34)
* feat(core): add requires_auth to collections; include name in entry-collection memberships

- Add `requires_auth INTEGER NOT NULL DEFAULT 1` column to the
  collections DDL and as an idempotent ALTER TABLE migration in
  initialize_schema (archive DB), not initialize_auth_schema.
- CollectionRecord and CollectionSummary gain `requires_auth: bool`.
- create_collection() and update_collection() accept the new field.
- get_entry_collection_memberships() now returns collection name as the
  third tuple element; EntryCollectionMembership gains a `name` field
  so the sidebar can show human-readable names instead of raw UIDs.

* feat(server): conditional auth for public collections; add requires_auth + original_url to API

- CreateCollectionBody gains requires_auth (default true).
- PatchCollectionBody gains requires_auth: Option<bool>.
- get_collection_handler: load record first, then skip auth.require_auth()
  when record.requires_auth == false so public collections are accessible
  to unauthenticated callers; caller_bits falls back to ROLE_GUEST (1)
  so only visibility_bits=3 entries are returned to guests.
- Collection JSON response includes requires_auth and each entry now
  includes original_url for use by the public collection page.
- list_collections_handler keeps require_auth (management UI).

* feat(frontend): public collection page at /c/:archiveId/:collUid

- Detect PUBLIC_COLL_ROUTE at module load time (like PREVIEW_ROUTE) and
  return <PublicCollectionPage> before any auth checks so unauthenticated
  users can view public collections without hitting the login gate.
- PublicCollectionPage fetches via getCollection() and renders the
  server-filtered entry list (no client-side bitmask filtering - the
  server already applies caller_bits=GUEST for unauthenticated requests).
  Entry titles link to original_url when present; fall back to plain
  text when original_url is null.
- api.js createCollection() gains requiresAuth param (default true),
  sent as requires_auth in the request body.
- Storybook story covers WithEntries, Empty, and LoadError states.

* feat(frontend): collections view improvements

- addVis in the 'Add entry' form now syncs to the selected collection's
  default_visibility_bits via useEffect on collDetail, so the default
  matches the collection's configured entry visibility.
- Rename 'Default visibility' label to 'Entries\' default visibility'
  in both the detail pane and the create form to distinguish it from
  the new collection-level access setting.
- Add 'Require authentication to view' checkbox in the detail pane
  backed by a PATCH to requires_auth; reads collDetail?.requires_auth
  with fallback to the list-level selected record.
- Create form gains a matching requires_auth checkbox (default: true),
  passed as 5th arg to createCollection().

* feat(frontend): context rail collection improvements

- Show collection name (c.name) instead of raw UID in the sidebar
  Collections section; names now come from the updated
  EntryCollectionMembership API response.
- Fix horizontal overflow on long collection names: coll-name gains
  overflow:hidden + text-overflow:ellipsis + white-space:nowrap +
  min-width:0; coll-row gets overflow:hidden.
- Single-entry 'Add to collection' UI: dropdown + button inside the
  Collections rail section lets users add the current entry to any
  non-default collection without multi-selecting. After add, the
  membership list refreshes automatically.
- Collections section now shows even when entryCollections is empty,
  as long as non-default collections exist (so the add form is
  accessible for un-membered entries).
- Bulk 'Add to collection' now uses the target collection's
  default_visibility_bits instead of hardcoded 2 (Users only).
- Both bulk and single-entry dropdowns filter out slug='_default_'
  to match the backend rejection in add_entry_to_collection_handler.
- Collections list is now fetched on archiveId change (not just on
  bulk mode entry) so it is available for single-entry mode too.

* build(frontend): update static assets

* feat(frontend): public collection link UX + app-styled public page

CollectionsView:
- When a collection has requires_auth=false, show a read-only URL input
  and Copy button below the auth checkbox so the public link is
  immediately discoverable. The input auto-selects on focus so manual
  copy always works. Copy button tries navigator.clipboard.writeText
  first; falls back to execCommand('copy') for HTTP deployments where
  the Clipboard API is unavailable in non-secure contexts.

PublicCollectionPage:
- Rewritten to use the app's CSS classes and variables instead of
  bare inline styles, so it visually matches the main archive UI.
  Dark topbar (.pub-coll-topbar) with brand + collection name, paper
  background body, entry list via .coll-entries-list / .coll-entry-row /
  .coll-entry-info / .coll-entry-kind — the same classes used in the
  authenticated Collections view.

styles.css:
- .coll-public-link-row / -wrap / -input / .coll-copy-btn for the
  new link field in CollectionsView detail pane.
- .pub-coll-* classes for the public page layout and typography.

* feat(core): add get_collection_by_slug; scope search to active collection

- get_collection_by_slug(): new function mirroring get_collection_by_uid
  but matching on slug, used to resolve the _default_ collection when no
  ?collection param is supplied.
- SearchEntriesQuery gains collection_id: Option<i64>. When set, the
  search SQL adds an EXISTS subquery that checks collection_entries cef
  for both membership (cef.collection_id = ?) and visibility bits in
  that specific collection — preventing cross-collection visibility
  leaks where an entry is public in one collection but private in the
  current one. Without collection_id the original cross-collection
  visibility fallback is kept.

* feat(server): collection-scoped entries/search with uniform auth gate

All entry listing and search now route through the active collection:

list_entries (?collection=<uid>|main|<omitted>):
- Resolves the target collection; omitted or 'main' resolves to _default_.
- Checks requires_auth on that collection; gates auth conditionally.
- Returns list_entries_for_collection() — same EntrySummary shape.

search_entries_handler:
- Same collection resolution + conditional auth as list_entries.
- Sets search_query.collection_id so SQL scopes membership + visibility
  to the specific collection, not cross-collection fallback.

list_collections_handler:
- Dropped require_auth() — collection summaries (name/slug/uid/
  requires_auth/default_visibility_bits) are public metadata needed for
  the guest collection-switcher dropdown.

Tests:
- list_collections_requires_auth → list_collections_is_public (200).
- list_entries_requires_auth and search coverage still pass.

* feat(frontend): integrate collection switching into main Archive view

Replaces the standalone /c/:archiveId/:collUid public page with a
unified main-view approach where all collection logic lives at /.

URL param:
- ?collection=<uid> selects a collection; omitted or 'main' = default.
- 'main' is normalized to null in parseLocation() so the dropdown shows
  'All entries' and the URL stays clean.

Collection switcher (Topbar):
- Dropdown always visible (guests need it to navigate public collections).
- Non-default collections only (All entries = no param = _default_).
- Guest selecting an auth-required collection calls onSignInClick().
- handleCollectionChange checks both named and _default_ requires_auth
  before proceeding, redirecting guests to login if needed.

listCollections fetched for all users (guests too) since the endpoint
is now public; used to populate the switcher without auth.

Public-session mode (authenticated state, no currentUser):
- Auth gate: fetchArchives() + fetchEntries() with collection param;
  401 falls through to login, 200 proceeds as guest.
- auth:expired suppressed when !currentUser.
- fetchEntryDetail skipped; ContextRail shows entry summary + sign-in prompt.
- ContextRail selection effect skips tag/collection API calls.
- runs/tags not fetched in guest mode.
- Child row expansion disabled in EntryRow (hasChildren = false).

api.js:
- fetchEntries/searchEntries both thread ?collection=<uid> to server.

Deleted: PublicCollectionPage.jsx, PublicCollectionPage.stories.jsx,
copy-link UI from CollectionsView, pub-coll-*/copy-link CSS.

* build(frontend): update static assets

* feat(core): add is_entry_publicly_accessible; checks entry+parent vs public collections

* feat(server): allow guests to fetch detail/children/artifacts for public entries

* feat(frontend): guest collection dropdown filtering; public entry detail without auth wall

* build(frontend): update static assets

* test(server): public entry detail/artifact/children contract for guests

* feat(server): filter auth-required collections from guest list_collections response
2026-07-24 20:16:17 +02:00
.github/workflows ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
crates feat(collections): public collections, per-collection auth, UX improvements (#34) 2026-07-24 20:16:17 +02:00
docker chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
docs feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
frontend feat(collections): public collections, per-collection auth, UX improvements (#34) 2026-07-24 20:16:17 +02:00
modules/nixos fix: address 6 code-review findings 2026-06-29 21:30:45 +02:00
vendor feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
.dockerignore chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
.gitignore ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
AGENTS.md feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
ARCHIVR-MENTAL-MODEL.md feat: add capturing w/ parent-child entries (#32) 2026-07-21 21:57:29 +02:00
Cargo.lock feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
Cargo.toml feat: share videos to TV (Chromecast + AirPlay) (#33) 2026-07-23 20:12:22 +02:00
docker-compose.yml chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
Dockerfile docker: add uBlock Lite + ISDCAC extensions, mirror flake.nix parity 2026-07-12 17:09:17 +02:00
flake.lock chore: yt-dlp 2026.06.09 → 2026.07.04 (#25) 2026-07-13 14:58:29 +02:00
flake.nix feat: uBlock Origin Lite + cookie consent extension + reader mode + ad placeholder cleanup (#21) 2026-07-08 23:26:48 +02:00

archivr

An open-source self-hosted archiving tool. Work in progress.

  • Archiving
    • Archiving media files from social media platforms
      • YouTube Videos
      • YouTube Playlists
      • YouTube Channels
      • Twitter Videos
      • Instagram
      • Facebook
      • TikTok
      • Reddit
      • Snapchat
      • YouTube Posts (postponed)
    • Archiving local files
    • Archiving Twitter Tweets, Threads, and Articles
    • Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
      • URLs
      • Google Drive
      • Dropbox
      • OneDrive
      • (Some of these could be postponed for later.)
    • Archive web pages (HTML, CSS, JS, images)
    • Archiving emails (???)
      • Gmail
      • Outlook
      • Yahoo Mail
  • Management
    • Deduplication
    • Tagging system
    • Search functionality
    • Categorization
    • Metadata extraction and storage
  • User Interface
    • Web-based UI
      • Authentication and login
      • Archive setup
      • Browse and view entries
      • Tag management and filtering
      • Search entries
      • View archive runs
      • Capture dialog
      • User settings and API tokens
      • Admin panel
  • Backup and Sync
    • Cloud backup (AWS S3, Google Cloud Storage)
    • Local backup

Motivation

There are two driving factors behind this project:

  • In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
  • I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.

This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.

Archive Inputs

archivr archive <path> currently accepts three kinds of inputs:

  • Local files via file://...
  • Direct platform URLs
  • Platform shorthand inputs such as tweet:..., yt:..., or instagram:...

Running Archivr

Archivr currently ships as two binaries:

  • archivr
    • The CLI for creating and writing to one archive.
    • Use this for init and archive.
  • archivr-server
    • The web server for reading one or more existing archives through the browser UI.
    • Use this after archives already exist.

With Nix, run the CLI with:

nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf

Run the web server with:

nix run .#archivr-server -- ./archivr-server.toml

The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.

Example:

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"

Then open:

http://127.0.0.1:8080

When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.

Security and Deployment

archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.

Changing the bind address

You can set the bind address in your TOML config:

# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"

Or override it with the ARCHIVR_BIND environment variable:

ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml

If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:

warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.

When will auth be added?

Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.

Supported Platforms

  • Local files: file:///absolute/path/to/file.ext
  • YouTube media: individual videos/shorts, playlists, and channels; standard URLs or shorthand video inputs. Playlists and channels archive as a container entry with each video stored as a child entry beneath it.
  • X/Twitter media from Tweets: normal Tweet URLs or the tweet:media:ID shorthand
  • X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in raw_tweets/)
  • Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to yt-dlp

Video quality and audio-only downloads

When capturing via the web UI, entering a URL for a yt-dlp-backed source (YouTube, Instagram, TikTok, Facebook, Reddit, Snapchat, X media) triggers a metadata probe via GET /api/archives/:id/captures/probe. The quality selector then shows only the heights actually available in that video plus Best quality (default). An Audio only option is appended whenever the probe confirms an audio track exists. UI behaviour by probe outcome:

qualities has_audio UI shows
["1080p", "720p", …] true Best / heights / Audio only
["1080p", …] false Best / heights
[] true Audio only (pre-selected, no Best option)
[] false "No media detected"
probe fails (502) — picker hidden, capture still submittable

The POST /api/archives/:id/captures endpoint accepts an optional quality field: "best", "audio", or any "NNNp" height string:

{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }

"audio" selects the most efficient native audio track without transcoding: Opus/WebM is preferred (smallest at equivalent quality), then AAC/M4A, then whatever yt-dlp considers best. The saved file's extension matches the native format (.webm for Opus, .m4a for AAC, etc.) — no ffmpeg re-encode, no size inflation. Any "NNNp" height is accepted; the server builds the yt-dlp format selector with an unconditional /best fallback so the download succeeds even if the exact height is unavailable. Omitting quality or passing "best" downloads at the highest available quality. Anything else is rejected with HTTP 400.

The probe endpoint (GET /api/archives/:id/captures/probe?locator=…) requires auth and returns 200 with:

{ "has_video": true,  "has_audio": true,  "qualities": ["1080p", "720p", "480p"] }
{ "has_video": false, "has_audio": true,  "qualities": [] }
{ "has_video": false, "has_audio": false, "qualities": [] }

has_video: false, has_audio: false means yt-dlp found no downloadable tracks (e.g. a tweet with no media). A 502 means yt-dlp itself failed (transient network error, rate-limit, unsupported extractor) — treat as inconclusive, not "no media."

YouTube playlists and channels

Capturing a YouTube playlist or channel URL creates a container entry for the playlist or channel, with each video archived as a child entry beneath it. Before downloading, the capture UI probes each video to fetch available quality options, letting you set quality per-video or apply a single quality to the whole playlist.

Incremental sync: When re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container entry is reused.

Excluding individual videos: In the expanded per-video list, each video has a remove button (×) to exclude it from the current capture. Removed videos are not downloaded; the rest proceed normally.

Hosting on NixOS

The flake exposes a nixosModules.default output. Add it to your system flake and enable the service:

# flake.nix (your system flake)
{
  inputs.archivr.url = "github:thegeneralist/archivr";

  outputs = { nixpkgs, archivr, ... }: {
    nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
      modules = [
        archivr.nixosModules.default
        {
          services.archivr-server = {
            enable = true;
            # listenAddress defaults to "127.0.0.1" (loopback only)
            # port defaults to 8080
            archives = [
              { id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
              { id = "work";     label = "Work";     path = "/srv/archivr/work/.archivr"; }
            ];
          };
        }
      ];
    };
  };
}

The module:

  • Creates an archivr system user and group.
  • Generates the TOML config from your options and stores the auth database under /var/lib/archivr-server/ (persists across upgrades).
  • Runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp, etc.). Archive directories are whitelisted for read-write access.
  • Restarts automatically on failure.

openFirewall — set to true to open the TCP port derived from bind. Only needed when binding to a non-loopback address:

services.archivr-server = {
  listenAddress = "0.0.0.0";
  port = 8080;            # explicit, though 8080 is the default
  openFirewall = true;
};

Archive directories must be readable and writable by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.

Hosting with Docker

A Dockerfile and docker-compose.yml are provided for self-hosting without Nix.

Quickstart

  1. Copy the example config and edit it:

    mkdir config
    cp docker/config.example.toml config/archivr-server.toml
    # edit config/archivr-server.toml — set archive id, label, and archive_path
    
  2. Initialize each archive on the persistent data volume before the first start. The image includes the archivr CLI for this purpose:

    docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"
    

    This creates /data/archives/main/.archivr/ with the metadata the server requires. A bare mkdir is not enough — the server reads name and store_path files that only archivr init writes.

  3. Start the server:

    docker compose up -d
    

    Then open http://localhost:8080.

Volumes

Mount Purpose
./config (read-only) Directory containing archivr-server.toml
archivr-data named volume Auth database (/data/archivr-auth.sqlite) and archive directories

Important: auth_db_path must be set explicitly in archivr-server.toml to a path on the writable data volume (e.g. /data/archivr-auth.sqlite). If left unset, the server defaults to writing the auth database next to the config file — which is on the read-only /config mount and will fail. The example config sets this correctly.

Twitter/X archiving

Supply a cookies file inside the config volume and set ARCHIVR_TWITTER_CREDENTIALS_FILE in docker-compose.yml:

environment:
  ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt

Building the image locally

docker build -t archivr-server .

The image compiles the Rust binary in a separate build stage so only the runtime dependencies (Chromium, Node.js, Python) land in the final layer.

Supported Shorthand Inputs

  • YouTube video/short media:
    • yt:video/ID
    • youtube:video/ID
    • yt:short/ID
    • yt:shorts/ID
    • youtube:shorts/ID
  • X/Twitter tweet JSON content:
    • tweet:ID
    • x:tweet:ID
    • x:x:ID
    • twitter:x:ID
    • twitter:tweet:ID
  • X/Twitter media/video download:
    • tweet:media:ID
  • X/Twitter thread JSON content:
    • x:thread:ID
    • twitter:thread:ID
  • Other platform shorthands:
    • instagram:ID
    • facebook:ID
    • tiktok:ID
    • reddit:ID
    • snapchat:ID

Environment Variables

  • ARCHIVR_BIND
    • Optional.
    • Overrides the bind address from the TOML config. Useful in Docker where you need 0.0.0.0:8080 without editing the config file. Default: 127.0.0.1:8080.
  • ARCHIVR_STATIC_DIR
    • Optional.
    • Path to the directory of pre-built frontend assets served by the web UI. Set automatically by the Nix wrapper and the Docker image. When running from source with cargo run, falls back to crates/archivr-server/static.
  • ARCHIVR_YT_DLP
    • Optional.
    • Overrides the yt-dlp binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
  • ARCHIVR_SINGLE_FILE
    • Optional.
    • Overrides the single-file binary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
  • ARCHIVR_CHROME
    • Optional.
    • Overrides the Chromium/Chrome executable passed to single-file via --browser-executable-path. Set automatically by the Nix wrapper and the Docker image. Default: chromium.
  • ARCHIVR_CHROME_ARGS
    • Optional.
    • Space-separated extra flags appended to Chromium's --browser-args. The Docker image sets this to --no-sandbox because Chromium refuses to run as root without it. Leave unset when running natively (Nix, Linux desktop). A --window-size=1920,1080 is always passed to provide a realistic desktop viewport (so responsive @media rules and styles are evaluated and preserved correctly). Supply your own --window-size=... here to override.
  • ARCHIVR_TWITTER_CREDENTIALS_FILE
    • Required for tweet/thread scraping inputs such as tweet:ID and x:thread:ID.
    • Must point to a cookies file for the vendored scraper.
  • ARCHIVR_TWEET_SCRAPER
    • Optional.
    • Overrides the tweet scraper script path. Default: vendor/twitter/scrape_user_tweet_contents.py.
  • ARCHIVR_TWEET_PYTHON
    • Optional.
    • Overrides the Python executable used to run the tweet scraper. Default: python3.

Current Limitations

  • Arbitrary http:// or https:// URLs that return HTML are archived as self-contained single-file HTML snapshots via single-file-cli (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires single-file and a Chromium binary on PATH, or the ARCHIVR_SINGLE_FILE / ARCHIVR_CHROME env vars set.
  • Local files currently need to be passed as file://... paths.

License

This project is licensed under the MIT License. See the LICENSE file for details.