1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-07-21 18:55:36 +02:00
No description
Find a file
TheGeneralist d610d37793
feat: entry previews (#24)
* feat: entry previews (video, tweet, article, iframe, image, audio bar)

- Add PreviewPanel dispatch hub routing by entity_kind + artifact extension
- VideoPreview: HTML5 <video> for YouTube/Instagram/TikTok/Reddit/X posts
- TweetPreview: tweet card, thread, and X article renderer (ported from x-article-renderer)
- IframePreview: sandboxed iframe for SingleFile web pages and PDFs
- ImagePreview: image viewer with click-to-open-fullsize
- AudioBar: persistent fixed-bottom player (Spotify-style) that survives entry navigation
- Lift entryDetail to App.jsx, shared between PreviewPanel and ContextRail
- 3-column layout (workspace | 300px preview | 340px rail) when preview active
- Stale-guard fixes: seq incremented before early returns in all async effects
- handleRearchive: capture startSeq/entryUid at call time, guard every async resume
- TweetPreview: reset loading/error/tweets before early-return branches

* feat: entry previews — tweet/thread/article/video/audio/image/iframe/pdf

- PreviewModal: modal overlay with new-tab link (↗) and keyboard close
- PreviewPanel: routes by entity_kind + primary_media extension to the
  correct viewer (tweet/video/audio/pdf/html/image/fallback)
- TweetPreview: full X-style tweet, thread, and article renderer with
  local artifact map for archived media (CDN fallback)
- AudioBar: persistent fixed bottom player, triggered via ContextRail
  Play button; body.has-audio-bar pads content above it
- VideoPreview, IframePreview, ImagePreview: inline viewers
- PreviewPage: standalone /preview/:archiveId/:entryUid route
- ContextRail: Play/Preview buttons; isAudio/isPreviewable detection
- App.jsx: preview modal state, currentAudio state, preview route guard,
  has-audio-bar body class effect
- routes.rs: CSP updated (media-src self blob https; frame-ancestors self;
  Google Fonts + external images/scripts whitelisted)
- styles.css: preview modal, tweet-wrap scroll (min-height:0), audio bar
  body padding, newtab button, preview panel flex layout

* style: tighten article/tweet preview spacing

- aMeta padding: 14px → 10px
- article title marginBottom: 10px → 8px
- aAuthorRow marginBottom: 10px → 8px
- bH1 top margin: 20px → 16px
- bH2 top margin: 18px → 14px
- bHr margin: 20px → 14px
- .preview-tweet-wrap padding: 20px → 12px (already committed)

More content visible above the fold in both modal and standalone views.

* feat: tweet/article preview quality pass

HTML entities: decode &gt; &lt; &amp; etc. on sliced segments only
(entity offsets index the stored string; decoding before slicing shifts them)

Image lightbox: click any tweet/thread/article image to open full-screen
viewer; cmd+click follows <a> to open in new tab; arrow-key + ‹ › nav;
Escape closes; 1/N counter; ↗ open-in-new-tab link

Multi-image grid: 2 photos → side-by-side (180px rows); 3 → left spans
both rows; 4 → 2×2 (140px rows); single image unchanged

Empty media grid ghost: build photos/videoItems arrays first, render
.mediaGrid div only when at least one item resolved (advisory: never gate
on raw media.length when map items can all return null)

QT indicator: ↻ QT badge on tweet.is_quote_status === true entries

Modal shrinks for short content: height: 88vh → max-height: 88vh;
.preview-modal-body gets max-height: calc(88vh - 52px) so long threads
still scroll (advisory: don't rely on flex:1 once parent has no fixed height)

Video scrollbar leak: .preview-modal-body overflow: auto → hidden; each
child (tweet-wrap, video-wrap, iframe) manages its own scroll surface

Styled thin scrollbar on .preview-tweet-wrap (matches workspace rail)

ArticleRenderer: cover image and body images are lightbox-clickable;
opts thread through renderBlocksJSX → renderBlockJSX → renderAtomicJSX

* fix: move artifact fetch to api.js; stop Escape propagation from lightbox

- Export fetchEntryArtifacts(archiveId, entryUid, indices) from api.js
  using Promise.all + getJson (follows project convention: all /api calls
  go through api.js, never inline fetch in components)
- TweetPreview: import fetchEntryArtifacts, replace inline Promise.all
- MediaLightbox keydown handler: stopPropagation + preventDefault for
  Escape/ArrowLeft/ArrowRight so the parent PreviewModal window listener
  does not also fire and close the modal behind the lightbox

* fix: iframe/page preview height chain and toolbar UX

Problem: changing .preview-modal from height to max-height broke iframe
previews - <iframe style='flex:1'> needs a concrete ancestor height, which
max-height alone doesn't supply when content is shorter than the cap.

Fix - CSS:
  .preview-modal--full { height: 88vh } applied to non-tweet modals
  .preview-modal--full .preview-modal-body { max-height: none }
  .preview-iframe-toolbar span: remove text-transform/letter-spacing
    (was uppercasing the URL/title in shouty caps)

Fix - PreviewModal: className adds --full when entity_kind is not
  tweet/tweet_thread; tweet previews keep shrink-to-fit behavior.

Fix - PreviewPanel: pass title + original_url from summary to IframePreview
  for both HTML and PDF; wrappers use flex:1/minHeight:0 not height:100%.

Fix - IframePreview:
  - Accept title + originalUrl props; show originalUrl in toolbar (falls
    back to artifact src only when original_url absent); show title above
    URL when available
  - flex:1 + minHeight:0 instead of height:100% on the wrap div
  - Single unified layout for page + pdf (both just show the iframe)

* feat: expand t.co links; linkify bare URLs in tweet and article text

Frontend:
- resolveEntityBounds: try multiple candidate strings in order (u.url
  first, since that's the t.co short URL that appears in full_text)
- normalizeUrlAnn: multi-candidate search; href = expanded > url,
  display = display_url > expanded > url
- linkifyText(): regex linkifier for entity-less bare URLs; trims
  trailing punctuation [.,;:!?)] before linking; used in both
  renderTweetTextJSX and renderInlineJSX including their early-return
  paths (anns.length === 0) that previously bypassed linkification
- renderInlineJSX: fix mention href mention.name → screen_name;
  replace t.co segment text with url.display when entity covers it

Scraper (vendor/twitter/scrape_user_tweet_contents.py):
- extract_tweet_data: when note_tweet text is used, pull urls/mentions/
  hashtags/symbols from note_result.entity_set (correct indices for the
  note text); keep media from legacy.entities (no note media downloads)

* feat: server-side t.co resolver + frontend augmentation

Server (routes.rs):
  POST /api/util/resolve-tco — unauthenticated, accepts JSON array of
  https://t.co/<alphanumeric> URLs only (strict regex validation, no SSRF
  via input), capped at 50 per batch, 3 s timeout, redirect(Policy::none)
  so the server only ever touches t.co itself. HEAD first, GET fallback if
  HEAD returns no Location. Location sanitized to http/https only —
  javascript:/data:/etc. fall back to the original t.co.

api.js:
  resolveTcoUrls(urls) — project-convention wrapper for the new endpoint.
  Returns {} on failure (callers degrade gracefully to bare t.co links).

TweetPreview.jsx:
  After tweet data loads, per-tweet range-based coverage detection:
  builds covered [start,end) from existing entity fromIndex/toIndex or
  indices fallback, then finds regex matches whose span is NOT covered.
  Resolves unique uncovered t.co URLs via resolveTcoUrls(), synthesises
  one entity per occurrence (with exact fromIndex/toIndex so normalizeUrlAnn
  gets correct bounds even for duplicate t.co URLs in the same tweet).
  Augments entities.urls before setTweets() so all rendering paths
  see expanded URLs.

* feat: suppress rendered media attachment URLs from tweet text

renderTweetTextJSX now accepts skipSpans=[] as third param.
Skip-span boundaries are added to the pts split set so a trailing
media t.co inside a plain segment still gets isolated and suppressed—
not re-linked by linkifyText. Early return only when both anns and
skipSpans are empty.

TweetCard computes mediaSkipSpans after building photos/videoItems:
for each rawMedia item whose src resolved (photo src match; any video),
resolveEntityBounds(m, ft, m.url) gives the precise [s,e] span using
indices/fromIndex first, indexOf fallback—then the span is passed to
renderTweetTextJSX so the t.co attachment URL is silently dropped.
2026-07-12 16:23:04 +02:00
.github/workflows ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
crates feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
docker chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
docs feat: add video quality selection for yt-dlp captures (#17) 2026-07-05 20:42:41 +02:00
frontend feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
modules/nixos fix: address 6 code-review findings 2026-06-29 21:30:45 +02:00
vendor feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
.dockerignore chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
.gitignore ci: weekly yt-dlp auto-update workflow 2026-07-03 18:28:18 +02:00
AGENTS.md docs: add AGENTS.md repository guide for coding agents 2026-07-03 15:40:49 +02:00
ARCHIVR-MENTAL-MODEL.md docs: remove completed planning docs and stale superpowers plan 2026-06-23 17:24:24 +02:00
Cargo.lock feat: entry previews (#24) 2026-07-12 16:23:04 +02:00
Cargo.toml security: add per-IP sliding-window rate limit on POST /api/auth/login 2026-06-29 20:06:41 +02:00
docker-compose.yml chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
Dockerfile chore: add Dockerfile, docker-compose, and Docker hosting docs (#12) 2026-06-30 14:35:43 +02:00
flake.lock chore: nix flake update — yt-dlp 2026.03.17 → 2026.06.09 2026-07-03 18:17:18 +02:00
flake.nix feat: uBlock Origin Lite + cookie consent extension + reader mode + ad placeholder cleanup (#21) 2026-07-08 23:26:48 +02:00
NEXT.md docs(settings): mark Track 7 done in NEXT.md 2026-06-26 17:20:06 +02:00

archivr

An open-source self-hosted archiving tool. Work in progress.

  • Archiving
    • Archiving media files from social media platforms
      • YouTube Videos
      • YouTube Playlists
      • YouTube Channels
      • Twitter Videos
      • Instagram
      • Facebook
      • TikTok
      • Reddit
      • Snapchat
      • YouTube Posts (postponed)
    • Archiving local files
    • Archiving Twitter Tweets, Threads, and Articles
    • Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
      • URLs
      • Google Drive
      • Dropbox
      • OneDrive
      • (Some of these could be postponed for later.)
    • Archive web pages (HTML, CSS, JS, images)
    • Archiving emails (???)
      • Gmail
      • Outlook
      • Yahoo Mail
  • Management
    • Deduplication
    • Tagging system
    • Search functionality
    • Categorization
    • Metadata extraction and storage
  • User Interface
    • Web-based UI
      • Authentication and login
      • Archive setup
      • Browse and view entries
      • Tag management and filtering
      • Search entries
      • View archive runs
      • Capture dialog
      • User settings and API tokens
      • Admin panel
  • Backup and Sync
    • Cloud backup (AWS S3, Google Cloud Storage)
    • Local backup

Motivation

There are two driving factors behind this project:

  • In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
  • I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.

This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.

Archive Inputs

archivr archive <path> currently accepts three kinds of inputs:

  • Local files via file://...
  • Direct platform URLs
  • Platform shorthand inputs such as tweet:..., yt:..., or instagram:...

Running Archivr

Archivr currently ships as two binaries:

  • archivr
    • The CLI for creating and writing to one archive.
    • Use this for init and archive.
  • archivr-server
    • The web server for reading one or more existing archives through the browser UI.
    • Use this after archives already exist.

With Nix, run the CLI with:

nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf

Run the web server with:

nix run .#archivr-server -- ./archivr-server.toml

The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.

Example:

[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"

Then open:

http://127.0.0.1:8080

When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.

Security and Deployment

archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.

Changing the bind address

You can set the bind address in your TOML config:

# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"

Or override it with the ARCHIVR_BIND environment variable:

ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml

If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:

warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.

When will auth be added?

Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.

Supported Platforms

  • Local files: file:///absolute/path/to/file.ext
  • YouTube media: standard video/short URLs, plus shorthand video inputs
  • X/Twitter media from Tweets: normal Tweet URLs or the tweet:media:ID shorthand
  • X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in raw_tweets/)
  • Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to yt-dlp

Video quality and audio-only downloads

When capturing via the web UI, entering a URL for a yt-dlp-backed source (YouTube, Instagram, TikTok, Facebook, Reddit, Snapchat, X media) triggers a metadata probe via GET /api/archives/:id/captures/probe. The quality selector then shows only the heights actually available in that video plus Best quality (default). An Audio only option is appended whenever the probe confirms an audio track exists. UI behaviour by probe outcome:

qualities has_audio UI shows
["1080p", "720p", …] true Best / heights / Audio only
["1080p", …] false Best / heights
[] true Audio only (pre-selected, no Best option)
[] false "No media detected"
probe fails (502) picker hidden, capture still submittable

The POST /api/archives/:id/captures endpoint accepts an optional quality field: "best", "audio", or any "NNNp" height string:

{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }

"audio" selects the most efficient native audio track without transcoding: Opus/WebM is preferred (smallest at equivalent quality), then AAC/M4A, then whatever yt-dlp considers best. The saved file's extension matches the native format (.webm for Opus, .m4a for AAC, etc.) — no ffmpeg re-encode, no size inflation. Any "NNNp" height is accepted; the server builds the yt-dlp format selector with an unconditional /best fallback so the download succeeds even if the exact height is unavailable. Omitting quality or passing "best" downloads at the highest available quality. Anything else is rejected with HTTP 400.

The probe endpoint (GET /api/archives/:id/captures/probe?locator=…) requires auth and returns 200 with:

{ "has_video": true,  "has_audio": true,  "qualities": ["1080p", "720p", "480p"] }
{ "has_video": false, "has_audio": true,  "qualities": [] }
{ "has_video": false, "has_audio": false, "qualities": [] }

has_video: false, has_audio: false means yt-dlp found no downloadable tracks (e.g. a tweet with no media). A 502 means yt-dlp itself failed (transient network error, rate-limit, unsupported extractor) — treat as inconclusive, not "no media."

Hosting on NixOS

The flake exposes a nixosModules.default output. Add it to your system flake and enable the service:

# flake.nix (your system flake)
{
  inputs.archivr.url = "github:thegeneralist/archivr";

  outputs = { nixpkgs, archivr, ... }: {
    nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
      modules = [
        archivr.nixosModules.default
        {
          services.archivr-server = {
            enable = true;
            # listenAddress defaults to "127.0.0.1" (loopback only)
            # port defaults to 8080
            archives = [
              { id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
              { id = "work";     label = "Work";     path = "/srv/archivr/work/.archivr"; }
            ];
          };
        }
      ];
    };
  };
}

The module:

  • Creates an archivr system user and group.
  • Generates the TOML config from your options and stores the auth database under /var/lib/archivr-server/ (persists across upgrades).
  • Runs under a hardened systemd unit (ProtectSystem = strict, NoNewPrivileges, PrivateTmp, etc.). Archive directories are whitelisted for read-write access.
  • Restarts automatically on failure.

openFirewall — set to true to open the TCP port derived from bind. Only needed when binding to a non-loopback address:

services.archivr-server = {
  listenAddress = "0.0.0.0";
  port = 8080;            # explicit, though 8080 is the default
  openFirewall = true;
};

Archive directories must be readable and writable by the archivr user. Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.

Hosting with Docker

A Dockerfile and docker-compose.yml are provided for self-hosting without Nix.

Quickstart

  1. Copy the example config and edit it:

    mkdir config
    cp docker/config.example.toml config/archivr-server.toml
    # edit config/archivr-server.toml — set archive id, label, and archive_path
    
  2. Initialize each archive on the persistent data volume before the first start. The image includes the archivr CLI for this purpose:

    docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"
    

    This creates /data/archives/main/.archivr/ with the metadata the server requires. A bare mkdir is not enough — the server reads name and store_path files that only archivr init writes.

  3. Start the server:

    docker compose up -d
    

    Then open http://localhost:8080.

Volumes

Mount Purpose
./config (read-only) Directory containing archivr-server.toml
archivr-data named volume Auth database (/data/archivr-auth.sqlite) and archive directories

Important: auth_db_path must be set explicitly in archivr-server.toml to a path on the writable data volume (e.g. /data/archivr-auth.sqlite). If left unset, the server defaults to writing the auth database next to the config file — which is on the read-only /config mount and will fail. The example config sets this correctly.

Twitter/X archiving

Supply a cookies file inside the config volume and set ARCHIVR_TWITTER_CREDENTIALS_FILE in docker-compose.yml:

environment:
  ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt

Building the image locally

docker build -t archivr-server .

The image compiles the Rust binary in a separate build stage so only the runtime dependencies (Chromium, Node.js, Python) land in the final layer.

Supported Shorthand Inputs

  • YouTube video/short media:
    • yt:video/ID
    • youtube:video/ID
    • yt:short/ID
    • yt:shorts/ID
    • youtube:shorts/ID
  • X/Twitter tweet JSON content:
    • tweet:ID
    • x:tweet:ID
    • x:x:ID
    • twitter:x:ID
    • twitter:tweet:ID
  • X/Twitter media/video download:
    • tweet:media:ID
  • X/Twitter thread JSON content:
    • x:thread:ID
    • twitter:thread:ID
  • Other platform shorthands:
    • instagram:ID
    • facebook:ID
    • tiktok:ID
    • reddit:ID
    • snapchat:ID

Environment Variables

  • ARCHIVR_BIND
    • Optional.
    • Overrides the bind address from the TOML config. Useful in Docker where you need 0.0.0.0:8080 without editing the config file. Default: 127.0.0.1:8080.
  • ARCHIVR_STATIC_DIR
    • Optional.
    • Path to the directory of pre-built frontend assets served by the web UI. Set automatically by the Nix wrapper and the Docker image. When running from source with cargo run, falls back to crates/archivr-server/static.
  • ARCHIVR_YT_DLP
    • Optional.
    • Overrides the yt-dlp binary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
  • ARCHIVR_SINGLE_FILE
    • Optional.
    • Overrides the single-file binary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
  • ARCHIVR_CHROME
    • Optional.
    • Overrides the Chromium/Chrome executable passed to single-file via --browser-executable-path. Set automatically by the Nix wrapper and the Docker image. Default: chromium.
  • ARCHIVR_CHROME_ARGS
    • Optional.
    • Space-separated extra flags appended to Chromium's --browser-args. The Docker image sets this to --no-sandbox because Chromium refuses to run as root without it. Leave unset when running natively (Nix, Linux desktop). A --window-size=1920,1080 is always passed to provide a realistic desktop viewport (so responsive @media rules and styles are evaluated and preserved correctly). Supply your own --window-size=... here to override.
  • ARCHIVR_TWITTER_CREDENTIALS_FILE
    • Required for tweet/thread scraping inputs such as tweet:ID and x:thread:ID.
    • Must point to a cookies file for the vendored scraper.
  • ARCHIVR_TWEET_SCRAPER
    • Optional.
    • Overrides the tweet scraper script path. Default: vendor/twitter/scrape_user_tweet_contents.py.
  • ARCHIVR_TWEET_PYTHON
    • Optional.
    • Overrides the Python executable used to run the tweet scraper. Default: python3.

Current Limitations

  • Arbitrary http:// or https:// URLs that return HTML are archived as self-contained single-file HTML snapshots via single-file-cli (requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requires single-file and a Chromium binary on PATH, or the ARCHIVR_SINGLE_FILE / ARCHIVR_CHROME env vars set.
  • Local files currently need to be passed as file://... paths.

License

This project is licensed under the MIT License. See the LICENSE file for details.