Rust toolchain (P1):
- Dockerfile: bump builder from rust:1.87 to rust:1.88; time@0.3.51,
time-core@0.1.9, and time-macros@0.2.30 (present in Cargo.lock) all
require MSRV 1.88, so the real cargo build --release step was failing
single-file-cli wait mode (P2):
- singlefile.rs: replace --browser-wait-until=networkidle2 with
networkAlmostIdle; the single-file-cli option only accepts
InteractiveTime/networkIdle/networkAlmostIdle/load/domContentLoaded
(verified in options.js); networkidle2 is a Puppeteer concept that the
CLI does not recognise, causing silent fallback to the earliest state
and incomplete captures. networkAlmostIdle is the closest equivalent
(<=2 open connections, matching Puppeteer's networkidle2 semantics)
Build context size (P3):
- .dockerignore: add target/, frontend/node_modules/, frontend/dist/;
these can reach 1.4G+ after a local dev build and are never read by
the Dockerfile, so sending them to the builder wastes time and memory
Chromium sandbox (P2):
- singlefile.rs: add ARCHIVR_CHROME_ARGS env var (space-separated flags
appended to Chromium's --browser-args JSON array); Dockerfile sets it
to --no-sandbox because Chromium refuses to start as root without it
Store-path outside volume (P1):
- README: pass explicit absolute store-path as the second positional arg
to `archivr init` so the blob store lands on /data instead of the
container layer (CLI default is ./.archivr/store, resolved from cwd,
which is / with no WORKDIR set)
ENTRYPOINT vs CMD (P2):
- Dockerfile: switch from ENTRYPOINT to CMD so `docker compose run
archivr archivr init …` overrides the full command instead of being
appended to the server invocation
ffmpeg missing (P2):
- Dockerfile: add ffmpeg to the apt-get install block (required by
yt-dlp --merge-output-format mp4 for bestvideo+bestaudio streams)
Node version (P2):
- Dockerfile: replace Debian bookworm's nodejs (18.x) with Node 20 via
the NodeSource setup script (single-file-cli declares engines.node >=20)
Build context secrets (P2):
- Add .dockerignore excluding config/ and docker/ from the build context
so runtime secrets (e.g. twitter-cookies.txt) are never sent to the builder
- Whitelist .dockerignore in .gitignore
docs:
- README: document ARCHIVR_CHROME_ARGS in the Environment Variables section
- capture.rs: add archive_id: Option<&str> to perform_capture; when Some,
call font_extractor::extract_and_rewrite before hashing HTML, register
each font as a deduplicated blob + 'font' artifact
- main.rs: pass None as archive_id (CLI keeps fonts embedded)
- routes.rs: add GET /api/archives/:id/blobs/:sha256 (serve_blob handler),
pass Some(&archive_id) to perform_capture in capture_handler,
add ApiError::internal constructor
- fix(hash): hash raw bytes instead of lossy UTF-8; add hash_bytes
- feat(database): add get_blob_by_sha256 lookup
- feat(font_extractor): extract embedded font data-URIs from archived HTML
- --user-agent: realistic Chrome UA so servers don't block headless string
- --browser-args=[--disable-web-security, --user-data-dir]: lets single-file
inline fonts from any cross-origin CDN (e.g. fonts.gstatic.com) regardless
of ACAO headers; user-data-dir required for --disable-web-security to take
effect in newer Chromium builds (otherwise silently ignored)
Font fidelity:
- --browser-wait-delay=2000: Cloudflare Fonts injects @font-face CSS after
HTML parse; the font hook needs extra time to see it after networkidle2
- --remove-unused-fonts=false: preserve @font-face rules even when fonts
haven't rendered yet (font-display:swap, off-screen text)
- --remove-alternative-fonts=false: preserve unicode-range subsets instead
of stripping them as 'alternatives'
ES module viewer error fix:
- Write a user script (sf-strip-scripts.js) that listens for
single-file-on-before-capture-start and removes all <script> elements
(except application/ld+json) from the live DOM before serialization.
Scripts still execute during capture for CSS fidelity; none end up in
the saved file, so no data:-URL base ES module resolution errors.
Defaults that were destroying CSS fidelity:
- --remove-unused-styles=true: strips CSS nesting rules (site uses & selector)
and any rule targeting JS-applied classes
- --remove-alternative-medias=true: deletes @media blocks that don't match
the capture viewport, breaking responsive layout
- --block-scripts=true: prevents JS from applying classes before CSS snapshot
All three now set to false.
http::download now returns (hash, extension, Option<title_hint>).
Title is derived from Content-Disposition filename header (RFC 5987
filename*= preferred over plain filename=), falling back to the last
path segment of the final URL after redirects (percent-decoded).
capture.rs Source::Url arm passes the title through to record_media_entry
instead of None, so entries like 'Facharbeit.pdf' or 'facharbeit' appear
in the Title column instead of the raw entry_uid.
Adds percent_decode(), title_from_content_disposition(), title_from_url()
helpers with 10 new unit tests (122 total, all green).
- perform_capture: fetch yt-dlp metadata (separate --dump-json call)
before download; derive local title from file:// path filename
- Compute entry_title before record_media_entry call; pass it through
instead of the temporary None placeholder
- record_tweet_entry: read tweet JSON before entry creation so title
can be set on NewEntry; add tweet_metadata_from_json helper
- Add PlatformMetadata struct and generate_entry_title() covering all 13
Source variants (YouTubeVideo, YouTubePlaylist, YouTubeChannel, X, Tweet,
TweetThread, Instagram, Facebook, TikTok, Reddit, Snapchat, Local, Other)
- Add downloader/metadata.rs: extract_from_ytdlp_json() parses yt-dlp
--dump-json output into PlatformMetadata; extracts Reddit subreddit
from webpage_url via regex
- Add ytdlp::fetch_metadata(): separate --dump-json invocation that does
not interfere with the actual download call
- Extend record_media_entry() with title: Option<String> param; wire
yt-dlp metadata fetch + local filename extraction in perform_capture
- Restructure record_tweet_entry() to read tweet JSON before entry
creation; extract title via tweet_metadata_from_json()
- 16 title-generation unit tests + 6 metadata extraction unit tests
expand_shorthand_to_url handled instagram:/facebook:/tiktok: etc. but
not yt:/youtube:. yt-dlp doesn't know the yt: scheme, so shorthands
like yt:video/ID failed with 'Unsupported url scheme'.
Add YouTube cases for video/, short/, shorts/, playlist/, channel/,
c/, user/, and @handle. Full https:// URLs pass through unchanged.
Add tests pinning all new expansions.
- Add archivr-core/src/capture.rs: Source enum, all capture helpers,
fail_run (replaces process::exit fail_archive_and_exit), and
perform_capture() as the public entry point
- Slim archivr-cli/src/main.rs to 103 lines: thin adapter over
archivr_core::capture::perform_capture
- Move all source-classification and tweet-entry tests from CLI to
archivr-core/src/capture.rs (they test core logic, belong in core)
- Add POST /api/archives/:archive_id/captures route in archivr-server:
validates non-empty locator (400), resolves archive (404), delegates
to perform_capture, returns {run_uid, status}
- Add capture dialog to browser UI: dialog markup in index.html,
showModal/submit/cancel/loading/error wiring in app.js, dialog
styles in styles.css
- 77 tests pass, clean build (0 warnings)