- capture.rs: add archive_id: Option<&str> to perform_capture; when Some,
call font_extractor::extract_and_rewrite before hashing HTML, register
each font as a deduplicated blob + 'font' artifact
- main.rs: pass None as archive_id (CLI keeps fonts embedded)
- routes.rs: add GET /api/archives/:id/blobs/:sha256 (serve_blob handler),
pass Some(&archive_id) to perform_capture in capture_handler,
add ApiError::internal constructor
- fix(hash): hash raw bytes instead of lossy UTF-8; add hash_bytes
- feat(database): add get_blob_by_sha256 lookup
- feat(font_extractor): extract embedded font data-URIs from archived HTML
- --user-agent: realistic Chrome UA so servers don't block headless string
- --browser-args=[--disable-web-security, --user-data-dir]: lets single-file
inline fonts from any cross-origin CDN (e.g. fonts.gstatic.com) regardless
of ACAO headers; user-data-dir required for --disable-web-security to take
effect in newer Chromium builds (otherwise silently ignored)
Font fidelity:
- --browser-wait-delay=2000: Cloudflare Fonts injects @font-face CSS after
HTML parse; the font hook needs extra time to see it after networkidle2
- --remove-unused-fonts=false: preserve @font-face rules even when fonts
haven't rendered yet (font-display:swap, off-screen text)
- --remove-alternative-fonts=false: preserve unicode-range subsets instead
of stripping them as 'alternatives'
ES module viewer error fix:
- Write a user script (sf-strip-scripts.js) that listens for
single-file-on-before-capture-start and removes all <script> elements
(except application/ld+json) from the live DOM before serialization.
Scripts still execute during capture for CSS fidelity; none end up in
the saved file, so no data:-URL base ES module resolution errors.
Defaults that were destroying CSS fidelity:
- --remove-unused-styles=true: strips CSS nesting rules (site uses & selector)
and any rule targeting JS-applied classes
- --remove-alternative-medias=true: deletes @media blocks that don't match
the capture viewport, breaking responsive layout
- --block-scripts=true: prevents JS from applying classes before CSS snapshot
All three now set to false.
http::download now returns (hash, extension, Option<title_hint>).
Title is derived from Content-Disposition filename header (RFC 5987
filename*= preferred over plain filename=), falling back to the last
path segment of the final URL after redirects (percent-decoded).
capture.rs Source::Url arm passes the title through to record_media_entry
instead of None, so entries like 'Facharbeit.pdf' or 'facharbeit' appear
in the Title column instead of the raw entry_uid.
Adds percent_decode(), title_from_content_disposition(), title_from_url()
helpers with 10 new unit tests (122 total, all green).
- perform_capture: fetch yt-dlp metadata (separate --dump-json call)
before download; derive local title from file:// path filename
- Compute entry_title before record_media_entry call; pass it through
instead of the temporary None placeholder
- record_tweet_entry: read tweet JSON before entry creation so title
can be set on NewEntry; add tweet_metadata_from_json helper
- Add PlatformMetadata struct and generate_entry_title() covering all 13
Source variants (YouTubeVideo, YouTubePlaylist, YouTubeChannel, X, Tweet,
TweetThread, Instagram, Facebook, TikTok, Reddit, Snapchat, Local, Other)
- Add downloader/metadata.rs: extract_from_ytdlp_json() parses yt-dlp
--dump-json output into PlatformMetadata; extracts Reddit subreddit
from webpage_url via regex
- Add ytdlp::fetch_metadata(): separate --dump-json invocation that does
not interfere with the actual download call
- Extend record_media_entry() with title: Option<String> param; wire
yt-dlp metadata fetch + local filename extraction in perform_capture
- Restructure record_tweet_entry() to read tweet JSON before entry
creation; extract title via tweet_metadata_from_json()
- 16 title-generation unit tests + 6 metadata extraction unit tests
- Delete NEXT.md (all six tracks implemented)
- Delete docs/superpowers/plans/track1 plan (track complete)
- Trim ARCHIVR-MENTAL-MODEL.md Key Documents table to only living files
- Add optional `bind` field to ServerRegistry (TOML + ARCHIVR_BIND env var)
- Default bind address remains 127.0.0.1:8080; non-loopback prints a warning
- Add route security classification comment block (READ/ADMIN/WRITE/STATIC)
- Add Security and Deployment section to docs/README.md
- Replace vague auth note in ARCHIVR-MENTAL-MODEL.md with concrete model description
- Add three registry tests covering bind field round-trip and defaults
expand_shorthand_to_url handled instagram:/facebook:/tiktok: etc. but
not yt:/youtube:. yt-dlp doesn't know the yt: scheme, so shorthands
like yt:video/ID failed with 'Unsupported url scheme'.
Add YouTube cases for video/, short/, shorts/, playlist/, channel/,
c/, user/, and @handle. Full https:// URLs pass through unchanged.
Add tests pinning all new expansions.
archivr_server now includes the Python scraper and sets
ARCHIVR_TWEET_PYTHON and ARCHIVR_TWEET_SCRAPER, matching what the
archivr CLI wrapper already does. Without this, browser capture of
tweets failed with 'No such file or directory' because the server
resolved the scraper relative to cwd.
- Add archivr-core/src/capture.rs: Source enum, all capture helpers,
fail_run (replaces process::exit fail_archive_and_exit), and
perform_capture() as the public entry point
- Slim archivr-cli/src/main.rs to 103 lines: thin adapter over
archivr_core::capture::perform_capture
- Move all source-classification and tweet-entry tests from CLI to
archivr-core/src/capture.rs (they test core logic, belong in core)
- Add POST /api/archives/:archive_id/captures route in archivr-server:
validates non-empty locator (400), resolves archive (404), delegates
to perform_capture, returns {run_uid, status}
- Add capture dialog to browser UI: dialog markup in index.html,
showModal/submit/cancel/loading/error wiring in app.js, dialog
styles in styles.css
- 77 tests pass, clean build (0 warnings)
- 5 tag routes (list tree, create, assign, remove, search with tag=)
- Tags nav, tag tree view, tag filter badge
- Entry tag pills with remove, assign-tag form in context rail