- capture.rs: add archive_id: Option<&str> to perform_capture; when Some,
call font_extractor::extract_and_rewrite before hashing HTML, register
each font as a deduplicated blob + 'font' artifact
- main.rs: pass None as archive_id (CLI keeps fonts embedded)
- routes.rs: add GET /api/archives/:id/blobs/:sha256 (serve_blob handler),
pass Some(&archive_id) to perform_capture in capture_handler,
add ApiError::internal constructor
- fix(hash): hash raw bytes instead of lossy UTF-8; add hash_bytes
- feat(database): add get_blob_by_sha256 lookup
- feat(font_extractor): extract embedded font data-URIs from archived HTML
- --user-agent: realistic Chrome UA so servers don't block headless string
- --browser-args=[--disable-web-security, --user-data-dir]: lets single-file
inline fonts from any cross-origin CDN (e.g. fonts.gstatic.com) regardless
of ACAO headers; user-data-dir required for --disable-web-security to take
effect in newer Chromium builds (otherwise silently ignored)
Font fidelity:
- --browser-wait-delay=2000: Cloudflare Fonts injects @font-face CSS after
HTML parse; the font hook needs extra time to see it after networkidle2
- --remove-unused-fonts=false: preserve @font-face rules even when fonts
haven't rendered yet (font-display:swap, off-screen text)
- --remove-alternative-fonts=false: preserve unicode-range subsets instead
of stripping them as 'alternatives'
ES module viewer error fix:
- Write a user script (sf-strip-scripts.js) that listens for
single-file-on-before-capture-start and removes all <script> elements
(except application/ld+json) from the live DOM before serialization.
Scripts still execute during capture for CSS fidelity; none end up in
the saved file, so no data:-URL base ES module resolution errors.
Defaults that were destroying CSS fidelity:
- --remove-unused-styles=true: strips CSS nesting rules (site uses & selector)
and any rule targeting JS-applied classes
- --remove-alternative-medias=true: deletes @media blocks that don't match
the capture viewport, breaking responsive layout
- --block-scripts=true: prevents JS from applying classes before CSS snapshot
All three now set to false.
http::download now returns (hash, extension, Option<title_hint>).
Title is derived from Content-Disposition filename header (RFC 5987
filename*= preferred over plain filename=), falling back to the last
path segment of the final URL after redirects (percent-decoded).
capture.rs Source::Url arm passes the title through to record_media_entry
instead of None, so entries like 'Facharbeit.pdf' or 'facharbeit' appear
in the Title column instead of the raw entry_uid.
Adds percent_decode(), title_from_content_disposition(), title_from_url()
helpers with 10 new unit tests (122 total, all green).
- perform_capture: fetch yt-dlp metadata (separate --dump-json call)
before download; derive local title from file:// path filename
- Compute entry_title before record_media_entry call; pass it through
instead of the temporary None placeholder
- record_tweet_entry: read tweet JSON before entry creation so title
can be set on NewEntry; add tweet_metadata_from_json helper