Add --window-size=1920,1080 to the Chromium flags passed via --browser-args.
This makes the existing --remove-unused-styles=false and
--remove-alternative-medias=false effective for real @media rules
and responsive styles (headless default is small).
Also document in ARCHIVR_CHROME_ARGS that users can override by
supplying their own --window-size in the env var.
CaptureDialog:
- Replace single textarea with multi-row inputs; + button adds rows
- Submit fires all pending rows in parallel, dialog stays open/usable
- Polling intervals live on a persistent ref (not cleared on close) so
toasts fire even after the dialog is dismissed
- archiveId stored per item at submit time; page-refresh reconnect uses
it.archiveId instead of the possibly-null prop
- Completed rows flash green then self-remove; failed rows show inline
error + retry button
- Cancel becomes Close while jobs are in flight
ToastStack (new component):
- Fixed bottom-right overlay with spring-in animation
- Error toast: truncated locator, View error / Hide toggle expanding
full error_text in a monospace pre block
- Auto-dismisses after 7 s; timer pauses while detail is expanded
RunsView:
- Failed rows are clickable and expand a full-width detail row showing
error_summary in a scrollable monospace block
capture.rs (archivr-core):
- Staging dir is now "{millis}-{uuid}" — parallel captures in the same
millisecond can no longer collide on temp paths
- create_archive_run moved before URL Content-Type probe so every
attempt appears in /runs regardless of outcome
- Probe failures now call create_archive_run_item with source_metadata
fallback then fail_run, recording error_text on the item and
error_summary on the run with correct failed_count
styles.css:
- Capture dialog: header row, multi-row layout, status dots, spinner,
add-row dashed button, per-row error text
- Toast stack: fixed overlay, error card with coloured left border,
monospace detail expansion
- Run error rows: clickable hover tint, expand hint chevron, detail pre
8082895 removed the ytdlp_metadata_json fetch, local_filename_title
derivation, and entry_title computation, replacing the title arg with
a None stub. Restores all three blocks so YouTube, Instagram, Reddit,
TikTok, Facebook, Snapchat, X, and local file entries receive proper
titles again.
Store how many bytes of each entry's artifacts are already on disk from
an earlier entry (content-addressed blob deduplication means shared
blobs are only stored once).
Design
------
- Add `cached_bytes INTEGER NOT NULL DEFAULT 0` to `archived_entries`
- Precompute at capture time via `database::refresh_entry_cached_bytes`
called after all artifacts are saved for every capture path
(web page, generic URL, tweet, yt-dlp/local)
- One-time migration in `initialize_schema`: detects missing column via
PRAGMA table_info, ALTERs the table, then back-fills all existing rows
with the correlated subquery
- `database::cascade_cached_bytes_after_delete` ready for when entry
deletion is implemented; designed to run asynchronously after the
delete is acknowledged to the user
- `cached_bytes` included in `EntrySummary` and all four SELECT paths
(list_root_entries, search_entries, list_entries_for_collection,
entries_for_tag) via the shared ENTRY_SELECT_COLS constant
Frontend
--------
- `EntryRow` shows a `% cached` sub-line under the size when non-zero,
with a tooltip showing the raw cached byte count
- No separate API endpoint or extra fetch — value rides in the existing
entries list response at zero extra query cost per read
* chore: add Dockerfile, docker-compose, and Docker docs
- Multi-stage Dockerfile: Rust builder stage + debian:bookworm-slim runtime
with Chromium, Node/single-file-cli, Python venv (yt-dlp + twitter-api-client)
- docker-compose.yml: wires ARCHIVR_BIND, config volume, and persistent data volume
- docker/config.example.toml: annotated TOML template for Docker deployments
- docs/README.md: add Hosting with Docker section; add ARCHIVR_BIND and
ARCHIVR_STATIC_DIR to the Environment Variables reference
* fix: address code review issues with Docker setup
- .gitignore: whitelist Dockerfile, docker-compose.yml, docker/ so they
are actually tracked (the * catch-all was silently dropping them)
- Dockerfile: build and ship the archivr CLI alongside archivr-server so
users can run `archivr init` inside the container on first setup
- docker/config.example.toml: fix archive_path to point at the .archivr
subdirectory that archivr init creates (not the parent directory), which
is what read_archive_paths expects
- docs/README.md: replace the bare mkdir quickstart step with
`archivr init`, explain why mkdir is insufficient; add a callout that
auth_db_path must be set explicitly to a writable path when the config
mount is read-only
* fix: address second round of Docker review issues
Chromium sandbox (P2):
- singlefile.rs: add ARCHIVR_CHROME_ARGS env var (space-separated flags
appended to Chromium's --browser-args JSON array); Dockerfile sets it
to --no-sandbox because Chromium refuses to start as root without it
Store-path outside volume (P1):
- README: pass explicit absolute store-path as the second positional arg
to `archivr init` so the blob store lands on /data instead of the
container layer (CLI default is ./.archivr/store, resolved from cwd,
which is / with no WORKDIR set)
ENTRYPOINT vs CMD (P2):
- Dockerfile: switch from ENTRYPOINT to CMD so `docker compose run
archivr archivr init …` overrides the full command instead of being
appended to the server invocation
ffmpeg missing (P2):
- Dockerfile: add ffmpeg to the apt-get install block (required by
yt-dlp --merge-output-format mp4 for bestvideo+bestaudio streams)
Node version (P2):
- Dockerfile: replace Debian bookworm's nodejs (18.x) with Node 20 via
the NodeSource setup script (single-file-cli declares engines.node >=20)
Build context secrets (P2):
- Add .dockerignore excluding config/ and docker/ from the build context
so runtime secrets (e.g. twitter-cookies.txt) are never sent to the builder
- Whitelist .dockerignore in .gitignore
docs:
- README: document ARCHIVR_CHROME_ARGS in the Environment Variables section
* fix: third round of Docker review issues
Rust toolchain (P1):
- Dockerfile: bump builder from rust:1.87 to rust:1.88; time@0.3.51,
time-core@0.1.9, and time-macros@0.2.30 (present in Cargo.lock) all
require MSRV 1.88, so the real cargo build --release step was failing
single-file-cli wait mode (P2):
- singlefile.rs: replace --browser-wait-until=networkidle2 with
networkAlmostIdle; the single-file-cli option only accepts
InteractiveTime/networkIdle/networkAlmostIdle/load/domContentLoaded
(verified in options.js); networkidle2 is a Puppeteer concept that the
CLI does not recognise, causing silent fallback to the earliest state
and incomplete captures. networkAlmostIdle is the closest equivalent
(<=2 open connections, matching Puppeteer's networkidle2 semantics)
Build context size (P3):
- .dockerignore: add target/, frontend/node_modules/, frontend/dist/;
these can reach 1.4G+ after a local dev build and are never read by
the Dockerfile, so sending them to the builder wastes time and memory
- capture.rs: add archive_id: Option<&str> to perform_capture; when Some,
call font_extractor::extract_and_rewrite before hashing HTML, register
each font as a deduplicated blob + 'font' artifact
- main.rs: pass None as archive_id (CLI keeps fonts embedded)
- routes.rs: add GET /api/archives/:id/blobs/:sha256 (serve_blob handler),
pass Some(&archive_id) to perform_capture in capture_handler,
add ApiError::internal constructor
- fix(hash): hash raw bytes instead of lossy UTF-8; add hash_bytes
- feat(database): add get_blob_by_sha256 lookup
- feat(font_extractor): extract embedded font data-URIs from archived HTML
- --user-agent: realistic Chrome UA so servers don't block headless string
- --browser-args=[--disable-web-security, --user-data-dir]: lets single-file
inline fonts from any cross-origin CDN (e.g. fonts.gstatic.com) regardless
of ACAO headers; user-data-dir required for --disable-web-security to take
effect in newer Chromium builds (otherwise silently ignored)
Font fidelity:
- --browser-wait-delay=2000: Cloudflare Fonts injects @font-face CSS after
HTML parse; the font hook needs extra time to see it after networkidle2
- --remove-unused-fonts=false: preserve @font-face rules even when fonts
haven't rendered yet (font-display:swap, off-screen text)
- --remove-alternative-fonts=false: preserve unicode-range subsets instead
of stripping them as 'alternatives'
ES module viewer error fix:
- Write a user script (sf-strip-scripts.js) that listens for
single-file-on-before-capture-start and removes all <script> elements
(except application/ld+json) from the live DOM before serialization.
Scripts still execute during capture for CSS fidelity; none end up in
the saved file, so no data:-URL base ES module resolution errors.
Defaults that were destroying CSS fidelity:
- --remove-unused-styles=true: strips CSS nesting rules (site uses & selector)
and any rule targeting JS-applied classes
- --remove-alternative-medias=true: deletes @media blocks that don't match
the capture viewport, breaking responsive layout
- --block-scripts=true: prevents JS from applying classes before CSS snapshot
All three now set to false.
http::download now returns (hash, extension, Option<title_hint>).
Title is derived from Content-Disposition filename header (RFC 5987
filename*= preferred over plain filename=), falling back to the last
path segment of the final URL after redirects (percent-decoded).
capture.rs Source::Url arm passes the title through to record_media_entry
instead of None, so entries like 'Facharbeit.pdf' or 'facharbeit' appear
in the Title column instead of the raw entry_uid.
Adds percent_decode(), title_from_content_disposition(), title_from_url()
helpers with 10 new unit tests (122 total, all green).
- perform_capture: fetch yt-dlp metadata (separate --dump-json call)
before download; derive local title from file:// path filename
- Compute entry_title before record_media_entry call; pass it through
instead of the temporary None placeholder
- record_tweet_entry: read tweet JSON before entry creation so title
can be set on NewEntry; add tweet_metadata_from_json helper
- Add PlatformMetadata struct and generate_entry_title() covering all 13
Source variants (YouTubeVideo, YouTubePlaylist, YouTubeChannel, X, Tweet,
TweetThread, Instagram, Facebook, TikTok, Reddit, Snapchat, Local, Other)
- Add downloader/metadata.rs: extract_from_ytdlp_json() parses yt-dlp
--dump-json output into PlatformMetadata; extracts Reddit subreddit
from webpage_url via regex
- Add ytdlp::fetch_metadata(): separate --dump-json invocation that does
not interfere with the actual download call
- Extend record_media_entry() with title: Option<String> param; wire
yt-dlp metadata fetch + local filename extraction in perform_capture
- Restructure record_tweet_entry() to read tweet JSON before entry
creation; extract title via tweet_metadata_from_json()
- 16 title-generation unit tests + 6 metadata extraction unit tests
expand_shorthand_to_url handled instagram:/facebook:/tiktok: etc. but
not yt:/youtube:. yt-dlp doesn't know the yt: scheme, so shorthands
like yt:video/ID failed with 'Unsupported url scheme'.
Add YouTube cases for video/, short/, shorts/, playlist/, channel/,
c/, user/, and @handle. Full https:// URLs pass through unchanged.
Add tests pinning all new expansions.
- Add archivr-core/src/capture.rs: Source enum, all capture helpers,
fail_run (replaces process::exit fail_archive_and_exit), and
perform_capture() as the public entry point
- Slim archivr-cli/src/main.rs to 103 lines: thin adapter over
archivr_core::capture::perform_capture
- Move all source-classification and tweet-entry tests from CLI to
archivr-core/src/capture.rs (they test core logic, belong in core)
- Add POST /api/archives/:archive_id/captures route in archivr-server:
validates non-empty locator (400), resolves archive (404), delegates
to perform_capture, returns {run_uid, status}
- Add capture dialog to browser UI: dialog markup in index.html,
showModal/submit/cancel/loading/error wiring in app.js, dialog
styles in styles.css
- 77 tests pass, clean build (0 warnings)