1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 21:03:17 +02:00
Commit graph

213 commits

Author SHA1 Message Date
253f779216 Add YouTube subtitles, local transcription, self-updating yt-dlp/Deno, X Article and thread titles
- Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles)
- Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error
- Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only)
- Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s)
- Settings > Instance > yt-dlp: status and in-app update without restart
- X Article titles from article.title, with idempotent startup backfill
- Thread title generation (single and bulk) with per-provider cheap models
- Per-provider title model settings in Settings > Instance
- Docs, mental model, AGENTS.md and transcription spec updated
2026-10-05 19:39:51 +02:00
4f3b2968b6 build: build frontend in Nix, remove static from git tracking
Add frontendDeps (FOD) and frontendStatic derivations to flake.nix so
that nix build handles the Vite bundle automatically. The archivr_server
derivation now copies from ${frontendStatic} instead of
${./crates/archivr-server/static}.

crates/archivr-server/static/ is excluded from git tracking (.gitignore);
the aarch64-darwin node_modules hash is included. Other systems use a
placeholder hash — run `nix build 2>&1 | grep 'got:'` to obtain theirs.

Update AGENTS.md and docs/README.md to reflect the new workflow.
2026-10-04 22:53:49 +02:00
699eb7f62d feat: reorder child entries with a configurable role permission
- Persist sibling order for child entries (archived_entries.position):
  backfilled to the previous archived_at order, new children appended
  (initial playlist order kept; sync appends after user order).
- PUT /api/archives/:id/entries/:uid/children/order takes the full child
  UID list in one IMMEDIATE transaction (401 guest, 403 role not allowed,
  404 unknown or invisible parent, 400 unless exactly the current children).
- Reorder permission is a role mask in the auth instance settings
  (reorder_children_role_bits, default Admin|Owner). Only the Owner can
  change it, built-in or custom roles, never Guest. /api/auth/me and login
  expose can_reorder_children. Settings PATCH is now one IMMEDIATE txn.
- Main page: drag handle for mouse/trackpad on viewports >640px, up/down
  arrows otherwise (pure CSS, arrows are the fallback), Alt+Up/Down
  everywhere; optimistic with revert. Settings > Instance > Permissions
  for the Owner (read-only for admins).
- Fix: renaming a child entry now updates its row immediately.
- Remove Storybook (deps, config, stories, docs).
2026-10-04 20:55:37 +02:00
8fc6754e28
feat(capture): route x.com status URLs to Source::Tweet
Full https://x.com/{user}/status/{id} URLs now go through the Twitter
JSON scraper (Source::Tweet) instead of yt-dlp (Source::X), consistent
with the x:tweet:ID shorthand. Non-status x.com URLs are unchanged.

tweet_id_from_archive_path extended to extract the numeric ID from full
status URLs, not just colon-separated shorthands.

x:media:ID remains the explicit opt-in for yt-dlp media-only capture.
2026-10-04 18:33:48 +02:00
c68469783c
chore(docker): pin yt-dlp and twitter-api-client versions to match Nix
- yt-dlp: pip install yt-dlp → yt-dlp==2026.8.19 (matches flake.nix pin)
- twitter-api-client: unpinned → 0.10.22 (matches flake.nix)
- docs: add Docker subsection to 'Keeping yt-dlp fresh' explaining that
  rebuilding the image is the Docker equivalent of the Nix/self-update
  paths, and that there is no in-container yt-dlp update path
2026-10-04 17:55:46 +02:00
archivr-qa
9b2ecb8d7b
fix: guard detailMatchesSelection against null detail/selectedEntry
When both detail and selectedEntry are null (initial render or mid-navigation
while a detail fetch is in flight), detail?.summary?.entry_uid and
selectedEntry?.entry_uid both evaluate to undefined, making the equality
true and causing a crash on detail.latest_summary.

Add selectedEntry?.entry_uid != null guard so the condition short-circuits
false whenever there is no real selection.
2026-08-28 20:25:41 +02:00
79ac44834e
feat: capture, summaries, search, and yt-dlp reliability (#38)
* ui: show spinner for pending captures

* feat(core): add text capture path with title + Markdown/plain body

- Add downloader/text.rs module with save() function that stages and hashes text content
- Support text/markdown and text/plain MIME types with .md and .txt extensions
- Add perform_text_capture() function for capturing user-supplied text
- Validates title (non-empty, max 500 chars) and body (non-empty, max 2 MiB)
- Creates blob records and entries with source_kind='text', entity_kind='document'
- Includes comprehensive unit tests for markdown, plain text, and validation

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(server): add POST /api/archives/:archive_id/captures/text

- Add CaptureTextBody struct for title, body, and optional MIME type
- Implement capture_text_handler with validation for empty fields and MIME type
- Route text submissions to perform_text_capture() in background
- Reuse existing capture job tracking and polling infrastructure
- Default MIME type to text/markdown when not specified
- Include route tests covering happy path, validation, auth, and error cases

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(frontend): add text-capture form to CaptureDialog

- Add submitTextCapture API client function with same error handling as submitCapture
- Create makeTextItem() factory for text capture state
- Implement CaptureTextRow component with title, body textarea, and MIME selector
- Add 'Add text' button in capture dialog toolbar
- Update handleArchive to filter and route text submissions
- Modify submitBgJob to detect and submit text items via submitTextCapture
- Skip probe and conflict checks for text items
- Reuse job tracking and batch settlement for text captures

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(core): add entry_summaries schema + summarizer trait/providers

Per-entry LLM summaries as a regenerable child record, not a column on
archived_entries and not an on-disk artifact: an entry may carry several
summaries (one per provider/model/prompt version), any of which can be
discarded and recomputed. Generation is manual-only — nothing in capture.rs
calls into this module.

- database.rs: entry_summaries table + index, EntrySummaryRecord, and
  upsert/update/find/latest helpers mirroring the capture_jobs style.
  provider_model is stored as '' rather than NULL because SQLite treats
  NULLs as distinct inside a UNIQUE index, which would stop the CLI
  providers (no model) from ever deduping on the cache key.
- summarizer.rs: SummaryProvider trait with four implementations —
  Anthropic Messages API, OpenAI-compatible chat completions, `claude -p`
  and `codex exec -`. Configuration comes from env vars only (never TOML),
  matching how yt-dlp / single-file / tweet-scraper are resolved, which
  also keeps API keys out of anything the archive persists.
- archive.rs: EntryDetail gains latest_summary, populated by one extra
  LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row
  rather than duplicating it.

Implementation notes:
- No tokio in core. CLI timeouts are enforced structurally: stdout is
  drained on its own thread and handed back over a channel so the calling
  thread can recv_timeout and kill an overrunning child; stdin is written
  on a third thread so a 48 KB prompt cannot deadlock against a child
  waiting for us to read.
- HTML is reduced with regex rather than a parser: html5ever is not in the
  tree, and a model tolerates imperfect whitespace. Paired tags are spelled
  out per tag because Rust's regex engine has no backreferences by design.
- reqwest is declared with only the `blocking` feature here, so bodies are
  serialized via .body(value.to_string()) instead of widening the
  workspace dependency for .json().
- input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one
  hashing primitive; the content is truncated to 48 KB *before* hashing so
  the cache key describes exactly the bytes the model saw.

Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for
one JSON shape is a poor trade, so the two halves that can actually break
are tested directly — request-body builders and response parsers — leaving
only reqwest's own transport uncovered. Plus schema idempotency, cache-key
dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML
and tweet extraction, output normalization, and the CLI runner's stdin
round-trip, timeout kill, and nonzero-exit paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(server): add GET/POST /api/archives/:id/entries/:uid/summary

GET is read-only and gated exactly like entry detail, so a guest can read a
summary only for an entry whose content they could already read. POST
requires ROLE_USER, matching capture / tags / patch / rearchive; no auth
roles change.

Both the provider config and the content extraction resolve on the request
thread, before spawn_blocking. That is what lets a missing env var come back
as a synchronous 400 naming the exact variable, and an unsummarizable
artifact (video, audio) as a 400 saying so, rather than becoming a
background job the caller must poll only to learn about a config typo.

The pending row is claimed before spawning so the 202 can name a summary_uid
the client can poll immediately. summarize_entry owns the
pending → running → completed/failed transitions for that same row — the
cache key is identical, so both upserts resolve to one row — leaving the
handler to catch only the case where it fails before recording anything.
When !force and an identical cache key already completed, the existing row
comes back as a 200 with no new work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(frontend): render Summary rail section + provider selector

New "Summary" rail section between the URL/Preview controls and .meta-list.
A completed summary renders as bold tl;dr, body paragraph, tag chips, and a
provider · model footer; missing or failed shows Generate; pending/running
shows an inline spinner and polls GET every 1500 ms until terminal.

- api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps
  ApiError's { "error": ... } body so the missing-env-var message reaches
  the user verbatim rather than as a bare status code.
- ContextRail.jsx: state seeds from detail.latest_summary so the section
  renders immediately on selection. Polling is anchored on the summary
  status rather than started inside the click handler, so a job still
  running when the user navigates away and back is picked up again. A
  transient poll failure is swallowed — the next tick retries, and a real
  failure arrives as status === 'failed'.
- Regenerate passes force:true only when a completed summary is already
  shown; otherwise the request can take the server's 200 cache-hit path.
- Provider choice persists in sessionStorage under archivr:summary:provider,
  with try/catch around both accessors for private-mode browsers.
- Public sessions never see the selector or the Generate button, and the
  section renders at all only when a completed summary made it through the
  server's visibility gate.
- styles.css: .rail-summary-* only; spacing and the action button reuse
  .rail-section and .rail-rearchive-btn. The spinner honours
  prefers-reduced-motion — the text alone conveys the state.
- AGENTS.md: document the summary env vars alongside the existing
  external-tool convention.

Smoke-tested end to end against a scratch archive with a seeded markdown
entry: claude_cli produced a real summary (pending → running → completed in
~11s); a local mock server exercised the openai_compatible transport and
confirmed the Bearer header, model, and system/user role split on the wire;
unconfigured providers return 400 naming the exact variable; a video entry
returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without
adding a row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): codex_cli — auto-discover binary + use --output-last-message

Two related fixes for the codex_cli summary provider:

1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but
   without it the code resolved to bare `codex` and relied on PATH.
   The ChatGPT desktop app installs codex at
   `/Applications/ChatGPT.app/Contents/Resources/codex` and does not
   put it on PATH, so users who only have the desktop app saw
   'No such file or directory' with no hint. `resolve_cli` now walks
   env override → a small set of well-known absolute paths → HOME
   /.local/bin/<bare> → bare fallback. Same treatment applied to
   claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/
   claude, HOME/.local/bin/claude).

2. Clean output. `codex exec -` writes a runtime header ("OpenAI
   Codex vX", session id, sandbox, model), the assistant reply, and a
   footer ("tokens used", replay of the reply) to stdout. The JSON
   extractor took the first '{' from the *user prompt echo* and the
   last '}' from the trailing replay, producing invalid text that
   fell through to the "raw text under summary" fallback path. Now
   uses `--output-last-message <tempfile>` and reads only the final
   assistant message. Fallback (positional prompt) uses the same
   flag. Tempfile is cleaned up on all paths, incl. spawn failure.

* fix(frontend): give text-capture row real CSS

The text row shipped with semantic classnames (`capture-text-inputs`,
`capture-text-title`, `capture-text-body`, `capture-text-mime`,
`capture-text-icon`) but no CSS rules. Falling through to the parent
`.capture-row-main` flex-row (`display: flex; align-items: center`)
meant the title, textarea, and mime-select stacked as intrinsic-width
boxes centered on the tall body, producing a layout where the body
floated to the top-right, the title box appeared BELOW it, and the mime
selector rendered as an unstyled OS dropdown.

Fix:
- `.capture-text-row .capture-row-main` uses `align-items: flex-start`
  so the leading icon and trailing × pin to the top of the block.
- `.capture-text-inputs` is now a full-width column-flex container with
  proper gaps.
- `.capture-text-title` reuses the 44px input height and typography of
  `.capture-input`; `.capture-text-body` gets a 140px min-height,
  vertical resize, and matching border/focus treatment.
- `.capture-text-mime` is styled as a small chip with a custom caret
  so it matches `.capture-quality` and stops looking like a raw
  `<select>`. Sits in a right-aligned footer under the body.
- `.capture-text-icon` gets a 44px column so it aligns with the title
  input; remove button gets a small top-margin for the same reason.

Rebuilt static bundle bumped as well (`index-BLxoi9rt.css`,
`index-CQcpPA_I.js`).

* chore(static): rebuild bundle after text-row CSS merge

* fix(core): summarize tweets + walk all tweets in a thread

Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.

Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.

Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
  role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
  `raw_tweet_json` artifacts (falling back to `primary_media` for
  archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
  markdown/html/json branches, then join thread pieces with a
  `---` separator so the model sees a real paragraph break between
  statuses instead of one flowing document.

Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.

* feat(frontend): preview text-capture entries (.md / .txt)

Text captures land as `.md` (Markdown) or `.txt` (plain) blobs, but
PreviewPanel only dispatched on video/audio/image/pdf/html extensions,
so opening a text entry hit the 'No preview available' fallback with
the raw artifact path exposed.

- New `TextPreview` component fetches the primary artifact as text,
  renders it in a monospace `<pre>` with word-wrap, and shows the
  entry title on top and the MIME as a small trailing tag. Handles
  loading/error states.
- `PreviewPanel` gains a `TEXT_EXTS` set + a branch that dispatches
  to `TextPreview` for `md` / `markdown` / `txt`.
- CSS is padded and centered to ~780px so a text note reads like a
  document rather than an edge-to-edge terminal dump.

v1 intentionally does NOT parse Markdown: keeping frontend deps at
react+react-dom only. Bump to a real Markdown renderer if we start
capturing Markdown-authored notes.

* fix(nix): pin yt-dlp from its own release + wire into server wrapper

Two independent problems, one commit:

1. Stale binary. nixpkgs-provided `pkgs.yt-dlp` on the pinned
   nixos-unstable rev is 2026.03.17 (Mar 2026). yt-dlp itself
   releases days-to-weeks, and YouTube frequently rotates the
   player-signature / client surfaces the older builds request
   (`android_vr` is the current casualty), which returns HTTP 403
   mid-download for the format specs archivr passes (`-f
   bestvideo+bestaudio/best`). Even bumping the nixpkgs input would
   leave us dependent on that channel's yt-dlp cadence.

   Fetch the upstream zipapp directly instead
   (github.com/yt-dlp/yt-dlp/releases/download/<ver>/yt-dlp), wrap so
   `python3` and `ffmpeg` are on PATH, and pin version+hash in one
   place. Bumping is: change version, replace hash from
   `nix hash file <url>`.

2. Missing pin in server wrapper. `archivr-cli` was already wrapped
   with `--set ARCHIVR_YT_DLP` + a PATH prefix; `archivr-server`
   was NOT — it only pinned single-file, chrome, and the tweet
   scraper, silently falling back to whatever `yt-dlp` the user
   happened to have on PATH. Server captures therefore inherited
   the user's (often stale) system yt-dlp regardless of the flake
   pin. Same wrapper flags now apply to both binaries.

devShell keeps `pkgs.yt-dlp` for now: the dev shell is a
convenience, not a release surface, and matching wouldn't fit in this
commit without duplicating the derivation across let-scopes.

* chore(static): rebuild bundle for round-3 fixes

* chore(nix): pin python 3.12 for yt-dlp zipapp (avoid py3.14 libffi crash on darwin/arm64)

* feat(core): resolve_yt_dlp picks the newer of pinned vs state-dir

The nix flake wrapper pins a yt-dlp via ARCHIVR_YT_DLP, but yt-dlp rots
fast — extractors break within weeks of a pin. Add a resolver that probes
`--version` on both the pinned binary and a user-installed copy under the
mutable state dir, and runs whichever is newer.

Version strings are YYYY.MM.DD, so plain string ordering is chronological.
Ties resolve toward the state dir: a user who installed it there did so
deliberately. ARCHIVR_YT_DLP_FORCE bypasses the comparison entirely, and
with no candidate at all we fall back to bare `yt-dlp` on PATH — exactly
the previous behaviour.

Resolution is cached in a OnceLock so `--version` costs one subprocess per
process, and all four inline env::var lookups now go through it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump yt-dlp from upstream releases, not nixpkgs

The flake no longer takes yt-dlp from nixpkgs; a dedicated `ytDlp`
derivation fetches the upstream release binary directly and pins both
`version` and an SRI `hash`. That makes the previous workflow inert: it
ran `nix flake update nixpkgs` and compared `nixpkgs#yt-dlp.version`
before and after, so it could churn the lockfile forever without ever
moving the version we actually ship.

The workflow now reads the pinned version straight out of the `ytDlp`
block in flake.nix, asks the GitHub API for yt-dlp's latest release tag,
short-circuits when they already match, downloads the new release to
recompute its SRI hash (required — the hash is part of the derivation's
identity, so the URL cannot be changed alone), and rewrites the three
pinned fields under a sed range address scoped to that block so sibling
pins like ublockLite and isdcac are untouched. It asserts only flake.nix
changed and that the new version appears exactly twice before opening
the PR.

* feat(cli): add `archivr yt-dlp update|status` subcommand

`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.

The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.

`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.

reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(maintainer): document summarizer, text capture, and yt-dlp lifecycle

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(readme): document LLM summaries, text notes, yt-dlp resolver + bump paths

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plan): specify X Article, image summaries, summary search

* fix: summarize X Article text

* feat: search completed summary tags

* fix: preserve X Article block order

* feat: model opt-in summary images

* feat: attach opted-in images to summaries

* feat: accept image summary requests

* feat: add summary image consent control

* docs: explain X Article, image summaries and summary search

* chore(static): rebuild bundle for summary image consent

* fix: show civilized unsupported summary errors

* chore(static): rebuild bundle for civilized summary errors

* fix: keep text previews and summary state scoped to entry

* fix: show forced yt-dlp candidate in status

* fix: infer trusted MIME for tweet images

* fix: preserve text capture bytes and hide synthetic URL

* fix: recover and preserve summary attempts

* fix: bound codex fallback and record resolved model

* fix: protect public summary diagnostics

* fix: guard summary callbacks during entry render

* fix: scope summary callbacks to selected entry

* test: cover terminal newline in text capture

* fix: preserve missing summary entry status

* test: cover tweet image summary selection

* docs: record summary lifecycle and review hardening

* chore(static): rebuild bundle for Sol review fixes

* fix: retain completed summary during regeneration display

* fix: hide superseded summary attempts

* chore(static): rebuild bundle after regeneration display fix

* fix: allow full-size text capture requests

* fix: allow escaped full-size text captures

* fix: preserve text draft whitespace in capture UI

* chore(static): rebuild bundle for text whitespace fix
2026-08-24 18:30:01 +02:00
github-actions[bot]
b4b4e67157
chore: yt-dlp 2026.07.04 → 2026.08.19 (#37)
Co-authored-by: thegeneralist01 <180094941+thegeneralist01@users.noreply.github.com>
2026-08-24 14:59:35 +02:00
bf5f95397f
fix: exclude child entries from collection listing
list_entries_for_collection was missing AND e.parent_entry_id IS NULL,
causing child entries (playlist/channel videos) to appear alongside
their parent in the flat archive view. list_root_entries and
search_entries already had this guard; the collection path did not.

Also removes garnix from flake.nix nixConfig.
2026-08-02 11:23:13 +02:00
95cd46978d
feat(core): add more Freedium domains + restrict Freedium routing to supported domains (#36)
Previously, any WebPage capture with via_freedium=true was routed
through the Freedium mirror regardless of the URL's host. This caused
non-paywall sites like borretti.me to be fetched through Freedium,
which is wrong — Freedium only knows how to handle a specific set of
publications.

Add is_freedium_supported_url() with a static allowlist of the 7 hosts
Freedium explicitly supports per its homepage announcement:

  Medium, NYT, WaPo, Bloomberg, Reuters, Economist, Financial Times

The gate matches on the bare domain or any subdomain (e.g.
towardsdatascience.medium.com). Unparseable URLs fall through to a
direct fetch. The existing freedium-mirror.cfd re-wrap guard is kept
as a belt-and-suspenders check after the new allowlist test.

Fixes: https://borretti.me/article/notes-on-managing-adhd archived via
Freedium despite not being a Medium article.
2026-07-29 17:40:07 +02:00
6a0f59ea94
docs: tighten wordmark tracking (24→10px, 0.14→0.06em) 2026-07-26 16:27:54 +02:00
eb5c774d0d
docs: scale banner to fill GitHub 838×279px display
Measured actual GitHub display: 838×279px at 1365px viewport.
Previous content (mark=216px, wordmark=132px) rendered at 83px/51px —
too small for the available space.

New target sizes calibrated to display dimensions:
  mark       216 → 285px target  →  110px at 838px display  (+32%)
  wordmark   132 → 170px target  →   55px height at display  (+35%)
  tagline     62 →  78px target  →   30px at display         (+26%)
  A in mark  160 → 210px target

Reduced inter-element gaps to keep content within 724px height;
fills ~84% of banner, leaving ~23px display margin each side.
2026-07-25 22:01:02 +02:00
cab544c76f
docs: scale up banner content for GitHub display size
At GitHub's ~840px README width, 2172×724 renders at 840×280px.
Previous element sizes (mark 130px, wordmark 80px) became 50px and 31px
at display — too small to read clearly.

Scale everything ~1.65× to fill ~80% of banner height:
  mark       130 → 216px at target  (≈ 83px at 840px display)
  wordmark    80 → 132px at target  (≈ 51px at display)
  tagline     38 →  62px at target  (≈ 24px at display)
  A in mark   96 → 160px at target

Proportional gap increases maintain the same visual rhythm.
2026-07-25 21:52:56 +02:00
13b22bc6f8
docs: redesign banners — logo mark, flat bg, no chrome artifacts
Previous version had three inelegancies visible on GitHub:
- Full-width ghost rule read as an artifact dividing the banner in half
- Top terracotta bar looked like site header chrome
- Radial gradient created compression banding at CDN quality

v5 design:
- Terracotta rounded-rect logo mark (130px) with Paper 'A' inside —
  the actual brand element per the style guide, not a floating letter
- Flat Shell / Paper background — compresses losslessly, zero banding
- No top bar, no ghost rule, no gradient gimmicks
- Same crisp Cormorant Garamond SemiBold + 2x supersample + Lanczos
- Both dark and light variants updated with identical layout
2026-07-25 21:42:49 +02:00
820260bc9c
docs: redesign banners with Cormorant Garamond + 2x supersample
Dark banner:
- Replace Georgia with Cormorant Garamond SemiBold (actual brand font)
- 2x supersample render → Lanczos downsample for crisp edges at all
  display densities
- Shell (#141d18) background with layered radial warmth
- Editorial stacked colophon layout: large 'A' / thin parchment rule /
  wide-tracked wordmark / amber italic tagline
- Ghost full-width hairline rule anchors the horizontal composition
- Terracotta dash ornaments flank the tagline (composited on separate
  RGBA layer for correct alpha blending)

Light banner (new PNG, replacing SVG):
- Same font, supersample, and layout
- Paper (#f5f0e7) background with subtle warm edge shadow
- Ink wordmark, Muted tagline — all brand palette

README: switch light-mode fallback from banner-light.svg to banner-light.png
2026-07-25 21:21:32 +02:00
1ff85f53c4
docs: redesign dark banner — editorial colophon layout
Replace app-icon-style horizontal lockup with centered stacked
composition: large standalone terracotta 'A' mark → thin parchment
rule → wide-tracked cream wordmark → amber italic tagline.

Edge-darkening vignette clears the center and frames the bokeh
atmosphere without muddying the text zone.
2026-07-25 21:01:38 +02:00
221f6cccce
docs: composite bokeh library background into dark-mode banner
Render generated-image.png (blurred library bokeh, prompt 3) as the
dark-mode README banner with logo mark, wordmark, tagline, and
terracotta top rule composited on top via Pillow.

Switch README dark-mode <source> from banner-dark.svg to banner-dark.png.
Light-mode banner stays as the clean Paper SVG.
2026-07-25 20:49:24 +02:00
0e18b294dd
docs: rewrite README with brand identity and proper project structure
- Add logo (SVG lockup), tagline, and brand-colored badges at top
- Replace dev checklist with prose Features section
- Add Architecture section with binary overview and on-disk layout
- Consolidate Supported Inputs into a table with inline shorthand examples
- Rewrite env vars as a scannable table
- Restructure into clear Quick Start → Architecture → Inputs → Config → Deployment → Development flow
- Remove personal motivation text and developer-facing in-progress notes
2026-07-25 20:36:16 +02:00
fa0340cef2
chore: add brand identity, style guide, and favicon (#35)
* Add Archivr brand identity and favicon

- Brand style guide (docs/branding/style-guide.html): logo, clear
  space, don'ts, color palette, typography, in-context mockup
- Logo assets (docs/branding/assets/): lockup SVGs for dark/light
  backgrounds, standalone mark, favicon.svg, multi-size favicon.ico
- Wire favicon into frontend/index.html (SVG + ICO fallback);
  frontend/public/ seeded so Vite copies assets on next build
- Revert display strings archivr back to Archivr in UI
  (Topbar, LoginPage, SetupPage, index.html titles, flake.nix)
- Update .gitignore to allow docs/branding/

* Build frontend with favicon links and assets

* Fix static file serving: serve all of static/ with SPA index.html fallback

Previously only /assets/* was served from disk; all other paths
(including /favicon.ico, /favicon.svg) hit the index.html fallback.
Replace nest_service("/assets") + fallback_service(index.html) with
ServeDir on the full static dir and not_found_service(index.html),
so any file that exists in static/ is served directly.
2026-07-25 15:31:25 +02:00
e1ee05bd41
feat(collections): public collections, per-collection auth, UX improvements (#34)
* feat(core): add requires_auth to collections; include name in entry-collection memberships

- Add `requires_auth INTEGER NOT NULL DEFAULT 1` column to the
  collections DDL and as an idempotent ALTER TABLE migration in
  initialize_schema (archive DB), not initialize_auth_schema.
- CollectionRecord and CollectionSummary gain `requires_auth: bool`.
- create_collection() and update_collection() accept the new field.
- get_entry_collection_memberships() now returns collection name as the
  third tuple element; EntryCollectionMembership gains a `name` field
  so the sidebar can show human-readable names instead of raw UIDs.

* feat(server): conditional auth for public collections; add requires_auth + original_url to API

- CreateCollectionBody gains requires_auth (default true).
- PatchCollectionBody gains requires_auth: Option<bool>.
- get_collection_handler: load record first, then skip auth.require_auth()
  when record.requires_auth == false so public collections are accessible
  to unauthenticated callers; caller_bits falls back to ROLE_GUEST (1)
  so only visibility_bits=3 entries are returned to guests.
- Collection JSON response includes requires_auth and each entry now
  includes original_url for use by the public collection page.
- list_collections_handler keeps require_auth (management UI).

* feat(frontend): public collection page at /c/:archiveId/:collUid

- Detect PUBLIC_COLL_ROUTE at module load time (like PREVIEW_ROUTE) and
  return <PublicCollectionPage> before any auth checks so unauthenticated
  users can view public collections without hitting the login gate.
- PublicCollectionPage fetches via getCollection() and renders the
  server-filtered entry list (no client-side bitmask filtering - the
  server already applies caller_bits=GUEST for unauthenticated requests).
  Entry titles link to original_url when present; fall back to plain
  text when original_url is null.
- api.js createCollection() gains requiresAuth param (default true),
  sent as requires_auth in the request body.
- Storybook story covers WithEntries, Empty, and LoadError states.

* feat(frontend): collections view improvements

- addVis in the 'Add entry' form now syncs to the selected collection's
  default_visibility_bits via useEffect on collDetail, so the default
  matches the collection's configured entry visibility.
- Rename 'Default visibility' label to 'Entries\' default visibility'
  in both the detail pane and the create form to distinguish it from
  the new collection-level access setting.
- Add 'Require authentication to view' checkbox in the detail pane
  backed by a PATCH to requires_auth; reads collDetail?.requires_auth
  with fallback to the list-level selected record.
- Create form gains a matching requires_auth checkbox (default: true),
  passed as 5th arg to createCollection().

* feat(frontend): context rail collection improvements

- Show collection name (c.name) instead of raw UID in the sidebar
  Collections section; names now come from the updated
  EntryCollectionMembership API response.
- Fix horizontal overflow on long collection names: coll-name gains
  overflow:hidden + text-overflow:ellipsis + white-space:nowrap +
  min-width:0; coll-row gets overflow:hidden.
- Single-entry 'Add to collection' UI: dropdown + button inside the
  Collections rail section lets users add the current entry to any
  non-default collection without multi-selecting. After add, the
  membership list refreshes automatically.
- Collections section now shows even when entryCollections is empty,
  as long as non-default collections exist (so the add form is
  accessible for un-membered entries).
- Bulk 'Add to collection' now uses the target collection's
  default_visibility_bits instead of hardcoded 2 (Users only).
- Both bulk and single-entry dropdowns filter out slug='_default_'
  to match the backend rejection in add_entry_to_collection_handler.
- Collections list is now fetched on archiveId change (not just on
  bulk mode entry) so it is available for single-entry mode too.

* build(frontend): update static assets

* feat(frontend): public collection link UX + app-styled public page

CollectionsView:
- When a collection has requires_auth=false, show a read-only URL input
  and Copy button below the auth checkbox so the public link is
  immediately discoverable. The input auto-selects on focus so manual
  copy always works. Copy button tries navigator.clipboard.writeText
  first; falls back to execCommand('copy') for HTTP deployments where
  the Clipboard API is unavailable in non-secure contexts.

PublicCollectionPage:
- Rewritten to use the app's CSS classes and variables instead of
  bare inline styles, so it visually matches the main archive UI.
  Dark topbar (.pub-coll-topbar) with brand + collection name, paper
  background body, entry list via .coll-entries-list / .coll-entry-row /
  .coll-entry-info / .coll-entry-kind — the same classes used in the
  authenticated Collections view.

styles.css:
- .coll-public-link-row / -wrap / -input / .coll-copy-btn for the
  new link field in CollectionsView detail pane.
- .pub-coll-* classes for the public page layout and typography.

* feat(core): add get_collection_by_slug; scope search to active collection

- get_collection_by_slug(): new function mirroring get_collection_by_uid
  but matching on slug, used to resolve the _default_ collection when no
  ?collection param is supplied.
- SearchEntriesQuery gains collection_id: Option<i64>. When set, the
  search SQL adds an EXISTS subquery that checks collection_entries cef
  for both membership (cef.collection_id = ?) and visibility bits in
  that specific collection — preventing cross-collection visibility
  leaks where an entry is public in one collection but private in the
  current one. Without collection_id the original cross-collection
  visibility fallback is kept.

* feat(server): collection-scoped entries/search with uniform auth gate

All entry listing and search now route through the active collection:

list_entries (?collection=<uid>|main|<omitted>):
- Resolves the target collection; omitted or 'main' resolves to _default_.
- Checks requires_auth on that collection; gates auth conditionally.
- Returns list_entries_for_collection() — same EntrySummary shape.

search_entries_handler:
- Same collection resolution + conditional auth as list_entries.
- Sets search_query.collection_id so SQL scopes membership + visibility
  to the specific collection, not cross-collection fallback.

list_collections_handler:
- Dropped require_auth() — collection summaries (name/slug/uid/
  requires_auth/default_visibility_bits) are public metadata needed for
  the guest collection-switcher dropdown.

Tests:
- list_collections_requires_auth → list_collections_is_public (200).
- list_entries_requires_auth and search coverage still pass.

* feat(frontend): integrate collection switching into main Archive view

Replaces the standalone /c/:archiveId/:collUid public page with a
unified main-view approach where all collection logic lives at /.

URL param:
- ?collection=<uid> selects a collection; omitted or 'main' = default.
- 'main' is normalized to null in parseLocation() so the dropdown shows
  'All entries' and the URL stays clean.

Collection switcher (Topbar):
- Dropdown always visible (guests need it to navigate public collections).
- Non-default collections only (All entries = no param = _default_).
- Guest selecting an auth-required collection calls onSignInClick().
- handleCollectionChange checks both named and _default_ requires_auth
  before proceeding, redirecting guests to login if needed.

listCollections fetched for all users (guests too) since the endpoint
is now public; used to populate the switcher without auth.

Public-session mode (authenticated state, no currentUser):
- Auth gate: fetchArchives() + fetchEntries() with collection param;
  401 falls through to login, 200 proceeds as guest.
- auth:expired suppressed when !currentUser.
- fetchEntryDetail skipped; ContextRail shows entry summary + sign-in prompt.
- ContextRail selection effect skips tag/collection API calls.
- runs/tags not fetched in guest mode.
- Child row expansion disabled in EntryRow (hasChildren = false).

api.js:
- fetchEntries/searchEntries both thread ?collection=<uid> to server.

Deleted: PublicCollectionPage.jsx, PublicCollectionPage.stories.jsx,
copy-link UI from CollectionsView, pub-coll-*/copy-link CSS.

* build(frontend): update static assets

* feat(core): add is_entry_publicly_accessible; checks entry+parent vs public collections

* feat(server): allow guests to fetch detail/children/artifacts for public entries

* feat(frontend): guest collection dropdown filtering; public entry detail without auth wall

* build(frontend): update static assets

* test(server): public entry detail/artifact/children contract for guests

* feat(server): filter auth-required collections from guest list_collections response
2026-07-24 20:16:17 +02:00
1af920eb63
feat: share videos to TV (Chromecast + AirPlay) (#33)
* server: add scoped media-token endpoint for Cast/AirPlay auth bypass

Chromecast and Apple TV fetch media URLs as independent HTTP clients
with no session cookie. The existing serve_artifact handler requires
auth_user.require_auth(), so those devices always received 401.

Changes:
- MediaToken struct stored in AppState (Arc<Mutex<HashMap>>), scoped to
  a single (archive_id, entry_uid, artifact_index) tuple with a 2-hour TTL
- POST /api/archives/:id/entries/:uid/artifacts/:idx/media-token
  requires an authenticated session, verifies the artifact exists,
  prunes expired tokens, mints a 43-char URL-safe token, and returns
  { url, expires_in_secs }
- serve_artifact now accepts an optional ?token= query param; a valid
  scoped token bypasses require_auth() while a missing/invalid/expired
  token falls through to the normal 401 path
- CSP script-src extended to include https://www.gstatic.com so the
  Cast sender SDK script (injected lazily by VideoPreview) is not blocked
- 4 new tests: bare-URL still 401, tokenized fetch succeeds without
  session cookie, bogus token 401, wrong-artifact-index 401

* frontend: Cast/AirPlay overlay in VideoPreview

When the user opens a video archive entry, VideoPreview now:

1. Issues a signed media token (POST .../artifacts/:idx/media-token) and
   uses the returned signed URL as <video src>. This ensures the video
   element's src is one that Cast devices and Apple TV can fetch without a
   session cookie.

2. Lazily injects the Google Cast SDK script (cast_sender.js from
   gstatic.com, now allowed by the updated CSP). Once the SDK reports
   available, a <google-cast-launcher> web component appears as an overlay
   button in the top-right corner of the video. Selecting a Cast device
   triggers loadMedia() with the signed URL and the artifact's MIME type.

3. Detects AirPlay support (webkitShowPlaybackTargetPicker on
   HTMLVideoElement) and shows an AirPlay icon button alongside Cast.
   The <video> element carries x-webkit-airplay='allow', so Safari's native
   controls also surface the AirPlay option. The explicit overlay button
   calls webkitShowPlaybackTargetPicker() for consistent placement.

Both buttons are hidden when the respective APIs are unavailable (HTTP
pages, non-Safari for AirPlay, no Cast extension/devices), so there is no
UI regression for users who don't cast.

PreviewPanel now passes contentType (derived from artifact extension) to
VideoPreview so Cast receives a correct MIME type.

New CSS: .video-tv-controls (absolute overlay), .video-tv-btn (frosted
glass icon button), .video-tv-loading (placeholder during token fetch).

* server: fix serve_artifact auth OR logic — bogus token falls back to session

Previously a request carrying ?token=<expired> was immediately rejected
with 401, even if the user held a valid session cookie. This broke
logged-in browser playback after the 2-hour signed-URL window expired,
because VideoPreview uses the signed URL as <video src>.

Fix: compute token_valid first; if the token is absent or invalid, fall
through to auth_user.require_auth() instead of returning early.
Effect: valid token skips session check, invalid/missing token checks
session, both invalid → 401 as before.

Updated the bogus-token-no-session test docstring to clarify it tests
the no-auth path specifically. Added new test:
  media_token_bogus_token_with_session_returns_200 — verifies a logged-in
  user can still fetch the artifact via a URL carrying a stale token.

* frontend: guard token-fetch effect against stale async resolution

A slow issueMediaToken() response for video A could resolve after the
user selected video B and call setSignedSrc(urlA), making the
preview/Cast play the wrong file.

Add a cancelled flag set in the effect cleanup; both .then and .catch
check it before touching state, so only the most recent src wins.

* frontend: load Cast media immediately if session already exists

Previously the effect only sent video to the TV on SESSION_STARTED /
SESSION_RESUMED events. Two gaps:

1. If a Cast session was already active when signedSrc became ready
   (e.g. the SDK resumed a session before the token fetch finished, or
   the user switches videos while already casting), nothing was sent.

2. Same gap if castReady fired after an already-established session.

Fix: extract loadMedia(session) and call it against
ctx.getCurrentSession() immediately when castReady + signedSrc are both
truthy, in addition to keeping the event listener for future connects.

* server: staged file-upload endpoint

POST /api/archives/:id/uploads streams a multipart body to a temp file
under the archive's store/temp/ directory and returns a staged_path the
capture pipeline can move into place.

- Routes: /api/archives/:id/uploads (POST, requires auth)
- Body cap: 10 GiB; chunk-streamed to disk, never buffered in memory
- Path-traversal sanitised on the filename field
- Temp files are cleaned up on error paths (disk-leak fix)
- main.rs wires the new route into the server startup
- Cargo: adds the multipart dependency

* frontend: file upload in Capture dialog

Drag-and-drop or 'Upload file' button stages files for archiving:

- File items sit alongside URL rows in the same list; each shows the
  original filename, a live progress bar during upload, and a check badge
  when ready.  The locator input is replaced entirely — no editable field.
- Archive button is disabled until all uploads finish; each file item
  contributes to the Archive N count once its upload is done.
- File items are excluded from sessionStorage persistence (they are
  transient — the staged server path would be invalid after a reload).

Staged-file cleanup is handled at every exit path so temp/uploads/ does
not accumulate:
  • removeRow on an in-progress item aborts the XHR; removeRow on a done
    item calls DELETE /archives/:id/uploads.
  • Dialog cancel (Escape / Cancel button) aborts all in-flight XHRs and
    DELETEs all completed staged files via the close-event handler.
  • handleArchive sets isSubmittingRef=true before dialog.close() so the
    close handler skips cleanup — the background capture job handles
    staged-file removal on success instead.
  • uploadFile() returns { promise, abort } so the component can cancel
    the XHR without any visible fetch.

api.js additions: uploadFile (XHR with progress + abort), deleteUpload.

(Static assets rebuilt from combined source to include screensharing
changes from this branch.)

* fix: collection enrollment with default_visibility_bits

Two related fixes from feat-file-uploading:

core: fix collection enrollment using default_visibility_bits instead of
entry.visibility — entries were being enrolled with the entry-level
visibility rather than the collection's configured default.

server: allow changing default_visibility_bits on the default collection
— the PATCH handler was incorrectly blocking updates to the default
collection's visibility configuration.

* server: fix unbounded staged-upload disk growth

Two review findings:

P2 — delete staged file on capture failure (routes.rs)
When perform_capture returns Err, the job was marked failed but
staged_upload_path was never removed. With a 10 GiB body cap a few
failed imports could exhaust archive storage before the next restart.
Mirror the success-path cleanup into the Err arm so the file is removed
immediately regardless of outcome.

P1 — periodic staged-upload pruning (main.rs)
The startup prune of temp/uploads/ only ran once, so uploads abandoned
mid-session (browser crash, navigation away) accumulated forever on a
long-running server. Folded the pruning logic into the existing 24 h
maintenance task alongside session cleanup, so stale dirs are swept
continuously without requiring a restart.

* server+frontend: fix staged-upload disk-growth and prune safety

Server (main.rs + routes.rs):

- Extract prune_stale_upload_dirs() helper called by both startup and
  the periodic 24h task, eliminating the duplicated loop.

- Sentinel (.uploading) created in the UUID dir before streaming begins;
  removed on successful completion; error path uses remove_dir_all so
  the partial file and sentinel are cleaned up together.
  The periodic prune skips any dir containing .uploading (active XHR).

- Startup prune passes cleanup_stale_sentinels=true: the server has not
  started accepting connections yet so any sentinel is a crash remnant —
  it is removed and the dir proceeds to the age check, preventing leaked
  dirs from a previous crash accumulating forever.

- Staleness measured from the newest non-sentinel child file mtime so a
  just-finished slow upload (dir mtime stale, file mtime fresh) is not
  pruned before the user can submit it for capture. Empty dirs fall back
  to dir mtime.

- Failed captures (Err branch in spawn_blocking) now also delete the
  staged file and UUID dir immediately, matching the success path.

Frontend (api.js + CaptureDialog.jsx):

- submitCapture attaches err.status = res.status on non-2xx responses
  so callers can distinguish a definite HTTP rejection from a network
  error where the response may have been lost.

- submitBgJob catch deletes the staged file only when e.status is set
  (server definitively rejected the POST /captures request). A network
  error leaves the file in place because the server may have accepted
  the job and the response was lost — deleting would race the capture.
2026-07-23 20:12:22 +02:00
6377daadae
feat: add capturing w/ parent-child entries (#32)
* feat(core): YouTube playlist/channel/YTM-playlist capture with parent–child entries

- ytdlp: add fetch_playlist_info() using yt-dlp -J --flat-playlist for
  reliable container title + shallow entry list; normalize item URLs via
  webpage_url → absolute url → id fallback (domain inferred from container
  URL so YTM stays on music.youtube.com)
- capture: add record_container_entry() (no blob, no primary_media artifact);
  extend record_media_entry() with parent_entry_id/root_entry_id params (all
  existing single-item call sites pass None, None)
- capture: implement YouTubePlaylist / YouTubeChannel / YouTubeMusicPlaylist
  capture path replacing the two not-implemented stubs: fetch playlist info →
  create container entry (reusing existing run + item) → per-child run items
  (parent_item_id = container item) → download each video/track as a child
  entry; per-child failures are non-fatal; perform_capture returns result.status
  reflecting actual run outcome so capture_handler marks the job correctly
- archive: add child_count i64 to EntrySummary (col 12 in all listing queries);
  add get_entry_summary() private helper; fix get_entry_detail() to use
  get_entry_summary() so child entries are resolvable via the detail endpoint;
  add list_child_entries(conn, uid, caller_bits) with the same
  admin/collection visibility predicate as list_root_entries

archive_runs.requested_count stays 1 (one user locator); discovered/
completed/failed_count reflect container item + N video items via
refresh_run_counters.

* feat(server,frontend): expose children endpoint + expand UI for container entries

server:
- add GET /api/archives/:id/entries/:uid/children → list_entry_children,
  calling list_child_entries with caller_bits so visibility model is enforced
- fix capture_handler: use result.status ("completed"/"failed") to set job
  status rather than always "completed", mirroring rearchive_handler; this
  surfaces partial playlist failures to the polling client

frontend:
- api.js: add fetchEntryChildren(archiveId, entryUid)
- EntryRow: one outer div.entry-row-outer (display:block) keeps nth-child
  striping correct; inner div.entry-row-main is the flex row with all column
  cells and event handling; .child-entries sits below inside the outer wrapper
- expand chevron appears when entry.child_count > 0; clicking fetches children
  lazily and renders ChildRow components reusing .col-* flex widths
- child-count badge shown next to title on container entries
- styles.css: scoped CSS with #entries-body > .entry-row-outer selectors
  (higher specificity than > div) to override flex on outer wrapper; inner row
  and column rules replicated at correct depth; nth-child, is-selected,
  is-multi-selected, url-cell hover all handled

* fix(frontend): make child entry rows interactive

ChildRow now receives onRowClick and selectedUids from EntryRow (which
receives selectedUids from EntriesView alongside the existing booleans).
Clicking a child row invokes onRowClick(child, e) so it flows through
handleRowClick → selectEntry → fetchEntryDetail exactly as a root entry
would. Shift-range selection gracefully degrades to single-select since
child entries are not in the root entries array.

Selected/multi-selected visual state is wired: .child-entry-row.is-selected
shows the same #eee2d2 background + accent outline as root rows; hover
restores full opacity. Frontend static assets rebuilt.

* feat(core): playlist per-item quality + incremental sync

ytdlp.rs:
- Add PlaylistItemProbe / PlaylistProbeResult (pub, serde::Serialize)
- Add private available_video_heights_from_value() helper for Value entries
- Add probe_playlist_qualities(): yt-dlp -J (full metadata, no flat flag)
  returns per-video quality lists in one subprocess call

capture.rs:
- Add per_item_quality: HashMap<String,String> and sync: bool to CaptureConfig
  (both Default; keyed by yt-dlp video ID, not URL)
- Add pub locator_to_playlist_url(): validates only the three playlist sources,
  expands shorthands; keeps locator_to_ytdlp_url's no-playlists contract
- Playlist capture block: sync-aware container resolution
  - sync + existing container → reuse it via complete_archive_run_item,
    skip already-archived children (by canonical URL) before creating
    run items so refresh_run_counters only counts new items
  - sync + no container → create normally (first sync run)
  - non-sync → always create fresh container (existing behaviour)
- Per-item quality: config.per_item_quality.get(id) falls back to child_quality

archive.rs:
- Add get_archived_playlist_child_urls(): returns HashSet of canonical URLs
  of all children under any container matching the playlist canonical URL
- Add find_container_entry_id_by_canonical_url(): returns most-recent
  container entry id (parent_entry_id IS NULL) for a given canonical URL

routes.rs: stub per_item_quality/sync on both CaptureConfig sites (server
agent will wire body fields in Phase 2)

* fix(core)+test: propagate sync query errors; cover new playlist/sync functions

archive.rs:
- get_archived_playlist_child_urls: collect() as rusqlite::Result<HashSet<_>>
  instead of filter_map(ok) so row-level errors surface rather than silently
  skipping and causing duplicate downloads

capture.rs:
- match get_archived_playlist_child_urls result and fail_run on error instead
  of unwrap_or_default, preventing silent re-downloads on DB failure

Tests added to capture.rs:
- locator_to_playlist_url_accepts_playlist_shorthands (yt:playlist/, ytm:playlist/, full URL)
- locator_to_playlist_url_accepts_channel_shorthands (yt:@handle)
- locator_to_playlist_url_rejects_non_playlist_sources (single video, tweet, web page)

Tests added to archive.rs (all use in-memory DB via make_tag_test_db):
- find_container_entry_id_returns_none_when_absent
- find_container_entry_id_returns_root_entry
- find_container_entry_id_ignores_child_entries (child with parent_entry_id set)
- get_archived_playlist_child_urls_empty_when_no_playlist
- get_archived_playlist_child_urls_returns_children
- get_archived_playlist_child_urls_excludes_other_playlists

* feat(server,frontend): playlist quality selector + per-video overrides + sync UI

routes.rs:
- CaptureBody gains per_item_quality (HashMap<String,String>, serde(default))
  and sync (bool, serde(default)); both validated before use
  - per_item_quality values validated against same quality predicate as top-level
    quality field ("best"|"audio"|"NNNp") so bad per-video values are rejected
    at the API boundary rather than silently falling through to quality_format
- capture_handler threads body.per_item_quality + body.sync into CaptureConfig
  (replaces hardcoded empty stubs); rearchive_handler keeps empty defaults
- New POST /api/archives/:id/captures/probe-playlist: calls
  probe_playlist_qualities via spawn_blocking; 400 for non-playlist locator,
  502 on yt-dlp failure, returns PlaylistProbeResult as JSON

api.js:
- probePlaylist(archiveId, locator): POST probe-playlist endpoint
- submitCapture: forwards per_item_quality (non-empty) and sync:true from
  extraExtensions param added to submitBgJob

CaptureDialog.jsx:
- isPlaylistSource(): detects yt:/youtube: playlist/@/channel, ytm:playlist/,
  YouTube/YTM HTTP(S) URLs with list= param or channel pathnames
- makeItem(): 6 new playlist state fields
- applyPlaylistQuality(): conflict logic — videos that can reach selected
  quality get it set; videos that can't and have no prior selection are left
  null (conflict); videos with a prior selection keep it when quality is raised
- hasConflict(): any playlistItems entry with quality===null
- updateLocator(): isPlaylistSource branch with 800ms debounce→probePlaylist;
  existing isVideoSource path unchanged
- Archive button disabled when anyConflict or any probe in flight
- Per-video expand list with individual quality selects, conflict badges,
  sync toggle (appears after probe completes)

styles.css: playlist expansion, conflict, sync toggle CSS

* fix(core): ignore per_item_quality for YTM playlist items

YouTube Music playlists force child_quality = Some("audio") because
yt-dlp can't download DRM-free audio-only tracks any other way. The
previous per_item_quality lookup could override this with e.g. "best",
defeating the invariant. Guard the lookup behind !is_audio so YTM items
are always downloaded as audio regardless of what the caller sends.

* fix(frontend): exclude /watch from isPlaylistSource

youtube.com/watch?v=...&list=... and music.youtube.com/watch are single
videos in the backend (Source::YouTubeVideo / YouTubeMusicTrack) regardless
of the list param. Previously isPlaylistSource returned true for these,
which would have triggered the playlist probe path while the video probe
was already running, and the render would attempt to show playlist UI on
an item whose playlistProbeState stays idle.

Guard: if pathname === '/watch', return false before the list-param check.

* fix(frontend): tighten isPlaylistSource to mirror backend routing exactly

Previous fix excluded /watch but still returned true for any youtube.com
URL with a ?list= param (e.g. /shorts/xxx?list=yyy). Backend determine_source
only routes to YouTubePlaylist on /playlist?list=... and to YouTubeChannel on
/@handle, /channel/, /c/, /user/ paths — everything else is a single item.

Rewrite the HTTP block to match:
- youtube.com: pathname==='/playlist' && list param → playlist
             : /@, /channel/, /c/, /user/ → channel
             : anything else (incl. /watch&list=, /shorts?list=) → false
- music.youtube.com: pathname==='/playlist' && list param → YTM playlist
                   : /watch → single track (falls through to false)

* fix(frontend): guard handleArchive against Enter-key bypass of disabled state

The Archive button is disabled when anyConflict || anyProbing, but
onKeyDown on the locator input calls onSubmit() → handleArchive()
directly, bypassing the button's disabled check entirely.

Add the same conditions as early returns inside handleArchive itself,
operating on toSubmit (the items that would actually be submitted) so
the guard is tight — items with no locator are already excluded by the
toSubmit filter.

* fix(frontend): drop m.youtube.com from isPlaylistSource

Backend determine_source playlist/channel regex only matches
(?:www\.)?youtube\.com — mobile URLs hitting m.youtube.com would be
probed as playlist in the UI but captured as Source::Url server-side.
Remove m.youtube.com from the detector to keep frontend and backend
in exact agreement. Add backend support when needed.

* fix(frontend): audio-only conflict handling in playlist quality selector

applyPlaylistQuality('audio'):
- Only sets quality='audio' on items where has_audio=true
- Items with has_audio=false: keep prior selection if set, else null
  (conflict) — same rule as unsupported height, blocks archive until
  user explicitly picks a quality for those items

Playlist-level 'Audio only' option:
- Changed hasAnyAudio → allHaveAudio (every item must have audio)
- When any item lacks audio, the option is hidden entirely so the
  selector can never create immediate conflicts just by appearing

* fix(frontend): add yt:user/ to isPlaylistSource shorthand detection

Backend determine_source routes yt:user/... (and youtube:user/...) to
YouTubeChannel — already covered by the yt: shorthand block for
playlist/, @, channel/, c/ but missing user/. Old-style user channel
URLs would capture correctly server-side but never show the playlist
quality/sync UI.

* fix(frontend): block playlist submission unless probe is done

Previous guard only blocked while playlistProbeState==='probing'.
Two remaining bypass paths:
- idle: 800ms debounce not yet fired after URL typed
- error: probe failed — no per-video quality data available

Change anyProbing and handleArchive guard to:
  isPlaylistSource(locator) && playlistProbeState !== 'done'

This means idle/probing/error all block submission for playlist items.
error is intentionally blocking — without quality data the per-video
requirement can't be satisfied; user must retry or remove the URL.

* fix(frontend): accurate error message when playlist probe fails

Previous text said 'using best quality' implying the capture would
proceed, but probe error now blocks submission. Replace with 'Probe
failed — edit URL to retry' in orange (capture-quality-hint--error)
so the disabled button and the message are consistent.

* fix(frontend): exact quality match in applyPlaylistQuality

Replace maxHeight >= newHeight (cap check) with item.qualities.includes(newQ)
(exact match). A video with [2160p, 1080p] does not support 1440p; the
previous logic would mark it as supporting any quality up to 2160p and
submit '1440p' which yt-dlp silently downloads as 1080p — misrepresenting
the selected quality.

With exact match, unsupported qualities correctly fall through to the
conflict path (keep prior selection or null), enforcing the same manual-
choice requirement as any other unsupported quality.

Per-row selects are unaffected: they already render only pi.qualities
(the video's actual available formats), no maxHeight logic involved.

* fix(frontend): move playlist expand chevron to left of input

User asked for the chevron to be on the left of the playlist input,
not tucked after the quality selector on the right.

- Remove chevron from qualityEl (it was between the quality select and
  the remove button)
- Add it as the first child of capture-row-main, before the <input>,
  when isPlaylistSource && playlistProbeState === 'done'
- Show a same-width placeholder span while probing/idle/error so the
  input does not jump left when the chevron appears after probe completes
- Add capture-playlist-toggle--left modifier (flex-shrink:0, tighter
  padding) and .capture-playlist-toggle-placeholder (fixed 22px width)

* doc(server): scope per_item_quality guarantee to the UI

The server validates per_item_quality value shapes but does not enforce
that every playlist item has an entry. Items without an override get
the global quality as a yt-dlp cap with graceful fallback.

The 'must choose quality for unsupported videos' invariant is a UI
constraint enforced by the frontend before submission. A direct API
caller bypassing the UI accepts yt-dlp's standard cap-and-fallback
behavior. Document this scope explicitly so the gap is intentional,
not accidental.

* fix(frontend): prevent 'reading some of undefined' crash from stale sessionStorage

Old captureItems entries saved before the playlist fields were added
have playlistItems=undefined (missing key). The hasConflict guard
checked !== null, which undefined passes, then called .some() on
undefined → TypeError.

Two-part fix:
1. sessionStorage restore: merge each saved item over makeItem() defaults
   so any missing fields (playlistItems, playlistProbeState, etc.) are
   filled with their correct initial values before the item is used
2. hasConflict: use Array.isArray() instead of !== null so undefined
   is also safely rejected — defence-in-depth for any future field gap

* fix(core): playlist total size includes children; per_item_quality as include-set

archive.rs: total_artifact_bytes for root entries now adds a correlated
subquery summing children's blob bytes so playlist/channel containers
show the real download size instead of 0.

capture.rs: non-empty per_item_quality map now acts as an include-set —
items whose yt-dlp ID is absent are skipped entirely. This wires the
UI's per-video delete button to actual capture exclusion. Empty map
preserves the existing behaviour (download everything).

routes.rs + capture.rs doc: comments updated to reflect both semantics
(empty = all / non-empty = only listed IDs) accurately.

* fix(frontend): QA fixes — playlist UX, selection stroke, URL expand

CaptureDialog.jsx:
- Placeholder no longer appears on idle playlist rows (only during
  probing); chevron shows only after probe completes — no left-padding
  while the user is still typing
- Per-video delete button added to expanded playlist list; removes item
  from playlistItems so its ID is absent from per_item_quality on submit
- anyEmptyPlaylist guard: Archive button disabled + handleArchive early-
  return when all videos have been deleted (empty map would otherwise
  silently download everything)

styles.css:
- Selection stroke switched from outline to box-shadow:inset everywhere
  (entry-row-outer, child-entry-row, legacy flat-div selector) —
  guaranteed inside element bounds, no layout interference
- URL cell overflow only on :hover; removed is-selected .url-cell rule
  that was expanding child URLs when their parent was selected
- Playlist items redesign: separator lines instead of background fills,
  amber left-border for conflicts, thin scrollbar, tighter padding
- Remove button: opacity 0.3 always-visible baseline; full opacity on
  hover/:focus-visible; forced to 1 on coarse-pointer (touch) devices

* fix(frontend): no left gap on playlist rows until chevron exists

Remove the probing-state placeholder span entirely. The chevron renders
only when playlistProbeState === 'done'; all other states (idle, probing,
error) render null. The small layout shift when the chevron appears after
probe is acceptable; blank padding while there is no chevron is not.

* fix(frontend): clear box-shadow on entry-row-outer to prevent double stroke

The flat selector '#entries-body > div.is-selected' already applies
box-shadow to .entry-row-outer (it IS a direct child div). The outer
override rule only cleared 'outline', so the box-shadow leaked through,
wrapping the entire parent+children block with a second stroke.

Add box-shadow: none to both .is-selected and .is-multi-selected on
.entry-row-outer so the stroke sits only on .entry-row-main.

* docs: document per-video exclude in README YouTube playlists section

* fix(frontend): scope selection background to entry-row-main only

Moving background:#eee2d2 off .entry-row-outer onto > .entry-row-main
so that expanded child entries don't inherit the selection highlight.

Outer wrapper now explicitly unsets background (cancelling the flat-div
cascade rule) and clears outline+box-shadow. Both the selection colour
and the inset stroke live on .entry-row-main only.

* fix(frontend): alternating stripe backgrounds on child entry rows

Child rows were inheriting the parent entry's stripe color, making the
expanded list look like one flat block. Apply the same odd/even palette
as root entries (var(--paper-3) / #f2ede5) so each video row is visually
distinct within the expanded group.

* fix(core): delete_entry correctly nulls FK for child entries

archive_run_items.produced_entry_id has no ON DELETE action so it must
be manually nulled before the entry row is deleted. The old query used
WHERE root_entry_id = entry_id, which finds descendants of a root but
returns nothing when entry_id IS a child (children have no sub-children,
so no row has root_entry_id = child_id). The DELETE then failed under
foreign_keys=ON.

Fix: add OR produced_entry_id = ?1 so the entry's own run_item FK is
always cleared before deletion, regardless of whether it is a root or a
child. The subtree subquery is kept for the root-deletion case where all
child run_items also need nulling.

* fix(core+frontend): child entry selection and delete correctness

database.rs — delete_entry:
- subtree_ids now uses WHERE id = ?1 OR root_entry_id = ?1 so the entry
  itself is always included; previously a child deletion passed an empty
  vec to cascade_cached_bytes_after_subtree_delete (no grandchildren
  exist), leaving cached_bytes stale on entries sharing that child's blobs

App.jsx — handleRowClick:
- Shift-range now queries DOM order (#entries-body [data-entry-uid])
  instead of entries.findIndex(); child rows are in the DOM but not in
  the root entries array, so findIndex always returned -1 for them
- Ctrl/meta branch computes next set before the state update so
  selectEntry can fire synchronously for child rows; auto-snap only
  restores root entries, so ctrl-clicking a lone child never loaded its
  detail panel — now calls selectEntry(entry) when the child ends up as
  the sole selection, selectEntry(null) on multi or deselect
- selectedUids added to handleRowClick's useCallback dep array

* fix(frontend): resolve detail entry for any remaining child after ctrl-deselect

Add entryCacheRef (uid→entry Map) populated on every row click. When
ctrl/meta-deselecting leaves exactly one other entry selected, look up
the remaining UID in the cache before falling back to the root entries
array. Without this, deselecting from a multi-selection where the
remaining entry is a child row left selectedEntry null (auto-snap only
searches root entries).

* fix(frontend): child row stripes, deleted-child visibility, selection completeness

EntryRow.jsx / ChildRow:
- Index-based light/dark classes (child-entry-row--light/dark) replace
  nth-child rules; parity comes from children.map idx so no sibling —
  including the loading div — can shift the stripe order
- Loading div moved outside .child-entries so it never affects child
  row ordering at all
- Accept deletedUids prop; filter expanded children array before render
  so deleted children disappear immediately without waiting for reload

EntriesView.jsx: thread deletedUids through to EntryRow

App.jsx:
- Add deletedUids state; handleEntryDeleted/handleBulkDeleted populate it
- isRoot/hasChildDelete computed from entries before setEntries (safe in
  StrictMode — no side-effects inside updater functions)
- Child delete triggers loadEntries to refresh stale parent child_count
  and total_artifact_bytes
- handleRowClick ctrl/meta cache-miss branch uses det.summary (not det)
  from fetchEntryDetail; archiveId added to dep array
- handleRowClick dep array includes archiveId

* fix(frontend): add inset stroke to child-entry-row.is-multi-selected

* fix(frontend): suppress mouse-click focus ring on entry expand button

* fix(frontend): index-based stripes for root and child rows

Root rows: EntriesView passes rowIndex from entries.map to EntryRow,
which applies entry-row-outer--light/dark. Retires both nth-child stripe
blocks so skeleton rows can never shift the first real entry to dark.
Skeleton rows keep a :not(.entry-row-outer) nth-child fallback.

Child rows: colours changed from the warm root palette (paper-3/#f2ede5)
to cooler near-whites (#fafaf8/#f2f0ec) so children are visually distinct
from their parent row regardless of which stripe the parent sits on.

* feat(core+frontend): bare yt:ID resolves to YouTube video

determine_source: yt:ID / youtube:ID with no prefix and exactly 11
chars [A-Za-z0-9_-] → YouTubeVideo via is_youtube_video_id helper.
Reserved prefixes (playlist/, channel/, c/, user/, @) still fire first
so they are unaffected by the fallback.

expand_shorthand_to_url: bare yt:ID expands to watch?v=ID using the
same predicate, consistent with how ytm:ID → music.youtube.com/watch.

isVideoSource (frontend): same 11-char /^[A-Za-z0-9_-]{11}$/ regex so
yt:ID triggers the quality probe and capture-row guards identically to
yt:video/ID.

Tests: test_is_youtube_video_id covers valid IDs (alphanumeric, with _
and -), too-short, too-long, and invalid-char cases using genuinely
invalid fixtures. test_youtube_sources adds bare-ID cases and confirms
reserved prefixes (playlist/, @) are not affected.

* fix(core+frontend): exclude avatar blobs from tweet % cached display

tweet/tweet_thread entries have avatar artifacts that are always
deduplicated from the first capture of each author. Counting them in
the cached-bytes percentage makes it artificially high (or incorrect
when the real media content is new but avatars are cached).

database.rs:
- refresh_entry_cached_bytes: AND ea.artifact_role != 'avatar'
- cascade_cached_bytes_after_delete: same filter
- cascade_cached_bytes_after_subtree_delete: same filter
- Initial cached_bytes migration: same filter
- Re-migration (else branch): recomputes cached_bytes for existing
  entries that have avatar artifacts, scoped to only those entries

archive.rs:
- EntrySummary gains cacheable_bytes: i64 — non-avatar total bytes,
  computed inline in every SQL query as the denominator for % cached
- ENTRY_SELECT_COLS adds cacheable_bytes at index 13 (with children
  subquery, same as total_artifact_bytes)
- list_root_entries adds same expression
- All 6 row-mapping closures include cacheable_bytes: row.get(13)?

EntryRow.jsx:
- SIZE display stays: formatBytes(entry.total_artifact_bytes)
- % cached badge uses cacheable_bytes as denominator:
  cached_bytes / cacheable_bytes * 100

* feat(frontend): j/k keyboard navigation for entries; test avatar cached_bytes

App.jsx — j/k handler:
- Fires on keydown when not focused on INPUT/TEXTAREA/SELECT/contenteditable
- Ignores meta/ctrl/alt modifier combos
- Uses DOM order (#entries-body [data-entry-uid]) so expanded child rows
  participate, matching the shift-range selection logic
- Resolves target entry via entryCacheRef → root entries array →
  fetchEntryDetail(..).summary on cache miss
- Scrolls target into view (block: nearest); updates lastAnchorIndexRef
  so subsequent shift-click ranges start from the keyboard-navigated row

archive.rs — cached_bytes_excludes_avatar_blobs test:
- Two tweet entries sharing an avatar blob (100B) and a media blob (900B)
- Asserts refresh_entry_cached_bytes sets cached_bytes = 900 (not 1000)
- Asserts list_root_entries summary: cached_bytes=900, cacheable_bytes=900,
  total_artifact_bytes=1000 — SIZE includes avatar, % cached denominator
  and numerator both exclude it

* fix(frontend): guard j/k uncached-child fetch with monotonic token

Rapid j/k over child rows that aren't in entryCacheRef triggers
fetchEntryDetail calls in parallel. Without a guard the last one to
settle wins, desyncing selectedUids (highlight) from selectedEntry
(detail panel/URL).

Fix:
- jkSeqRef (monotonic counter) incremented before each server fetch;
  the .then() guard tok === jkSeqRef.current drops results from
  superseded navigations
- handleRowClick increments jkSeqRef.current so any click also cancels
  an in-flight j/k fetch

* fix(frontend): guard ctrl/meta cache-miss fetch; add / search shortcut

App.jsx ctrl/meta branch: cache-miss fetchEntryDetail now captures
tok = ++jkSeqRef.current before the fetch and gates selectEntry on
tok === jkSeqRef.current — same pattern as the j/k handler — so a
slow response after a later click/navigate can't overwrite selection.

/ key: reuses the Cmd+K/Ctrl+K handler path (focus+select search input,
or pendingSearchFocus + archive view switch). Guards: editable targets
(INPUT/TEXTAREA/SELECT/contentEditable) and modifier keys are checked
before preventDefault() so typing / in inputs is untouched.

* feat(core): attempt SpotifyTrack via yt-dlp instead of hard-failing

Previously all Spotify sources returned an error claiming yt-dlp cannot
download DRM-protected audio. yt-dlp does have a Spotify extractor
(experimental, content-dependent), so refusing upfront is worse than
trying and letting it report the real failure.

SpotifyTrack now goes through the same yt-dlp metadata + audio-quality
download path as YouTubeMusicTrack. SpotifyAlbum and SpotifyPlaylist
still return an explicit error — fetch_playlist_info has YouTube-specific
URL fallback logic that would produce bogus child URLs for Spotify flat
entries; those sources need dedicated container handling first.

* chore: remove docs/superpowers from repo and gitignore whitelist

* feat(core+frontend): Spotify album/playlist capture via yt-dlp container path

Previously SpotifyAlbum and SpotifyPlaylist hard-failed with a DRM error.
yt-dlp does support Spotify (experimentally), so they now go through the
same probe/container/child path as YouTube playlists.

ytdlp.rs:
- Extract normalize_item_url() helper used by both fetch_playlist_info
  and probe_playlist_qualities; eliminates the duplicate URL-normalization
  blocks and the YouTube-specific fallback_host variable
- Fallback logic is now platform-aware: YouTube/YTM bare IDs → watch URL,
  Spotify → open.spotify.com/track/{id}, unknown → skip with warning

capture.rs:
- locator_to_playlist_url: accept SpotifyAlbum | SpotifyPlaylist so the
  probe-playlist endpoint accepts Spotify album/playlist URLs
- Container branch: add SpotifyAlbum | SpotifyPlaylist to the matches!
- is_audio: true for SpotifyAlbum | SpotifyPlaylist (audio-only, like YTM)
- child_source: SpotifyTrack for Spotify containers (not YouTubeMusicTrack)
- generate_entry_title and record_media_entry both use the correct child
  source so entity_kind, source_kind, and representation_kind are right

CaptureDialog.jsx:
- isPlaylistSource: recognise open.spotify.com/album/ and /playlist/,
  and spotify:album:ID / spotify:playlist:ID shorthands, so Spotify
  containers get the probe UI, per-track excludes, and sync toggle

* fix(core+server): partial playlist refresh + child visibility inheritance

database.rs:
- finish_archive_run: reverted 'partial' status (violates CHECK constraint);
  restored binary completed/failed
- get_run_completed_count(): new helper for callers that need to distinguish
  partial success without touching the DB status enum

capture.rs:
- CaptureResult gains completed_count: i64 (0 for single-item captures;
  populated from get_run_completed_count for container/playlist captures)
- All CaptureResult construction sites updated

routes.rs:
- Playlist job_status mapping: mark job 'completed' when completed_count > 0
  (even if status == 'failed'), so partially-successful playlist captures
  trigger onCaptured and show the archived entries without a manual reload.
  Truly zero-success runs (completed_count == 0, status == 'failed') stay
  failed as expected.
- probe_playlist_handler error message updated to mention Spotify

archive.rs (list_child_entries):
- Add third OR arm: children are visible when their parent container is in a
  collection visible to the caller. Fixes empty child list for non-admin
  users with access to a playlist container but not its newly-created
  children (which are all in the default 'private' collection).

* fix(server): use completed_count to distinguish partial from total failure

finish_archive_run is binary (completed/failed) — a playlist where some
tracks succeed and some fail returns status='failed'. Without this fix,
the job mapping treated any failed status as a job failure, causing
onCaptured to never fire and leaving successfully archived entries hidden
until a manual reload.

Now: job is 'failed' only when status='failed' && completed_count==0.
Partial runs (completed_count > 0) map to a completed job so the UI
refreshes and shows the entries that were successfully captured.

* fix(core+server): exclude container from child success count; rename field; tests

database.rs:
- get_run_completed_child_count(): renamed from get_run_completed_count and
  scoped to child items only (parent_item_id IS NOT NULL). The container
  run item is always completed first, so the old function inflated the count
  by at least 1 for every playlist, making all-video-failed runs appear as
  partial successes.
- New regression test: completed_child_count_excludes_container — completes
  both a root and a child item, asserts DB completed_count==2 while
  get_run_completed_child_count==1.

capture.rs:
- CaptureResult.completed_count renamed to completed_child_count to match
  the function and make the semantics unambiguous at the call site.

routes.rs:
- job_status decision now uses result.completed_child_count == 0 so that a
  playlist where every video fails (child_count==0) is correctly reported
  as a failed job, not a completed one.

archive.rs:
- New regression test: list_child_entries_inherits_parent_visibility —
  enrolls only the container in a USER-visible collection, asserts the
  child is visible to a USER caller and invisible to a GUEST caller.
2026-07-21 21:57:29 +02:00
d202e177e1
feat(frontend): add async capture UX with skeleton entries for in-progress captures (#31)
* feat(frontend): add SkeletonEntryRow with shimmer animation

Adds a new SkeletonEntryRow component that renders animated shimmer
placeholder cells matching the exact column layout of EntryRow
(col-added, col-title with icon circle, col-type pill, col-size,
col-url). CSS appended to styles.css using existing design tokens
(--paper-2, --line-soft, --paper-3) for the warm-toned shimmer.

* feat(frontend): async capture UX — reset dialog on submit, show skeleton rows

When the user presses Archive, CaptureDialog now immediately resets its
form to a fresh empty state and closes. Background captures continue
polling via intervals that survive the dialog close.

Skeleton rows appear at the top of EntriesView (filtered to the active
archive) for each in-flight job, giving visual feedback that something
is being processed. On completion the skeleton is removed and the entry
list refreshes; on failure the skeleton is removed and an error toast
fires — matching the existing toast behaviour.

Architecture:
- App owns pending-capture state (pendingCaptures) as the single source
  of truth, persisted to sessionStorage['pendingCaptures']. On page
  refresh, App seeds the list and passes it to CaptureDialog as
  activeJobs so polling reconnects without a second sessionStorage read.
- CaptureDialog emits onJobStarted({id,jobUid,locator,archiveId}) only
  after submitCapture() returns a job_uid — never before, never on
  failure — ensuring no orphan skeletons can persist after a refresh.
- onJobSettled(id) is called by startPolling on any terminal state
  (completed, failed, or network error), removing the skeleton.
- EntriesView filters pendingCaptures by archiveId so a capture in
  archive A never shows a skeleton in archive B.

Removed from CaptureDialog: submitItem, resetRow, hasActiveJobs,
anyActive, CapStatusDot. CaptureRow is simplified to idle-only display
with no disabled state, no retry button, no status dot.

* fix(frontend): skeleton replacement ordering and col-check alignment

- Await the entry list refresh before removing the skeleton so the
  real row arrives before the placeholder disappears. handleCaptured
  now returns its Promise.all; startPolling awaits it before calling
  onJobSettled on the success path.
- Add col-check as the first child of SkeletonEntryRow to match
  EntryRow's DOM structure; without it, columns misalign on touch
  devices where col-check becomes display:flex.

* fix(frontend): use Promise.allSettled in handleCaptured to prevent runs-fetch failure from poisoning successful captures

Promise.all rejects on the first failure. If the /runs refresh threw,
the rejection would propagate through startPolling's await and land in
the outer catch block, firing an error toast for a capture that had
already succeeded. Promise.allSettled settles unconditionally so a
transient runs fetch error is silently absorbed while the entry list
still refreshes.
2026-07-20 14:27:15 +02:00
14fd091d1f
fix(core): tighten reader-mode typography and spacing
- h1: explicit font-size:2em, font-weight:700, line-height:1.2
- byline: margin:.3em top, font-size:15px, line-height:1.5
- siteName: margin:.3em top, font-size:13px
- add p{margin:0 0 1.15em} for consistent paragraph spacing
- add figcaption rule: 14px italic gray with .5em top margin
- bump figure/img margin to 1.5em, blockquote margin to 1.2em
2026-07-19 21:07:19 +02:00
00e3b966ee
fix(core): tighten Freedium download-button cleanup
Replace the broad previousElementSibling/HEADER removal with a
precise selector that matches only the .flex.justify-end wrapper
containing the 'Download article' button. The old heuristic was
removing article headers on NYT/WaPo captures.
2026-07-19 20:32:21 +02:00
64f94c6eda
feat: archive paywalled articles through Freedium (#30)
* feat(core): add via_freedium to CaptureConfig; route WebPage captures through Freedium mirror

When via_freedium is true and the locator is not already a Freedium URL,
perform_capture passes https://freedium-mirror.cfd/<original-url> to
singlefile::save() so paywalled articles are fetched via the mirror.
The original locator is kept for requested_locator and canonical_locator
in the DB entry. An empty cookie map is used for the mirror fetch to
prevent original-domain credentials from being sent to freedium-mirror.cfd.

* feat(server): expose via_freedium in capture API (default on)

CaptureBody gains via_freedium: Option<bool>; absent defaults to true.
The rearchive handler sets it to false — existing entries should not be
silently re-fetched through a mirror.

* feat(frontend): add Freedium mirror toggle to capture advanced options

- freediumEnabled state defaults to true (on by default)
- via_freedium forwarded through submitCapture to the capture API
- Toggle rendered last in the advanced panel, matching existing rows
- Built frontend static assets included

* fix(core): Freedium capture fixes — title, notifications, reader images

- extract_html_title: take().read_to_end() up to 256 KiB (single read() can
  short-read, leaving title at byte ~106 K undiscovered); regression test added
- Strip " - Freedium" suffix before storing entry title
- is_freedium_fetch keys off actual fetch_url host
- Remove Freedium toast overlay ([data-sonner-toaster]) before capture
- Resolve lazy images (data-zoom-src etc.) before Readability in reader mode

* fix(core): extract HTML title after font stripping, not before

SingleFile embeds fonts as base64 data URIs in <style> blocks in the
<head>, pushing the <title> tag to ~1.2 MB in the raw temp file.
The 256 KiB read window in extract_html_title missed it.

Font extraction rewrites the HTML in-place (2.4 MB → 1.26 MB) before
hashing, so the title is accessible at byte ~106 KB in that content.

Fix: add extract_html_title_str() that operates on a &str; call it on
the in-memory rewritten string after font extraction (server path).
CLI path (no font extraction) falls back to result.title as before.

* fix(core): inline reader-mode images via Rust post-processing

SingleFile cannot inline resources added to the DOM at before-capture
time — only resources tracked during the page's initial load cycle get
embedded. Article images in Readability output fall into this gap when
the page framework has already resolved their lazy src to a CDN URL
that isn't in SingleFile's resource cache for the new DOM elements.

Fix in three parts:
- Browser script: after body.innerHTML = article.content, stamp
  data-archivr-src=<absolute-proxy-url> on any image whose src is
  not an already-inlined large data URI. Remove loading attr. Don't
  touch src (let SingleFile try; Rust handles the rest).
- Rust (save_with): after SingleFile writes the file and before hashing,
  call inline_archivr_img_srcs() to scan for data-archivr-src markers.
- inline_archivr_img_srcs(): fetches each marked URL with blocking
  reqwest (10 s timeout, 5 redirects, image/* Content-Type guard,
  20 MiB cap), base64-encodes, replaces src with the data URI, and
  removes the marker attr. Non-fatal — fetch failures are logged.

Also: Freedium UI cleanup now removes footer, #progress, and empty
data-nosnippet wrappers in addition to the existing nav/toaster removal.

* fix(core): clean up Freedium chrome and drop reader-mode debug counters

Two fixes:

1. freedium_cleanup: remove remaining Freedium article chrome that
   survived the existing nav/footer/toaster pass:
   - <header class="p-6 bg-gray-50 ..."> — author/metadata bar with
     profile pictures and byline, a Freedium wrapper around the article
   - <section> containing [data-slot="dropdown-menu-trigger"] or
     [aria-haspopup="menu"] — the "Download article" dropdown

   Selectors target stable Tailwind utility class prefixes (p-6, bg-gray-50,
   bg-zinc-800) for the header and the WAI-ARIA menu role for the button,
   both resilient to minor Freedium UI updates.

2. Reader-mode meta tag: strip per-capture debug counters.
   _archivrReaderMark now writes 'applied' instead of
   'applied:pre_s=N,...:post_s=N,...' — the counters were useful during
   development but have no place as permanent archive content.

* fix(core): prevent lazy-resolver double-processing on Freedium reader images

_archivrResolveLazyImgs is called twice: before Readability (_pre) and
after the post-body stamp pass (_post). The stamp pass sets data-archivr-src
but previously left data-src/data-zoom-src/etc intact, so _post's
'lazySrc && _isPlaceholder' condition fired on the same images and
rewrote src to the CDN URL — which SingleFile then tried and failed to
fetch from the Freedium proxy context, redundantly.

Fix: strip all lazy attrs (data-src, data-lazy-src, data-zoom-src,
data-original, data-lazy) when stamping data-archivr-src. _post then
finds no lazySrc on those images and skips them cleanly. Rust owns
them via data-archivr-src. _post continues to handle any remaining
placeholder images not covered by the stamp pass (non-Freedium reader
captures).

* fix(core): address Codex review findings in reader image post-processor

P1 — Enforce size cap before buffering (singlefile.rs fetch_image_as_data_uri):
resp.bytes() buffered the entire response before checking MAX_BYTES, enabling
OOM on oversized or attacker-controlled images. Fix: reject via Content-Length
header when present, then stream at most max_bytes+1 bytes with Read::take so
the cap is enforced without materialising the full body first.

P2 — Decode HTML entities in extracted image URLs (inline_archivr_img_srcs):
The browser's HTML serialiser encodes & as &amp; in attribute values, so CDN
signed URLs with query parameters (e.g. ?a=1&b=2) arrived as &amp;-escaped
strings. reqwest sent the wrong URL, breaking signed or transformed CDN
images. Fix: unescape &amp; &lt; &gt; &quot; before passing to the fetcher.

P2 — Preserve authentication for same-origin lazy images (inline_archivr_img_srcs):
The post-processor created a bare reqwest client with no cookies, causing 401/403
for reader-mode captures of auth-gated sites whose lazy images are same-origin.
Fix: pass capture_url and cookies into inline_archivr_img_srcs; attach a Cookie
header only when domain_from_url(img_url) == domain_from_url(capture_url) and
cookies is non-empty. Third-party image hosts and Freedium fetches (which receive
empty cookies in capture.rs) are unaffected.

* test(core): extract helpers and add unit tests for Codex review fixes

Extract two testable pure functions from inline_archivr_img_srcs:
- html_attr_decode: decodes HTML character references in attribute values.
  Fixes decode order: &amp; runs LAST so &amp;lt; → &lt; (one layer
  removed), not < (two layers). Previous order caused double-decoding.
- same_origin_cookie_header: returns a Cookie header value only when
  img_url's domain matches capture_url's domain.

Add 11 unit tests covering:
- html_attr_decode: plain URL no-op, &amp; in CDN query params, single-layer
  decode (&amp;lt; → &lt; not <), double-encoded amp (&amp;amp; → &amp;),
  direct &lt;/&gt;/&quot; decode
- same_origin_cookie_header: same host attaches cookies, third-party returns
  None, empty cookies returns None, Freedium empty-cookies no-op
- bounded_read: Read::take stops at max+1 bytes (guard fires), allows
  exactly-at-limit payloads (guard does not fire)

* docs: refresh AGENTS.md and mental model; drop NEXT.md

AGENTS.md: mention Freedium mirror in the overview and capture flow,
add vendor/readability/ to Key Directories, drop the NEXT.md pointer.

ARCHIVR-MENTAL-MODEL.md: fix stale references to legacy static/ paths
in Where To Edit (frontend now lives in frontend/src/); add capture.rs
and auth.rs rows; replace the outdated 'Current Limitations' section
(which claimed no capture and no auth) with a Server Capabilities
section reflecting async capture jobs, the auth model, search, and
admin scope; add a Web Capture Pipeline section covering the Freedium
mirror, SingleFile+Chromium, vendored Readability, cleanup, and Rust
post-processing.

NEXT.md: remove; the roadmap tracking is stale and unused.
2026-07-19 17:45:07 +02:00
15660e6532
docs: refresh AGENTS.md and mental model; drop NEXT.md
AGENTS.md: mention Freedium mirror in the overview and capture flow,
add vendor/readability/ to Key Directories, drop the NEXT.md pointer.

ARCHIVR-MENTAL-MODEL.md: fix stale references to legacy static/ paths
in Where To Edit (frontend now lives in frontend/src/); add capture.rs
and auth.rs rows; replace the outdated 'Current Limitations' section
(which claimed no capture and no auth) with a Server Capabilities
section reflecting async capture jobs, the auth model, search, and
admin scope; add a Web Capture Pipeline section covering the Freedium
mirror, SingleFile+Chromium, vendored Readability, cleanup, and Rust
post-processing.

NEXT.md: remove; the roadmap tracking is stale and unused.
2026-07-19 17:42:40 +02:00
e5fb6ea9bb
test(core): extract helpers and add unit tests for Codex review fixes
Extract two testable pure functions from inline_archivr_img_srcs:
- html_attr_decode: decodes HTML character references in attribute values.
  Fixes decode order: &amp; runs LAST so &amp;lt; → &lt; (one layer
  removed), not < (two layers). Previous order caused double-decoding.
- same_origin_cookie_header: returns a Cookie header value only when
  img_url's domain matches capture_url's domain.

Add 11 unit tests covering:
- html_attr_decode: plain URL no-op, &amp; in CDN query params, single-layer
  decode (&amp;lt; → &lt; not <), double-encoded amp (&amp;amp; → &amp;),
  direct &lt;/&gt;/&quot; decode
- same_origin_cookie_header: same host attaches cookies, third-party returns
  None, empty cookies returns None, Freedium empty-cookies no-op
- bounded_read: Read::take stops at max+1 bytes (guard fires), allows
  exactly-at-limit payloads (guard does not fire)
2026-07-19 15:18:16 +02:00
8c19d90e94
fix(core): address Codex review findings in reader image post-processor
P1 — Enforce size cap before buffering (singlefile.rs fetch_image_as_data_uri):
resp.bytes() buffered the entire response before checking MAX_BYTES, enabling
OOM on oversized or attacker-controlled images. Fix: reject via Content-Length
header when present, then stream at most max_bytes+1 bytes with Read::take so
the cap is enforced without materialising the full body first.

P2 — Decode HTML entities in extracted image URLs (inline_archivr_img_srcs):
The browser's HTML serialiser encodes & as &amp; in attribute values, so CDN
signed URLs with query parameters (e.g. ?a=1&b=2) arrived as &amp;-escaped
strings. reqwest sent the wrong URL, breaking signed or transformed CDN
images. Fix: unescape &amp; &lt; &gt; &quot; before passing to the fetcher.

P2 — Preserve authentication for same-origin lazy images (inline_archivr_img_srcs):
The post-processor created a bare reqwest client with no cookies, causing 401/403
for reader-mode captures of auth-gated sites whose lazy images are same-origin.
Fix: pass capture_url and cookies into inline_archivr_img_srcs; attach a Cookie
header only when domain_from_url(img_url) == domain_from_url(capture_url) and
cookies is non-empty. Third-party image hosts and Freedium fetches (which receive
empty cookies in capture.rs) are unaffected.
2026-07-19 15:09:38 +02:00
571da72649
fix(core): prevent lazy-resolver double-processing on Freedium reader images
_archivrResolveLazyImgs is called twice: before Readability (_pre) and
after the post-body stamp pass (_post). The stamp pass sets data-archivr-src
but previously left data-src/data-zoom-src/etc intact, so _post's
'lazySrc && _isPlaceholder' condition fired on the same images and
rewrote src to the CDN URL — which SingleFile then tried and failed to
fetch from the Freedium proxy context, redundantly.

Fix: strip all lazy attrs (data-src, data-lazy-src, data-zoom-src,
data-original, data-lazy) when stamping data-archivr-src. _post then
finds no lazySrc on those images and skips them cleanly. Rust owns
them via data-archivr-src. _post continues to handle any remaining
placeholder images not covered by the stamp pass (non-Freedium reader
captures).
2026-07-19 14:50:59 +02:00
3d95df8d5a
fix(core): clean up Freedium chrome and drop reader-mode debug counters
Two fixes:

1. freedium_cleanup: remove remaining Freedium article chrome that
   survived the existing nav/footer/toaster pass:
   - <header class="p-6 bg-gray-50 ..."> — author/metadata bar with
     profile pictures and byline, a Freedium wrapper around the article
   - <section> containing [data-slot="dropdown-menu-trigger"] or
     [aria-haspopup="menu"] — the "Download article" dropdown

   Selectors target stable Tailwind utility class prefixes (p-6, bg-gray-50,
   bg-zinc-800) for the header and the WAI-ARIA menu role for the button,
   both resilient to minor Freedium UI updates.

2. Reader-mode meta tag: strip per-capture debug counters.
   _archivrReaderMark now writes 'applied' instead of
   'applied:pre_s=N,...:post_s=N,...' — the counters were useful during
   development but have no place as permanent archive content.
2026-07-19 14:46:22 +02:00
453a8c6487
fix(core): inline reader-mode images via Rust post-processing
SingleFile cannot inline resources added to the DOM at before-capture
time — only resources tracked during the page's initial load cycle get
embedded. Article images in Readability output fall into this gap when
the page framework has already resolved their lazy src to a CDN URL
that isn't in SingleFile's resource cache for the new DOM elements.

Fix in three parts:
- Browser script: after body.innerHTML = article.content, stamp
  data-archivr-src=<absolute-proxy-url> on any image whose src is
  not an already-inlined large data URI. Remove loading attr. Don't
  touch src (let SingleFile try; Rust handles the rest).
- Rust (save_with): after SingleFile writes the file and before hashing,
  call inline_archivr_img_srcs() to scan for data-archivr-src markers.
- inline_archivr_img_srcs(): fetches each marked URL with blocking
  reqwest (10 s timeout, 5 redirects, image/* Content-Type guard,
  20 MiB cap), base64-encodes, replaces src with the data URI, and
  removes the marker attr. Non-fatal — fetch failures are logged.

Also: Freedium UI cleanup now removes footer, #progress, and empty
data-nosnippet wrappers in addition to the existing nav/toaster removal.
2026-07-19 14:06:54 +02:00
331fe7fd61
fix(core): extract HTML title after font stripping, not before
SingleFile embeds fonts as base64 data URIs in <style> blocks in the
<head>, pushing the <title> tag to ~1.2 MB in the raw temp file.
The 256 KiB read window in extract_html_title missed it.

Font extraction rewrites the HTML in-place (2.4 MB → 1.26 MB) before
hashing, so the title is accessible at byte ~106 KB in that content.

Fix: add extract_html_title_str() that operates on a &str; call it on
the in-memory rewritten string after font extraction (server path).
CLI path (no font extraction) falls back to result.title as before.
2026-07-19 13:15:35 +02:00
f9759c08a3
fix(core): Freedium capture fixes — title, notifications, reader images
- extract_html_title: take().read_to_end() up to 256 KiB (single read() can
  short-read, leaving title at byte ~106 K undiscovered); regression test added
- Strip " - Freedium" suffix before storing entry title
- is_freedium_fetch keys off actual fetch_url host
- Remove Freedium toast overlay ([data-sonner-toaster]) before capture
- Resolve lazy images (data-zoom-src etc.) before Readability in reader mode
2026-07-19 12:49:18 +02:00
5db18122f7
feat(frontend): add Freedium mirror toggle to capture advanced options
- freediumEnabled state defaults to true (on by default)
- via_freedium forwarded through submitCapture to the capture API
- Toggle rendered last in the advanced panel, matching existing rows
- Built frontend static assets included
2026-07-19 12:09:42 +02:00
b33bf8dbd7
feat(server): expose via_freedium in capture API (default on)
CaptureBody gains via_freedium: Option<bool>; absent defaults to true.
The rearchive handler sets it to false — existing entries should not be
silently re-fetched through a mirror.
2026-07-19 12:09:37 +02:00
a926d7cf2d
feat(core): add via_freedium to CaptureConfig; route WebPage captures through Freedium mirror
When via_freedium is true and the locator is not already a Freedium URL,
perform_capture passes https://freedium-mirror.cfd/<original-url> to
singlefile::save() so paywalled articles are fetched via the mirror.
The original locator is kept for requested_locator and canonical_locator
in the DB entry. An empty cookie map is used for the mirror fetch to
prevent original-domain credentials from being sent to freedium-mirror.cfd.
2026-07-19 12:09:30 +02:00
289037235c
feat(tags): revamp tags tab (#29)
feat(tags): revamp tags tab — tooltips, entry counts, Create/Move flows, Esc handling (#29)
2026-07-19 11:02:57 +02:00
a4de506495
frontend: Esc deselects entry; group font artifacts in rail 2026-07-18 21:38:22 +02:00
3407122303
feat: style X article standalone preview to match x.com
- Apply X article typography (font-size, line-height, font-family) and
  layout to the standalone preview tab (PreviewPage)
- Match scrollbar styling to x.com measurements
- Tighten body line-height to 1.5 (25.5 px at 17 px base)
- Refactor TweetPreview and PreviewPanel to share the updated styles
2026-07-18 20:20:27 +02:00
cd463d2810
feat: multi-select entries with bulk delete, tag, and collection actions (#28)
* feat: multi-select entry rows (shift/ctrl+click, mobile checkbox)
* feat: bulk-action panel for multi-selected entries
2026-07-15 16:50:37 +02:00
278dc928df
feat: integrate Modal Closer as injected browser script (#27)
Port the modalcloser abx-plugin as a SingleFile --browser-script rather
than a spawned Puppeteer daemon. No external extension download required.

Core behavior (singlefile.rs):
- MODAL_CLOSER_DIALOG_OVERRIDES: main-world <script> bridge (best-effort;
  blocked by strict script-src CSP) that overrides window.alert/confirm/
  prompt/print, nulls window.onbeforeunload, traps its setter, and wraps
  window.addEventListener to no-op 'beforeunload' registrations.
- MODAL_CLOSER_POLLING_SETUP: defines _archivr_mc_run() which runs two
  passes on every tick (immediate first run, then setInterval every 500 ms
  matching MODALCLOSER_POLL_INTERVAL default):
    Pass 1 (best-effort, CSP-sensitive): main-world <script> bridge calls
    Bootstrap/jQuery/jQuery UI/SweetAlert teardown APIs.
    Pass 2 (always, CSP-immune): isolated-world DOM mutations — Escape-key
    dispatch (Radix/Headless UI/Angular Material), backdrop clicks, full
    CSS selector hiding for 40+ named consent/overlay vendors, body
    scroll-lock reset. Direct DOM mutations need no inline script execution.
- resolve_modal_closer_config(): reads ARCHIVR_MODAL_CLOSER env var
  (default true); no external resource required.
- modal_closer_enabled: Option<bool> in CaptureConfig so Default::default()
  yields None (follow env var) rather than false.

Schema (database.rs):
- modal_closer_enabled column in auth DB instance_settings (default 1).
- Idempotent ALTER TABLE migration for existing databases.
- get/update_instance_settings updated (SELECT col 6, UPDATE param ?7).

HTTP (routes.rs):
- modal_closer_enabled in CaptureBody and UpdateInstanceSettingsBody.
- Per-capture body overrides global setting which overrides env var.

Frontend:
- SettingsView ExtensionsTab: third ext-card for Modal & Dialog Closer.
- CaptureDialog: per-capture toggle in Advanced Options, state initialized
  from server default, included in submitCapture payload.
- api.js: submitCapture forwards modal_closer_enabled to capture body.

Behavior vs original modalcloser daemon:
- Functionally identical on non-strict-CSP pages (the large majority).
- Gap: <script> bridge is subject to page CSP; page.evaluate() is not.
  On strict-CSP pages framework teardown and dialog overrides are no-ops;
  CSS selector hiding and Escape dispatch (Pass 2) always work.
- Gap: alert/confirm/prompt overridden to instant no-op vs CDP timed
  dialog.accept() with 1250 ms delay; irrelevant for archival in practice.
2026-07-14 13:47:47 +02:00
github-actions[bot]
b0d02034ea
chore: yt-dlp 2026.06.09 → 2026.07.04 (#25)
Co-authored-by: thegeneralist01 <180094941+thegeneralist01@users.noreply.github.com>
2026-07-13 14:58:29 +02:00
39765ef893
feat(frontend): small dashboard revamp (#26)
Add ⌘K search focus, URL-param'd search/tag/entry, persistent selection, extensions grid

* feat(frontend): focus search input on ⌘K/Ctrl+K

Add a global keydown handler that focuses (and selects) the search
input when Meta+K or Ctrl+K is pressed. If the current view is not
'archive', switch to it first and focus the input after the next
render frame via a pendingSearchFocus ref + requestAnimationFrame.

The ⌘K hint badge in the toolbar was already present; this wires
the behaviour behind it.

* feat(frontend): persist search query and tag filter in URL params

Extend parseLocation() to read ?q= and ?tag= search parameters.
Initialise searchQuery and tagFilter from the URL on mount so that
sharing or refreshing a filtered view restores the exact same
results.

- replaceState on every q/tag change (no extra history entries)
- pushState on view navigation now preserves existing search params
- popstate handler restores q and tag on back/forward
- firstArchiveLoad ref prevents the archive-change effect from
  wiping URL-initialised filters on the first mount

* feat(frontend): persist selected entry in URL params

Extend parseLocation() to read ?entry= and initialise
selectedEntryUid from it on mount.

When entries load (or reload), a new effect checks whether
selectedEntryUid is set without a corresponding selectedEntry
object and restores it by finding the matching entry in the list.
This covers page refresh, URL sharing, and back/forward navigation.

The URL params sync effect now includes selectedEntryUid alongside
q and tag, and the popstate handler restores all three.

Also includes rebuilt frontend static assets.

* fix(frontend): add cookies and extensions to settings URL routing

SETTINGS_TABS was missing 'cookies' and 'extensions', so navigating
to /settings/cookies or /settings/extensions silently fell back to
the profile tab. Both tabs already existed in SettingsView; they
just weren't recognised by parseLocation().

Includes rebuilt frontend static assets.

* style(frontend): ext cards in auto-fill grid

* fix(frontend): URL state regressions from Codex review

Two issues:

1. Re-selection stall when popstate lands on the same entry UID.
   The restoration effect only had [entries, selectedEntryUid] as
   deps, so setting selectedEntry=null without changing either dep
   left the rail blank. Adding selectedEntry to the dep array lets
   the null→restore cycle complete.

2. tag/entry params leaking onto non-archive URLs.
   pushState on view changes carried window.location.search verbatim,
   so /settings?tag=foo was possible; reloading it triggered the
   tagFilter effect and force-switched back to archive.
   Fix: parseLocation() only extracts tag/entry when view=archive,
   and the params-sync effect only emits them on archive too.
2026-07-13 14:50:52 +02:00
72cf7ff642
docker: add uBlock Lite + ISDCAC extensions, mirror flake.nix parity
- Install unzip in runtime stage (needed to extract extension zips)
- Download uBlock Origin Lite (v2026.705.2152) and I Still Don't Care
  About Cookies (v1.1.9) during build, matching flake.nix versions
- Verify each zip against its SHA-256 (converted from Nix SRI hashes)
  before extraction so a changed GitHub release fails fast, same as Nix
- Check manifest.json at ISDCAC extension root post-extraction, mirroring
  the flake installPhase guard
- Set ARCHIVR_UBLOCK_EXT and ARCHIVR_COOKIE_EXT env vars to the extracted
  extension directories
2026-07-12 17:09:17 +02:00
d610d37793
feat: entry previews (#24)
* feat: entry previews (video, tweet, article, iframe, image, audio bar)

- Add PreviewPanel dispatch hub routing by entity_kind + artifact extension
- VideoPreview: HTML5 <video> for YouTube/Instagram/TikTok/Reddit/X posts
- TweetPreview: tweet card, thread, and X article renderer (ported from x-article-renderer)
- IframePreview: sandboxed iframe for SingleFile web pages and PDFs
- ImagePreview: image viewer with click-to-open-fullsize
- AudioBar: persistent fixed-bottom player (Spotify-style) that survives entry navigation
- Lift entryDetail to App.jsx, shared between PreviewPanel and ContextRail
- 3-column layout (workspace | 300px preview | 340px rail) when preview active
- Stale-guard fixes: seq incremented before early returns in all async effects
- handleRearchive: capture startSeq/entryUid at call time, guard every async resume
- TweetPreview: reset loading/error/tweets before early-return branches

* feat: entry previews — tweet/thread/article/video/audio/image/iframe/pdf

- PreviewModal: modal overlay with new-tab link (↗) and keyboard close
- PreviewPanel: routes by entity_kind + primary_media extension to the
  correct viewer (tweet/video/audio/pdf/html/image/fallback)
- TweetPreview: full X-style tweet, thread, and article renderer with
  local artifact map for archived media (CDN fallback)
- AudioBar: persistent fixed bottom player, triggered via ContextRail
  Play button; body.has-audio-bar pads content above it
- VideoPreview, IframePreview, ImagePreview: inline viewers
- PreviewPage: standalone /preview/:archiveId/:entryUid route
- ContextRail: Play/Preview buttons; isAudio/isPreviewable detection
- App.jsx: preview modal state, currentAudio state, preview route guard,
  has-audio-bar body class effect
- routes.rs: CSP updated (media-src self blob https; frame-ancestors self;
  Google Fonts + external images/scripts whitelisted)
- styles.css: preview modal, tweet-wrap scroll (min-height:0), audio bar
  body padding, newtab button, preview panel flex layout

* style: tighten article/tweet preview spacing

- aMeta padding: 14px → 10px
- article title marginBottom: 10px → 8px
- aAuthorRow marginBottom: 10px → 8px
- bH1 top margin: 20px → 16px
- bH2 top margin: 18px → 14px
- bHr margin: 20px → 14px
- .preview-tweet-wrap padding: 20px → 12px (already committed)

More content visible above the fold in both modal and standalone views.

* feat: tweet/article preview quality pass

HTML entities: decode &gt; &lt; &amp; etc. on sliced segments only
(entity offsets index the stored string; decoding before slicing shifts them)

Image lightbox: click any tweet/thread/article image to open full-screen
viewer; cmd+click follows <a> to open in new tab; arrow-key + ‹ › nav;
Escape closes; 1/N counter; ↗ open-in-new-tab link

Multi-image grid: 2 photos → side-by-side (180px rows); 3 → left spans
both rows; 4 → 2×2 (140px rows); single image unchanged

Empty media grid ghost: build photos/videoItems arrays first, render
.mediaGrid div only when at least one item resolved (advisory: never gate
on raw media.length when map items can all return null)

QT indicator: ↻ QT badge on tweet.is_quote_status === true entries

Modal shrinks for short content: height: 88vh → max-height: 88vh;
.preview-modal-body gets max-height: calc(88vh - 52px) so long threads
still scroll (advisory: don't rely on flex:1 once parent has no fixed height)

Video scrollbar leak: .preview-modal-body overflow: auto → hidden; each
child (tweet-wrap, video-wrap, iframe) manages its own scroll surface

Styled thin scrollbar on .preview-tweet-wrap (matches workspace rail)

ArticleRenderer: cover image and body images are lightbox-clickable;
opts thread through renderBlocksJSX → renderBlockJSX → renderAtomicJSX

* fix: move artifact fetch to api.js; stop Escape propagation from lightbox

- Export fetchEntryArtifacts(archiveId, entryUid, indices) from api.js
  using Promise.all + getJson (follows project convention: all /api calls
  go through api.js, never inline fetch in components)
- TweetPreview: import fetchEntryArtifacts, replace inline Promise.all
- MediaLightbox keydown handler: stopPropagation + preventDefault for
  Escape/ArrowLeft/ArrowRight so the parent PreviewModal window listener
  does not also fire and close the modal behind the lightbox

* fix: iframe/page preview height chain and toolbar UX

Problem: changing .preview-modal from height to max-height broke iframe
previews - <iframe style='flex:1'> needs a concrete ancestor height, which
max-height alone doesn't supply when content is shorter than the cap.

Fix - CSS:
  .preview-modal--full { height: 88vh } applied to non-tweet modals
  .preview-modal--full .preview-modal-body { max-height: none }
  .preview-iframe-toolbar span: remove text-transform/letter-spacing
    (was uppercasing the URL/title in shouty caps)

Fix - PreviewModal: className adds --full when entity_kind is not
  tweet/tweet_thread; tweet previews keep shrink-to-fit behavior.

Fix - PreviewPanel: pass title + original_url from summary to IframePreview
  for both HTML and PDF; wrappers use flex:1/minHeight:0 not height:100%.

Fix - IframePreview:
  - Accept title + originalUrl props; show originalUrl in toolbar (falls
    back to artifact src only when original_url absent); show title above
    URL when available
  - flex:1 + minHeight:0 instead of height:100% on the wrap div
  - Single unified layout for page + pdf (both just show the iframe)

* feat: expand t.co links; linkify bare URLs in tweet and article text

Frontend:
- resolveEntityBounds: try multiple candidate strings in order (u.url
  first, since that's the t.co short URL that appears in full_text)
- normalizeUrlAnn: multi-candidate search; href = expanded > url,
  display = display_url > expanded > url
- linkifyText(): regex linkifier for entity-less bare URLs; trims
  trailing punctuation [.,;:!?)] before linking; used in both
  renderTweetTextJSX and renderInlineJSX including their early-return
  paths (anns.length === 0) that previously bypassed linkification
- renderInlineJSX: fix mention href mention.name → screen_name;
  replace t.co segment text with url.display when entity covers it

Scraper (vendor/twitter/scrape_user_tweet_contents.py):
- extract_tweet_data: when note_tweet text is used, pull urls/mentions/
  hashtags/symbols from note_result.entity_set (correct indices for the
  note text); keep media from legacy.entities (no note media downloads)

* feat: server-side t.co resolver + frontend augmentation

Server (routes.rs):
  POST /api/util/resolve-tco — unauthenticated, accepts JSON array of
  https://t.co/<alphanumeric> URLs only (strict regex validation, no SSRF
  via input), capped at 50 per batch, 3 s timeout, redirect(Policy::none)
  so the server only ever touches t.co itself. HEAD first, GET fallback if
  HEAD returns no Location. Location sanitized to http/https only —
  javascript:/data:/etc. fall back to the original t.co.

api.js:
  resolveTcoUrls(urls) — project-convention wrapper for the new endpoint.
  Returns {} on failure (callers degrade gracefully to bare t.co links).

TweetPreview.jsx:
  After tweet data loads, per-tweet range-based coverage detection:
  builds covered [start,end) from existing entity fromIndex/toIndex or
  indices fallback, then finds regex matches whose span is NOT covered.
  Resolves unique uncovered t.co URLs via resolveTcoUrls(), synthesises
  one entity per occurrence (with exact fromIndex/toIndex so normalizeUrlAnn
  gets correct bounds even for duplicate t.co URLs in the same tweet).
  Augments entities.urls before setTweets() so all rendering paths
  see expanded URLs.

* feat: suppress rendered media attachment URLs from tweet text

renderTweetTextJSX now accepts skipSpans=[] as third param.
Skip-span boundaries are added to the pts split set so a trailing
media t.co inside a plain segment still gets isolated and suppressed—
not re-linked by linkifyText. Early return only when both anns and
skipSpans are empty.

TweetCard computes mediaSkipSpans after building photos/videoItems:
for each rawMedia item whose src resolved (photo src match; any video),
resolveEntityBounds(m, ft, m.url) gives the precise [s,e] span using
indices/fromIndex first, indexOf fallback—then the span is passed to
renderTweetTextJSX so the t.co attachment URL is silently dropped.
2026-07-12 16:23:04 +02:00
2779afee2d
feat(tweets): add re-archive button and fix thread-tweet orphan cleanup (#23)
Orphan cleanup bug: archiving x🧵A downloaded D/C/B/A JSONs and
media, but only registered artifacts for A. D/C/B files had no
entry_artifacts rows and were deleted as orphans.

Fix (staged scraper output, precise touched set):
- tweets::archive() stages all scraper output in temp/{ts}/tweet_stage/,
  validates, then renames JSONs to raw_tweets/. Return type changed from
  Result<bool> to Result<Vec<String>> (store-relative relpaths of every
  produced tweet JSON, i.e. the exact touched set).
- tweets::rearchive() (new): same staged approach but always runs the
  scraper. On scraper failure (tweet deleted/private), errors before
  touching raw_tweets/ so existing data is preserved.
- register_tweet_artifacts() (new private helper in capture.rs): registers
  every JSON in the touched set as a raw_tweet_json artifact, parses each
  for media blobs, registers those too. JSON read failure is a hard error
  with context, not a silent skip.
- record_tweet_entry() now accepts tweet_json_relpaths: &[String] and
  delegates artifact registration to register_tweet_artifacts().
- perform_capture() passes the returned vec from tweets::archive().

Re-archive feature:
- capture::perform_rearchive(): looks up entry by uid, validates
  tweet/tweet_thread, runs tweets::rearchive(), atomically swaps
  entry_artifacts in a DB transaction. archived_at, title, tags,
  collections are untouched.
- database: add get_entry_for_rearchive() and delete_entry_artifacts().
- POST /api/archives/:id/entries/:uid/rearchive: requires ROLE_USER,
  creates capture job, returns 202 + job_uid, runs perform_rearchive in
  spawn_blocking.
- Frontend: re-archive button in ContextRail for tweet/tweet_thread
  entries; polls job at 500ms; refreshes entry detail on success; shows
  error text on failure. Poll interval cleared before early-return on
  entry deselect to prevent stale updates.
2026-07-11 16:14:58 +02:00
03390362c5
capture: close dialog on Archive; rich per-URL and batch toast notifications (#22)
* http: realistic UA; fall back to WebPage on probe failure

Replace bare 'archivr/0.1' user agent with a full Chrome 131 UA string
(archivr/0.1 token retained at the end) in both probe_url_kind and download.

More importantly, stop hard-failing when probe_url_kind returns an error.
Sites behind Cloudflare's managed JS challenge (e.g. Medium) return 403
before any header tuning can help — a plain HTTP client cannot pass the
challenge. Instead, log a warning and fall back to Source::WebPage so that
the SingleFile/Chromium path gets a chance; a real browser can solve the
challenge transparently.

* capture: close dialog on Archive; rich per-URL and batch toast notifications

UX changes:
- Pressing Archive closes the capture dialog immediately; jobs continue
  polling in the background (component stays mounted).
- Probe failures (e.g. Cloudflare JS challenge, HTTP 403) now fall back to
  Source::WebPage so SingleFile/Chromium can attempt the capture instead of
  hard-failing at the probe stage.

Toast notifications:
- Single URL: green 'Archived' on success, amber 'Archived with warnings'
  (with expandable detail) when uBlock/cookie-ext was skipped, red 'Capture
  failed' with expandable error text on failure. Per-item warning/error toasts
  include the locator so the user knows which URL was affected.
- Multi-URL batch: per-item failure and warning toasts still fire with
  locators; per-item success toasts are suppressed. Once all jobs settle a
  single summary toast fires: 'N archived', 'N archived (M with warnings)',
  'N archived (M with warnings), F failed', or 'N failed'. The summary Detail
  section lists the exact URLs that failed or warned, so the user retains that
  information after per-item toasts auto-dismiss.
- Batch summary color: green = all clean; amber = any warnings or failures
  present; red = all failed.
- handleIgnoreUblock now only removes per-item warning toasts (those with a
  locator) and persists the ignore flag; batch summary warnings (locator=null)
  are not swept.

ToastStack improvements:
- All three branches (success/warning/error) support a toast.headline field
  so batch summaries can set their own copy.
- Warning branch: Details button conditional on toast.text; Ignore button
  conditional on toast.locator (uBlock-specific, not shown on batch summaries).
- Locator display uses hostname/…/last-segment for URLs (preserves domain for
  context, tail for identity) and tail-truncation for shorthands; full locator
  in title attribute for hover.
- Icon colors: green for success (✓), amber for warning (⚠), red for error (✕).
2026-07-11 15:29:10 +02:00
2e8820a0da
feat: uBlock Origin Lite + cookie consent extension + reader mode + ad placeholder cleanup (#21)
* feat: uBlock Origin Lite integration for ad-blocking during WebPage captures

- singlefile.rs: when ARCHIVR_UBLOCK=true and ARCHIVR_UBLOCK_EXT is set,
  archivr owns Chrome's lifecycle (--headless=new, --remote-debugging-port,
  --load-extension); single-file connects via --browser-server instead of
  launching its own Chrome. Falls back to old behaviour with ublock_skipped=true
  when the ext path is missing or invalid.
- capture.rs: thread ublock_skipped through CaptureResult
- database.rs: add notes_json TEXT column to capture_jobs (DDL + idempotent
  ALTER TABLE migration); update_capture_job_status gains notes_json param
- archive.rs: expose notes_json in CaptureJobSummary
- routes.rs: store {"ublock_skipped":true} in notes_json on completed captures
- ToastStack.jsx: warning toast variant (toast--warning) with Details expander
  and Ignore button
- CaptureDialog.jsx: fire warning toast when poll result has ublock_skipped
- App.jsx: sessionStorage-backed Ignore suppression for ublock warnings
- styles.css: .toast--warning (amber left border) + .toast-warning-detail
- flake.nix: ublockLite derivation fetches uBOLite_2026.705.2152.chromium.zip
  (pinned SHA256) from uBlockOrigin/uBOL-home; sets ARCHIVR_UBLOCK_EXT in both
  archivr and archivr-server wrappers

Env vars:
  ARCHIVR_UBLOCK=true (default) — enable uBlock during WebPage captures
  ARCHIVR_UBLOCK_EXT — path to unpacked uBOL extension dir (set by Nix)

* feat: Extensions settings tab + capture dialog redesign with Advanced options

Settings/Extensions tab (admin-only):
- New 'Extensions' tab between Cookies and Storage
- ExtensionsTab component: shows uBlock Origin Lite card with pill toggle
- Reads ublock_enabled from instance settings; patch via existing PATCH endpoint
- Shows ublock_ext_available status from server (whether ARCHIVR_UBLOCK_EXT is set)

Instance settings:
- Add ublock_enabled BOOLEAN (default true) to instance_settings auth DB table
- Idempotent ALTER TABLE migration in initialize_auth_schema()
- get/update_instance_settings include ublock_enabled
- GET /api/admin/instance-settings now also returns ublock_ext_available (computed
  from ARCHIVR_UBLOCK_EXT env var at request time)
- PATCH /api/admin/instance-settings accepts ublock_enabled

Per-capture override:
- CaptureBody gains ublock_enabled: Option<bool>
- CaptureConfig gains ublock_enabled: Option<bool>
- singlefile::save() gains ublock_enabled_override: Option<bool> param
- Capture handler resolves: body override > global instance setting > env var
- submitCapture(aid, loc, qual, extensions) in api.js passes ublock_enabled

Capture dialog redesign:
- Archive button: full-width, 13px padding, min-width 220px, primary CTA
- Cancel: full-width but text-style, below Archive
- ‹Advanced options› chevron toggle (rotates on open)
- Expanded panel shows uBlock toggle for this capture session
- Loads global ublock_enabled default from instance settings on mount

Styles:
- .ext-toggle pill switch (44×24 and 36×20 small variant)
- .ext-card for Settings Extensions tab
- .capture-advanced + .capture-advanced-panel + .capture-chevron
- .capture-ext-row / .capture-ext-label / .capture-ext-name / .capture-ext-desc
- .form-hint utility class

* fix: remove ublock_enabled from INSERT OR IGNORE in DDL batch

The INSERT ran before the ALTER TABLE migration added the column,
causing 'table instance_settings has no column named ublock_enabled'
on existing databases. The INSERT OR IGNORE for the default row only
needs the original columns; the migration's DEFAULT 1 handles the
new column for existing and new rows alike.

* feat: Reader mode via Mozilla Readability.js

Adds an opt-in 'Reader mode' advanced option to the capture dialog.
When enabled, Readability.js is injected as a browser script during
SingleFile capture; it fires on single-file-on-before-capture-start,
replaces the page body with the distilled article content, injects a
clean typographic stylesheet, and adds a header with title/byline/site.
Falls back silently if Readability fails (e.g. non-article pages).

- vendor/readability/Readability.js  Apache 2.0, Mozilla, v0.6.0
- singlefile.rs: embed READABILITY_JS + READER_MODE_WRAPPER_JS via
  include_str!; write both to temp dir when reader_mode is true;
  base_single_file_cmd now accepts &[&Path] for multiple --browser-script
- capture.rs: CaptureConfig.reader_mode: bool
- routes.rs: CaptureBody.reader_mode: Option<bool> (defaults false)
- api.js: submitCapture passes reader_mode in payload
- CaptureDialog.jsx: Reader mode toggle in Advanced options (off by default)

* fix: diagnose single-file no-output-file error + prevent stdout dumping

- Add --dump-content=false to every single-file invocation to prevent
  the Docker-detection heuristic from routing HTML to stdout instead of
  the output file (the heuristic can trigger in some macOS environments)
- Improve the no-output-file error message to include: temp dir contents,
  stderr, and first 200 chars of stdout — this gives enough context to
  diagnose any remaining cause without re-running

* fix: switch uBlock loading from --browser-server to --browser-args

The --browser-server (CDP) path caused 'Unexpected server response: 404'
on macOS Chrome because simple-cdp's WebSocket upgrade to the debugger
endpoint failed after Chrome started — likely a version-specific CDP
endpoint shape mismatch.

New approach: single-file always manages Chrome. When ARCHIVR_UBLOCK_EXT
is set, --headless=new, --load-extension, and --disable-extensions-except
are injected via --browser-args. single-file's browser.js prefix-strips
its own conflicting flags before appending ours, so --headless=new
overrides the default --headless (enabling extension support in headless).

Removes allocate_free_port, wait_for_chrome_ready, run_single_file_with_server
(all dead code now). Docblock updated to reflect actual behaviour and notes
the --single-process caveat: uBOL's declarativeNetRequest static rulesets
are expected to work (network-stack level, not service-worker), but this
has not been mechanically verified under --single-process.

Smoke tested on macOS (this machine): capture with --load-extension + all
three browser-scripts (strip, Readability, reader-mode wrapper) produces
output file correctly. Ad-blocking verification deferred to manual test
with a tracker-heavy URL.

* fix: use correct single-file hook event (single-file-on-before-capture-request)

Prior scripts listened on 'single-file-on-before-capture-start' which
does not exist in single-file-core 1.1.49.  The real hook is:

  single-file-on-before-capture-request  (dispatched by initUserScriptHandler
  after receiving single-file-user-script-init; userScriptEnabled defaults
  to true in args.js so it always fires when --browser-script is passed)

Changes:
- strip-scripts: -start -> -request (no preventDefault needed; synchronous)
- READER_MODE_SCRIPT: -start -> -request; add 'installed' meta marker at
  script-evaluation time so artifact inspection can distinguish 'script
  not injected' / 'hook never fired' / 'Readability parse failed'

* fix: correct singlefile.rs docstring (scripts.js concatenates, not isolates)

* fix: dispatch single-file-user-script-init so request hook fires

single-file's initUserScriptHandler (in single-file-bootstrap.js) listens
for 'single-file-user-script-init' and only then installs
_singleFile_waitForUserScript.  Without that dispatch our scripts'
'single-file-on-before-capture-request' listeners were never reached,
so neither strip-scripts nor reader-mode Readability applied.

Dispatch the init event at the top of strip-scripts (always present) and
redundantly in READER_MODE_SCRIPT.  Verified end-to-end: artifact for
run_b3181d6d276e4e56a1a6c356ef9bbe8f has
  meta content="applied", max-width:680px CSS, 0 script tags.

* feat: cookie consent extension support (ARCHIVR_COOKIE_EXT)

Mirrors the uBlock Origin Lite integration exactly:

Backend:
- singlefile.rs: resolve_cookie_ext_config() reads ARCHIVR_COOKIE_CONSENT
  (default true) + ARCHIVR_COOKIE_EXT path; extension paths comma-joined
  into --load-extension / --disable-extensions-except so uBlock and cookie
  ext can coexist; SaveResult.cookie_ext_skipped tracks miss
- database.rs: cookie_ext_enabled column on instance_settings (DEFAULT 1);
  idempotent ALTER TABLE migration; get/update wired through
- capture.rs: CaptureConfig.cookie_ext_enabled: Option<bool>; threaded to
  singlefile::save(); cookie_ext_skipped surfaced in CaptureResult
- routes.rs: CaptureBody + UpdateInstanceSettingsBody get cookie_ext_enabled;
  capture handler resolves effective value (body overrides global); notes_json
  only includes skipped fields that are true; GET instance-settings includes
  cookie_ext_available from env path check

Frontend:
- api.js: submitCapture forwards cookie_ext_enabled
- SettingsView.jsx: 'I Still Don't Care About Cookies' card in Extensions
  tab; always-active toggle (user can disable even when ext not installed);
  amber 'Not configured' hint + ARCHIVR_COOKIE_EXT guidance when unavailable
- CaptureDialog.jsx: 'Block cookie banners' toggle in Advanced options;
  always shown with amber hint when ext not configured; defaults from
  global setting

Operator setup: download + unzip the extension from GitHub releases, set
ARCHIVR_COOKIE_EXT=/path/to/unpacked/ext. No Node daemon needed.

* fix: surface cookie_ext_skipped warning toast in CaptureDialog

* feat: package istilldontcareaboutcookies in flake, wire ARCHIVR_COOKIE_EXT

Add isdcac derivation mirroring ublockLite:
- Fetches ISDCAC-chrome-source.zip v1.1.9 from GitHub releases
- Validates manifest.json at extension root before install (guard against
  nested-folder zip regressions in future releases)
- Sets ARCHIVR_COOKIE_EXT in both archivr and archivr_server wrappers

Verified: nix build .#archivr-server and .#archivr both succeed;
wrapper scripts export correct store paths; manifest.json present at root.

* fix: gate consent-overlay cleanup on cookie_ext; reset overflow; narrow selectors

- Strip overflow:hidden from body/html only when cookie_ext is active for
  the capture — prevents mutating legitimate pages when the feature is off
- Remove .fc-dialog (Google Funding Choices), .qc-cmp2-*, .sp-message-container,
  #sp-cc, #usercentrics-root as fallback for CMPs the extension misses
- Removed overbroad [class^="uc-"] and [id^="usercentrics"] selectors
  that could match real page content

* fix: remove ad placeholders when uBlock active; kept height causes blank gap

uBlock Origin Lite blocks ad network requests but first-party placeholder
elements (ins.adsbygoogle, #aswift_* iframe hosts) retain their computed
height (e.g. 280px for a top banner), leaving a large blank space at the
top of captured pages.

Gate cleanup on ublock_ext.is_some(): remove ins.adsbygoogle, aswift_*
iframes, and google_ads_* iframes before SingleFile serialises. Also
collapse the parent container if it becomes empty after removal.

* fix: walk up to .top-ad/.google-auto-placed ancestor before removing ad slot

Removing only the inner ins.adsbygoogle left the outer .container.top-ad
wrapper (with pb-4 padding) in the layout, preserving the blank gap.
Now walk up via closest() to the nearest ad-slot container class before
removal so the whole slot including padding collapses.
2026-07-08 23:26:48 +02:00
dae61e585d
feat: add user-configurable cookie rules (#20)
Adds per-instance cookie rules (admin-only) that are injected into
every network touchpoint during capture.

Storage:
- New cookie_rules table in the auth DB (idempotent migration)
- Rules have pattern_kind (global/wildcard/regex), optional url_pattern,
  and cookies_json (validated as string-only JSON object)

Matching (resolve_cookies_for_url):
- Global rules always apply
- Wildcard: * and ? with full metacharacter escaping; matched against
  hostname via reqwest::Url when pattern has no ://, full URL otherwise
- Regex: matched against the full URL
- Later rules in ordinal order override earlier ones per cookie name

All six network touchpoints receive resolved cookies:
- http::probe_url_kind and http::download: Cookie request header
- singlefile::save: Netscape cookie file -> --browser-cookies-file
- ytdlp::fetch_metadata and ytdlp::download: Netscape cookie file -> --cookies
- tweets::archive: semicolon credentials file -> --credentials-file
  (only when both ct0 and auth_token are present; otherwise falls back
  to ARCHIVR_TWITTER_CREDENTIALS_FILE)

Security:
- Cookie files written 0o600 (owner read/write only)
- Exact parsed hostname used as cookie domain (no PSL stripping)
- Files deleted unconditionally before any error propagates,
  including spawn failures (hold-result-then-delete pattern)
- No cookie values in process args (no --add-header exposure)

API: GET/POST /api/admin/cookie-rules, PATCH/DELETE /api/admin/cookie-rules/:uid
Frontend: Cookies tab in Settings (admin only) with rule list,
  inline edit, pattern-type selector, client-side JSON validation
CLI: CaptureConfig::default() - no behaviour change

254 tests passing (4 new cookie-rule handler tests)
2026-07-06 19:01:34 +02:00