mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 21:03:17 +02:00
14 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
e97a31c0d1
|
feat(frontend): add text-capture form to CaptureDialog
- Add submitTextCapture API client function with same error handling as submitCapture - Create makeTextItem() factory for text capture state - Implement CaptureTextRow component with title, body textarea, and MIME selector - Add 'Add text' button in capture dialog toolbar - Update handleArchive to filter and route text submissions - Modify submitBgJob to detect and submit text items via submitTextCapture - Skip probe and conflict checks for text items - Reuse job tracking and batch settlement for text captures Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> |
|||
|
1af920eb63
|
feat: share videos to TV (Chromecast + AirPlay) (#33)
* server: add scoped media-token endpoint for Cast/AirPlay auth bypass
Chromecast and Apple TV fetch media URLs as independent HTTP clients
with no session cookie. The existing serve_artifact handler requires
auth_user.require_auth(), so those devices always received 401.
Changes:
- MediaToken struct stored in AppState (Arc<Mutex<HashMap>>), scoped to
a single (archive_id, entry_uid, artifact_index) tuple with a 2-hour TTL
- POST /api/archives/:id/entries/:uid/artifacts/:idx/media-token
requires an authenticated session, verifies the artifact exists,
prunes expired tokens, mints a 43-char URL-safe token, and returns
{ url, expires_in_secs }
- serve_artifact now accepts an optional ?token= query param; a valid
scoped token bypasses require_auth() while a missing/invalid/expired
token falls through to the normal 401 path
- CSP script-src extended to include https://www.gstatic.com so the
Cast sender SDK script (injected lazily by VideoPreview) is not blocked
- 4 new tests: bare-URL still 401, tokenized fetch succeeds without
session cookie, bogus token 401, wrong-artifact-index 401
* frontend: Cast/AirPlay overlay in VideoPreview
When the user opens a video archive entry, VideoPreview now:
1. Issues a signed media token (POST .../artifacts/:idx/media-token) and
uses the returned signed URL as <video src>. This ensures the video
element's src is one that Cast devices and Apple TV can fetch without a
session cookie.
2. Lazily injects the Google Cast SDK script (cast_sender.js from
gstatic.com, now allowed by the updated CSP). Once the SDK reports
available, a <google-cast-launcher> web component appears as an overlay
button in the top-right corner of the video. Selecting a Cast device
triggers loadMedia() with the signed URL and the artifact's MIME type.
3. Detects AirPlay support (webkitShowPlaybackTargetPicker on
HTMLVideoElement) and shows an AirPlay icon button alongside Cast.
The <video> element carries x-webkit-airplay='allow', so Safari's native
controls also surface the AirPlay option. The explicit overlay button
calls webkitShowPlaybackTargetPicker() for consistent placement.
Both buttons are hidden when the respective APIs are unavailable (HTTP
pages, non-Safari for AirPlay, no Cast extension/devices), so there is no
UI regression for users who don't cast.
PreviewPanel now passes contentType (derived from artifact extension) to
VideoPreview so Cast receives a correct MIME type.
New CSS: .video-tv-controls (absolute overlay), .video-tv-btn (frosted
glass icon button), .video-tv-loading (placeholder during token fetch).
* server: fix serve_artifact auth OR logic — bogus token falls back to session
Previously a request carrying ?token=<expired> was immediately rejected
with 401, even if the user held a valid session cookie. This broke
logged-in browser playback after the 2-hour signed-URL window expired,
because VideoPreview uses the signed URL as <video src>.
Fix: compute token_valid first; if the token is absent or invalid, fall
through to auth_user.require_auth() instead of returning early.
Effect: valid token skips session check, invalid/missing token checks
session, both invalid → 401 as before.
Updated the bogus-token-no-session test docstring to clarify it tests
the no-auth path specifically. Added new test:
media_token_bogus_token_with_session_returns_200 — verifies a logged-in
user can still fetch the artifact via a URL carrying a stale token.
* frontend: guard token-fetch effect against stale async resolution
A slow issueMediaToken() response for video A could resolve after the
user selected video B and call setSignedSrc(urlA), making the
preview/Cast play the wrong file.
Add a cancelled flag set in the effect cleanup; both .then and .catch
check it before touching state, so only the most recent src wins.
* frontend: load Cast media immediately if session already exists
Previously the effect only sent video to the TV on SESSION_STARTED /
SESSION_RESUMED events. Two gaps:
1. If a Cast session was already active when signedSrc became ready
(e.g. the SDK resumed a session before the token fetch finished, or
the user switches videos while already casting), nothing was sent.
2. Same gap if castReady fired after an already-established session.
Fix: extract loadMedia(session) and call it against
ctx.getCurrentSession() immediately when castReady + signedSrc are both
truthy, in addition to keeping the event listener for future connects.
* server: staged file-upload endpoint
POST /api/archives/:id/uploads streams a multipart body to a temp file
under the archive's store/temp/ directory and returns a staged_path the
capture pipeline can move into place.
- Routes: /api/archives/:id/uploads (POST, requires auth)
- Body cap: 10 GiB; chunk-streamed to disk, never buffered in memory
- Path-traversal sanitised on the filename field
- Temp files are cleaned up on error paths (disk-leak fix)
- main.rs wires the new route into the server startup
- Cargo: adds the multipart dependency
* frontend: file upload in Capture dialog
Drag-and-drop or 'Upload file' button stages files for archiving:
- File items sit alongside URL rows in the same list; each shows the
original filename, a live progress bar during upload, and a check badge
when ready. The locator input is replaced entirely — no editable field.
- Archive button is disabled until all uploads finish; each file item
contributes to the Archive N count once its upload is done.
- File items are excluded from sessionStorage persistence (they are
transient — the staged server path would be invalid after a reload).
Staged-file cleanup is handled at every exit path so temp/uploads/ does
not accumulate:
• removeRow on an in-progress item aborts the XHR; removeRow on a done
item calls DELETE /archives/:id/uploads.
• Dialog cancel (Escape / Cancel button) aborts all in-flight XHRs and
DELETEs all completed staged files via the close-event handler.
• handleArchive sets isSubmittingRef=true before dialog.close() so the
close handler skips cleanup — the background capture job handles
staged-file removal on success instead.
• uploadFile() returns { promise, abort } so the component can cancel
the XHR without any visible fetch.
api.js additions: uploadFile (XHR with progress + abort), deleteUpload.
(Static assets rebuilt from combined source to include screensharing
changes from this branch.)
* fix: collection enrollment with default_visibility_bits
Two related fixes from feat-file-uploading:
core: fix collection enrollment using default_visibility_bits instead of
entry.visibility — entries were being enrolled with the entry-level
visibility rather than the collection's configured default.
server: allow changing default_visibility_bits on the default collection
— the PATCH handler was incorrectly blocking updates to the default
collection's visibility configuration.
* server: fix unbounded staged-upload disk growth
Two review findings:
P2 — delete staged file on capture failure (routes.rs)
When perform_capture returns Err, the job was marked failed but
staged_upload_path was never removed. With a 10 GiB body cap a few
failed imports could exhaust archive storage before the next restart.
Mirror the success-path cleanup into the Err arm so the file is removed
immediately regardless of outcome.
P1 — periodic staged-upload pruning (main.rs)
The startup prune of temp/uploads/ only ran once, so uploads abandoned
mid-session (browser crash, navigation away) accumulated forever on a
long-running server. Folded the pruning logic into the existing 24 h
maintenance task alongside session cleanup, so stale dirs are swept
continuously without requiring a restart.
* server+frontend: fix staged-upload disk-growth and prune safety
Server (main.rs + routes.rs):
- Extract prune_stale_upload_dirs() helper called by both startup and
the periodic 24h task, eliminating the duplicated loop.
- Sentinel (.uploading) created in the UUID dir before streaming begins;
removed on successful completion; error path uses remove_dir_all so
the partial file and sentinel are cleaned up together.
The periodic prune skips any dir containing .uploading (active XHR).
- Startup prune passes cleanup_stale_sentinels=true: the server has not
started accepting connections yet so any sentinel is a crash remnant —
it is removed and the dir proceeds to the age check, preventing leaked
dirs from a previous crash accumulating forever.
- Staleness measured from the newest non-sentinel child file mtime so a
just-finished slow upload (dir mtime stale, file mtime fresh) is not
pruned before the user can submit it for capture. Empty dirs fall back
to dir mtime.
- Failed captures (Err branch in spawn_blocking) now also delete the
staged file and UUID dir immediately, matching the success path.
Frontend (api.js + CaptureDialog.jsx):
- submitCapture attaches err.status = res.status on non-2xx responses
so callers can distinguish a definite HTTP rejection from a network
error where the response may have been lost.
- submitBgJob catch deletes the staged file only when e.status is set
(server definitively rejected the POST /captures request). A network
error leaves the file in place because the server may have accepted
the job and the response was lost — deleting would race the capture.
|
|||
|
6377daadae
|
feat: add capturing w/ parent-child entries (#32)
* feat(core): YouTube playlist/channel/YTM-playlist capture with parent–child entries
- ytdlp: add fetch_playlist_info() using yt-dlp -J --flat-playlist for
reliable container title + shallow entry list; normalize item URLs via
webpage_url → absolute url → id fallback (domain inferred from container
URL so YTM stays on music.youtube.com)
- capture: add record_container_entry() (no blob, no primary_media artifact);
extend record_media_entry() with parent_entry_id/root_entry_id params (all
existing single-item call sites pass None, None)
- capture: implement YouTubePlaylist / YouTubeChannel / YouTubeMusicPlaylist
capture path replacing the two not-implemented stubs: fetch playlist info →
create container entry (reusing existing run + item) → per-child run items
(parent_item_id = container item) → download each video/track as a child
entry; per-child failures are non-fatal; perform_capture returns result.status
reflecting actual run outcome so capture_handler marks the job correctly
- archive: add child_count i64 to EntrySummary (col 12 in all listing queries);
add get_entry_summary() private helper; fix get_entry_detail() to use
get_entry_summary() so child entries are resolvable via the detail endpoint;
add list_child_entries(conn, uid, caller_bits) with the same
admin/collection visibility predicate as list_root_entries
archive_runs.requested_count stays 1 (one user locator); discovered/
completed/failed_count reflect container item + N video items via
refresh_run_counters.
* feat(server,frontend): expose children endpoint + expand UI for container entries
server:
- add GET /api/archives/:id/entries/:uid/children → list_entry_children,
calling list_child_entries with caller_bits so visibility model is enforced
- fix capture_handler: use result.status ("completed"/"failed") to set job
status rather than always "completed", mirroring rearchive_handler; this
surfaces partial playlist failures to the polling client
frontend:
- api.js: add fetchEntryChildren(archiveId, entryUid)
- EntryRow: one outer div.entry-row-outer (display:block) keeps nth-child
striping correct; inner div.entry-row-main is the flex row with all column
cells and event handling; .child-entries sits below inside the outer wrapper
- expand chevron appears when entry.child_count > 0; clicking fetches children
lazily and renders ChildRow components reusing .col-* flex widths
- child-count badge shown next to title on container entries
- styles.css: scoped CSS with #entries-body > .entry-row-outer selectors
(higher specificity than > div) to override flex on outer wrapper; inner row
and column rules replicated at correct depth; nth-child, is-selected,
is-multi-selected, url-cell hover all handled
* fix(frontend): make child entry rows interactive
ChildRow now receives onRowClick and selectedUids from EntryRow (which
receives selectedUids from EntriesView alongside the existing booleans).
Clicking a child row invokes onRowClick(child, e) so it flows through
handleRowClick → selectEntry → fetchEntryDetail exactly as a root entry
would. Shift-range selection gracefully degrades to single-select since
child entries are not in the root entries array.
Selected/multi-selected visual state is wired: .child-entry-row.is-selected
shows the same #eee2d2 background + accent outline as root rows; hover
restores full opacity. Frontend static assets rebuilt.
* feat(core): playlist per-item quality + incremental sync
ytdlp.rs:
- Add PlaylistItemProbe / PlaylistProbeResult (pub, serde::Serialize)
- Add private available_video_heights_from_value() helper for Value entries
- Add probe_playlist_qualities(): yt-dlp -J (full metadata, no flat flag)
returns per-video quality lists in one subprocess call
capture.rs:
- Add per_item_quality: HashMap<String,String> and sync: bool to CaptureConfig
(both Default; keyed by yt-dlp video ID, not URL)
- Add pub locator_to_playlist_url(): validates only the three playlist sources,
expands shorthands; keeps locator_to_ytdlp_url's no-playlists contract
- Playlist capture block: sync-aware container resolution
- sync + existing container → reuse it via complete_archive_run_item,
skip already-archived children (by canonical URL) before creating
run items so refresh_run_counters only counts new items
- sync + no container → create normally (first sync run)
- non-sync → always create fresh container (existing behaviour)
- Per-item quality: config.per_item_quality.get(id) falls back to child_quality
archive.rs:
- Add get_archived_playlist_child_urls(): returns HashSet of canonical URLs
of all children under any container matching the playlist canonical URL
- Add find_container_entry_id_by_canonical_url(): returns most-recent
container entry id (parent_entry_id IS NULL) for a given canonical URL
routes.rs: stub per_item_quality/sync on both CaptureConfig sites (server
agent will wire body fields in Phase 2)
* fix(core)+test: propagate sync query errors; cover new playlist/sync functions
archive.rs:
- get_archived_playlist_child_urls: collect() as rusqlite::Result<HashSet<_>>
instead of filter_map(ok) so row-level errors surface rather than silently
skipping and causing duplicate downloads
capture.rs:
- match get_archived_playlist_child_urls result and fail_run on error instead
of unwrap_or_default, preventing silent re-downloads on DB failure
Tests added to capture.rs:
- locator_to_playlist_url_accepts_playlist_shorthands (yt:playlist/, ytm:playlist/, full URL)
- locator_to_playlist_url_accepts_channel_shorthands (yt:@handle)
- locator_to_playlist_url_rejects_non_playlist_sources (single video, tweet, web page)
Tests added to archive.rs (all use in-memory DB via make_tag_test_db):
- find_container_entry_id_returns_none_when_absent
- find_container_entry_id_returns_root_entry
- find_container_entry_id_ignores_child_entries (child with parent_entry_id set)
- get_archived_playlist_child_urls_empty_when_no_playlist
- get_archived_playlist_child_urls_returns_children
- get_archived_playlist_child_urls_excludes_other_playlists
* feat(server,frontend): playlist quality selector + per-video overrides + sync UI
routes.rs:
- CaptureBody gains per_item_quality (HashMap<String,String>, serde(default))
and sync (bool, serde(default)); both validated before use
- per_item_quality values validated against same quality predicate as top-level
quality field ("best"|"audio"|"NNNp") so bad per-video values are rejected
at the API boundary rather than silently falling through to quality_format
- capture_handler threads body.per_item_quality + body.sync into CaptureConfig
(replaces hardcoded empty stubs); rearchive_handler keeps empty defaults
- New POST /api/archives/:id/captures/probe-playlist: calls
probe_playlist_qualities via spawn_blocking; 400 for non-playlist locator,
502 on yt-dlp failure, returns PlaylistProbeResult as JSON
api.js:
- probePlaylist(archiveId, locator): POST probe-playlist endpoint
- submitCapture: forwards per_item_quality (non-empty) and sync:true from
extraExtensions param added to submitBgJob
CaptureDialog.jsx:
- isPlaylistSource(): detects yt:/youtube: playlist/@/channel, ytm:playlist/,
YouTube/YTM HTTP(S) URLs with list= param or channel pathnames
- makeItem(): 6 new playlist state fields
- applyPlaylistQuality(): conflict logic — videos that can reach selected
quality get it set; videos that can't and have no prior selection are left
null (conflict); videos with a prior selection keep it when quality is raised
- hasConflict(): any playlistItems entry with quality===null
- updateLocator(): isPlaylistSource branch with 800ms debounce→probePlaylist;
existing isVideoSource path unchanged
- Archive button disabled when anyConflict or any probe in flight
- Per-video expand list with individual quality selects, conflict badges,
sync toggle (appears after probe completes)
styles.css: playlist expansion, conflict, sync toggle CSS
* fix(core): ignore per_item_quality for YTM playlist items
YouTube Music playlists force child_quality = Some("audio") because
yt-dlp can't download DRM-free audio-only tracks any other way. The
previous per_item_quality lookup could override this with e.g. "best",
defeating the invariant. Guard the lookup behind !is_audio so YTM items
are always downloaded as audio regardless of what the caller sends.
* fix(frontend): exclude /watch from isPlaylistSource
youtube.com/watch?v=...&list=... and music.youtube.com/watch are single
videos in the backend (Source::YouTubeVideo / YouTubeMusicTrack) regardless
of the list param. Previously isPlaylistSource returned true for these,
which would have triggered the playlist probe path while the video probe
was already running, and the render would attempt to show playlist UI on
an item whose playlistProbeState stays idle.
Guard: if pathname === '/watch', return false before the list-param check.
* fix(frontend): tighten isPlaylistSource to mirror backend routing exactly
Previous fix excluded /watch but still returned true for any youtube.com
URL with a ?list= param (e.g. /shorts/xxx?list=yyy). Backend determine_source
only routes to YouTubePlaylist on /playlist?list=... and to YouTubeChannel on
/@handle, /channel/, /c/, /user/ paths — everything else is a single item.
Rewrite the HTTP block to match:
- youtube.com: pathname==='/playlist' && list param → playlist
: /@, /channel/, /c/, /user/ → channel
: anything else (incl. /watch&list=, /shorts?list=) → false
- music.youtube.com: pathname==='/playlist' && list param → YTM playlist
: /watch → single track (falls through to false)
* fix(frontend): guard handleArchive against Enter-key bypass of disabled state
The Archive button is disabled when anyConflict || anyProbing, but
onKeyDown on the locator input calls onSubmit() → handleArchive()
directly, bypassing the button's disabled check entirely.
Add the same conditions as early returns inside handleArchive itself,
operating on toSubmit (the items that would actually be submitted) so
the guard is tight — items with no locator are already excluded by the
toSubmit filter.
* fix(frontend): drop m.youtube.com from isPlaylistSource
Backend determine_source playlist/channel regex only matches
(?:www\.)?youtube\.com — mobile URLs hitting m.youtube.com would be
probed as playlist in the UI but captured as Source::Url server-side.
Remove m.youtube.com from the detector to keep frontend and backend
in exact agreement. Add backend support when needed.
* fix(frontend): audio-only conflict handling in playlist quality selector
applyPlaylistQuality('audio'):
- Only sets quality='audio' on items where has_audio=true
- Items with has_audio=false: keep prior selection if set, else null
(conflict) — same rule as unsupported height, blocks archive until
user explicitly picks a quality for those items
Playlist-level 'Audio only' option:
- Changed hasAnyAudio → allHaveAudio (every item must have audio)
- When any item lacks audio, the option is hidden entirely so the
selector can never create immediate conflicts just by appearing
* fix(frontend): add yt:user/ to isPlaylistSource shorthand detection
Backend determine_source routes yt:user/... (and youtube:user/...) to
YouTubeChannel — already covered by the yt: shorthand block for
playlist/, @, channel/, c/ but missing user/. Old-style user channel
URLs would capture correctly server-side but never show the playlist
quality/sync UI.
* fix(frontend): block playlist submission unless probe is done
Previous guard only blocked while playlistProbeState==='probing'.
Two remaining bypass paths:
- idle: 800ms debounce not yet fired after URL typed
- error: probe failed — no per-video quality data available
Change anyProbing and handleArchive guard to:
isPlaylistSource(locator) && playlistProbeState !== 'done'
This means idle/probing/error all block submission for playlist items.
error is intentionally blocking — without quality data the per-video
requirement can't be satisfied; user must retry or remove the URL.
* fix(frontend): accurate error message when playlist probe fails
Previous text said 'using best quality' implying the capture would
proceed, but probe error now blocks submission. Replace with 'Probe
failed — edit URL to retry' in orange (capture-quality-hint--error)
so the disabled button and the message are consistent.
* fix(frontend): exact quality match in applyPlaylistQuality
Replace maxHeight >= newHeight (cap check) with item.qualities.includes(newQ)
(exact match). A video with [2160p, 1080p] does not support 1440p; the
previous logic would mark it as supporting any quality up to 2160p and
submit '1440p' which yt-dlp silently downloads as 1080p — misrepresenting
the selected quality.
With exact match, unsupported qualities correctly fall through to the
conflict path (keep prior selection or null), enforcing the same manual-
choice requirement as any other unsupported quality.
Per-row selects are unaffected: they already render only pi.qualities
(the video's actual available formats), no maxHeight logic involved.
* fix(frontend): move playlist expand chevron to left of input
User asked for the chevron to be on the left of the playlist input,
not tucked after the quality selector on the right.
- Remove chevron from qualityEl (it was between the quality select and
the remove button)
- Add it as the first child of capture-row-main, before the <input>,
when isPlaylistSource && playlistProbeState === 'done'
- Show a same-width placeholder span while probing/idle/error so the
input does not jump left when the chevron appears after probe completes
- Add capture-playlist-toggle--left modifier (flex-shrink:0, tighter
padding) and .capture-playlist-toggle-placeholder (fixed 22px width)
* doc(server): scope per_item_quality guarantee to the UI
The server validates per_item_quality value shapes but does not enforce
that every playlist item has an entry. Items without an override get
the global quality as a yt-dlp cap with graceful fallback.
The 'must choose quality for unsupported videos' invariant is a UI
constraint enforced by the frontend before submission. A direct API
caller bypassing the UI accepts yt-dlp's standard cap-and-fallback
behavior. Document this scope explicitly so the gap is intentional,
not accidental.
* fix(frontend): prevent 'reading some of undefined' crash from stale sessionStorage
Old captureItems entries saved before the playlist fields were added
have playlistItems=undefined (missing key). The hasConflict guard
checked !== null, which undefined passes, then called .some() on
undefined → TypeError.
Two-part fix:
1. sessionStorage restore: merge each saved item over makeItem() defaults
so any missing fields (playlistItems, playlistProbeState, etc.) are
filled with their correct initial values before the item is used
2. hasConflict: use Array.isArray() instead of !== null so undefined
is also safely rejected — defence-in-depth for any future field gap
* fix(core): playlist total size includes children; per_item_quality as include-set
archive.rs: total_artifact_bytes for root entries now adds a correlated
subquery summing children's blob bytes so playlist/channel containers
show the real download size instead of 0.
capture.rs: non-empty per_item_quality map now acts as an include-set —
items whose yt-dlp ID is absent are skipped entirely. This wires the
UI's per-video delete button to actual capture exclusion. Empty map
preserves the existing behaviour (download everything).
routes.rs + capture.rs doc: comments updated to reflect both semantics
(empty = all / non-empty = only listed IDs) accurately.
* fix(frontend): QA fixes — playlist UX, selection stroke, URL expand
CaptureDialog.jsx:
- Placeholder no longer appears on idle playlist rows (only during
probing); chevron shows only after probe completes — no left-padding
while the user is still typing
- Per-video delete button added to expanded playlist list; removes item
from playlistItems so its ID is absent from per_item_quality on submit
- anyEmptyPlaylist guard: Archive button disabled + handleArchive early-
return when all videos have been deleted (empty map would otherwise
silently download everything)
styles.css:
- Selection stroke switched from outline to box-shadow:inset everywhere
(entry-row-outer, child-entry-row, legacy flat-div selector) —
guaranteed inside element bounds, no layout interference
- URL cell overflow only on :hover; removed is-selected .url-cell rule
that was expanding child URLs when their parent was selected
- Playlist items redesign: separator lines instead of background fills,
amber left-border for conflicts, thin scrollbar, tighter padding
- Remove button: opacity 0.3 always-visible baseline; full opacity on
hover/:focus-visible; forced to 1 on coarse-pointer (touch) devices
* fix(frontend): no left gap on playlist rows until chevron exists
Remove the probing-state placeholder span entirely. The chevron renders
only when playlistProbeState === 'done'; all other states (idle, probing,
error) render null. The small layout shift when the chevron appears after
probe is acceptable; blank padding while there is no chevron is not.
* fix(frontend): clear box-shadow on entry-row-outer to prevent double stroke
The flat selector '#entries-body > div.is-selected' already applies
box-shadow to .entry-row-outer (it IS a direct child div). The outer
override rule only cleared 'outline', so the box-shadow leaked through,
wrapping the entire parent+children block with a second stroke.
Add box-shadow: none to both .is-selected and .is-multi-selected on
.entry-row-outer so the stroke sits only on .entry-row-main.
* docs: document per-video exclude in README YouTube playlists section
* fix(frontend): scope selection background to entry-row-main only
Moving background:#eee2d2 off .entry-row-outer onto > .entry-row-main
so that expanded child entries don't inherit the selection highlight.
Outer wrapper now explicitly unsets background (cancelling the flat-div
cascade rule) and clears outline+box-shadow. Both the selection colour
and the inset stroke live on .entry-row-main only.
* fix(frontend): alternating stripe backgrounds on child entry rows
Child rows were inheriting the parent entry's stripe color, making the
expanded list look like one flat block. Apply the same odd/even palette
as root entries (var(--paper-3) / #f2ede5) so each video row is visually
distinct within the expanded group.
* fix(core): delete_entry correctly nulls FK for child entries
archive_run_items.produced_entry_id has no ON DELETE action so it must
be manually nulled before the entry row is deleted. The old query used
WHERE root_entry_id = entry_id, which finds descendants of a root but
returns nothing when entry_id IS a child (children have no sub-children,
so no row has root_entry_id = child_id). The DELETE then failed under
foreign_keys=ON.
Fix: add OR produced_entry_id = ?1 so the entry's own run_item FK is
always cleared before deletion, regardless of whether it is a root or a
child. The subtree subquery is kept for the root-deletion case where all
child run_items also need nulling.
* fix(core+frontend): child entry selection and delete correctness
database.rs — delete_entry:
- subtree_ids now uses WHERE id = ?1 OR root_entry_id = ?1 so the entry
itself is always included; previously a child deletion passed an empty
vec to cascade_cached_bytes_after_subtree_delete (no grandchildren
exist), leaving cached_bytes stale on entries sharing that child's blobs
App.jsx — handleRowClick:
- Shift-range now queries DOM order (#entries-body [data-entry-uid])
instead of entries.findIndex(); child rows are in the DOM but not in
the root entries array, so findIndex always returned -1 for them
- Ctrl/meta branch computes next set before the state update so
selectEntry can fire synchronously for child rows; auto-snap only
restores root entries, so ctrl-clicking a lone child never loaded its
detail panel — now calls selectEntry(entry) when the child ends up as
the sole selection, selectEntry(null) on multi or deselect
- selectedUids added to handleRowClick's useCallback dep array
* fix(frontend): resolve detail entry for any remaining child after ctrl-deselect
Add entryCacheRef (uid→entry Map) populated on every row click. When
ctrl/meta-deselecting leaves exactly one other entry selected, look up
the remaining UID in the cache before falling back to the root entries
array. Without this, deselecting from a multi-selection where the
remaining entry is a child row left selectedEntry null (auto-snap only
searches root entries).
* fix(frontend): child row stripes, deleted-child visibility, selection completeness
EntryRow.jsx / ChildRow:
- Index-based light/dark classes (child-entry-row--light/dark) replace
nth-child rules; parity comes from children.map idx so no sibling —
including the loading div — can shift the stripe order
- Loading div moved outside .child-entries so it never affects child
row ordering at all
- Accept deletedUids prop; filter expanded children array before render
so deleted children disappear immediately without waiting for reload
EntriesView.jsx: thread deletedUids through to EntryRow
App.jsx:
- Add deletedUids state; handleEntryDeleted/handleBulkDeleted populate it
- isRoot/hasChildDelete computed from entries before setEntries (safe in
StrictMode — no side-effects inside updater functions)
- Child delete triggers loadEntries to refresh stale parent child_count
and total_artifact_bytes
- handleRowClick ctrl/meta cache-miss branch uses det.summary (not det)
from fetchEntryDetail; archiveId added to dep array
- handleRowClick dep array includes archiveId
* fix(frontend): add inset stroke to child-entry-row.is-multi-selected
* fix(frontend): suppress mouse-click focus ring on entry expand button
* fix(frontend): index-based stripes for root and child rows
Root rows: EntriesView passes rowIndex from entries.map to EntryRow,
which applies entry-row-outer--light/dark. Retires both nth-child stripe
blocks so skeleton rows can never shift the first real entry to dark.
Skeleton rows keep a :not(.entry-row-outer) nth-child fallback.
Child rows: colours changed from the warm root palette (paper-3/#f2ede5)
to cooler near-whites (#fafaf8/#f2f0ec) so children are visually distinct
from their parent row regardless of which stripe the parent sits on.
* feat(core+frontend): bare yt:ID resolves to YouTube video
determine_source: yt:ID / youtube:ID with no prefix and exactly 11
chars [A-Za-z0-9_-] → YouTubeVideo via is_youtube_video_id helper.
Reserved prefixes (playlist/, channel/, c/, user/, @) still fire first
so they are unaffected by the fallback.
expand_shorthand_to_url: bare yt:ID expands to watch?v=ID using the
same predicate, consistent with how ytm:ID → music.youtube.com/watch.
isVideoSource (frontend): same 11-char /^[A-Za-z0-9_-]{11}$/ regex so
yt:ID triggers the quality probe and capture-row guards identically to
yt:video/ID.
Tests: test_is_youtube_video_id covers valid IDs (alphanumeric, with _
and -), too-short, too-long, and invalid-char cases using genuinely
invalid fixtures. test_youtube_sources adds bare-ID cases and confirms
reserved prefixes (playlist/, @) are not affected.
* fix(core+frontend): exclude avatar blobs from tweet % cached display
tweet/tweet_thread entries have avatar artifacts that are always
deduplicated from the first capture of each author. Counting them in
the cached-bytes percentage makes it artificially high (or incorrect
when the real media content is new but avatars are cached).
database.rs:
- refresh_entry_cached_bytes: AND ea.artifact_role != 'avatar'
- cascade_cached_bytes_after_delete: same filter
- cascade_cached_bytes_after_subtree_delete: same filter
- Initial cached_bytes migration: same filter
- Re-migration (else branch): recomputes cached_bytes for existing
entries that have avatar artifacts, scoped to only those entries
archive.rs:
- EntrySummary gains cacheable_bytes: i64 — non-avatar total bytes,
computed inline in every SQL query as the denominator for % cached
- ENTRY_SELECT_COLS adds cacheable_bytes at index 13 (with children
subquery, same as total_artifact_bytes)
- list_root_entries adds same expression
- All 6 row-mapping closures include cacheable_bytes: row.get(13)?
EntryRow.jsx:
- SIZE display stays: formatBytes(entry.total_artifact_bytes)
- % cached badge uses cacheable_bytes as denominator:
cached_bytes / cacheable_bytes * 100
* feat(frontend): j/k keyboard navigation for entries; test avatar cached_bytes
App.jsx — j/k handler:
- Fires on keydown when not focused on INPUT/TEXTAREA/SELECT/contenteditable
- Ignores meta/ctrl/alt modifier combos
- Uses DOM order (#entries-body [data-entry-uid]) so expanded child rows
participate, matching the shift-range selection logic
- Resolves target entry via entryCacheRef → root entries array →
fetchEntryDetail(..).summary on cache miss
- Scrolls target into view (block: nearest); updates lastAnchorIndexRef
so subsequent shift-click ranges start from the keyboard-navigated row
archive.rs — cached_bytes_excludes_avatar_blobs test:
- Two tweet entries sharing an avatar blob (100B) and a media blob (900B)
- Asserts refresh_entry_cached_bytes sets cached_bytes = 900 (not 1000)
- Asserts list_root_entries summary: cached_bytes=900, cacheable_bytes=900,
total_artifact_bytes=1000 — SIZE includes avatar, % cached denominator
and numerator both exclude it
* fix(frontend): guard j/k uncached-child fetch with monotonic token
Rapid j/k over child rows that aren't in entryCacheRef triggers
fetchEntryDetail calls in parallel. Without a guard the last one to
settle wins, desyncing selectedUids (highlight) from selectedEntry
(detail panel/URL).
Fix:
- jkSeqRef (monotonic counter) incremented before each server fetch;
the .then() guard tok === jkSeqRef.current drops results from
superseded navigations
- handleRowClick increments jkSeqRef.current so any click also cancels
an in-flight j/k fetch
* fix(frontend): guard ctrl/meta cache-miss fetch; add / search shortcut
App.jsx ctrl/meta branch: cache-miss fetchEntryDetail now captures
tok = ++jkSeqRef.current before the fetch and gates selectEntry on
tok === jkSeqRef.current — same pattern as the j/k handler — so a
slow response after a later click/navigate can't overwrite selection.
/ key: reuses the Cmd+K/Ctrl+K handler path (focus+select search input,
or pendingSearchFocus + archive view switch). Guards: editable targets
(INPUT/TEXTAREA/SELECT/contentEditable) and modifier keys are checked
before preventDefault() so typing / in inputs is untouched.
* feat(core): attempt SpotifyTrack via yt-dlp instead of hard-failing
Previously all Spotify sources returned an error claiming yt-dlp cannot
download DRM-protected audio. yt-dlp does have a Spotify extractor
(experimental, content-dependent), so refusing upfront is worse than
trying and letting it report the real failure.
SpotifyTrack now goes through the same yt-dlp metadata + audio-quality
download path as YouTubeMusicTrack. SpotifyAlbum and SpotifyPlaylist
still return an explicit error — fetch_playlist_info has YouTube-specific
URL fallback logic that would produce bogus child URLs for Spotify flat
entries; those sources need dedicated container handling first.
* chore: remove docs/superpowers from repo and gitignore whitelist
* feat(core+frontend): Spotify album/playlist capture via yt-dlp container path
Previously SpotifyAlbum and SpotifyPlaylist hard-failed with a DRM error.
yt-dlp does support Spotify (experimentally), so they now go through the
same probe/container/child path as YouTube playlists.
ytdlp.rs:
- Extract normalize_item_url() helper used by both fetch_playlist_info
and probe_playlist_qualities; eliminates the duplicate URL-normalization
blocks and the YouTube-specific fallback_host variable
- Fallback logic is now platform-aware: YouTube/YTM bare IDs → watch URL,
Spotify → open.spotify.com/track/{id}, unknown → skip with warning
capture.rs:
- locator_to_playlist_url: accept SpotifyAlbum | SpotifyPlaylist so the
probe-playlist endpoint accepts Spotify album/playlist URLs
- Container branch: add SpotifyAlbum | SpotifyPlaylist to the matches!
- is_audio: true for SpotifyAlbum | SpotifyPlaylist (audio-only, like YTM)
- child_source: SpotifyTrack for Spotify containers (not YouTubeMusicTrack)
- generate_entry_title and record_media_entry both use the correct child
source so entity_kind, source_kind, and representation_kind are right
CaptureDialog.jsx:
- isPlaylistSource: recognise open.spotify.com/album/ and /playlist/,
and spotify:album:ID / spotify:playlist:ID shorthands, so Spotify
containers get the probe UI, per-track excludes, and sync toggle
* fix(core+server): partial playlist refresh + child visibility inheritance
database.rs:
- finish_archive_run: reverted 'partial' status (violates CHECK constraint);
restored binary completed/failed
- get_run_completed_count(): new helper for callers that need to distinguish
partial success without touching the DB status enum
capture.rs:
- CaptureResult gains completed_count: i64 (0 for single-item captures;
populated from get_run_completed_count for container/playlist captures)
- All CaptureResult construction sites updated
routes.rs:
- Playlist job_status mapping: mark job 'completed' when completed_count > 0
(even if status == 'failed'), so partially-successful playlist captures
trigger onCaptured and show the archived entries without a manual reload.
Truly zero-success runs (completed_count == 0, status == 'failed') stay
failed as expected.
- probe_playlist_handler error message updated to mention Spotify
archive.rs (list_child_entries):
- Add third OR arm: children are visible when their parent container is in a
collection visible to the caller. Fixes empty child list for non-admin
users with access to a playlist container but not its newly-created
children (which are all in the default 'private' collection).
* fix(server): use completed_count to distinguish partial from total failure
finish_archive_run is binary (completed/failed) — a playlist where some
tracks succeed and some fail returns status='failed'. Without this fix,
the job mapping treated any failed status as a job failure, causing
onCaptured to never fire and leaving successfully archived entries hidden
until a manual reload.
Now: job is 'failed' only when status='failed' && completed_count==0.
Partial runs (completed_count > 0) map to a completed job so the UI
refreshes and shows the entries that were successfully captured.
* fix(core+server): exclude container from child success count; rename field; tests
database.rs:
- get_run_completed_child_count(): renamed from get_run_completed_count and
scoped to child items only (parent_item_id IS NOT NULL). The container
run item is always completed first, so the old function inflated the count
by at least 1 for every playlist, making all-video-failed runs appear as
partial successes.
- New regression test: completed_child_count_excludes_container — completes
both a root and a child item, asserts DB completed_count==2 while
get_run_completed_child_count==1.
capture.rs:
- CaptureResult.completed_count renamed to completed_child_count to match
the function and make the semantics unambiguous at the call site.
routes.rs:
- job_status decision now uses result.completed_child_count == 0 so that a
playlist where every video fails (child_count==0) is correctly reported
as a failed job, not a completed one.
archive.rs:
- New regression test: list_child_entries_inherits_parent_visibility —
enrolls only the container in a USER-visible collection, asserts the
child is visible to a USER caller and invisible to a GUEST caller.
|
|||
|
d202e177e1
|
feat(frontend): add async capture UX with skeleton entries for in-progress captures (#31)
* feat(frontend): add SkeletonEntryRow with shimmer animation
Adds a new SkeletonEntryRow component that renders animated shimmer
placeholder cells matching the exact column layout of EntryRow
(col-added, col-title with icon circle, col-type pill, col-size,
col-url). CSS appended to styles.css using existing design tokens
(--paper-2, --line-soft, --paper-3) for the warm-toned shimmer.
* feat(frontend): async capture UX — reset dialog on submit, show skeleton rows
When the user presses Archive, CaptureDialog now immediately resets its
form to a fresh empty state and closes. Background captures continue
polling via intervals that survive the dialog close.
Skeleton rows appear at the top of EntriesView (filtered to the active
archive) for each in-flight job, giving visual feedback that something
is being processed. On completion the skeleton is removed and the entry
list refreshes; on failure the skeleton is removed and an error toast
fires — matching the existing toast behaviour.
Architecture:
- App owns pending-capture state (pendingCaptures) as the single source
of truth, persisted to sessionStorage['pendingCaptures']. On page
refresh, App seeds the list and passes it to CaptureDialog as
activeJobs so polling reconnects without a second sessionStorage read.
- CaptureDialog emits onJobStarted({id,jobUid,locator,archiveId}) only
after submitCapture() returns a job_uid — never before, never on
failure — ensuring no orphan skeletons can persist after a refresh.
- onJobSettled(id) is called by startPolling on any terminal state
(completed, failed, or network error), removing the skeleton.
- EntriesView filters pendingCaptures by archiveId so a capture in
archive A never shows a skeleton in archive B.
Removed from CaptureDialog: submitItem, resetRow, hasActiveJobs,
anyActive, CapStatusDot. CaptureRow is simplified to idle-only display
with no disabled state, no retry button, no status dot.
* fix(frontend): skeleton replacement ordering and col-check alignment
- Await the entry list refresh before removing the skeleton so the
real row arrives before the placeholder disappears. handleCaptured
now returns its Promise.all; startPolling awaits it before calling
onJobSettled on the success path.
- Add col-check as the first child of SkeletonEntryRow to match
EntryRow's DOM structure; without it, columns misalign on touch
devices where col-check becomes display:flex.
* fix(frontend): use Promise.allSettled in handleCaptured to prevent runs-fetch failure from poisoning successful captures
Promise.all rejects on the first failure. If the /runs refresh threw,
the rejection would propagate through startPolling's await and land in
the outer catch block, firing an error toast for a capture that had
already succeeded. Promise.allSettled settles unconditionally so a
transient runs fetch error is silently absorbed while the entry list
still refreshes.
|
|||
|
5db18122f7
|
feat(frontend): add Freedium mirror toggle to capture advanced options
- freediumEnabled state defaults to true (on by default) - via_freedium forwarded through submitCapture to the capture API - Toggle rendered last in the advanced panel, matching existing rows - Built frontend static assets included |
|||
|
278dc928df
|
feat: integrate Modal Closer as injected browser script (#27)
Port the modalcloser abx-plugin as a SingleFile --browser-script rather
than a spawned Puppeteer daemon. No external extension download required.
Core behavior (singlefile.rs):
- MODAL_CLOSER_DIALOG_OVERRIDES: main-world <script> bridge (best-effort;
blocked by strict script-src CSP) that overrides window.alert/confirm/
prompt/print, nulls window.onbeforeunload, traps its setter, and wraps
window.addEventListener to no-op 'beforeunload' registrations.
- MODAL_CLOSER_POLLING_SETUP: defines _archivr_mc_run() which runs two
passes on every tick (immediate first run, then setInterval every 500 ms
matching MODALCLOSER_POLL_INTERVAL default):
Pass 1 (best-effort, CSP-sensitive): main-world <script> bridge calls
Bootstrap/jQuery/jQuery UI/SweetAlert teardown APIs.
Pass 2 (always, CSP-immune): isolated-world DOM mutations — Escape-key
dispatch (Radix/Headless UI/Angular Material), backdrop clicks, full
CSS selector hiding for 40+ named consent/overlay vendors, body
scroll-lock reset. Direct DOM mutations need no inline script execution.
- resolve_modal_closer_config(): reads ARCHIVR_MODAL_CLOSER env var
(default true); no external resource required.
- modal_closer_enabled: Option<bool> in CaptureConfig so Default::default()
yields None (follow env var) rather than false.
Schema (database.rs):
- modal_closer_enabled column in auth DB instance_settings (default 1).
- Idempotent ALTER TABLE migration for existing databases.
- get/update_instance_settings updated (SELECT col 6, UPDATE param ?7).
HTTP (routes.rs):
- modal_closer_enabled in CaptureBody and UpdateInstanceSettingsBody.
- Per-capture body overrides global setting which overrides env var.
Frontend:
- SettingsView ExtensionsTab: third ext-card for Modal & Dialog Closer.
- CaptureDialog: per-capture toggle in Advanced Options, state initialized
from server default, included in submitCapture payload.
- api.js: submitCapture forwards modal_closer_enabled to capture body.
Behavior vs original modalcloser daemon:
- Functionally identical on non-strict-CSP pages (the large majority).
- Gap: <script> bridge is subject to page CSP; page.evaluate() is not.
On strict-CSP pages framework teardown and dialog overrides are no-ops;
CSS selector hiding and Escape dispatch (Pass 2) always work.
- Gap: alert/confirm/prompt overridden to instant no-op vs CDP timed
dialog.accept() with 1250 ms delay; irrelevant for archival in practice.
|
|||
|
03390362c5
|
capture: close dialog on Archive; rich per-URL and batch toast notifications (#22)
* http: realistic UA; fall back to WebPage on probe failure Replace bare 'archivr/0.1' user agent with a full Chrome 131 UA string (archivr/0.1 token retained at the end) in both probe_url_kind and download. More importantly, stop hard-failing when probe_url_kind returns an error. Sites behind Cloudflare's managed JS challenge (e.g. Medium) return 403 before any header tuning can help — a plain HTTP client cannot pass the challenge. Instead, log a warning and fall back to Source::WebPage so that the SingleFile/Chromium path gets a chance; a real browser can solve the challenge transparently. * capture: close dialog on Archive; rich per-URL and batch toast notifications UX changes: - Pressing Archive closes the capture dialog immediately; jobs continue polling in the background (component stays mounted). - Probe failures (e.g. Cloudflare JS challenge, HTTP 403) now fall back to Source::WebPage so SingleFile/Chromium can attempt the capture instead of hard-failing at the probe stage. Toast notifications: - Single URL: green 'Archived' on success, amber 'Archived with warnings' (with expandable detail) when uBlock/cookie-ext was skipped, red 'Capture failed' with expandable error text on failure. Per-item warning/error toasts include the locator so the user knows which URL was affected. - Multi-URL batch: per-item failure and warning toasts still fire with locators; per-item success toasts are suppressed. Once all jobs settle a single summary toast fires: 'N archived', 'N archived (M with warnings)', 'N archived (M with warnings), F failed', or 'N failed'. The summary Detail section lists the exact URLs that failed or warned, so the user retains that information after per-item toasts auto-dismiss. - Batch summary color: green = all clean; amber = any warnings or failures present; red = all failed. - handleIgnoreUblock now only removes per-item warning toasts (those with a locator) and persists the ignore flag; batch summary warnings (locator=null) are not swept. ToastStack improvements: - All three branches (success/warning/error) support a toast.headline field so batch summaries can set their own copy. - Warning branch: Details button conditional on toast.text; Ignore button conditional on toast.locator (uBlock-specific, not shown on batch summaries). - Locator display uses hostname/…/last-segment for URLs (preserves domain for context, tail for identity) and tail-truncation for shorthands; full locator in title attribute for hover. - Icon colors: green for success (✓), amber for warning (⚠), red for error (✕). |
|||
|
2e8820a0da
|
feat: uBlock Origin Lite + cookie consent extension + reader mode + ad placeholder cleanup (#21)
* feat: uBlock Origin Lite integration for ad-blocking during WebPage captures
- singlefile.rs: when ARCHIVR_UBLOCK=true and ARCHIVR_UBLOCK_EXT is set,
archivr owns Chrome's lifecycle (--headless=new, --remote-debugging-port,
--load-extension); single-file connects via --browser-server instead of
launching its own Chrome. Falls back to old behaviour with ublock_skipped=true
when the ext path is missing or invalid.
- capture.rs: thread ublock_skipped through CaptureResult
- database.rs: add notes_json TEXT column to capture_jobs (DDL + idempotent
ALTER TABLE migration); update_capture_job_status gains notes_json param
- archive.rs: expose notes_json in CaptureJobSummary
- routes.rs: store {"ublock_skipped":true} in notes_json on completed captures
- ToastStack.jsx: warning toast variant (toast--warning) with Details expander
and Ignore button
- CaptureDialog.jsx: fire warning toast when poll result has ublock_skipped
- App.jsx: sessionStorage-backed Ignore suppression for ublock warnings
- styles.css: .toast--warning (amber left border) + .toast-warning-detail
- flake.nix: ublockLite derivation fetches uBOLite_2026.705.2152.chromium.zip
(pinned SHA256) from uBlockOrigin/uBOL-home; sets ARCHIVR_UBLOCK_EXT in both
archivr and archivr-server wrappers
Env vars:
ARCHIVR_UBLOCK=true (default) — enable uBlock during WebPage captures
ARCHIVR_UBLOCK_EXT — path to unpacked uBOL extension dir (set by Nix)
* feat: Extensions settings tab + capture dialog redesign with Advanced options
Settings/Extensions tab (admin-only):
- New 'Extensions' tab between Cookies and Storage
- ExtensionsTab component: shows uBlock Origin Lite card with pill toggle
- Reads ublock_enabled from instance settings; patch via existing PATCH endpoint
- Shows ublock_ext_available status from server (whether ARCHIVR_UBLOCK_EXT is set)
Instance settings:
- Add ublock_enabled BOOLEAN (default true) to instance_settings auth DB table
- Idempotent ALTER TABLE migration in initialize_auth_schema()
- get/update_instance_settings include ublock_enabled
- GET /api/admin/instance-settings now also returns ublock_ext_available (computed
from ARCHIVR_UBLOCK_EXT env var at request time)
- PATCH /api/admin/instance-settings accepts ublock_enabled
Per-capture override:
- CaptureBody gains ublock_enabled: Option<bool>
- CaptureConfig gains ublock_enabled: Option<bool>
- singlefile::save() gains ublock_enabled_override: Option<bool> param
- Capture handler resolves: body override > global instance setting > env var
- submitCapture(aid, loc, qual, extensions) in api.js passes ublock_enabled
Capture dialog redesign:
- Archive button: full-width, 13px padding, min-width 220px, primary CTA
- Cancel: full-width but text-style, below Archive
- ‹Advanced options› chevron toggle (rotates on open)
- Expanded panel shows uBlock toggle for this capture session
- Loads global ublock_enabled default from instance settings on mount
Styles:
- .ext-toggle pill switch (44×24 and 36×20 small variant)
- .ext-card for Settings Extensions tab
- .capture-advanced + .capture-advanced-panel + .capture-chevron
- .capture-ext-row / .capture-ext-label / .capture-ext-name / .capture-ext-desc
- .form-hint utility class
* fix: remove ublock_enabled from INSERT OR IGNORE in DDL batch
The INSERT ran before the ALTER TABLE migration added the column,
causing 'table instance_settings has no column named ublock_enabled'
on existing databases. The INSERT OR IGNORE for the default row only
needs the original columns; the migration's DEFAULT 1 handles the
new column for existing and new rows alike.
* feat: Reader mode via Mozilla Readability.js
Adds an opt-in 'Reader mode' advanced option to the capture dialog.
When enabled, Readability.js is injected as a browser script during
SingleFile capture; it fires on single-file-on-before-capture-start,
replaces the page body with the distilled article content, injects a
clean typographic stylesheet, and adds a header with title/byline/site.
Falls back silently if Readability fails (e.g. non-article pages).
- vendor/readability/Readability.js Apache 2.0, Mozilla, v0.6.0
- singlefile.rs: embed READABILITY_JS + READER_MODE_WRAPPER_JS via
include_str!; write both to temp dir when reader_mode is true;
base_single_file_cmd now accepts &[&Path] for multiple --browser-script
- capture.rs: CaptureConfig.reader_mode: bool
- routes.rs: CaptureBody.reader_mode: Option<bool> (defaults false)
- api.js: submitCapture passes reader_mode in payload
- CaptureDialog.jsx: Reader mode toggle in Advanced options (off by default)
* fix: diagnose single-file no-output-file error + prevent stdout dumping
- Add --dump-content=false to every single-file invocation to prevent
the Docker-detection heuristic from routing HTML to stdout instead of
the output file (the heuristic can trigger in some macOS environments)
- Improve the no-output-file error message to include: temp dir contents,
stderr, and first 200 chars of stdout — this gives enough context to
diagnose any remaining cause without re-running
* fix: switch uBlock loading from --browser-server to --browser-args
The --browser-server (CDP) path caused 'Unexpected server response: 404'
on macOS Chrome because simple-cdp's WebSocket upgrade to the debugger
endpoint failed after Chrome started — likely a version-specific CDP
endpoint shape mismatch.
New approach: single-file always manages Chrome. When ARCHIVR_UBLOCK_EXT
is set, --headless=new, --load-extension, and --disable-extensions-except
are injected via --browser-args. single-file's browser.js prefix-strips
its own conflicting flags before appending ours, so --headless=new
overrides the default --headless (enabling extension support in headless).
Removes allocate_free_port, wait_for_chrome_ready, run_single_file_with_server
(all dead code now). Docblock updated to reflect actual behaviour and notes
the --single-process caveat: uBOL's declarativeNetRequest static rulesets
are expected to work (network-stack level, not service-worker), but this
has not been mechanically verified under --single-process.
Smoke tested on macOS (this machine): capture with --load-extension + all
three browser-scripts (strip, Readability, reader-mode wrapper) produces
output file correctly. Ad-blocking verification deferred to manual test
with a tracker-heavy URL.
* fix: use correct single-file hook event (single-file-on-before-capture-request)
Prior scripts listened on 'single-file-on-before-capture-start' which
does not exist in single-file-core 1.1.49. The real hook is:
single-file-on-before-capture-request (dispatched by initUserScriptHandler
after receiving single-file-user-script-init; userScriptEnabled defaults
to true in args.js so it always fires when --browser-script is passed)
Changes:
- strip-scripts: -start -> -request (no preventDefault needed; synchronous)
- READER_MODE_SCRIPT: -start -> -request; add 'installed' meta marker at
script-evaluation time so artifact inspection can distinguish 'script
not injected' / 'hook never fired' / 'Readability parse failed'
* fix: correct singlefile.rs docstring (scripts.js concatenates, not isolates)
* fix: dispatch single-file-user-script-init so request hook fires
single-file's initUserScriptHandler (in single-file-bootstrap.js) listens
for 'single-file-user-script-init' and only then installs
_singleFile_waitForUserScript. Without that dispatch our scripts'
'single-file-on-before-capture-request' listeners were never reached,
so neither strip-scripts nor reader-mode Readability applied.
Dispatch the init event at the top of strip-scripts (always present) and
redundantly in READER_MODE_SCRIPT. Verified end-to-end: artifact for
run_b3181d6d276e4e56a1a6c356ef9bbe8f has
meta content="applied", max-width:680px CSS, 0 script tags.
* feat: cookie consent extension support (ARCHIVR_COOKIE_EXT)
Mirrors the uBlock Origin Lite integration exactly:
Backend:
- singlefile.rs: resolve_cookie_ext_config() reads ARCHIVR_COOKIE_CONSENT
(default true) + ARCHIVR_COOKIE_EXT path; extension paths comma-joined
into --load-extension / --disable-extensions-except so uBlock and cookie
ext can coexist; SaveResult.cookie_ext_skipped tracks miss
- database.rs: cookie_ext_enabled column on instance_settings (DEFAULT 1);
idempotent ALTER TABLE migration; get/update wired through
- capture.rs: CaptureConfig.cookie_ext_enabled: Option<bool>; threaded to
singlefile::save(); cookie_ext_skipped surfaced in CaptureResult
- routes.rs: CaptureBody + UpdateInstanceSettingsBody get cookie_ext_enabled;
capture handler resolves effective value (body overrides global); notes_json
only includes skipped fields that are true; GET instance-settings includes
cookie_ext_available from env path check
Frontend:
- api.js: submitCapture forwards cookie_ext_enabled
- SettingsView.jsx: 'I Still Don't Care About Cookies' card in Extensions
tab; always-active toggle (user can disable even when ext not installed);
amber 'Not configured' hint + ARCHIVR_COOKIE_EXT guidance when unavailable
- CaptureDialog.jsx: 'Block cookie banners' toggle in Advanced options;
always shown with amber hint when ext not configured; defaults from
global setting
Operator setup: download + unzip the extension from GitHub releases, set
ARCHIVR_COOKIE_EXT=/path/to/unpacked/ext. No Node daemon needed.
* fix: surface cookie_ext_skipped warning toast in CaptureDialog
* feat: package istilldontcareaboutcookies in flake, wire ARCHIVR_COOKIE_EXT
Add isdcac derivation mirroring ublockLite:
- Fetches ISDCAC-chrome-source.zip v1.1.9 from GitHub releases
- Validates manifest.json at extension root before install (guard against
nested-folder zip regressions in future releases)
- Sets ARCHIVR_COOKIE_EXT in both archivr and archivr_server wrappers
Verified: nix build .#archivr-server and .#archivr both succeed;
wrapper scripts export correct store paths; manifest.json present at root.
* fix: gate consent-overlay cleanup on cookie_ext; reset overflow; narrow selectors
- Strip overflow:hidden from body/html only when cookie_ext is active for
the capture — prevents mutating legitimate pages when the feature is off
- Remove .fc-dialog (Google Funding Choices), .qc-cmp2-*, .sp-message-container,
#sp-cc, #usercentrics-root as fallback for CMPs the extension misses
- Removed overbroad [class^="uc-"] and [id^="usercentrics"] selectors
that could match real page content
* fix: remove ad placeholders when uBlock active; kept height causes blank gap
uBlock Origin Lite blocks ad network requests but first-party placeholder
elements (ins.adsbygoogle, #aswift_* iframe hosts) retain their computed
height (e.g. 280px for a top banner), leaving a large blank space at the
top of captured pages.
Gate cleanup on ublock_ext.is_some(): remove ins.adsbygoogle, aswift_*
iframes, and google_ads_* iframes before SingleFile serialises. Also
collapse the parent container if it becomes empty after removal.
* fix: walk up to .top-ad/.google-auto-placed ancestor before removing ad slot
Removing only the inner ins.adsbygoogle left the outer .container.top-ad
wrapper (with pb-4 padding) in the layout, preserving the blank gap.
Now walk up via closest() to the nearest ad-slot container class before
removal so the whole slot including padding collapses.
|
|||
|
21b11c211f
|
feat: YouTube Music audio capture (ytm: shorthand, Spotify detection, stalled job recovery) (#19)
* feat: add YouTube Music and Spotify source detection - Add Source variants: YouTubeMusicTrack, YouTubeMusicPlaylist, SpotifyTrack, SpotifyAlbum, SpotifyPlaylist - ytm:ID shorthand → music.youtube.com/watch?v=ID (audio-only, forced in core regardless of caller quality hint) - ytm:playlist/ID and music.youtube.com/playlist URLs detected but fail with 'not yet implemented' via fail_run - Spotify URLs/shorthands detected and fail fast with clear DRM error via fail_run (after run item created, so status is visible in /runs) - source_metadata: youtube_music/music/audio and spotify/music/audio (entity_kind='music' for UI pill, representation_kind='audio' stored) - locator_to_ytdlp_url includes YouTubeMusicTrack for probe endpoint - generate_entry_title: 'Title — Artist' for YTM tracks - Frontend: isVideoSource handles ytm: and music.youtube.com/watch; Spotify returns false (no probe, clear server error on submit) - Placeholder updated to include ytm:ID - SOURCE_ICONS: youtube_music (red disc) and spotify (green waves) - 14 new tests covering all new sources (163 total, all pass) * fix: prevent yt-dlp playlist expansion and stalled run recovery - Add --no-playlist to ytdlp::download and fetch_metadata: URLs with a list= parameter (e.g. music.youtube.com/watch?v=ID&list=RDAMVM…) no longer cause yt-dlp to expand the full playlist and hang; both the metadata probe and the download are now single-item only - Fix fail_stalled_capture_jobs to also recover archive_runs and archive_run_items: capture_jobs.run_uid is NULL at crash time so a join is unreliable; instead fail all archive_runs/items still in_progress directly, then recount failed_count via subquery. Startup recovery now makes the Runs UI reflect the correct failed state after a hard shutdown - Expand fail_stalled_jobs_on_restart test to assert archive_run and archive_run_item rows are also marked failed, not just capture_jobs * fix: use play triangle for youtube_music icon |
|||
|
b8e496457f
|
feat: add video quality selection for yt-dlp captures (#17)
- ytdlp::download() accepts quality: Option<&str>; quality_format() maps best/1080p/720p/480p/360p to yt-dlp -f format strings - perform_capture() threads quality through to the downloader - CaptureBody gains optional quality field; capture_handler validates it against the allowlist (400 on unknown values) before spawning - CLI passes None (preserves existing best-quality behaviour) - Frontend: isVideoSource() mirrors determine_source() exactly — shows quality picker only for yt-dlp-backed sources, excludes playlist/channel shorthands and tweet/thread paths - submitCapture(archiveId, locator, quality) sends quality in POST body - CSS: .capture-quality styles the inline select to fit the capture row - Tests: quality_format unit tests in ytdlp.rs; two new route tests (valid quality accepted, invalid quality rejected with 400) - Docs: video quality section added under Supported Platforms |
|||
|
fb1115a409
|
feat: non-blocking batch capture dialog, toasts on failure, fix /runs errors
CaptureDialog:
- Replace single textarea with multi-row inputs; + button adds rows
- Submit fires all pending rows in parallel, dialog stays open/usable
- Polling intervals live on a persistent ref (not cleared on close) so
toasts fire even after the dialog is dismissed
- archiveId stored per item at submit time; page-refresh reconnect uses
it.archiveId instead of the possibly-null prop
- Completed rows flash green then self-remove; failed rows show inline
error + retry button
- Cancel becomes Close while jobs are in flight
ToastStack (new component):
- Fixed bottom-right overlay with spring-in animation
- Error toast: truncated locator, View error / Hide toggle expanding
full error_text in a monospace pre block
- Auto-dismisses after 7 s; timer pauses while detail is expanded
RunsView:
- Failed rows are clickable and expand a full-width detail row showing
error_summary in a scrollable monospace block
capture.rs (archivr-core):
- Staging dir is now "{millis}-{uuid}" — parallel captures in the same
millisecond can no longer collide on temp paths
- create_archive_run moved before URL Content-Type probe so every
attempt appears in /runs regardless of outcome
- Probe failures now call create_archive_run_item with source_metadata
fallback then fail_run, recording error_text on the item and
error_summary on the run with correct failed_count
styles.css:
- Capture dialog: header row, multi-row layout, status dots, spinner,
add-row dashed button, per-row error text
- Toast stack: fixed overlay, error card with coloured left border,
monospace detail expansion
- Run error rows: clickable hover tint, expand hint chevron, detail pre
|
|||
|
d6b52ba06c
|
fix: capture popup non-persistance
- Save and restore dialog open/closed state in sessionStorage (App.jsx) - Persist form data: locator, error, busy, jobStatus, jobUid (CaptureDialog.jsx) - Auto-resume polling if capture job was in progress before page refresh - Only clear form on fresh user click, not when restoring from refresh - Clean up sessionStorage when capture completes successfully Fixes: Capture pop-up disappears on page refresh with unsaved data |
|||
|
9ac28459ed
|
feat(capture): async polling in CaptureDialog + pollCaptureJob api helper | |||
|
4458f17b13
|
feat(ui): rewrite frontend in React with Vite
- Scaffold Vite+React project in frontend/ (bun, react 18, @vitejs/plugin-react) - vite.config.js outputs to crates/archivr-server/static/ directly - src/utils.js: formatBytes, valueText, formatTimestamp, SOURCE_ICONS, sourceIconSvg - src/api.js: typed fetch wrappers for all API endpoints - App.jsx: full state management, archive switching, debounced search, tag filter, view routing, capture dialog orchestration - components/Topbar.jsx: archive switcher, nav, capture button - components/CaptureDialog.jsx: native <dialog> ref with showModal/close, Escape key support, locator validation - components/EntriesView.jsx + EntryRow.jsx: flex table with source icons, type pills, keyboard selection - components/ContextRail.jsx: parallel detail+tags fetch, stale-race guard via selectSeqRef, inline tag assign/remove - components/RunsView.jsx: runs kept as <table> (matches original) - components/AdminView.jsx: mounted archives list - components/TagsView.jsx: recursive tag tree with active state - Remove old app.js; styles.css moved to frontend/src/ (Vite bundles it) |