* feat(core): YouTube playlist/channel/YTM-playlist capture with parent–child entries
- ytdlp: add fetch_playlist_info() using yt-dlp -J --flat-playlist for
reliable container title + shallow entry list; normalize item URLs via
webpage_url → absolute url → id fallback (domain inferred from container
URL so YTM stays on music.youtube.com)
- capture: add record_container_entry() (no blob, no primary_media artifact);
extend record_media_entry() with parent_entry_id/root_entry_id params (all
existing single-item call sites pass None, None)
- capture: implement YouTubePlaylist / YouTubeChannel / YouTubeMusicPlaylist
capture path replacing the two not-implemented stubs: fetch playlist info →
create container entry (reusing existing run + item) → per-child run items
(parent_item_id = container item) → download each video/track as a child
entry; per-child failures are non-fatal; perform_capture returns result.status
reflecting actual run outcome so capture_handler marks the job correctly
- archive: add child_count i64 to EntrySummary (col 12 in all listing queries);
add get_entry_summary() private helper; fix get_entry_detail() to use
get_entry_summary() so child entries are resolvable via the detail endpoint;
add list_child_entries(conn, uid, caller_bits) with the same
admin/collection visibility predicate as list_root_entries
archive_runs.requested_count stays 1 (one user locator); discovered/
completed/failed_count reflect container item + N video items via
refresh_run_counters.
* feat(server,frontend): expose children endpoint + expand UI for container entries
server:
- add GET /api/archives/:id/entries/:uid/children → list_entry_children,
calling list_child_entries with caller_bits so visibility model is enforced
- fix capture_handler: use result.status ("completed"/"failed") to set job
status rather than always "completed", mirroring rearchive_handler; this
surfaces partial playlist failures to the polling client
frontend:
- api.js: add fetchEntryChildren(archiveId, entryUid)
- EntryRow: one outer div.entry-row-outer (display:block) keeps nth-child
striping correct; inner div.entry-row-main is the flex row with all column
cells and event handling; .child-entries sits below inside the outer wrapper
- expand chevron appears when entry.child_count > 0; clicking fetches children
lazily and renders ChildRow components reusing .col-* flex widths
- child-count badge shown next to title on container entries
- styles.css: scoped CSS with #entries-body > .entry-row-outer selectors
(higher specificity than > div) to override flex on outer wrapper; inner row
and column rules replicated at correct depth; nth-child, is-selected,
is-multi-selected, url-cell hover all handled
* fix(frontend): make child entry rows interactive
ChildRow now receives onRowClick and selectedUids from EntryRow (which
receives selectedUids from EntriesView alongside the existing booleans).
Clicking a child row invokes onRowClick(child, e) so it flows through
handleRowClick → selectEntry → fetchEntryDetail exactly as a root entry
would. Shift-range selection gracefully degrades to single-select since
child entries are not in the root entries array.
Selected/multi-selected visual state is wired: .child-entry-row.is-selected
shows the same #eee2d2 background + accent outline as root rows; hover
restores full opacity. Frontend static assets rebuilt.
* feat(core): playlist per-item quality + incremental sync
ytdlp.rs:
- Add PlaylistItemProbe / PlaylistProbeResult (pub, serde::Serialize)
- Add private available_video_heights_from_value() helper for Value entries
- Add probe_playlist_qualities(): yt-dlp -J (full metadata, no flat flag)
returns per-video quality lists in one subprocess call
capture.rs:
- Add per_item_quality: HashMap<String,String> and sync: bool to CaptureConfig
(both Default; keyed by yt-dlp video ID, not URL)
- Add pub locator_to_playlist_url(): validates only the three playlist sources,
expands shorthands; keeps locator_to_ytdlp_url's no-playlists contract
- Playlist capture block: sync-aware container resolution
- sync + existing container → reuse it via complete_archive_run_item,
skip already-archived children (by canonical URL) before creating
run items so refresh_run_counters only counts new items
- sync + no container → create normally (first sync run)
- non-sync → always create fresh container (existing behaviour)
- Per-item quality: config.per_item_quality.get(id) falls back to child_quality
archive.rs:
- Add get_archived_playlist_child_urls(): returns HashSet of canonical URLs
of all children under any container matching the playlist canonical URL
- Add find_container_entry_id_by_canonical_url(): returns most-recent
container entry id (parent_entry_id IS NULL) for a given canonical URL
routes.rs: stub per_item_quality/sync on both CaptureConfig sites (server
agent will wire body fields in Phase 2)
* fix(core)+test: propagate sync query errors; cover new playlist/sync functions
archive.rs:
- get_archived_playlist_child_urls: collect() as rusqlite::Result<HashSet<_>>
instead of filter_map(ok) so row-level errors surface rather than silently
skipping and causing duplicate downloads
capture.rs:
- match get_archived_playlist_child_urls result and fail_run on error instead
of unwrap_or_default, preventing silent re-downloads on DB failure
Tests added to capture.rs:
- locator_to_playlist_url_accepts_playlist_shorthands (yt:playlist/, ytm:playlist/, full URL)
- locator_to_playlist_url_accepts_channel_shorthands (yt:@handle)
- locator_to_playlist_url_rejects_non_playlist_sources (single video, tweet, web page)
Tests added to archive.rs (all use in-memory DB via make_tag_test_db):
- find_container_entry_id_returns_none_when_absent
- find_container_entry_id_returns_root_entry
- find_container_entry_id_ignores_child_entries (child with parent_entry_id set)
- get_archived_playlist_child_urls_empty_when_no_playlist
- get_archived_playlist_child_urls_returns_children
- get_archived_playlist_child_urls_excludes_other_playlists
* feat(server,frontend): playlist quality selector + per-video overrides + sync UI
routes.rs:
- CaptureBody gains per_item_quality (HashMap<String,String>, serde(default))
and sync (bool, serde(default)); both validated before use
- per_item_quality values validated against same quality predicate as top-level
quality field ("best"|"audio"|"NNNp") so bad per-video values are rejected
at the API boundary rather than silently falling through to quality_format
- capture_handler threads body.per_item_quality + body.sync into CaptureConfig
(replaces hardcoded empty stubs); rearchive_handler keeps empty defaults
- New POST /api/archives/:id/captures/probe-playlist: calls
probe_playlist_qualities via spawn_blocking; 400 for non-playlist locator,
502 on yt-dlp failure, returns PlaylistProbeResult as JSON
api.js:
- probePlaylist(archiveId, locator): POST probe-playlist endpoint
- submitCapture: forwards per_item_quality (non-empty) and sync:true from
extraExtensions param added to submitBgJob
CaptureDialog.jsx:
- isPlaylistSource(): detects yt:/youtube: playlist/@/channel, ytm:playlist/,
YouTube/YTM HTTP(S) URLs with list= param or channel pathnames
- makeItem(): 6 new playlist state fields
- applyPlaylistQuality(): conflict logic — videos that can reach selected
quality get it set; videos that can't and have no prior selection are left
null (conflict); videos with a prior selection keep it when quality is raised
- hasConflict(): any playlistItems entry with quality===null
- updateLocator(): isPlaylistSource branch with 800ms debounce→probePlaylist;
existing isVideoSource path unchanged
- Archive button disabled when anyConflict or any probe in flight
- Per-video expand list with individual quality selects, conflict badges,
sync toggle (appears after probe completes)
styles.css: playlist expansion, conflict, sync toggle CSS
* fix(core): ignore per_item_quality for YTM playlist items
YouTube Music playlists force child_quality = Some("audio") because
yt-dlp can't download DRM-free audio-only tracks any other way. The
previous per_item_quality lookup could override this with e.g. "best",
defeating the invariant. Guard the lookup behind !is_audio so YTM items
are always downloaded as audio regardless of what the caller sends.
* fix(frontend): exclude /watch from isPlaylistSource
youtube.com/watch?v=...&list=... and music.youtube.com/watch are single
videos in the backend (Source::YouTubeVideo / YouTubeMusicTrack) regardless
of the list param. Previously isPlaylistSource returned true for these,
which would have triggered the playlist probe path while the video probe
was already running, and the render would attempt to show playlist UI on
an item whose playlistProbeState stays idle.
Guard: if pathname === '/watch', return false before the list-param check.
* fix(frontend): tighten isPlaylistSource to mirror backend routing exactly
Previous fix excluded /watch but still returned true for any youtube.com
URL with a ?list= param (e.g. /shorts/xxx?list=yyy). Backend determine_source
only routes to YouTubePlaylist on /playlist?list=... and to YouTubeChannel on
/@handle, /channel/, /c/, /user/ paths — everything else is a single item.
Rewrite the HTTP block to match:
- youtube.com: pathname==='/playlist' && list param → playlist
: /@, /channel/, /c/, /user/ → channel
: anything else (incl. /watch&list=, /shorts?list=) → false
- music.youtube.com: pathname==='/playlist' && list param → YTM playlist
: /watch → single track (falls through to false)
* fix(frontend): guard handleArchive against Enter-key bypass of disabled state
The Archive button is disabled when anyConflict || anyProbing, but
onKeyDown on the locator input calls onSubmit() → handleArchive()
directly, bypassing the button's disabled check entirely.
Add the same conditions as early returns inside handleArchive itself,
operating on toSubmit (the items that would actually be submitted) so
the guard is tight — items with no locator are already excluded by the
toSubmit filter.
* fix(frontend): drop m.youtube.com from isPlaylistSource
Backend determine_source playlist/channel regex only matches
(?:www\.)?youtube\.com — mobile URLs hitting m.youtube.com would be
probed as playlist in the UI but captured as Source::Url server-side.
Remove m.youtube.com from the detector to keep frontend and backend
in exact agreement. Add backend support when needed.
* fix(frontend): audio-only conflict handling in playlist quality selector
applyPlaylistQuality('audio'):
- Only sets quality='audio' on items where has_audio=true
- Items with has_audio=false: keep prior selection if set, else null
(conflict) — same rule as unsupported height, blocks archive until
user explicitly picks a quality for those items
Playlist-level 'Audio only' option:
- Changed hasAnyAudio → allHaveAudio (every item must have audio)
- When any item lacks audio, the option is hidden entirely so the
selector can never create immediate conflicts just by appearing
* fix(frontend): add yt:user/ to isPlaylistSource shorthand detection
Backend determine_source routes yt:user/... (and youtube:user/...) to
YouTubeChannel — already covered by the yt: shorthand block for
playlist/, @, channel/, c/ but missing user/. Old-style user channel
URLs would capture correctly server-side but never show the playlist
quality/sync UI.
* fix(frontend): block playlist submission unless probe is done
Previous guard only blocked while playlistProbeState==='probing'.
Two remaining bypass paths:
- idle: 800ms debounce not yet fired after URL typed
- error: probe failed — no per-video quality data available
Change anyProbing and handleArchive guard to:
isPlaylistSource(locator) && playlistProbeState !== 'done'
This means idle/probing/error all block submission for playlist items.
error is intentionally blocking — without quality data the per-video
requirement can't be satisfied; user must retry or remove the URL.
* fix(frontend): accurate error message when playlist probe fails
Previous text said 'using best quality' implying the capture would
proceed, but probe error now blocks submission. Replace with 'Probe
failed — edit URL to retry' in orange (capture-quality-hint--error)
so the disabled button and the message are consistent.
* fix(frontend): exact quality match in applyPlaylistQuality
Replace maxHeight >= newHeight (cap check) with item.qualities.includes(newQ)
(exact match). A video with [2160p, 1080p] does not support 1440p; the
previous logic would mark it as supporting any quality up to 2160p and
submit '1440p' which yt-dlp silently downloads as 1080p — misrepresenting
the selected quality.
With exact match, unsupported qualities correctly fall through to the
conflict path (keep prior selection or null), enforcing the same manual-
choice requirement as any other unsupported quality.
Per-row selects are unaffected: they already render only pi.qualities
(the video's actual available formats), no maxHeight logic involved.
* fix(frontend): move playlist expand chevron to left of input
User asked for the chevron to be on the left of the playlist input,
not tucked after the quality selector on the right.
- Remove chevron from qualityEl (it was between the quality select and
the remove button)
- Add it as the first child of capture-row-main, before the <input>,
when isPlaylistSource && playlistProbeState === 'done'
- Show a same-width placeholder span while probing/idle/error so the
input does not jump left when the chevron appears after probe completes
- Add capture-playlist-toggle--left modifier (flex-shrink:0, tighter
padding) and .capture-playlist-toggle-placeholder (fixed 22px width)
* doc(server): scope per_item_quality guarantee to the UI
The server validates per_item_quality value shapes but does not enforce
that every playlist item has an entry. Items without an override get
the global quality as a yt-dlp cap with graceful fallback.
The 'must choose quality for unsupported videos' invariant is a UI
constraint enforced by the frontend before submission. A direct API
caller bypassing the UI accepts yt-dlp's standard cap-and-fallback
behavior. Document this scope explicitly so the gap is intentional,
not accidental.
* fix(frontend): prevent 'reading some of undefined' crash from stale sessionStorage
Old captureItems entries saved before the playlist fields were added
have playlistItems=undefined (missing key). The hasConflict guard
checked !== null, which undefined passes, then called .some() on
undefined → TypeError.
Two-part fix:
1. sessionStorage restore: merge each saved item over makeItem() defaults
so any missing fields (playlistItems, playlistProbeState, etc.) are
filled with their correct initial values before the item is used
2. hasConflict: use Array.isArray() instead of !== null so undefined
is also safely rejected — defence-in-depth for any future field gap
* fix(core): playlist total size includes children; per_item_quality as include-set
archive.rs: total_artifact_bytes for root entries now adds a correlated
subquery summing children's blob bytes so playlist/channel containers
show the real download size instead of 0.
capture.rs: non-empty per_item_quality map now acts as an include-set —
items whose yt-dlp ID is absent are skipped entirely. This wires the
UI's per-video delete button to actual capture exclusion. Empty map
preserves the existing behaviour (download everything).
routes.rs + capture.rs doc: comments updated to reflect both semantics
(empty = all / non-empty = only listed IDs) accurately.
* fix(frontend): QA fixes — playlist UX, selection stroke, URL expand
CaptureDialog.jsx:
- Placeholder no longer appears on idle playlist rows (only during
probing); chevron shows only after probe completes — no left-padding
while the user is still typing
- Per-video delete button added to expanded playlist list; removes item
from playlistItems so its ID is absent from per_item_quality on submit
- anyEmptyPlaylist guard: Archive button disabled + handleArchive early-
return when all videos have been deleted (empty map would otherwise
silently download everything)
styles.css:
- Selection stroke switched from outline to box-shadow:inset everywhere
(entry-row-outer, child-entry-row, legacy flat-div selector) —
guaranteed inside element bounds, no layout interference
- URL cell overflow only on :hover; removed is-selected .url-cell rule
that was expanding child URLs when their parent was selected
- Playlist items redesign: separator lines instead of background fills,
amber left-border for conflicts, thin scrollbar, tighter padding
- Remove button: opacity 0.3 always-visible baseline; full opacity on
hover/:focus-visible; forced to 1 on coarse-pointer (touch) devices
* fix(frontend): no left gap on playlist rows until chevron exists
Remove the probing-state placeholder span entirely. The chevron renders
only when playlistProbeState === 'done'; all other states (idle, probing,
error) render null. The small layout shift when the chevron appears after
probe is acceptable; blank padding while there is no chevron is not.
* fix(frontend): clear box-shadow on entry-row-outer to prevent double stroke
The flat selector '#entries-body > div.is-selected' already applies
box-shadow to .entry-row-outer (it IS a direct child div). The outer
override rule only cleared 'outline', so the box-shadow leaked through,
wrapping the entire parent+children block with a second stroke.
Add box-shadow: none to both .is-selected and .is-multi-selected on
.entry-row-outer so the stroke sits only on .entry-row-main.
* docs: document per-video exclude in README YouTube playlists section
* fix(frontend): scope selection background to entry-row-main only
Moving background:#eee2d2 off .entry-row-outer onto > .entry-row-main
so that expanded child entries don't inherit the selection highlight.
Outer wrapper now explicitly unsets background (cancelling the flat-div
cascade rule) and clears outline+box-shadow. Both the selection colour
and the inset stroke live on .entry-row-main only.
* fix(frontend): alternating stripe backgrounds on child entry rows
Child rows were inheriting the parent entry's stripe color, making the
expanded list look like one flat block. Apply the same odd/even palette
as root entries (var(--paper-3) / #f2ede5) so each video row is visually
distinct within the expanded group.
* fix(core): delete_entry correctly nulls FK for child entries
archive_run_items.produced_entry_id has no ON DELETE action so it must
be manually nulled before the entry row is deleted. The old query used
WHERE root_entry_id = entry_id, which finds descendants of a root but
returns nothing when entry_id IS a child (children have no sub-children,
so no row has root_entry_id = child_id). The DELETE then failed under
foreign_keys=ON.
Fix: add OR produced_entry_id = ?1 so the entry's own run_item FK is
always cleared before deletion, regardless of whether it is a root or a
child. The subtree subquery is kept for the root-deletion case where all
child run_items also need nulling.
* fix(core+frontend): child entry selection and delete correctness
database.rs — delete_entry:
- subtree_ids now uses WHERE id = ?1 OR root_entry_id = ?1 so the entry
itself is always included; previously a child deletion passed an empty
vec to cascade_cached_bytes_after_subtree_delete (no grandchildren
exist), leaving cached_bytes stale on entries sharing that child's blobs
App.jsx — handleRowClick:
- Shift-range now queries DOM order (#entries-body [data-entry-uid])
instead of entries.findIndex(); child rows are in the DOM but not in
the root entries array, so findIndex always returned -1 for them
- Ctrl/meta branch computes next set before the state update so
selectEntry can fire synchronously for child rows; auto-snap only
restores root entries, so ctrl-clicking a lone child never loaded its
detail panel — now calls selectEntry(entry) when the child ends up as
the sole selection, selectEntry(null) on multi or deselect
- selectedUids added to handleRowClick's useCallback dep array
* fix(frontend): resolve detail entry for any remaining child after ctrl-deselect
Add entryCacheRef (uid→entry Map) populated on every row click. When
ctrl/meta-deselecting leaves exactly one other entry selected, look up
the remaining UID in the cache before falling back to the root entries
array. Without this, deselecting from a multi-selection where the
remaining entry is a child row left selectedEntry null (auto-snap only
searches root entries).
* fix(frontend): child row stripes, deleted-child visibility, selection completeness
EntryRow.jsx / ChildRow:
- Index-based light/dark classes (child-entry-row--light/dark) replace
nth-child rules; parity comes from children.map idx so no sibling —
including the loading div — can shift the stripe order
- Loading div moved outside .child-entries so it never affects child
row ordering at all
- Accept deletedUids prop; filter expanded children array before render
so deleted children disappear immediately without waiting for reload
EntriesView.jsx: thread deletedUids through to EntryRow
App.jsx:
- Add deletedUids state; handleEntryDeleted/handleBulkDeleted populate it
- isRoot/hasChildDelete computed from entries before setEntries (safe in
StrictMode — no side-effects inside updater functions)
- Child delete triggers loadEntries to refresh stale parent child_count
and total_artifact_bytes
- handleRowClick ctrl/meta cache-miss branch uses det.summary (not det)
from fetchEntryDetail; archiveId added to dep array
- handleRowClick dep array includes archiveId
* fix(frontend): add inset stroke to child-entry-row.is-multi-selected
* fix(frontend): suppress mouse-click focus ring on entry expand button
* fix(frontend): index-based stripes for root and child rows
Root rows: EntriesView passes rowIndex from entries.map to EntryRow,
which applies entry-row-outer--light/dark. Retires both nth-child stripe
blocks so skeleton rows can never shift the first real entry to dark.
Skeleton rows keep a :not(.entry-row-outer) nth-child fallback.
Child rows: colours changed from the warm root palette (paper-3/#f2ede5)
to cooler near-whites (#fafaf8/#f2f0ec) so children are visually distinct
from their parent row regardless of which stripe the parent sits on.
* feat(core+frontend): bare yt:ID resolves to YouTube video
determine_source: yt:ID / youtube:ID with no prefix and exactly 11
chars [A-Za-z0-9_-] → YouTubeVideo via is_youtube_video_id helper.
Reserved prefixes (playlist/, channel/, c/, user/, @) still fire first
so they are unaffected by the fallback.
expand_shorthand_to_url: bare yt:ID expands to watch?v=ID using the
same predicate, consistent with how ytm:ID → music.youtube.com/watch.
isVideoSource (frontend): same 11-char /^[A-Za-z0-9_-]{11}$/ regex so
yt:ID triggers the quality probe and capture-row guards identically to
yt:video/ID.
Tests: test_is_youtube_video_id covers valid IDs (alphanumeric, with _
and -), too-short, too-long, and invalid-char cases using genuinely
invalid fixtures. test_youtube_sources adds bare-ID cases and confirms
reserved prefixes (playlist/, @) are not affected.
* fix(core+frontend): exclude avatar blobs from tweet % cached display
tweet/tweet_thread entries have avatar artifacts that are always
deduplicated from the first capture of each author. Counting them in
the cached-bytes percentage makes it artificially high (or incorrect
when the real media content is new but avatars are cached).
database.rs:
- refresh_entry_cached_bytes: AND ea.artifact_role != 'avatar'
- cascade_cached_bytes_after_delete: same filter
- cascade_cached_bytes_after_subtree_delete: same filter
- Initial cached_bytes migration: same filter
- Re-migration (else branch): recomputes cached_bytes for existing
entries that have avatar artifacts, scoped to only those entries
archive.rs:
- EntrySummary gains cacheable_bytes: i64 — non-avatar total bytes,
computed inline in every SQL query as the denominator for % cached
- ENTRY_SELECT_COLS adds cacheable_bytes at index 13 (with children
subquery, same as total_artifact_bytes)
- list_root_entries adds same expression
- All 6 row-mapping closures include cacheable_bytes: row.get(13)?
EntryRow.jsx:
- SIZE display stays: formatBytes(entry.total_artifact_bytes)
- % cached badge uses cacheable_bytes as denominator:
cached_bytes / cacheable_bytes * 100
* feat(frontend): j/k keyboard navigation for entries; test avatar cached_bytes
App.jsx — j/k handler:
- Fires on keydown when not focused on INPUT/TEXTAREA/SELECT/contenteditable
- Ignores meta/ctrl/alt modifier combos
- Uses DOM order (#entries-body [data-entry-uid]) so expanded child rows
participate, matching the shift-range selection logic
- Resolves target entry via entryCacheRef → root entries array →
fetchEntryDetail(..).summary on cache miss
- Scrolls target into view (block: nearest); updates lastAnchorIndexRef
so subsequent shift-click ranges start from the keyboard-navigated row
archive.rs — cached_bytes_excludes_avatar_blobs test:
- Two tweet entries sharing an avatar blob (100B) and a media blob (900B)
- Asserts refresh_entry_cached_bytes sets cached_bytes = 900 (not 1000)
- Asserts list_root_entries summary: cached_bytes=900, cacheable_bytes=900,
total_artifact_bytes=1000 — SIZE includes avatar, % cached denominator
and numerator both exclude it
* fix(frontend): guard j/k uncached-child fetch with monotonic token
Rapid j/k over child rows that aren't in entryCacheRef triggers
fetchEntryDetail calls in parallel. Without a guard the last one to
settle wins, desyncing selectedUids (highlight) from selectedEntry
(detail panel/URL).
Fix:
- jkSeqRef (monotonic counter) incremented before each server fetch;
the .then() guard tok === jkSeqRef.current drops results from
superseded navigations
- handleRowClick increments jkSeqRef.current so any click also cancels
an in-flight j/k fetch
* fix(frontend): guard ctrl/meta cache-miss fetch; add / search shortcut
App.jsx ctrl/meta branch: cache-miss fetchEntryDetail now captures
tok = ++jkSeqRef.current before the fetch and gates selectEntry on
tok === jkSeqRef.current — same pattern as the j/k handler — so a
slow response after a later click/navigate can't overwrite selection.
/ key: reuses the Cmd+K/Ctrl+K handler path (focus+select search input,
or pendingSearchFocus + archive view switch). Guards: editable targets
(INPUT/TEXTAREA/SELECT/contentEditable) and modifier keys are checked
before preventDefault() so typing / in inputs is untouched.
* feat(core): attempt SpotifyTrack via yt-dlp instead of hard-failing
Previously all Spotify sources returned an error claiming yt-dlp cannot
download DRM-protected audio. yt-dlp does have a Spotify extractor
(experimental, content-dependent), so refusing upfront is worse than
trying and letting it report the real failure.
SpotifyTrack now goes through the same yt-dlp metadata + audio-quality
download path as YouTubeMusicTrack. SpotifyAlbum and SpotifyPlaylist
still return an explicit error — fetch_playlist_info has YouTube-specific
URL fallback logic that would produce bogus child URLs for Spotify flat
entries; those sources need dedicated container handling first.
* chore: remove docs/superpowers from repo and gitignore whitelist
* feat(core+frontend): Spotify album/playlist capture via yt-dlp container path
Previously SpotifyAlbum and SpotifyPlaylist hard-failed with a DRM error.
yt-dlp does support Spotify (experimentally), so they now go through the
same probe/container/child path as YouTube playlists.
ytdlp.rs:
- Extract normalize_item_url() helper used by both fetch_playlist_info
and probe_playlist_qualities; eliminates the duplicate URL-normalization
blocks and the YouTube-specific fallback_host variable
- Fallback logic is now platform-aware: YouTube/YTM bare IDs → watch URL,
Spotify → open.spotify.com/track/{id}, unknown → skip with warning
capture.rs:
- locator_to_playlist_url: accept SpotifyAlbum | SpotifyPlaylist so the
probe-playlist endpoint accepts Spotify album/playlist URLs
- Container branch: add SpotifyAlbum | SpotifyPlaylist to the matches!
- is_audio: true for SpotifyAlbum | SpotifyPlaylist (audio-only, like YTM)
- child_source: SpotifyTrack for Spotify containers (not YouTubeMusicTrack)
- generate_entry_title and record_media_entry both use the correct child
source so entity_kind, source_kind, and representation_kind are right
CaptureDialog.jsx:
- isPlaylistSource: recognise open.spotify.com/album/ and /playlist/,
and spotify:album:ID / spotify:playlist:ID shorthands, so Spotify
containers get the probe UI, per-track excludes, and sync toggle
* fix(core+server): partial playlist refresh + child visibility inheritance
database.rs:
- finish_archive_run: reverted 'partial' status (violates CHECK constraint);
restored binary completed/failed
- get_run_completed_count(): new helper for callers that need to distinguish
partial success without touching the DB status enum
capture.rs:
- CaptureResult gains completed_count: i64 (0 for single-item captures;
populated from get_run_completed_count for container/playlist captures)
- All CaptureResult construction sites updated
routes.rs:
- Playlist job_status mapping: mark job 'completed' when completed_count > 0
(even if status == 'failed'), so partially-successful playlist captures
trigger onCaptured and show the archived entries without a manual reload.
Truly zero-success runs (completed_count == 0, status == 'failed') stay
failed as expected.
- probe_playlist_handler error message updated to mention Spotify
archive.rs (list_child_entries):
- Add third OR arm: children are visible when their parent container is in a
collection visible to the caller. Fixes empty child list for non-admin
users with access to a playlist container but not its newly-created
children (which are all in the default 'private' collection).
* fix(server): use completed_count to distinguish partial from total failure
finish_archive_run is binary (completed/failed) — a playlist where some
tracks succeed and some fail returns status='failed'. Without this fix,
the job mapping treated any failed status as a job failure, causing
onCaptured to never fire and leaving successfully archived entries hidden
until a manual reload.
Now: job is 'failed' only when status='failed' && completed_count==0.
Partial runs (completed_count > 0) map to a completed job so the UI
refreshes and shows the entries that were successfully captured.
* fix(core+server): exclude container from child success count; rename field; tests
database.rs:
- get_run_completed_child_count(): renamed from get_run_completed_count and
scoped to child items only (parent_item_id IS NOT NULL). The container
run item is always completed first, so the old function inflated the count
by at least 1 for every playlist, making all-video-failed runs appear as
partial successes.
- New regression test: completed_child_count_excludes_container — completes
both a root and a child item, asserts DB completed_count==2 while
get_run_completed_child_count==1.
capture.rs:
- CaptureResult.completed_count renamed to completed_child_count to match
the function and make the semantics unambiguous at the call site.
routes.rs:
- job_status decision now uses result.completed_child_count == 0 so that a
playlist where every video fails (child_count==0) is correctly reported
as a failed job, not a completed one.
archive.rs:
- New regression test: list_child_entries_inherits_parent_visibility —
enrolls only the container in a USER-visible collection, asserts the
child is visible to a USER caller and invisible to a GUEST caller.
|
||
|---|---|---|
| .. | ||
| LICENSE.md | ||
| README.md | ||
archivr
An open-source self-hosted archiving tool. Work in progress.
- Archiving
- Archiving media files from social media platforms
- YouTube Videos
- YouTube Playlists
- YouTube Channels
- Twitter Videos
- TikTok
- Snapchat
- YouTube Posts (postponed)
- Archiving local files
- Archiving Twitter Tweets, Threads, and Articles
- Archiving files from cloud storage services (Google Drive, Dropbox, OneDrive) and from URLs
- URLs
- Google Drive
- Dropbox
- OneDrive
- (Some of these could be postponed for later.)
- Archive web pages (HTML, CSS, JS, images)
- Archiving emails (???)
- Gmail
- Outlook
- Yahoo Mail
- Archiving media files from social media platforms
- Management
- Deduplication
- Tagging system
- Search functionality
- Categorization
- Metadata extraction and storage
- User Interface
- Web-based UI
- Authentication and login
- Archive setup
- Browse and view entries
- Tag management and filtering
- Search entries
- View archive runs
- Capture dialog
- User settings and API tokens
- Admin panel
- Web-based UI
- Backup and Sync
- Cloud backup (AWS S3, Google Cloud Storage)
- Local backup
Motivation
There are two driving factors behind this project:
- In the age of information, all data is ephemeral. Social media platforms frequently delete content, and cloud storage services can become inaccessible and unreliable. Being able to archive important data is very important for preserving personal memories and digital history.
- I will be creating a small encyclopedia for my future family and kids. Therefore, I want to make sure that all the information I gather is preserved and accessible for future reference.
This project aims to provide a reliable solution for archiving important data from various sources, ensuring that users can preserve their digital assets for the long term.
Archive Inputs
archivr archive <path> currently accepts three kinds of inputs:
- Local files via
file://... - Direct platform URLs
- Platform shorthand inputs such as
tweet:...,yt:..., orinstagram:...
Running Archivr
Archivr currently ships as two binaries:
archivr- The CLI for creating and writing to one archive.
- Use this for
initandarchive.
archivr-server- The web server for reading one or more existing archives through the browser UI.
- Use this after archives already exist.
With Nix, run the CLI with:
nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
Run the web server with:
nix run .#archivr-server -- ./archivr-server.toml
The server expects a TOML registry file. If no path is passed, it reads ./archivr-server.toml.
Example:
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
Then open:
http://127.0.0.1:8080
When installed through Nix, archivr-server is wrapped so it can find the static web UI assets automatically. The wrapper sets ARCHIVR_STATIC_DIR to the installed static asset directory. Running from source with cargo run -p archivr-server falls back to crates/archivr-server/static.
Security and Deployment
archivr-server is a local-only tool by default. It binds to 127.0.0.1:8080 and has no authentication or access control. Do not expose it to a public network or a shared LAN without understanding the risks.
Changing the bind address
You can set the bind address in your TOML config:
# Optional. Default: 127.0.0.1:8080
# Only change this if you know what you are doing — the server has no authentication.
bind = "127.0.0.1:9090"
Or override it with the ARCHIVR_BIND environment variable:
ARCHIVR_BIND=127.0.0.1:9090 nix run .#archivr-server -- ./archivr-server.toml
If the server is started with a non-loopback address (e.g. 0.0.0.0), it prints a warning to stderr:
warn: archivr-server is bound to 0.0.0.0:8080 — this server has no authentication. Only expose it on a trusted network.
When will auth be added?
Auth and session handling will be designed when remote or public hosting becomes a real requirement. Until then, keep the server on loopback. See crates/archivr-server/src/routes.rs for the route classification that will guide where middleware is applied.
Supported Platforms
- Local files:
file:///absolute/path/to/file.ext - YouTube media: individual videos/shorts, playlists, and channels; standard URLs or shorthand video inputs. Playlists and channels archive as a container entry with each video stored as a child entry beneath it.
- X/Twitter media from Tweets: normal Tweet URLs or the
tweet:media:IDshorthand - X/Twitter Tweet content scrape: Tweet and Thread shorthands. (These are saved as JSON files in
raw_tweets/) - Instagram, Facebook, TikTok, Reddit, Snapchat: direct URLs or platform-prefixed shorthand passed through to
yt-dlp
Video quality and audio-only downloads
When capturing via the web UI, entering a URL for a yt-dlp-backed source (YouTube, Instagram, TikTok, Facebook, Reddit, Snapchat, X media) triggers a metadata probe via GET /api/archives/:id/captures/probe. The quality selector then shows only the heights actually available in that video plus Best quality (default). An Audio only option is appended whenever the probe confirms an audio track exists. UI behaviour by probe outcome:
qualities |
has_audio |
UI shows |
|---|---|---|
["1080p", "720p", …] |
true |
Best / heights / Audio only |
["1080p", …] |
false |
Best / heights |
[] |
true |
Audio only (pre-selected, no Best option) |
[] |
false |
"No media detected" |
| probe fails (502) | — | picker hidden, capture still submittable |
The POST /api/archives/:id/captures endpoint accepts an optional quality field: "best", "audio", or any "NNNp" height string:
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "720p" }
{ "locator": "https://www.youtube.com/watch?v=...", "quality": "audio" }
"audio" selects the most efficient native audio track without transcoding: Opus/WebM is preferred (smallest at equivalent quality), then AAC/M4A, then whatever yt-dlp considers best. The saved file's extension matches the native format (.webm for Opus, .m4a for AAC, etc.) — no ffmpeg re-encode, no size inflation. Any "NNNp" height is accepted; the server builds the yt-dlp format selector with an unconditional /best fallback so the download succeeds even if the exact height is unavailable. Omitting quality or passing "best" downloads at the highest available quality. Anything else is rejected with HTTP 400.
The probe endpoint (GET /api/archives/:id/captures/probe?locator=…) requires auth and returns 200 with:
{ "has_video": true, "has_audio": true, "qualities": ["1080p", "720p", "480p"] }
{ "has_video": false, "has_audio": true, "qualities": [] }
{ "has_video": false, "has_audio": false, "qualities": [] }
has_video: false, has_audio: false means yt-dlp found no downloadable tracks (e.g. a tweet with no media). A 502 means yt-dlp itself failed (transient network error, rate-limit, unsupported extractor) — treat as inconclusive, not "no media."
YouTube playlists and channels
Capturing a YouTube playlist or channel URL creates a container entry for the playlist or channel, with each video archived as a child entry beneath it. Before downloading, the capture UI probes each video to fetch available quality options, letting you set quality per-video or apply a single quality to the whole playlist.
Incremental sync: When re-archiving a playlist or channel, enable sync mode in the capture dialog to skip videos that are already in the archive. Only new videos are downloaded; the existing container entry is reused.
Excluding individual videos: In the expanded per-video list, each video has a remove button (×) to exclude it from the current capture. Removed videos are not downloaded; the rest proceed normally.
Hosting on NixOS
The flake exposes a nixosModules.default output. Add it to your system flake and
enable the service:
# flake.nix (your system flake)
{
inputs.archivr.url = "github:thegeneralist/archivr";
outputs = { nixpkgs, archivr, ... }: {
nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
modules = [
archivr.nixosModules.default
{
services.archivr-server = {
enable = true;
# listenAddress defaults to "127.0.0.1" (loopback only)
# port defaults to 8080
archives = [
{ id = "personal"; label = "Personal"; path = "/srv/archivr/personal/.archivr"; }
{ id = "work"; label = "Work"; path = "/srv/archivr/work/.archivr"; }
];
};
}
];
};
};
}
The module:
- Creates an
archivrsystem user and group. - Generates the TOML config from your options and stores the auth database under
/var/lib/archivr-server/(persists across upgrades). - Runs under a hardened systemd unit (
ProtectSystem = strict,NoNewPrivileges,PrivateTmp, etc.). Archive directories are whitelisted for read-write access. - Restarts automatically on failure.
openFirewall — set to true to open the TCP port derived from bind.
Only needed when binding to a non-loopback address:
services.archivr-server = {
listenAddress = "0.0.0.0";
port = 8080; # explicit, though 8080 is the default
openFirewall = true;
};
Archive directories must be readable and writable by the archivr user.
Initialise them with archivr init first, then chown -R archivr:archivr /srv/archivr.
Hosting with Docker
A Dockerfile and docker-compose.yml are provided for self-hosting without Nix.
Quickstart
-
Copy the example config and edit it:
mkdir config cp docker/config.example.toml config/archivr-server.toml # edit config/archivr-server.toml — set archive id, label, and archive_path -
Initialize each archive on the persistent data volume before the first start. The image includes the
archivrCLI for this purpose:docker compose run --rm archivr archivr init /data/archives/main /data/archives/main/.archivr/store --name "Main Archive"This creates
/data/archives/main/.archivr/with the metadata the server requires. A baremkdiris not enough — the server readsnameandstore_pathfiles that onlyarchivr initwrites. -
Start the server:
docker compose up -dThen open
http://localhost:8080.
Volumes
| Mount | Purpose |
|---|---|
./config (read-only) |
Directory containing archivr-server.toml |
archivr-data named volume |
Auth database (/data/archivr-auth.sqlite) and archive directories |
Important:
auth_db_pathmust be set explicitly inarchivr-server.tomlto a path on the writable data volume (e.g./data/archivr-auth.sqlite). If left unset, the server defaults to writing the auth database next to the config file — which is on the read-only/configmount and will fail. The example config sets this correctly.
Twitter/X archiving
Supply a cookies file inside the config volume and set ARCHIVR_TWITTER_CREDENTIALS_FILE in docker-compose.yml:
environment:
ARCHIVR_TWITTER_CREDENTIALS_FILE: /config/twitter-cookies.txt
Building the image locally
docker build -t archivr-server .
The image compiles the Rust binary in a separate build stage so only the runtime dependencies (Chromium, Node.js, Python) land in the final layer.
Supported Shorthand Inputs
- YouTube video/short media:
yt:video/IDyoutube:video/IDyt:short/IDyt:shorts/IDyoutube:shorts/ID
- X/Twitter tweet JSON content:
tweet:IDx:tweet:IDx:x:IDtwitter:x:IDtwitter:tweet:ID
- X/Twitter media/video download:
tweet:media:ID
- X/Twitter thread JSON content:
x:thread:IDtwitter:thread:ID
- Other platform shorthands:
instagram:IDfacebook:IDtiktok:IDreddit:IDsnapchat:ID
Environment Variables
ARCHIVR_BIND- Optional.
- Overrides the bind address from the TOML config. Useful in Docker where you need
0.0.0.0:8080without editing the config file. Default:127.0.0.1:8080.
ARCHIVR_STATIC_DIR- Optional.
- Path to the directory of pre-built frontend assets served by the web UI.
Set automatically by the Nix wrapper and the Docker image. When running from
source with
cargo run, falls back tocrates/archivr-server/static.
ARCHIVR_YT_DLP- Optional.
- Overrides the
yt-dlpbinary used for YouTube, X media posts, Instagram, Facebook, TikTok, Reddit, and Snapchat downloads.
ARCHIVR_SINGLE_FILE- Optional.
- Overrides the
single-filebinary used for web page archiving. Set automatically by the Nix wrapper and the Docker image.
ARCHIVR_CHROME- Optional.
- Overrides the Chromium/Chrome executable passed to
single-filevia--browser-executable-path. Set automatically by the Nix wrapper and the Docker image. Default:chromium.
ARCHIVR_CHROME_ARGS- Optional.
- Space-separated extra flags appended to Chromium's
--browser-args. The Docker image sets this to--no-sandboxbecause Chromium refuses to run as root without it. Leave unset when running natively (Nix, Linux desktop). A--window-size=1920,1080is always passed to provide a realistic desktop viewport (so responsive @media rules and styles are evaluated and preserved correctly). Supply your own--window-size=...here to override.
ARCHIVR_TWITTER_CREDENTIALS_FILE- Required for tweet/thread scraping inputs such as
tweet:IDandx:thread:ID. - Must point to a cookies file for the vendored scraper.
- Required for tweet/thread scraping inputs such as
ARCHIVR_TWEET_SCRAPER- Optional.
- Overrides the tweet scraper script path. Default:
vendor/twitter/scrape_user_tweet_contents.py.
ARCHIVR_TWEET_PYTHON- Optional.
- Overrides the Python executable used to run the tweet scraper. Default:
python3.
Current Limitations
- Arbitrary
http://orhttps://URLs that return HTML are archived as self-contained single-file HTML snapshots viasingle-file-cli(requires Chromium). Plain file URLs (PDFs, images, zips, etc.) are downloaded directly. Requiressingle-fileand a Chromium binary on PATH, or theARCHIVR_SINGLE_FILE/ARCHIVR_CHROMEenv vars set. - Local files currently need to be passed as
file://...paths.
License
This project is licensed under the MIT License. See the LICENSE file for details.