1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
Commit graph

16 commits

Author SHA1 Message Date
archivr-qa
55a74ea3c2
fix: preserve X Article block order 2026-08-23 20:51:19 +02:00
archivr-qa
585d0a049b
feat(cli): add archivr yt-dlp update|status subcommand
`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.

The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.

`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.

reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 19:22:45 +02:00
1af920eb63
feat: share videos to TV (Chromecast + AirPlay) (#33)
* server: add scoped media-token endpoint for Cast/AirPlay auth bypass

Chromecast and Apple TV fetch media URLs as independent HTTP clients
with no session cookie. The existing serve_artifact handler requires
auth_user.require_auth(), so those devices always received 401.

Changes:
- MediaToken struct stored in AppState (Arc<Mutex<HashMap>>), scoped to
  a single (archive_id, entry_uid, artifact_index) tuple with a 2-hour TTL
- POST /api/archives/:id/entries/:uid/artifacts/:idx/media-token
  requires an authenticated session, verifies the artifact exists,
  prunes expired tokens, mints a 43-char URL-safe token, and returns
  { url, expires_in_secs }
- serve_artifact now accepts an optional ?token= query param; a valid
  scoped token bypasses require_auth() while a missing/invalid/expired
  token falls through to the normal 401 path
- CSP script-src extended to include https://www.gstatic.com so the
  Cast sender SDK script (injected lazily by VideoPreview) is not blocked
- 4 new tests: bare-URL still 401, tokenized fetch succeeds without
  session cookie, bogus token 401, wrong-artifact-index 401

* frontend: Cast/AirPlay overlay in VideoPreview

When the user opens a video archive entry, VideoPreview now:

1. Issues a signed media token (POST .../artifacts/:idx/media-token) and
   uses the returned signed URL as <video src>. This ensures the video
   element's src is one that Cast devices and Apple TV can fetch without a
   session cookie.

2. Lazily injects the Google Cast SDK script (cast_sender.js from
   gstatic.com, now allowed by the updated CSP). Once the SDK reports
   available, a <google-cast-launcher> web component appears as an overlay
   button in the top-right corner of the video. Selecting a Cast device
   triggers loadMedia() with the signed URL and the artifact's MIME type.

3. Detects AirPlay support (webkitShowPlaybackTargetPicker on
   HTMLVideoElement) and shows an AirPlay icon button alongside Cast.
   The <video> element carries x-webkit-airplay='allow', so Safari's native
   controls also surface the AirPlay option. The explicit overlay button
   calls webkitShowPlaybackTargetPicker() for consistent placement.

Both buttons are hidden when the respective APIs are unavailable (HTTP
pages, non-Safari for AirPlay, no Cast extension/devices), so there is no
UI regression for users who don't cast.

PreviewPanel now passes contentType (derived from artifact extension) to
VideoPreview so Cast receives a correct MIME type.

New CSS: .video-tv-controls (absolute overlay), .video-tv-btn (frosted
glass icon button), .video-tv-loading (placeholder during token fetch).

* server: fix serve_artifact auth OR logic — bogus token falls back to session

Previously a request carrying ?token=<expired> was immediately rejected
with 401, even if the user held a valid session cookie. This broke
logged-in browser playback after the 2-hour signed-URL window expired,
because VideoPreview uses the signed URL as <video src>.

Fix: compute token_valid first; if the token is absent or invalid, fall
through to auth_user.require_auth() instead of returning early.
Effect: valid token skips session check, invalid/missing token checks
session, both invalid → 401 as before.

Updated the bogus-token-no-session test docstring to clarify it tests
the no-auth path specifically. Added new test:
  media_token_bogus_token_with_session_returns_200 — verifies a logged-in
  user can still fetch the artifact via a URL carrying a stale token.

* frontend: guard token-fetch effect against stale async resolution

A slow issueMediaToken() response for video A could resolve after the
user selected video B and call setSignedSrc(urlA), making the
preview/Cast play the wrong file.

Add a cancelled flag set in the effect cleanup; both .then and .catch
check it before touching state, so only the most recent src wins.

* frontend: load Cast media immediately if session already exists

Previously the effect only sent video to the TV on SESSION_STARTED /
SESSION_RESUMED events. Two gaps:

1. If a Cast session was already active when signedSrc became ready
   (e.g. the SDK resumed a session before the token fetch finished, or
   the user switches videos while already casting), nothing was sent.

2. Same gap if castReady fired after an already-established session.

Fix: extract loadMedia(session) and call it against
ctx.getCurrentSession() immediately when castReady + signedSrc are both
truthy, in addition to keeping the event listener for future connects.

* server: staged file-upload endpoint

POST /api/archives/:id/uploads streams a multipart body to a temp file
under the archive's store/temp/ directory and returns a staged_path the
capture pipeline can move into place.

- Routes: /api/archives/:id/uploads (POST, requires auth)
- Body cap: 10 GiB; chunk-streamed to disk, never buffered in memory
- Path-traversal sanitised on the filename field
- Temp files are cleaned up on error paths (disk-leak fix)
- main.rs wires the new route into the server startup
- Cargo: adds the multipart dependency

* frontend: file upload in Capture dialog

Drag-and-drop or 'Upload file' button stages files for archiving:

- File items sit alongside URL rows in the same list; each shows the
  original filename, a live progress bar during upload, and a check badge
  when ready.  The locator input is replaced entirely — no editable field.
- Archive button is disabled until all uploads finish; each file item
  contributes to the Archive N count once its upload is done.
- File items are excluded from sessionStorage persistence (they are
  transient — the staged server path would be invalid after a reload).

Staged-file cleanup is handled at every exit path so temp/uploads/ does
not accumulate:
  • removeRow on an in-progress item aborts the XHR; removeRow on a done
    item calls DELETE /archives/:id/uploads.
  • Dialog cancel (Escape / Cancel button) aborts all in-flight XHRs and
    DELETEs all completed staged files via the close-event handler.
  • handleArchive sets isSubmittingRef=true before dialog.close() so the
    close handler skips cleanup — the background capture job handles
    staged-file removal on success instead.
  • uploadFile() returns { promise, abort } so the component can cancel
    the XHR without any visible fetch.

api.js additions: uploadFile (XHR with progress + abort), deleteUpload.

(Static assets rebuilt from combined source to include screensharing
changes from this branch.)

* fix: collection enrollment with default_visibility_bits

Two related fixes from feat-file-uploading:

core: fix collection enrollment using default_visibility_bits instead of
entry.visibility — entries were being enrolled with the entry-level
visibility rather than the collection's configured default.

server: allow changing default_visibility_bits on the default collection
— the PATCH handler was incorrectly blocking updates to the default
collection's visibility configuration.

* server: fix unbounded staged-upload disk growth

Two review findings:

P2 — delete staged file on capture failure (routes.rs)
When perform_capture returns Err, the job was marked failed but
staged_upload_path was never removed. With a 10 GiB body cap a few
failed imports could exhaust archive storage before the next restart.
Mirror the success-path cleanup into the Err arm so the file is removed
immediately regardless of outcome.

P1 — periodic staged-upload pruning (main.rs)
The startup prune of temp/uploads/ only ran once, so uploads abandoned
mid-session (browser crash, navigation away) accumulated forever on a
long-running server. Folded the pruning logic into the existing 24 h
maintenance task alongside session cleanup, so stale dirs are swept
continuously without requiring a restart.

* server+frontend: fix staged-upload disk-growth and prune safety

Server (main.rs + routes.rs):

- Extract prune_stale_upload_dirs() helper called by both startup and
  the periodic 24h task, eliminating the duplicated loop.

- Sentinel (.uploading) created in the UUID dir before streaming begins;
  removed on successful completion; error path uses remove_dir_all so
  the partial file and sentinel are cleaned up together.
  The periodic prune skips any dir containing .uploading (active XHR).

- Startup prune passes cleanup_stale_sentinels=true: the server has not
  started accepting connections yet so any sentinel is a crash remnant —
  it is removed and the dir proceeds to the age check, preventing leaked
  dirs from a previous crash accumulating forever.

- Staleness measured from the newest non-sentinel child file mtime so a
  just-finished slow upload (dir mtime stale, file mtime fresh) is not
  pruned before the user can submit it for capture. Empty dirs fall back
  to dir mtime.

- Failed captures (Err branch in spawn_blocking) now also delete the
  staged file and UUID dir immediately, matching the success path.

Frontend (api.js + CaptureDialog.jsx):

- submitCapture attaches err.status = res.status on non-2xx responses
  so callers can distinguish a definite HTTP rejection from a network
  error where the response may have been lost.

- submitBgJob catch deletes the staged file only when e.status is set
  (server definitively rejected the POST /captures request). A network
  error leaves the file in place because the server may have accepted
  the job and the response was lost — deleting would race the capture.
2026-07-23 20:12:22 +02:00
d610d37793
feat: entry previews (#24)
* feat: entry previews (video, tweet, article, iframe, image, audio bar)

- Add PreviewPanel dispatch hub routing by entity_kind + artifact extension
- VideoPreview: HTML5 <video> for YouTube/Instagram/TikTok/Reddit/X posts
- TweetPreview: tweet card, thread, and X article renderer (ported from x-article-renderer)
- IframePreview: sandboxed iframe for SingleFile web pages and PDFs
- ImagePreview: image viewer with click-to-open-fullsize
- AudioBar: persistent fixed-bottom player (Spotify-style) that survives entry navigation
- Lift entryDetail to App.jsx, shared between PreviewPanel and ContextRail
- 3-column layout (workspace | 300px preview | 340px rail) when preview active
- Stale-guard fixes: seq incremented before early returns in all async effects
- handleRearchive: capture startSeq/entryUid at call time, guard every async resume
- TweetPreview: reset loading/error/tweets before early-return branches

* feat: entry previews — tweet/thread/article/video/audio/image/iframe/pdf

- PreviewModal: modal overlay with new-tab link (↗) and keyboard close
- PreviewPanel: routes by entity_kind + primary_media extension to the
  correct viewer (tweet/video/audio/pdf/html/image/fallback)
- TweetPreview: full X-style tweet, thread, and article renderer with
  local artifact map for archived media (CDN fallback)
- AudioBar: persistent fixed bottom player, triggered via ContextRail
  Play button; body.has-audio-bar pads content above it
- VideoPreview, IframePreview, ImagePreview: inline viewers
- PreviewPage: standalone /preview/:archiveId/:entryUid route
- ContextRail: Play/Preview buttons; isAudio/isPreviewable detection
- App.jsx: preview modal state, currentAudio state, preview route guard,
  has-audio-bar body class effect
- routes.rs: CSP updated (media-src self blob https; frame-ancestors self;
  Google Fonts + external images/scripts whitelisted)
- styles.css: preview modal, tweet-wrap scroll (min-height:0), audio bar
  body padding, newtab button, preview panel flex layout

* style: tighten article/tweet preview spacing

- aMeta padding: 14px → 10px
- article title marginBottom: 10px → 8px
- aAuthorRow marginBottom: 10px → 8px
- bH1 top margin: 20px → 16px
- bH2 top margin: 18px → 14px
- bHr margin: 20px → 14px
- .preview-tweet-wrap padding: 20px → 12px (already committed)

More content visible above the fold in both modal and standalone views.

* feat: tweet/article preview quality pass

HTML entities: decode &gt; &lt; &amp; etc. on sliced segments only
(entity offsets index the stored string; decoding before slicing shifts them)

Image lightbox: click any tweet/thread/article image to open full-screen
viewer; cmd+click follows <a> to open in new tab; arrow-key + ‹ › nav;
Escape closes; 1/N counter; ↗ open-in-new-tab link

Multi-image grid: 2 photos → side-by-side (180px rows); 3 → left spans
both rows; 4 → 2×2 (140px rows); single image unchanged

Empty media grid ghost: build photos/videoItems arrays first, render
.mediaGrid div only when at least one item resolved (advisory: never gate
on raw media.length when map items can all return null)

QT indicator: ↻ QT badge on tweet.is_quote_status === true entries

Modal shrinks for short content: height: 88vh → max-height: 88vh;
.preview-modal-body gets max-height: calc(88vh - 52px) so long threads
still scroll (advisory: don't rely on flex:1 once parent has no fixed height)

Video scrollbar leak: .preview-modal-body overflow: auto → hidden; each
child (tweet-wrap, video-wrap, iframe) manages its own scroll surface

Styled thin scrollbar on .preview-tweet-wrap (matches workspace rail)

ArticleRenderer: cover image and body images are lightbox-clickable;
opts thread through renderBlocksJSX → renderBlockJSX → renderAtomicJSX

* fix: move artifact fetch to api.js; stop Escape propagation from lightbox

- Export fetchEntryArtifacts(archiveId, entryUid, indices) from api.js
  using Promise.all + getJson (follows project convention: all /api calls
  go through api.js, never inline fetch in components)
- TweetPreview: import fetchEntryArtifacts, replace inline Promise.all
- MediaLightbox keydown handler: stopPropagation + preventDefault for
  Escape/ArrowLeft/ArrowRight so the parent PreviewModal window listener
  does not also fire and close the modal behind the lightbox

* fix: iframe/page preview height chain and toolbar UX

Problem: changing .preview-modal from height to max-height broke iframe
previews - <iframe style='flex:1'> needs a concrete ancestor height, which
max-height alone doesn't supply when content is shorter than the cap.

Fix - CSS:
  .preview-modal--full { height: 88vh } applied to non-tweet modals
  .preview-modal--full .preview-modal-body { max-height: none }
  .preview-iframe-toolbar span: remove text-transform/letter-spacing
    (was uppercasing the URL/title in shouty caps)

Fix - PreviewModal: className adds --full when entity_kind is not
  tweet/tweet_thread; tweet previews keep shrink-to-fit behavior.

Fix - PreviewPanel: pass title + original_url from summary to IframePreview
  for both HTML and PDF; wrappers use flex:1/minHeight:0 not height:100%.

Fix - IframePreview:
  - Accept title + originalUrl props; show originalUrl in toolbar (falls
    back to artifact src only when original_url absent); show title above
    URL when available
  - flex:1 + minHeight:0 instead of height:100% on the wrap div
  - Single unified layout for page + pdf (both just show the iframe)

* feat: expand t.co links; linkify bare URLs in tweet and article text

Frontend:
- resolveEntityBounds: try multiple candidate strings in order (u.url
  first, since that's the t.co short URL that appears in full_text)
- normalizeUrlAnn: multi-candidate search; href = expanded > url,
  display = display_url > expanded > url
- linkifyText(): regex linkifier for entity-less bare URLs; trims
  trailing punctuation [.,;:!?)] before linking; used in both
  renderTweetTextJSX and renderInlineJSX including their early-return
  paths (anns.length === 0) that previously bypassed linkification
- renderInlineJSX: fix mention href mention.name → screen_name;
  replace t.co segment text with url.display when entity covers it

Scraper (vendor/twitter/scrape_user_tweet_contents.py):
- extract_tweet_data: when note_tweet text is used, pull urls/mentions/
  hashtags/symbols from note_result.entity_set (correct indices for the
  note text); keep media from legacy.entities (no note media downloads)

* feat: server-side t.co resolver + frontend augmentation

Server (routes.rs):
  POST /api/util/resolve-tco — unauthenticated, accepts JSON array of
  https://t.co/<alphanumeric> URLs only (strict regex validation, no SSRF
  via input), capped at 50 per batch, 3 s timeout, redirect(Policy::none)
  so the server only ever touches t.co itself. HEAD first, GET fallback if
  HEAD returns no Location. Location sanitized to http/https only —
  javascript:/data:/etc. fall back to the original t.co.

api.js:
  resolveTcoUrls(urls) — project-convention wrapper for the new endpoint.
  Returns {} on failure (callers degrade gracefully to bare t.co links).

TweetPreview.jsx:
  After tweet data loads, per-tweet range-based coverage detection:
  builds covered [start,end) from existing entity fromIndex/toIndex or
  indices fallback, then finds regex matches whose span is NOT covered.
  Resolves unique uncovered t.co URLs via resolveTcoUrls(), synthesises
  one entity per occurrence (with exact fromIndex/toIndex so normalizeUrlAnn
  gets correct bounds even for duplicate t.co URLs in the same tweet).
  Augments entities.urls before setTweets() so all rendering paths
  see expanded URLs.

* feat: suppress rendered media attachment URLs from tweet text

renderTweetTextJSX now accepts skipSpans=[] as third param.
Skip-span boundaries are added to the pts split set so a trailing
media t.co inside a plain segment still gets isolated and suppressed—
not re-linked by linkifyText. Early return only when both anns and
skipSpans are empty.

TweetCard computes mediaSkipSpans after building photos/videoItems:
for each rawMedia item whose src resolved (photo src match; any video),
resolveEntityBounds(m, ft, m.url) gives the precise [s,e] span using
indices/fromIndex first, indexOf fallback—then the span is passed to
renderTweetTextJSX so the t.co attachment URL is silently dropped.
2026-07-12 16:23:04 +02:00
dae61e585d
feat: add user-configurable cookie rules (#20)
Adds per-instance cookie rules (admin-only) that are injected into
every network touchpoint during capture.

Storage:
- New cookie_rules table in the auth DB (idempotent migration)
- Rules have pattern_kind (global/wildcard/regex), optional url_pattern,
  and cookies_json (validated as string-only JSON object)

Matching (resolve_cookies_for_url):
- Global rules always apply
- Wildcard: * and ? with full metacharacter escaping; matched against
  hostname via reqwest::Url when pattern has no ://, full URL otherwise
- Regex: matched against the full URL
- Later rules in ordinal order override earlier ones per cookie name

All six network touchpoints receive resolved cookies:
- http::probe_url_kind and http::download: Cookie request header
- singlefile::save: Netscape cookie file -> --browser-cookies-file
- ytdlp::fetch_metadata and ytdlp::download: Netscape cookie file -> --cookies
- tweets::archive: semicolon credentials file -> --credentials-file
  (only when both ct0 and auth_token are present; otherwise falls back
  to ARCHIVR_TWITTER_CREDENTIALS_FILE)

Security:
- Cookie files written 0o600 (owner read/write only)
- Exact parsed hostname used as cookie domain (no PSL stripping)
- Files deleted unconditionally before any error propagates,
  including spawn failures (hold-result-then-delete pattern)
- No cookie values in process args (no --add-header exposure)

API: GET/POST /api/admin/cookie-rules, PATCH/DELETE /api/admin/cookie-rules/:uid
Frontend: Cookies tab in Settings (admin only) with rule list,
  inline edit, pattern-type selector, client-side JSON validation
CLI: CaptureConfig::default() - no behaviour change

254 tests passing (4 new cookie-rule handler tests)
2026-07-06 19:01:34 +02:00
7ed7cb882f
chore: update Cargo.lock after parking_lot dep addition 2026-06-29 20:13:59 +02:00
4d0daf2dc0
chore: update Cargo.lock after collections feature 2026-06-26 17:07:48 +02:00
7cebf06124
chore: update Cargo.lock after frontend build 2026-06-26 12:03:30 +02:00
57fc48d73c
feat(auth): add argon2, rand, axum-extra dependencies 2026-06-26 11:28:51 +02:00
ee697625fb
feat(singlefile): add SaveResult, favicon extraction, wait for networkidle2 2026-06-24 19:16:21 +02:00
fd06632073
feat(singlefile): add single-file-cli downloader module 2026-06-24 18:38:57 +02:00
03abfb4d18
feat(core): generic HTTP/S file URL capture (Track 1)
- Add crates/archivr-core/src/downloader/http.rs
  - download(url, store_path, timestamp) -> Result<(hash, extension)>
  - Rejects text/html responses with a clear error
  - Derives extension from URL path or Content-Type header
  - Follows redirects (capped at 10), user-agent archivr/0.1
  - 10 unit tests for extension/content-type helpers
- Add Source::Url variant to capture::Source enum
- determine_source: unmatched http/https URLs route to Source::Url
- source_metadata: Source::Url => ("web", "file", "file")
- generate_entry_title: Source::Url arm -> "Downloaded File" fallback
- perform_capture: Source::Url arm calls http::download, uses existing
  temp -> hash_exists -> move_temp_to_raw -> record_media_entry pipeline
- Update test expectations: 3 plain https:// cases now expect Source::Url
- Add reqwest 0.12 (blocking) to workspace and archivr-core deps
- Mark URLs milestone done in docs/README.md
- Update NEXT.md Track 1 status
2026-06-24 14:37:08 +02:00
5803f2119b
feat(server+ui): tag management API routes and browser UI
- 5 tag routes (list tree, create, assign, remove, search with tag=)
- Tags nav, tag tree view, tag filter badge
- Entry tag pills with remove, assign-tag form in context rail
2026-06-22 14:48:47 +02:00
b56c969624
feat: add db and multi-archive web UI foundation (#8)
* Add SQLite metadata database support

* Implement archive metadata database

* chore: let's guess cargoHash because there's something wrong with nixpkgs!

* Gate test-only database helpers behind cfg(test)

* Fix archive database row identity

* Use serde for archive metadata JSON

* Finalize archive runs at command level

* Handle archive command errors without panics

* Cover tweet entry metadata recording

* Document static regex invariants

* docs: add web UI design spec

* docs: add web UI implementation plan

* chore: move cli into workspace crate

* chore: track workspace crates directory

* refactor: extract archive core crate

* refactor: add core archive opening APIs

* refactor: rename taxonomy model to tags

* feat: add archive query APIs

* feat: add web server registry

* feat: expose archive server APIs

* feat: add archive table web UI

* fix: complete web UI smoke path

* docs: add architecture mental model

* docs: remove private superpowers plans

* nix: split cli and server packages

* chore: remove PLAN.md
2026-06-14 00:27:16 +02:00
2d59ab0af5
feat: add archiving of platform media files (#1)
* chore: specify non-ignored `.md` files

* refactor: rename youtube downloader to ytdlp

More generic name since yt-dlp supports many sites beyond YouTube.

* feat: add local file downloader

Supports file:// URLs for archiving local files.

* deps: add regex crate for URL pattern matching

* feat: expand source detection with granular YouTube types

- Split Source::YouTube into YouTubeVideo, YouTubePlaylist, YouTubeChannel
- Add Source::X for Twitter/X posts
- Add Source::Local for file:// URLs
- Add regex-based URL pattern matching for YouTube URLs
- Add shorthand schemes (yt:video/ID, youtube:playlist/ID, etc.)
- Add comprehensive tests for all URL patterns

* docs: update README milestones

Mark YouTube videos, Twitter videos, and local files as done.

* chore: update flake.lock

* feat: add shorthand schemes for X/Twitter media

* chore: move docs into docs dir

* Remove temp file using timestamp path

Delete the temp entry at store_path/temp/<timestamp> in both
the hash-exists and success paths. Stop constructing the full filename
with extension and remove the early process::exit to de-duplicate
cleanup.

* Add Nix caches and default flake package

* Add social platform source detection and update milestones

* Tighten social URL matching to avoid false positives

* Mark media archiving milestone complete
2026-03-31 12:39:35 +02:00
01cc7826bf
feat: finish YouTube downloading; primitive "raw" archiving 2025-10-15 00:29:08 +02:00