1
Fork 0
mirror of https://github.com/thegeneralist01/archivr synced 2026-10-09 12:55:00 +02:00
Commit graph

13 commits

Author SHA1 Message Date
79ac44834e
feat: capture, summaries, search, and yt-dlp reliability (#38)
* ui: show spinner for pending captures

* feat(core): add text capture path with title + Markdown/plain body

- Add downloader/text.rs module with save() function that stages and hashes text content
- Support text/markdown and text/plain MIME types with .md and .txt extensions
- Add perform_text_capture() function for capturing user-supplied text
- Validates title (non-empty, max 500 chars) and body (non-empty, max 2 MiB)
- Creates blob records and entries with source_kind='text', entity_kind='document'
- Includes comprehensive unit tests for markdown, plain text, and validation

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(server): add POST /api/archives/:archive_id/captures/text

- Add CaptureTextBody struct for title, body, and optional MIME type
- Implement capture_text_handler with validation for empty fields and MIME type
- Route text submissions to perform_text_capture() in background
- Reuse existing capture job tracking and polling infrastructure
- Default MIME type to text/markdown when not specified
- Include route tests covering happy path, validation, auth, and error cases

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(frontend): add text-capture form to CaptureDialog

- Add submitTextCapture API client function with same error handling as submitCapture
- Create makeTextItem() factory for text capture state
- Implement CaptureTextRow component with title, body textarea, and MIME selector
- Add 'Add text' button in capture dialog toolbar
- Update handleArchive to filter and route text submissions
- Modify submitBgJob to detect and submit text items via submitTextCapture
- Skip probe and conflict checks for text items
- Reuse job tracking and batch settlement for text captures

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* feat(core): add entry_summaries schema + summarizer trait/providers

Per-entry LLM summaries as a regenerable child record, not a column on
archived_entries and not an on-disk artifact: an entry may carry several
summaries (one per provider/model/prompt version), any of which can be
discarded and recomputed. Generation is manual-only — nothing in capture.rs
calls into this module.

- database.rs: entry_summaries table + index, EntrySummaryRecord, and
  upsert/update/find/latest helpers mirroring the capture_jobs style.
  provider_model is stored as '' rather than NULL because SQLite treats
  NULLs as distinct inside a UNIQUE index, which would stop the CLI
  providers (no model) from ever deduping on the cache key.
- summarizer.rs: SummaryProvider trait with four implementations —
  Anthropic Messages API, OpenAI-compatible chat completions, `claude -p`
  and `codex exec -`. Configuration comes from env vars only (never TOML),
  matching how yt-dlp / single-file / tweet-scraper are resolved, which
  also keeps API keys out of anything the archive persists.
- archive.rs: EntryDetail gains latest_summary, populated by one extra
  LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row
  rather than duplicating it.

Implementation notes:
- No tokio in core. CLI timeouts are enforced structurally: stdout is
  drained on its own thread and handed back over a channel so the calling
  thread can recv_timeout and kill an overrunning child; stdin is written
  on a third thread so a 48 KB prompt cannot deadlock against a child
  waiting for us to read.
- HTML is reduced with regex rather than a parser: html5ever is not in the
  tree, and a model tolerates imperfect whitespace. Paired tags are spelled
  out per tag because Rust's regex engine has no backreferences by design.
- reqwest is declared with only the `blocking` feature here, so bodies are
  serialized via .body(value.to_string()) instead of widening the
  workspace dependency for .json().
- input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one
  hashing primitive; the content is truncated to 48 KB *before* hashing so
  the cache key describes exactly the bytes the model saw.

Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for
one JSON shape is a poor trade, so the two halves that can actually break
are tested directly — request-body builders and response parsers — leaving
only reqwest's own transport uncovered. Plus schema idempotency, cache-key
dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML
and tweet extraction, output normalization, and the CLI runner's stdin
round-trip, timeout kill, and nonzero-exit paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(server): add GET/POST /api/archives/:id/entries/:uid/summary

GET is read-only and gated exactly like entry detail, so a guest can read a
summary only for an entry whose content they could already read. POST
requires ROLE_USER, matching capture / tags / patch / rearchive; no auth
roles change.

Both the provider config and the content extraction resolve on the request
thread, before spawn_blocking. That is what lets a missing env var come back
as a synchronous 400 naming the exact variable, and an unsummarizable
artifact (video, audio) as a 400 saying so, rather than becoming a
background job the caller must poll only to learn about a config typo.

The pending row is claimed before spawning so the 202 can name a summary_uid
the client can poll immediately. summarize_entry owns the
pending → running → completed/failed transitions for that same row — the
cache key is identical, so both upserts resolve to one row — leaving the
handler to catch only the case where it fails before recording anything.
When !force and an identical cache key already completed, the existing row
comes back as a 200 with no new work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(frontend): render Summary rail section + provider selector

New "Summary" rail section between the URL/Preview controls and .meta-list.
A completed summary renders as bold tl;dr, body paragraph, tag chips, and a
provider · model footer; missing or failed shows Generate; pending/running
shows an inline spinner and polls GET every 1500 ms until terminal.

- api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps
  ApiError's { "error": ... } body so the missing-env-var message reaches
  the user verbatim rather than as a bare status code.
- ContextRail.jsx: state seeds from detail.latest_summary so the section
  renders immediately on selection. Polling is anchored on the summary
  status rather than started inside the click handler, so a job still
  running when the user navigates away and back is picked up again. A
  transient poll failure is swallowed — the next tick retries, and a real
  failure arrives as status === 'failed'.
- Regenerate passes force:true only when a completed summary is already
  shown; otherwise the request can take the server's 200 cache-hit path.
- Provider choice persists in sessionStorage under archivr:summary:provider,
  with try/catch around both accessors for private-mode browsers.
- Public sessions never see the selector or the Generate button, and the
  section renders at all only when a completed summary made it through the
  server's visibility gate.
- styles.css: .rail-summary-* only; spacing and the action button reuse
  .rail-section and .rail-rearchive-btn. The spinner honours
  prefers-reduced-motion — the text alone conveys the state.
- AGENTS.md: document the summary env vars alongside the existing
  external-tool convention.

Smoke-tested end to end against a scratch archive with a seeded markdown
entry: claude_cli produced a real summary (pending → running → completed in
~11s); a local mock server exercised the openai_compatible transport and
confirmed the Bearer header, model, and system/user role split on the wire;
unconfigured providers return 400 naming the exact variable; a video entry
returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without
adding a row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): codex_cli — auto-discover binary + use --output-last-message

Two related fixes for the codex_cli summary provider:

1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but
   without it the code resolved to bare `codex` and relied on PATH.
   The ChatGPT desktop app installs codex at
   `/Applications/ChatGPT.app/Contents/Resources/codex` and does not
   put it on PATH, so users who only have the desktop app saw
   'No such file or directory' with no hint. `resolve_cli` now walks
   env override → a small set of well-known absolute paths → HOME
   /.local/bin/<bare> → bare fallback. Same treatment applied to
   claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/
   claude, HOME/.local/bin/claude).

2. Clean output. `codex exec -` writes a runtime header ("OpenAI
   Codex vX", session id, sandbox, model), the assistant reply, and a
   footer ("tokens used", replay of the reply) to stdout. The JSON
   extractor took the first '{' from the *user prompt echo* and the
   last '}' from the trailing replay, producing invalid text that
   fell through to the "raw text under summary" fallback path. Now
   uses `--output-last-message <tempfile>` and reads only the final
   assistant message. Fallback (positional prompt) uses the same
   flag. Tempfile is cleaned up on all paths, incl. spawn failure.

* fix(frontend): give text-capture row real CSS

The text row shipped with semantic classnames (`capture-text-inputs`,
`capture-text-title`, `capture-text-body`, `capture-text-mime`,
`capture-text-icon`) but no CSS rules. Falling through to the parent
`.capture-row-main` flex-row (`display: flex; align-items: center`)
meant the title, textarea, and mime-select stacked as intrinsic-width
boxes centered on the tall body, producing a layout where the body
floated to the top-right, the title box appeared BELOW it, and the mime
selector rendered as an unstyled OS dropdown.

Fix:
- `.capture-text-row .capture-row-main` uses `align-items: flex-start`
  so the leading icon and trailing × pin to the top of the block.
- `.capture-text-inputs` is now a full-width column-flex container with
  proper gaps.
- `.capture-text-title` reuses the 44px input height and typography of
  `.capture-input`; `.capture-text-body` gets a 140px min-height,
  vertical resize, and matching border/focus treatment.
- `.capture-text-mime` is styled as a small chip with a custom caret
  so it matches `.capture-quality` and stops looking like a raw
  `<select>`. Sits in a right-aligned footer under the body.
- `.capture-text-icon` gets a 44px column so it aligns with the title
  input; remove button gets a small top-margin for the same reason.

Rebuilt static bundle bumped as well (`index-BLxoi9rt.css`,
`index-CQcpPA_I.js`).

* chore(static): rebuild bundle after text-row CSS merge

* fix(core): summarize tweets + walk all tweets in a thread

Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.

Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.

Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
  role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
  `raw_tweet_json` artifacts (falling back to `primary_media` for
  archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
  markdown/html/json branches, then join thread pieces with a
  `---` separator so the model sees a real paragraph break between
  statuses instead of one flowing document.

Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.

* feat(frontend): preview text-capture entries (.md / .txt)

Text captures land as `.md` (Markdown) or `.txt` (plain) blobs, but
PreviewPanel only dispatched on video/audio/image/pdf/html extensions,
so opening a text entry hit the 'No preview available' fallback with
the raw artifact path exposed.

- New `TextPreview` component fetches the primary artifact as text,
  renders it in a monospace `<pre>` with word-wrap, and shows the
  entry title on top and the MIME as a small trailing tag. Handles
  loading/error states.
- `PreviewPanel` gains a `TEXT_EXTS` set + a branch that dispatches
  to `TextPreview` for `md` / `markdown` / `txt`.
- CSS is padded and centered to ~780px so a text note reads like a
  document rather than an edge-to-edge terminal dump.

v1 intentionally does NOT parse Markdown: keeping frontend deps at
react+react-dom only. Bump to a real Markdown renderer if we start
capturing Markdown-authored notes.

* fix(nix): pin yt-dlp from its own release + wire into server wrapper

Two independent problems, one commit:

1. Stale binary. nixpkgs-provided `pkgs.yt-dlp` on the pinned
   nixos-unstable rev is 2026.03.17 (Mar 2026). yt-dlp itself
   releases days-to-weeks, and YouTube frequently rotates the
   player-signature / client surfaces the older builds request
   (`android_vr` is the current casualty), which returns HTTP 403
   mid-download for the format specs archivr passes (`-f
   bestvideo+bestaudio/best`). Even bumping the nixpkgs input would
   leave us dependent on that channel's yt-dlp cadence.

   Fetch the upstream zipapp directly instead
   (github.com/yt-dlp/yt-dlp/releases/download/<ver>/yt-dlp), wrap so
   `python3` and `ffmpeg` are on PATH, and pin version+hash in one
   place. Bumping is: change version, replace hash from
   `nix hash file <url>`.

2. Missing pin in server wrapper. `archivr-cli` was already wrapped
   with `--set ARCHIVR_YT_DLP` + a PATH prefix; `archivr-server`
   was NOT — it only pinned single-file, chrome, and the tweet
   scraper, silently falling back to whatever `yt-dlp` the user
   happened to have on PATH. Server captures therefore inherited
   the user's (often stale) system yt-dlp regardless of the flake
   pin. Same wrapper flags now apply to both binaries.

devShell keeps `pkgs.yt-dlp` for now: the dev shell is a
convenience, not a release surface, and matching wouldn't fit in this
commit without duplicating the derivation across let-scopes.

* chore(static): rebuild bundle for round-3 fixes

* chore(nix): pin python 3.12 for yt-dlp zipapp (avoid py3.14 libffi crash on darwin/arm64)

* feat(core): resolve_yt_dlp picks the newer of pinned vs state-dir

The nix flake wrapper pins a yt-dlp via ARCHIVR_YT_DLP, but yt-dlp rots
fast — extractors break within weeks of a pin. Add a resolver that probes
`--version` on both the pinned binary and a user-installed copy under the
mutable state dir, and runs whichever is newer.

Version strings are YYYY.MM.DD, so plain string ordering is chronological.
Ties resolve toward the state dir: a user who installed it there did so
deliberately. ARCHIVR_YT_DLP_FORCE bypasses the comparison entirely, and
with no candidate at all we fall back to bare `yt-dlp` on PATH — exactly
the previous behaviour.

Resolution is cached in a OnceLock so `--version` costs one subprocess per
process, and all four inline env::var lookups now go through it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump yt-dlp from upstream releases, not nixpkgs

The flake no longer takes yt-dlp from nixpkgs; a dedicated `ytDlp`
derivation fetches the upstream release binary directly and pins both
`version` and an SRI `hash`. That makes the previous workflow inert: it
ran `nix flake update nixpkgs` and compared `nixpkgs#yt-dlp.version`
before and after, so it could churn the lockfile forever without ever
moving the version we actually ship.

The workflow now reads the pinned version straight out of the `ytDlp`
block in flake.nix, asks the GitHub API for yt-dlp's latest release tag,
short-circuits when they already match, downloads the new release to
recompute its SRI hash (required — the hash is part of the derivation's
identity, so the URL cannot be changed alone), and rewrites the three
pinned fields under a sed range address scoped to that block so sibling
pins like ublockLite and isdcac are untouched. It asserts only flake.nix
changed and that the new version appears exactly twice before opening
the PR.

* feat(cli): add `archivr yt-dlp update|status` subcommand

`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.

The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.

`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.

reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(maintainer): document summarizer, text capture, and yt-dlp lifecycle

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(readme): document LLM summaries, text notes, yt-dlp resolver + bump paths

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plan): specify X Article, image summaries, summary search

* fix: summarize X Article text

* feat: search completed summary tags

* fix: preserve X Article block order

* feat: model opt-in summary images

* feat: attach opted-in images to summaries

* feat: accept image summary requests

* feat: add summary image consent control

* docs: explain X Article, image summaries and summary search

* chore(static): rebuild bundle for summary image consent

* fix: show civilized unsupported summary errors

* chore(static): rebuild bundle for civilized summary errors

* fix: keep text previews and summary state scoped to entry

* fix: show forced yt-dlp candidate in status

* fix: infer trusted MIME for tweet images

* fix: preserve text capture bytes and hide synthetic URL

* fix: recover and preserve summary attempts

* fix: bound codex fallback and record resolved model

* fix: protect public summary diagnostics

* fix: guard summary callbacks during entry render

* fix: scope summary callbacks to selected entry

* test: cover terminal newline in text capture

* fix: preserve missing summary entry status

* test: cover tweet image summary selection

* docs: record summary lifecycle and review hardening

* chore(static): rebuild bundle for Sol review fixes

* fix: retain completed summary during regeneration display

* fix: hide superseded summary attempts

* chore(static): rebuild bundle after regeneration display fix

* fix: allow full-size text capture requests

* fix: allow escaped full-size text captures

* fix: preserve text draft whitespace in capture UI

* chore(static): rebuild bundle for text whitespace fix
2026-08-24 18:30:01 +02:00
e1ee05bd41
feat(collections): public collections, per-collection auth, UX improvements (#34)
* feat(core): add requires_auth to collections; include name in entry-collection memberships

- Add `requires_auth INTEGER NOT NULL DEFAULT 1` column to the
  collections DDL and as an idempotent ALTER TABLE migration in
  initialize_schema (archive DB), not initialize_auth_schema.
- CollectionRecord and CollectionSummary gain `requires_auth: bool`.
- create_collection() and update_collection() accept the new field.
- get_entry_collection_memberships() now returns collection name as the
  third tuple element; EntryCollectionMembership gains a `name` field
  so the sidebar can show human-readable names instead of raw UIDs.

* feat(server): conditional auth for public collections; add requires_auth + original_url to API

- CreateCollectionBody gains requires_auth (default true).
- PatchCollectionBody gains requires_auth: Option<bool>.
- get_collection_handler: load record first, then skip auth.require_auth()
  when record.requires_auth == false so public collections are accessible
  to unauthenticated callers; caller_bits falls back to ROLE_GUEST (1)
  so only visibility_bits=3 entries are returned to guests.
- Collection JSON response includes requires_auth and each entry now
  includes original_url for use by the public collection page.
- list_collections_handler keeps require_auth (management UI).

* feat(frontend): public collection page at /c/:archiveId/:collUid

- Detect PUBLIC_COLL_ROUTE at module load time (like PREVIEW_ROUTE) and
  return <PublicCollectionPage> before any auth checks so unauthenticated
  users can view public collections without hitting the login gate.
- PublicCollectionPage fetches via getCollection() and renders the
  server-filtered entry list (no client-side bitmask filtering - the
  server already applies caller_bits=GUEST for unauthenticated requests).
  Entry titles link to original_url when present; fall back to plain
  text when original_url is null.
- api.js createCollection() gains requiresAuth param (default true),
  sent as requires_auth in the request body.
- Storybook story covers WithEntries, Empty, and LoadError states.

* feat(frontend): collections view improvements

- addVis in the 'Add entry' form now syncs to the selected collection's
  default_visibility_bits via useEffect on collDetail, so the default
  matches the collection's configured entry visibility.
- Rename 'Default visibility' label to 'Entries\' default visibility'
  in both the detail pane and the create form to distinguish it from
  the new collection-level access setting.
- Add 'Require authentication to view' checkbox in the detail pane
  backed by a PATCH to requires_auth; reads collDetail?.requires_auth
  with fallback to the list-level selected record.
- Create form gains a matching requires_auth checkbox (default: true),
  passed as 5th arg to createCollection().

* feat(frontend): context rail collection improvements

- Show collection name (c.name) instead of raw UID in the sidebar
  Collections section; names now come from the updated
  EntryCollectionMembership API response.
- Fix horizontal overflow on long collection names: coll-name gains
  overflow:hidden + text-overflow:ellipsis + white-space:nowrap +
  min-width:0; coll-row gets overflow:hidden.
- Single-entry 'Add to collection' UI: dropdown + button inside the
  Collections rail section lets users add the current entry to any
  non-default collection without multi-selecting. After add, the
  membership list refreshes automatically.
- Collections section now shows even when entryCollections is empty,
  as long as non-default collections exist (so the add form is
  accessible for un-membered entries).
- Bulk 'Add to collection' now uses the target collection's
  default_visibility_bits instead of hardcoded 2 (Users only).
- Both bulk and single-entry dropdowns filter out slug='_default_'
  to match the backend rejection in add_entry_to_collection_handler.
- Collections list is now fetched on archiveId change (not just on
  bulk mode entry) so it is available for single-entry mode too.

* build(frontend): update static assets

* feat(frontend): public collection link UX + app-styled public page

CollectionsView:
- When a collection has requires_auth=false, show a read-only URL input
  and Copy button below the auth checkbox so the public link is
  immediately discoverable. The input auto-selects on focus so manual
  copy always works. Copy button tries navigator.clipboard.writeText
  first; falls back to execCommand('copy') for HTTP deployments where
  the Clipboard API is unavailable in non-secure contexts.

PublicCollectionPage:
- Rewritten to use the app's CSS classes and variables instead of
  bare inline styles, so it visually matches the main archive UI.
  Dark topbar (.pub-coll-topbar) with brand + collection name, paper
  background body, entry list via .coll-entries-list / .coll-entry-row /
  .coll-entry-info / .coll-entry-kind — the same classes used in the
  authenticated Collections view.

styles.css:
- .coll-public-link-row / -wrap / -input / .coll-copy-btn for the
  new link field in CollectionsView detail pane.
- .pub-coll-* classes for the public page layout and typography.

* feat(core): add get_collection_by_slug; scope search to active collection

- get_collection_by_slug(): new function mirroring get_collection_by_uid
  but matching on slug, used to resolve the _default_ collection when no
  ?collection param is supplied.
- SearchEntriesQuery gains collection_id: Option<i64>. When set, the
  search SQL adds an EXISTS subquery that checks collection_entries cef
  for both membership (cef.collection_id = ?) and visibility bits in
  that specific collection — preventing cross-collection visibility
  leaks where an entry is public in one collection but private in the
  current one. Without collection_id the original cross-collection
  visibility fallback is kept.

* feat(server): collection-scoped entries/search with uniform auth gate

All entry listing and search now route through the active collection:

list_entries (?collection=<uid>|main|<omitted>):
- Resolves the target collection; omitted or 'main' resolves to _default_.
- Checks requires_auth on that collection; gates auth conditionally.
- Returns list_entries_for_collection() — same EntrySummary shape.

search_entries_handler:
- Same collection resolution + conditional auth as list_entries.
- Sets search_query.collection_id so SQL scopes membership + visibility
  to the specific collection, not cross-collection fallback.

list_collections_handler:
- Dropped require_auth() — collection summaries (name/slug/uid/
  requires_auth/default_visibility_bits) are public metadata needed for
  the guest collection-switcher dropdown.

Tests:
- list_collections_requires_auth → list_collections_is_public (200).
- list_entries_requires_auth and search coverage still pass.

* feat(frontend): integrate collection switching into main Archive view

Replaces the standalone /c/:archiveId/:collUid public page with a
unified main-view approach where all collection logic lives at /.

URL param:
- ?collection=<uid> selects a collection; omitted or 'main' = default.
- 'main' is normalized to null in parseLocation() so the dropdown shows
  'All entries' and the URL stays clean.

Collection switcher (Topbar):
- Dropdown always visible (guests need it to navigate public collections).
- Non-default collections only (All entries = no param = _default_).
- Guest selecting an auth-required collection calls onSignInClick().
- handleCollectionChange checks both named and _default_ requires_auth
  before proceeding, redirecting guests to login if needed.

listCollections fetched for all users (guests too) since the endpoint
is now public; used to populate the switcher without auth.

Public-session mode (authenticated state, no currentUser):
- Auth gate: fetchArchives() + fetchEntries() with collection param;
  401 falls through to login, 200 proceeds as guest.
- auth:expired suppressed when !currentUser.
- fetchEntryDetail skipped; ContextRail shows entry summary + sign-in prompt.
- ContextRail selection effect skips tag/collection API calls.
- runs/tags not fetched in guest mode.
- Child row expansion disabled in EntryRow (hasChildren = false).

api.js:
- fetchEntries/searchEntries both thread ?collection=<uid> to server.

Deleted: PublicCollectionPage.jsx, PublicCollectionPage.stories.jsx,
copy-link UI from CollectionsView, pub-coll-*/copy-link CSS.

* build(frontend): update static assets

* feat(core): add is_entry_publicly_accessible; checks entry+parent vs public collections

* feat(server): allow guests to fetch detail/children/artifacts for public entries

* feat(frontend): guest collection dropdown filtering; public entry detail without auth wall

* build(frontend): update static assets

* test(server): public entry detail/artifact/children contract for guests

* feat(server): filter auth-required collections from guest list_collections response
2026-07-24 20:16:17 +02:00
a4de506495
frontend: Esc deselects entry; group font artifacts in rail 2026-07-18 21:38:22 +02:00
cd463d2810
feat: multi-select entries with bulk delete, tag, and collection actions (#28)
* feat: multi-select entry rows (shift/ctrl+click, mobile checkbox)
* feat: bulk-action panel for multi-selected entries
2026-07-15 16:50:37 +02:00
d610d37793
feat: entry previews (#24)
* feat: entry previews (video, tweet, article, iframe, image, audio bar)

- Add PreviewPanel dispatch hub routing by entity_kind + artifact extension
- VideoPreview: HTML5 <video> for YouTube/Instagram/TikTok/Reddit/X posts
- TweetPreview: tweet card, thread, and X article renderer (ported from x-article-renderer)
- IframePreview: sandboxed iframe for SingleFile web pages and PDFs
- ImagePreview: image viewer with click-to-open-fullsize
- AudioBar: persistent fixed-bottom player (Spotify-style) that survives entry navigation
- Lift entryDetail to App.jsx, shared between PreviewPanel and ContextRail
- 3-column layout (workspace | 300px preview | 340px rail) when preview active
- Stale-guard fixes: seq incremented before early returns in all async effects
- handleRearchive: capture startSeq/entryUid at call time, guard every async resume
- TweetPreview: reset loading/error/tweets before early-return branches

* feat: entry previews — tweet/thread/article/video/audio/image/iframe/pdf

- PreviewModal: modal overlay with new-tab link (↗) and keyboard close
- PreviewPanel: routes by entity_kind + primary_media extension to the
  correct viewer (tweet/video/audio/pdf/html/image/fallback)
- TweetPreview: full X-style tweet, thread, and article renderer with
  local artifact map for archived media (CDN fallback)
- AudioBar: persistent fixed bottom player, triggered via ContextRail
  Play button; body.has-audio-bar pads content above it
- VideoPreview, IframePreview, ImagePreview: inline viewers
- PreviewPage: standalone /preview/:archiveId/:entryUid route
- ContextRail: Play/Preview buttons; isAudio/isPreviewable detection
- App.jsx: preview modal state, currentAudio state, preview route guard,
  has-audio-bar body class effect
- routes.rs: CSP updated (media-src self blob https; frame-ancestors self;
  Google Fonts + external images/scripts whitelisted)
- styles.css: preview modal, tweet-wrap scroll (min-height:0), audio bar
  body padding, newtab button, preview panel flex layout

* style: tighten article/tweet preview spacing

- aMeta padding: 14px → 10px
- article title marginBottom: 10px → 8px
- aAuthorRow marginBottom: 10px → 8px
- bH1 top margin: 20px → 16px
- bH2 top margin: 18px → 14px
- bHr margin: 20px → 14px
- .preview-tweet-wrap padding: 20px → 12px (already committed)

More content visible above the fold in both modal and standalone views.

* feat: tweet/article preview quality pass

HTML entities: decode &gt; &lt; &amp; etc. on sliced segments only
(entity offsets index the stored string; decoding before slicing shifts them)

Image lightbox: click any tweet/thread/article image to open full-screen
viewer; cmd+click follows <a> to open in new tab; arrow-key + ‹ › nav;
Escape closes; 1/N counter; ↗ open-in-new-tab link

Multi-image grid: 2 photos → side-by-side (180px rows); 3 → left spans
both rows; 4 → 2×2 (140px rows); single image unchanged

Empty media grid ghost: build photos/videoItems arrays first, render
.mediaGrid div only when at least one item resolved (advisory: never gate
on raw media.length when map items can all return null)

QT indicator: ↻ QT badge on tweet.is_quote_status === true entries

Modal shrinks for short content: height: 88vh → max-height: 88vh;
.preview-modal-body gets max-height: calc(88vh - 52px) so long threads
still scroll (advisory: don't rely on flex:1 once parent has no fixed height)

Video scrollbar leak: .preview-modal-body overflow: auto → hidden; each
child (tweet-wrap, video-wrap, iframe) manages its own scroll surface

Styled thin scrollbar on .preview-tweet-wrap (matches workspace rail)

ArticleRenderer: cover image and body images are lightbox-clickable;
opts thread through renderBlocksJSX → renderBlockJSX → renderAtomicJSX

* fix: move artifact fetch to api.js; stop Escape propagation from lightbox

- Export fetchEntryArtifacts(archiveId, entryUid, indices) from api.js
  using Promise.all + getJson (follows project convention: all /api calls
  go through api.js, never inline fetch in components)
- TweetPreview: import fetchEntryArtifacts, replace inline Promise.all
- MediaLightbox keydown handler: stopPropagation + preventDefault for
  Escape/ArrowLeft/ArrowRight so the parent PreviewModal window listener
  does not also fire and close the modal behind the lightbox

* fix: iframe/page preview height chain and toolbar UX

Problem: changing .preview-modal from height to max-height broke iframe
previews - <iframe style='flex:1'> needs a concrete ancestor height, which
max-height alone doesn't supply when content is shorter than the cap.

Fix - CSS:
  .preview-modal--full { height: 88vh } applied to non-tweet modals
  .preview-modal--full .preview-modal-body { max-height: none }
  .preview-iframe-toolbar span: remove text-transform/letter-spacing
    (was uppercasing the URL/title in shouty caps)

Fix - PreviewModal: className adds --full when entity_kind is not
  tweet/tweet_thread; tweet previews keep shrink-to-fit behavior.

Fix - PreviewPanel: pass title + original_url from summary to IframePreview
  for both HTML and PDF; wrappers use flex:1/minHeight:0 not height:100%.

Fix - IframePreview:
  - Accept title + originalUrl props; show originalUrl in toolbar (falls
    back to artifact src only when original_url absent); show title above
    URL when available
  - flex:1 + minHeight:0 instead of height:100% on the wrap div
  - Single unified layout for page + pdf (both just show the iframe)

* feat: expand t.co links; linkify bare URLs in tweet and article text

Frontend:
- resolveEntityBounds: try multiple candidate strings in order (u.url
  first, since that's the t.co short URL that appears in full_text)
- normalizeUrlAnn: multi-candidate search; href = expanded > url,
  display = display_url > expanded > url
- linkifyText(): regex linkifier for entity-less bare URLs; trims
  trailing punctuation [.,;:!?)] before linking; used in both
  renderTweetTextJSX and renderInlineJSX including their early-return
  paths (anns.length === 0) that previously bypassed linkification
- renderInlineJSX: fix mention href mention.name → screen_name;
  replace t.co segment text with url.display when entity covers it

Scraper (vendor/twitter/scrape_user_tweet_contents.py):
- extract_tweet_data: when note_tweet text is used, pull urls/mentions/
  hashtags/symbols from note_result.entity_set (correct indices for the
  note text); keep media from legacy.entities (no note media downloads)

* feat: server-side t.co resolver + frontend augmentation

Server (routes.rs):
  POST /api/util/resolve-tco — unauthenticated, accepts JSON array of
  https://t.co/<alphanumeric> URLs only (strict regex validation, no SSRF
  via input), capped at 50 per batch, 3 s timeout, redirect(Policy::none)
  so the server only ever touches t.co itself. HEAD first, GET fallback if
  HEAD returns no Location. Location sanitized to http/https only —
  javascript:/data:/etc. fall back to the original t.co.

api.js:
  resolveTcoUrls(urls) — project-convention wrapper for the new endpoint.
  Returns {} on failure (callers degrade gracefully to bare t.co links).

TweetPreview.jsx:
  After tweet data loads, per-tweet range-based coverage detection:
  builds covered [start,end) from existing entity fromIndex/toIndex or
  indices fallback, then finds regex matches whose span is NOT covered.
  Resolves unique uncovered t.co URLs via resolveTcoUrls(), synthesises
  one entity per occurrence (with exact fromIndex/toIndex so normalizeUrlAnn
  gets correct bounds even for duplicate t.co URLs in the same tweet).
  Augments entities.urls before setTweets() so all rendering paths
  see expanded URLs.

* feat: suppress rendered media attachment URLs from tweet text

renderTweetTextJSX now accepts skipSpans=[] as third param.
Skip-span boundaries are added to the pts split set so a trailing
media t.co inside a plain segment still gets isolated and suppressed—
not re-linked by linkifyText. Early return only when both anns and
skipSpans are empty.

TweetCard computes mediaSkipSpans after building photos/videoItems:
for each rawMedia item whose src resolved (photo src match; any video),
resolveEntityBounds(m, ft, m.url) gives the precise [s,e] span using
indices/fromIndex first, indexOf fallback—then the span is passed to
renderTweetTextJSX so the t.co attachment URL is silently dropped.
2026-07-12 16:23:04 +02:00
2779afee2d
feat(tweets): add re-archive button and fix thread-tweet orphan cleanup (#23)
Orphan cleanup bug: archiving x🧵A downloaded D/C/B/A JSONs and
media, but only registered artifacts for A. D/C/B files had no
entry_artifacts rows and were deleted as orphans.

Fix (staged scraper output, precise touched set):
- tweets::archive() stages all scraper output in temp/{ts}/tweet_stage/,
  validates, then renames JSONs to raw_tweets/. Return type changed from
  Result<bool> to Result<Vec<String>> (store-relative relpaths of every
  produced tweet JSON, i.e. the exact touched set).
- tweets::rearchive() (new): same staged approach but always runs the
  scraper. On scraper failure (tweet deleted/private), errors before
  touching raw_tweets/ so existing data is preserved.
- register_tweet_artifacts() (new private helper in capture.rs): registers
  every JSON in the touched set as a raw_tweet_json artifact, parses each
  for media blobs, registers those too. JSON read failure is a hard error
  with context, not a silent skip.
- record_tweet_entry() now accepts tweet_json_relpaths: &[String] and
  delegates artifact registration to register_tweet_artifacts().
- perform_capture() passes the returned vec from tweets::archive().

Re-archive feature:
- capture::perform_rearchive(): looks up entry by uid, validates
  tweet/tweet_thread, runs tweets::rearchive(), atomically swaps
  entry_artifacts in a DB transaction. archived_at, title, tags,
  collections are untouched.
- database: add get_entry_for_rearchive() and delete_entry_artifacts().
- POST /api/archives/:id/entries/:uid/rearchive: requires ROLE_USER,
  creates capture job, returns 202 + job_uid, runs perform_rearchive in
  spawn_blocking.
- Frontend: re-archive button in ContextRail for tweet/tweet_thread
  entries; polls job at 500ms; refreshes entry detail on success; shows
  error text on failure. Poll interval cleared before early-return on
  entry deselect to prevent stale updates.
2026-07-11 16:14:58 +02:00
ed1f883ff1
feat: implement entry deletion
- database: cascade_cached_bytes_after_subtree_delete — one-pass set-aware
  SQL that excludes the entire subtree simultaneously, avoiding sibling-blob
  cross-counting bug
- database: delete_entry — collects subtree IDs, runs set-aware cascade,
  NULLs archive_run_items.produced_entry_id (FK blocker), deletes children
  then root; ON DELETE CASCADE handles artifacts/tags/collections
- server: DELETE /api/archives/:archive_id/entries/:entry_uid route,
  wrapped in a transaction for atomicity
- frontend: deleteEntry API call, handleEntryDeleted callback in App.jsx
  (optimistic list removal + selection clear), Delete entry button in
  ContextRail with confirm guard
- tests: 4 database tests (unknown uid, subtree removal, run_item nulling,
  cached_bytes recalculation) + 3 route tests (204+gone, 404, 401)
2026-07-04 13:46:19 +02:00
55f85134df
feat: tag delete, rename, and per-user humanize-tags display setting (#16)
* feat: add tag delete and rename (backend + frontend)

- DELETE /api/archives/:id/tags/:tag_uid — deletes tag subtree via
  recursive CTE; entry_tag_assignments cascade automatically
- PATCH  /api/archives/:id/tags/:tag_uid { name } — renames a tag
  segment (case-preserving slug), cascades full_path to all descendants
  in one transaction using hierarchy CTE (not LIKE), returns updated Tag
- database: rename_tag + delete_tag with 7 unit tests covering subtree
  cascade, collision detection, descendant path rewrite, slug stripping
- TagsView: inline rename (pencil icon / double-click → input), × delete
  with confirmation dialog (warns about child tags)
- App.jsx: handleTagRenamed rewrites tagFilter for exact + descendant
  paths; handleTagDeleted clears filter for deleted subtree
- api.js: renameTag (PATCH, returns Tag) + deleteTag (DELETE)
- ContextRail: tag pills show displayPath(full_path) (humanized) with
  raw full_path in title tooltip; humanize-tags setting coming next

* feat: per-user humanize-tags display setting

- auth DB: idempotent migration adds users.humanize_slugs INTEGER DEFAULT 0
- GET /api/auth/me: returns humanize_slugs bool
- PATCH /api/auth/me: accepts { humanize_slugs: bool }, persisted via
  database::update_user_humanize_slugs
- displayPath() helper moved to utils.js (was local to ContextRail)
- TagsView: node label shows tag.name vs tag.slug based on humanizeTags
- ContextRail: pill label applies displayPath conditionally
- App.jsx: derives humanizeTags from currentUser.humanize_slugs, passes
  to TagsView + ContextRail; filter badge label humanized when on
- Settings > Profile: Display Preferences section with checkbox toggle,
  updates backend and currentUser immediately via setCurrentUser
- api.js: patchMe(patch) for generic PATCH /api/auth/me
- Tests: auth_me default=false, PATCH persists=true (2 route tests)
2026-07-03 14:26:45 +02:00
2502de45b6
feat: inline entry title renaming (#14)
* feat: inline entry title renaming

- database.rs: add update_entry_title(conn, entry_uid, title) -> Result<bool>
  targets archived_entries table; returns false for unknown uid
- routes.rs: PATCH /api/archives/:archive_id/entries/:entry_uid (ROLE_USER)
  maps false -> 404; 3 tests (401 unauthed, 204+persists, 404 unknown uid)
- api.js: updateEntryTitle(archiveId, entryUid, title)
- ContextRail.jsx: click title h2 -> inline input; Enter/blur commits,
  Escape cancels (cancel ref prevents double-save on blur after Escape);
  pencil icon fades in on hover as edit affordance
- App.jsx: handleEntryTitleChange mutates entries + selectedEntry in-place
- styles.css: rail-title--editable cursor/hover, pencil icon show/hide,
  rail-title-input underline style
- rebuild frontend bundle

* chore: remove + prefix from Capture button label
2026-07-03 13:01:34 +02:00
4311e85f95
ui: comprehensive UI polish pass
- LoginPage/SetupPage: centered card layout with display font, styled fields and submit button
- Topbar: fix duplicate Settings nav item; add styled user-menu with username + logout button
- RunsView: format ISO timestamps to readable dates, add colored status badges (completed/failed/running)
- AdminView: view-tabs system, styled admin-table, admin-input, status-badge (active/disabled), btn-primary
- SettingsView: replace all inline styles with form-section/form-field/field-input/btn-primary/btn-danger
- CollectionsView: restructure create form with proper field labels and btn-primary
- TagsView: add Tags section heading with separator
- ContextRail: show visibility as human-readable label not raw number; fix assign-error class
- styles.css: add auth-loading, view-tabs, form utilities, btn variants, status badges,
  run-status pills, token-banner/row, checkbox-row, coll-create-form, tag-tree-header CSS
2026-06-28 22:14:08 +02:00
3ccfcce87b
feat(ui): add Storybook design system refinements 2026-06-28 21:01:50 +02:00
763cb8e17f
feat(collections): frontend — CollectionsView, nav, visibility in ContextRail 2026-06-26 17:04:49 +02:00
4458f17b13
feat(ui): rewrite frontend in React with Vite
- Scaffold Vite+React project in frontend/ (bun, react 18, @vitejs/plugin-react)
- vite.config.js outputs to crates/archivr-server/static/ directly
- src/utils.js: formatBytes, valueText, formatTimestamp, SOURCE_ICONS, sourceIconSvg
- src/api.js: typed fetch wrappers for all API endpoints
- App.jsx: full state management, archive switching, debounced search,
  tag filter, view routing, capture dialog orchestration
- components/Topbar.jsx: archive switcher, nav, capture button
- components/CaptureDialog.jsx: native <dialog> ref with showModal/close,
  Escape key support, locator validation
- components/EntriesView.jsx + EntryRow.jsx: flex table with source icons,
  type pills, keyboard selection
- components/ContextRail.jsx: parallel detail+tags fetch, stale-race guard
  via selectSeqRef, inline tag assign/remove
- components/RunsView.jsx: runs kept as <table> (matches original)
- components/AdminView.jsx: mounted archives list
- components/TagsView.jsx: recursive tag tree with active state
- Remove old app.js; styles.css moved to frontend/src/ (Vite bundles it)
2026-06-24 12:24:34 +02:00