From 79ac44834ec6ddbd1bdd743e0a1596d0ca6494a5 Mon Sep 17 00:00:00 2001 From: TheGeneralist <180094941+thegeneralist01@users.noreply.github.com> Date: Mon, 24 Aug 2026 18:30:01 +0200 Subject: [PATCH] feat: capture, summaries, search, and yt-dlp reliability (#38) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * ui: show spinner for pending captures * feat(core): add text capture path with title + Markdown/plain body - Add downloader/text.rs module with save() function that stages and hashes text content - Support text/markdown and text/plain MIME types with .md and .txt extensions - Add perform_text_capture() function for capturing user-supplied text - Validates title (non-empty, max 500 chars) and body (non-empty, max 2 MiB) - Creates blob records and entries with source_kind='text', entity_kind='document' - Includes comprehensive unit tests for markdown, plain text, and validation Co-Authored-By: Claude Haiku 4.5 * feat(server): add POST /api/archives/:archive_id/captures/text - Add CaptureTextBody struct for title, body, and optional MIME type - Implement capture_text_handler with validation for empty fields and MIME type - Route text submissions to perform_text_capture() in background - Reuse existing capture job tracking and polling infrastructure - Default MIME type to text/markdown when not specified - Include route tests covering happy path, validation, auth, and error cases Co-Authored-By: Claude Haiku 4.5 * feat(frontend): add text-capture form to CaptureDialog - Add submitTextCapture API client function with same error handling as submitCapture - Create makeTextItem() factory for text capture state - Implement CaptureTextRow component with title, body textarea, and MIME selector - Add 'Add text' button in capture dialog toolbar - Update handleArchive to filter and route text submissions - Modify submitBgJob to detect and submit text items via submitTextCapture - Skip probe and conflict checks for text items - Reuse job tracking and batch settlement for text captures Co-Authored-By: Claude Haiku 4.5 * feat(core): add entry_summaries schema + summarizer trait/providers Per-entry LLM summaries as a regenerable child record, not a column on archived_entries and not an on-disk artifact: an entry may carry several summaries (one per provider/model/prompt version), any of which can be discarded and recomputed. Generation is manual-only — nothing in capture.rs calls into this module. - database.rs: entry_summaries table + index, EntrySummaryRecord, and upsert/update/find/latest helpers mirroring the capture_jobs style. provider_model is stored as '' rather than NULL because SQLite treats NULLs as distinct inside a UNIQUE index, which would stop the CLI providers (no model) from ever deduping on the cache key. - summarizer.rs: SummaryProvider trait with four implementations — Anthropic Messages API, OpenAI-compatible chat completions, `claude -p` and `codex exec -`. Configuration comes from env vars only (never TOML), matching how yt-dlp / single-file / tweet-scraper are resolved, which also keeps API keys out of anything the archive persists. - archive.rs: EntryDetail gains latest_summary, populated by one extra LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row rather than duplicating it. Implementation notes: - No tokio in core. CLI timeouts are enforced structurally: stdout is drained on its own thread and handed back over a channel so the calling thread can recv_timeout and kill an overrunning child; stdin is written on a third thread so a 48 KB prompt cannot deadlock against a child waiting for us to read. - HTML is reduced with regex rather than a parser: html5ever is not in the tree, and a model tolerates imperfect whitespace. Paired tags are spelled out per tag because Rust's regex engine has no backreferences by design. - reqwest is declared with only the `blocking` feature here, so bodies are serialized via .body(value.to_string()) instead of widening the workspace dependency for .json(). - input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one hashing primitive; the content is truncated to 48 KB *before* hashing so the cache key describes exactly the bytes the model saw. Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for one JSON shape is a poor trade, so the two halves that can actually break are tested directly — request-body builders and response parsers — leaving only reqwest's own transport uncovered. Plus schema idempotency, cache-key dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML and tweet extraction, output normalization, and the CLI runner's stdin round-trip, timeout kill, and nonzero-exit paths. Co-Authored-By: Claude Opus 5 * feat(server): add GET/POST /api/archives/:id/entries/:uid/summary GET is read-only and gated exactly like entry detail, so a guest can read a summary only for an entry whose content they could already read. POST requires ROLE_USER, matching capture / tags / patch / rearchive; no auth roles change. Both the provider config and the content extraction resolve on the request thread, before spawn_blocking. That is what lets a missing env var come back as a synchronous 400 naming the exact variable, and an unsummarizable artifact (video, audio) as a 400 saying so, rather than becoming a background job the caller must poll only to learn about a config typo. The pending row is claimed before spawning so the 202 can name a summary_uid the client can poll immediately. summarize_entry owns the pending → running → completed/failed transitions for that same row — the cache key is identical, so both upserts resolve to one row — leaving the handler to catch only the case where it fails before recording anything. When !force and an identical cache key already completed, the existing row comes back as a 200 with no new work. Co-Authored-By: Claude Opus 5 * feat(frontend): render Summary rail section + provider selector New "Summary" rail section between the URL/Preview controls and .meta-list. A completed summary renders as bold tl;dr, body paragraph, tag chips, and a provider · model footer; missing or failed shows Generate; pending/running shows an inline spinner and polls GET every 1500 ms until terminal. - api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps ApiError's { "error": ... } body so the missing-env-var message reaches the user verbatim rather than as a bare status code. - ContextRail.jsx: state seeds from detail.latest_summary so the section renders immediately on selection. Polling is anchored on the summary status rather than started inside the click handler, so a job still running when the user navigates away and back is picked up again. A transient poll failure is swallowed — the next tick retries, and a real failure arrives as status === 'failed'. - Regenerate passes force:true only when a completed summary is already shown; otherwise the request can take the server's 200 cache-hit path. - Provider choice persists in sessionStorage under archivr:summary:provider, with try/catch around both accessors for private-mode browsers. - Public sessions never see the selector or the Generate button, and the section renders at all only when a completed summary made it through the server's visibility gate. - styles.css: .rail-summary-* only; spacing and the action button reuse .rail-section and .rail-rearchive-btn. The spinner honours prefers-reduced-motion — the text alone conveys the state. - AGENTS.md: document the summary env vars alongside the existing external-tool convention. Smoke-tested end to end against a scratch archive with a seeded markdown entry: claude_cli produced a real summary (pending → running → completed in ~11s); a local mock server exercised the openai_compatible transport and confirmed the Bearer header, model, and system/user role split on the wire; unconfigured providers return 400 naming the exact variable; a video entry returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without adding a row. Co-Authored-By: Claude Opus 5 * fix(core): codex_cli — auto-discover binary + use --output-last-message Two related fixes for the codex_cli summary provider: 1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but without it the code resolved to bare `codex` and relied on PATH. The ChatGPT desktop app installs codex at `/Applications/ChatGPT.app/Contents/Resources/codex` and does not put it on PATH, so users who only have the desktop app saw 'No such file or directory' with no hint. `resolve_cli` now walks env override → a small set of well-known absolute paths → HOME /.local/bin/ → bare fallback. Same treatment applied to claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/ claude, HOME/.local/bin/claude). 2. Clean output. `codex exec -` writes a runtime header ("OpenAI Codex vX", session id, sandbox, model), the assistant reply, and a footer ("tokens used", replay of the reply) to stdout. The JSON extractor took the first '{' from the *user prompt echo* and the last '}' from the trailing replay, producing invalid text that fell through to the "raw text under summary" fallback path. Now uses `--output-last-message ` and reads only the final assistant message. Fallback (positional prompt) uses the same flag. Tempfile is cleaned up on all paths, incl. spawn failure. * fix(frontend): give text-capture row real CSS The text row shipped with semantic classnames (`capture-text-inputs`, `capture-text-title`, `capture-text-body`, `capture-text-mime`, `capture-text-icon`) but no CSS rules. Falling through to the parent `.capture-row-main` flex-row (`display: flex; align-items: center`) meant the title, textarea, and mime-select stacked as intrinsic-width boxes centered on the tall body, producing a layout where the body floated to the top-right, the title box appeared BELOW it, and the mime selector rendered as an unstyled OS dropdown. Fix: - `.capture-text-row .capture-row-main` uses `align-items: flex-start` so the leading icon and trailing × pin to the top of the block. - `.capture-text-inputs` is now a full-width column-flex container with proper gaps. - `.capture-text-title` reuses the 44px input height and typography of `.capture-input`; `.capture-text-body` gets a 140px min-height, vertical resize, and matching border/focus treatment. - `.capture-text-mime` is styled as a small chip with a custom caret so it matches `.capture-quality` and stops looking like a raw ``, render a labeled checkbox with exact visible label `Include attached images`. Its help text must say that selected archived images are sent to the chosen provider and that only up to four supported images (5 MiB each, 12 MiB total) can be attached; unsupported or oversized artifacts are skipped. When `summaryProvider === 'claude_cli'`, render the checkbox disabled, force `includeSummaryImages` to false via an effect or provider-change handler, and show `Claude CLI cannot attach local images. Choose an HTTP provider or Codex CLI.` Do not submit a silently dropped image choice. + +- [ ] **Step 5: Add scoped plain-CSS rules.** + + Add `.rail-summary-image-option`, `.rail-summary-image-option__label`, `.rail-summary-image-option__note`, and `.rail-summary-image-option--disabled` under the existing summary rail CSS. Use the project variables (`--muted`, `--line`, `--paper`) and preserve keyboard focus and normal checkbox semantics; do not use a generic row class or inline layout styles. + +- [ ] **Step 6: Run the manual success script and build verification.** + + Run: `bun run build` + + Expected: successful production bundle in `crates/archivr-server/static`. Then manually verify: unchecked generation sends `include_images:false`; checked Anthropic/OpenAI/Codex generation sends `true`; changing to Claude unchecks/disables the control and shows the exact explanation; server errors still appear through existing `summaryError` handling. + +- [ ] **Step 7: Commit source files, not generated static output.** + + ```bash + git add frontend/src/api.js frontend/src/components/ContextRail.jsx frontend/src/styles.css + git commit -m "feat: add summary image consent control" + ``` + +### Task 6: Search latest completed summary JSON without changing API shape + +**Files:** + +- Modify: `crates/archivr-core/src/archive.rs` +- Test: `crates/archivr-core/src/archive.rs` (`#[cfg(test)] mod tests`) + +- [ ] **Step 1: Write failing search tests with real cache rows.** + + Extend `make_test_db_with_entries` or add a focused fixture helper that inserts summary rows through `database::upsert_pending_entry_summary` and `database::update_entry_summary_status`. Add assertions for all of the following: + + ```rust + // A completed JSON string with {"tags":["skincare","dermatology"]} matches skincare. + // An unrelated query yields no result. + // Two completed rows: the newer completed row is searched; the older one is not. + // A completed older row remains matched while a newer row is pending or failed. + // source:, entity:, url:, title:, after:, before:, tag: and collection/visibility scope keep their current behavior. + ``` + + Set distinct `updated_at` values (or insert/transition rows in distinct timestamp order) so "latest completed" is unambiguous. The tag assertion must match `summary_text` itself, not `entry_tag_assignments`. + +- [ ] **Step 2: Run the focused tests and observe failure.** + + Run: `cargo test -p archivr-core search_.*summary` + + Expected: FAIL because free text only checks entry and source identity fields. + +- [ ] **Step 3: Add a parameter-bound latest-completed summary predicate.** + + In the existing unqualified `query.q` block in `search_entries`, preserve every present `LOWER(...) LIKE ?{n}` condition and add this clause using the same single bound `term`: + + ```sql + OR LOWER(COALESCE(( + SELECT s.summary_text + FROM entry_summaries s + WHERE s.entry_id = e.id + AND s.status = 'completed' + AND s.summary_text IS NOT NULL + ORDER BY s.completed_at DESC, s.updated_at DESC, s.id DESC + LIMIT 1 + ), '')) LIKE ?N + ``` + + Use the existing numbered parameter construction (`?{n}`) and push `term` once, so user text is never concatenated into SQL. `completed_at` ordering means only completed rows participate, and a later pending/failed row cannot displace an older completed row. Do not add an archive method, schema column, migration, route parameter, or frontend response field. + +- [ ] **Step 4: Run core search tests.** + + Run: `cargo test -p archivr-core search_` + + Expected: PASS, including old prefix-filter behavior and JSON tag substring matches. + +- [ ] **Step 5: Verify the existing API transport needs no change.** + + Inspect `crates/archivr-server/src/routes.rs::search_entries_handler`, `frontend/src/api.js::searchEntries`, and `frontend/src/App.jsx` search call sites. Confirm the server still calls `archive::search_entries` with the same `SearchEntriesQuery` and returns `Vec`; record no source edit for these files unless the inspection reveals a type break. Add no client-side filtering. + +- [ ] **Step 6: Commit the search change.** + + ```bash + git add crates/archivr-core/src/archive.rs + git commit -m "feat: search completed summary tags" + ``` + +### Task 7: Document, bundle, and verify the completed feature set + +**Files:** + +- Modify: `docs/README.md` +- Modify: `AGENTS.md` +- Modify: `ARCHIVR-MENTAL-MODEL.md` +- Generated (do not hand-edit): `crates/archivr-server/static/` + +- [ ] **Step 1: Update user documentation.** + + In `docs/README.md`, document that summaries are manual, text-only by default, and the explicit `Include attached images` option sends at most four eligible local images to the selected provider. State the allowed formats and byte limits, identify Anthropic/OpenAI-compatible/Codex support, state Claude CLI cannot attach local images, and state free-text search includes the latest completed summary text and its generated tags. + +- [ ] **Step 2: Update contributor constraints.** + + In `AGENTS.md`, record the `SummaryBuildOptions`/digest rule, the image candidate role and limits, the provider capability matrix, and the latest-completed-only search semantic. Preserve the rule that core remains synchronous and that generated static files are not hand-edited. + +- [ ] **Step 3: Update the architectural data-flow documentation.** + + In `ARCHIVR-MENTAL-MODEL.md`, extend the LLM Summary section to show explicit UI consent flowing into image selection, cache hashing, provider transport, and the existing row lifecycle. Add that entry search reads only the latest completed `summary_text`, retaining a prior completed result while newer work is pending or failed. + +- [ ] **Step 4: Build and test the final implementation.** + + Run: + + ```bash + cargo test + bun --cwd frontend run build + cargo build + ``` + + Expected: all Rust tests pass, frontend build succeeds, and the generated static bundle contains the checkbox UI. Do not hand-edit generated files; include them in a commit only if this repository currently tracks frontend bundle changes after `bun run build`. + +- [ ] **Step 5: Run the end-to-end manual smoke test.** + + Start the server with a test archive and verify: an X Article whose normal tweet text is only a t.co URL summarizes article body text; a text-only generation remains unchanged; a checked vision-capable request attaches only bounded eligible images; Claude has a disabled explanatory option and a direct API request is rejected; a `skincare` search finds completed summary JSON tags; pending/failed rows do not hide a previous completed match. + +- [ ] **Step 6: Commit documentation and any tracked generated bundle.** + + ```bash + git add docs/README.md AGENTS.md ARCHIVR-MENTAL-MODEL.md crates/archivr-server/static + git commit -m "docs: explain image summaries and summary search" + ``` + +### Task 8: Final implementation review before integration + +**Files:** + +- Review: `crates/archivr-core/src/summarizer.rs` +- Review: `crates/archivr-core/src/archive.rs` +- Review: `crates/archivr-server/src/routes.rs` +- Review: `frontend/src/api.js` +- Review: `frontend/src/components/ContextRail.jsx` +- Review: `frontend/src/styles.css` +- Review: `docs/README.md`, `AGENTS.md`, `ARCHIVR-MENTAL-MODEL.md` + +- [ ] **Step 1: Perform the approved-design coverage review.** + + Verify A is covered by Tasks 1 and 2 (plain text, blocks, preview/summary, normal tweet fallback, and every thread JSON artifact); B by Tasks 2–5 (explicit default-off consent, selection policy/caps, input hash, all four provider outcomes, POST flag, and compatible GET); and C by Task 6 (latest completed summary JSON/tags, pending/failed semantics, prefix-filter preservation, server-side architecture). + +- [ ] **Step 2: Scan the plan and implementation for unfinished markers and type drift.** + + Run: `rg -n -i '\\bt[o]do\\b|\\bt[b]d\\b|placehold[e]r|implement[[:space:]]later' docs/superpowers/plans/2026-08-23-x-article-vision-search.md crates/archivr-core/src/summarizer.rs crates/archivr-core/src/archive.rs crates/archivr-server/src/routes.rs frontend/src` + + Expected: no newly introduced unfinished markers in the changed feature code or plan. Confirm every use of `SummaryBuildOptions`, `SummaryImage`, options-aware `build_summary_input`, and options-aware `summarize_entry` matches the Task 2 definitions. + +- [ ] **Step 3: Review commits and working tree.** + + Run: `git log --oneline --decorate -8` and `git status --short` + + Expected: atomic commits cover the reducer, core image model/providers, server API, frontend control, search, and docs; no unintended artifacts or source edits remain. diff --git a/flake.nix b/flake.nix index 95f177b..16384e5 100644 --- a/flake.nix +++ b/flake.nix @@ -92,6 +92,35 @@ cp -r . $out/ ''; }; + # yt-dlp — pinned to a specific GitHub release rather than pulled through + # nixpkgs. Rationale: YouTube frequently rotates player-signature/API + # surfaces, and yt-dlp ships updates on a days-to-weeks cadence; even + # nixos-unstable often lags by months. When the binary is stale, + # captures fail with HTTP 403 on formats the old client can't + # authenticate. Fetching the zipapp directly (a Python zipapp with a + # `#!/usr/bin/env python3` shebang) lets us bump the version + hash in + # one place without waiting on nixpkgs. Wrapped so `python3` and + # `ffmpeg` — the two runtime deps for muxed downloads — are always on + # PATH regardless of the caller's environment. + # + # Bumping: replace `version`, then run `nix hash file ` on the + # new zipapp URL and paste the sri output into `hash`. + ytDlp = pkgs.stdenv.mkDerivation { + pname = "yt-dlp"; + version = "2026.08.19"; + src = pkgs.fetchurl { + url = "https://github.com/yt-dlp/yt-dlp/releases/download/2026.08.19/yt-dlp"; + hash = "sha256-H6ZzPDfqb7Ucma2P54Xnt+XzJGybmAIwMp1Pty7Y1NY="; + }; + dontUnpack = true; + nativeBuildInputs = [ pkgs.makeWrapper ]; + installPhase = '' + mkdir -p $out/bin + install -m 0755 $src $out/bin/yt-dlp + wrapProgram $out/bin/yt-dlp \ + --prefix PATH : ${lib.makeBinPath [ pkgs.python312 pkgs.ffmpeg ]} + ''; + }; version = "0.1.0"; src = pkgs.lib.cleanSource ./.; cargoLock = { @@ -139,7 +168,7 @@ version = "0.1.0"; nativeBuildInputs = [ pkgs.makeWrapper ]; buildInputs = [ - pkgs.yt-dlp + ytDlp pkgs.single-file-cli tweetPython ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ]; @@ -150,7 +179,7 @@ cp ${./vendor/twitter/scrape_user_tweet_contents.py} $out/libexec/archivr/scrape_user_tweet_contents.py chmod +x $out/libexec/archivr/scrape_user_tweet_contents.py makeWrapper $out/libexec/archivr/archivr $out/bin/archivr \ - --set ARCHIVR_YT_DLP ${pkgs.yt-dlp}/bin/yt-dlp \ + --set ARCHIVR_YT_DLP ${ytDlp}/bin/yt-dlp \ --set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \ ${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \ --set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \ @@ -159,7 +188,7 @@ --set ARCHIVR_COOKIE_EXT ${isdcac} \ --prefix PATH : ${ lib.makeBinPath ([ - pkgs.yt-dlp + ytDlp pkgs.single-file-cli tweetPython ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ]) @@ -170,7 +199,7 @@ pname = "archivr-server-wrapped"; inherit version; nativeBuildInputs = [ pkgs.makeWrapper ]; - buildInputs = [ tweetPython pkgs.single-file-cli ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ]; + buildInputs = [ ytDlp tweetPython pkgs.single-file-cli ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ]; phases = [ "installPhase" ]; installPhase = '' mkdir -p $out/bin $out/libexec/archivr-server $out/share/archivr-server/static @@ -180,12 +209,14 @@ cp -r ${./crates/archivr-server/static}/* $out/share/archivr-server/static/ makeWrapper $out/libexec/archivr-server/archivr-server $out/bin/archivr-server \ --set ARCHIVR_STATIC_DIR $out/share/archivr-server/static \ + --set ARCHIVR_YT_DLP ${ytDlp}/bin/yt-dlp \ --set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \ ${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \ --set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \ --set ARCHIVR_TWEET_SCRAPER $out/libexec/archivr-server/scrape_user_tweet_contents.py \ --set ARCHIVR_UBLOCK_EXT ${ublockLite} \ - --set ARCHIVR_COOKIE_EXT ${isdcac} + --set ARCHIVR_COOKIE_EXT ${isdcac} \ + --prefix PATH : ${lib.makeBinPath ([ ytDlp pkgs.single-file-cli tweetPython ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ])} ''; }; archivr-all = pkgs.symlinkJoin { diff --git a/frontend/src/api.js b/frontend/src/api.js index 3bbed1f..2e73e34 100644 --- a/frontend/src/api.js +++ b/frontend/src/api.js @@ -1,5 +1,5 @@ -async function getJson(url) { - const response = await fetch(url); +async function getJson(url, options) { + const response = await fetch(url, options); if (!response.ok) { throw new Error(`${response.status} ${response.statusText}`); } @@ -29,6 +29,50 @@ export async function fetchEntryDetail(archiveId, entryUid) { return getJson(`/api/archives/${archiveId}/entries/${entryUid}`); } +// ── Entry summaries ──────────────────────────────────────────────────────── +// Summaries are generated on demand, never at capture time. GET is safe for +// public sessions (the server applies the same visibility gate as entry detail). + +export async function fetchEntrySummary(archiveId, entryUid, { signal } = {}) { + return getJson(`/api/archives/${archiveId}/entries/${entryUid}/summary`, { signal }); +} + +// Kicks off generation. Resolves to either an existing completed summary (200) +// or a freshly claimed pending row (202) — both carry a summary_uid, so the +// caller polls fetchEntrySummary either way. +// The server returns 400 with the exact missing env var name when a provider is +// unconfigured, so its body is surfaced verbatim rather than replaced. +export async function requestEntrySummary(archiveId, entryUid, { provider, force = false, includeImages = false, signal } = {}) { + const resp = await fetch( + `/api/archives/${archiveId}/entries/${entryUid}/summary`, + { + method: "POST", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify({ provider, force, include_images: includeImages }), + signal, + } + ); + if (!resp.ok) { + // ApiError renders as { "error": "..." }; that message is the useful part + // (e.g. "missing required environment variable: ARCHIVR_ANTHROPIC_API_KEY"), + // so surface it verbatim instead of a generic status string. + const detail = await resp.text(); + let message = detail.trim(); + try { message = JSON.parse(detail).error || message } catch { /* non-JSON body */ } + throw new Error(message || `Summary request failed (${resp.status})`); + } + return resp.json(); +} + +// Text artifacts are served by the same entry-artifact endpoint as previews. +// Keep credentials explicit because this helper is also used by public/private +// archive views, and preserve the previous concise HTTP error contract. +export async function fetchArtifactText(src, { signal } = {}) { + const response = await fetch(src, { credentials: 'same-origin', signal }); + if (!response.ok) throw new Error(`HTTP ${response.status}`); + return response.text(); +} + export async function fetchEntryChildren(archiveId, entryUid) { return getJson(`/api/archives/${archiveId}/entries/${entryUid}/children`); } @@ -174,6 +218,24 @@ export async function submitCapture(archiveId, locator, quality = null, extensio return res.json(); // { job_uid, status: "pending" } } +export async function submitTextCapture(archiveId, {title, body, mime = 'text/markdown'}) { + const payload = { title, body }; + if (mime && mime !== 'text/markdown') payload.mime = mime; + + const res = await fetch(`/api/archives/${archiveId}/captures/text`, { + method: "POST", + headers: { "Content-Type": "application/json" }, + body: JSON.stringify(payload), + }); + if (!res.ok) { + const body = await res.json().catch(() => ({})); + const err = new Error(body.error || `HTTP ${res.status}`); + err.status = res.status; + throw err; + } + return res.json(); // { job_uid, status: "pending" } +} + // Returns { has_video: bool, qualities: string[] } e.g. { has_video: true, qualities: ["1080p","720p","480p"] } // Throws on network error; returns { has_video: false, qualities: [] } on non-video locators. export async function probeCapture(archiveId, locator) { diff --git a/frontend/src/components/CaptureDialog.jsx b/frontend/src/components/CaptureDialog.jsx index 937bdd3..46a5be4 100644 --- a/frontend/src/components/CaptureDialog.jsx +++ b/frontend/src/components/CaptureDialog.jsx @@ -1,5 +1,5 @@ import { useRef, useEffect, useState, useCallback } from 'react' -import { submitCapture, pollCaptureJob, probeCapture, probePlaylist, getInstanceSettings, uploadFile, deleteUpload } from '../api' +import { submitCapture, submitTextCapture, pollCaptureJob, probeCapture, probePlaylist, getInstanceSettings, uploadFile, deleteUpload } from '../api' let nextItemId = 1 @@ -157,6 +157,30 @@ function makeFileItem(filename) { } } +function makeTextItem() { + return { + id: nextItemId++, + kind: 'text', + title: '', + body: '', + mime: 'text/markdown', + // Fields present for submission-logic compatibility + locator: '', + quality: 'best', + probeState: 'idle', + probeQualities: null, + probeHasAudio: false, + playlistProbeState: 'idle', + playlistInfo: null, + playlistItems: null, + playlistQuality: null, + playlistExpanded: false, + syncEnabled: false, + error: null, + status: 'idle', + } +} + function applyPlaylistQuality(newQ, currentItems) { if (newQ === 'best') { return currentItems.map(item => ({ ...item, quality: 'best' })) @@ -430,9 +454,30 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on onToastRef.current(text, null, type, headline) } - async function submitBgJob(locator, quality, batchId, extraExtensions = {}) { + async function submitBgJob(submission, batchId) { const aid = archiveIdRef.current const id = crypto.randomUUID?.() ?? `job-${Date.now()}-${Math.random()}` + + // Text submission + if (submission.type === 'text') { + try { + const job = await submitTextCapture(aid, { title: submission.title, body: submission.body, mime: submission.mime }) + const locator = `text:${submission.title}` + // Notify App to add skeleton + persist + onJobStartedRef.current?.({ id, jobUid: job.job_uid, locator, archiveId: aid }) + startPolling(id, job.job_uid, locator, aid, batchId) + } catch (e) { + const msg = e.message || 'Submission failed.' + onToastRef.current(msg, `text:${submission.title}`) + settleBatch(batchId, 'failed', `text:${submission.title}`) + } + return + } + + // URL/file submission + const locator = submission.locator + const quality = submission.quality + const extraExtensions = submission.extraExtensions || {} // Capture session options at call time (synchronous — before first await) const extensions = { ublock_enabled: ublockEnabled, @@ -466,16 +511,18 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on // Guard against the Enter-key shortcut in CaptureRow bypassing the // disabled button — uploads must be complete before archiving starts. if (items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading')) return - const toSubmit = items.filter(it => - it.kind === 'file' ? (it.uploadStatus === 'done' && it.uploadLocator) : it.locator.trim() - ) + const toSubmit = items.filter(it => { + if (it.kind === 'file') return it.uploadStatus === 'done' && it.uploadLocator + if (it.kind === 'text') return it.title.trim() && it.body.trim() + return it.locator.trim() + }) if (toSubmit.length === 0) return - if (toSubmit.some(it => it.kind !== 'file' && hasConflict(it))) return - if (toSubmit.some(it => it.kind !== 'file' && ( + if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && hasConflict(it))) return + if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && ( it.probeState === 'probing' || (isPlaylistSource(it.locator) && it.playlistProbeState !== 'done')))) return - if (toSubmit.some(it => it.kind !== 'file' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0)) return + if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0)) return const batchId = toSubmit.length > 1 ? (crypto.randomUUID?.() ?? `batch-${Date.now()}`) : null @@ -485,9 +532,13 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on // Capture all submission data before any state changes const submissions = toSubmit.map(it => { if (it.kind === 'file') { - return { locator: it.uploadLocator, quality: 'best', extraExtensions: {} } + return { type: 'file', locator: it.uploadLocator, quality: 'best', extraExtensions: {} } + } + if (it.kind === 'text') { + return { type: 'text', title: it.title.trim(), body: it.body, mime: it.mime } } return { + type: 'url', locator: it.locator.trim(), quality: it.playlistItems !== null ? null : (it.quality || 'best'), extraExtensions: it.playlistItems !== null @@ -502,8 +553,8 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on setItems([makeItem()]) dialogRef.current?.close() // Submit each in background - submissions.forEach(({ locator, quality, extraExtensions }) => - submitBgJob(locator, quality, batchId, extraExtensions) + submissions.forEach(submission => + submitBgJob(submission, batchId) ) } @@ -630,8 +681,10 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on files.forEach(file => { const newItem = makeFileItem(file.name) setItems(prev => { - // Replace a sole empty URL row with the file item; otherwise append - if (prev.length === 1 && prev[0].kind !== 'file' && !prev[0].locator.trim()) { + // Only a normal URL row can be replaced. Text drafts deliberately use + // an empty compatibility locator, but their title/body must survive a + // file attachment and remain independently archivable. + if (prev.length === 1 && !prev[0].kind && !prev[0].locator.trim()) { return [newItem] } return [...prev, newItem] @@ -687,16 +740,18 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on const anyUploading = items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading') - const pendingCount = items.filter(it => - it.kind === 'file' ? (it.uploadStatus === 'done' && it.uploadLocator) : it.locator.trim() - ).length - const anyConflict = items.some(it => it.kind !== 'file' && hasConflict(it)) + const pendingCount = items.filter(it => { + if (it.kind === 'file') return it.uploadStatus === 'done' && it.uploadLocator + if (it.kind === 'text') return it.title.trim() && it.body.trim() + return it.locator.trim() + }).length + const anyConflict = items.some(it => it.kind !== 'file' && it.kind !== 'text' && hasConflict(it)) // True if any playlist row has had all its videos deleted — archive would be a no-op. const anyEmptyPlaylist = items.some(it => - it.kind !== 'file' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0 + it.kind !== 'file' && it.kind !== 'text' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0 ) const anyProbing = items.some(it => - it.kind !== 'file' && ( + it.kind !== 'file' && it.kind !== 'text' && ( it.probeState === 'probing' || // For playlist sources block unless probe completed successfully: // idle = debounce not yet fired; probing = in flight; error = no quality data. @@ -735,6 +790,17 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on item={item} onRemove={() => removeRow(item.id)} /> + ) : item.kind === 'text' ? ( + setItems(prev => prev.map(it => it.id === item.id ? { ...it, title: val } : it))} + onBodyChange={val => setItems(prev => prev.map(it => it.id === item.id ? { ...it, body: val } : it))} + onMimeChange={val => setItems(prev => prev.map(it => it.id === item.id ? { ...it, mime: val } : it))} + onRemove={() => removeRow(item.id)} + onSubmit={handleArchive} + /> ) : ( Upload file + {/* ── Advanced options ────────────────────────────── */} @@ -1139,3 +1211,63 @@ function CaptureFileRow({ item, onRemove }) { ) } + +function CaptureTextRow({ item, autoFocus, onTitleChange, onBodyChange, onMimeChange, onRemove, onSubmit }) { + const titleInputRef = useRef(null) + + useEffect(() => { + if (autoFocus) { + titleInputRef.current?.focus() + } + }, [autoFocus]) // eslint-disable-line react-hooks/exhaustive-deps + + return ( +
+
+ +
+ onTitleChange(e.target.value)} + maxLength={500} + /> +