mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 12:55:00 +02:00
feat: capture, summaries, search, and yt-dlp reliability (#38)
* ui: show spinner for pending captures
* feat(core): add text capture path with title + Markdown/plain body
- Add downloader/text.rs module with save() function that stages and hashes text content
- Support text/markdown and text/plain MIME types with .md and .txt extensions
- Add perform_text_capture() function for capturing user-supplied text
- Validates title (non-empty, max 500 chars) and body (non-empty, max 2 MiB)
- Creates blob records and entries with source_kind='text', entity_kind='document'
- Includes comprehensive unit tests for markdown, plain text, and validation
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(server): add POST /api/archives/:archive_id/captures/text
- Add CaptureTextBody struct for title, body, and optional MIME type
- Implement capture_text_handler with validation for empty fields and MIME type
- Route text submissions to perform_text_capture() in background
- Reuse existing capture job tracking and polling infrastructure
- Default MIME type to text/markdown when not specified
- Include route tests covering happy path, validation, auth, and error cases
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(frontend): add text-capture form to CaptureDialog
- Add submitTextCapture API client function with same error handling as submitCapture
- Create makeTextItem() factory for text capture state
- Implement CaptureTextRow component with title, body textarea, and MIME selector
- Add 'Add text' button in capture dialog toolbar
- Update handleArchive to filter and route text submissions
- Modify submitBgJob to detect and submit text items via submitTextCapture
- Skip probe and conflict checks for text items
- Reuse job tracking and batch settlement for text captures
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* feat(core): add entry_summaries schema + summarizer trait/providers
Per-entry LLM summaries as a regenerable child record, not a column on
archived_entries and not an on-disk artifact: an entry may carry several
summaries (one per provider/model/prompt version), any of which can be
discarded and recomputed. Generation is manual-only — nothing in capture.rs
calls into this module.
- database.rs: entry_summaries table + index, EntrySummaryRecord, and
upsert/update/find/latest helpers mirroring the capture_jobs style.
provider_model is stored as '' rather than NULL because SQLite treats
NULLs as distinct inside a UNIQUE index, which would stop the CLI
providers (no model) from ever deduping on the cache key.
- summarizer.rs: SummaryProvider trait with four implementations —
Anthropic Messages API, OpenAI-compatible chat completions, `claude -p`
and `codex exec -`. Configuration comes from env vars only (never TOML),
matching how yt-dlp / single-file / tweet-scraper are resolved, which
also keeps API keys out of anything the archive persists.
- archive.rs: EntryDetail gains latest_summary, populated by one extra
LIMIT 1 query in get_entry_detail. EntrySummaryView aliases the DB row
rather than duplicating it.
Implementation notes:
- No tokio in core. CLI timeouts are enforced structurally: stdout is
drained on its own thread and handed back over a channel so the calling
thread can recv_timeout and kill an overrunning child; stdin is written
on a third thread so a 48 KB prompt cannot deadlock against a child
waiting for us to read.
- HTML is reduced with regex rather than a parser: html5ever is not in the
tree, and a model tolerates imperfect whitespace. Paired tags are spelled
out per tag because Rust's regex engine has no backreferences by design.
- reqwest is declared with only the `blocking` feature here, so bodies are
serialized via .body(value.to_string()) instead of widening the
workspace dependency for .json().
- input_sha256 holds a SHA3-256 digest via hash::hash_bytes, the tree's one
hashing primitive; the content is truncated to 48 KB *before* hashing so
the cache key describes exactly the bytes the model saw.
Tests: no mockito/wiremock in dev-deps, and adding a mock HTTP server for
one JSON shape is a poor trade, so the two halves that can actually break
are tested directly — request-body builders and response parsers — leaving
only reqwest's own transport uncovered. Plus schema idempotency, cache-key
dedupe, cascade-on-delete, provider_from_env happy/missing-var paths, HTML
and tweet extraction, output normalization, and the CLI runner's stdin
round-trip, timeout kill, and nonzero-exit paths.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(server): add GET/POST /api/archives/:id/entries/:uid/summary
GET is read-only and gated exactly like entry detail, so a guest can read a
summary only for an entry whose content they could already read. POST
requires ROLE_USER, matching capture / tags / patch / rearchive; no auth
roles change.
Both the provider config and the content extraction resolve on the request
thread, before spawn_blocking. That is what lets a missing env var come back
as a synchronous 400 naming the exact variable, and an unsummarizable
artifact (video, audio) as a 400 saying so, rather than becoming a
background job the caller must poll only to learn about a config typo.
The pending row is claimed before spawning so the 202 can name a summary_uid
the client can poll immediately. summarize_entry owns the
pending → running → completed/failed transitions for that same row — the
cache key is identical, so both upserts resolve to one row — leaving the
handler to catch only the case where it fails before recording anything.
When !force and an identical cache key already completed, the existing row
comes back as a 200 with no new work.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(frontend): render Summary rail section + provider selector
New "Summary" rail section between the URL/Preview controls and .meta-list.
A completed summary renders as bold tl;dr, body paragraph, tag chips, and a
provider · model footer; missing or failed shows Generate; pending/running
shows an inline spinner and polls GET every 1500 ms until terminal.
- api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps
ApiError's { "error": ... } body so the missing-env-var message reaches
the user verbatim rather than as a bare status code.
- ContextRail.jsx: state seeds from detail.latest_summary so the section
renders immediately on selection. Polling is anchored on the summary
status rather than started inside the click handler, so a job still
running when the user navigates away and back is picked up again. A
transient poll failure is swallowed — the next tick retries, and a real
failure arrives as status === 'failed'.
- Regenerate passes force:true only when a completed summary is already
shown; otherwise the request can take the server's 200 cache-hit path.
- Provider choice persists in sessionStorage under archivr:summary:provider,
with try/catch around both accessors for private-mode browsers.
- Public sessions never see the selector or the Generate button, and the
section renders at all only when a completed summary made it through the
server's visibility gate.
- styles.css: .rail-summary-* only; spacing and the action button reuse
.rail-section and .rail-rearchive-btn. The spinner honours
prefers-reduced-motion — the text alone conveys the state.
- AGENTS.md: document the summary env vars alongside the existing
external-tool convention.
Smoke-tested end to end against a scratch archive with a seeded markdown
entry: claude_cli produced a real summary (pending → running → completed in
~11s); a local mock server exercised the openai_compatible transport and
confirmed the Bearer header, model, and system/user role split on the wire;
unconfigured providers return 400 naming the exact variable; a video entry
returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without
adding a row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): codex_cli — auto-discover binary + use --output-last-message
Two related fixes for the codex_cli summary provider:
1. Executable discovery. `ARCHIVR_CODEX_CLI` was already respected, but
without it the code resolved to bare `codex` and relied on PATH.
The ChatGPT desktop app installs codex at
`/Applications/ChatGPT.app/Contents/Resources/codex` and does not
put it on PATH, so users who only have the desktop app saw
'No such file or directory' with no hint. `resolve_cli` now walks
env override → a small set of well-known absolute paths → HOME
/.local/bin/<bare> → bare fallback. Same treatment applied to
claude_cli for symmetry (/opt/homebrew/bin/claude, /usr/local/bin/
claude, HOME/.local/bin/claude).
2. Clean output. `codex exec -` writes a runtime header ("OpenAI
Codex vX", session id, sandbox, model), the assistant reply, and a
footer ("tokens used", replay of the reply) to stdout. The JSON
extractor took the first '{' from the *user prompt echo* and the
last '}' from the trailing replay, producing invalid text that
fell through to the "raw text under summary" fallback path. Now
uses `--output-last-message <tempfile>` and reads only the final
assistant message. Fallback (positional prompt) uses the same
flag. Tempfile is cleaned up on all paths, incl. spawn failure.
* fix(frontend): give text-capture row real CSS
The text row shipped with semantic classnames (`capture-text-inputs`,
`capture-text-title`, `capture-text-body`, `capture-text-mime`,
`capture-text-icon`) but no CSS rules. Falling through to the parent
`.capture-row-main` flex-row (`display: flex; align-items: center`)
meant the title, textarea, and mime-select stacked as intrinsic-width
boxes centered on the tall body, producing a layout where the body
floated to the top-right, the title box appeared BELOW it, and the mime
selector rendered as an unstyled OS dropdown.
Fix:
- `.capture-text-row .capture-row-main` uses `align-items: flex-start`
so the leading icon and trailing × pin to the top of the block.
- `.capture-text-inputs` is now a full-width column-flex container with
proper gaps.
- `.capture-text-title` reuses the 44px input height and typography of
`.capture-input`; `.capture-text-body` gets a 140px min-height,
vertical resize, and matching border/focus treatment.
- `.capture-text-mime` is styled as a small chip with a custom caret
so it matches `.capture-quality` and stops looking like a raw
`<select>`. Sits in a right-aligned footer under the body.
- `.capture-text-icon` gets a 44px column so it aligns with the title
input; remove button gets a small top-margin for the same reason.
Rebuilt static bundle bumped as well (`index-BLxoi9rt.css`,
`index-CQcpPA_I.js`).
* chore(static): rebuild bundle after text-row CSS merge
* fix(core): summarize tweets + walk all tweets in a thread
Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.
Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.
Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
`raw_tweet_json` artifacts (falling back to `primary_media` for
archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
markdown/html/json branches, then join thread pieces with a
`---` separator so the model sees a real paragraph break between
statuses instead of one flowing document.
Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.
* feat(frontend): preview text-capture entries (.md / .txt)
Text captures land as `.md` (Markdown) or `.txt` (plain) blobs, but
PreviewPanel only dispatched on video/audio/image/pdf/html extensions,
so opening a text entry hit the 'No preview available' fallback with
the raw artifact path exposed.
- New `TextPreview` component fetches the primary artifact as text,
renders it in a monospace `<pre>` with word-wrap, and shows the
entry title on top and the MIME as a small trailing tag. Handles
loading/error states.
- `PreviewPanel` gains a `TEXT_EXTS` set + a branch that dispatches
to `TextPreview` for `md` / `markdown` / `txt`.
- CSS is padded and centered to ~780px so a text note reads like a
document rather than an edge-to-edge terminal dump.
v1 intentionally does NOT parse Markdown: keeping frontend deps at
react+react-dom only. Bump to a real Markdown renderer if we start
capturing Markdown-authored notes.
* fix(nix): pin yt-dlp from its own release + wire into server wrapper
Two independent problems, one commit:
1. Stale binary. nixpkgs-provided `pkgs.yt-dlp` on the pinned
nixos-unstable rev is 2026.03.17 (Mar 2026). yt-dlp itself
releases days-to-weeks, and YouTube frequently rotates the
player-signature / client surfaces the older builds request
(`android_vr` is the current casualty), which returns HTTP 403
mid-download for the format specs archivr passes (`-f
bestvideo+bestaudio/best`). Even bumping the nixpkgs input would
leave us dependent on that channel's yt-dlp cadence.
Fetch the upstream zipapp directly instead
(github.com/yt-dlp/yt-dlp/releases/download/<ver>/yt-dlp), wrap so
`python3` and `ffmpeg` are on PATH, and pin version+hash in one
place. Bumping is: change version, replace hash from
`nix hash file <url>`.
2. Missing pin in server wrapper. `archivr-cli` was already wrapped
with `--set ARCHIVR_YT_DLP` + a PATH prefix; `archivr-server`
was NOT — it only pinned single-file, chrome, and the tweet
scraper, silently falling back to whatever `yt-dlp` the user
happened to have on PATH. Server captures therefore inherited
the user's (often stale) system yt-dlp regardless of the flake
pin. Same wrapper flags now apply to both binaries.
devShell keeps `pkgs.yt-dlp` for now: the dev shell is a
convenience, not a release surface, and matching wouldn't fit in this
commit without duplicating the derivation across let-scopes.
* chore(static): rebuild bundle for round-3 fixes
* chore(nix): pin python 3.12 for yt-dlp zipapp (avoid py3.14 libffi crash on darwin/arm64)
* feat(core): resolve_yt_dlp picks the newer of pinned vs state-dir
The nix flake wrapper pins a yt-dlp via ARCHIVR_YT_DLP, but yt-dlp rots
fast — extractors break within weeks of a pin. Add a resolver that probes
`--version` on both the pinned binary and a user-installed copy under the
mutable state dir, and runs whichever is newer.
Version strings are YYYY.MM.DD, so plain string ordering is chronological.
Ties resolve toward the state dir: a user who installed it there did so
deliberately. ARCHIVR_YT_DLP_FORCE bypasses the comparison entirely, and
with no candidate at all we fall back to bare `yt-dlp` on PATH — exactly
the previous behaviour.
Resolution is cached in a OnceLock so `--version` costs one subprocess per
process, and all four inline env::var lookups now go through it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: bump yt-dlp from upstream releases, not nixpkgs
The flake no longer takes yt-dlp from nixpkgs; a dedicated `ytDlp`
derivation fetches the upstream release binary directly and pins both
`version` and an SRI `hash`. That makes the previous workflow inert: it
ran `nix flake update nixpkgs` and compared `nixpkgs#yt-dlp.version`
before and after, so it could churn the lockfile forever without ever
moving the version we actually ship.
The workflow now reads the pinned version straight out of the `ytDlp`
block in flake.nix, asks the GitHub API for yt-dlp's latest release tag,
short-circuits when they already match, downloads the new release to
recompute its SRI hash (required — the hash is part of the derivation's
identity, so the URL cannot be changed alone), and rewrites the three
pinned fields under a sed range address scoped to that block so sibling
pins like ublockLite and isdcac are untouched. It asserts only flake.nix
changed and that the new version appears exactly twice before opening
the PR.
* feat(cli): add `archivr yt-dlp update|status` subcommand
`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.
The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.
`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.
reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(maintainer): document summarizer, text capture, and yt-dlp lifecycle
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(readme): document LLM summaries, text notes, yt-dlp resolver + bump paths
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(plan): specify X Article, image summaries, summary search
* fix: summarize X Article text
* feat: search completed summary tags
* fix: preserve X Article block order
* feat: model opt-in summary images
* feat: attach opted-in images to summaries
* feat: accept image summary requests
* feat: add summary image consent control
* docs: explain X Article, image summaries and summary search
* chore(static): rebuild bundle for summary image consent
* fix: show civilized unsupported summary errors
* chore(static): rebuild bundle for civilized summary errors
* fix: keep text previews and summary state scoped to entry
* fix: show forced yt-dlp candidate in status
* fix: infer trusted MIME for tweet images
* fix: preserve text capture bytes and hide synthetic URL
* fix: recover and preserve summary attempts
* fix: bound codex fallback and record resolved model
* fix: protect public summary diagnostics
* fix: guard summary callbacks during entry render
* fix: scope summary callbacks to selected entry
* test: cover terminal newline in text capture
* fix: preserve missing summary entry status
* test: cover tweet image summary selection
* docs: record summary lifecycle and review hardening
* chore(static): rebuild bundle for Sol review fixes
* fix: retain completed summary during regeneration display
* fix: hide superseded summary attempts
* chore(static): rebuild bundle after regeneration display fix
* fix: allow full-size text capture requests
* fix: allow escaped full-size text captures
* fix: preserve text draft whitespace in capture UI
* chore(static): rebuild bundle for text whitespace fix
This commit is contained in:
parent
b4b4e67157
commit
79ac44834e
35 changed files with 7201 additions and 218 deletions
131
.github/workflows/update-ytdlp.yml
vendored
131
.github/workflows/update-ytdlp.yml
vendored
|
|
@ -13,39 +13,126 @@ jobs:
|
||||||
pull-requests: write
|
pull-requests: write
|
||||||
|
|
||||||
steps:
|
steps:
|
||||||
- uses: actions/checkout@v4
|
- name: Check out repository
|
||||||
|
uses: actions/checkout@v4
|
||||||
|
|
||||||
- uses: DeterminateSystems/nix-installer-action@main
|
- name: Install Nix
|
||||||
|
uses: DeterminateSystems/nix-installer-action@main
|
||||||
|
|
||||||
- uses: DeterminateSystems/magic-nix-cache-action@main
|
- name: Enable Nix cache
|
||||||
|
uses: DeterminateSystems/magic-nix-cache-action@main
|
||||||
|
|
||||||
- name: Get current yt-dlp version
|
- name: Read currently pinned yt-dlp version
|
||||||
id: before
|
id: current
|
||||||
run: |
|
run: |
|
||||||
rev=$(jq -r '.nodes.nixpkgs.locked.rev' flake.lock)
|
set -euo pipefail
|
||||||
version=$(nix eval --raw "github:nixos/nixpkgs/${rev}#yt-dlp.version")
|
current=$(sed -n '/ytDlp = pkgs.stdenv.mkDerivation/,/^ };$/{ s/^ *version = "\([^"]*\)";/\1/p; }' flake.nix | head -1)
|
||||||
echo "version=${version}" >> "$GITHUB_OUTPUT"
|
if [ -z "$current" ]; then
|
||||||
|
echo "::error::Could not read the pinned yt-dlp version from flake.nix. Did the ytDlp derivation move or get renamed?"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "Currently pinned yt-dlp: $current"
|
||||||
|
echo "version=${current}" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
- name: Update nixpkgs
|
- name: Query latest yt-dlp release
|
||||||
run: nix flake update nixpkgs
|
id: latest
|
||||||
|
env:
|
||||||
- name: Get new yt-dlp version
|
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||||
id: after
|
|
||||||
run: |
|
run: |
|
||||||
rev=$(jq -r '.nodes.nixpkgs.locked.rev' flake.lock)
|
set -euo pipefail
|
||||||
version=$(nix eval --raw "github:nixos/nixpkgs/${rev}#yt-dlp.version")
|
latest=$(curl -sSL \
|
||||||
echo "version=${version}" >> "$GITHUB_OUTPUT"
|
-H "Authorization: Bearer $GITHUB_TOKEN" \
|
||||||
|
-H "Accept: application/vnd.github+json" \
|
||||||
|
https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest | jq -r .tag_name)
|
||||||
|
if [ -z "$latest" ] || [ "$latest" = "null" ]; then
|
||||||
|
echo "::error::Could not determine the latest yt-dlp release tag from the GitHub API."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "Latest yt-dlp release: $latest"
|
||||||
|
echo "version=${latest}" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
- name: Open PR if yt-dlp was updated
|
- name: Decide whether an update is needed
|
||||||
if: steps.before.outputs.version != steps.after.outputs.version
|
id: check
|
||||||
|
env:
|
||||||
|
CURRENT: ${{ steps.current.outputs.version }}
|
||||||
|
LATEST: ${{ steps.latest.outputs.version }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
if [ "$CURRENT" = "$LATEST" ]; then
|
||||||
|
echo "Already at $CURRENT"
|
||||||
|
echo "changed=false" >> "$GITHUB_OUTPUT"
|
||||||
|
else
|
||||||
|
echo "Update available: $CURRENT -> $LATEST"
|
||||||
|
echo "changed=true" >> "$GITHUB_OUTPUT"
|
||||||
|
fi
|
||||||
|
|
||||||
|
- name: Compute SRI hash of the new release
|
||||||
|
id: hash
|
||||||
|
if: steps.check.outputs.changed == 'true'
|
||||||
|
env:
|
||||||
|
LATEST: ${{ steps.latest.outputs.version }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
curl -sSL --fail \
|
||||||
|
"https://github.com/yt-dlp/yt-dlp/releases/download/${LATEST}/yt-dlp" \
|
||||||
|
-o /tmp/yt-dlp
|
||||||
|
hash=$(nix hash file --sri --type sha256 /tmp/yt-dlp)
|
||||||
|
case "$hash" in
|
||||||
|
sha256-*) ;;
|
||||||
|
*)
|
||||||
|
echo "::error::Computed hash '${hash}' is not an SRI sha256 hash."
|
||||||
|
exit 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
echo "SRI hash: $hash"
|
||||||
|
echo "hash=${hash}" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
|
- name: Rewrite the yt-dlp pin in flake.nix
|
||||||
|
if: steps.check.outputs.changed == 'true'
|
||||||
|
env:
|
||||||
|
CURRENT: ${{ steps.current.outputs.version }}
|
||||||
|
LATEST: ${{ steps.latest.outputs.version }}
|
||||||
|
HASH: ${{ steps.hash.outputs.hash }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
sed -i \
|
||||||
|
-e "/ytDlp = pkgs.stdenv.mkDerivation/,/^ };\$/{ s|^\( *version = \"\)[^\"]*\(\";\)|\1${LATEST}\2|; }" \
|
||||||
|
-e "/ytDlp = pkgs.stdenv.mkDerivation/,/^ };\$/{ s|\(url = \"https://github.com/yt-dlp/yt-dlp/releases/download/\)[^/]*\(/yt-dlp\";\)|\1${LATEST}\2|; }" \
|
||||||
|
-e "/ytDlp = pkgs.stdenv.mkDerivation/,/^ };\$/{ s|^\( *hash = \"\)sha256-[^\"]*\(\";\)|\1${HASH}\2|; }" \
|
||||||
|
flake.nix
|
||||||
|
|
||||||
|
echo "--- git diff --stat ---"
|
||||||
|
git diff --stat flake.nix
|
||||||
|
|
||||||
|
changed_files=$(git diff --name-only)
|
||||||
|
if [ "$changed_files" != "flake.nix" ]; then
|
||||||
|
echo "::error::Expected only flake.nix to change, got: ${changed_files}"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
occurrences=$(grep -c "$LATEST" flake.nix || true)
|
||||||
|
if [ "$occurrences" -ne 2 ]; then
|
||||||
|
echo "::error::Expected the new version ${LATEST} to appear twice in flake.nix (version line + URL), found ${occurrences}."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
if ! grep -q "$HASH" flake.nix; then
|
||||||
|
echo "::error::New SRI hash was not written into flake.nix."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "Rewrote yt-dlp pin: ${CURRENT} -> ${LATEST}"
|
||||||
|
|
||||||
|
- name: Open pull request
|
||||||
|
if: steps.check.outputs.changed == 'true'
|
||||||
uses: peter-evans/create-pull-request@v6
|
uses: peter-evans/create-pull-request@v6
|
||||||
with:
|
with:
|
||||||
branch: auto/yt-dlp-update
|
branch: auto/yt-dlp-update
|
||||||
delete-branch: true
|
delete-branch: true
|
||||||
commit-message: "chore: yt-dlp ${{ steps.before.outputs.version }} → ${{ steps.after.outputs.version }}"
|
commit-message: "chore(nix): yt-dlp ${{ steps.current.outputs.version }} → ${{ steps.latest.outputs.version }}"
|
||||||
title: "chore: yt-dlp ${{ steps.before.outputs.version }} → ${{ steps.after.outputs.version }}"
|
title: "chore(nix): yt-dlp ${{ steps.current.outputs.version }} → ${{ steps.latest.outputs.version }}"
|
||||||
body: |
|
body: |
|
||||||
Automated `flake.lock` update. yt-dlp bumped from `${{ steps.before.outputs.version }}` to `${{ steps.after.outputs.version }}`.
|
Automated bump of the pinned yt-dlp release. Old: `${{ steps.current.outputs.version }}`. New: `${{ steps.latest.outputs.version }}`. SRI hash: `${{ steps.hash.outputs.hash }}`.
|
||||||
|
|
||||||
Triggered by the weekly nixpkgs check.
|
Upstream release: https://github.com/yt-dlp/yt-dlp/releases/tag/${{ steps.latest.outputs.version }}
|
||||||
labels: dependencies
|
labels: dependencies
|
||||||
|
|
|
||||||
53
AGENTS.md
53
AGENTS.md
|
|
@ -18,6 +18,31 @@ Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture
|
||||||
|
|
||||||
YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the playlist probe (fetching per-video quality metadata before archiving), the multi-video download loop, and sync mode (skipping already-archived videos when re-archiving a playlist or channel).
|
YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the playlist probe (fetching per-video quality metadata before archiving), the multi-video download loop, and sync mode (skipping already-archived videos when re-archiving a playlist or channel).
|
||||||
|
|
||||||
|
Pasted text takes a much shorter path: `perform_text_capture()` (`capture.rs`) skips source detection
|
||||||
|
and every downloader shell-out — `downloader/text.rs` stages the body under `temp/`, hashes it, and the
|
||||||
|
blob lands in `raw/` like any other artifact. Entrypoint is
|
||||||
|
`POST /api/archives/:archive_id/captures/text`. Its body is byte-preserving, its entry has no fabricated
|
||||||
|
`original_url`, and its normal text preview opens from the entry rail.
|
||||||
|
|
||||||
|
LLM summaries are a post-capture, manual-only subsystem: `crates/archivr-core/src/summarizer.rs` behind
|
||||||
|
`GET`/`POST /api/archives/:archive_id/entries/:entry_uid/summary`, cached in the `entry_summaries`
|
||||||
|
table per (entry, provider, model, prompt version, input hash). See `ARCHIVR-MENTAL-MODEL.md` for the
|
||||||
|
provider set and the status lifecycle.
|
||||||
|
|
||||||
|
The requested provider model is the cache identity; a provider-returned resolved model is display attribution.
|
||||||
|
At startup, pending/running attempts interrupted by shutdown are failed. A regeneration keeps the previous completed
|
||||||
|
summary visible until its replacement completes; public readers receive completed content only, never diagnostics.
|
||||||
|
|
||||||
|
`SummaryBuildOptions` keeps summaries text-only unless `include_images` is set. The input digest includes that flag and
|
||||||
|
the selected blobs' SHA-256, MIME types, and sizes, so a distinct image selection cannot reuse a text-only cache row.
|
||||||
|
Candidates are `media` artifacts only: `jpg`/`jpeg`, `png`, `webp`, `gif`, and `avif`, capped at four images, 5 MiB each,
|
||||||
|
and 12 MiB in aggregate. Anthropic HTTP, OpenAI-compatible HTTP, and Codex support images; Claude CLI does not. The core
|
||||||
|
remains synchronous: the server puts provider work in its blocking boundary rather than introducing async to
|
||||||
|
`archivr-core`.
|
||||||
|
|
||||||
|
Entry free-text search includes summary text (and generated JSON tags inside it) from the latest completed summary only.
|
||||||
|
Pending and failed rows do not match, and a newer pending or failed request does not hide an older completed summary.
|
||||||
|
|
||||||
Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8).
|
Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8).
|
||||||
|
|
||||||
The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`.
|
The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`.
|
||||||
|
|
@ -32,6 +57,7 @@ The server mounts multiple archives from a TOML registry (`crates/archivr-server
|
||||||
| `frontend/src/` | React app: `App.jsx` (root state + custom routing), `api.js` (fetch client), `components/`, `styles.css` |
|
| `frontend/src/` | React app: `App.jsx` (root state + custom routing), `api.js` (fetch client), `components/`, `styles.css` |
|
||||||
| `docs/` | User docs (`README.md`), `superpowers/plans/` and `superpowers/specs/` (dated design docs — write plans there before large features) |
|
| `docs/` | User docs (`README.md`), `superpowers/plans/` and `superpowers/specs/` (dated design docs — write plans there before large features) |
|
||||||
| `modules/nixos/` | NixOS module (`services.archivr-server`) |
|
| `modules/nixos/` | NixOS module (`services.archivr-server`) |
|
||||||
|
| `.github/workflows/` | `update-ytdlp.yml` — weekly cron that PRs a yt-dlp version bump into `flake.nix` |
|
||||||
| `vendor/twitter/` | Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually. |
|
| `vendor/twitter/` | Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually. |
|
||||||
| `vendor/readability/` | Mozilla `Readability.js`, concatenated into the SingleFile reader-mode browser script by `downloader/singlefile.rs`. |
|
| `vendor/readability/` | Mozilla `Readability.js`, concatenated into the SingleFile reader-mode browser script by `downloader/singlefile.rs`. |
|
||||||
| `testing/` | Legacy scraping scripts + sample data. `testing/creds.txt` holds real tokens — never read, commit, or print it. |
|
| `testing/` | Legacy scraping scripts + sample data. `testing/creds.txt` holds real tokens — never read, commit, or print it. |
|
||||||
|
|
@ -72,7 +98,14 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
|
||||||
- **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2.
|
- **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2.
|
||||||
- **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal.
|
- **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal.
|
||||||
- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`. Downloaders shell out to subprocesses; resolve binaries through these vars.
|
- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`. Downloaders shell out to subprocesses; resolve binaries through these vars.
|
||||||
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`.
|
- **yt-dlp is resolved, not just read**: `ARCHIVR_YT_DLP` (set by the flake wrappers) is only the
|
||||||
|
*pinned candidate* handed to `resolve_yt_dlp()` (`downloader/ytdlp.rs`), which compares it against a
|
||||||
|
self-updated copy in the state dir. `ARCHIVR_YT_DLP_FORCE` (absolute path) bypasses that comparison
|
||||||
|
entirely; `archivr yt-dlp status` shows and chooses that forced candidate when it applies; `ARCHIVR_STATE_DIR`
|
||||||
|
relocates the state dir. Never spawn bare `yt-dlp` — call the resolver.
|
||||||
|
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them. The two CLI vars are optional overrides: unset, `resolve_cli()` auto-discovers well-known absolute installs first (`/opt/homebrew/bin/claude`, `/usr/local/bin/claude`; `/Applications/ChatGPT.app/Contents/Resources/codex`, `/opt/homebrew/bin/codex`, `/usr/local/bin/codex`), then `$HOME/.local/bin/<name>`, then the bare name on PATH — the absolute defaults matter because the ChatGPT desktop app ships `codex` off PATH. Frontend static output is generated; never hand-edit `crates/archivr-server/static/`.
|
||||||
|
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`. Summary polling and generate callbacks must remain scoped to the currently selected entry; image-inclusion behavior is unchanged. In-progress captures render through `SkeletonEntryRow.jsx`, which is a compact spinner + locator + "Archiving…" line (with a playlist/channel hint), **not** a grey skeleton block — don't reintroduce placeholder shimmer. Layout comes from semantic classes (e.g. `.capture-text-row` in `styles.css`), never from fallthrough on a generic row class; give a new row shape its own class.
|
||||||
|
- **CLI providers parse a file, not stdout**: `codex` is invoked as `codex exec --output-last-message <tempfile> -` (prompt on stdin) and the reply is read back from that file — raw stdout carries a runtime header, an echo of the user prompt, and a `tokens used` footer that the JSON extractor will happily mistake for the answer. There is a positional-prompt fallback for older builds that reject `-`. Keep any new CLI provider on the same "give me only the final message" contract.
|
||||||
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
|
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
|
||||||
|
|
||||||
## Important Files
|
## Important Files
|
||||||
|
|
@ -80,9 +113,16 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
|
||||||
- `crates/archivr-server/src/main.rs` — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup).
|
- `crates/archivr-server/src/main.rs` — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup).
|
||||||
- `crates/archivr-server/src/routes.rs` — all HTTP handlers and the router; grep here first for API work.
|
- `crates/archivr-server/src/routes.rs` — all HTTP handlers and the router; grep here first for API work.
|
||||||
- `crates/archivr-core/src/capture.rs` — `perform_capture()`, `Source` enum, shorthand parsing.
|
- `crates/archivr-core/src/capture.rs` — `perform_capture()`, `Source` enum, shorthand parsing.
|
||||||
- `crates/archivr-core/src/downloader/ytdlp.rs` — yt-dlp integration; YouTube playlist/channel probe and download, sync mode logic.
|
- `crates/archivr-core/src/downloader/ytdlp.rs` — every yt-dlp shell-out (playlist/channel probe and download, sync mode) **plus** the binary resolver: `resolve_yt_dlp()`, `state_dir()`, `probe_version()`.
|
||||||
|
- `crates/archivr-core/src/downloader/text.rs` — pasted-text staging + hashing (`save()` → `StagedText`); accepts only `text/plain` and `text/markdown`.
|
||||||
|
- `crates/archivr-core/src/summarizer.rs` — the `SummaryProvider` trait and its four implementations (Anthropic HTTP, OpenAI-compatible HTTP, `claude` CLI, `codex` CLI), `PROMPT_VERSION`, `resolve_cli()`, prompt assembly, and `build_summary_input()` (artifact selection + HTML/text/JSON reduction).
|
||||||
|
- `crates/archivr-cli/src/main.rs` — CLI entry point, including the `archivr yt-dlp update|status` subcommand (staged, atomic zipapp install into the state dir).
|
||||||
|
- `.github/workflows/update-ytdlp.yml` — weekly (`0 6 * * 1`) + manual auto-bump of the `flake.nix` yt-dlp pin.
|
||||||
- `crates/archivr-core/src/database.rs` — single source of truth for all SQLite schema and queries (both archive and auth DBs).
|
- `crates/archivr-core/src/database.rs` — single source of truth for all SQLite schema and queries (both archive and auth DBs).
|
||||||
- `frontend/src/App.jsx` / `frontend/src/api.js` — frontend root state and API surface.
|
- `frontend/src/App.jsx` / `frontend/src/api.js` — frontend root state and API surface.
|
||||||
|
- `frontend/src/components/CaptureDialog.jsx` — `CaptureRow` (locator input, playlist quality selectors) and `CaptureTextRow` (the "Add text" flow), plus job polling.
|
||||||
|
- `frontend/src/components/ContextRail.jsx` — the Summary section: provider selector, generate/regenerate, and summary polling.
|
||||||
|
- `frontend/src/components/TextPreview.jsx` — preview renderer for text/markdown entries.
|
||||||
- `docker/config.example.toml` — server config schema: `bind`, `auth_db_path`, repeated `[[archives]]` (`id`, `label`, `archive_path`).
|
- `docker/config.example.toml` — server config schema: `bind`, `auth_db_path`, repeated `[[archives]]` (`id`, `label`, `archive_path`).
|
||||||
- `flake.nix`, `modules/nixos/archivr-server.nix`, `Dockerfile`, `docker-compose.yml` — deployment surfaces; config schema changes must be reflected in all of them plus `docs/README.md`.
|
- `flake.nix`, `modules/nixos/archivr-server.nix`, `Dockerfile`, `docker-compose.yml` — deployment surfaces; config schema changes must be reflected in all of them plus `docs/README.md`.
|
||||||
|
|
||||||
|
|
@ -93,9 +133,16 @@ No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy
|
||||||
- Runtime binaries the app expects on PATH or via env vars: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
|
- Runtime binaries the app expects on PATH or via env vars: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
|
||||||
- `.gitignore` is **default-deny with an allowlist** — new top-level files/dirs are invisible to git until explicitly allowed there.
|
- `.gitignore` is **default-deny with an allowlist** — new top-level files/dirs are invisible to git until explicitly allowed there.
|
||||||
- Frontend build output (`crates/archivr-server/static/`) is generated; never hand-edit it.
|
- Frontend build output (`crates/archivr-server/static/`) is generated; never hand-edit it.
|
||||||
|
- **yt-dlp is pinned to a specific GitHub release** in the `ytDlp` derivation in `flake.nix` (zipapp
|
||||||
|
fetched from `github.com/yt-dlp/yt-dlp/releases`, wrapped with `python312` + `ffmpeg`) — not taken
|
||||||
|
from nixpkgs. Both the `archivr` and `archivr-server` wrappers set `ARCHIVR_YT_DLP` from it. Three
|
||||||
|
ways to bump: the weekly `Update yt-dlp` workflow (automatic PR), `archivr yt-dlp update`
|
||||||
|
(per-machine, into the state dir), or editing the three fields of the `ytDlp` block by hand. Which
|
||||||
|
binary actually runs is decided at runtime by `resolve_yt_dlp()` in `downloader/ytdlp.rs`; use
|
||||||
|
`archivr yt-dlp status` to see the candidates and the winner.
|
||||||
|
|
||||||
## Testing & QA
|
## Testing & QA
|
||||||
|
|
||||||
- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p <crate>`.
|
- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`, and newer: `summarizer.rs` — 23 tests over provider construction, CLI/env resolution, output extraction, HTML/text/JSON input reduction and tweet-thread joining; `downloader/ytdlp.rs` — 13 tests, several covering resolver priority and version tie-breaking). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p <crate>`.
|
||||||
- **Frontend**: no test framework. Storybook (`bun run storybook`) is the component QA surface — stories are colocated `*.stories.jsx` files; add one when adding a nontrivial component.
|
- **Frontend**: no test framework. Storybook (`bun run storybook`) is the component QA surface — stories are colocated `*.stories.jsx` files; add one when adding a nontrivial component.
|
||||||
- Manual smoke test for server changes: build frontend, `cargo run -p archivr-server -- <config.toml>`, exercise `/api/*`.
|
- Manual smoke test for server changes: build frontend, `cargo run -p archivr-server -- <config.toml>`, exercise `/api/*`.
|
||||||
|
|
|
||||||
|
|
@ -81,7 +81,7 @@ There are two user-facing binaries:
|
||||||
|
|
||||||
| Binary | Purpose |
|
| Binary | Purpose |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `archivr` | CLI for initializing archives and capturing material into one archive |
|
| `archivr` | CLI for initializing archives and capturing material into one archive (also `yt-dlp status\|update`) |
|
||||||
| `archivr-server` | Web server for browsing one or more existing archives |
|
| `archivr-server` | Web server for browsing one or more existing archives |
|
||||||
|
|
||||||
The CLI writes archive data:
|
The CLI writes archive data:
|
||||||
|
|
@ -128,6 +128,16 @@ sequenceDiagram
|
||||||
CLI->>User: terminal result
|
CLI->>User: terminal result
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Pasted text short-circuits most of that.** `perform_text_capture` (`capture.rs`) is not a `Source`
|
||||||
|
route: there is no locator to classify, no URL probe, and no downloader subprocess. It validates the
|
||||||
|
title (non-empty, ≤ 500 chars), the body (non-empty, ≤ 2 MiB) and the MIME type (`text/plain` or
|
||||||
|
`text/markdown` only), then calls `downloader/text.rs` to write the bytes into `store/temp/<timestamp>/`
|
||||||
|
and hash them. From there it rejoins the normal path — dedup into `raw/A/B/HASH.EXT`, then run, entry
|
||||||
|
and artifact rows. The server exposes it as `POST /api/archives/:archive_id/captures/text`, and the UI
|
||||||
|
drives it from `CaptureTextRow` in `CaptureDialog.jsx`. The body is preserved byte-for-byte, the entry's
|
||||||
|
`original_url` remains empty (no fabricated `text:` URL), and its normal entry-rail preview renders through
|
||||||
|
`TextPreview.jsx`.
|
||||||
|
|
||||||
## Web Capture Pipeline
|
## Web Capture Pipeline
|
||||||
|
|
||||||
Web pages (`Source::WebPage`) take a longer path than yt-dlp or tweets:
|
Web pages (`Source::WebPage`) take a longer path than yt-dlp or tweets:
|
||||||
|
|
@ -158,6 +168,82 @@ sequenceDiagram
|
||||||
Server-->>Browser: JSON
|
Server-->>Browser: JSON
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## LLM Summaries
|
||||||
|
|
||||||
|
Summaries are a **post-capture, manually triggered** subsystem. Nothing in `capture.rs` calls the
|
||||||
|
summarizer; a summary exists only because someone pressed generate in the Summary section of
|
||||||
|
`ContextRail.jsx`.
|
||||||
|
|
||||||
|
`crates/archivr-core/src/summarizer.rs` defines one `SummaryProvider` trait with four implementations:
|
||||||
|
Anthropic HTTP, OpenAI-compatible HTTP, the local `claude` CLI, and the local `codex` CLI. Each is
|
||||||
|
built purely from environment variables (`provider_from_env`), so no key or model name is ever written
|
||||||
|
into archive data. `PROMPT_VERSION` in the same file stamps every row, so changing the prompt
|
||||||
|
invalidates the cache instead of silently mixing generations.
|
||||||
|
|
||||||
|
`entry_summaries` (schema in `database.rs`) is that cache, unique on
|
||||||
|
(`entry_id`, `provider_kind`, `provider_model`, `prompt_version`, `input_sha256`) — the requested provider
|
||||||
|
model is the cache identity, so the same entry summarised by two providers, two requested models, or after a
|
||||||
|
prompt change yields distinct rows, while a repeat request with identical inputs reuses one. When a provider
|
||||||
|
returns its concrete resolved model, it is stored separately and displayed as attribution without changing that
|
||||||
|
identity. Rows move `pending` → `running` → `completed` | `failed`, mirroring how capture jobs are tracked;
|
||||||
|
the frontend polls only for the currently selected entry, and its generate callbacks are scoped to that same
|
||||||
|
selection. Image-selection behavior remains unchanged.
|
||||||
|
|
||||||
|
On startup the server marks interrupted `pending` or `running` attempts failed. Regeneration is non-destructive:
|
||||||
|
the prior completed summary stays visible until a replacement completes successfully. Public readers receive only
|
||||||
|
completed summary content, never pending/failed state or diagnostic error text.
|
||||||
|
|
||||||
|
The summary path is deliberately explicit: UI consent (`Include attached images`) → core selection → input digest and
|
||||||
|
cache lookup → provider transport → `pending`/`running`/`completed` lifecycle. Text is the default. When consent is
|
||||||
|
present, `SummaryBuildOptions::include_images` admits only bounded `media` image candidates and the digest includes both
|
||||||
|
the flag and selected blob identity, MIME type, and size. The core is still synchronous; the server owns the blocking
|
||||||
|
boundary. Anthropic HTTP, OpenAI-compatible HTTP, and Codex can transport the selected image data; Claude CLI receives
|
||||||
|
text only.
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
UI["ContextRail Summary"] -->|POST .../summary| Server
|
||||||
|
Server --> Input["build_summary_input()"]
|
||||||
|
Input --> Artifacts["entry artifacts on disk"]
|
||||||
|
Server --> Row["entry_summaries: pending → running"]
|
||||||
|
Server --> Provider["SummaryProvider (HTTP or CLI)"]
|
||||||
|
Provider --> Row2["completed / failed"]
|
||||||
|
UI -->|GET .../summary poll| Row2
|
||||||
|
```
|
||||||
|
|
||||||
|
**X Articles and tweet threads are why the artifact lookup is special.** For most entries `build_summary_input` reads
|
||||||
|
the single `primary_media` artifact. A `tweet` or `tweet_thread` entry has no `primary_media` — it has
|
||||||
|
N `raw_tweet_json` artifacts, one per status in the thread. So the summarizer selects on the
|
||||||
|
`raw_tweet_json` role instead, loads **all** matching artifacts in order, and joins them with
|
||||||
|
`\n\n---\n\n`; a `---` line reads as a hard paragraph break to every model, keeping individual
|
||||||
|
statuses from bleeding into one another. For an X Article, the reducer prefers article text over the tweet's body:
|
||||||
|
`plain_text`, then flattened ordered blocks, then `preview_text`, then `summary_text`; only then does it fall back to
|
||||||
|
`full_text`/`text`/`content`/`body`. Any change to article reduction or thread status storage must be mirrored here.
|
||||||
|
|
||||||
|
## yt-dlp Lifecycle
|
||||||
|
|
||||||
|
There is no single yt-dlp. Up to three can exist on one machine:
|
||||||
|
|
||||||
|
1. **The flake pin** — the `ytDlp` derivation in `flake.nix` fetches an exact release zipapp from
|
||||||
|
`github.com/yt-dlp/yt-dlp/releases` and wraps it with `python312` + `ffmpeg`. Both the `archivr` and
|
||||||
|
`archivr-server` wrappers export it as `ARCHIVR_YT_DLP`.
|
||||||
|
2. **A state-dir install** — `archivr yt-dlp update` downloads the latest zipapp and installs it
|
||||||
|
atomically (staged file, then rename) at `<state_dir>/yt-dlp/yt-dlp` with a sibling `.version`
|
||||||
|
sentinel that lets repeat runs skip the download.
|
||||||
|
3. **Whatever is on PATH** — the historical behaviour, and the last-resort fallback.
|
||||||
|
|
||||||
|
`resolve_yt_dlp()` in `downloader/ytdlp.rs` picks between them once per process (cached in a
|
||||||
|
`OnceLock`): `ARCHIVR_YT_DLP_FORCE` wins outright if it points at a real file; otherwise the pinned and
|
||||||
|
state-dir candidates are probed with `--version` and the newest wins — yt-dlp versions are `YYYY.MM.DD`,
|
||||||
|
so plain string ordering is chronological — with exact ties going to the state-dir copy the user
|
||||||
|
deliberately installed. If neither exists, it falls back to bare `yt-dlp`. `archivr yt-dlp status`
|
||||||
|
prints every candidate, its version, and the winner; when the force variable applies, it includes that
|
||||||
|
forced candidate and selects it as the winner.
|
||||||
|
|
||||||
|
Three ways to move the version forward: the weekly `.github/workflows/update-ytdlp.yml` cron (reads the
|
||||||
|
current pin, queries the GitHub releases API, re-hashes with `nix hash file --sri`, rewrites the `ytDlp`
|
||||||
|
block and opens a PR), `archivr yt-dlp update` for one machine, or editing `flake.nix` by hand.
|
||||||
|
|
||||||
## Where To Edit
|
## Where To Edit
|
||||||
|
|
||||||
| Feature kind | Edit here |
|
| Feature kind | Edit here |
|
||||||
|
|
@ -167,6 +253,10 @@ sequenceDiagram
|
||||||
| Archive opening, listing entries, entry detail, runs | `crates/archivr-core/src/archive.rs` |
|
| Archive opening, listing entries, entry detail, runs | `crates/archivr-core/src/archive.rs` |
|
||||||
| Download/save behavior | `crates/archivr-core/src/downloader/` |
|
| Download/save behavior | `crates/archivr-core/src/downloader/` |
|
||||||
| YouTube playlist/channel download, playlist probe, sync mode | `crates/archivr-core/src/downloader/ytdlp.rs` and `capture.rs` |
|
| YouTube playlist/channel download, playlist probe, sync mode | `crates/archivr-core/src/downloader/ytdlp.rs` and `capture.rs` |
|
||||||
|
| Which yt-dlp binary runs (resolver, state dir, version probe) | `crates/archivr-core/src/downloader/ytdlp.rs` |
|
||||||
|
| Pasted-text capture (staging, hashing, MIME allowlist) | `crates/archivr-core/src/downloader/text.rs` and `capture.rs` |
|
||||||
|
| LLM summary providers, prompt, `PROMPT_VERSION`, input building | `crates/archivr-core/src/summarizer.rs` |
|
||||||
|
| `entry_summaries` schema and summary CRUD | `crates/archivr-core/src/database.rs` |
|
||||||
| CLI commands, argument parsing, terminal output | `crates/archivr-cli/src/main.rs` |
|
| CLI commands, argument parsing, terminal output | `crates/archivr-cli/src/main.rs` |
|
||||||
| Server API routes | `crates/archivr-server/src/routes.rs` |
|
| Server API routes | `crates/archivr-server/src/routes.rs` |
|
||||||
| Auth model (users, sessions, tokens, roles) | `crates/archivr-server/src/auth.rs` |
|
| Auth model (users, sessions, tokens, roles) | `crates/archivr-server/src/auth.rs` |
|
||||||
|
|
@ -174,6 +264,9 @@ sequenceDiagram
|
||||||
| Frontend root state + routing | `frontend/src/App.jsx` |
|
| Frontend root state + routing | `frontend/src/App.jsx` |
|
||||||
| Frontend API client | `frontend/src/api.js` |
|
| Frontend API client | `frontend/src/api.js` |
|
||||||
| Frontend components | `frontend/src/components/` |
|
| Frontend components | `frontend/src/components/` |
|
||||||
|
| Summary UI (provider selector, generate, polling) | `frontend/src/components/ContextRail.jsx` |
|
||||||
|
| Text/Markdown entry preview | `frontend/src/components/TextPreview.jsx` |
|
||||||
|
| "Add text" capture row | `frontend/src/components/CaptureDialog.jsx` |
|
||||||
| Frontend styling | `frontend/src/styles.css` |
|
| Frontend styling | `frontend/src/styles.css` |
|
||||||
|
|
||||||
## Practical Feature Rule
|
## Practical Feature Rule
|
||||||
|
|
@ -196,7 +289,8 @@ The server both reads and writes archive data. Capture jobs are asynchronous: `P
|
||||||
|
|
||||||
**Auth model.** A separate `archivr-auth.sqlite` (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are `u32` flags (`GUEST`, `USER`, `ADMIN`, `OWNER`) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is `setup_guard` → `login_rate_limit` → `security_headers`; route families are classified `READ / ADMIN / WRITE / STATIC` in `routes.rs`.
|
**Auth model.** A separate `archivr-auth.sqlite` (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are `u32` flags (`GUEST`, `USER`, `ADMIN`, `OWNER`) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is `setup_guard` → `login_rate_limit` → `security_headers`; route families are classified `READ / ADMIN / WRITE / STATIC` in `routes.rs`.
|
||||||
|
|
||||||
**Search** is client-side filtering over entries the frontend has already fetched.
|
**Search** is server-side free-text filtering over entry fields and the latest completed summary. The summary JSON is
|
||||||
|
searched as text, so generated `tags` participate. Older completed summaries stay searchable while a newer request is
|
||||||
|
pending or failed; rows with no completed summary contribute no summary-derived match.
|
||||||
|
|
||||||
**Admin view** covers mounted archives, users, sessions, and API tokens.
|
**Admin view** covers mounted archives, users, sessions, and API tokens.
|
||||||
|
|
||||||
|
|
|
||||||
3
Cargo.lock
generated
3
Cargo.lock
generated
|
|
@ -97,8 +97,10 @@ dependencies = [
|
||||||
"chrono",
|
"chrono",
|
||||||
"clap",
|
"clap",
|
||||||
"regex",
|
"regex",
|
||||||
|
"reqwest",
|
||||||
"rusqlite",
|
"rusqlite",
|
||||||
"serde_json",
|
"serde_json",
|
||||||
|
"tempfile",
|
||||||
]
|
]
|
||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
|
|
@ -1542,6 +1544,7 @@ version = "1.0.150"
|
||||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||||
checksum = "e8014e44b4736ed0538adeecded0fce2a272f22dc9578a7eb6b2d9993c74cfb9"
|
checksum = "e8014e44b4736ed0538adeecded0fce2a272f22dc9578a7eb6b2d9993c74cfb9"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
|
"indexmap",
|
||||||
"itoa",
|
"itoa",
|
||||||
"memchr",
|
"memchr",
|
||||||
"serde",
|
"serde",
|
||||||
|
|
|
||||||
|
|
@ -19,7 +19,7 @@ hex = "0.4.3"
|
||||||
regex = "1.12.2"
|
regex = "1.12.2"
|
||||||
rusqlite = { version = "0.32.1", features = ["bundled"] }
|
rusqlite = { version = "0.32.1", features = ["bundled"] }
|
||||||
serde = { version = "1.0.228", features = ["derive"] }
|
serde = { version = "1.0.228", features = ["derive"] }
|
||||||
serde_json = "1.0.132"
|
serde_json = { version = "1.0.132", features = ["preserve_order"] }
|
||||||
sha3 = "0.10.8"
|
sha3 = "0.10.8"
|
||||||
tempfile = "3.13.0"
|
tempfile = "3.13.0"
|
||||||
tokio = { version = "1.41.1", features = ["macros", "rt-multi-thread", "net", "fs", "io-util"] }
|
tokio = { version = "1.41.1", features = ["macros", "rt-multi-thread", "net", "fs", "io-util"] }
|
||||||
|
|
|
||||||
|
|
@ -15,3 +15,7 @@ clap.workspace = true
|
||||||
regex.workspace = true
|
regex.workspace = true
|
||||||
rusqlite.workspace = true
|
rusqlite.workspace = true
|
||||||
serde_json.workspace = true
|
serde_json.workspace = true
|
||||||
|
reqwest.workspace = true
|
||||||
|
|
||||||
|
[dev-dependencies]
|
||||||
|
tempfile.workspace = true
|
||||||
|
|
|
||||||
|
|
@ -1,12 +1,27 @@
|
||||||
use anyhow::{Context, Result};
|
use anyhow::{bail, Context, Result};
|
||||||
use archivr_core::{archive, capture::CaptureConfig};
|
use archivr_core::{
|
||||||
|
archive,
|
||||||
|
capture::CaptureConfig,
|
||||||
|
downloader::ytdlp::{
|
||||||
|
forced_yt_dlp, pinned_yt_dlp, probe_version, resolve_yt_dlp, state_dir, state_dir_yt_dlp,
|
||||||
|
},
|
||||||
|
};
|
||||||
use clap::{Parser, Subcommand};
|
use clap::{Parser, Subcommand};
|
||||||
use std::{
|
use std::{
|
||||||
env,
|
env,
|
||||||
path::Path,
|
path::{Path, PathBuf},
|
||||||
process,
|
process,
|
||||||
|
process::Command as ProcCommand,
|
||||||
};
|
};
|
||||||
|
|
||||||
|
/// GitHub release metadata endpoint for the upstream yt-dlp project.
|
||||||
|
const YT_DLP_LATEST_RELEASE: &str =
|
||||||
|
"https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest";
|
||||||
|
|
||||||
|
/// Every python zipapp starts with this shebang; used as a sanity check that we
|
||||||
|
/// downloaded the artifact and not an HTML error page or an LFS pointer.
|
||||||
|
const ZIPAPP_SHEBANG: &[u8] = b"#!/usr/bin/env python3";
|
||||||
|
|
||||||
#[derive(Parser, Debug)]
|
#[derive(Parser, Debug)]
|
||||||
#[command(version, about, long_about = None)]
|
#[command(version, about, long_about = None)]
|
||||||
struct Args {
|
struct Args {
|
||||||
|
|
@ -48,6 +63,25 @@ enum Command {
|
||||||
#[arg(long = "force-with-info-removal")]
|
#[arg(long = "force-with-info-removal")]
|
||||||
force_with_info_removal: bool,
|
force_with_info_removal: bool,
|
||||||
},
|
},
|
||||||
|
|
||||||
|
/// Inspect or update the yt-dlp binary archivr runs
|
||||||
|
#[command(name = "yt-dlp")]
|
||||||
|
YtDlp {
|
||||||
|
#[command(subcommand)]
|
||||||
|
subcmd: YtDlpCmd,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
#[derive(Subcommand, Debug)]
|
||||||
|
enum YtDlpCmd {
|
||||||
|
/// Download the latest yt-dlp zipapp into archivr's state directory
|
||||||
|
Update {
|
||||||
|
/// Install this exact release tag instead of the latest (e.g. 2026.09.15)
|
||||||
|
#[arg(long)]
|
||||||
|
version: Option<String>,
|
||||||
|
},
|
||||||
|
/// Show every yt-dlp candidate, its version, and which one wins
|
||||||
|
Status,
|
||||||
}
|
}
|
||||||
|
|
||||||
fn main() -> Result<()> {
|
fn main() -> Result<()> {
|
||||||
|
|
@ -96,7 +130,217 @@ fn main() -> Result<()> {
|
||||||
);
|
);
|
||||||
|
|
||||||
Ok(())
|
Ok(())
|
||||||
} // _ => eprintln!("Unknown command: {:?}", args.command),
|
}
|
||||||
|
|
||||||
|
Command::YtDlp { subcmd } => match subcmd {
|
||||||
|
YtDlpCmd::Update { version } => yt_dlp_update(version.as_deref()),
|
||||||
|
YtDlpCmd::Status => yt_dlp_status(),
|
||||||
|
},
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
/// Resolves `<state_dir>/yt-dlp/`, erroring out if there is no usable HOME.
|
||||||
|
fn yt_dlp_state_dir() -> Result<PathBuf> {
|
||||||
|
state_dir()
|
||||||
|
.map(|d| d.join("yt-dlp"))
|
||||||
|
.context("could not determine a state directory (is $HOME set?)")
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Formats one `status` row. Missing candidates show an em dash.
|
||||||
|
fn format_status_row(role: &str, path: Option<&Path>, chosen: &Path) -> String {
|
||||||
|
match path {
|
||||||
|
Some(p) => {
|
||||||
|
let version = probe_version(p).unwrap_or_else(|| "—".to_string());
|
||||||
|
let star = if p == chosen { "*" } else { "" };
|
||||||
|
format!("{role}\t{}\t{version}\t{star}", p.display())
|
||||||
|
}
|
||||||
|
None => format!("{role}\t—\t—\t"),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Prints one `status` row. Missing candidates show an em dash.
|
||||||
|
fn status_row(role: &str, path: Option<&Path>, chosen: &Path) {
|
||||||
|
println!("{}", format_status_row(role, path, chosen));
|
||||||
|
}
|
||||||
|
|
||||||
|
fn yt_dlp_status() -> Result<()> {
|
||||||
|
let chosen = resolve_yt_dlp();
|
||||||
|
|
||||||
|
println!("role\tpath\tversion\tchosen");
|
||||||
|
status_row(
|
||||||
|
"force (ARCHIVR_YT_DLP_FORCE)",
|
||||||
|
forced_yt_dlp().as_deref(),
|
||||||
|
&chosen,
|
||||||
|
);
|
||||||
|
status_row("env (ARCHIVR_YT_DLP)", pinned_yt_dlp().as_deref(), &chosen);
|
||||||
|
|
||||||
|
// Show the state-dir slot even when empty, so users can see where an
|
||||||
|
// `archivr yt-dlp update` would land.
|
||||||
|
let state_candidate = state_dir_yt_dlp().filter(|p| p.is_file());
|
||||||
|
status_row("state-dir", state_candidate.as_deref(), &chosen);
|
||||||
|
|
||||||
|
status_row(
|
||||||
|
"path-fallback (yt-dlp)",
|
||||||
|
Some(Path::new("yt-dlp")),
|
||||||
|
&chosen,
|
||||||
|
);
|
||||||
|
|
||||||
|
if let Ok(dir) = yt_dlp_state_dir() {
|
||||||
|
if state_dir_yt_dlp().is_none_or(|p| !p.is_file()) {
|
||||||
|
println!("\nNo state-dir install yet; `archivr yt-dlp update` would write to {}", dir.join("yt-dlp").display());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Asks the GitHub API for the newest yt-dlp release tag.
|
||||||
|
fn latest_yt_dlp_version(client: &reqwest::blocking::Client) -> Result<String> {
|
||||||
|
let body = client
|
||||||
|
.get(YT_DLP_LATEST_RELEASE)
|
||||||
|
.send()
|
||||||
|
.context("failed to reach the GitHub releases API")?
|
||||||
|
.error_for_status()
|
||||||
|
.context("GitHub releases API returned an error")?
|
||||||
|
.text()
|
||||||
|
.context("failed to read the GitHub releases API response")?;
|
||||||
|
|
||||||
|
let json: serde_json::Value =
|
||||||
|
serde_json::from_str(&body).context("GitHub releases API returned invalid JSON")?;
|
||||||
|
|
||||||
|
json.get("tag_name")
|
||||||
|
.and_then(|t| t.as_str())
|
||||||
|
.map(str::to_string)
|
||||||
|
.context("GitHub releases API response had no tag_name")
|
||||||
|
}
|
||||||
|
|
||||||
|
fn yt_dlp_update(requested_version: Option<&str>) -> Result<()> {
|
||||||
|
let dir = yt_dlp_state_dir()?;
|
||||||
|
let target = dir.join("yt-dlp");
|
||||||
|
let staging = dir.join("yt-dlp.new");
|
||||||
|
let version_file = dir.join(".version");
|
||||||
|
|
||||||
|
let client = reqwest::blocking::Client::builder()
|
||||||
|
.user_agent(concat!("archivr-cli/", env!("CARGO_PKG_VERSION")))
|
||||||
|
.build()
|
||||||
|
.context("failed to build an HTTP client")?;
|
||||||
|
|
||||||
|
let version = match requested_version {
|
||||||
|
Some(v) => v.to_string(),
|
||||||
|
None => latest_yt_dlp_version(&client)?,
|
||||||
|
};
|
||||||
|
|
||||||
|
// The sibling .version file is what lets us skip a ~3MB download on a
|
||||||
|
// no-op update; the binary itself is a zipapp with no cheap version probe
|
||||||
|
// that doesn't cost a python startup.
|
||||||
|
let installed = std::fs::read_to_string(&version_file).ok();
|
||||||
|
if target.is_file() && installed.as_deref().map(str::trim) == Some(version.as_str()) {
|
||||||
|
println!("yt-dlp {version} is already installed at {}", target.display());
|
||||||
|
return Ok(());
|
||||||
|
}
|
||||||
|
|
||||||
|
println!("Downloading yt-dlp {version}…");
|
||||||
|
let url = format!("https://github.com/yt-dlp/yt-dlp/releases/download/{version}/yt-dlp");
|
||||||
|
let bytes = client
|
||||||
|
.get(&url)
|
||||||
|
.send()
|
||||||
|
.with_context(|| format!("failed to download {url}"))?
|
||||||
|
.error_for_status()
|
||||||
|
.with_context(|| format!("download failed — is {version} a real release tag?"))?
|
||||||
|
.bytes()
|
||||||
|
.context("failed to read the downloaded yt-dlp body")?;
|
||||||
|
|
||||||
|
if !bytes.starts_with(ZIPAPP_SHEBANG) {
|
||||||
|
bail!(
|
||||||
|
"downloaded artifact from {url} is not a python zipapp \
|
||||||
|
(expected it to start with `{}`) — refusing to install it",
|
||||||
|
String::from_utf8_lossy(ZIPAPP_SHEBANG)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
std::fs::create_dir_all(&dir)
|
||||||
|
.with_context(|| format!("failed to create {}", dir.display()))?;
|
||||||
|
std::fs::write(&staging, &bytes)
|
||||||
|
.with_context(|| format!("failed to write {}", staging.display()))?;
|
||||||
|
|
||||||
|
#[cfg(unix)]
|
||||||
|
{
|
||||||
|
use std::os::unix::fs::PermissionsExt;
|
||||||
|
std::fs::set_permissions(&staging, std::fs::Permissions::from_mode(0o755))
|
||||||
|
.with_context(|| format!("failed to chmod +x {}", staging.display()))?;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Atomic swap: a concurrently-running archivr sees either the whole old
|
||||||
|
// binary or the whole new one, never a half-written file.
|
||||||
|
std::fs::rename(&staging, &target)
|
||||||
|
.with_context(|| format!("failed to install {}", target.display()))?;
|
||||||
|
std::fs::write(&version_file, format!("{version}\n"))
|
||||||
|
.with_context(|| format!("failed to record version in {}", version_file.display()))?;
|
||||||
|
|
||||||
|
// The zipapp is python source, not a native binary — installing it on a
|
||||||
|
// host without python3 is legal (the server may run under a nix wrapper
|
||||||
|
// with its own PATH) but worth flagging loudly.
|
||||||
|
let has_python = ProcCommand::new("python3")
|
||||||
|
.arg("--version")
|
||||||
|
.output()
|
||||||
|
.map(|o| o.status.success())
|
||||||
|
.unwrap_or(false);
|
||||||
|
if !has_python {
|
||||||
|
eprintln!(
|
||||||
|
"warning: python3 was not found on PATH — the yt-dlp zipapp just installed \
|
||||||
|
at {} will not run until python3 is available",
|
||||||
|
target.display()
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
println!("Installed yt-dlp {version} to {}", target.display());
|
||||||
|
println!("archivr will now prefer it whenever it is newer than the pinned binary (ARCHIVR_YT_DLP).");
|
||||||
|
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::format_status_row;
|
||||||
|
use archivr_core::downloader::ytdlp::{
|
||||||
|
forced_yt_dlp, resolve_yt_dlp_uncached, YT_DLP_FORCE_ENV,
|
||||||
|
};
|
||||||
|
use std::path::Path;
|
||||||
|
|
||||||
|
fn fake_yt_dlp(path: &Path, version: &str) {
|
||||||
|
std::fs::create_dir_all(path.parent().unwrap()).unwrap();
|
||||||
|
std::fs::write(path, format!("#!/bin/sh\necho {version}\n")).unwrap();
|
||||||
|
#[cfg(unix)]
|
||||||
|
{
|
||||||
|
use std::os::unix::fs::PermissionsExt;
|
||||||
|
std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o755)).unwrap();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn forced_candidate_is_rendered_and_selected() {
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
let forced = tmp.path().join("forced/yt-dlp");
|
||||||
|
fake_yt_dlp(&forced, "2020.01.01");
|
||||||
|
unsafe { std::env::set_var(YT_DLP_FORCE_ENV, &forced) };
|
||||||
|
|
||||||
|
let candidate = forced_yt_dlp();
|
||||||
|
assert_eq!(candidate.as_deref(), Some(forced.as_path()));
|
||||||
|
let chosen = resolve_yt_dlp_uncached();
|
||||||
|
assert_eq!(chosen, forced);
|
||||||
|
assert_eq!(
|
||||||
|
format_status_row(
|
||||||
|
"force (ARCHIVR_YT_DLP_FORCE)",
|
||||||
|
candidate.as_deref(),
|
||||||
|
&chosen,
|
||||||
|
),
|
||||||
|
format!(
|
||||||
|
"force (ARCHIVR_YT_DLP_FORCE)\t{}\t2020.01.01\t*",
|
||||||
|
forced.display()
|
||||||
|
)
|
||||||
|
);
|
||||||
|
|
||||||
|
unsafe { std::env::remove_var(YT_DLP_FORCE_ENV) };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
|
||||||
|
|
@ -36,6 +36,14 @@ pub struct EntrySummary {
|
||||||
pub cacheable_bytes: i64,
|
pub cacheable_bytes: i64,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// One stored LLM summary, as exposed over the API.
|
||||||
|
///
|
||||||
|
/// Aliased rather than redefined: the DB row is already the exact shape the
|
||||||
|
/// frontend needs, and a second near-identical struct would only add a mapping
|
||||||
|
/// step to keep in sync. The `View` name exists because `EntrySummary` in this
|
||||||
|
/// module is the *entry listing* row, an unrelated thing.
|
||||||
|
pub use crate::database::EntrySummaryRecord as EntrySummaryView;
|
||||||
|
|
||||||
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
|
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
|
||||||
pub struct EntryDetail {
|
pub struct EntryDetail {
|
||||||
pub summary: EntrySummary,
|
pub summary: EntrySummary,
|
||||||
|
|
@ -43,6 +51,12 @@ pub struct EntryDetail {
|
||||||
pub source_metadata_json: String,
|
pub source_metadata_json: String,
|
||||||
pub display_metadata_json: Option<String>,
|
pub display_metadata_json: Option<String>,
|
||||||
pub artifacts: Vec<EntryArtifactSummary>,
|
pub artifacts: Vec<EntryArtifactSummary>,
|
||||||
|
/// Most recent completed summary for this entry. Always `None` on a fresh
|
||||||
|
/// capture — summarization is manual.
|
||||||
|
pub latest_summary: Option<EntrySummaryView>,
|
||||||
|
/// Latest non-completed generation attempt, kept separate so a replacement
|
||||||
|
/// never displaces readable completed content.
|
||||||
|
pub summary_attempt: Option<EntrySummaryView>,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
|
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
|
||||||
|
|
@ -343,12 +357,17 @@ pub fn get_entry_detail(
|
||||||
})?
|
})?
|
||||||
.collect::<rusqlite::Result<Vec<_>>>()?;
|
.collect::<rusqlite::Result<Vec<_>>>()?;
|
||||||
|
|
||||||
|
let latest_summary = database::latest_completed_entry_summary(conn, entry_id)?;
|
||||||
|
let summary_attempt = database::latest_entry_summary_attempt(conn, entry_id)?;
|
||||||
|
|
||||||
Ok(Some(EntryDetail {
|
Ok(Some(EntryDetail {
|
||||||
summary,
|
summary,
|
||||||
structured_root_relpath,
|
structured_root_relpath,
|
||||||
source_metadata_json,
|
source_metadata_json,
|
||||||
display_metadata_json,
|
display_metadata_json,
|
||||||
artifacts,
|
artifacts,
|
||||||
|
latest_summary,
|
||||||
|
summary_attempt,
|
||||||
}))
|
}))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -759,7 +778,14 @@ pub fn search_entries(
|
||||||
sql.push_str(&format!(
|
sql.push_str(&format!(
|
||||||
" AND (LOWER(e.title) LIKE ?{n} OR LOWER(si.canonical_url) LIKE ?{n} \
|
" AND (LOWER(e.title) LIKE ?{n} OR LOWER(si.canonical_url) LIKE ?{n} \
|
||||||
OR LOWER(e.entry_uid) LIKE ?{n} OR LOWER(e.source_kind) LIKE ?{n} \
|
OR LOWER(e.entry_uid) LIKE ?{n} OR LOWER(e.source_kind) LIKE ?{n} \
|
||||||
OR LOWER(e.entity_kind) LIKE ?{n} OR LOWER(e.visibility) LIKE ?{n})"
|
OR LOWER(e.entity_kind) LIKE ?{n} OR LOWER(e.visibility) LIKE ?{n} \
|
||||||
|
OR LOWER(COALESCE((\
|
||||||
|
SELECT s.summary_text FROM entry_summaries s \
|
||||||
|
WHERE s.entry_id = e.id AND s.status = 'completed' \
|
||||||
|
AND s.summary_text IS NOT NULL \
|
||||||
|
ORDER BY s.completed_at DESC, s.updated_at DESC, s.id DESC \
|
||||||
|
LIMIT 1\
|
||||||
|
), '')) LIKE ?{n})"
|
||||||
));
|
));
|
||||||
params.push(term);
|
params.push(term);
|
||||||
}
|
}
|
||||||
|
|
@ -1483,6 +1509,192 @@ mod tests {
|
||||||
assert_eq!(results.len(), 1);
|
assert_eq!(results.len(), 1);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn entry_id_by_title(conn: &rusqlite::Connection, title: &str) -> i64 {
|
||||||
|
conn.query_row(
|
||||||
|
"SELECT id FROM archived_entries WHERE title = ?1",
|
||||||
|
[title],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn complete_summary(
|
||||||
|
conn: &rusqlite::Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
cache_key: &str,
|
||||||
|
summary_text: &str,
|
||||||
|
) -> String {
|
||||||
|
let uid = database::upsert_pending_entry_summary(
|
||||||
|
conn,
|
||||||
|
entry_id,
|
||||||
|
"test_provider",
|
||||||
|
None,
|
||||||
|
"v1",
|
||||||
|
cache_key,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
database::update_entry_summary_status(conn, &uid, "completed", Some(summary_text), None)
|
||||||
|
.unwrap();
|
||||||
|
uid
|
||||||
|
}
|
||||||
|
|
||||||
|
fn set_summary_timestamps(
|
||||||
|
conn: &rusqlite::Connection,
|
||||||
|
summary_uid: &str,
|
||||||
|
timestamp: &str,
|
||||||
|
) {
|
||||||
|
conn.execute(
|
||||||
|
"UPDATE entry_summaries SET completed_at = ?1, updated_at = ?1 WHERE summary_uid = ?2",
|
||||||
|
rusqlite::params![timestamp, summary_uid],
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn search_summary_json_tags_match_and_unrelated_text_is_absent() {
|
||||||
|
let conn = make_test_db_with_entries();
|
||||||
|
let entry_id = entry_id_by_title(&conn, "Resume Templates");
|
||||||
|
complete_summary(
|
||||||
|
&conn,
|
||||||
|
entry_id,
|
||||||
|
"tags",
|
||||||
|
r#"{"tags":["skincare","dermatology"]}"#,
|
||||||
|
);
|
||||||
|
|
||||||
|
let matches = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
q: Some("skincare".to_string()),
|
||||||
|
..Default::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(matches.len(), 1);
|
||||||
|
assert_eq!(matches[0].title.as_deref(), Some("Resume Templates"));
|
||||||
|
|
||||||
|
let unrelated = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
q: Some("neurology".to_string()),
|
||||||
|
..Default::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert!(unrelated.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn search_summary_uses_only_the_newest_completed_row() {
|
||||||
|
let conn = make_test_db_with_entries();
|
||||||
|
let entry_id = entry_id_by_title(&conn, "Resume Templates");
|
||||||
|
let older = complete_summary(&conn, entry_id, "older", "legacy-skincare-term");
|
||||||
|
set_summary_timestamps(&conn, &older, "2026-01-01T00:00:00Z");
|
||||||
|
let newer = complete_summary(&conn, entry_id, "newer", "current-dermatology-term");
|
||||||
|
set_summary_timestamps(&conn, &newer, "2026-02-01T00:00:00Z");
|
||||||
|
|
||||||
|
let old_matches = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
q: Some("legacy-skincare-term".to_string()),
|
||||||
|
..Default::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert!(old_matches.is_empty());
|
||||||
|
|
||||||
|
let current_matches = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
q: Some("current-dermatology-term".to_string()),
|
||||||
|
..Default::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(current_matches.len(), 1);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn search_summary_keeps_latest_completed_when_newer_rows_are_pending_or_failed() {
|
||||||
|
let conn = make_test_db_with_entries();
|
||||||
|
let entry_id = entry_id_by_title(&conn, "Resume Templates");
|
||||||
|
let completed = complete_summary(&conn, entry_id, "completed", "retained-skincare-term");
|
||||||
|
set_summary_timestamps(&conn, &completed, "2026-01-01T00:00:00Z");
|
||||||
|
|
||||||
|
let pending = database::upsert_pending_entry_summary(
|
||||||
|
&conn,
|
||||||
|
entry_id,
|
||||||
|
"test_provider",
|
||||||
|
None,
|
||||||
|
"v1",
|
||||||
|
"pending",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
set_summary_timestamps(&conn, &pending, "2026-03-01T00:00:00Z");
|
||||||
|
let failed = database::upsert_pending_entry_summary(
|
||||||
|
&conn,
|
||||||
|
entry_id,
|
||||||
|
"test_provider",
|
||||||
|
None,
|
||||||
|
"v1",
|
||||||
|
"failed",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
database::update_entry_summary_status(&conn, &failed, "failed", None, Some("boom")).unwrap();
|
||||||
|
set_summary_timestamps(&conn, &failed, "2026-04-01T00:00:00Z");
|
||||||
|
|
||||||
|
let matches = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
q: Some("retained-skincare-term".to_string()),
|
||||||
|
..Default::default()
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(matches.len(), 1);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn search_summary_preserves_prefix_collection_and_visibility_scope() {
|
||||||
|
let conn = make_test_db_with_entries();
|
||||||
|
let entry_id = entry_id_by_title(&conn, "Polymarket tweet");
|
||||||
|
complete_summary(&conn, entry_id, "scoped", "scoped-skincare-term");
|
||||||
|
let tag = create_tag(&conn, "/summary-scope").unwrap();
|
||||||
|
database::assign_entry_to_tag(
|
||||||
|
&conn,
|
||||||
|
entry_id,
|
||||||
|
database::get_tag_by_uid(&conn, &tag.tag_uid).unwrap().unwrap().id,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
let collection = database::create_collection(&conn, "Summary scope", "summary-scope", 2, false)
|
||||||
|
.unwrap();
|
||||||
|
database::add_entry_to_collection(&conn, collection.id, entry_id, 2).unwrap();
|
||||||
|
|
||||||
|
let query = SearchEntriesQuery {
|
||||||
|
q: Some("scoped-skincare-term".to_string()),
|
||||||
|
source_kind: Some("x".to_string()),
|
||||||
|
entity_kind: Some("tweet".to_string()),
|
||||||
|
url: Some("x.com".to_string()),
|
||||||
|
title: Some("polymarket".to_string()),
|
||||||
|
after: Some("2020-01-01T00:00:00Z".to_string()),
|
||||||
|
before: Some("9999-01-01T00:00:00Z".to_string()),
|
||||||
|
tag: Some("/summary-scope".to_string()),
|
||||||
|
caller_bits: 1,
|
||||||
|
collection_id: Some(collection.id),
|
||||||
|
};
|
||||||
|
assert!(search_entries(&conn, &query).unwrap().is_empty());
|
||||||
|
|
||||||
|
let matches = search_entries(
|
||||||
|
&conn,
|
||||||
|
&SearchEntriesQuery {
|
||||||
|
caller_bits: 2,
|
||||||
|
..query
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(matches.len(), 1);
|
||||||
|
assert_eq!(matches[0].title.as_deref(), Some("Polymarket tweet"));
|
||||||
|
}
|
||||||
|
|
||||||
// ---- tag API tests ----
|
// ---- tag API tests ----
|
||||||
|
|
||||||
fn make_tag_test_db() -> (rusqlite::Connection, i64, i64) {
|
fn make_tag_test_db() -> (rusqlite::Connection, i64, i64) {
|
||||||
|
|
|
||||||
|
|
@ -951,7 +951,10 @@ fn register_tweet_artifacts(
|
||||||
})?;
|
})?;
|
||||||
for (role, raw_relpath) in tweet_raw_artifacts(&json_str)? {
|
for (role, raw_relpath) in tweet_raw_artifacts(&json_str)? {
|
||||||
let raw_path = PathBuf::from(&raw_relpath);
|
let raw_path = PathBuf::from(&raw_relpath);
|
||||||
let blob = blob_record_for_raw_relpath(store_path, &raw_path)?;
|
let mut blob = blob_record_for_raw_relpath(store_path, &raw_path)?;
|
||||||
|
if role == "media" {
|
||||||
|
blob.mime_type = tweet_media_image_mime(blob.extension.as_deref());
|
||||||
|
}
|
||||||
let blob_id = database::upsert_blob(conn, &blob)?;
|
let blob_id = database::upsert_blob(conn, &blob)?;
|
||||||
database::add_entry_artifact(
|
database::add_entry_artifact(
|
||||||
conn,
|
conn,
|
||||||
|
|
@ -1030,6 +1033,22 @@ fn record_tweet_entry(
|
||||||
Ok(entry)
|
Ok(entry)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Trusted image MIME types emitted by the X downloader's media paths.
|
||||||
|
///
|
||||||
|
/// Tweet JSON has no MIME field for ordinary downloaded media. Restricting this
|
||||||
|
/// inference to the explicit image extensions keeps binary video/audio and
|
||||||
|
/// unknown extensions out of multimodal summary input.
|
||||||
|
fn tweet_media_image_mime(extension: Option<&str>) -> Option<String> {
|
||||||
|
match extension?.to_ascii_lowercase().as_str() {
|
||||||
|
"jpg" | "jpeg" => Some("image/jpeg".to_string()),
|
||||||
|
"png" => Some("image/png".to_string()),
|
||||||
|
"webp" => Some("image/webp".to_string()),
|
||||||
|
"gif" => Some("image/gif".to_string()),
|
||||||
|
"avif" => Some("image/avif".to_string()),
|
||||||
|
_ => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
fn tweet_raw_artifacts(tweet_json: &str) -> Result<Vec<(String, String)>> {
|
fn tweet_raw_artifacts(tweet_json: &str) -> Result<Vec<(String, String)>> {
|
||||||
let regex = regex::Regex::new(r#""(avatar_local_path|local_path)": "([^"\n]+)""#)?;
|
let regex = regex::Regex::new(r#""(avatar_local_path|local_path)": "([^"\n]+)""#)?;
|
||||||
let mut seen = HashSet::new();
|
let mut seen = HashSet::new();
|
||||||
|
|
@ -1562,8 +1581,8 @@ pub fn perform_capture(
|
||||||
let (rewritten, fonts) =
|
let (rewritten, fonts) =
|
||||||
downloader::font_extractor::extract_and_rewrite(&content, store_path, aid)
|
downloader::font_extractor::extract_and_rewrite(&content, store_path, aid)
|
||||||
.unwrap_or_else(|_| (content.clone(), vec![])); // non-fatal
|
.unwrap_or_else(|_| (content.clone(), vec![])); // non-fatal
|
||||||
// Extract title after font-stripping so the title tag is not buried
|
// Extract title after font-stripping so the title tag is not buried
|
||||||
// behind multi-MB embedded font data that would exceed the 256 KiB window.
|
// behind multi-MB embedded font data that would exceed the 256 KiB window.
|
||||||
let title = downloader::singlefile::extract_html_title_str(&rewritten);
|
let title = downloader::singlefile::extract_html_title_str(&rewritten);
|
||||||
fs::write(&temp_html, rewritten.as_bytes())
|
fs::write(&temp_html, rewritten.as_bytes())
|
||||||
.with_context(|| "failed to write rewritten HTML")?;
|
.with_context(|| "failed to write rewritten HTML")?;
|
||||||
|
|
@ -1918,6 +1937,171 @@ pub fn perform_capture(
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Archives user-supplied plain text or Markdown content.
|
||||||
|
///
|
||||||
|
/// # Arguments
|
||||||
|
/// * `archive_paths` - Path configuration for the archive
|
||||||
|
/// * `title` - User-supplied title (non-empty, trimmed, capped at 500 chars)
|
||||||
|
/// * `body` - Text content (non-empty, capped at 2 MiB)
|
||||||
|
/// * `mime` - MIME type: "text/plain" or "text/markdown"
|
||||||
|
/// * `archive_id` - Optional archive ID (used for job tracking if provided)
|
||||||
|
///
|
||||||
|
/// # Returns
|
||||||
|
/// * `CaptureResult` with the run UID and status
|
||||||
|
///
|
||||||
|
/// # Errors
|
||||||
|
/// * Empty or oversized title/body
|
||||||
|
/// * Unsupported MIME type
|
||||||
|
/// * Database or file system errors
|
||||||
|
pub fn perform_text_capture(
|
||||||
|
archive_paths: &ArchivePaths,
|
||||||
|
title: &str,
|
||||||
|
body: &str,
|
||||||
|
mime: &str,
|
||||||
|
_archive_id: Option<&str>,
|
||||||
|
) -> Result<CaptureResult> {
|
||||||
|
// Validate title
|
||||||
|
let title = title.trim();
|
||||||
|
if title.is_empty() {
|
||||||
|
anyhow::bail!("title must not be empty");
|
||||||
|
}
|
||||||
|
if title.len() > 500 {
|
||||||
|
anyhow::bail!("title must not exceed 500 characters");
|
||||||
|
}
|
||||||
|
|
||||||
|
// Validate body
|
||||||
|
if body.trim().is_empty() {
|
||||||
|
anyhow::bail!("body must not be empty");
|
||||||
|
}
|
||||||
|
if body.len() > 2 * 1024 * 1024 {
|
||||||
|
anyhow::bail!("body must not exceed 2 MiB");
|
||||||
|
}
|
||||||
|
|
||||||
|
// Validate MIME type
|
||||||
|
if mime != "text/plain" && mime != "text/markdown" {
|
||||||
|
anyhow::bail!("unsupported MIME type: {mime}. Must be 'text/plain' or 'text/markdown'");
|
||||||
|
}
|
||||||
|
|
||||||
|
// Generate timestamp
|
||||||
|
let timestamp = format!(
|
||||||
|
"{}-{}",
|
||||||
|
Local::now().format("%Y-%m-%dT%H-%M-%S%.3f"),
|
||||||
|
Uuid::new_v4().simple(),
|
||||||
|
);
|
||||||
|
let store_path = &archive_paths.store_path;
|
||||||
|
|
||||||
|
// Initialize database
|
||||||
|
let conn = database::open_or_initialize(&archive_paths.archive_path)?;
|
||||||
|
let user_id = database::ensure_default_user(&conn)?;
|
||||||
|
|
||||||
|
// Create run and item
|
||||||
|
let run = database::create_archive_run(&conn, user_id, 1)?;
|
||||||
|
let source_kind = "text";
|
||||||
|
let entity_kind = "document";
|
||||||
|
let item = database::create_archive_run_item(
|
||||||
|
&conn,
|
||||||
|
run.id,
|
||||||
|
None,
|
||||||
|
0,
|
||||||
|
&format!("text:{}", title),
|
||||||
|
None,
|
||||||
|
source_kind,
|
||||||
|
entity_kind,
|
||||||
|
)?;
|
||||||
|
|
||||||
|
// Stage the text content
|
||||||
|
let staged_text = downloader::text::save(body.as_bytes(), mime, store_path, ×tamp)?;
|
||||||
|
|
||||||
|
// Check if hash already exists
|
||||||
|
let file_extension = format!(".{}", staged_text.extension);
|
||||||
|
let hash_exists = hash_exists(&staged_text.hash, &file_extension, store_path)?;
|
||||||
|
|
||||||
|
if !hash_exists {
|
||||||
|
// Move staged file to raw storage
|
||||||
|
move_temp_to_raw(&staged_text.staged_path, &staged_text.hash, store_path)?;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Clean up temp directory
|
||||||
|
let _ = fs::remove_dir_all(store_path.join("temp").join(×tamp));
|
||||||
|
|
||||||
|
// Create blob record
|
||||||
|
let raw_relpath = raw_relative_path_from_hash(&staged_text.hash, &file_extension)?;
|
||||||
|
let blob = database::BlobRecord {
|
||||||
|
sha256: staged_text.hash.clone(),
|
||||||
|
byte_size: staged_text.byte_size as i64,
|
||||||
|
mime_type: Some(mime.to_string()),
|
||||||
|
extension: Some(staged_text.extension.clone()),
|
||||||
|
raw_relpath: path_to_store_string(&raw_relpath),
|
||||||
|
};
|
||||||
|
let blob_id = database::upsert_blob(&conn, &blob)?;
|
||||||
|
|
||||||
|
// Create source identity
|
||||||
|
let canonical_locator = format!("text:{}", staged_text.hash);
|
||||||
|
let source_identity_id = database::upsert_source_identity(
|
||||||
|
&conn,
|
||||||
|
source_kind,
|
||||||
|
entity_kind,
|
||||||
|
None,
|
||||||
|
None,
|
||||||
|
&canonical_locator,
|
||||||
|
)?;
|
||||||
|
|
||||||
|
// Create entry
|
||||||
|
let entry = database::create_archived_entry(
|
||||||
|
&conn,
|
||||||
|
&database::NewEntry {
|
||||||
|
source_identity_id,
|
||||||
|
archive_run_id: run.id,
|
||||||
|
parent_entry_id: None,
|
||||||
|
root_entry_id: None,
|
||||||
|
created_by_user_id: user_id,
|
||||||
|
owned_by_user_id: user_id,
|
||||||
|
source_kind: source_kind.to_string(),
|
||||||
|
entity_kind: entity_kind.to_string(),
|
||||||
|
title: Some(title.to_string()),
|
||||||
|
visibility: "private".to_string(),
|
||||||
|
representation_kind: "text".to_string(),
|
||||||
|
source_metadata_json: json!({
|
||||||
|
"requested_locator": format!("text:{}", title),
|
||||||
|
"canonical_locator": canonical_locator,
|
||||||
|
"mime_type": mime
|
||||||
|
})
|
||||||
|
.to_string(),
|
||||||
|
display_metadata_json: None,
|
||||||
|
},
|
||||||
|
)?;
|
||||||
|
|
||||||
|
// Create structured root directory
|
||||||
|
create_structured_root(store_path, &entry)?;
|
||||||
|
|
||||||
|
// Create primary_media artifact
|
||||||
|
database::add_entry_artifact(
|
||||||
|
&conn,
|
||||||
|
&database::NewArtifact {
|
||||||
|
entry_id: entry.id,
|
||||||
|
artifact_role: "primary_media".to_string(),
|
||||||
|
storage_area: "raw".to_string(),
|
||||||
|
relpath: blob.raw_relpath,
|
||||||
|
blob_id: Some(blob_id),
|
||||||
|
logical_path: None,
|
||||||
|
metadata_json: None,
|
||||||
|
},
|
||||||
|
)?;
|
||||||
|
|
||||||
|
// Complete the run item
|
||||||
|
database::complete_archive_run_item(&conn, item.id, entry.id)?;
|
||||||
|
database::refresh_entry_cached_bytes(&conn, entry.id)?;
|
||||||
|
database::finish_archive_run(&conn, run.id)?;
|
||||||
|
|
||||||
|
Ok(CaptureResult {
|
||||||
|
run_uid: run.run_uid.clone(),
|
||||||
|
status: "completed".to_string(),
|
||||||
|
completed_child_count: 0,
|
||||||
|
ublock_skipped: false,
|
||||||
|
cookie_ext_skipped: false,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
/// Result of a tweet re-archive operation.
|
/// Result of a tweet re-archive operation.
|
||||||
#[derive(Debug, serde::Serialize)]
|
#[derive(Debug, serde::Serialize)]
|
||||||
pub struct RearchiveResult {
|
pub struct RearchiveResult {
|
||||||
|
|
@ -2628,6 +2812,325 @@ mod tests {
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_markdown() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
|
||||||
|
// Initialize archive structure
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(
|
||||||
|
archive_path.join("store_path"),
|
||||||
|
store_path.to_str().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path: archive_path.clone(),
|
||||||
|
store_path: store_path.clone(),
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
let title = "My Markdown Note";
|
||||||
|
let body = "# Heading\n\nSome **bold** text.";
|
||||||
|
let mime = "text/markdown";
|
||||||
|
|
||||||
|
let result = perform_text_capture(&archive_paths, title, body, mime, None).unwrap();
|
||||||
|
|
||||||
|
assert_eq!(result.status, "completed");
|
||||||
|
assert_eq!(result.completed_child_count, 0);
|
||||||
|
assert!(!result.ublock_skipped);
|
||||||
|
assert!(!result.cookie_ext_skipped);
|
||||||
|
|
||||||
|
// Verify entry was created
|
||||||
|
let conn = database::open_or_initialize(&archive_path).unwrap();
|
||||||
|
let default_coll_id = database::ensure_default_collection(&conn).unwrap();
|
||||||
|
let entries =
|
||||||
|
archive::list_entries_for_collection(&conn, default_coll_id, 0xFFFFFFFF).unwrap();
|
||||||
|
assert_eq!(entries.len(), 1);
|
||||||
|
let entry = &entries[0];
|
||||||
|
assert_eq!(entry.title, Some(title.to_string()));
|
||||||
|
assert_eq!(entry.source_kind, "text");
|
||||||
|
assert_eq!(entry.entity_kind, "document");
|
||||||
|
|
||||||
|
// Clean up
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_plain() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-plain-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
|
||||||
|
// Initialize archive structure
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(
|
||||||
|
archive_path.join("store_path"),
|
||||||
|
store_path.to_str().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path: archive_path.clone(),
|
||||||
|
store_path: store_path.clone(),
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
let title = "Plain Text Note";
|
||||||
|
let body = "Just plain text content.";
|
||||||
|
let mime = "text/plain";
|
||||||
|
|
||||||
|
let result = perform_text_capture(&archive_paths, title, body, mime, None).unwrap();
|
||||||
|
|
||||||
|
assert_eq!(result.status, "completed");
|
||||||
|
|
||||||
|
// Verify entry was created
|
||||||
|
let conn = database::open_or_initialize(&archive_path).unwrap();
|
||||||
|
let default_coll_id = database::ensure_default_collection(&conn).unwrap();
|
||||||
|
let entries =
|
||||||
|
archive::list_entries_for_collection(&conn, default_coll_id, 0xFFFFFFFF).unwrap();
|
||||||
|
assert_eq!(entries.len(), 1);
|
||||||
|
let entry = &entries[0];
|
||||||
|
assert_eq!(entry.title, Some(title.to_string()));
|
||||||
|
|
||||||
|
// Clean up
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_preserves_intentional_whitespace_in_stored_artifact() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-whitespace-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(archive_path.join("store_path"), store_path.to_str().unwrap()).unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path: archive_path.clone(),
|
||||||
|
store_path: store_path.clone(),
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
let body = " \n# Heading\n\nContent with a final newline\n\t \n";
|
||||||
|
|
||||||
|
perform_text_capture(
|
||||||
|
&archive_paths,
|
||||||
|
"Whitespace Note",
|
||||||
|
body,
|
||||||
|
"text/markdown",
|
||||||
|
None,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let conn = database::open_or_initialize(&archive_path).unwrap();
|
||||||
|
let raw_relpath: String = conn
|
||||||
|
.query_row(
|
||||||
|
"SELECT b.raw_relpath
|
||||||
|
FROM entry_artifacts ea
|
||||||
|
JOIN blobs b ON b.id = ea.blob_id
|
||||||
|
WHERE ea.artifact_role = 'primary_media'",
|
||||||
|
[],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(fs::read(store_path.join(raw_relpath)).unwrap(), body.as_bytes());
|
||||||
|
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_hides_synthetic_url_but_reuses_source_identity() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-identity-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(archive_path.join("store_path"), store_path.to_str().unwrap()).unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path: archive_path.clone(),
|
||||||
|
store_path,
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
let body = "Same body, same text identity.";
|
||||||
|
|
||||||
|
perform_text_capture(&archive_paths, "First title", body, "text/plain", None).unwrap();
|
||||||
|
perform_text_capture(&archive_paths, "Second title", body, "text/plain", None).unwrap();
|
||||||
|
|
||||||
|
let conn = database::open_or_initialize(&archive_path).unwrap();
|
||||||
|
let default_coll_id = database::ensure_default_collection(&conn).unwrap();
|
||||||
|
let entries = archive::list_entries_for_collection(&conn, default_coll_id, 0xFFFFFFFF).unwrap();
|
||||||
|
assert_eq!(entries.len(), 2);
|
||||||
|
assert!(entries.iter().all(|entry| entry.original_url.is_none()));
|
||||||
|
|
||||||
|
let expected_locator = format!("text:{}", crate::hash::hash_bytes(body.as_bytes()));
|
||||||
|
let (canonical_url, normalized_locator): (Option<String>, String) = conn
|
||||||
|
.query_row(
|
||||||
|
"SELECT canonical_url, normalized_locator
|
||||||
|
FROM source_identities
|
||||||
|
WHERE source_kind = 'text' AND normalized_locator = ?1",
|
||||||
|
[&expected_locator],
|
||||||
|
|row| Ok((row.get(0)?, row.get(1)?)),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(canonical_url, None);
|
||||||
|
assert_eq!(normalized_locator, expected_locator);
|
||||||
|
|
||||||
|
let source_identity_count: i64 = conn
|
||||||
|
.query_row(
|
||||||
|
"SELECT COUNT(*) FROM source_identities WHERE source_kind = 'text'",
|
||||||
|
[],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(source_identity_count, 1);
|
||||||
|
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_rejects_empty_title() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-empty-title-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(
|
||||||
|
archive_path.join("store_path"),
|
||||||
|
store_path.to_str().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path,
|
||||||
|
store_path,
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
let result = perform_text_capture(&archive_paths, "", "Some body", "text/plain", None);
|
||||||
|
assert!(result.is_err());
|
||||||
|
assert!(result
|
||||||
|
.unwrap_err()
|
||||||
|
.to_string()
|
||||||
|
.contains("title must not be empty"));
|
||||||
|
|
||||||
|
// Clean up
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_rejects_empty_body() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-empty-body-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(
|
||||||
|
archive_path.join("store_path"),
|
||||||
|
store_path.to_str().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path,
|
||||||
|
store_path,
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
let result = perform_text_capture(&archive_paths, "Some Title", "", "text/plain", None);
|
||||||
|
assert!(result.is_err());
|
||||||
|
assert!(result
|
||||||
|
.unwrap_err()
|
||||||
|
.to_string()
|
||||||
|
.contains("body must not be empty"));
|
||||||
|
|
||||||
|
// Clean up
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_text_capture_rejects_bad_mime() {
|
||||||
|
let base_path = env::temp_dir().join(format!(
|
||||||
|
"archivr-text-bad-mime-test-{}",
|
||||||
|
Local::now().format("%Y%m%d%H%M%S%3f")
|
||||||
|
));
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
fs::create_dir_all(&base_path).unwrap();
|
||||||
|
|
||||||
|
let store_path = base_path.join("store");
|
||||||
|
let archive_path = base_path.join(".archivr");
|
||||||
|
archive::initialize_store_directories(&store_path).unwrap();
|
||||||
|
fs::create_dir_all(&archive_path).unwrap();
|
||||||
|
fs::write(archive_path.join("name"), "test-archive").unwrap();
|
||||||
|
fs::write(
|
||||||
|
archive_path.join("store_path"),
|
||||||
|
store_path.to_str().unwrap(),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
let archive_paths = ArchivePaths {
|
||||||
|
archive_path,
|
||||||
|
store_path,
|
||||||
|
name: "test-archive".to_string(),
|
||||||
|
};
|
||||||
|
|
||||||
|
let result = perform_text_capture(&archive_paths, "Title", "Body", "text/html", None);
|
||||||
|
assert!(result.is_err());
|
||||||
|
assert!(result
|
||||||
|
.unwrap_err()
|
||||||
|
.to_string()
|
||||||
|
.contains("unsupported MIME type"));
|
||||||
|
|
||||||
|
// Clean up
|
||||||
|
let _ = fs::remove_dir_all(&base_path);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_initialize_store_directories() {
|
fn test_initialize_store_directories() {
|
||||||
let store_path = env::temp_dir().join(format!(
|
let store_path = env::temp_dir().join(format!(
|
||||||
|
|
@ -2648,12 +3151,15 @@ mod tests {
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_record_tweet_entry_links_json_and_raw_artifacts() {
|
fn test_record_tweet_entry_links_json_and_raw_artifacts() {
|
||||||
let store_path = env::temp_dir().join(format!(
|
let temp = tempfile::tempdir().unwrap();
|
||||||
"archivr-tweet-db-test-{}",
|
let archive_paths = archive::initialize_archive(
|
||||||
Local::now().format("%Y%m%d%H%M%S%3f")
|
temp.path(),
|
||||||
));
|
&temp.path().join("store"),
|
||||||
let _ = fs::remove_dir_all(&store_path);
|
"Tweet summary test",
|
||||||
archive::initialize_store_directories(&store_path).unwrap();
|
false,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
let store_path = &archive_paths.store_path;
|
||||||
fs::create_dir_all(store_path.join("raw").join("a").join("b")).unwrap();
|
fs::create_dir_all(store_path.join("raw").join("a").join("b")).unwrap();
|
||||||
fs::create_dir_all(store_path.join("raw").join("c").join("d")).unwrap();
|
fs::create_dir_all(store_path.join("raw").join("c").join("d")).unwrap();
|
||||||
fs::write(
|
fs::write(
|
||||||
|
|
@ -2670,21 +3176,21 @@ mod tests {
|
||||||
.join("raw")
|
.join("raw")
|
||||||
.join("c")
|
.join("c")
|
||||||
.join("d")
|
.join("d")
|
||||||
.join("cdef01.mp4"),
|
.join("cdef01.jpg"),
|
||||||
b"media",
|
b"media",
|
||||||
)
|
)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
fs::write(
|
fs::write(
|
||||||
store_path.join("raw_tweets").join("tweet-123.json"),
|
store_path.join("raw_tweets").join("tweet-123.json"),
|
||||||
r#"{
|
r#"{
|
||||||
|
"full_text": "Tweet body for summary selection.",
|
||||||
"author": { "avatar_local_path": "raw/a/b/abcdef.jpg" },
|
"author": { "avatar_local_path": "raw/a/b/abcdef.jpg" },
|
||||||
"entities": { "media": [{ "local_path": "raw/c/d/cdef01.mp4" }] }
|
"entities": { "media": [{ "local_path": "raw/c/d/cdef01.jpg" }] }
|
||||||
}"#,
|
}"#,
|
||||||
)
|
)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
|
|
||||||
let conn = rusqlite::Connection::open_in_memory().unwrap();
|
let conn = database::open_or_initialize(&archive_paths.archive_path).unwrap();
|
||||||
database::initialize_schema(&conn).unwrap();
|
|
||||||
let user_id = database::ensure_default_user(&conn).unwrap();
|
let user_id = database::ensure_default_user(&conn).unwrap();
|
||||||
let run = database::create_archive_run(&conn, user_id, 1).unwrap();
|
let run = database::create_archive_run(&conn, user_id, 1).unwrap();
|
||||||
let item = database::create_archive_run_item(
|
let item = database::create_archive_run_item(
|
||||||
|
|
@ -2733,10 +3239,26 @@ mod tests {
|
||||||
|
|
||||||
assert_eq!(artifact_count, 3);
|
assert_eq!(artifact_count, 3);
|
||||||
assert_eq!(blob_count, 2);
|
assert_eq!(blob_count, 2);
|
||||||
|
let media_mime: Option<String> = conn
|
||||||
|
.query_row(
|
||||||
|
"SELECT b.mime_type FROM entry_artifacts ea JOIN blobs b ON b.id = ea.blob_id WHERE ea.entry_id = ?1 AND ea.artifact_role = 'media'",
|
||||||
|
[entry.id],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(media_mime.as_deref(), Some("image/jpeg"));
|
||||||
|
let summary_input = crate::summarizer::build_summary_input(
|
||||||
|
&archive_paths,
|
||||||
|
&entry.entry_uid,
|
||||||
|
crate::summarizer::SummaryBuildOptions {
|
||||||
|
include_images: true,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(summary_input.request.images.len(), 1);
|
||||||
|
assert_eq!(summary_input.request.images[0].mime_type, "image/jpeg");
|
||||||
assert_eq!(run_status, "completed");
|
assert_eq!(run_status, "completed");
|
||||||
assert!(store_path.join(&entry.structured_root_relpath).is_dir());
|
assert!(store_path.join(&entry.structured_root_relpath).is_dir());
|
||||||
|
|
||||||
let _ = fs::remove_dir_all(store_path);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
mod title_tests {
|
mod title_tests {
|
||||||
|
|
|
||||||
|
|
@ -110,6 +110,30 @@ pub struct CaptureJobRecord {
|
||||||
pub updated_at: String,
|
pub updated_at: String,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// One row of `entry_summaries` — a regenerable LLM summary of an entry.
|
||||||
|
///
|
||||||
|
/// `provider_model` is stored as `''` (not NULL) when a provider has no explicit
|
||||||
|
/// model, because SQLite treats NULLs as distinct inside a UNIQUE index and a
|
||||||
|
/// NULL model would defeat the `(entry_id, provider_kind, provider_model,
|
||||||
|
/// prompt_version, input_sha256)` dedupe key. Readers map `''` back to `None`.
|
||||||
|
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
|
||||||
|
pub struct EntrySummaryRecord {
|
||||||
|
pub summary_uid: String,
|
||||||
|
pub entry_uid: String,
|
||||||
|
pub provider_kind: String,
|
||||||
|
/// Model selected by the provider after resolving the requested cache-key alias.
|
||||||
|
pub resolved_model: Option<String>,
|
||||||
|
pub provider_model: Option<String>,
|
||||||
|
pub prompt_version: String,
|
||||||
|
pub input_sha256: String,
|
||||||
|
pub status: String,
|
||||||
|
pub summary_text: Option<String>,
|
||||||
|
pub error_text: Option<String>,
|
||||||
|
pub created_at: String,
|
||||||
|
pub updated_at: String,
|
||||||
|
pub completed_at: Option<String>,
|
||||||
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, serde::Serialize)]
|
#[derive(Debug, Clone, serde::Serialize)]
|
||||||
pub struct UserSummary {
|
pub struct UserSummary {
|
||||||
pub user_uid: String,
|
pub user_uid: String,
|
||||||
|
|
@ -322,6 +346,26 @@ pub fn initialize_schema(conn: &Connection) -> Result<()> {
|
||||||
created_at TEXT NOT NULL,
|
created_at TEXT NOT NULL,
|
||||||
updated_at TEXT NOT NULL
|
updated_at TEXT NOT NULL
|
||||||
);
|
);
|
||||||
|
CREATE TABLE IF NOT EXISTS entry_summaries (
|
||||||
|
id INTEGER PRIMARY KEY,
|
||||||
|
summary_uid TEXT NOT NULL UNIQUE,
|
||||||
|
entry_id INTEGER NOT NULL REFERENCES archived_entries(id) ON DELETE CASCADE,
|
||||||
|
provider_kind TEXT NOT NULL,
|
||||||
|
provider_model TEXT NOT NULL DEFAULT '',
|
||||||
|
resolved_model TEXT,
|
||||||
|
prompt_version TEXT NOT NULL,
|
||||||
|
input_sha256 TEXT NOT NULL,
|
||||||
|
status TEXT NOT NULL CHECK(status IN ('pending','running','completed','failed')),
|
||||||
|
summary_text TEXT,
|
||||||
|
error_text TEXT,
|
||||||
|
created_at TEXT NOT NULL,
|
||||||
|
updated_at TEXT NOT NULL,
|
||||||
|
completed_at TEXT
|
||||||
|
);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_entry_summaries_entry_updated
|
||||||
|
ON entry_summaries(entry_id, updated_at DESC);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_entry_summaries_cache_lookup
|
||||||
|
ON entry_summaries(entry_id, provider_kind, provider_model, prompt_version, input_sha256, status);
|
||||||
CREATE INDEX IF NOT EXISTS idx_capture_jobs_status ON capture_jobs(status);
|
CREATE INDEX IF NOT EXISTS idx_capture_jobs_status ON capture_jobs(status);
|
||||||
CREATE INDEX IF NOT EXISTS idx_archive_run_items_run_id ON archive_run_items(run_id);
|
CREATE INDEX IF NOT EXISTS idx_archive_run_items_run_id ON archive_run_items(run_id);
|
||||||
CREATE INDEX IF NOT EXISTS idx_archived_entries_source_identity_id ON archived_entries(source_identity_id);
|
CREATE INDEX IF NOT EXISTS idx_archived_entries_source_identity_id ON archived_entries(source_identity_id);
|
||||||
|
|
@ -436,6 +480,12 @@ pub fn initialize_schema(conn: &Connection) -> Result<()> {
|
||||||
// Migration: add notes_json column to existing capture_jobs tables.
|
// Migration: add notes_json column to existing capture_jobs tables.
|
||||||
// Silently ignored when the column already exists (idempotent).
|
// Silently ignored when the column already exists (idempotent).
|
||||||
let _ = conn.execute("ALTER TABLE capture_jobs ADD COLUMN notes_json TEXT", []);
|
let _ = conn.execute("ALTER TABLE capture_jobs ADD COLUMN notes_json TEXT", []);
|
||||||
|
// Provider responses may resolve a requested alias to a concrete model.
|
||||||
|
// Keep that display-only value outside the cache key.
|
||||||
|
let _ = conn.execute(
|
||||||
|
"ALTER TABLE entry_summaries ADD COLUMN resolved_model TEXT",
|
||||||
|
[],
|
||||||
|
);
|
||||||
// Migration: add requires_auth column to existing collections tables.
|
// Migration: add requires_auth column to existing collections tables.
|
||||||
// Silently ignored when the column already exists (idempotent).
|
// Silently ignored when the column already exists (idempotent).
|
||||||
let _ = conn.execute(
|
let _ = conn.execute(
|
||||||
|
|
@ -443,6 +493,54 @@ pub fn initialize_schema(conn: &Connection) -> Result<()> {
|
||||||
[],
|
[],
|
||||||
);
|
);
|
||||||
|
|
||||||
|
// Summary attempts used to be unique by cache key, which meant forced
|
||||||
|
// regeneration erased the last completed result. Rebuild that small table
|
||||||
|
// without the cache-key constraint while retaining all existing rows.
|
||||||
|
let summary_table_sql: Option<String> = conn
|
||||||
|
.query_row(
|
||||||
|
"SELECT sql FROM sqlite_master WHERE type = 'table' AND name = 'entry_summaries'",
|
||||||
|
[],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.optional()?;
|
||||||
|
if summary_table_sql.as_deref().is_some_and(|sql| {
|
||||||
|
sql.contains(
|
||||||
|
"UNIQUE(entry_id, provider_kind, provider_model, prompt_version, input_sha256)",
|
||||||
|
)
|
||||||
|
}) {
|
||||||
|
conn.execute_batch(
|
||||||
|
"BEGIN;
|
||||||
|
CREATE TABLE entry_summaries_rebuilt (
|
||||||
|
id INTEGER PRIMARY KEY,
|
||||||
|
summary_uid TEXT NOT NULL UNIQUE,
|
||||||
|
entry_id INTEGER NOT NULL REFERENCES archived_entries(id) ON DELETE CASCADE,
|
||||||
|
provider_kind TEXT NOT NULL,
|
||||||
|
provider_model TEXT NOT NULL DEFAULT '',
|
||||||
|
resolved_model TEXT,
|
||||||
|
prompt_version TEXT NOT NULL,
|
||||||
|
input_sha256 TEXT NOT NULL,
|
||||||
|
status TEXT NOT NULL CHECK(status IN ('pending','running','completed','failed')),
|
||||||
|
summary_text TEXT,
|
||||||
|
error_text TEXT,
|
||||||
|
created_at TEXT NOT NULL,
|
||||||
|
updated_at TEXT NOT NULL,
|
||||||
|
completed_at TEXT
|
||||||
|
);
|
||||||
|
INSERT INTO entry_summaries_rebuilt
|
||||||
|
SELECT id, summary_uid, entry_id, provider_kind, COALESCE(provider_model, ''),
|
||||||
|
resolved_model, prompt_version, input_sha256, status, summary_text, error_text,
|
||||||
|
created_at, updated_at, completed_at
|
||||||
|
FROM entry_summaries;
|
||||||
|
DROP TABLE entry_summaries;
|
||||||
|
ALTER TABLE entry_summaries_rebuilt RENAME TO entry_summaries;
|
||||||
|
CREATE INDEX idx_entry_summaries_entry_updated
|
||||||
|
ON entry_summaries(entry_id, updated_at DESC);
|
||||||
|
CREATE INDEX idx_entry_summaries_cache_lookup
|
||||||
|
ON entry_summaries(entry_id, provider_kind, provider_model, prompt_version, input_sha256, status);
|
||||||
|
COMMIT;",
|
||||||
|
)?;
|
||||||
|
}
|
||||||
|
|
||||||
Ok(())
|
Ok(())
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -1354,6 +1452,265 @@ pub fn fail_stalled_capture_jobs(conn: &Connection) -> Result<usize> {
|
||||||
Ok(n)
|
Ok(n)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Marks summary attempts left pending or running by a previous process as
|
||||||
|
/// failed. The message is deliberately stable and does not expose internals.
|
||||||
|
pub fn fail_stalled_entry_summaries(conn: &Connection) -> Result<usize> {
|
||||||
|
let now = now_timestamp();
|
||||||
|
conn.execute(
|
||||||
|
"UPDATE entry_summaries
|
||||||
|
SET status = 'failed',
|
||||||
|
error_text = 'Summary generation was interrupted by a server restart.',
|
||||||
|
updated_at = ?1
|
||||||
|
WHERE status IN ('pending', 'running')",
|
||||||
|
[now],
|
||||||
|
)
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── entry_summaries ────────────────────────────────────────────────────────
|
||||||
|
//
|
||||||
|
// Summaries are a regenerable child record of an entry, never a column on
|
||||||
|
// `archived_entries`: an entry can carry several (one per provider/model/prompt
|
||||||
|
// version), and any of them can be thrown away and recomputed. Generation is
|
||||||
|
// manual-only — nothing in `capture.rs` writes here.
|
||||||
|
|
||||||
|
/// `SELECT` list shared by every `entry_summaries` read, so all readers build
|
||||||
|
/// an identical `EntrySummaryRecord` from the same column ordering.
|
||||||
|
const ENTRY_SUMMARY_COLS: &str =
|
||||||
|
"SELECT s.summary_uid, e.entry_uid, s.provider_kind, s.provider_model, s.resolved_model,
|
||||||
|
s.prompt_version, s.input_sha256, s.status, s.summary_text,
|
||||||
|
s.error_text, s.created_at, s.updated_at, s.completed_at
|
||||||
|
FROM entry_summaries s
|
||||||
|
JOIN archived_entries e ON e.id = s.entry_id";
|
||||||
|
|
||||||
|
fn map_entry_summary(row: &rusqlite::Row<'_>) -> rusqlite::Result<EntrySummaryRecord> {
|
||||||
|
Ok(EntrySummaryRecord {
|
||||||
|
summary_uid: row.get(0)?,
|
||||||
|
entry_uid: row.get(1)?,
|
||||||
|
provider_kind: row.get(2)?,
|
||||||
|
// '' is the stored stand-in for "this provider has no model"; see the
|
||||||
|
// doc comment on EntrySummaryRecord for why it is not NULL.
|
||||||
|
provider_model: row.get::<_, Option<String>>(3)?.filter(|m| !m.is_empty()),
|
||||||
|
resolved_model: row.get(4)?,
|
||||||
|
prompt_version: row.get(5)?,
|
||||||
|
input_sha256: row.get(6)?,
|
||||||
|
status: row.get(7)?,
|
||||||
|
summary_text: row.get(8)?,
|
||||||
|
error_text: row.get(9)?,
|
||||||
|
created_at: row.get(10)?,
|
||||||
|
updated_at: row.get(11)?,
|
||||||
|
completed_at: row.get(12)?,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Resolves `entry_uid` to its integer primary key. `Ok(None)` if absent.
|
||||||
|
pub fn entry_id_for_uid(conn: &Connection, entry_uid: &str) -> Result<Option<i64>> {
|
||||||
|
conn.query_row(
|
||||||
|
"SELECT id FROM archived_entries WHERE entry_uid = ?1",
|
||||||
|
[entry_uid],
|
||||||
|
|row| row.get(0),
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Creates a fresh pending summary attempt for one cache key.
|
||||||
|
///
|
||||||
|
/// Attempts are intentionally not unique by cache key: a forced regeneration
|
||||||
|
/// must leave an older completed result available while the new attempt runs.
|
||||||
|
pub fn upsert_pending_entry_summary(
|
||||||
|
conn: &Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
provider_kind: &str,
|
||||||
|
provider_model: Option<&str>,
|
||||||
|
prompt_version: &str,
|
||||||
|
input_sha256: &str,
|
||||||
|
) -> Result<String> {
|
||||||
|
let summary_uid = format!("sum_{}", &Uuid::new_v4().simple().to_string()[..10]);
|
||||||
|
let now = now_timestamp();
|
||||||
|
let model = provider_model.unwrap_or("");
|
||||||
|
conn.execute(
|
||||||
|
"INSERT INTO entry_summaries
|
||||||
|
(summary_uid, entry_id, provider_kind, provider_model, prompt_version,
|
||||||
|
input_sha256, status, summary_text, error_text, created_at, updated_at, completed_at)
|
||||||
|
VALUES (?1, ?2, ?3, ?4, ?5, ?6, 'pending', NULL, NULL, ?7, ?7, NULL)",
|
||||||
|
rusqlite::params![
|
||||||
|
summary_uid,
|
||||||
|
entry_id,
|
||||||
|
provider_kind,
|
||||||
|
model,
|
||||||
|
prompt_version,
|
||||||
|
input_sha256,
|
||||||
|
now
|
||||||
|
],
|
||||||
|
)?;
|
||||||
|
Ok(summary_uid)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Completes a summary while retaining the provider's concrete response model
|
||||||
|
/// for display. The requested provider model remains the cache-key identity.
|
||||||
|
pub fn update_entry_summary_completed(
|
||||||
|
conn: &Connection,
|
||||||
|
summary_uid: &str,
|
||||||
|
summary_text: &str,
|
||||||
|
resolved_model: Option<&str>,
|
||||||
|
) -> Result<()> {
|
||||||
|
let now = now_timestamp();
|
||||||
|
conn.execute(
|
||||||
|
"UPDATE entry_summaries
|
||||||
|
SET status = 'completed', summary_text = ?1, error_text = NULL,
|
||||||
|
resolved_model = ?2, completed_at = ?3, updated_at = ?3
|
||||||
|
WHERE summary_uid = ?4",
|
||||||
|
rusqlite::params![summary_text, resolved_model, now, summary_uid],
|
||||||
|
)?;
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Moves a summary row through `running` → `completed` / `failed`.
|
||||||
|
/// `completed_at` is stamped only on the terminal `completed` transition.
|
||||||
|
pub fn update_entry_summary_status(
|
||||||
|
conn: &Connection,
|
||||||
|
summary_uid: &str,
|
||||||
|
status: &str,
|
||||||
|
summary_text: Option<&str>,
|
||||||
|
error_text: Option<&str>,
|
||||||
|
) -> Result<()> {
|
||||||
|
let now = now_timestamp();
|
||||||
|
let completed_at = if status == "completed" {
|
||||||
|
Some(now.clone())
|
||||||
|
} else {
|
||||||
|
None
|
||||||
|
};
|
||||||
|
conn.execute(
|
||||||
|
"UPDATE entry_summaries
|
||||||
|
SET status = ?1, summary_text = ?2, error_text = ?3,
|
||||||
|
completed_at = ?4, updated_at = ?5
|
||||||
|
WHERE summary_uid = ?6",
|
||||||
|
rusqlite::params![
|
||||||
|
status,
|
||||||
|
summary_text,
|
||||||
|
error_text,
|
||||||
|
completed_at,
|
||||||
|
now,
|
||||||
|
summary_uid
|
||||||
|
],
|
||||||
|
)?;
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Returns one summary by its public uid.
|
||||||
|
pub fn get_entry_summary_by_uid(
|
||||||
|
conn: &Connection,
|
||||||
|
summary_uid: &str,
|
||||||
|
) -> Result<Option<EntrySummaryRecord>> {
|
||||||
|
conn.query_row(
|
||||||
|
&format!("{ENTRY_SUMMARY_COLS} WHERE s.summary_uid = ?1"),
|
||||||
|
[summary_uid],
|
||||||
|
map_entry_summary,
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Looks up the row for one exact cache key — used to short-circuit a POST when
|
||||||
|
/// a completed summary for identical input already exists and `force` is false.
|
||||||
|
pub fn find_entry_summary(
|
||||||
|
conn: &Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
provider_kind: &str,
|
||||||
|
provider_model: Option<&str>,
|
||||||
|
prompt_version: &str,
|
||||||
|
input_sha256: &str,
|
||||||
|
) -> Result<Option<EntrySummaryRecord>> {
|
||||||
|
conn.query_row(
|
||||||
|
&format!(
|
||||||
|
"{ENTRY_SUMMARY_COLS}
|
||||||
|
WHERE s.entry_id = ?1 AND s.provider_kind = ?2 AND s.provider_model = ?3
|
||||||
|
AND s.prompt_version = ?4 AND s.input_sha256 = ?5
|
||||||
|
ORDER BY CASE WHEN s.status = 'completed' THEN 0 ELSE 1 END,
|
||||||
|
s.completed_at DESC, s.updated_at DESC, s.id DESC"
|
||||||
|
),
|
||||||
|
rusqlite::params![
|
||||||
|
entry_id,
|
||||||
|
provider_kind,
|
||||||
|
provider_model.unwrap_or(""),
|
||||||
|
prompt_version,
|
||||||
|
input_sha256
|
||||||
|
],
|
||||||
|
map_entry_summary,
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Most recently touched summary for an entry, whatever its status.
|
||||||
|
/// Used where a caller explicitly needs the most recent attempt regardless of
|
||||||
|
/// whether it has completed.
|
||||||
|
pub fn latest_entry_summary(
|
||||||
|
conn: &Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
) -> Result<Option<EntrySummaryRecord>> {
|
||||||
|
conn.query_row(
|
||||||
|
&format!(
|
||||||
|
"{ENTRY_SUMMARY_COLS} WHERE s.entry_id = ?1
|
||||||
|
ORDER BY s.updated_at DESC, s.id DESC LIMIT 1"
|
||||||
|
),
|
||||||
|
[entry_id],
|
||||||
|
map_entry_summary,
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The most recent completed summary, excluding in-flight and failed attempts.
|
||||||
|
/// This is the stable summary shown in entry detail and used by free-text search.
|
||||||
|
pub fn latest_completed_entry_summary(
|
||||||
|
conn: &Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
) -> Result<Option<EntrySummaryRecord>> {
|
||||||
|
conn.query_row(
|
||||||
|
&format!(
|
||||||
|
"{ENTRY_SUMMARY_COLS} WHERE s.entry_id = ?1 AND s.status = 'completed'
|
||||||
|
ORDER BY s.completed_at DESC, s.updated_at DESC, s.id DESC LIMIT 1"
|
||||||
|
),
|
||||||
|
[entry_id],
|
||||||
|
map_entry_summary,
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The newest non-completed replacement attempt after the retained completed
|
||||||
|
/// result. This includes failed rows so authenticated callers can show a recent
|
||||||
|
/// failure beside readable content, but suppresses historical failures after a
|
||||||
|
/// newer successful regeneration.
|
||||||
|
pub fn latest_entry_summary_attempt(
|
||||||
|
conn: &Connection,
|
||||||
|
entry_id: i64,
|
||||||
|
) -> Result<Option<EntrySummaryRecord>> {
|
||||||
|
conn.query_row(
|
||||||
|
&format!(
|
||||||
|
"{ENTRY_SUMMARY_COLS} WHERE s.entry_id = ?1 AND s.status != 'completed'
|
||||||
|
AND (
|
||||||
|
NOT EXISTS (
|
||||||
|
SELECT 1 FROM entry_summaries c
|
||||||
|
WHERE c.entry_id = s.entry_id AND c.status = 'completed'
|
||||||
|
)
|
||||||
|
OR (s.updated_at, s.id) > (
|
||||||
|
SELECT c.updated_at, c.id FROM entry_summaries c
|
||||||
|
WHERE c.entry_id = s.entry_id AND c.status = 'completed'
|
||||||
|
ORDER BY c.completed_at DESC, c.updated_at DESC, c.id DESC LIMIT 1
|
||||||
|
)
|
||||||
|
)
|
||||||
|
ORDER BY s.updated_at DESC, s.id DESC LIMIT 1"
|
||||||
|
),
|
||||||
|
[entry_id],
|
||||||
|
map_entry_summary,
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
|
}
|
||||||
|
|
||||||
pub fn create_archive_run(
|
pub fn create_archive_run(
|
||||||
conn: &Connection,
|
conn: &Connection,
|
||||||
created_by_user_id: i64,
|
created_by_user_id: i64,
|
||||||
|
|
@ -1445,7 +1802,11 @@ pub fn finish_archive_run(conn: &Connection, run_id: i64) -> Result<()> {
|
||||||
[run_id],
|
[run_id],
|
||||||
|row| row.get(0),
|
|row| row.get(0),
|
||||||
)?;
|
)?;
|
||||||
let status = if failed_count > 0 { "failed" } else { "completed" };
|
let status = if failed_count > 0 {
|
||||||
|
"failed"
|
||||||
|
} else {
|
||||||
|
"completed"
|
||||||
|
};
|
||||||
conn.execute(
|
conn.execute(
|
||||||
"UPDATE archive_runs SET status = ?1, finished_at = ?2 WHERE id = ?3",
|
"UPDATE archive_runs SET status = ?1, finished_at = ?2 WHERE id = ?3",
|
||||||
params![status, now_timestamp(), run_id],
|
params![status, now_timestamp(), run_id],
|
||||||
|
|
@ -1681,9 +2042,8 @@ pub fn delete_entry(conn: &Connection, entry_uid: &str) -> Result<bool> {
|
||||||
// (no grandchildren), so without `id = ?1` the set would be empty and
|
// (no grandchildren), so without `id = ?1` the set would be empty and
|
||||||
// cascade_cached_bytes_after_subtree_delete would not recalculate shared-blob totals.
|
// cascade_cached_bytes_after_subtree_delete would not recalculate shared-blob totals.
|
||||||
let subtree_ids: Vec<i64> = {
|
let subtree_ids: Vec<i64> = {
|
||||||
let mut stmt = conn.prepare(
|
let mut stmt =
|
||||||
"SELECT id FROM archived_entries WHERE id = ?1 OR root_entry_id = ?1",
|
conn.prepare("SELECT id FROM archived_entries WHERE id = ?1 OR root_entry_id = ?1")?;
|
||||||
)?;
|
|
||||||
stmt.query_map([entry_id], |row| row.get(0))?
|
stmt.query_map([entry_id], |row| row.get(0))?
|
||||||
.collect::<rusqlite::Result<_>>()?
|
.collect::<rusqlite::Result<_>>()?
|
||||||
};
|
};
|
||||||
|
|
@ -2334,7 +2694,6 @@ pub fn visibility_to_bits(visibility: &str) -> u32 {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
/// Returns the id of the '_default_' collection, creating it if absent.
|
/// Returns the id of the '_default_' collection, creating it if absent.
|
||||||
pub fn ensure_default_collection(conn: &Connection) -> Result<i64> {
|
pub fn ensure_default_collection(conn: &Connection) -> Result<i64> {
|
||||||
let now = now_timestamp();
|
let now = now_timestamp();
|
||||||
|
|
@ -2437,16 +2796,20 @@ pub fn get_collection_by_slug(conn: &Connection, slug: &str) -> Result<Option<Co
|
||||||
"SELECT id, collection_uid, name, slug, default_visibility_bits, created_at, requires_auth \
|
"SELECT id, collection_uid, name, slug, default_visibility_bits, created_at, requires_auth \
|
||||||
FROM collections WHERE slug = ?1",
|
FROM collections WHERE slug = ?1",
|
||||||
[slug],
|
[slug],
|
||||||
|row| Ok(CollectionRecord {
|
|row| {
|
||||||
id: row.get(0)?,
|
Ok(CollectionRecord {
|
||||||
collection_uid: row.get(1)?,
|
id: row.get(0)?,
|
||||||
name: row.get(2)?,
|
collection_uid: row.get(1)?,
|
||||||
slug: row.get(3)?,
|
name: row.get(2)?,
|
||||||
default_visibility_bits: row.get::<_, i64>(4)? as u32,
|
slug: row.get(3)?,
|
||||||
created_at: row.get(5)?,
|
default_visibility_bits: row.get::<_, i64>(4)? as u32,
|
||||||
requires_auth: row.get::<_, i64>(6)? != 0,
|
created_at: row.get(5)?,
|
||||||
}),
|
requires_auth: row.get::<_, i64>(6)? != 0,
|
||||||
).optional().map_err(Into::into)
|
})
|
||||||
|
},
|
||||||
|
)
|
||||||
|
.optional()
|
||||||
|
.map_err(Into::into)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Adds an entry to a collection with given visibility_bits. Idempotent (INSERT OR IGNORE).
|
/// Adds an entry to a collection with given visibility_bits. Idempotent (INSERT OR IGNORE).
|
||||||
|
|
@ -2505,7 +2868,12 @@ pub fn get_entry_collection_memberships(
|
||||||
WHERE ce.entry_id = ?1",
|
WHERE ce.entry_id = ?1",
|
||||||
)?;
|
)?;
|
||||||
stmt.query_map([entry_id], |row| {
|
stmt.query_map([entry_id], |row| {
|
||||||
Ok((row.get(0)?, row.get(1)?, row.get(2)?, row.get::<_, i64>(3)? as u32))
|
Ok((
|
||||||
|
row.get(0)?,
|
||||||
|
row.get(1)?,
|
||||||
|
row.get(2)?,
|
||||||
|
row.get::<_, i64>(3)? as u32,
|
||||||
|
))
|
||||||
})?
|
})?
|
||||||
.collect::<Result<_, _>>()
|
.collect::<Result<_, _>>()
|
||||||
.map_err(Into::into)
|
.map_err(Into::into)
|
||||||
|
|
@ -3063,7 +3431,10 @@ mod tests {
|
||||||
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
||||||
)
|
)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
assert_eq!(default_bits, 2, "default collection should start USER-visible");
|
assert_eq!(
|
||||||
|
default_bits, 2,
|
||||||
|
"default collection should start USER-visible"
|
||||||
|
);
|
||||||
|
|
||||||
// Create an entry with visibility = "private" (what capture always passes).
|
// Create an entry with visibility = "private" (what capture always passes).
|
||||||
let entry = create_entry_fixture(&conn, "private", None, None);
|
let entry = create_entry_fixture(&conn, "private", None, None);
|
||||||
|
|
@ -3111,7 +3482,10 @@ mod tests {
|
||||||
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
||||||
)
|
)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
assert_eq!(public_bits, 3, "collection default=public should produce bits=3");
|
assert_eq!(
|
||||||
|
public_bits, 3,
|
||||||
|
"collection default=public should produce bits=3"
|
||||||
|
);
|
||||||
|
|
||||||
// Child entries must NOT use the collection default — they keep
|
// Child entries must NOT use the collection default — they keep
|
||||||
// visibility_to_bits(entry.visibility) so parent-child visibility
|
// visibility_to_bits(entry.visibility) so parent-child visibility
|
||||||
|
|
@ -3124,7 +3498,10 @@ mod tests {
|
||||||
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
|row| row.get::<_, i64>(0).map(|v| v as u32),
|
||||||
)
|
)
|
||||||
.unwrap();
|
.unwrap();
|
||||||
assert_eq!(child_bits, 0, "child entries must use visibility_to_bits, not collection default");
|
assert_eq!(
|
||||||
|
child_bits, 0,
|
||||||
|
"child entries must use visibility_to_bits, not collection default"
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
|
|
@ -4153,27 +4530,414 @@ mod tests {
|
||||||
|
|
||||||
// Root item (parent_item_id IS NULL) — mirrors what record_container_entry does.
|
// Root item (parent_item_id IS NULL) — mirrors what record_container_entry does.
|
||||||
let root_item = create_archive_run_item(
|
let root_item = create_archive_run_item(
|
||||||
&c, run.id, None, 0, "https://example.com/pl", None, "youtube", "playlist",
|
&c,
|
||||||
).unwrap();
|
run.id,
|
||||||
|
None,
|
||||||
|
0,
|
||||||
|
"https://example.com/pl",
|
||||||
|
None,
|
||||||
|
"youtube",
|
||||||
|
"playlist",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
// Child item (parent_item_id IS NOT NULL).
|
// Child item (parent_item_id IS NOT NULL).
|
||||||
let child_item = create_archive_run_item(
|
let child_item = create_archive_run_item(
|
||||||
&c, run.id, Some(root_item.id), 1, "https://example.com/v1", None, "youtube", "video",
|
&c,
|
||||||
).unwrap();
|
run.id,
|
||||||
|
Some(root_item.id),
|
||||||
|
1,
|
||||||
|
"https://example.com/v1",
|
||||||
|
None,
|
||||||
|
"youtube",
|
||||||
|
"video",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
// Complete both — marks archive_runs.completed_count = 2.
|
// Complete both — marks archive_runs.completed_count = 2.
|
||||||
c.execute(
|
c.execute(
|
||||||
"UPDATE archive_run_items SET status = 'completed' WHERE id IN (?1, ?2)",
|
"UPDATE archive_run_items SET status = 'completed' WHERE id IN (?1, ?2)",
|
||||||
rusqlite::params![root_item.id, child_item.id],
|
rusqlite::params![root_item.id, child_item.id],
|
||||||
).unwrap();
|
)
|
||||||
|
.unwrap();
|
||||||
refresh_run_counters(&c, run.id).unwrap();
|
refresh_run_counters(&c, run.id).unwrap();
|
||||||
|
|
||||||
let total: i64 = c.query_row(
|
let total: i64 = c
|
||||||
"SELECT completed_count FROM archive_runs WHERE id = ?1", [run.id], |r| r.get(0),
|
.query_row(
|
||||||
).unwrap();
|
"SELECT completed_count FROM archive_runs WHERE id = ?1",
|
||||||
|
[run.id],
|
||||||
|
|r| r.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
assert_eq!(total, 2, "both items completed: DB counter must be 2");
|
assert_eq!(total, 2, "both items completed: DB counter must be 2");
|
||||||
|
|
||||||
let child_count = get_run_completed_child_count(&c, run.id).unwrap();
|
let child_count = get_run_completed_child_count(&c, run.id).unwrap();
|
||||||
assert_eq!(child_count, 1, "only the child item must be counted");
|
assert_eq!(child_count, 1, "only the child item must be counted");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── entry_summaries ────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn initialize_schema_is_idempotent_for_entry_summaries() {
|
||||||
|
// initialize_schema runs on *every* open_or_initialize, so re-running it
|
||||||
|
// against a populated DB must be a no-op, not an error.
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "abc").unwrap();
|
||||||
|
|
||||||
|
initialize_schema(&c).unwrap();
|
||||||
|
initialize_schema(&c).unwrap();
|
||||||
|
|
||||||
|
let exists: i64 = c
|
||||||
|
.query_row(
|
||||||
|
"SELECT COUNT(*) FROM sqlite_master WHERE type='table' AND name='entry_summaries'",
|
||||||
|
[],
|
||||||
|
|r| r.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(exists, 1);
|
||||||
|
assert!(latest_entry_summary(&c, entry.id).unwrap().is_some());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn initialize_schema_migrates_resolved_model_for_existing_summary_table() {
|
||||||
|
let c = Connection::open_in_memory().unwrap();
|
||||||
|
c.execute_batch(
|
||||||
|
"CREATE TABLE entry_summaries (
|
||||||
|
id INTEGER PRIMARY KEY,
|
||||||
|
summary_uid TEXT NOT NULL UNIQUE,
|
||||||
|
entry_id INTEGER NOT NULL,
|
||||||
|
provider_kind TEXT NOT NULL,
|
||||||
|
provider_model TEXT,
|
||||||
|
prompt_version TEXT NOT NULL,
|
||||||
|
input_sha256 TEXT NOT NULL,
|
||||||
|
status TEXT NOT NULL,
|
||||||
|
summary_text TEXT,
|
||||||
|
error_text TEXT,
|
||||||
|
created_at TEXT NOT NULL,
|
||||||
|
updated_at TEXT NOT NULL,
|
||||||
|
completed_at TEXT,
|
||||||
|
UNIQUE(entry_id, provider_kind, provider_model, prompt_version, input_sha256)
|
||||||
|
);",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
initialize_schema(&c).unwrap();
|
||||||
|
|
||||||
|
let has_resolved_model: i64 = c
|
||||||
|
.query_row(
|
||||||
|
"SELECT COUNT(*) FROM pragma_table_info('entry_summaries') WHERE name = 'resolved_model'",
|
||||||
|
[],
|
||||||
|
|r| r.get(0),
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(has_resolved_model, 1);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn entry_summary_lifecycle_pending_running_completed() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let uid =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "anthropic_http", Some("m1"), "v1", "sha1")
|
||||||
|
.unwrap();
|
||||||
|
assert!(uid.starts_with("sum_"), "got {uid}");
|
||||||
|
assert_eq!(uid.len(), "sum_".len() + 10);
|
||||||
|
|
||||||
|
let rec = get_entry_summary_by_uid(&c, &uid).unwrap().unwrap();
|
||||||
|
assert_eq!(rec.status, "pending");
|
||||||
|
assert_eq!(rec.entry_uid, entry.entry_uid);
|
||||||
|
assert_eq!(rec.provider_model.as_deref(), Some("m1"));
|
||||||
|
assert!(rec.completed_at.is_none());
|
||||||
|
|
||||||
|
update_entry_summary_status(&c, &uid, "running", None, None).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
get_entry_summary_by_uid(&c, &uid).unwrap().unwrap().status,
|
||||||
|
"running"
|
||||||
|
);
|
||||||
|
|
||||||
|
update_entry_summary_status(&c, &uid, "completed", Some("{\"summary\":\"s\"}"), None)
|
||||||
|
.unwrap();
|
||||||
|
let rec = get_entry_summary_by_uid(&c, &uid).unwrap().unwrap();
|
||||||
|
assert_eq!(rec.status, "completed");
|
||||||
|
assert_eq!(rec.summary_text.as_deref(), Some("{\"summary\":\"s\"}"));
|
||||||
|
assert!(rec.completed_at.is_some(), "completed rows must be stamped");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn entry_summary_keeps_requested_model_for_cache_and_resolved_model_for_display() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let uid = upsert_pending_entry_summary(
|
||||||
|
&c,
|
||||||
|
entry.id,
|
||||||
|
"anthropic_http",
|
||||||
|
Some("claude-3-5-sonnet-latest"),
|
||||||
|
"v1",
|
||||||
|
"sha1",
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
|
||||||
|
update_entry_summary_completed(&c, &uid, "summary", Some("claude-3-5-sonnet-20241022"))
|
||||||
|
.unwrap();
|
||||||
|
let rec = get_entry_summary_by_uid(&c, &uid).unwrap().unwrap();
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
rec.provider_model.as_deref(),
|
||||||
|
Some("claude-3-5-sonnet-latest")
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
rec.resolved_model.as_deref(),
|
||||||
|
Some("claude-3-5-sonnet-20241022")
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
find_entry_summary(
|
||||||
|
&c,
|
||||||
|
entry.id,
|
||||||
|
"anthropic_http",
|
||||||
|
Some("claude-3-5-sonnet-latest"),
|
||||||
|
"v1",
|
||||||
|
"sha1",
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap()
|
||||||
|
.summary_uid,
|
||||||
|
uid,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn entry_summary_failure_records_error_and_no_completed_at() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let uid = upsert_pending_entry_summary(&c, entry.id, "codex_cli", None, "v1", "s").unwrap();
|
||||||
|
update_entry_summary_status(&c, &uid, "failed", None, Some("boom")).unwrap();
|
||||||
|
let rec = get_entry_summary_by_uid(&c, &uid).unwrap().unwrap();
|
||||||
|
assert_eq!(rec.status, "failed");
|
||||||
|
assert_eq!(rec.error_text.as_deref(), Some("boom"));
|
||||||
|
assert!(rec.completed_at.is_none());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn regenerating_a_completed_summary_creates_a_new_attempt_and_preserves_completion() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let first =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "sha").unwrap();
|
||||||
|
update_entry_summary_status(&c, &first, "completed", Some("old"), None).unwrap();
|
||||||
|
|
||||||
|
// Regeneration must retain the finished result while a new attempt runs.
|
||||||
|
let second =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "sha").unwrap();
|
||||||
|
assert_ne!(first, second, "regeneration needs a distinct attempt uid");
|
||||||
|
|
||||||
|
let n: i64 = c
|
||||||
|
.query_row("SELECT COUNT(*) FROM entry_summaries", [], |r| r.get(0))
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(n, 2);
|
||||||
|
|
||||||
|
let completed = get_entry_summary_by_uid(&c, &first).unwrap().unwrap();
|
||||||
|
assert_eq!(completed.status, "completed");
|
||||||
|
assert_eq!(completed.summary_text.as_deref(), Some("old"));
|
||||||
|
let rec = get_entry_summary_by_uid(&c, &second).unwrap().unwrap();
|
||||||
|
assert_eq!(rec.status, "pending");
|
||||||
|
update_entry_summary_status(&c, &second, "failed", None, Some("boom")).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
latest_completed_entry_summary(&c, entry.id)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap()
|
||||||
|
.summary_uid,
|
||||||
|
first,
|
||||||
|
"a failed regeneration must not replace the previous completed result"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn latest_summary_attempt_is_separate_from_the_retained_completed_summary() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let completed =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "old").unwrap();
|
||||||
|
update_entry_summary_status(&c, &completed, "completed", Some("previous"), None).unwrap();
|
||||||
|
let pending =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "new").unwrap();
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
latest_completed_entry_summary(&c, entry.id).unwrap().unwrap().summary_uid,
|
||||||
|
completed
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
latest_entry_summary_attempt(&c, entry.id).unwrap().unwrap().summary_uid,
|
||||||
|
pending
|
||||||
|
);
|
||||||
|
|
||||||
|
update_entry_summary_status(&c, &pending, "failed", None, Some("boom")).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
latest_completed_entry_summary(&c, entry.id).unwrap().unwrap().summary_text.as_deref(),
|
||||||
|
Some("previous")
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
latest_entry_summary_attempt(&c, entry.id).unwrap().unwrap().status,
|
||||||
|
"failed"
|
||||||
|
);
|
||||||
|
|
||||||
|
let successful_replacement =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "newer").unwrap();
|
||||||
|
update_entry_summary_status(
|
||||||
|
&c,
|
||||||
|
&successful_replacement,
|
||||||
|
"completed",
|
||||||
|
Some("replacement"),
|
||||||
|
None,
|
||||||
|
)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
latest_completed_entry_summary(&c, entry.id)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap()
|
||||||
|
.summary_uid,
|
||||||
|
successful_replacement
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
latest_entry_summary_attempt(&c, entry.id).unwrap().is_none(),
|
||||||
|
"a failed attempt predating a successful replacement must not remain visible"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn fail_stalled_entry_summaries_marks_pending_and_running_with_restart_message() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let pending =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "a").unwrap();
|
||||||
|
let running =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "codex_cli", None, "v1", "b").unwrap();
|
||||||
|
let completed =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "anthropic_http", None, "v1", "c").unwrap();
|
||||||
|
update_entry_summary_status(&c, &running, "running", None, None).unwrap();
|
||||||
|
update_entry_summary_status(&c, &completed, "completed", Some("done"), None).unwrap();
|
||||||
|
|
||||||
|
assert_eq!(fail_stalled_entry_summaries(&c).unwrap(), 2);
|
||||||
|
for uid in [&pending, &running] {
|
||||||
|
let row = get_entry_summary_by_uid(&c, uid).unwrap().unwrap();
|
||||||
|
assert_eq!(row.status, "failed");
|
||||||
|
assert_eq!(
|
||||||
|
row.error_text.as_deref(),
|
||||||
|
Some("Summary generation was interrupted by a server restart.")
|
||||||
|
);
|
||||||
|
assert!(!row.updated_at.is_empty());
|
||||||
|
}
|
||||||
|
assert_eq!(
|
||||||
|
get_entry_summary_by_uid(&c, &completed)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap()
|
||||||
|
.status,
|
||||||
|
"completed"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn entry_summaries_with_no_model_preserve_none_across_attempts() {
|
||||||
|
// CLI providers have no explicit model, so the stored empty-string
|
||||||
|
// sentinel must always map back to None even when attempts accumulate.
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "sha").unwrap();
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "sha").unwrap();
|
||||||
|
let n: i64 = c
|
||||||
|
.query_row("SELECT COUNT(*) FROM entry_summaries", [], |r| r.get(0))
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(n, 2);
|
||||||
|
// …and it reads back as None, not as an empty-string model.
|
||||||
|
let rec = latest_entry_summary(&c, entry.id).unwrap().unwrap();
|
||||||
|
assert_eq!(rec.provider_model, None);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn different_cache_keys_produce_separate_summaries() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "sha").unwrap();
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "codex_cli", None, "v1", "sha").unwrap();
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v2", "sha").unwrap();
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "other").unwrap();
|
||||||
|
let n: i64 = c
|
||||||
|
.query_row("SELECT COUNT(*) FROM entry_summaries", [], |r| r.get(0))
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(n, 4);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_entry_summary_matches_only_the_exact_cache_key() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let uid =
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "openai_compatible", Some("gpt"), "v1", "s")
|
||||||
|
.unwrap();
|
||||||
|
let hit = find_entry_summary(&c, entry.id, "openai_compatible", Some("gpt"), "v1", "s")
|
||||||
|
.unwrap()
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(hit.summary_uid, uid);
|
||||||
|
assert!(
|
||||||
|
find_entry_summary(
|
||||||
|
&c,
|
||||||
|
entry.id,
|
||||||
|
"openai_compatible",
|
||||||
|
Some("gpt"),
|
||||||
|
"v1",
|
||||||
|
"changed"
|
||||||
|
)
|
||||||
|
.unwrap()
|
||||||
|
.is_none(),
|
||||||
|
"a changed input digest must miss the cache"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn latest_entry_summary_returns_the_most_recently_updated_row() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
let a = upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "a").unwrap();
|
||||||
|
let b = upsert_pending_entry_summary(&c, entry.id, "codex_cli", None, "v1", "b").unwrap();
|
||||||
|
update_entry_summary_status(&c, &a, "completed", Some("first"), None).unwrap();
|
||||||
|
update_entry_summary_status(&c, &b, "completed", Some("second"), None).unwrap();
|
||||||
|
// Same-timestamp ties break on id DESC, so the later insert wins either way.
|
||||||
|
assert_eq!(
|
||||||
|
latest_entry_summary(&c, entry.id)
|
||||||
|
.unwrap()
|
||||||
|
.unwrap()
|
||||||
|
.summary_uid,
|
||||||
|
b
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn latest_entry_summary_is_none_for_an_unsummarized_entry() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
assert!(latest_entry_summary(&c, entry.id).unwrap().is_none());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn deleting_an_entry_cascades_to_its_summaries() {
|
||||||
|
let c = conn();
|
||||||
|
c.pragma_update(None, "foreign_keys", "ON").unwrap();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
upsert_pending_entry_summary(&c, entry.id, "claude_cli", None, "v1", "a").unwrap();
|
||||||
|
assert!(delete_entry(&c, &entry.entry_uid).unwrap());
|
||||||
|
let n: i64 = c
|
||||||
|
.query_row("SELECT COUNT(*) FROM entry_summaries", [], |r| r.get(0))
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(n, 0);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn entry_id_for_uid_resolves_and_misses() {
|
||||||
|
let c = conn();
|
||||||
|
let entry = create_entry_fixture(&c, "private", None, None);
|
||||||
|
assert_eq!(
|
||||||
|
entry_id_for_uid(&c, &entry.entry_uid).unwrap(),
|
||||||
|
Some(entry.id)
|
||||||
|
);
|
||||||
|
assert_eq!(entry_id_for_uid(&c, "ent_nope").unwrap(), None);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -7,3 +7,4 @@ pub mod metadata;
|
||||||
pub mod http;
|
pub mod http;
|
||||||
pub mod singlefile;
|
pub mod singlefile;
|
||||||
pub mod font_extractor;
|
pub mod font_extractor;
|
||||||
|
pub mod text;
|
||||||
|
|
|
||||||
118
crates/archivr-core/src/downloader/text.rs
Normal file
118
crates/archivr-core/src/downloader/text.rs
Normal file
|
|
@ -0,0 +1,118 @@
|
||||||
|
use anyhow::{bail, Result};
|
||||||
|
use std::path::{Path, PathBuf};
|
||||||
|
|
||||||
|
use crate::hash::hash_bytes;
|
||||||
|
|
||||||
|
/// Represents a staged text file ready to be moved into the raw store.
|
||||||
|
#[derive(Debug)]
|
||||||
|
pub struct StagedText {
|
||||||
|
pub staged_path: PathBuf,
|
||||||
|
pub hash: String,
|
||||||
|
pub extension: String,
|
||||||
|
pub byte_size: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Stages a text body (plain or Markdown) in the temp directory and computes its hash.
|
||||||
|
///
|
||||||
|
/// # Arguments
|
||||||
|
/// * `body` - The raw bytes of the text content
|
||||||
|
/// * `mime` - MIME type, must be "text/plain" or "text/markdown"
|
||||||
|
/// * `store_path` - Root store path where temp/ subdirectory will be created
|
||||||
|
/// * `timestamp` - Timestamp string used in the staged file name
|
||||||
|
///
|
||||||
|
/// # Returns
|
||||||
|
/// * `StagedText` with the staged path, hash, extension, and byte size
|
||||||
|
///
|
||||||
|
/// # Errors
|
||||||
|
/// * Rejects MIME types other than "text/plain" or "text/markdown"
|
||||||
|
/// * IO errors during directory creation or file writing
|
||||||
|
pub fn save(body: &[u8], mime: &str, store_path: &Path, timestamp: &str) -> Result<StagedText> {
|
||||||
|
// Validate MIME type
|
||||||
|
let extension = match mime {
|
||||||
|
"text/markdown" => ".md",
|
||||||
|
"text/plain" => ".txt",
|
||||||
|
_ => bail!("unsupported MIME type: {mime}. Must be 'text/plain' or 'text/markdown'"),
|
||||||
|
};
|
||||||
|
|
||||||
|
// Create temp directory
|
||||||
|
let temp_dir = store_path.join("temp").join(timestamp);
|
||||||
|
std::fs::create_dir_all(&temp_dir)?;
|
||||||
|
|
||||||
|
// Stage under temp/<timestamp>/<timestamp><ext>
|
||||||
|
let staged_path = temp_dir.join(format!("{timestamp}{extension}"));
|
||||||
|
|
||||||
|
// Write the content
|
||||||
|
std::fs::write(&staged_path, body)?;
|
||||||
|
|
||||||
|
// Compute SHA3 hash
|
||||||
|
let hash = hash_bytes(body);
|
||||||
|
let byte_size = body.len() as u64;
|
||||||
|
let extension_str = extension.trim_start_matches('.').to_string();
|
||||||
|
|
||||||
|
Ok(StagedText {
|
||||||
|
staged_path,
|
||||||
|
hash,
|
||||||
|
extension: extension_str,
|
||||||
|
byte_size,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
use tempfile::TempDir;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_save_markdown() {
|
||||||
|
let temp_dir = TempDir::new().unwrap();
|
||||||
|
let store_path = temp_dir.path();
|
||||||
|
let content = b"# Hello\n\nThis is markdown.";
|
||||||
|
let mime = "text/markdown";
|
||||||
|
|
||||||
|
let result = save(content, mime, store_path, "2024-01-01T12-00-00.000-abc123").unwrap();
|
||||||
|
|
||||||
|
assert_eq!(result.extension, "md");
|
||||||
|
assert_eq!(result.byte_size, content.len() as u64);
|
||||||
|
assert!(result.staged_path.exists());
|
||||||
|
assert_eq!(std::fs::read(&result.staged_path).unwrap(), content);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_save_plain_text() {
|
||||||
|
let temp_dir = TempDir::new().unwrap();
|
||||||
|
let store_path = temp_dir.path();
|
||||||
|
let content = b"Plain text content";
|
||||||
|
let mime = "text/plain";
|
||||||
|
|
||||||
|
let result = save(content, mime, store_path, "2024-01-01T12-00-00.000-abc123").unwrap();
|
||||||
|
|
||||||
|
assert_eq!(result.extension, "txt");
|
||||||
|
assert_eq!(result.byte_size, content.len() as u64);
|
||||||
|
assert!(result.staged_path.exists());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_save_rejects_unsupported_mime() {
|
||||||
|
let temp_dir = TempDir::new().unwrap();
|
||||||
|
let store_path = temp_dir.path();
|
||||||
|
let content = b"test";
|
||||||
|
|
||||||
|
let result = save(content, "text/html", store_path, "2024-01-01T12-00-00.000-abc123");
|
||||||
|
assert!(result.is_err());
|
||||||
|
assert!(result.unwrap_err().to_string().contains("unsupported MIME type"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_save_hash_is_consistent() {
|
||||||
|
let temp_dir = TempDir::new().unwrap();
|
||||||
|
let store_path = temp_dir.path();
|
||||||
|
let content = b"archivr text";
|
||||||
|
|
||||||
|
let result1 = save(content, "text/plain", store_path, "2024-01-01T12-00-00.000-abc123").unwrap();
|
||||||
|
|
||||||
|
let temp_dir2 = TempDir::new().unwrap();
|
||||||
|
let result2 = save(content, "text/plain", temp_dir2.path(), "2024-01-01T12-00-01.000-def456").unwrap();
|
||||||
|
|
||||||
|
assert_eq!(result1.hash, result2.hash);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
@ -4,6 +4,7 @@ use std::{
|
||||||
env,
|
env,
|
||||||
path::{Path, PathBuf},
|
path::{Path, PathBuf},
|
||||||
process::Command,
|
process::Command,
|
||||||
|
sync::OnceLock,
|
||||||
};
|
};
|
||||||
use uuid::Uuid;
|
use uuid::Uuid;
|
||||||
use serde_json;
|
use serde_json;
|
||||||
|
|
@ -11,6 +12,112 @@ use serde_json;
|
||||||
use crate::downloader::cookies::{domain_from_url, write_netscape_cookie_file};
|
use crate::downloader::cookies::{domain_from_url, write_netscape_cookie_file};
|
||||||
use crate::hash::hash_file;
|
use crate::hash::hash_file;
|
||||||
|
|
||||||
|
/// Env var that force-pins a specific yt-dlp binary, bypassing version comparison.
|
||||||
|
pub const YT_DLP_FORCE_ENV: &str = "ARCHIVR_YT_DLP_FORCE";
|
||||||
|
/// Env var set by the nix flake wrapper, pointing at the pinned yt-dlp.
|
||||||
|
pub const YT_DLP_ENV: &str = "ARCHIVR_YT_DLP";
|
||||||
|
/// Override for the mutable state directory (used by `archivr yt-dlp` and tests).
|
||||||
|
pub const STATE_DIR_ENV: &str = "ARCHIVR_STATE_DIR";
|
||||||
|
|
||||||
|
static RESOLVED_YT_DLP: OnceLock<PathBuf> = OnceLock::new();
|
||||||
|
|
||||||
|
/// Mutable per-user state directory for archivr.
|
||||||
|
///
|
||||||
|
/// `ARCHIVR_STATE_DIR` wins if set. Otherwise this mirrors what `dirs::state_dir()`
|
||||||
|
/// would give us without taking on the dependency: `~/Library/Application Support`
|
||||||
|
/// on macOS, `$XDG_STATE_HOME` (default `~/.local/state`) elsewhere.
|
||||||
|
pub fn state_dir() -> Option<PathBuf> {
|
||||||
|
if let Some(dir) = env::var_os(STATE_DIR_ENV) {
|
||||||
|
if !dir.is_empty() {
|
||||||
|
return Some(PathBuf::from(dir));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
let home = PathBuf::from(env::var_os("HOME").filter(|h| !h.is_empty())?);
|
||||||
|
|
||||||
|
if cfg!(target_os = "macos") {
|
||||||
|
Some(home.join("Library").join("Application Support").join("archivr"))
|
||||||
|
} else {
|
||||||
|
let base = env::var_os("XDG_STATE_HOME")
|
||||||
|
.filter(|d| !d.is_empty())
|
||||||
|
.map(PathBuf::from)
|
||||||
|
.unwrap_or_else(|| home.join(".local").join("state"));
|
||||||
|
Some(base.join("archivr"))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Path of the user-installed (self-updated) yt-dlp inside the state dir.
|
||||||
|
pub fn state_dir_yt_dlp() -> Option<PathBuf> {
|
||||||
|
state_dir().map(|d| d.join("yt-dlp").join("yt-dlp"))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The explicit yt-dlp override, if it points to a file on disk.
|
||||||
|
pub fn forced_yt_dlp() -> Option<PathBuf> {
|
||||||
|
let p = PathBuf::from(env::var_os(YT_DLP_FORCE_ENV).filter(|v| !v.is_empty())?);
|
||||||
|
p.is_file().then_some(p)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The nix-pinned yt-dlp advertised via `ARCHIVR_YT_DLP`, if it exists on disk.
|
||||||
|
pub fn pinned_yt_dlp() -> Option<PathBuf> {
|
||||||
|
let p = PathBuf::from(env::var_os(YT_DLP_ENV).filter(|v| !v.is_empty())?);
|
||||||
|
p.is_file().then_some(p)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Runs `<binary> --version` and returns the trimmed stdout.
|
||||||
|
///
|
||||||
|
/// yt-dlp versions are `YYYY.MM.DD`, so plain string ordering is chronological
|
||||||
|
/// ordering — no semver parsing needed.
|
||||||
|
pub fn probe_version(binary: &Path) -> Option<String> {
|
||||||
|
let out = Command::new(binary).arg("--version").output().ok()?;
|
||||||
|
if !out.status.success() {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let version = String::from_utf8_lossy(&out.stdout).trim().to_string();
|
||||||
|
(!version.is_empty()).then_some(version)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The candidate yt-dlp binaries, in priority order for tie-breaking
|
||||||
|
/// (later entries win ties, so the deliberately-installed state-dir copy is last).
|
||||||
|
pub fn yt_dlp_candidates() -> Vec<(&'static str, PathBuf)> {
|
||||||
|
let mut candidates = Vec::new();
|
||||||
|
if let Some(p) = pinned_yt_dlp() {
|
||||||
|
candidates.push(("env (ARCHIVR_YT_DLP)", p));
|
||||||
|
}
|
||||||
|
if let Some(p) = state_dir_yt_dlp() {
|
||||||
|
if p.is_file() {
|
||||||
|
candidates.push(("state-dir", p));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
candidates
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Picks the yt-dlp binary to run, without consulting the process-wide cache.
|
||||||
|
///
|
||||||
|
/// Priority: `ARCHIVR_YT_DLP_FORCE` > newest of (pinned, state-dir) by version
|
||||||
|
/// string > bare `yt-dlp` (PATH lookup, the historical behaviour).
|
||||||
|
pub fn resolve_yt_dlp_uncached() -> PathBuf {
|
||||||
|
if let Some(forced) = forced_yt_dlp() {
|
||||||
|
return forced;
|
||||||
|
}
|
||||||
|
|
||||||
|
yt_dlp_candidates()
|
||||||
|
.into_iter()
|
||||||
|
.filter_map(|(_, path)| probe_version(&path).map(|v| (v, path)))
|
||||||
|
// `max_by` keeps the *last* maximum, and the state-dir candidate is last,
|
||||||
|
// so an exact version tie resolves in favour of the user's own install.
|
||||||
|
.max_by(|(a, _), (b, _)| a.cmp(b))
|
||||||
|
.map(|(_, path)| path)
|
||||||
|
.unwrap_or_else(|| PathBuf::from("yt-dlp"))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Cached [`resolve_yt_dlp_uncached`] — `--version` is spawned at most once
|
||||||
|
/// per process no matter how many yt-dlp calls the run makes.
|
||||||
|
pub fn resolve_yt_dlp() -> PathBuf {
|
||||||
|
RESOLVED_YT_DLP
|
||||||
|
.get_or_init(resolve_yt_dlp_uncached)
|
||||||
|
.clone()
|
||||||
|
}
|
||||||
|
|
||||||
/// A single item in a flat playlist listing from `fetch_playlist_info`.
|
/// A single item in a flat playlist listing from `fetch_playlist_info`.
|
||||||
#[derive(Debug)]
|
#[derive(Debug)]
|
||||||
pub struct PlaylistItem {
|
pub struct PlaylistItem {
|
||||||
|
|
@ -164,7 +271,7 @@ pub fn download(
|
||||||
) -> Result<(String, String)> {
|
) -> Result<(String, String)> {
|
||||||
println!("Downloading with yt-dlp: {path}");
|
println!("Downloading with yt-dlp: {path}");
|
||||||
|
|
||||||
let ytdlp = env::var("ARCHIVR_YT_DLP").unwrap_or_else(|_| "yt-dlp".to_string());
|
let ytdlp = resolve_yt_dlp();
|
||||||
let is_audio = quality == Some("audio");
|
let is_audio = quality == Some("audio");
|
||||||
|
|
||||||
let temp_dir = store_path.join("temp").join(timestamp);
|
let temp_dir = store_path.join("temp").join(timestamp);
|
||||||
|
|
@ -207,7 +314,7 @@ pub fn download(
|
||||||
.arg("-o")
|
.arg("-o")
|
||||||
.arg(&out_template)
|
.arg(&out_template)
|
||||||
.output()
|
.output()
|
||||||
.with_context(|| format!("failed to spawn {ytdlp} process"));
|
.with_context(|| format!("failed to spawn {} process", ytdlp.display()));
|
||||||
|
|
||||||
// Remove cookie file immediately regardless of outcome.
|
// Remove cookie file immediately regardless of outcome.
|
||||||
if let Some(cf) = &cookie_file {
|
if let Some(cf) = &cookie_file {
|
||||||
|
|
@ -253,7 +360,7 @@ fn find_downloaded_file(temp_dir: &Path, timestamp: &str) -> Result<PathBuf> {
|
||||||
/// On failure (non-zero exit or no stdout), prints the captured stderr
|
/// On failure (non-zero exit or no stdout), prints the captured stderr
|
||||||
/// to stderr (for debugging) then returns `None` so callers can proceed.
|
/// to stderr (for debugging) then returns `None` so callers can proceed.
|
||||||
pub fn fetch_metadata(path: &str, cookies: &HashMap<String, String>) -> Option<String> {
|
pub fn fetch_metadata(path: &str, cookies: &HashMap<String, String>) -> Option<String> {
|
||||||
let ytdlp = std::env::var("ARCHIVR_YT_DLP").unwrap_or_else(|_| "yt-dlp".to_string());
|
let ytdlp = resolve_yt_dlp();
|
||||||
|
|
||||||
// Write a temp cookie file if needed; UUID-named to avoid collisions.
|
// Write a temp cookie file if needed; UUID-named to avoid collisions.
|
||||||
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
||||||
|
|
@ -342,7 +449,7 @@ fn normalize_item_url(
|
||||||
/// Returns an error if yt-dlp fails, the output is not valid JSON, or
|
/// Returns an error if yt-dlp fails, the output is not valid JSON, or
|
||||||
/// the root `_type` is not `"playlist"`.
|
/// the root `_type` is not `"playlist"`.
|
||||||
pub fn fetch_playlist_info(url: &str, cookies: &HashMap<String, String>) -> Result<PlaylistInfo> {
|
pub fn fetch_playlist_info(url: &str, cookies: &HashMap<String, String>) -> Result<PlaylistInfo> {
|
||||||
let ytdlp = std::env::var("ARCHIVR_YT_DLP").unwrap_or_else(|_| "yt-dlp".to_string());
|
let ytdlp = resolve_yt_dlp();
|
||||||
|
|
||||||
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
||||||
let domain = domain_from_url(url);
|
let domain = domain_from_url(url);
|
||||||
|
|
@ -366,7 +473,7 @@ pub fn fetch_playlist_info(url: &str, cookies: &HashMap<String, String>) -> Resu
|
||||||
if let Some(cf) = &cookie_file {
|
if let Some(cf) = &cookie_file {
|
||||||
let _ = std::fs::remove_file(cf);
|
let _ = std::fs::remove_file(cf);
|
||||||
}
|
}
|
||||||
let out = out.with_context(|| format!("failed to spawn {ytdlp}"))?;
|
let out = out.with_context(|| format!("failed to spawn {}", ytdlp.display()))?;
|
||||||
if !out.status.success() {
|
if !out.status.success() {
|
||||||
let stderr = String::from_utf8_lossy(&out.stderr);
|
let stderr = String::from_utf8_lossy(&out.stderr);
|
||||||
bail!("yt-dlp -J --flat-playlist failed for {url}: {stderr}");
|
bail!("yt-dlp -J --flat-playlist failed for {url}: {stderr}");
|
||||||
|
|
@ -425,7 +532,7 @@ pub fn probe_playlist_qualities(
|
||||||
url: &str,
|
url: &str,
|
||||||
cookies: &HashMap<String, String>,
|
cookies: &HashMap<String, String>,
|
||||||
) -> Result<PlaylistProbeResult> {
|
) -> Result<PlaylistProbeResult> {
|
||||||
let ytdlp = std::env::var("ARCHIVR_YT_DLP").unwrap_or_else(|_| "yt-dlp".to_string());
|
let ytdlp = resolve_yt_dlp();
|
||||||
|
|
||||||
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
let cookie_file: Option<PathBuf> = if !cookies.is_empty() {
|
||||||
let domain = domain_from_url(url);
|
let domain = domain_from_url(url);
|
||||||
|
|
@ -449,7 +556,7 @@ pub fn probe_playlist_qualities(
|
||||||
if let Some(cf) = &cookie_file {
|
if let Some(cf) = &cookie_file {
|
||||||
let _ = std::fs::remove_file(cf);
|
let _ = std::fs::remove_file(cf);
|
||||||
}
|
}
|
||||||
let out = out.with_context(|| format!("failed to spawn {ytdlp}"))?;
|
let out = out.with_context(|| format!("failed to spawn {}", ytdlp.display()))?;
|
||||||
if !out.status.success() {
|
if !out.status.success() {
|
||||||
let stderr = String::from_utf8_lossy(&out.stderr);
|
let stderr = String::from_utf8_lossy(&out.stderr);
|
||||||
bail!("yt-dlp -J failed for {url}: {stderr}");
|
bail!("yt-dlp -J failed for {url}: {stderr}");
|
||||||
|
|
@ -501,7 +608,131 @@ pub fn probe_playlist_qualities(
|
||||||
|
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
mod tests {
|
mod tests {
|
||||||
use super::{available_video_heights, has_audio_track, quality_format};
|
use super::{
|
||||||
|
available_video_heights, has_audio_track, quality_format, resolve_yt_dlp_uncached,
|
||||||
|
state_dir, STATE_DIR_ENV, YT_DLP_ENV, YT_DLP_FORCE_ENV,
|
||||||
|
};
|
||||||
|
use std::path::{Path, PathBuf};
|
||||||
|
use std::sync::{Mutex, MutexGuard};
|
||||||
|
|
||||||
|
/// Env vars are process-global, so resolver tests take turns.
|
||||||
|
static ENV_LOCK: Mutex<()> = Mutex::new(());
|
||||||
|
|
||||||
|
/// Clears every env var the resolver reads and hands back the serialising guard.
|
||||||
|
fn env_guard() -> MutexGuard<'static, ()> {
|
||||||
|
let guard = ENV_LOCK.lock().unwrap_or_else(|e| e.into_inner());
|
||||||
|
for key in [YT_DLP_FORCE_ENV, YT_DLP_ENV, STATE_DIR_ENV] {
|
||||||
|
unsafe { std::env::remove_var(key) };
|
||||||
|
}
|
||||||
|
guard
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Writes an executable stub that reports `version` when asked for `--version`.
|
||||||
|
fn fake_yt_dlp(path: &Path, version: &str) {
|
||||||
|
std::fs::create_dir_all(path.parent().unwrap()).unwrap();
|
||||||
|
std::fs::write(path, format!("#!/bin/sh\necho {version}\n")).unwrap();
|
||||||
|
#[cfg(unix)]
|
||||||
|
{
|
||||||
|
use std::os::unix::fs::PermissionsExt;
|
||||||
|
std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o755)).unwrap();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn resolve_yt_dlp_prefers_state_dir_when_newer() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
|
||||||
|
let pinned = tmp.path().join("nix/yt-dlp");
|
||||||
|
fake_yt_dlp(&pinned, "2026.08.19");
|
||||||
|
|
||||||
|
let state = tmp.path().join("state");
|
||||||
|
fake_yt_dlp(&state.join("yt-dlp/yt-dlp"), "2026.09.01");
|
||||||
|
|
||||||
|
unsafe {
|
||||||
|
std::env::set_var(YT_DLP_ENV, &pinned);
|
||||||
|
std::env::set_var(STATE_DIR_ENV, &state);
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(resolve_yt_dlp_uncached(), state.join("yt-dlp/yt-dlp"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn resolve_yt_dlp_prefers_pinned_when_newer() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
|
||||||
|
let pinned = tmp.path().join("nix/yt-dlp");
|
||||||
|
fake_yt_dlp(&pinned, "2026.09.15");
|
||||||
|
|
||||||
|
let state = tmp.path().join("state");
|
||||||
|
fake_yt_dlp(&state.join("yt-dlp/yt-dlp"), "2026.08.19");
|
||||||
|
|
||||||
|
unsafe {
|
||||||
|
std::env::set_var(YT_DLP_ENV, &pinned);
|
||||||
|
std::env::set_var(STATE_DIR_ENV, &state);
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(resolve_yt_dlp_uncached(), pinned);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn resolve_yt_dlp_breaks_version_ties_toward_state_dir() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
|
||||||
|
let pinned = tmp.path().join("nix/yt-dlp");
|
||||||
|
fake_yt_dlp(&pinned, "2026.09.01");
|
||||||
|
|
||||||
|
let state = tmp.path().join("state");
|
||||||
|
fake_yt_dlp(&state.join("yt-dlp/yt-dlp"), "2026.09.01");
|
||||||
|
|
||||||
|
unsafe {
|
||||||
|
std::env::set_var(YT_DLP_ENV, &pinned);
|
||||||
|
std::env::set_var(STATE_DIR_ENV, &state);
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(resolve_yt_dlp_uncached(), state.join("yt-dlp/yt-dlp"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn resolve_yt_dlp_honours_force_override_regardless_of_version() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
|
||||||
|
let forced = tmp.path().join("forced/yt-dlp");
|
||||||
|
fake_yt_dlp(&forced, "2020.01.01");
|
||||||
|
|
||||||
|
let state = tmp.path().join("state");
|
||||||
|
fake_yt_dlp(&state.join("yt-dlp/yt-dlp"), "2026.09.01");
|
||||||
|
|
||||||
|
unsafe {
|
||||||
|
std::env::set_var(YT_DLP_FORCE_ENV, &forced);
|
||||||
|
std::env::set_var(STATE_DIR_ENV, &state);
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(resolve_yt_dlp_uncached(), forced);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn resolve_yt_dlp_falls_back_to_bare_when_no_candidate_exists() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
let tmp = tempfile::tempdir().unwrap();
|
||||||
|
|
||||||
|
unsafe {
|
||||||
|
std::env::set_var(YT_DLP_ENV, tmp.path().join("missing/yt-dlp"));
|
||||||
|
std::env::set_var(STATE_DIR_ENV, tmp.path().join("empty-state"));
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(resolve_yt_dlp_uncached(), PathBuf::from("yt-dlp"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn state_dir_override_wins_over_platform_default() {
|
||||||
|
let _guard = env_guard();
|
||||||
|
unsafe { std::env::set_var(STATE_DIR_ENV, "/tmp/archivr-state-override") };
|
||||||
|
assert_eq!(state_dir(), Some(PathBuf::from("/tmp/archivr-state-override")));
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn quality_format_audio() {
|
fn quality_format_audio() {
|
||||||
|
|
|
||||||
|
|
@ -4,3 +4,4 @@ pub mod database;
|
||||||
pub mod downloader;
|
pub mod downloader;
|
||||||
pub mod hash;
|
pub mod hash;
|
||||||
pub mod twitter;
|
pub mod twitter;
|
||||||
|
pub mod summarizer;
|
||||||
|
|
|
||||||
2177
crates/archivr-core/src/summarizer.rs
Normal file
2177
crates/archivr-core/src/summarizer.rs
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -98,6 +98,17 @@ async fn main() -> Result<()> {
|
||||||
),
|
),
|
||||||
_ => {}
|
_ => {}
|
||||||
}
|
}
|
||||||
|
match archivr_core::database::fail_stalled_entry_summaries(&conn) {
|
||||||
|
Ok(n) if n > 0 => eprintln!(
|
||||||
|
"info: marked {n} stalled summary attempt(s) as failed in '{}'",
|
||||||
|
archive.id
|
||||||
|
),
|
||||||
|
Err(e) => eprintln!(
|
||||||
|
"warn: stalled summary cleanup failed for '{}': {e:#}",
|
||||||
|
archive.id
|
||||||
|
),
|
||||||
|
_ => {}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
File diff suppressed because it is too large
Load diff
1
crates/archivr-server/static/assets/index-1h0SqIvL.css
Normal file
1
crates/archivr-server/static/assets/index-1h0SqIvL.css
Normal file
File diff suppressed because one or more lines are too long
49
crates/archivr-server/static/assets/index-9VVEGfc7.js
Normal file
49
crates/archivr-server/static/assets/index-9VVEGfc7.js
Normal file
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -6,8 +6,8 @@
|
||||||
<title>Archivr</title>
|
<title>Archivr</title>
|
||||||
<link rel="icon" type="image/svg+xml" href="/favicon.svg">
|
<link rel="icon" type="image/svg+xml" href="/favicon.svg">
|
||||||
<link rel="icon" type="image/x-icon" href="/favicon.ico">
|
<link rel="icon" type="image/x-icon" href="/favicon.ico">
|
||||||
<script type="module" crossorigin src="/assets/index-B67momER.js"></script>
|
<script type="module" crossorigin src="/assets/index-9VVEGfc7.js"></script>
|
||||||
<link rel="stylesheet" crossorigin href="/assets/index-DicH9MNh.css">
|
<link rel="stylesheet" crossorigin href="/assets/index-1h0SqIvL.css">
|
||||||
</head>
|
</head>
|
||||||
<body>
|
<body>
|
||||||
<div id="root"></div>
|
<div id="root"></div>
|
||||||
|
|
|
||||||
128
docs/README.md
128
docs/README.md
|
|
@ -25,9 +25,12 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y
|
||||||
- [Supported Inputs](#supported-inputs)
|
- [Supported Inputs](#supported-inputs)
|
||||||
- [YouTube playlists and channels](#youtube-playlists-and-channels)
|
- [YouTube playlists and channels](#youtube-playlists-and-channels)
|
||||||
- [Video quality and audio-only](#video-quality-and-audio-only)
|
- [Video quality and audio-only](#video-quality-and-audio-only)
|
||||||
|
- [Text notes](#text-notes)
|
||||||
- [Configuration](#configuration)
|
- [Configuration](#configuration)
|
||||||
- [TOML config file](#toml-config-file)
|
- [TOML config file](#toml-config-file)
|
||||||
- [Environment variables](#environment-variables)
|
- [Environment variables](#environment-variables)
|
||||||
|
- [LLM providers](#llm-providers)
|
||||||
|
- [Keeping yt-dlp fresh](#keeping-yt-dlp-fresh)
|
||||||
- [Deployment](#deployment)
|
- [Deployment](#deployment)
|
||||||
- [Security](#security)
|
- [Security](#security)
|
||||||
- [NixOS](#hosting-on-nixos)
|
- [NixOS](#hosting-on-nixos)
|
||||||
|
|
@ -41,10 +44,13 @@ Archivr is a self-hosted tool for capturing and preserving digital content — Y
|
||||||
- **Web pages** — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
|
- **Web pages** — full self-contained HTML snapshots via SingleFile + Chromium; optional Freedium mirror for paywalled articles; reader mode
|
||||||
- **Local files** — import any file from disk by `file://` path
|
- **Local files** — import any file from disk by `file://` path
|
||||||
- **Deduplication** — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
|
- **Deduplication** — SHA3-256 content-addressed blob store shared across all captures; identical files are stored once
|
||||||
- **Tags and search** — hierarchical tag tree, full-text search, filterable entry list
|
- **Tags and search** — hierarchical tag tree, full-text search (including the latest completed summary and its generated JSON tags), filterable entry list
|
||||||
- **Multiple archives** — the server mounts any number of separate archives from a single TOML config
|
- **Multiple archives** — the server mounts any number of separate archives from a single TOML config
|
||||||
- **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
|
- **Role-based auth** — Guest / User / Admin / Owner roles; session cookies and API tokens; Argon2 passwords
|
||||||
- **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
|
- **Quality selection** — choose video quality or audio-only per capture; a live metadata probe populates the selector before download
|
||||||
|
- **LLM summaries** — regenerable per-entry summary via the Anthropic HTTP API, an OpenAI-compatible HTTP API, a local `claude` CLI, or a local `codex` CLI; triggered manually from the entry rail, never automatically on capture; text-only by default, with an explicit `Include attached images` option
|
||||||
|
- **Text notes** — capture a plain-text or Markdown note with a title and no URL; the byte-preserving note is stored as a normal deduplicated blob and opens in the usual entry-rail preview
|
||||||
|
- **In-progress capture indicator** — running captures appear as a compact spinner row in the entries list until they finish, replacing the earlier grey skeleton block
|
||||||
|
|
||||||
## Quick Start
|
## Quick Start
|
||||||
|
|
||||||
|
|
@ -161,6 +167,18 @@ The `POST /api/archives/:id/captures` endpoint accepts an optional `quality` fie
|
||||||
|
|
||||||
`"audio"` selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting `quality` or passing `"best"` downloads at the highest available quality.
|
`"audio"` selects the most efficient native audio track without re-encoding (Opus/WebM preferred, then AAC/M4A). Omitting `quality` or passing `"best"` downloads at the highest available quality.
|
||||||
|
|
||||||
|
### Text notes
|
||||||
|
|
||||||
|
Not every capture has a URL. **Add text** in the capture dialog takes a title and a body and turns them into a
|
||||||
|
self-contained entry — useful for a scrap of prose, a quote, or a note attached to the surrounding archive. The
|
||||||
|
body is stored verbatim; no network fetch happens.
|
||||||
|
|
||||||
|
Two body types are accepted: `text/markdown` (saved as `.md`) and `text/plain` (saved as `.txt`). Anything else is
|
||||||
|
rejected. The body lands in `store/raw/…` under its SHA3-256 content hash, exactly like every other capture, so an
|
||||||
|
identical note captured twice is stored once.
|
||||||
|
|
||||||
|
Text notes have no synthetic source URL: the original-URL field stays empty rather than inventing a `text:` locator.
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
### TOML config file
|
### TOML config file
|
||||||
|
|
@ -191,7 +209,8 @@ See `docker/config.example.toml` for a complete annotated example.
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `ARCHIVR_BIND` | `127.0.0.1:8080` | Bind address; overrides `bind` in TOML |
|
| `ARCHIVR_BIND` | `127.0.0.1:8080` | Bind address; overrides `bind` in TOML |
|
||||||
| `ARCHIVR_STATIC_DIR` | `crates/archivr-server/static` | Pre-built frontend asset directory |
|
| `ARCHIVR_STATIC_DIR` | `crates/archivr-server/static` | Pre-built frontend asset directory |
|
||||||
| `ARCHIVR_YT_DLP` | `yt-dlp` | yt-dlp binary used for video and social downloads |
|
| `ARCHIVR_YT_DLP` | `yt-dlp` | yt-dlp binary used for video and social downloads; the Nix wrappers point this at the pinned release |
|
||||||
|
| `ARCHIVR_YT_DLP_FORCE` | — | Absolute path to a yt-dlp binary that MUST be used, bypassing the resolver. Prefer `ARCHIVR_YT_DLP` unless you are overriding for a specific run |
|
||||||
| `ARCHIVR_SINGLE_FILE` | `single-file` | single-file-cli binary for web page archiving |
|
| `ARCHIVR_SINGLE_FILE` | `single-file` | single-file-cli binary for web page archiving |
|
||||||
| `ARCHIVR_CHROME` | `chromium` | Chromium executable passed to single-file |
|
| `ARCHIVR_CHROME` | `chromium` | Chromium executable passed to single-file |
|
||||||
| `ARCHIVR_CHROME_ARGS` | — | Extra space-separated Chromium flags (Docker sets `--no-sandbox`) |
|
| `ARCHIVR_CHROME_ARGS` | — | Extra space-separated Chromium flags (Docker sets `--no-sandbox`) |
|
||||||
|
|
@ -201,6 +220,105 @@ See `docker/config.example.toml` for a complete annotated example.
|
||||||
|
|
||||||
The Nix wrapper and Docker image set `ARCHIVR_STATIC_DIR`, `ARCHIVR_SINGLE_FILE`, and `ARCHIVR_CHROME` automatically.
|
The Nix wrapper and Docker image set `ARCHIVR_STATIC_DIR`, `ARCHIVR_SINGLE_FILE`, and `ARCHIVR_CHROME` automatically.
|
||||||
|
|
||||||
|
#### LLM providers
|
||||||
|
|
||||||
|
Summaries are manual and provider-agnostic. Only the variables for the provider you actually select are read; the two
|
||||||
|
HTTP providers refuse to start without their API key. They are text-only by default. Selecting `Include attached images`
|
||||||
|
explicitly sends eligible archived image data to the chosen provider; it is never attached automatically.
|
||||||
|
|
||||||
|
| Provider | Attached images |
|
||||||
|
|---|---|
|
||||||
|
| Anthropic HTTP | Supported |
|
||||||
|
| OpenAI-compatible HTTP | Supported |
|
||||||
|
| Codex CLI | Supported |
|
||||||
|
| Claude CLI | Not supported |
|
||||||
|
|
||||||
|
Image inclusion considers only `media` artifacts with `jpg`, `jpeg`, `png`, `webp`, `gif`, or `avif` files. At most four
|
||||||
|
images are sent, each no larger than 5 MiB and no more than 12 MiB in total.
|
||||||
|
|
||||||
|
Free-text entry search also matches the latest completed summary text and its generated JSON tags. Entries with no
|
||||||
|
summary, or only a pending or failed summary, get no summary-derived match.
|
||||||
|
|
||||||
|
Each request is cached under the provider and the **requested** model identifier. If a provider reports a more precise
|
||||||
|
resolved model (for example, an alias's concrete version), the UI displays that resolved name as attribution without
|
||||||
|
changing the cache identity.
|
||||||
|
|
||||||
|
Summary attempts move from `pending` to `running` and then to `completed` or `failed`. On server startup, interrupted
|
||||||
|
pending or running attempts are marked failed. Regenerating does not replace an earlier completed summary until the
|
||||||
|
replacement succeeds, and public readers receive completed content only—never pending state or diagnostic errors.
|
||||||
|
|
||||||
|
| Variable | Default | Description |
|
||||||
|
|---|---|---|
|
||||||
|
| `ARCHIVR_ANTHROPIC_API_KEY` | *(required for `anthropic_http`)* | API key for the Anthropic Messages API |
|
||||||
|
| `ARCHIVR_ANTHROPIC_URL` | `https://api.anthropic.com/v1/messages` | Endpoint override, e.g. an internal proxy |
|
||||||
|
| `ARCHIVR_ANTHROPIC_MODEL` | `claude-3-5-sonnet-latest` | Model id used for Anthropic summaries |
|
||||||
|
| `ARCHIVR_OPENAI_API_KEY` | *(required for `openai_compatible`)* | API key for any OpenAI-compatible endpoint |
|
||||||
|
| `ARCHIVR_OPENAI_URL` | `https://api.openai.com/v1/chat/completions` | Endpoint override; point this at a local server to run offline |
|
||||||
|
| `ARCHIVR_OPENAI_MODEL` | `gpt-4o-mini` | Model id used for OpenAI-compatible summaries |
|
||||||
|
| `ARCHIVR_CLAUDE_CLI` | *(auto-discovered)* | Path to a local `claude` binary |
|
||||||
|
| `ARCHIVR_CLAUDE_MODEL` | *(the CLI's own default)* | Optional model override for the local Claude CLI |
|
||||||
|
| `ARCHIVR_CODEX_CLI` | *(auto-discovered)* | Path to a local `codex` binary |
|
||||||
|
| `ARCHIVR_CODEX_MODEL` | *(the CLI's own default)* | Optional model override for the local Codex CLI |
|
||||||
|
| `ARCHIVR_SUMMARY_HTTP_TIMEOUT` | `120` | Seconds before an HTTP-provider summary is killed |
|
||||||
|
| `ARCHIVR_SUMMARY_CLI_TIMEOUT` | `300` | Seconds before a CLI-provider summary is killed |
|
||||||
|
|
||||||
|
When `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CODEX_CLI` is unset the binary is auto-discovered, in this order: the well-known
|
||||||
|
absolute paths, then `$HOME/.local/bin/<name>`, then the bare name resolved through `PATH`. Note that `PATH` is
|
||||||
|
consulted **last** — if a stale binary sits at one of the well-known paths it wins over a newer one on `PATH`, so set
|
||||||
|
the variable explicitly when you have both. The well-known paths are `/opt/homebrew/bin/claude` and
|
||||||
|
`/usr/local/bin/claude` for Claude, and `/Applications/ChatGPT.app/Contents/Resources/codex`,
|
||||||
|
`/opt/homebrew/bin/codex`, and `/usr/local/bin/codex` for Codex.
|
||||||
|
|
||||||
|
## Keeping yt-dlp fresh
|
||||||
|
|
||||||
|
yt-dlp is the download engine behind every video and social capture. YouTube rotates its player-signature and API
|
||||||
|
surfaces on a days-to-weeks cadence, so a binary that worked last month starts returning HTTP 403 on downloads. Keeping
|
||||||
|
it current is ordinary maintenance, not an emergency.
|
||||||
|
|
||||||
|
**What ships.** `flake.nix` pins a specific yt-dlp release fetched straight from `github.com/yt-dlp/yt-dlp/releases`,
|
||||||
|
not from nixpkgs — that channel usually lags months behind. The `archivr-server` and `archivr` wrappers set
|
||||||
|
`ARCHIVR_YT_DLP` to that pinned binary.
|
||||||
|
|
||||||
|
**How the resolver picks.** At runtime archivr probes `--version` on each candidate — the pinned binary from
|
||||||
|
`ARCHIVR_YT_DLP` and any user-installed binary at `<state_dir>/yt-dlp/yt-dlp` — and runs the newest. An exact version
|
||||||
|
tie resolves in favour of your own install. Setting `ARCHIVR_YT_DLP_FORCE=/path/to/yt-dlp` bypasses the comparison
|
||||||
|
entirely. The state dir is `~/Library/Application Support/archivr` on macOS, and `$XDG_STATE_HOME/archivr` (default
|
||||||
|
`~/.local/state/archivr`) elsewhere.
|
||||||
|
|
||||||
|
There are three ways to get a fresh version, cheapest first.
|
||||||
|
|
||||||
|
**1. Self-update — no rebuild required.**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
archivr yt-dlp status # every candidate, its version, and which one wins
|
||||||
|
archivr yt-dlp update # download the latest zipapp into the state dir
|
||||||
|
archivr yt-dlp update --version 2026.09.15 # pin a specific release tag
|
||||||
|
```
|
||||||
|
|
||||||
|
When `ARCHIVR_YT_DLP_FORCE` applies, `status` shows that forced candidate and selects it as the winner.
|
||||||
|
|
||||||
|
The released artifact is a Python zipapp, so this path needs `python3` on `PATH` at run time.
|
||||||
|
|
||||||
|
**2. Automatic weekly bump.** `.github/workflows/update-ytdlp.yml` runs every Monday at 06:00 UTC, queries GitHub for
|
||||||
|
the latest release, and opens a PR bumping `version` and `hash` in `flake.nix` via `peter-evans/create-pull-request`.
|
||||||
|
It also accepts `workflow_dispatch` for an on-demand run.
|
||||||
|
|
||||||
|
**3. Manual bump**, when you need it now and do not want to wait for the weekly:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
NEW=$(curl -s https://api.github.com/repos/yt-dlp/yt-dlp/releases/latest | jq -r .tag_name)
|
||||||
|
HASH=$(nix hash file --sri --type sha256 <(curl -sL "https://github.com/yt-dlp/yt-dlp/releases/download/${NEW}/yt-dlp"))
|
||||||
|
|
||||||
|
# In flake.nix, inside the `ytDlp = pkgs.stdenv.mkDerivation { … }` block:
|
||||||
|
# version = "OLD"; → version = "$NEW";
|
||||||
|
# url = ".../download/OLD/yt-dlp"; → .../download/$NEW/yt-dlp
|
||||||
|
# hash = "sha256-OLD…"; → hash = "$HASH";
|
||||||
|
|
||||||
|
nix build .#archivr-server
|
||||||
|
./result/bin/archivr yt-dlp status # the env row should report the new version
|
||||||
|
git commit -am "chore(nix): yt-dlp OLD → $NEW"
|
||||||
|
```
|
||||||
|
|
||||||
## Deployment
|
## Deployment
|
||||||
|
|
||||||
### Security
|
### Security
|
||||||
|
|
@ -286,6 +404,11 @@ The image compiles the Rust binary in a separate build stage; only runtime depen
|
||||||
|
|
||||||
Runtime dependencies beyond Rust and Node: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
|
Runtime dependencies beyond Rust and Node: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
|
||||||
|
|
||||||
|
Entry summaries are served by one of four interchangeable providers — `anthropic_http`, `openai_compatible`,
|
||||||
|
`claude_cli`, or `codex_cli` — each configured entirely through the environment; see
|
||||||
|
[LLM providers](#llm-providers) for the full variable list. The `archivr` CLI itself exposes `archive`, `init`, and
|
||||||
|
`yt-dlp status` / `yt-dlp update`; summaries are triggered from the web UI rather than the command line.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
# Rust (workspace root)
|
# Rust (workspace root)
|
||||||
cargo build
|
cargo build
|
||||||
|
|
@ -307,3 +430,4 @@ nix build .#archivr-server
|
||||||
## License
|
## License
|
||||||
|
|
||||||
MIT — see [LICENSE](../LICENSE.md).
|
MIT — see [LICENSE](../LICENSE.md).
|
||||||
|
\n
|
||||||
|
|
|
||||||
474
docs/superpowers/plans/2026-08-23-x-article-vision-search.md
Normal file
474
docs/superpowers/plans/2026-08-23-x-article-vision-search.md
Normal file
|
|
@ -0,0 +1,474 @@
|
||||||
|
# X Article, Vision Summaries, and Summary Search Implementation Plan
|
||||||
|
|
||||||
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||||
|
|
||||||
|
**Goal:** Make X Articles summarize their archived article text, let a user explicitly attach eligible archived images to a requested summary, and search the latest completed summary text (including its JSON tags).
|
||||||
|
|
||||||
|
**Architecture:** Keep `archivr-core` synchronous and make the summary input the single carrier of both reduced text and an opt-in, bounded list of image descriptors. The image-selection policy is serialized into the existing `input_sha256` preimage, preserving the existing `entry_summaries` uniqueness key without a migration. Extend the existing server-side entry search query with a correlated latest-completed-summary predicate; the endpoint response and frontend search transport stay unchanged.
|
||||||
|
|
||||||
|
**Tech Stack:** Rust 2024, `anyhow`, `rusqlite`, `reqwest` blocking HTTP, `serde_json`, Axum, React JSX, and plain CSS.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## File map and interfaces
|
||||||
|
|
||||||
|
| File | Responsibility |
|
||||||
|
| --- | --- |
|
||||||
|
| `crates/archivr-core/src/summarizer.rs` | X Article reducer, image candidate policy, `SummaryBuildOptions`, `SummaryImage`, cache digest preimage, provider request payloads, Codex invocation, and unit tests. |
|
||||||
|
| `crates/archivr-server/src/routes.rs` | Parse `include_images`, reject an unsupported Claude CLI vision request before a job is created, and pass build options to both cache lookup and the background worker. |
|
||||||
|
| `frontend/src/api.js` | Send the explicit `include_images` boolean in the existing summary POST. |
|
||||||
|
| `frontend/src/components/ContextRail.jsx` | Per-generation checkbox, provider-specific disabled Claude state, privacy/cap warning, and request wiring. |
|
||||||
|
| `frontend/src/styles.css` | Dedicated summary-image option layout and disabled-note treatment. |
|
||||||
|
| `crates/archivr-core/src/archive.rs` | Correlated SQL predicate for the latest completed summary and search tests. |
|
||||||
|
| `docs/README.md`, `AGENTS.md`, `ARCHIVR-MENTAL-MODEL.md` | User, contributor, and architectural documentation after the implementation is complete. |
|
||||||
|
|
||||||
|
Define the following core interfaces before server or UI tasks use them. `archive_file` is an absolute local path derived from `ArchivePaths.store_path` plus the stored artifact `relpath`; it is never returned from an API.
|
||||||
|
|
||||||
|
```rust
|
||||||
|
pub const MAX_SUMMARY_IMAGES: usize = 4;
|
||||||
|
pub const MAX_SUMMARY_IMAGE_BYTES: u64 = 5 * 1024 * 1024;
|
||||||
|
pub const MAX_SUMMARY_IMAGE_TOTAL_BYTES: u64 = 12 * 1024 * 1024;
|
||||||
|
|
||||||
|
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
|
||||||
|
pub struct SummaryBuildOptions {
|
||||||
|
pub include_images: bool,
|
||||||
|
}
|
||||||
|
|
||||||
|
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||||
|
pub struct SummaryImage {
|
||||||
|
pub sha256: String,
|
||||||
|
pub mime_type: String,
|
||||||
|
pub byte_size: u64,
|
||||||
|
pub archive_file: PathBuf,
|
||||||
|
}
|
||||||
|
|
||||||
|
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||||
|
pub struct SummaryRequest {
|
||||||
|
pub entry_uid: String,
|
||||||
|
pub title: Option<String>,
|
||||||
|
pub source_kind: String,
|
||||||
|
pub entity_kind: String,
|
||||||
|
pub content: String,
|
||||||
|
pub images: Vec<SummaryImage>,
|
||||||
|
}
|
||||||
|
|
||||||
|
pub fn build_summary_input(
|
||||||
|
paths: &ArchivePaths,
|
||||||
|
entry_uid: &str,
|
||||||
|
options: SummaryBuildOptions,
|
||||||
|
) -> Result<SummaryInput>;
|
||||||
|
|
||||||
|
pub fn summarize_entry(
|
||||||
|
archive_paths: &ArchivePaths,
|
||||||
|
entry_uid: &str,
|
||||||
|
options: SummaryBuildOptions,
|
||||||
|
provider: &dyn SummaryProvider,
|
||||||
|
prompt_version: &str,
|
||||||
|
) -> Result<database::EntrySummaryRecord>;
|
||||||
|
```
|
||||||
|
|
||||||
|
Use one deterministic digest preimage: `content` bytes, then `"\0images="`, then `include_images` as `"0"` or `"1"`, followed by each selected image in query order as `"\0" + sha256 + "\0" + mime_type + "\0" + byte_size`. Hash that complete byte sequence with the existing `hash::hash_bytes`. This means an unchecked request has no images and a different digest from a checked request even when no candidate qualifies.
|
||||||
|
|
||||||
|
### Task 1: Add the X Article reducer with test-first precedence
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `crates/archivr-core/src/summarizer.rs`
|
||||||
|
- Test: `crates/archivr-core/src/summarizer.rs` (`#[cfg(test)] mod tests`)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing tests for each per-status X Article precedence rule.**
|
||||||
|
|
||||||
|
Add assertions against `extract_tweet_text` using values that retain a link-only top-level tweet body:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
#[test]
|
||||||
|
fn extract_tweet_text_prefers_x_article_plain_text_over_tco_body() {
|
||||||
|
let tweet = serde_json::json!({
|
||||||
|
"full_text": "https://t.co/article",
|
||||||
|
"article": { "title": "Skin guide", "plain_text": "Use sunscreen daily." }
|
||||||
|
});
|
||||||
|
assert_eq!(extract_tweet_text(&tweet).as_deref(),
|
||||||
|
Some("Skin guide\n\nUse sunscreen daily."));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn extract_tweet_text_uses_article_blocks_when_plain_text_is_empty() {
|
||||||
|
let tweet = serde_json::json!({"article": {
|
||||||
|
"title": "Blocks", "plain_text": " ",
|
||||||
|
"blocks": [{"text": "First"}, {"children": [{"text": "Second"}]}]
|
||||||
|
}});
|
||||||
|
assert_eq!(extract_tweet_text(&tweet).as_deref(), Some("Blocks\n\nFirst\n\nSecond"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn extract_tweet_text_falls_back_from_article_to_link_only_tweet_body() {
|
||||||
|
let tweet = serde_json::json!({
|
||||||
|
"full_text": "https://t.co/fallback",
|
||||||
|
"article": {"title": "Preview", "preview_text": "Preview copy", "summary_text": "Later"}
|
||||||
|
});
|
||||||
|
assert_eq!(extract_tweet_text(&tweet).as_deref(), Some("Preview\n\nPreview copy"));
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the focused test target and observe it fail.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core extract_tweet_text_`
|
||||||
|
|
||||||
|
Expected: FAIL because the current reducer returns the top-level `full_text` or does not descend into `article.blocks`.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement `article_text` and deterministic block flattening.**
|
||||||
|
|
||||||
|
Add private helpers before `extract_tweet_text`:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
fn nonempty_string(v: &serde_json::Value, key: &str) -> Option<String>;
|
||||||
|
fn flatten_article_blocks(v: &serde_json::Value, out: &mut Vec<String>);
|
||||||
|
fn article_text(status: &serde_json::Value) -> Option<String>;
|
||||||
|
```
|
||||||
|
|
||||||
|
`article_text` must inspect the status's `article` object before ordinary fields. With a nonempty title, format each successful source as `title + "\n\n" + body`; if title is empty, return only `body`. Select the body in this exact order: nonblank `plain_text`; recursive text leaves from `blocks` in JSON array/object encounter order; nonblank `preview_text`; nonblank `summary_text`. `flatten_article_blocks` must collect only textual scalar values from conventional textual keys (`text`, `plain_text`, `content`, `body`, `title`, `heading`) and recursively visit arrays and objects; it must not stringify IDs, URLs, booleans, media metadata, or arbitrary scalar fields. Join block leaves with `"\n\n"`.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Integrate the helper into all tweet-status paths.**
|
||||||
|
|
||||||
|
Make `one(v)` call `article_text(v).or_else(|| normal_tweet_text(v))`, where `normal_tweet_text` preserves the existing `full_text`, `text`, `content`, `body` sequence. Keep support for a top-level `{ "tweet": ... }` wrapper and for embedded `thread`, `tweets`, and `replies` members.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Add the thread-artifact regression test and run the focused tests.**
|
||||||
|
|
||||||
|
Add a fixture archive with two `raw_tweet_json` artifacts, where each JSON status has article text, then assert `build_summary_input(..., SummaryBuildOptions::default())?.request.content` contains both article bodies separated by `"\n\n---\n\n"`. Run: `cargo test -p archivr-core extract_tweet_text_ build_summary_input_`
|
||||||
|
|
||||||
|
Expected: PASS, including the existing wrapped/thread tweet tests.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit the atomic reducer change.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add crates/archivr-core/src/summarizer.rs
|
||||||
|
git commit -m "fix: summarize X Article text"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 2: Model and select explicit image inputs in core
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `crates/archivr-core/src/summarizer.rs`
|
||||||
|
- Test: `crates/archivr-core/src/summarizer.rs` (`#[cfg(test)] mod tests`)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing candidate-selection and digest tests.**
|
||||||
|
|
||||||
|
Build a temporary archive entry containing `media` artifacts for valid `jpg`, `png`, `webp`, `gif`, and `avif`, plus `avatar`, `video`, `audio`, unsupported `svg`, one 5 MiB + 1 byte image, and enough valid images to exceed both the four-image and 12 MiB limits. Assert only role `media`, allowed MIME/extension pairs, at most four descriptors, no descriptor over 5 MiB, and total selected bytes at most 12 MiB. Also assert:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
let text_only = build_summary_input(&paths, &uid, SummaryBuildOptions { include_images: false })?;
|
||||||
|
let visual = build_summary_input(&paths, &uid, SummaryBuildOptions { include_images: true })?;
|
||||||
|
assert!(text_only.request.images.is_empty());
|
||||||
|
assert!(!visual.request.images.is_empty());
|
||||||
|
assert_ne!(text_only.input_sha256, visual.input_sha256);
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the new core tests and observe failure.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core summary_image_`
|
||||||
|
|
||||||
|
Expected: FAIL because `SummaryRequest` has no images and input construction has no image-selection mode.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Define the shared image model and query candidates from existing blobs.**
|
||||||
|
|
||||||
|
Add the constants and `SummaryBuildOptions`, `SummaryImage`, and `SummaryRequest.images` definitions from the file map. Add a private `load_summary_image_candidates(conn, entry_id) -> Result<Vec<SummaryImage>>` querying `entry_artifacts ea JOIN blobs b` for `ea.entry_id = ?1 AND ea.artifact_role = 'media'`, ordered by `ea.id ASC`, selecting `b.sha256`, `b.mime_type`, `b.extension`, `b.byte_size`, and `ea.relpath`.
|
||||||
|
|
||||||
|
Accept a candidate only when both of the following are true: its extension is one of `jpg`, `jpeg`, `png`, `webp`, `gif`, `avif`, and its MIME is the matching `image/jpeg`, `image/png`, `image/webp`, `image/gif`, or `image/avif` family. Resolve `archive_file` under the configured store path and reject a candidate whose canonicalized/normalized path escapes that store root. Stop at the first candidate that would exceed either image count, per-image, or aggregate byte limit; continue scanning later candidates so a too-large or unsupported early artifact cannot hide a valid later one.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Build options-aware input and a complete cache digest.**
|
||||||
|
|
||||||
|
Change every current `build_summary_input` call to pass `SummaryBuildOptions::default()` until Task 4 changes the server. Populate `request.images` only if `options.include_images` is true; text extraction and `MAX_INPUT_CHARS` handling remain identical. Replace `hash_bytes(content.as_bytes())` with a private `summary_input_digest(content, include_images, images)` implementing the stated NUL-delimited preimage so the existing database uniqueness constraint continues to distinguish all modes. Do not alter `database.rs`: `input_sha256` already participates in the cache key.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run core regression tests.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core summary_image_ build_summary_input_`
|
||||||
|
|
||||||
|
Expected: PASS; the text-only request has zero image descriptors, and selection is deterministic by artifact insertion order.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit the core input model.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add crates/archivr-core/src/summarizer.rs
|
||||||
|
git commit -m "feat: model opt-in summary images"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 3: Make provider transports honor image descriptors
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `crates/archivr-core/src/summarizer.rs`
|
||||||
|
- Test: `crates/archivr-core/src/summarizer.rs` (`#[cfg(test)] mod tests`)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing payload, command, and capability tests.**
|
||||||
|
|
||||||
|
Construct a `SummaryRequest` with one tiny fixture `SummaryImage` and assert `anthropic_request_body` puts a text block and `{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } }` in `messages[0].content`. Assert `openai_request_body` emits a text content part plus `{ "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } }`. Unit-test a pure Codex argument builder so its primary arguments contain `exec`, `--image`, the fixture path, `--output-last-message`, output path, and final `-`; test its positional fallback also keeps `--image`. Assert `ClaudeCliProvider::summarize` returns an error containing `Claude CLI cannot attach local images` when `request.images` is nonempty.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the provider tests and observe failure.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core "anthropic_request_body|openai_request_body|codex.*image|claude.*images"`
|
||||||
|
|
||||||
|
Expected: FAIL because HTTP bodies are string-only, Codex has no `--image`, and Claude silently accepts the request.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Encode images for each HTTP protocol.**
|
||||||
|
|
||||||
|
Add `read_image_base64(image: &SummaryImage) -> Result<String>` which reads only the already bounded selected file and uses `base64::Engine` with the existing dependency or workspace dependency. Make `anthropic_request_body` and `openai_request_body` return their present text-only JSON shapes when `request.images.is_empty()` and their documented content-part arrays otherwise. Preserve `SYSTEM_PROMPT`, `build_user_prompt`, provider URL, headers, and response parsing.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Extend Codex safely and reject Claude at the provider boundary.**
|
||||||
|
|
||||||
|
Refactor `codex::run` to accept `&[SummaryImage]`; add `--image <archive_file>` once per selected image before `--output-last-message` in both primary stdin and positional-prompt forms. Keep the existing last-message temporary-file contract and cleanup behavior. Have `ClaudeCliProvider::summarize` `bail!("Claude CLI cannot attach local images; choose an HTTP provider or Codex CLI")` before spawning when images are supplied.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Thread image-aware options through synchronous orchestration.**
|
||||||
|
|
||||||
|
Change `summarize_entry` to accept `SummaryBuildOptions` and call the options-aware builder before provider invocation. This is a synchronous core function; do not introduce Tokio, async traits, or new database fields.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Run provider and existing summary tests.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core summarizer::tests`
|
||||||
|
|
||||||
|
Expected: PASS, with all four providers retaining their text-only behavior when `images` is empty.
|
||||||
|
|
||||||
|
- [ ] **Step 7: Commit the provider implementation.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add crates/archivr-core/src/summarizer.rs
|
||||||
|
git commit -m "feat: attach opted-in images to summaries"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 4: Expose the opt-in flag through the server API
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `crates/archivr-server/src/routes.rs`
|
||||||
|
- Test: `crates/archivr-server/src/routes.rs` (`#[cfg(test)] mod tests`)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing route tests.**
|
||||||
|
|
||||||
|
Add authenticated POST tests that deserialize a request without `include_images` and assert it takes the text-only build path, and with `{ "provider": "claude_cli", "include_images": true }` assert status `400 BAD_REQUEST` and an error containing `Claude CLI cannot attach local images`. Add a successful non-Claude request test with `include_images: true` using a configured test provider and assert the returned cache key differs from the equivalent text-only request. Keep GET assertions unchanged: it remains `{ "entry_uid", "summary" }`.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the focused route tests and observe failure.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-server "summary.*include_images|claude.*images"`
|
||||||
|
|
||||||
|
Expected: FAIL because the request body does not accept the field and no capability check occurs.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Add the backward-compatible request field and capability guard.**
|
||||||
|
|
||||||
|
Extend `SummaryRequestBody` exactly as follows:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
#[derive(Debug, serde::Deserialize)]
|
||||||
|
struct SummaryRequestBody {
|
||||||
|
provider: String,
|
||||||
|
#[serde(default)]
|
||||||
|
force: bool,
|
||||||
|
#[serde(default)]
|
||||||
|
include_images: bool,
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
After resolving `provider_cfg` and before cache lookup/upsert/spawn, return `ApiError::bad_request("Claude CLI cannot attach local images; choose an HTTP provider or Codex CLI")` when `body.include_images && matches!(provider_cfg, ProviderConfig::ClaudeCli(_))`. Construct `SummaryBuildOptions { include_images: body.include_images }` once and pass it to the preflight builder and cloned into `spawn_blocking` for `summarize_entry`.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Verify cache and lifecycle consistency.**
|
||||||
|
|
||||||
|
Ensure preflight `build_summary_input` and background `summarize_entry` receive the same options, so `find_entry_summary` and `upsert_pending_entry_summary` use the same digest. Do not change `entry_summaries`, `latest_entry_summary`, or the GET route; the existing input-hash uniqueness constraint is sufficient.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run the focused server tests.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-server "summary.*include_images|claude.*images"`
|
||||||
|
|
||||||
|
Expected: PASS; omitting the new field is text-only and a Claude image request produces no pending summary row.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit the API wiring.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add crates/archivr-server/src/routes.rs
|
||||||
|
git commit -m "feat: accept image summary requests"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 5: Add the explicit, provider-aware image consent control
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `frontend/src/api.js`
|
||||||
|
- Modify: `frontend/src/components/ContextRail.jsx`
|
||||||
|
- Modify: `frontend/src/styles.css`
|
||||||
|
- Test: manual browser smoke test (no frontend test harness exists)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Inspect the existing provider selector and write the manual failure script.**
|
||||||
|
|
||||||
|
In a locally authenticated entry detail, select an HTTP provider and verify the Summary section currently has no `Include attached images` checkbox; select Claude CLI and verify there is no capability explanation. Record this as the observed pre-implementation failure. Do not add a frontend test framework.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Change the API client contract.**
|
||||||
|
|
||||||
|
Change the function signature and POST body only:
|
||||||
|
|
||||||
|
```js
|
||||||
|
export async function requestEntrySummary(
|
||||||
|
archiveId, entryUid, { provider, force = false, includeImages = false } = {}
|
||||||
|
) {
|
||||||
|
// existing fetch and error parsing
|
||||||
|
body: JSON.stringify({ provider, force, include_images: includeImages })
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Keep all `fetch` calls inside `frontend/src/api.js`; do not add an inline fetch in the component.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement the local per-generation control.**
|
||||||
|
|
||||||
|
Add `const [includeSummaryImages, setIncludeSummaryImages] = useState(false)` beside the summary provider state. In `handleGenerateSummary`, pass `includeImages: includeSummaryImages`. Reset this state to `false` whenever `detail?.summary?.entry_uid` changes, so a consent choice cannot carry to another entry. Do not persist the checkbox in `sessionStorage`; the consent is per generation and defaults off.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Render clear consent, scope, and Claude capability states.**
|
||||||
|
|
||||||
|
In `.rail-summary-controls`, below the provider `<select>`, render a labeled checkbox with exact visible label `Include attached images`. Its help text must say that selected archived images are sent to the chosen provider and that only up to four supported images (5 MiB each, 12 MiB total) can be attached; unsupported or oversized artifacts are skipped. When `summaryProvider === 'claude_cli'`, render the checkbox disabled, force `includeSummaryImages` to false via an effect or provider-change handler, and show `Claude CLI cannot attach local images. Choose an HTTP provider or Codex CLI.` Do not submit a silently dropped image choice.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Add scoped plain-CSS rules.**
|
||||||
|
|
||||||
|
Add `.rail-summary-image-option`, `.rail-summary-image-option__label`, `.rail-summary-image-option__note`, and `.rail-summary-image-option--disabled` under the existing summary rail CSS. Use the project variables (`--muted`, `--line`, `--paper`) and preserve keyboard focus and normal checkbox semantics; do not use a generic row class or inline layout styles.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Run the manual success script and build verification.**
|
||||||
|
|
||||||
|
Run: `bun run build`
|
||||||
|
|
||||||
|
Expected: successful production bundle in `crates/archivr-server/static`. Then manually verify: unchecked generation sends `include_images:false`; checked Anthropic/OpenAI/Codex generation sends `true`; changing to Claude unchecks/disables the control and shows the exact explanation; server errors still appear through existing `summaryError` handling.
|
||||||
|
|
||||||
|
- [ ] **Step 7: Commit source files, not generated static output.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add frontend/src/api.js frontend/src/components/ContextRail.jsx frontend/src/styles.css
|
||||||
|
git commit -m "feat: add summary image consent control"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 6: Search latest completed summary JSON without changing API shape
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `crates/archivr-core/src/archive.rs`
|
||||||
|
- Test: `crates/archivr-core/src/archive.rs` (`#[cfg(test)] mod tests`)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write failing search tests with real cache rows.**
|
||||||
|
|
||||||
|
Extend `make_test_db_with_entries` or add a focused fixture helper that inserts summary rows through `database::upsert_pending_entry_summary` and `database::update_entry_summary_status`. Add assertions for all of the following:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// A completed JSON string with {"tags":["skincare","dermatology"]} matches skincare.
|
||||||
|
// An unrelated query yields no result.
|
||||||
|
// Two completed rows: the newer completed row is searched; the older one is not.
|
||||||
|
// A completed older row remains matched while a newer row is pending or failed.
|
||||||
|
// source:, entity:, url:, title:, after:, before:, tag: and collection/visibility scope keep their current behavior.
|
||||||
|
```
|
||||||
|
|
||||||
|
Set distinct `updated_at` values (or insert/transition rows in distinct timestamp order) so "latest completed" is unambiguous. The tag assertion must match `summary_text` itself, not `entry_tag_assignments`.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Run the focused tests and observe failure.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core search_.*summary`
|
||||||
|
|
||||||
|
Expected: FAIL because free text only checks entry and source identity fields.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Add a parameter-bound latest-completed summary predicate.**
|
||||||
|
|
||||||
|
In the existing unqualified `query.q` block in `search_entries`, preserve every present `LOWER(...) LIKE ?{n}` condition and add this clause using the same single bound `term`:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
OR LOWER(COALESCE((
|
||||||
|
SELECT s.summary_text
|
||||||
|
FROM entry_summaries s
|
||||||
|
WHERE s.entry_id = e.id
|
||||||
|
AND s.status = 'completed'
|
||||||
|
AND s.summary_text IS NOT NULL
|
||||||
|
ORDER BY s.completed_at DESC, s.updated_at DESC, s.id DESC
|
||||||
|
LIMIT 1
|
||||||
|
), '')) LIKE ?N
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the existing numbered parameter construction (`?{n}`) and push `term` once, so user text is never concatenated into SQL. `completed_at` ordering means only completed rows participate, and a later pending/failed row cannot displace an older completed row. Do not add an archive method, schema column, migration, route parameter, or frontend response field.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Run core search tests.**
|
||||||
|
|
||||||
|
Run: `cargo test -p archivr-core search_`
|
||||||
|
|
||||||
|
Expected: PASS, including old prefix-filter behavior and JSON tag substring matches.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Verify the existing API transport needs no change.**
|
||||||
|
|
||||||
|
Inspect `crates/archivr-server/src/routes.rs::search_entries_handler`, `frontend/src/api.js::searchEntries`, and `frontend/src/App.jsx` search call sites. Confirm the server still calls `archive::search_entries` with the same `SearchEntriesQuery` and returns `Vec<EntrySummary>`; record no source edit for these files unless the inspection reveals a type break. Add no client-side filtering.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit the search change.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add crates/archivr-core/src/archive.rs
|
||||||
|
git commit -m "feat: search completed summary tags"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 7: Document, bundle, and verify the completed feature set
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Modify: `docs/README.md`
|
||||||
|
- Modify: `AGENTS.md`
|
||||||
|
- Modify: `ARCHIVR-MENTAL-MODEL.md`
|
||||||
|
- Generated (do not hand-edit): `crates/archivr-server/static/`
|
||||||
|
|
||||||
|
- [ ] **Step 1: Update user documentation.**
|
||||||
|
|
||||||
|
In `docs/README.md`, document that summaries are manual, text-only by default, and the explicit `Include attached images` option sends at most four eligible local images to the selected provider. State the allowed formats and byte limits, identify Anthropic/OpenAI-compatible/Codex support, state Claude CLI cannot attach local images, and state free-text search includes the latest completed summary text and its generated tags.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Update contributor constraints.**
|
||||||
|
|
||||||
|
In `AGENTS.md`, record the `SummaryBuildOptions`/digest rule, the image candidate role and limits, the provider capability matrix, and the latest-completed-only search semantic. Preserve the rule that core remains synchronous and that generated static files are not hand-edited.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Update the architectural data-flow documentation.**
|
||||||
|
|
||||||
|
In `ARCHIVR-MENTAL-MODEL.md`, extend the LLM Summary section to show explicit UI consent flowing into image selection, cache hashing, provider transport, and the existing row lifecycle. Add that entry search reads only the latest completed `summary_text`, retaining a prior completed result while newer work is pending or failed.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Build and test the final implementation.**
|
||||||
|
|
||||||
|
Run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cargo test
|
||||||
|
bun --cwd frontend run build
|
||||||
|
cargo build
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: all Rust tests pass, frontend build succeeds, and the generated static bundle contains the checkbox UI. Do not hand-edit generated files; include them in a commit only if this repository currently tracks frontend bundle changes after `bun run build`.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Run the end-to-end manual smoke test.**
|
||||||
|
|
||||||
|
Start the server with a test archive and verify: an X Article whose normal tweet text is only a t.co URL summarizes article body text; a text-only generation remains unchanged; a checked vision-capable request attaches only bounded eligible images; Claude has a disabled explanatory option and a direct API request is rejected; a `skincare` search finds completed summary JSON tags; pending/failed rows do not hide a previous completed match.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit documentation and any tracked generated bundle.**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add docs/README.md AGENTS.md ARCHIVR-MENTAL-MODEL.md crates/archivr-server/static
|
||||||
|
git commit -m "docs: explain image summaries and summary search"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Task 8: Final implementation review before integration
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
|
||||||
|
- Review: `crates/archivr-core/src/summarizer.rs`
|
||||||
|
- Review: `crates/archivr-core/src/archive.rs`
|
||||||
|
- Review: `crates/archivr-server/src/routes.rs`
|
||||||
|
- Review: `frontend/src/api.js`
|
||||||
|
- Review: `frontend/src/components/ContextRail.jsx`
|
||||||
|
- Review: `frontend/src/styles.css`
|
||||||
|
- Review: `docs/README.md`, `AGENTS.md`, `ARCHIVR-MENTAL-MODEL.md`
|
||||||
|
|
||||||
|
- [ ] **Step 1: Perform the approved-design coverage review.**
|
||||||
|
|
||||||
|
Verify A is covered by Tasks 1 and 2 (plain text, blocks, preview/summary, normal tweet fallback, and every thread JSON artifact); B by Tasks 2–5 (explicit default-off consent, selection policy/caps, input hash, all four provider outcomes, POST flag, and compatible GET); and C by Task 6 (latest completed summary JSON/tags, pending/failed semantics, prefix-filter preservation, server-side architecture).
|
||||||
|
|
||||||
|
- [ ] **Step 2: Scan the plan and implementation for unfinished markers and type drift.**
|
||||||
|
|
||||||
|
Run: `rg -n -i '\\bt[o]do\\b|\\bt[b]d\\b|placehold[e]r|implement[[:space:]]later' docs/superpowers/plans/2026-08-23-x-article-vision-search.md crates/archivr-core/src/summarizer.rs crates/archivr-core/src/archive.rs crates/archivr-server/src/routes.rs frontend/src`
|
||||||
|
|
||||||
|
Expected: no newly introduced unfinished markers in the changed feature code or plan. Confirm every use of `SummaryBuildOptions`, `SummaryImage`, options-aware `build_summary_input`, and options-aware `summarize_entry` matches the Task 2 definitions.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Review commits and working tree.**
|
||||||
|
|
||||||
|
Run: `git log --oneline --decorate -8` and `git status --short`
|
||||||
|
|
||||||
|
Expected: atomic commits cover the reducer, core image model/providers, server API, frontend control, search, and docs; no unintended artifacts or source edits remain.
|
||||||
41
flake.nix
41
flake.nix
|
|
@ -92,6 +92,35 @@
|
||||||
cp -r . $out/
|
cp -r . $out/
|
||||||
'';
|
'';
|
||||||
};
|
};
|
||||||
|
# yt-dlp — pinned to a specific GitHub release rather than pulled through
|
||||||
|
# nixpkgs. Rationale: YouTube frequently rotates player-signature/API
|
||||||
|
# surfaces, and yt-dlp ships updates on a days-to-weeks cadence; even
|
||||||
|
# nixos-unstable often lags by months. When the binary is stale,
|
||||||
|
# captures fail with HTTP 403 on formats the old client can't
|
||||||
|
# authenticate. Fetching the zipapp directly (a Python zipapp with a
|
||||||
|
# `#!/usr/bin/env python3` shebang) lets us bump the version + hash in
|
||||||
|
# one place without waiting on nixpkgs. Wrapped so `python3` and
|
||||||
|
# `ffmpeg` — the two runtime deps for muxed downloads — are always on
|
||||||
|
# PATH regardless of the caller's environment.
|
||||||
|
#
|
||||||
|
# Bumping: replace `version`, then run `nix hash file <url>` on the
|
||||||
|
# new zipapp URL and paste the sri output into `hash`.
|
||||||
|
ytDlp = pkgs.stdenv.mkDerivation {
|
||||||
|
pname = "yt-dlp";
|
||||||
|
version = "2026.08.19";
|
||||||
|
src = pkgs.fetchurl {
|
||||||
|
url = "https://github.com/yt-dlp/yt-dlp/releases/download/2026.08.19/yt-dlp";
|
||||||
|
hash = "sha256-H6ZzPDfqb7Ucma2P54Xnt+XzJGybmAIwMp1Pty7Y1NY=";
|
||||||
|
};
|
||||||
|
dontUnpack = true;
|
||||||
|
nativeBuildInputs = [ pkgs.makeWrapper ];
|
||||||
|
installPhase = ''
|
||||||
|
mkdir -p $out/bin
|
||||||
|
install -m 0755 $src $out/bin/yt-dlp
|
||||||
|
wrapProgram $out/bin/yt-dlp \
|
||||||
|
--prefix PATH : ${lib.makeBinPath [ pkgs.python312 pkgs.ffmpeg ]}
|
||||||
|
'';
|
||||||
|
};
|
||||||
version = "0.1.0";
|
version = "0.1.0";
|
||||||
src = pkgs.lib.cleanSource ./.;
|
src = pkgs.lib.cleanSource ./.;
|
||||||
cargoLock = {
|
cargoLock = {
|
||||||
|
|
@ -139,7 +168,7 @@
|
||||||
version = "0.1.0";
|
version = "0.1.0";
|
||||||
nativeBuildInputs = [ pkgs.makeWrapper ];
|
nativeBuildInputs = [ pkgs.makeWrapper ];
|
||||||
buildInputs = [
|
buildInputs = [
|
||||||
pkgs.yt-dlp
|
ytDlp
|
||||||
pkgs.single-file-cli
|
pkgs.single-file-cli
|
||||||
tweetPython
|
tweetPython
|
||||||
] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ];
|
] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ];
|
||||||
|
|
@ -150,7 +179,7 @@
|
||||||
cp ${./vendor/twitter/scrape_user_tweet_contents.py} $out/libexec/archivr/scrape_user_tweet_contents.py
|
cp ${./vendor/twitter/scrape_user_tweet_contents.py} $out/libexec/archivr/scrape_user_tweet_contents.py
|
||||||
chmod +x $out/libexec/archivr/scrape_user_tweet_contents.py
|
chmod +x $out/libexec/archivr/scrape_user_tweet_contents.py
|
||||||
makeWrapper $out/libexec/archivr/archivr $out/bin/archivr \
|
makeWrapper $out/libexec/archivr/archivr $out/bin/archivr \
|
||||||
--set ARCHIVR_YT_DLP ${pkgs.yt-dlp}/bin/yt-dlp \
|
--set ARCHIVR_YT_DLP ${ytDlp}/bin/yt-dlp \
|
||||||
--set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \
|
--set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \
|
||||||
${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \
|
${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \
|
||||||
--set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \
|
--set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \
|
||||||
|
|
@ -159,7 +188,7 @@
|
||||||
--set ARCHIVR_COOKIE_EXT ${isdcac} \
|
--set ARCHIVR_COOKIE_EXT ${isdcac} \
|
||||||
--prefix PATH : ${
|
--prefix PATH : ${
|
||||||
lib.makeBinPath ([
|
lib.makeBinPath ([
|
||||||
pkgs.yt-dlp
|
ytDlp
|
||||||
pkgs.single-file-cli
|
pkgs.single-file-cli
|
||||||
tweetPython
|
tweetPython
|
||||||
] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ])
|
] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ])
|
||||||
|
|
@ -170,7 +199,7 @@
|
||||||
pname = "archivr-server-wrapped";
|
pname = "archivr-server-wrapped";
|
||||||
inherit version;
|
inherit version;
|
||||||
nativeBuildInputs = [ pkgs.makeWrapper ];
|
nativeBuildInputs = [ pkgs.makeWrapper ];
|
||||||
buildInputs = [ tweetPython pkgs.single-file-cli ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ];
|
buildInputs = [ ytDlp tweetPython pkgs.single-file-cli ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ];
|
||||||
phases = [ "installPhase" ];
|
phases = [ "installPhase" ];
|
||||||
installPhase = ''
|
installPhase = ''
|
||||||
mkdir -p $out/bin $out/libexec/archivr-server $out/share/archivr-server/static
|
mkdir -p $out/bin $out/libexec/archivr-server $out/share/archivr-server/static
|
||||||
|
|
@ -180,12 +209,14 @@
|
||||||
cp -r ${./crates/archivr-server/static}/* $out/share/archivr-server/static/
|
cp -r ${./crates/archivr-server/static}/* $out/share/archivr-server/static/
|
||||||
makeWrapper $out/libexec/archivr-server/archivr-server $out/bin/archivr-server \
|
makeWrapper $out/libexec/archivr-server/archivr-server $out/bin/archivr-server \
|
||||||
--set ARCHIVR_STATIC_DIR $out/share/archivr-server/static \
|
--set ARCHIVR_STATIC_DIR $out/share/archivr-server/static \
|
||||||
|
--set ARCHIVR_YT_DLP ${ytDlp}/bin/yt-dlp \
|
||||||
--set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \
|
--set ARCHIVR_SINGLE_FILE ${pkgs.single-file-cli}/bin/single-file \
|
||||||
${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \
|
${lib.optionalString pkgs.stdenv.isLinux "--set ARCHIVR_CHROME ${pkgs.chromium}/bin/chromium"} \
|
||||||
--set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \
|
--set ARCHIVR_TWEET_PYTHON ${tweetPython}/bin/python3 \
|
||||||
--set ARCHIVR_TWEET_SCRAPER $out/libexec/archivr-server/scrape_user_tweet_contents.py \
|
--set ARCHIVR_TWEET_SCRAPER $out/libexec/archivr-server/scrape_user_tweet_contents.py \
|
||||||
--set ARCHIVR_UBLOCK_EXT ${ublockLite} \
|
--set ARCHIVR_UBLOCK_EXT ${ublockLite} \
|
||||||
--set ARCHIVR_COOKIE_EXT ${isdcac}
|
--set ARCHIVR_COOKIE_EXT ${isdcac} \
|
||||||
|
--prefix PATH : ${lib.makeBinPath ([ ytDlp pkgs.single-file-cli tweetPython ] ++ lib.optionals pkgs.stdenv.isLinux [ pkgs.chromium ])}
|
||||||
'';
|
'';
|
||||||
};
|
};
|
||||||
archivr-all = pkgs.symlinkJoin {
|
archivr-all = pkgs.symlinkJoin {
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,5 @@
|
||||||
async function getJson(url) {
|
async function getJson(url, options) {
|
||||||
const response = await fetch(url);
|
const response = await fetch(url, options);
|
||||||
if (!response.ok) {
|
if (!response.ok) {
|
||||||
throw new Error(`${response.status} ${response.statusText}`);
|
throw new Error(`${response.status} ${response.statusText}`);
|
||||||
}
|
}
|
||||||
|
|
@ -29,6 +29,50 @@ export async function fetchEntryDetail(archiveId, entryUid) {
|
||||||
return getJson(`/api/archives/${archiveId}/entries/${entryUid}`);
|
return getJson(`/api/archives/${archiveId}/entries/${entryUid}`);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── Entry summaries ────────────────────────────────────────────────────────
|
||||||
|
// Summaries are generated on demand, never at capture time. GET is safe for
|
||||||
|
// public sessions (the server applies the same visibility gate as entry detail).
|
||||||
|
|
||||||
|
export async function fetchEntrySummary(archiveId, entryUid, { signal } = {}) {
|
||||||
|
return getJson(`/api/archives/${archiveId}/entries/${entryUid}/summary`, { signal });
|
||||||
|
}
|
||||||
|
|
||||||
|
// Kicks off generation. Resolves to either an existing completed summary (200)
|
||||||
|
// or a freshly claimed pending row (202) — both carry a summary_uid, so the
|
||||||
|
// caller polls fetchEntrySummary either way.
|
||||||
|
// The server returns 400 with the exact missing env var name when a provider is
|
||||||
|
// unconfigured, so its body is surfaced verbatim rather than replaced.
|
||||||
|
export async function requestEntrySummary(archiveId, entryUid, { provider, force = false, includeImages = false, signal } = {}) {
|
||||||
|
const resp = await fetch(
|
||||||
|
`/api/archives/${archiveId}/entries/${entryUid}/summary`,
|
||||||
|
{
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": "application/json" },
|
||||||
|
body: JSON.stringify({ provider, force, include_images: includeImages }),
|
||||||
|
signal,
|
||||||
|
}
|
||||||
|
);
|
||||||
|
if (!resp.ok) {
|
||||||
|
// ApiError renders as { "error": "..." }; that message is the useful part
|
||||||
|
// (e.g. "missing required environment variable: ARCHIVR_ANTHROPIC_API_KEY"),
|
||||||
|
// so surface it verbatim instead of a generic status string.
|
||||||
|
const detail = await resp.text();
|
||||||
|
let message = detail.trim();
|
||||||
|
try { message = JSON.parse(detail).error || message } catch { /* non-JSON body */ }
|
||||||
|
throw new Error(message || `Summary request failed (${resp.status})`);
|
||||||
|
}
|
||||||
|
return resp.json();
|
||||||
|
}
|
||||||
|
|
||||||
|
// Text artifacts are served by the same entry-artifact endpoint as previews.
|
||||||
|
// Keep credentials explicit because this helper is also used by public/private
|
||||||
|
// archive views, and preserve the previous concise HTTP error contract.
|
||||||
|
export async function fetchArtifactText(src, { signal } = {}) {
|
||||||
|
const response = await fetch(src, { credentials: 'same-origin', signal });
|
||||||
|
if (!response.ok) throw new Error(`HTTP ${response.status}`);
|
||||||
|
return response.text();
|
||||||
|
}
|
||||||
|
|
||||||
export async function fetchEntryChildren(archiveId, entryUid) {
|
export async function fetchEntryChildren(archiveId, entryUid) {
|
||||||
return getJson(`/api/archives/${archiveId}/entries/${entryUid}/children`);
|
return getJson(`/api/archives/${archiveId}/entries/${entryUid}/children`);
|
||||||
}
|
}
|
||||||
|
|
@ -174,6 +218,24 @@ export async function submitCapture(archiveId, locator, quality = null, extensio
|
||||||
return res.json(); // { job_uid, status: "pending" }
|
return res.json(); // { job_uid, status: "pending" }
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export async function submitTextCapture(archiveId, {title, body, mime = 'text/markdown'}) {
|
||||||
|
const payload = { title, body };
|
||||||
|
if (mime && mime !== 'text/markdown') payload.mime = mime;
|
||||||
|
|
||||||
|
const res = await fetch(`/api/archives/${archiveId}/captures/text`, {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": "application/json" },
|
||||||
|
body: JSON.stringify(payload),
|
||||||
|
});
|
||||||
|
if (!res.ok) {
|
||||||
|
const body = await res.json().catch(() => ({}));
|
||||||
|
const err = new Error(body.error || `HTTP ${res.status}`);
|
||||||
|
err.status = res.status;
|
||||||
|
throw err;
|
||||||
|
}
|
||||||
|
return res.json(); // { job_uid, status: "pending" }
|
||||||
|
}
|
||||||
|
|
||||||
// Returns { has_video: bool, qualities: string[] } e.g. { has_video: true, qualities: ["1080p","720p","480p"] }
|
// Returns { has_video: bool, qualities: string[] } e.g. { has_video: true, qualities: ["1080p","720p","480p"] }
|
||||||
// Throws on network error; returns { has_video: false, qualities: [] } on non-video locators.
|
// Throws on network error; returns { has_video: false, qualities: [] } on non-video locators.
|
||||||
export async function probeCapture(archiveId, locator) {
|
export async function probeCapture(archiveId, locator) {
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,5 @@
|
||||||
import { useRef, useEffect, useState, useCallback } from 'react'
|
import { useRef, useEffect, useState, useCallback } from 'react'
|
||||||
import { submitCapture, pollCaptureJob, probeCapture, probePlaylist, getInstanceSettings, uploadFile, deleteUpload } from '../api'
|
import { submitCapture, submitTextCapture, pollCaptureJob, probeCapture, probePlaylist, getInstanceSettings, uploadFile, deleteUpload } from '../api'
|
||||||
|
|
||||||
let nextItemId = 1
|
let nextItemId = 1
|
||||||
|
|
||||||
|
|
@ -157,6 +157,30 @@ function makeFileItem(filename) {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function makeTextItem() {
|
||||||
|
return {
|
||||||
|
id: nextItemId++,
|
||||||
|
kind: 'text',
|
||||||
|
title: '',
|
||||||
|
body: '',
|
||||||
|
mime: 'text/markdown',
|
||||||
|
// Fields present for submission-logic compatibility
|
||||||
|
locator: '',
|
||||||
|
quality: 'best',
|
||||||
|
probeState: 'idle',
|
||||||
|
probeQualities: null,
|
||||||
|
probeHasAudio: false,
|
||||||
|
playlistProbeState: 'idle',
|
||||||
|
playlistInfo: null,
|
||||||
|
playlistItems: null,
|
||||||
|
playlistQuality: null,
|
||||||
|
playlistExpanded: false,
|
||||||
|
syncEnabled: false,
|
||||||
|
error: null,
|
||||||
|
status: 'idle',
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
function applyPlaylistQuality(newQ, currentItems) {
|
function applyPlaylistQuality(newQ, currentItems) {
|
||||||
if (newQ === 'best') {
|
if (newQ === 'best') {
|
||||||
return currentItems.map(item => ({ ...item, quality: 'best' }))
|
return currentItems.map(item => ({ ...item, quality: 'best' }))
|
||||||
|
|
@ -430,9 +454,30 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
onToastRef.current(text, null, type, headline)
|
onToastRef.current(text, null, type, headline)
|
||||||
}
|
}
|
||||||
|
|
||||||
async function submitBgJob(locator, quality, batchId, extraExtensions = {}) {
|
async function submitBgJob(submission, batchId) {
|
||||||
const aid = archiveIdRef.current
|
const aid = archiveIdRef.current
|
||||||
const id = crypto.randomUUID?.() ?? `job-${Date.now()}-${Math.random()}`
|
const id = crypto.randomUUID?.() ?? `job-${Date.now()}-${Math.random()}`
|
||||||
|
|
||||||
|
// Text submission
|
||||||
|
if (submission.type === 'text') {
|
||||||
|
try {
|
||||||
|
const job = await submitTextCapture(aid, { title: submission.title, body: submission.body, mime: submission.mime })
|
||||||
|
const locator = `text:${submission.title}`
|
||||||
|
// Notify App to add skeleton + persist
|
||||||
|
onJobStartedRef.current?.({ id, jobUid: job.job_uid, locator, archiveId: aid })
|
||||||
|
startPolling(id, job.job_uid, locator, aid, batchId)
|
||||||
|
} catch (e) {
|
||||||
|
const msg = e.message || 'Submission failed.'
|
||||||
|
onToastRef.current(msg, `text:${submission.title}`)
|
||||||
|
settleBatch(batchId, 'failed', `text:${submission.title}`)
|
||||||
|
}
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// URL/file submission
|
||||||
|
const locator = submission.locator
|
||||||
|
const quality = submission.quality
|
||||||
|
const extraExtensions = submission.extraExtensions || {}
|
||||||
// Capture session options at call time (synchronous — before first await)
|
// Capture session options at call time (synchronous — before first await)
|
||||||
const extensions = {
|
const extensions = {
|
||||||
ublock_enabled: ublockEnabled,
|
ublock_enabled: ublockEnabled,
|
||||||
|
|
@ -466,16 +511,18 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
// Guard against the Enter-key shortcut in CaptureRow bypassing the
|
// Guard against the Enter-key shortcut in CaptureRow bypassing the
|
||||||
// disabled button — uploads must be complete before archiving starts.
|
// disabled button — uploads must be complete before archiving starts.
|
||||||
if (items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading')) return
|
if (items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading')) return
|
||||||
const toSubmit = items.filter(it =>
|
const toSubmit = items.filter(it => {
|
||||||
it.kind === 'file' ? (it.uploadStatus === 'done' && it.uploadLocator) : it.locator.trim()
|
if (it.kind === 'file') return it.uploadStatus === 'done' && it.uploadLocator
|
||||||
)
|
if (it.kind === 'text') return it.title.trim() && it.body.trim()
|
||||||
|
return it.locator.trim()
|
||||||
|
})
|
||||||
if (toSubmit.length === 0) return
|
if (toSubmit.length === 0) return
|
||||||
if (toSubmit.some(it => it.kind !== 'file' && hasConflict(it))) return
|
if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && hasConflict(it))) return
|
||||||
if (toSubmit.some(it => it.kind !== 'file' && (
|
if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && (
|
||||||
it.probeState === 'probing' ||
|
it.probeState === 'probing' ||
|
||||||
(isPlaylistSource(it.locator) && it.playlistProbeState !== 'done'))))
|
(isPlaylistSource(it.locator) && it.playlistProbeState !== 'done'))))
|
||||||
return
|
return
|
||||||
if (toSubmit.some(it => it.kind !== 'file' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0)) return
|
if (toSubmit.some(it => it.kind !== 'file' && it.kind !== 'text' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0)) return
|
||||||
const batchId = toSubmit.length > 1
|
const batchId = toSubmit.length > 1
|
||||||
? (crypto.randomUUID?.() ?? `batch-${Date.now()}`)
|
? (crypto.randomUUID?.() ?? `batch-${Date.now()}`)
|
||||||
: null
|
: null
|
||||||
|
|
@ -485,9 +532,13 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
// Capture all submission data before any state changes
|
// Capture all submission data before any state changes
|
||||||
const submissions = toSubmit.map(it => {
|
const submissions = toSubmit.map(it => {
|
||||||
if (it.kind === 'file') {
|
if (it.kind === 'file') {
|
||||||
return { locator: it.uploadLocator, quality: 'best', extraExtensions: {} }
|
return { type: 'file', locator: it.uploadLocator, quality: 'best', extraExtensions: {} }
|
||||||
|
}
|
||||||
|
if (it.kind === 'text') {
|
||||||
|
return { type: 'text', title: it.title.trim(), body: it.body, mime: it.mime }
|
||||||
}
|
}
|
||||||
return {
|
return {
|
||||||
|
type: 'url',
|
||||||
locator: it.locator.trim(),
|
locator: it.locator.trim(),
|
||||||
quality: it.playlistItems !== null ? null : (it.quality || 'best'),
|
quality: it.playlistItems !== null ? null : (it.quality || 'best'),
|
||||||
extraExtensions: it.playlistItems !== null
|
extraExtensions: it.playlistItems !== null
|
||||||
|
|
@ -502,8 +553,8 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
setItems([makeItem()])
|
setItems([makeItem()])
|
||||||
dialogRef.current?.close()
|
dialogRef.current?.close()
|
||||||
// Submit each in background
|
// Submit each in background
|
||||||
submissions.forEach(({ locator, quality, extraExtensions }) =>
|
submissions.forEach(submission =>
|
||||||
submitBgJob(locator, quality, batchId, extraExtensions)
|
submitBgJob(submission, batchId)
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -630,8 +681,10 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
files.forEach(file => {
|
files.forEach(file => {
|
||||||
const newItem = makeFileItem(file.name)
|
const newItem = makeFileItem(file.name)
|
||||||
setItems(prev => {
|
setItems(prev => {
|
||||||
// Replace a sole empty URL row with the file item; otherwise append
|
// Only a normal URL row can be replaced. Text drafts deliberately use
|
||||||
if (prev.length === 1 && prev[0].kind !== 'file' && !prev[0].locator.trim()) {
|
// an empty compatibility locator, but their title/body must survive a
|
||||||
|
// file attachment and remain independently archivable.
|
||||||
|
if (prev.length === 1 && !prev[0].kind && !prev[0].locator.trim()) {
|
||||||
return [newItem]
|
return [newItem]
|
||||||
}
|
}
|
||||||
return [...prev, newItem]
|
return [...prev, newItem]
|
||||||
|
|
@ -687,16 +740,18 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
|
|
||||||
|
|
||||||
const anyUploading = items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading')
|
const anyUploading = items.some(it => it.kind === 'file' && it.uploadStatus === 'uploading')
|
||||||
const pendingCount = items.filter(it =>
|
const pendingCount = items.filter(it => {
|
||||||
it.kind === 'file' ? (it.uploadStatus === 'done' && it.uploadLocator) : it.locator.trim()
|
if (it.kind === 'file') return it.uploadStatus === 'done' && it.uploadLocator
|
||||||
).length
|
if (it.kind === 'text') return it.title.trim() && it.body.trim()
|
||||||
const anyConflict = items.some(it => it.kind !== 'file' && hasConflict(it))
|
return it.locator.trim()
|
||||||
|
}).length
|
||||||
|
const anyConflict = items.some(it => it.kind !== 'file' && it.kind !== 'text' && hasConflict(it))
|
||||||
// True if any playlist row has had all its videos deleted — archive would be a no-op.
|
// True if any playlist row has had all its videos deleted — archive would be a no-op.
|
||||||
const anyEmptyPlaylist = items.some(it =>
|
const anyEmptyPlaylist = items.some(it =>
|
||||||
it.kind !== 'file' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0
|
it.kind !== 'file' && it.kind !== 'text' && Array.isArray(it.playlistItems) && it.playlistItems.length === 0
|
||||||
)
|
)
|
||||||
const anyProbing = items.some(it =>
|
const anyProbing = items.some(it =>
|
||||||
it.kind !== 'file' && (
|
it.kind !== 'file' && it.kind !== 'text' && (
|
||||||
it.probeState === 'probing' ||
|
it.probeState === 'probing' ||
|
||||||
// For playlist sources block unless probe completed successfully:
|
// For playlist sources block unless probe completed successfully:
|
||||||
// idle = debounce not yet fired; probing = in flight; error = no quality data.
|
// idle = debounce not yet fired; probing = in flight; error = no quality data.
|
||||||
|
|
@ -735,6 +790,17 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
item={item}
|
item={item}
|
||||||
onRemove={() => removeRow(item.id)}
|
onRemove={() => removeRow(item.id)}
|
||||||
/>
|
/>
|
||||||
|
) : item.kind === 'text' ? (
|
||||||
|
<CaptureTextRow
|
||||||
|
key={item.id}
|
||||||
|
item={item}
|
||||||
|
autoFocus={idx === items.length - 1}
|
||||||
|
onTitleChange={val => setItems(prev => prev.map(it => it.id === item.id ? { ...it, title: val } : it))}
|
||||||
|
onBodyChange={val => setItems(prev => prev.map(it => it.id === item.id ? { ...it, body: val } : it))}
|
||||||
|
onMimeChange={val => setItems(prev => prev.map(it => it.id === item.id ? { ...it, mime: val } : it))}
|
||||||
|
onRemove={() => removeRow(item.id)}
|
||||||
|
onSubmit={handleArchive}
|
||||||
|
/>
|
||||||
) : (
|
) : (
|
||||||
<CaptureRow
|
<CaptureRow
|
||||||
key={item.id}
|
key={item.id}
|
||||||
|
|
@ -781,6 +847,12 @@ export default function CaptureDialog({ open, archiveId, onClose, onCaptured, on
|
||||||
</svg>
|
</svg>
|
||||||
Upload file
|
Upload file
|
||||||
</button>
|
</button>
|
||||||
|
<button type="button" className="capture-add-row capture-add-text" onClick={() => setItems(prev => [...prev, makeTextItem()])}>
|
||||||
|
<svg viewBox="0 0 16 16" fill="none" stroke="currentColor" strokeWidth="1.75" strokeLinecap="round" strokeLinejoin="round">
|
||||||
|
<path d="M2 3h12M2 7h12M2 11h8"/>
|
||||||
|
</svg>
|
||||||
|
Add text
|
||||||
|
</button>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
{/* ── Advanced options ────────────────────────────── */}
|
{/* ── Advanced options ────────────────────────────── */}
|
||||||
|
|
@ -1139,3 +1211,63 @@ function CaptureFileRow({ item, onRemove }) {
|
||||||
</div>
|
</div>
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function CaptureTextRow({ item, autoFocus, onTitleChange, onBodyChange, onMimeChange, onRemove, onSubmit }) {
|
||||||
|
const titleInputRef = useRef(null)
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (autoFocus) {
|
||||||
|
titleInputRef.current?.focus()
|
||||||
|
}
|
||||||
|
}, [autoFocus]) // eslint-disable-line react-hooks/exhaustive-deps
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="capture-row capture-text-row">
|
||||||
|
<div className="capture-row-main">
|
||||||
|
<span className="capture-text-icon" aria-hidden="true">
|
||||||
|
<svg viewBox="0 0 16 16" fill="none" stroke="currentColor" strokeWidth="1.75" strokeLinecap="round" strokeLinejoin="round" style={{ width: 14, height: 14 }}>
|
||||||
|
<path d="M2 3h12M2 7h12M2 11h8"/>
|
||||||
|
</svg>
|
||||||
|
</span>
|
||||||
|
<div className="capture-text-inputs">
|
||||||
|
<input
|
||||||
|
ref={titleInputRef}
|
||||||
|
className="capture-text-input capture-text-title"
|
||||||
|
type="text"
|
||||||
|
placeholder="Title"
|
||||||
|
value={item.title}
|
||||||
|
onChange={e => onTitleChange(e.target.value)}
|
||||||
|
maxLength={500}
|
||||||
|
/>
|
||||||
|
<textarea
|
||||||
|
className="capture-text-input capture-text-body"
|
||||||
|
placeholder="Body (markdown or plain text)"
|
||||||
|
value={item.body}
|
||||||
|
onChange={e => onBodyChange(e.target.value)}
|
||||||
|
rows={6}
|
||||||
|
/>
|
||||||
|
<div className="capture-text-footer">
|
||||||
|
<select
|
||||||
|
className="capture-text-mime"
|
||||||
|
value={item.mime}
|
||||||
|
onChange={e => onMimeChange(e.target.value)}
|
||||||
|
>
|
||||||
|
<option value="text/markdown">Markdown</option>
|
||||||
|
<option value="text/plain">Plain text</option>
|
||||||
|
</select>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="capture-row-action capture-row-remove"
|
||||||
|
onClick={onRemove}
|
||||||
|
aria-label="Remove"
|
||||||
|
>
|
||||||
|
<svg viewBox="0 0 16 16" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round">
|
||||||
|
<line x1="3" y1="3" x2="13" y2="13"/><line x1="13" y1="3" x2="3" y2="13"/>
|
||||||
|
</svg>
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
|
||||||
|
|
@ -1,9 +1,43 @@
|
||||||
import { useState, useEffect, useRef } from 'react'
|
import { useState, useEffect, useLayoutEffect, useRef } from 'react'
|
||||||
import { fetchEntryTags, assignTag, removeTag, listEntryCollections, listCollections, addEntryToCollection, updateEntryTitle, deleteEntry, rearchiveEntry, pollCaptureJob } from '../api'
|
import { fetchEntryTags, assignTag, removeTag, listEntryCollections, listCollections, addEntryToCollection, updateEntryTitle, deleteEntry, rearchiveEntry, pollCaptureJob, fetchEntrySummary, requestEntrySummary } from '../api'
|
||||||
import { formatTimestamp, formatBytes, valueText, sourceIconSvg, displayPath } from '../utils'
|
import { formatTimestamp, formatBytes, valueText, sourceIconSvg, displayPath } from '../utils'
|
||||||
|
|
||||||
const VIS_LABEL = { 0: 'Private', 1: 'Public', 2: 'Users only', 3: 'Public' }
|
const VIS_LABEL = { 0: 'Private', 1: 'Public', 2: 'Users only', 3: 'Public' }
|
||||||
|
|
||||||
|
// Provider labels are display-only; the values are the provider_kind strings
|
||||||
|
// the server persists in entry_summaries.provider_kind.
|
||||||
|
const SUMMARY_PROVIDERS = [
|
||||||
|
{ value: 'anthropic_http', label: 'Anthropic API' },
|
||||||
|
{ value: 'openai_compatible', label: 'OpenAI-compatible API' },
|
||||||
|
{ value: 'claude_cli', label: 'Claude CLI' },
|
||||||
|
{ value: 'codex_cli', label: 'Codex CLI' },
|
||||||
|
]
|
||||||
|
const PROVIDER_LABEL = Object.fromEntries(SUMMARY_PROVIDERS.map(p => [p.value, p.label]))
|
||||||
|
const SUMMARY_PROVIDER_KEY = 'archivr:summary:provider'
|
||||||
|
const SUMMARY_POLL_MS = 1500
|
||||||
|
const UNSUPPORTED_SUMMARY_CONTENT_HEADING = 'This entry can’t be summarized yet.'
|
||||||
|
const UNSUPPORTED_SUMMARY_CONTENT_DETAIL = 'It doesn’t contain archived text that a summary provider can read. Summaries currently support text notes, web pages, X posts and threads, and X Articles. Video, audio, and image-only entries need a transcript or text source.'
|
||||||
|
const UNSUPPORTED_SUMMARY_CONTENT_MESSAGE = `${UNSUPPORTED_SUMMARY_CONTENT_HEADING}\n\n${UNSUPPORTED_SUMMARY_CONTENT_DETAIL}`
|
||||||
|
|
||||||
|
// Summaries are stored as the raw JSON string the model produced (normalized
|
||||||
|
// server-side to {tldr, summary, tags}). Parsing can still fail for rows written
|
||||||
|
// by an older prompt version, so fall back to showing the text as-is rather than
|
||||||
|
// hiding a summary the user can perfectly well read.
|
||||||
|
function parseSummaryText(text) {
|
||||||
|
if (!text) return null
|
||||||
|
try {
|
||||||
|
const parsed = JSON.parse(text)
|
||||||
|
if (parsed && typeof parsed === 'object') {
|
||||||
|
return {
|
||||||
|
tldr: typeof parsed.tldr === 'string' ? parsed.tldr : '',
|
||||||
|
summary: typeof parsed.summary === 'string' ? parsed.summary : '',
|
||||||
|
tags: Array.isArray(parsed.tags) ? parsed.tags.filter(t => typeof t === 'string') : [],
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} catch { /* not JSON — fall through */ }
|
||||||
|
return { tldr: '', summary: text, tags: [] }
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
const ExternalIcon = () => (
|
const ExternalIcon = () => (
|
||||||
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round">
|
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round">
|
||||||
|
|
@ -26,6 +60,30 @@ export default function ContextRail({ archiveId, selectedEntry, selectedUids, se
|
||||||
const [fontsOpen, setFontsOpen] = useState(false)
|
const [fontsOpen, setFontsOpen] = useState(false)
|
||||||
useEffect(() => { setFontsOpen(false) }, [detail?.summary?.entry_uid])
|
useEffect(() => { setFontsOpen(false) }, [detail?.summary?.entry_uid])
|
||||||
|
|
||||||
|
// ── Summary state ───────────────────────────────────────────────────────
|
||||||
|
// A completed summary and its replacement attempt are intentionally separate:
|
||||||
|
// regeneration must not blank or overwrite readable content while it runs.
|
||||||
|
const [summary, setSummary] = useState(null)
|
||||||
|
const [summaryAttempt, setSummaryAttempt] = useState(null)
|
||||||
|
const [summaryError, setSummaryError] = useState('')
|
||||||
|
const [summaryBusy, setSummaryBusy] = useState(false)
|
||||||
|
const [summaryProvider, setSummaryProvider] = useState(() => {
|
||||||
|
try {
|
||||||
|
return sessionStorage.getItem(SUMMARY_PROVIDER_KEY) || SUMMARY_PROVIDERS[0].value
|
||||||
|
} catch { return SUMMARY_PROVIDERS[0].value }
|
||||||
|
})
|
||||||
|
const [includeSummaryImages, setIncludeSummaryImages] = useState(false)
|
||||||
|
const summaryPollRef = useRef(null)
|
||||||
|
const summaryPollAbortRef = useRef(null)
|
||||||
|
const summaryGenerateAbortRef = useRef(null)
|
||||||
|
const summarySelectionRef = useRef(null)
|
||||||
|
// Update before effects run from the list selection, not detail: detail can
|
||||||
|
// briefly describe the previously selected entry while its replacement loads.
|
||||||
|
const summarySelectionKey = archiveId && selectedEntry?.entry_uid
|
||||||
|
? `${archiveId}:${selectedEntry.entry_uid}`
|
||||||
|
: null
|
||||||
|
summarySelectionRef.current = summarySelectionKey
|
||||||
|
|
||||||
// ── Bulk-panel state ────────────────────────────────────────────────────
|
// ── Bulk-panel state ────────────────────────────────────────────────────
|
||||||
const isBulk = selectedUids?.size >= 2
|
const isBulk = selectedUids?.size >= 2
|
||||||
const [bulkTagInput, setBulkTagInput] = useState('')
|
const [bulkTagInput, setBulkTagInput] = useState('')
|
||||||
|
|
@ -76,6 +134,116 @@ export default function ContextRail({ archiveId, selectedEntry, selectedUids, se
|
||||||
}
|
}
|
||||||
}, [])
|
}, [])
|
||||||
|
|
||||||
|
// Seed the summary from the entry detail payload and stop any poll left over
|
||||||
|
// from the previously selected entry.
|
||||||
|
useLayoutEffect(() => {
|
||||||
|
clearInterval(summaryPollRef.current)
|
||||||
|
summaryPollRef.current = null
|
||||||
|
summaryPollAbortRef.current?.abort()
|
||||||
|
summaryPollAbortRef.current = null
|
||||||
|
summaryGenerateAbortRef.current?.abort()
|
||||||
|
summaryGenerateAbortRef.current = null
|
||||||
|
const detailMatchesSelection = detail?.summary?.entry_uid === selectedEntry?.entry_uid
|
||||||
|
setSummary(detailMatchesSelection ? detail.latest_summary ?? null : null)
|
||||||
|
setSummaryAttempt(detailMatchesSelection ? detail.summary_attempt ?? null : null)
|
||||||
|
setSummaryError('')
|
||||||
|
setSummaryBusy(false)
|
||||||
|
setIncludeSummaryImages(false)
|
||||||
|
}, [archiveId, selectedEntry?.entry_uid, detail?.summary?.entry_uid])
|
||||||
|
|
||||||
|
// Poll only while a replacement attempt is non-terminal. Anchoring the effect
|
||||||
|
// on its status means a job still running when the user navigates away and
|
||||||
|
// back is picked up again without displacing completed content.
|
||||||
|
const summaryAttemptStatus = summaryAttempt?.status
|
||||||
|
useEffect(() => {
|
||||||
|
clearInterval(summaryPollRef.current)
|
||||||
|
summaryPollRef.current = null
|
||||||
|
if (summaryAttemptStatus !== 'pending' && summaryAttemptStatus !== 'running') return
|
||||||
|
if (!archiveId || !detail?.summary?.entry_uid) return
|
||||||
|
const entryUid = detail.summary.entry_uid
|
||||||
|
const selectionKey = `${archiveId}:${entryUid}`
|
||||||
|
if (summarySelectionKey !== selectionKey) return
|
||||||
|
const controller = new AbortController()
|
||||||
|
summaryPollAbortRef.current = controller
|
||||||
|
const poll = async () => {
|
||||||
|
try {
|
||||||
|
const res = await fetchEntrySummary(archiveId, entryUid, { signal: controller.signal })
|
||||||
|
if (controller.signal.aborted || summarySelectionRef.current !== selectionKey) return
|
||||||
|
setSummary(res.summary ?? null)
|
||||||
|
setSummaryAttempt(res.attempt ?? null)
|
||||||
|
const st = res.attempt?.status
|
||||||
|
if (st !== 'pending' && st !== 'running') {
|
||||||
|
clearInterval(intervalId)
|
||||||
|
if (summaryPollRef.current === intervalId) summaryPollRef.current = null
|
||||||
|
setSummaryBusy(false)
|
||||||
|
if (st === 'completed' && summarySelectionRef.current === selectionKey) onDetailRefresh?.()
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
if (controller.signal.aborted || summarySelectionRef.current !== selectionKey) return
|
||||||
|
// A transient poll failure is not worth tearing the section down; the
|
||||||
|
// next tick retries, and a real failure lands as status === 'failed'.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const intervalId = setInterval(poll, SUMMARY_POLL_MS)
|
||||||
|
summaryPollRef.current = intervalId
|
||||||
|
return () => {
|
||||||
|
clearInterval(intervalId)
|
||||||
|
if (summaryPollRef.current === intervalId) summaryPollRef.current = null
|
||||||
|
controller.abort()
|
||||||
|
if (summaryPollAbortRef.current === controller) summaryPollAbortRef.current = null
|
||||||
|
}
|
||||||
|
}, [summaryAttemptStatus, archiveId, selectedEntry?.entry_uid, detail?.summary?.entry_uid, summarySelectionKey])
|
||||||
|
|
||||||
|
useEffect(() => () => {
|
||||||
|
clearInterval(summaryPollRef.current)
|
||||||
|
summaryPollAbortRef.current?.abort()
|
||||||
|
summaryGenerateAbortRef.current?.abort()
|
||||||
|
}, [])
|
||||||
|
|
||||||
|
async function handleGenerateSummary(force = false) {
|
||||||
|
if (!archiveId || !detail?.summary?.entry_uid || summaryBusy) return
|
||||||
|
const entryUid = detail.summary.entry_uid
|
||||||
|
const selectionKey = `${archiveId}:${entryUid}`
|
||||||
|
if (summarySelectionRef.current !== selectionKey) return
|
||||||
|
const controller = new AbortController()
|
||||||
|
summaryGenerateAbortRef.current?.abort()
|
||||||
|
summaryGenerateAbortRef.current = controller
|
||||||
|
setSummaryBusy(true)
|
||||||
|
setSummaryError('')
|
||||||
|
try {
|
||||||
|
const res = await requestEntrySummary(archiveId, entryUid, {
|
||||||
|
provider: summaryProvider,
|
||||||
|
force,
|
||||||
|
includeImages: includeSummaryImages,
|
||||||
|
signal: controller.signal,
|
||||||
|
})
|
||||||
|
if (controller.signal.aborted || summarySelectionRef.current !== selectionKey) return
|
||||||
|
if (res.status === 'completed') {
|
||||||
|
// 200 cache hit: the response *is* the row, no polling needed.
|
||||||
|
setSummary(res)
|
||||||
|
setSummaryAttempt(null)
|
||||||
|
setSummaryBusy(false)
|
||||||
|
if (summarySelectionRef.current === selectionKey) onDetailRefresh?.()
|
||||||
|
} else {
|
||||||
|
// 202: seed a local pending row so the poll effect starts immediately
|
||||||
|
// rather than waiting a tick for the first GET.
|
||||||
|
setSummaryAttempt({ ...(res ?? {}), status: 'pending' })
|
||||||
|
}
|
||||||
|
} catch (e) {
|
||||||
|
if (controller.signal.aborted || summarySelectionRef.current !== selectionKey) return
|
||||||
|
setSummaryError(e.message || 'Summary request failed')
|
||||||
|
setSummaryBusy(false)
|
||||||
|
} finally {
|
||||||
|
if (summaryGenerateAbortRef.current === controller) summaryGenerateAbortRef.current = null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function handleProviderChange(value) {
|
||||||
|
setSummaryProvider(value)
|
||||||
|
if (value === 'claude_cli') setIncludeSummaryImages(false)
|
||||||
|
try { sessionStorage.setItem(SUMMARY_PROVIDER_KEY, value) } catch { /* private mode */ }
|
||||||
|
}
|
||||||
|
|
||||||
// Fetch available collections whenever archiveId is available
|
// Fetch available collections whenever archiveId is available
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!archiveId) { setCollections([]); return }
|
if (!archiveId) { setCollections([]); return }
|
||||||
|
|
@ -276,7 +444,7 @@ export default function ContextRail({ archiveId, selectedEntry, selectedUids, se
|
||||||
] : []
|
] : []
|
||||||
|
|
||||||
const AUDIO_EXTS = new Set(['mp3','ogg','m4a','opus','wav','flac','aac'])
|
const AUDIO_EXTS = new Set(['mp3','ogg','m4a','opus','wav','flac','aac'])
|
||||||
const PREVIEW_EXTS = new Set(['mp4','webm','mov','mkv','avi','m4v','ogv','pdf','html','htm','jpg','jpeg','png','gif','webp','avif','svg','bmp'])
|
const PREVIEW_EXTS = new Set(['mp4','webm','mov','mkv','avi','m4v','ogv','pdf','html','htm','md','markdown','txt','jpg','jpeg','png','gif','webp','avif','svg','bmp'])
|
||||||
const primaryMediaIdx = detail ? detail.artifacts.findIndex(a => a.artifact_role === 'primary_media') : -1
|
const primaryMediaIdx = detail ? detail.artifacts.findIndex(a => a.artifact_role === 'primary_media') : -1
|
||||||
const primaryMedia = primaryMediaIdx >= 0 ? detail.artifacts[primaryMediaIdx] : null
|
const primaryMedia = primaryMediaIdx >= 0 ? detail.artifacts[primaryMediaIdx] : null
|
||||||
const pmExt = primaryMedia ? primaryMedia.relpath.split('.').pop().toLowerCase() : ''
|
const pmExt = primaryMedia ? primaryMedia.relpath.split('.').pop().toLowerCase() : ''
|
||||||
|
|
@ -435,6 +603,107 @@ export default function ContextRail({ archiveId, selectedEntry, selectedUids, se
|
||||||
</button>
|
</button>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
|
{(() => {
|
||||||
|
// Public sessions get read-only treatment: the completed text if the
|
||||||
|
// server's visibility gate let the detail through at all, and never
|
||||||
|
// the provider selector or Generate button.
|
||||||
|
const parsed = summary?.status === 'completed'
|
||||||
|
? parseSummaryText(summary.summary_text)
|
||||||
|
: null
|
||||||
|
const running = summaryAttempt?.status === 'pending' || summaryAttempt?.status === 'running'
|
||||||
|
const unsupportedContent =
|
||||||
|
(summaryAttempt?.status === 'failed' && summaryAttempt.error_text === UNSUPPORTED_SUMMARY_CONTENT_MESSAGE) ||
|
||||||
|
summaryError === UNSUPPORTED_SUMMARY_CONTENT_MESSAGE
|
||||||
|
if (isPublicSession && !parsed) return null
|
||||||
|
return (
|
||||||
|
<div className="rail-section rail-summary">
|
||||||
|
<div className="rail-section-heading">Summary</div>
|
||||||
|
|
||||||
|
{parsed && (
|
||||||
|
<div className="rail-summary-body">
|
||||||
|
{parsed.tldr && <p className="rail-summary-tldr">{parsed.tldr}</p>}
|
||||||
|
{parsed.summary && <p className="rail-summary-text">{parsed.summary}</p>}
|
||||||
|
{parsed.tags.length > 0 && (
|
||||||
|
<div className="rail-summary-tags">
|
||||||
|
{parsed.tags.map(t => (
|
||||||
|
<span key={t} className="rail-summary-tag">{t}</span>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
<p className="rail-summary-provider">
|
||||||
|
{PROVIDER_LABEL[summary.provider_kind] || summary.provider_kind}
|
||||||
|
{summary.resolved_model || summary.provider_model
|
||||||
|
? ` \u00b7 ${summary.resolved_model || summary.provider_model}`
|
||||||
|
: ''}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{running && (
|
||||||
|
<p className="rail-summary-status">
|
||||||
|
<span className="rail-summary-spinner" aria-hidden="true" />
|
||||||
|
{'Generating\u2026'}
|
||||||
|
</p>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{unsupportedContent && !isPublicSession && (
|
||||||
|
<div className="rail-summary-info" role="status">
|
||||||
|
<p className="rail-summary-info__heading">{UNSUPPORTED_SUMMARY_CONTENT_HEADING}</p>
|
||||||
|
<p className="rail-summary-info__detail">{UNSUPPORTED_SUMMARY_CONTENT_DETAIL}</p>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
{summaryAttempt?.status === 'failed' && summaryAttempt.error_text && !unsupportedContent && !isPublicSession && (
|
||||||
|
<p className="form-msg form-msg--err rail-summary-error">
|
||||||
|
{summaryAttempt.error_text}
|
||||||
|
</p>
|
||||||
|
)}
|
||||||
|
{summaryError && !unsupportedContent && (
|
||||||
|
<p className="form-msg form-msg--err rail-summary-error">
|
||||||
|
{summaryError}
|
||||||
|
</p>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{!isPublicSession && !running && (
|
||||||
|
<div className="rail-summary-controls">
|
||||||
|
<select
|
||||||
|
className="rail-summary-select"
|
||||||
|
value={summaryProvider}
|
||||||
|
onChange={e => handleProviderChange(e.target.value)}
|
||||||
|
aria-label="Summary provider"
|
||||||
|
>
|
||||||
|
{SUMMARY_PROVIDERS.map(p => (
|
||||||
|
<option key={p.value} value={p.value}>{p.label}</option>
|
||||||
|
))}
|
||||||
|
</select>
|
||||||
|
<div className={`rail-summary-image-option${summaryProvider === 'claude_cli' ? ' rail-summary-image-option--disabled' : ''}`}>
|
||||||
|
<label className="rail-summary-image-option__label">
|
||||||
|
<input
|
||||||
|
type="checkbox"
|
||||||
|
checked={includeSummaryImages}
|
||||||
|
disabled={summaryProvider === 'claude_cli'}
|
||||||
|
onChange={e => setIncludeSummaryImages(e.target.checked)}
|
||||||
|
/>
|
||||||
|
Include attached images
|
||||||
|
</label>
|
||||||
|
<p className="rail-summary-image-option__note">
|
||||||
|
{summaryProvider === 'claude_cli'
|
||||||
|
? 'Claude CLI cannot attach local images. Choose an HTTP provider or Codex CLI.'
|
||||||
|
: 'Selected archived images are sent to the chosen provider. Up to 4 supported images (5 MiB each, 12 MiB total) can be attached; unsupported or oversized artifacts are skipped.'}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<button
|
||||||
|
className="rail-rearchive-btn"
|
||||||
|
onClick={() => handleGenerateSummary(!!parsed)}
|
||||||
|
disabled={summaryBusy}
|
||||||
|
>
|
||||||
|
{summaryBusy ? '\u2026' : parsed ? 'Regenerate' : 'Generate'}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
)
|
||||||
|
})()}
|
||||||
|
|
||||||
<div className="meta-list">
|
<div className="meta-list">
|
||||||
{metaRows.filter(([, v]) => v != null && v !== '').map(([label, value]) => (
|
{metaRows.filter(([, v]) => v != null && v !== '').map(([label, value]) => (
|
||||||
<div key={label} className="meta-item">
|
<div key={label} className="meta-item">
|
||||||
|
|
|
||||||
|
|
@ -16,7 +16,7 @@ export default function EntriesView({ entries, selectedUids, onRowClick, archive
|
||||||
</div>
|
</div>
|
||||||
<div id="entries-body">
|
<div id="entries-body">
|
||||||
{pendingCaptures.filter(c => c.archiveId === archiveId).reverse().map(cap => (
|
{pendingCaptures.filter(c => c.archiveId === archiveId).reverse().map(cap => (
|
||||||
<SkeletonEntryRow key={cap.id} />
|
<SkeletonEntryRow key={cap.id} locator={cap.locator} />
|
||||||
))}
|
))}
|
||||||
{entries.map((entry, idx) => (
|
{entries.map((entry, idx) => (
|
||||||
<EntryRow
|
<EntryRow
|
||||||
|
|
|
||||||
|
|
@ -2,10 +2,12 @@ import VideoPreview from './VideoPreview';
|
||||||
import IframePreview from './IframePreview';
|
import IframePreview from './IframePreview';
|
||||||
import ImagePreview from './ImagePreview';
|
import ImagePreview from './ImagePreview';
|
||||||
import TweetPreview from './TweetPreview';
|
import TweetPreview from './TweetPreview';
|
||||||
|
import TextPreview from './TextPreview';
|
||||||
|
|
||||||
const VIDEO_EXTS = new Set(['mp4', 'webm', 'mov', 'mkv', 'avi', 'm4v', 'ogv']);
|
const VIDEO_EXTS = new Set(['mp4', 'webm', 'mov', 'mkv', 'avi', 'm4v', 'ogv']);
|
||||||
const AUDIO_EXTS = new Set(['mp3', 'ogg', 'm4a', 'opus', 'wav', 'flac', 'aac']);
|
const AUDIO_EXTS = new Set(['mp3', 'ogg', 'm4a', 'opus', 'wav', 'flac', 'aac']);
|
||||||
const IMAGE_EXTS = new Set(['jpg', 'jpeg', 'png', 'gif', 'webp', 'avif', 'svg', 'bmp']);
|
const IMAGE_EXTS = new Set(['jpg', 'jpeg', 'png', 'gif', 'webp', 'avif', 'svg', 'bmp']);
|
||||||
|
const TEXT_EXTS = new Set(['md', 'markdown', 'txt']);
|
||||||
const CONTENT_TYPES = {
|
const CONTENT_TYPES = {
|
||||||
mp4: 'video/mp4', webm: 'video/webm', mov: 'video/quicktime',
|
mp4: 'video/mp4', webm: 'video/webm', mov: 'video/quicktime',
|
||||||
mkv: 'video/x-matroska', avi: 'video/x-msvideo', m4v: 'video/mp4', ogv: 'video/ogg',
|
mkv: 'video/x-matroska', avi: 'video/x-msvideo', m4v: 'video/mp4', ogv: 'video/ogg',
|
||||||
|
|
@ -131,7 +133,20 @@ export default function PreviewPanel({ archiveId, entry, detail, fullPage, onXAr
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
// 7. Image
|
// 6b. Plain text / Markdown
|
||||||
|
if (TEXT_EXTS.has(ext)) {
|
||||||
|
return (
|
||||||
|
<div className="preview-panel" style={{ flex: 1, minHeight: 0, overflow: 'auto' }}>
|
||||||
|
<TextPreview
|
||||||
|
src={primaryMediaUrl}
|
||||||
|
mime={primaryArtifact.mime_type}
|
||||||
|
title={summary.title}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// 7. Image
|
||||||
if (IMAGE_EXTS.has(ext)) {
|
if (IMAGE_EXTS.has(ext)) {
|
||||||
return (
|
return (
|
||||||
<div className="preview-panel" style={{ height: '100%' }}>
|
<div className="preview-panel" style={{ height: '100%' }}>
|
||||||
|
|
|
||||||
|
|
@ -1,23 +1,37 @@
|
||||||
export default function SkeletonEntryRow() {
|
const COLLECTION_MARKERS = [
|
||||||
return (
|
'list=',
|
||||||
<div className="skeleton-row">
|
'/playlist/',
|
||||||
<div className="col-check" aria-hidden="true" />
|
'/channel/',
|
||||||
<div className="col-added">
|
'yt:playlist:',
|
||||||
<span className="skeleton-cell" style={{ width: 108, height: 13 }} />
|
'yt:channel:',
|
||||||
</div>
|
'ytm:playlist:',
|
||||||
<div className="col-title" style={{ gap: '0.42em', display: 'flex', alignItems: 'center' }}>
|
'spotify:playlist:',
|
||||||
<span className="skeleton-cell" style={{ width: 14, height: 14, borderRadius: '50%', flexShrink: 0 }} />
|
'spotify:album:',
|
||||||
<span className="skeleton-cell" style={{ width: '58%', height: 13 }} />
|
];
|
||||||
</div>
|
|
||||||
<div className="col-type">
|
function isCollectionLocator(locator) {
|
||||||
<span className="skeleton-cell" style={{ width: 58, height: 20, borderRadius: 99 }} />
|
const normalized = locator.toLowerCase();
|
||||||
</div>
|
return COLLECTION_MARKERS.some(marker => normalized.includes(marker));
|
||||||
<div className="col-size">
|
}
|
||||||
<span className="skeleton-cell" style={{ width: 44, height: 12 }} />
|
|
||||||
</div>
|
function truncateLocator(locator) {
|
||||||
<div className="col-url">
|
return locator.length > 80 ? `${locator.slice(0, 79)}…` : locator;
|
||||||
<span className="skeleton-cell" style={{ width: '65%', height: 12 }} />
|
}
|
||||||
</div>
|
|
||||||
</div>
|
export default function SkeletonEntryRow({ locator = '' }) {
|
||||||
)
|
const locatorText = String(locator);
|
||||||
|
const isCollection = isCollectionLocator(locatorText);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="in-progress-entry-row" role="status" aria-live="polite">
|
||||||
|
<span className="cap-spinner in-progress-entry-row__spinner" aria-hidden="true" />
|
||||||
|
<span className="in-progress-entry-row__locator" title={locatorText}>
|
||||||
|
{truncateLocator(locatorText)}
|
||||||
|
</span>
|
||||||
|
<span className="in-progress-entry-row__status">
|
||||||
|
Archiving…
|
||||||
|
{isCollection && <span className="in-progress-entry-row__kind">(playlist)</span>}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
|
||||||
31
frontend/src/components/SkeletonEntryRow.stories.jsx
Normal file
31
frontend/src/components/SkeletonEntryRow.stories.jsx
Normal file
|
|
@ -0,0 +1,31 @@
|
||||||
|
import SkeletonEntryRow from './SkeletonEntryRow';
|
||||||
|
|
||||||
|
export default {
|
||||||
|
component: SkeletonEntryRow,
|
||||||
|
tags: ['autodocs'],
|
||||||
|
parameters: { layout: 'padded' },
|
||||||
|
};
|
||||||
|
|
||||||
|
function PendingRows({ locators }) {
|
||||||
|
return (
|
||||||
|
<div className="entry-table">
|
||||||
|
<div id="entries-body">
|
||||||
|
{locators.map(locator => (
|
||||||
|
<SkeletonEntryRow key={locator} locator={locator} />
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
export const Examples = {
|
||||||
|
render: () => (
|
||||||
|
<PendingRows
|
||||||
|
locators={[
|
||||||
|
'https://example.com/articles/a-small-and-useful-page',
|
||||||
|
'https://www.youtube.com/watch?v=abc123&list=PL1234567890',
|
||||||
|
'tweet:1891234567890123456',
|
||||||
|
]}
|
||||||
|
/>
|
||||||
|
),
|
||||||
|
};
|
||||||
43
frontend/src/components/SkeletonEntryRow.test.jsx
Normal file
43
frontend/src/components/SkeletonEntryRow.test.jsx
Normal file
|
|
@ -0,0 +1,43 @@
|
||||||
|
import { describe, expect, test } from 'bun:test';
|
||||||
|
import { renderToStaticMarkup } from 'react-dom/server';
|
||||||
|
|
||||||
|
import SkeletonEntryRow from './SkeletonEntryRow';
|
||||||
|
|
||||||
|
describe('SkeletonEntryRow', () => {
|
||||||
|
test('renders a compact archiving status row for a locator', () => {
|
||||||
|
const markup = renderToStaticMarkup(
|
||||||
|
<SkeletonEntryRow locator="tweet:1891234567890123456" />,
|
||||||
|
);
|
||||||
|
|
||||||
|
expect(markup).toContain('in-progress-entry-row__spinner');
|
||||||
|
expect(markup).toContain('tweet:1891234567890123456');
|
||||||
|
expect(markup).toContain('Archiving…');
|
||||||
|
expect(markup).not.toContain('(playlist)');
|
||||||
|
});
|
||||||
|
|
||||||
|
test('labels playlist and channel locators', () => {
|
||||||
|
const collectionLocators = [
|
||||||
|
'https://www.youtube.com/watch?v=abc123&list=PL123',
|
||||||
|
'https://www.youtube.com/playlist/PL123',
|
||||||
|
'https://www.youtube.com/channel/UC123',
|
||||||
|
'yt:playlist:PL123',
|
||||||
|
'yt:channel:UC123',
|
||||||
|
'ytm:playlist:PL123',
|
||||||
|
'spotify:playlist:123',
|
||||||
|
'spotify:album:123',
|
||||||
|
];
|
||||||
|
|
||||||
|
for (const locator of collectionLocators) {
|
||||||
|
const markup = renderToStaticMarkup(<SkeletonEntryRow locator={locator} />);
|
||||||
|
expect(markup).toContain('(playlist)');
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('truncates long locators to 80 displayed characters', () => {
|
||||||
|
const locator = 'x'.repeat(100);
|
||||||
|
const markup = renderToStaticMarkup(<SkeletonEntryRow locator={locator} />);
|
||||||
|
const visibleLocator = markup.match(/in-progress-entry-row__locator"[^>]*>(.*?)<\/span>/)?.[1];
|
||||||
|
|
||||||
|
expect(visibleLocator).toBe(`${'x'.repeat(79)}…`);
|
||||||
|
});
|
||||||
|
});
|
||||||
45
frontend/src/components/TextPreview.jsx
Normal file
45
frontend/src/components/TextPreview.jsx
Normal file
|
|
@ -0,0 +1,45 @@
|
||||||
|
import { useEffect, useState } from 'react'
|
||||||
|
import { fetchArtifactText } from '../api'
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Renders the primary_media artifact of a text/document entry as plain text.
|
||||||
|
*
|
||||||
|
* v1 intentionally does NOT render Markdown — we don't want a Markdown parser
|
||||||
|
* dep just to unblock the "no preview available" fallback, and monospace text
|
||||||
|
* with visible fences reads fine for the note-length content this feature
|
||||||
|
* captures. Bump to a real renderer if/when Markdown-authored entries grow.
|
||||||
|
*/
|
||||||
|
export default function TextPreview({ src, mime, title }) {
|
||||||
|
const [text, setText] = useState(null)
|
||||||
|
const [error, setError] = useState(null)
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
const controller = new AbortController()
|
||||||
|
setText(null)
|
||||||
|
setError(null)
|
||||||
|
fetchArtifactText(src, { signal: controller.signal })
|
||||||
|
.then(setText)
|
||||||
|
.catch(e => {
|
||||||
|
if (e.name !== 'AbortError') setError(e.message || String(e))
|
||||||
|
})
|
||||||
|
return () => controller.abort()
|
||||||
|
}, [src])
|
||||||
|
|
||||||
|
if (error) {
|
||||||
|
return (
|
||||||
|
<div className="text-preview text-preview--error">
|
||||||
|
Failed to load text: {error}
|
||||||
|
</div>
|
||||||
|
)
|
||||||
|
}
|
||||||
|
if (text === null) {
|
||||||
|
return <div className="text-preview text-preview--loading">Loading…</div>
|
||||||
|
}
|
||||||
|
return (
|
||||||
|
<div className="text-preview">
|
||||||
|
{title && <h1 className="text-preview__title">{title}</h1>}
|
||||||
|
<pre className="text-preview__body">{text}</pre>
|
||||||
|
{mime && <div className="text-preview__mime">{mime}</div>}
|
||||||
|
</div>
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
@ -302,10 +302,10 @@ select {
|
||||||
border-bottom: 1px solid var(--line-soft);
|
border-bottom: 1px solid var(--line-soft);
|
||||||
}
|
}
|
||||||
#entries-body > div > div { padding: 7px 10px; flex-shrink: 0; overflow: hidden; }
|
#entries-body > div > div { padding: 7px 10px; flex-shrink: 0; overflow: hidden; }
|
||||||
/* Skeleton rows (no index class) fall back to nth-child; real rows use explicit classes. */
|
/* Pending capture rows (no index class) fall back to nth-child; real rows use explicit classes. */
|
||||||
#entries-body > div:not(.entry-row-outer):nth-child(even) { background: #f2ede5; }
|
#entries-body > div:not(.entry-row-outer):nth-child(even) { background: #f2ede5; }
|
||||||
#entries-body > div:not(.entry-row-outer):nth-child(odd) { background: var(--paper-3); }
|
#entries-body > div:not(.entry-row-outer):nth-child(odd) { background: var(--paper-3); }
|
||||||
/* Index-based stripes for real entry rows — immune to skeleton sibling count. */
|
/* Index-based stripes for real entry rows — immune to pending-capture sibling count. */
|
||||||
#entries-body > .entry-row-outer--light { background: var(--paper-3); }
|
#entries-body > .entry-row-outer--light { background: var(--paper-3); }
|
||||||
#entries-body > .entry-row-outer--dark { background: #f2ede5; }
|
#entries-body > .entry-row-outer--dark { background: #f2ede5; }
|
||||||
#entries-body > div.is-selected {
|
#entries-body > div.is-selected {
|
||||||
|
|
@ -952,6 +952,104 @@ select {
|
||||||
line-height: 1.45;
|
line-height: 1.45;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ── Text-capture row ───────────────────────────────────────────────────
|
||||||
|
A text row breaks the [dot][one-line-input][×] shape used by URL and
|
||||||
|
file rows: it stacks a title, a multi-line body, and a small footer
|
||||||
|
(mime selector). Everything lives inside .capture-text-inputs so the
|
||||||
|
outer flex-row still gives us [icon][content][×] alignment. */
|
||||||
|
|
||||||
|
.capture-text-row > .capture-row-main {
|
||||||
|
/* Icon and remove button sit at the top of the block, not centered on
|
||||||
|
the tall textarea. */
|
||||||
|
align-items: flex-start;
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-icon {
|
||||||
|
flex-shrink: 0;
|
||||||
|
width: 20px;
|
||||||
|
height: 44px; /* matches .capture-text-title height */
|
||||||
|
display: grid;
|
||||||
|
place-items: center;
|
||||||
|
color: var(--muted);
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-inputs {
|
||||||
|
flex: 1 1 auto;
|
||||||
|
min-width: 0; /* let the textarea shrink inside flex */
|
||||||
|
display: flex;
|
||||||
|
flex-direction: column;
|
||||||
|
gap: 8px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-input {
|
||||||
|
width: 100%;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
background: var(--field);
|
||||||
|
color: var(--ink);
|
||||||
|
border-radius: var(--r2);
|
||||||
|
outline: none;
|
||||||
|
transition: border-color .15s ease, box-shadow .15s ease;
|
||||||
|
font-family: inherit;
|
||||||
|
}
|
||||||
|
.capture-text-input:focus {
|
||||||
|
border-color: var(--accent);
|
||||||
|
box-shadow: 0 0 0 3px color-mix(in srgb, var(--accent) 14%, transparent);
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-title {
|
||||||
|
height: 44px;
|
||||||
|
padding: 0 14px;
|
||||||
|
font-size: 15px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-body {
|
||||||
|
min-height: 140px;
|
||||||
|
padding: 10px 14px;
|
||||||
|
font-size: 14px;
|
||||||
|
line-height: 1.55;
|
||||||
|
resize: vertical;
|
||||||
|
}
|
||||||
|
.capture-text-body::placeholder,
|
||||||
|
.capture-text-title::placeholder {
|
||||||
|
color: var(--muted);
|
||||||
|
}
|
||||||
|
|
||||||
|
.capture-text-footer {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 8px;
|
||||||
|
justify-content: flex-end; /* mime selector sits under the body */
|
||||||
|
}
|
||||||
|
|
||||||
|
/* Mime selector: reuse the small chip styling of .capture-quality. */
|
||||||
|
.capture-text-mime {
|
||||||
|
flex-shrink: 0;
|
||||||
|
height: 28px;
|
||||||
|
padding: 0 26px 0 10px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
border-radius: var(--r);
|
||||||
|
background: var(--paper);
|
||||||
|
color: var(--ink);
|
||||||
|
font-size: 12px;
|
||||||
|
cursor: pointer;
|
||||||
|
outline: none;
|
||||||
|
transition: border-color .15s;
|
||||||
|
/* Custom caret so it doesn't look like an unstyled OS dropdown. */
|
||||||
|
appearance: none;
|
||||||
|
-webkit-appearance: none;
|
||||||
|
background-image: url("data:image/svg+xml;utf8,<svg xmlns='http://www.w3.org/2000/svg' width='12' height='12' viewBox='0 0 12 12' fill='none' stroke='%23888' stroke-width='1.75' stroke-linecap='round' stroke-linejoin='round'><path d='M3 4.5l3 3 3-3'/></svg>");
|
||||||
|
background-repeat: no-repeat;
|
||||||
|
background-position: right 8px center;
|
||||||
|
}
|
||||||
|
.capture-text-mime:hover { border-color: var(--accent-2); }
|
||||||
|
.capture-text-mime:focus { border-color: var(--accent); }
|
||||||
|
|
||||||
|
/* Remove button on text row: pin to top so tall bodies don't push it. */
|
||||||
|
.capture-text-row .capture-row-action {
|
||||||
|
margin-top: 8px;
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
/* Add-another button */
|
/* Add-another button */
|
||||||
.capture-add-row {
|
.capture-add-row {
|
||||||
display: flex;
|
display: flex;
|
||||||
|
|
@ -2836,36 +2934,44 @@ google-cast-launcher.video-tv-btn {
|
||||||
/* Push content above the fixed AudioBar when it is visible */
|
/* Push content above the fixed AudioBar when it is visible */
|
||||||
body.has-audio-bar { padding-bottom: 56px; }
|
body.has-audio-bar { padding-bottom: 56px; }
|
||||||
|
|
||||||
/* ── Skeleton entry row ─────────────────────────────────────────────────── */
|
/* ── In-progress entry row ──────────────────────────────────────────────── */
|
||||||
@keyframes skeleton-shimmer {
|
.in-progress-entry-row {
|
||||||
0% { background-position: 200% center; }
|
box-sizing: border-box;
|
||||||
100% { background-position: -200% center; }
|
min-width: 0;
|
||||||
|
padding: 7px 22px;
|
||||||
|
gap: 9px;
|
||||||
|
color: var(--muted);
|
||||||
|
white-space: nowrap;
|
||||||
}
|
}
|
||||||
|
.in-progress-entry-row__spinner {
|
||||||
.skeleton-cell {
|
width: 16px;
|
||||||
display: inline-block;
|
height: 16px;
|
||||||
border-radius: 3px;
|
border-width: 2px;
|
||||||
background: linear-gradient(90deg, var(--paper-2) 25%, var(--line-soft) 50%, var(--paper-2) 75%);
|
color: var(--accent);
|
||||||
background-size: 200% 100%;
|
|
||||||
animation: skeleton-shimmer 1.8s ease-in-out infinite;
|
|
||||||
}
|
|
||||||
|
|
||||||
/* Row container — flex row matching #entries-body > div layout */
|
|
||||||
.skeleton-row {
|
|
||||||
display: flex;
|
|
||||||
align-items: center;
|
|
||||||
background: var(--paper-3);
|
|
||||||
}
|
|
||||||
.skeleton-row > div {
|
|
||||||
padding: 7px 10px;
|
|
||||||
flex-shrink: 0;
|
flex-shrink: 0;
|
||||||
overflow: hidden;
|
|
||||||
}
|
}
|
||||||
.skeleton-row .col-added { padding-left: 22px; }
|
.in-progress-entry-row__locator {
|
||||||
.skeleton-row > div:last-child { padding-right: 22px; }
|
min-width: 0;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
font-family: ui-monospace, "SF Mono", Menlo, monospace;
|
||||||
|
font-size: 11.5px;
|
||||||
|
}
|
||||||
|
.in-progress-entry-row__status {
|
||||||
|
display: inline-flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 5px;
|
||||||
|
flex-shrink: 0;
|
||||||
|
color: var(--muted-2);
|
||||||
|
font-size: 11.5px;
|
||||||
|
}
|
||||||
|
.in-progress-entry-row__kind {
|
||||||
|
font-size: 10.5px;
|
||||||
|
opacity: 0.85;
|
||||||
|
}
|
||||||
|
|
||||||
@media (pointer: coarse) {
|
@media (pointer: coarse) {
|
||||||
.skeleton-row .col-added { padding-left: 10px; }
|
.in-progress-entry-row { padding-inline: 10px; }
|
||||||
}
|
}
|
||||||
|
|
||||||
/* ── Child entry expansion ───────────────────────────────────────────────── */
|
/* ── Child entry expansion ───────────────────────────────────────────────── */
|
||||||
|
|
@ -3118,3 +3224,172 @@ body.has-audio-bar { padding-bottom: 56px; }
|
||||||
cursor: pointer;
|
cursor: pointer;
|
||||||
}
|
}
|
||||||
.capture-sync-row input[type=checkbox] { cursor: pointer; }
|
.capture-sync-row input[type=checkbox] { cursor: pointer; }
|
||||||
|
|
||||||
|
/* ── Summary rail section ────────────────────────────────────────────────── */
|
||||||
|
/* Reuses .rail-section spacing and .rail-rearchive-btn for the action button;
|
||||||
|
only the summary-specific typography and the provider row are new here. */
|
||||||
|
.rail-summary-body { margin-bottom: 10px; }
|
||||||
|
.rail-summary-tldr {
|
||||||
|
margin: 0 0 8px;
|
||||||
|
font-size: 13.5px;
|
||||||
|
font-weight: 600;
|
||||||
|
color: var(--ink);
|
||||||
|
line-height: 1.45;
|
||||||
|
}
|
||||||
|
.rail-summary-text {
|
||||||
|
margin: 0 0 8px;
|
||||||
|
font-size: 13px;
|
||||||
|
color: var(--ink);
|
||||||
|
line-height: 1.55;
|
||||||
|
}
|
||||||
|
.rail-summary-tags { display: flex; flex-wrap: wrap; gap: 5px; margin-bottom: 8px; }
|
||||||
|
.rail-summary-tag {
|
||||||
|
font-size: 11px;
|
||||||
|
padding: 2px 7px;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
border-radius: 999px;
|
||||||
|
color: var(--muted);
|
||||||
|
}
|
||||||
|
.rail-summary-provider {
|
||||||
|
margin: 0;
|
||||||
|
font-size: 11px;
|
||||||
|
color: var(--muted-2);
|
||||||
|
letter-spacing: 0.02em;
|
||||||
|
}
|
||||||
|
.rail-summary-status {
|
||||||
|
display: flex; align-items: center; gap: 7px;
|
||||||
|
margin: 0 0 8px;
|
||||||
|
font-size: 12.5px;
|
||||||
|
color: var(--muted);
|
||||||
|
}
|
||||||
|
.rail-summary-info {
|
||||||
|
margin: 0 0 8px;
|
||||||
|
padding: 8px;
|
||||||
|
color: var(--muted);
|
||||||
|
background: var(--paper-2);
|
||||||
|
border: 1px solid var(--line-soft);
|
||||||
|
border-radius: 4px;
|
||||||
|
}
|
||||||
|
.rail-summary-info__heading {
|
||||||
|
margin: 0 0 4px;
|
||||||
|
color: var(--ink);
|
||||||
|
font-size: 12.5px;
|
||||||
|
font-weight: 600;
|
||||||
|
}
|
||||||
|
.rail-summary-info__detail {
|
||||||
|
margin: 0;
|
||||||
|
font-size: 12px;
|
||||||
|
line-height: 1.45;
|
||||||
|
}
|
||||||
|
.rail-summary-error { margin: 0 0 8px; }
|
||||||
|
.rail-summary-spinner {
|
||||||
|
width: 11px; height: 11px;
|
||||||
|
border: 1.5px solid var(--line);
|
||||||
|
border-top-color: var(--muted);
|
||||||
|
border-radius: 50%;
|
||||||
|
animation: rail-summary-spin 0.7s linear infinite;
|
||||||
|
flex-shrink: 0;
|
||||||
|
}
|
||||||
|
@keyframes rail-summary-spin { to { transform: rotate(360deg); } }
|
||||||
|
/* Respect a reduced-motion preference: the text alone still conveys the state. */
|
||||||
|
@media (prefers-reduced-motion: reduce) {
|
||||||
|
.rail-summary-spinner { animation: none; }
|
||||||
|
}
|
||||||
|
.rail-summary-controls { display: flex; flex-direction: column; gap: 6px; }
|
||||||
|
.rail-summary-select {
|
||||||
|
width: 100%;
|
||||||
|
padding: 5px 8px;
|
||||||
|
font-size: 12.5px;
|
||||||
|
color: var(--ink);
|
||||||
|
background: var(--paper);
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
border-radius: 4px;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option {
|
||||||
|
padding: 7px 8px;
|
||||||
|
color: var(--muted);
|
||||||
|
background: var(--paper);
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
border-radius: 4px;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option__label {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 6px;
|
||||||
|
color: var(--ink);
|
||||||
|
font-size: 12.5px;
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option__label input[type="checkbox"] {
|
||||||
|
margin: 0;
|
||||||
|
accent-color: var(--accent);
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option__label input[type="checkbox"]:focus-visible {
|
||||||
|
outline: 2px solid var(--accent);
|
||||||
|
outline-offset: 2px;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option__note {
|
||||||
|
margin: 5px 0 0;
|
||||||
|
font-size: 11px;
|
||||||
|
line-height: 1.4;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option--disabled {
|
||||||
|
color: var(--muted-2);
|
||||||
|
background: var(--paper);
|
||||||
|
}
|
||||||
|
.rail-summary-image-option--disabled .rail-summary-image-option__label {
|
||||||
|
color: var(--muted-2);
|
||||||
|
cursor: not-allowed;
|
||||||
|
}
|
||||||
|
.rail-summary-image-option--disabled .rail-summary-image-option__label input[type="checkbox"] {
|
||||||
|
cursor: not-allowed;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── Text / Markdown preview ────────────────────────────────────────────── */
|
||||||
|
.text-preview {
|
||||||
|
padding: 32px 40px 48px;
|
||||||
|
max-width: 780px;
|
||||||
|
margin: 0 auto;
|
||||||
|
font-family: var(--sans);
|
||||||
|
color: var(--ink);
|
||||||
|
overflow: auto;
|
||||||
|
height: 100%;
|
||||||
|
box-sizing: border-box;
|
||||||
|
}
|
||||||
|
.text-preview--loading,
|
||||||
|
.text-preview--error {
|
||||||
|
padding: 24px;
|
||||||
|
color: var(--muted);
|
||||||
|
font-size: 13px;
|
||||||
|
}
|
||||||
|
.text-preview--error { color: var(--accent); }
|
||||||
|
.text-preview__title {
|
||||||
|
font-family: var(--serif, var(--sans));
|
||||||
|
font-size: 22px;
|
||||||
|
font-weight: 600;
|
||||||
|
margin: 0 0 20px;
|
||||||
|
line-height: 1.25;
|
||||||
|
color: var(--ink);
|
||||||
|
}
|
||||||
|
.text-preview__body {
|
||||||
|
white-space: pre-wrap;
|
||||||
|
overflow-wrap: anywhere;
|
||||||
|
word-break: break-word;
|
||||||
|
font-family: ui-monospace, "SF Mono", Menlo, Consolas, monospace;
|
||||||
|
font-size: 13.5px;
|
||||||
|
line-height: 1.65;
|
||||||
|
color: var(--ink);
|
||||||
|
background: var(--paper);
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
border-radius: var(--r2);
|
||||||
|
padding: 16px 18px;
|
||||||
|
margin: 0;
|
||||||
|
}
|
||||||
|
.text-preview__mime {
|
||||||
|
margin-top: 12px;
|
||||||
|
font-size: 11px;
|
||||||
|
color: var(--muted);
|
||||||
|
letter-spacing: .04em;
|
||||||
|
text-align: right;
|
||||||
|
}
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue