`update` fetches the latest release tag from the GitHub API (or takes
--version), downloads the cross-platform python zipapp, and installs it
into archivr's state dir. The install is atomic — staged as yt-dlp.new,
chmod +x'd, then renamed over the target — so a concurrently running
capture never sees a half-written binary. A sibling .version file makes a
repeat update a no-op instead of a 3MB re-download.
The download is checked for the python3 shebang before install, which
catches the usual failure mode of getting an HTML error page back. python3
itself is only warned about, not required: the server may run under a nix
wrapper with its own PATH.
`status` prints all three candidates (env / state-dir / PATH fallback) with
their versions and stars whichever the resolver picks, so it is obvious
which yt-dlp a capture will actually use.
reqwest is pulled from the existing workspace dependency; the GitHub JSON is
parsed with serde_json so the "json" feature is not needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The flake no longer takes yt-dlp from nixpkgs; a dedicated `ytDlp`
derivation fetches the upstream release binary directly and pins both
`version` and an SRI `hash`. That makes the previous workflow inert: it
ran `nix flake update nixpkgs` and compared `nixpkgs#yt-dlp.version`
before and after, so it could churn the lockfile forever without ever
moving the version we actually ship.
The workflow now reads the pinned version straight out of the `ytDlp`
block in flake.nix, asks the GitHub API for yt-dlp's latest release tag,
short-circuits when they already match, downloads the new release to
recompute its SRI hash (required — the hash is part of the derivation's
identity, so the URL cannot be changed alone), and rewrites the three
pinned fields under a sed range address scoped to that block so sibling
pins like ublockLite and isdcac are untouched. It asserts only flake.nix
changed and that the new version appears exactly twice before opening
the PR.
The nix flake wrapper pins a yt-dlp via ARCHIVR_YT_DLP, but yt-dlp rots
fast — extractors break within weeks of a pin. Add a resolver that probes
`--version` on both the pinned binary and a user-installed copy under the
mutable state dir, and runs whichever is newer.
Version strings are YYYY.MM.DD, so plain string ordering is chronological.
Ties resolve toward the state dir: a user who installed it there did so
deliberately. ARCHIVR_YT_DLP_FORCE bypasses the comparison entirely, and
with no candidate at all we fall back to bare `yt-dlp` on PATH — exactly
the previous behaviour.
Resolution is cached in a OnceLock so `--version` costs one subprocess per
process, and all four inline env::var lookups now go through it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two independent problems, one commit:
1. Stale binary. nixpkgs-provided `pkgs.yt-dlp` on the pinned
nixos-unstable rev is 2026.03.17 (Mar 2026). yt-dlp itself
releases days-to-weeks, and YouTube frequently rotates the
player-signature / client surfaces the older builds request
(`android_vr` is the current casualty), which returns HTTP 403
mid-download for the format specs archivr passes (`-f
bestvideo+bestaudio/best`). Even bumping the nixpkgs input would
leave us dependent on that channel's yt-dlp cadence.
Fetch the upstream zipapp directly instead
(github.com/yt-dlp/yt-dlp/releases/download/<ver>/yt-dlp), wrap so
`python3` and `ffmpeg` are on PATH, and pin version+hash in one
place. Bumping is: change version, replace hash from
`nix hash file <url>`.
2. Missing pin in server wrapper. `archivr-cli` was already wrapped
with `--set ARCHIVR_YT_DLP` + a PATH prefix; `archivr-server`
was NOT — it only pinned single-file, chrome, and the tweet
scraper, silently falling back to whatever `yt-dlp` the user
happened to have on PATH. Server captures therefore inherited
the user's (often stale) system yt-dlp regardless of the flake
pin. Same wrapper flags now apply to both binaries.
devShell keeps `pkgs.yt-dlp` for now: the dev shell is a
convenience, not a release surface, and matching wouldn't fit in this
commit without duplicating the derivation across let-scopes.
Text captures land as `.md` (Markdown) or `.txt` (plain) blobs, but
PreviewPanel only dispatched on video/audio/image/pdf/html extensions,
so opening a text entry hit the 'No preview available' fallback with
the raw artifact path exposed.
- New `TextPreview` component fetches the primary artifact as text,
renders it in a monospace `<pre>` with word-wrap, and shows the
entry title on top and the MIME as a small trailing tag. Handles
loading/error states.
- `PreviewPanel` gains a `TEXT_EXTS` set + a branch that dispatches
to `TextPreview` for `md` / `markdown` / `txt`.
- CSS is padded and centered to ~780px so a text note reads like a
document rather than an edge-to-edge terminal dump.
v1 intentionally does NOT parse Markdown: keeping frontend deps at
react+react-dom only. Bump to a real Markdown renderer if we start
capturing Markdown-authored notes.
Tweet and tweet_thread entries store their payload under artifact_role
`raw_tweet_json`, not `primary_media`. `build_summary_input` filtered
strictly for `primary_media LIMIT 1`, so both cases silently failed
with 'entry X has no primary_media artifact to summarize'.
Threads compound the problem: the tweet scraper writes ONE json file
per status, so even a fixed lookup that took the first row would
summarize only the initial tweet and lose the rest of the conversation.
Fixes:
- New `load_summary_artifacts` helper returns every artifact for a
role in insertion order.
- For entity_kind `tweet` / `tweet_thread`, load all
`raw_tweet_json` artifacts (falling back to `primary_media` for
archives predating that role convention).
- Iterate artifacts, extract text per file with the existing
markdown/html/json branches, then join thread pieces with a
`---` separator so the model sees a real paragraph break between
statuses instead of one flowing document.
Single-tweet entries produce one piece and the separator never
renders. Non-tweet entries behave exactly as before.
The text row shipped with semantic classnames (`capture-text-inputs`,
`capture-text-title`, `capture-text-body`, `capture-text-mime`,
`capture-text-icon`) but no CSS rules. Falling through to the parent
`.capture-row-main` flex-row (`display: flex; align-items: center`)
meant the title, textarea, and mime-select stacked as intrinsic-width
boxes centered on the tall body, producing a layout where the body
floated to the top-right, the title box appeared BELOW it, and the mime
selector rendered as an unstyled OS dropdown.
Fix:
- `.capture-text-row .capture-row-main` uses `align-items: flex-start`
so the leading icon and trailing × pin to the top of the block.
- `.capture-text-inputs` is now a full-width column-flex container with
proper gaps.
- `.capture-text-title` reuses the 44px input height and typography of
`.capture-input`; `.capture-text-body` gets a 140px min-height,
vertical resize, and matching border/focus treatment.
- `.capture-text-mime` is styled as a small chip with a custom caret
so it matches `.capture-quality` and stops looking like a raw
`<select>`. Sits in a right-aligned footer under the body.
- `.capture-text-icon` gets a 44px column so it aligns with the title
input; remove button gets a small top-margin for the same reason.
Rebuilt static bundle bumped as well (`index-BLxoi9rt.css`,
`index-CQcpPA_I.js`).