mirror of
https://github.com/thegeneralist01/archivr
synced 2026-10-09 21:03:17 +02:00
New "Summary" rail section between the URL/Preview controls and .meta-list.
A completed summary renders as bold tl;dr, body paragraph, tag chips, and a
provider · model footer; missing or failed shows Generate; pending/running
shows an inline spinner and polls GET every 1500 ms until terminal.
- api.js: fetchEntrySummary + requestEntrySummary. The POST helper unwraps
ApiError's { "error": ... } body so the missing-env-var message reaches
the user verbatim rather than as a bare status code.
- ContextRail.jsx: state seeds from detail.latest_summary so the section
renders immediately on selection. Polling is anchored on the summary
status rather than started inside the click handler, so a job still
running when the user navigates away and back is picked up again. A
transient poll failure is swallowed — the next tick retries, and a real
failure arrives as status === 'failed'.
- Regenerate passes force:true only when a completed summary is already
shown; otherwise the request can take the server's 200 cache-hit path.
- Provider choice persists in sessionStorage under archivr:summary:provider,
with try/catch around both accessors for private-mode browsers.
- Public sessions never see the selector or the Generate button, and the
section renders at all only when a completed summary made it through the
server's visibility gate.
- styles.css: .rail-summary-* only; spacing and the action button reuse
.rail-section and .rail-rearchive-btn. The spinner honours
prefers-reduced-motion — the text alone conveys the state.
- AGENTS.md: document the summary env vars alongside the existing
external-tool convention.
Smoke-tested end to end against a scratch archive with a seeded markdown
entry: claude_cli produced a real summary (pending → running → completed in
~11s); a local mock server exercised the openai_compatible transport and
confirmed the Bearer header, model, and system/user role split on the wire;
unconfigured providers return 400 naming the exact variable; a video entry
returns 400 "v1 unsupported"; a repeat POST returns 200 from cache without
adding a row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
102 lines
9.3 KiB
Markdown
102 lines
9.3 KiB
Markdown
# Repository Guidelines
|
|
|
|
## Project Overview
|
|
|
|
archivr is a self-hosted archival tool that captures and preserves digital content — YouTube/Twitter/Instagram/TikTok/Reddit posts, arbitrary URLs, full web pages (via SingleFile + Chromium, optionally through a Freedium mirror for paywalled articles), and local files — into self-contained, SQLite-backed archive directories with blob deduplication, hierarchical tags, collections, and role-based auth. Rust workspace + React frontend.
|
|
|
|
Read `ARCHIVR-MENTAL-MODEL.md` before making structural changes.
|
|
|
|
## Architecture & Data Flow
|
|
|
|
Three crates with a strict ownership split — **core owns truth; CLI and server are adapters**:
|
|
|
|
- `crates/archivr-core` — domain library: capture orchestration, SQLite schema/CRUD, downloaders, hashing. New archive features start here.
|
|
- `crates/archivr-server` — Axum HTTP API + auth + static frontend serving.
|
|
- `crates/archivr-cli` — clap-based CLI (`archivr` binary): `init`, `archive` subcommands.
|
|
|
|
Capture flow: locator → `determine_source()` (`crates/archivr-core/src/capture.rs`) routes by platform/shorthand (`yt:`, `x:`, `tweet:` …) → platform downloader (`downloader/ytdlp.rs`, `tweets.rs`, `singlefile.rs`, `http.rs`, `local.rs`) stages into `temp/` → SHA3-256 dedup (`hash.rs`, `downloader/store.rs`) moves blobs to `raw/A/B/HASH.EXT` → rows written to `archivr.sqlite` (runs, entries, artifacts, blobs) → served via `/api/archives/:id/...`. `CaptureConfig` carries per-request toggles (uBlock, reader mode, Freedium mirror, etc.); when `via_freedium` is set, the fetch URL is rewritten through `freedium-mirror.cfd` while the canonical DB URL stays the original locator.
|
|
|
|
YouTube playlists and channels produce a **parent container entry** with each video captured as a child entry. `downloader/ytdlp.rs` handles the playlist probe (fetching per-video quality metadata before archiving), the multi-video download loop, and sync mode (skipping already-archived videos when re-archiving a playlist or channel).
|
|
|
|
Per-archive layout (created by `archivr init`): `.archivr/` (name, store_path, `archivr.sqlite`) + sibling `store/` (`raw/`, `raw_tweets/`, `structured/`, `temp/`). Server-level auth lives in a **separate** `archivr-auth.sqlite` (users, sessions, API tokens, role bits GUEST=1/USER=2/ADMIN=4/OWNER=8).
|
|
|
|
The server mounts multiple archives from a TOML registry (`crates/archivr-server/src/registry.rs`); routes are parameterized by `:archive_id`.
|
|
|
|
## Key Directories
|
|
|
|
| Path | Purpose |
|
|
|---|---|
|
|
| `crates/archivr-core/src/` | Domain logic: `capture.rs`, `archive.rs`, `database.rs` (~2700 lines, all schema/CRUD), `downloader/` |
|
|
| `crates/archivr-server/src/` | `main.rs` (bootstrap), `routes.rs` (~3000 lines: AppState, handlers, middleware), `auth.rs`, `registry.rs` |
|
|
| `crates/archivr-cli/src/` | CLI entry point |
|
|
| `frontend/src/` | React app: `App.jsx` (root state + custom routing), `api.js` (fetch client), `components/`, `styles.css` |
|
|
| `docs/` | User docs (`README.md`), `superpowers/plans/` and `superpowers/specs/` (dated design docs — write plans there before large features) |
|
|
| `modules/nixos/` | NixOS module (`services.archivr-server`) |
|
|
| `vendor/twitter/` | Vendored Twitter scraper (active; the Python the server shells out to). Don't refactor casually. |
|
|
| `vendor/readability/` | Mozilla `Readability.js`, concatenated into the SingleFile reader-mode browser script by `downloader/singlefile.rs`. |
|
|
| `testing/` | Legacy scraping scripts + sample data. `testing/creds.txt` holds real tokens — never read, commit, or print it. |
|
|
|
|
## Development Commands
|
|
|
|
```bash
|
|
# Rust
|
|
cargo build # whole workspace
|
|
cargo test # all unit tests
|
|
cargo test -p archivr-core # single crate
|
|
cargo build --release -p archivr-server
|
|
|
|
# Frontend (Bun, from frontend/)
|
|
bun install
|
|
bun run dev # Vite dev server
|
|
bun run build # → ../crates/archivr-server/static (served by the server)
|
|
bun run storybook # Storybook on :6006
|
|
|
|
# Nix
|
|
nix develop # devshell: yt-dlp, nushell, uv, twitter-api-client
|
|
nix build .#archivr-server # also .#archivr-cli, .#archivr-all
|
|
|
|
# Docker
|
|
docker compose up -d # port 8080; config in ./config/archivr-server.toml (see docker/config.example.toml)
|
|
|
|
# Run server locally
|
|
cargo run -p archivr-server -- path/to/archivr-server.toml
|
|
```
|
|
|
|
No CI is configured; no rustfmt.toml/clippy.toml — default `cargo fmt`/`clippy` settings apply.
|
|
|
|
## Code Conventions & Common Patterns
|
|
|
|
- **Errors**: `anyhow::Result<T>` everywhere in core; no custom error enums. The server converts to `ApiError { status, message }` (`routes.rs`) with helpers `not_found`/`bad_request`/`unauthorized`/`forbidden`; any `anyhow` error maps to 500.
|
|
- **Async**: Tokio only at the server boundary (`#[tokio::main]`, async handlers, `tokio::spawn` for the 24h session-cleanup task). **Core is synchronous** — blocking `reqwest`, subprocess downloaders, sync `rusqlite`. Keep it that way; don't introduce async into archivr-core.
|
|
- **State**: Axum `AppState { registry, auth_db_path, login_attempts }` via `State` extractor; middleware stack = `setup_guard` (503 until owner exists) → `login_rate_limit` (5/15min per IP) → `security_headers`.
|
|
- **Auth extraction**: `AuthUser` implements `FromRequestParts` — session cookie (`session`) or `Authorization: Bearer` token (stored SHA3-256-hashed). Passwords are Argon2.
|
|
- **Logging**: `eprintln!` with `info:`/`warn:` prefixes. No `tracing`/`log` — don't add structured logging piecemeal.
|
|
- **External tools by env var**: `ARCHIVR_YT_DLP`, `ARCHIVR_CHROME`, `ARCHIVR_SINGLE_FILE`, `ARCHIVR_TWEET_PYTHON`, `ARCHIVR_TWEET_SCRAPER`, `ARCHIVR_STATIC_DIR`, `ARCHIVR_BIND`. Downloaders shell out to subprocesses; resolve binaries through these vars.
|
|
- **LLM summaries by env var**: `ARCHIVR_ANTHROPIC_API_KEY` / `ARCHIVR_ANTHROPIC_URL` / `ARCHIVR_ANTHROPIC_MODEL`, `ARCHIVR_OPENAI_API_KEY` / `ARCHIVR_OPENAI_URL` / `ARCHIVR_OPENAI_MODEL`, `ARCHIVR_CLAUDE_CLI` / `ARCHIVR_CLAUDE_MODEL`, `ARCHIVR_CODEX_CLI` / `ARCHIVR_CODEX_MODEL`, plus `ARCHIVR_SUMMARY_HTTP_TIMEOUT` (default 120s) and `ARCHIVR_SUMMARY_CLI_TIMEOUT` (default 300s). Same convention as above — never TOML, which also keeps API keys out of anything the archive persists. Summaries are manual-only: nothing in `capture.rs` triggers them.
|
|
- **Frontend**: JSX (no TypeScript), PascalCase components in `frontend/src/components/`, kebab-case CSS classes, plain CSS with custom properties in `styles.css` (no Tailwind/CSS-in-JS). No router — `App.jsx` parses `window.location.pathname` + `history.pushState`. State = `useState` + one `AuthContext`; `sessionStorage` for refresh-resilient dialog state (see `CaptureDialog.jsx` job polling, 500ms). All API calls through `frontend/src/api.js` with relative `/api/*` URLs — add new endpoints there, not inline `fetch`.
|
|
- **Naming (Rust)**: standard snake_case/PascalCase; visibility and roles are bitflag `u32`s, not enums.
|
|
|
|
## Important Files
|
|
|
|
- `crates/archivr-server/src/main.rs` — server bootstrap: config load, archive mounting, auth DB init, stalled-job recovery (running → failed on startup).
|
|
- `crates/archivr-server/src/routes.rs` — all HTTP handlers and the router; grep here first for API work.
|
|
- `crates/archivr-core/src/capture.rs` — `perform_capture()`, `Source` enum, shorthand parsing.
|
|
- `crates/archivr-core/src/downloader/ytdlp.rs` — yt-dlp integration; YouTube playlist/channel probe and download, sync mode logic.
|
|
- `crates/archivr-core/src/database.rs` — single source of truth for all SQLite schema and queries (both archive and auth DBs).
|
|
- `frontend/src/App.jsx` / `frontend/src/api.js` — frontend root state and API surface.
|
|
- `docker/config.example.toml` — server config schema: `bind`, `auth_db_path`, repeated `[[archives]]` (`id`, `label`, `archive_path`).
|
|
- `flake.nix`, `modules/nixos/archivr-server.nix`, `Dockerfile`, `docker-compose.yml` — deployment surfaces; config schema changes must be reflected in all of them plus `docs/README.md`.
|
|
|
|
## Runtime/Tooling Preferences
|
|
|
|
- **Rust edition 2024** (root `Cargo.toml`); shared deps live in `[workspace.dependencies]` — add new deps there and reference with `workspace = true`.
|
|
- **Bun** is the frontend package manager (`frontend/bun.lock`); use `bun`, not npm/yarn.
|
|
- Runtime binaries the app expects on PATH or via env vars: `yt-dlp`, Chromium, `single-file` (Node), Python 3 with `twitter-api-client`, `ffmpeg`. `nix develop` provides the dev subset.
|
|
- `.gitignore` is **default-deny with an allowlist** — new top-level files/dirs are invisible to git until explicitly allowed there.
|
|
- Frontend build output (`crates/archivr-server/static/`) is generated; never hand-edit it.
|
|
|
|
## Testing & QA
|
|
|
|
- **Rust**: unit tests only, in `#[cfg(test)]` modules inside source files (e.g. `capture.rs`, `database.rs`, `registry.rs`, `routes.rs`, `hash.rs`). No `tests/` integration dir. Patterns: `tempfile` for scratch archives, config round-trip assertions, regex/parser validation. Run `cargo test` or `cargo test -p <crate>`.
|
|
- **Frontend**: no test framework. Storybook (`bun run storybook`) is the component QA surface — stories are colocated `*.stories.jsx` files; add one when adding a nontrivial component.
|
|
- Manual smoke test for server changes: build frontend, `cargo run -p archivr-server -- <config.toml>`, exercise `/api/*`.
|