AGENTS.md: mention Freedium mirror in the overview and capture flow, add vendor/readability/ to Key Directories, drop the NEXT.md pointer. ARCHIVR-MENTAL-MODEL.md: fix stale references to legacy static/ paths in Where To Edit (frontend now lives in frontend/src/); add capture.rs and auth.rs rows; replace the outdated 'Current Limitations' section (which claimed no capture and no auth) with a Server Capabilities section reflecting async capture jobs, the auth model, search, and admin scope; add a Web Capture Pipeline section covering the Freedium mirror, SingleFile+Chromium, vendored Readability, cleanup, and Rust post-processing. NEXT.md: remove; the roadmap tracking is stale and unused.
6.7 KiB
Archivr Mental Model
This document explains the current project shape after the workspace refactor.
Key Documents
| Document | Role |
|---|---|
ARCHIVR-MENTAL-MODEL.md |
This file. Current architecture, data flows, and where to edit. |
docs/README.md |
User-facing docs: how to run the tool, supported inputs, environment variables. |
The Big Model
Archivr is now a Rust workspace with three crates:
flowchart LR
CLI["archivr-cli"] --> Core["archivr-core"]
Server["archivr-server"] --> Core
UI["static web UI"] --> Server
ServerConfig["server TOML registry"] --> Server
Core --> DB["archive/.archivr/archivr.sqlite"]
Core --> Store["archive store: raw, raw_tweets, structured, temp"]
The key rule:
archivr-coreowns archive behavior.archivr-cliandarchivr-serverare adapters.
Crates
| Crate | Responsibility |
|---|---|
archivr-core |
Archive/domain logic, database schema, queries, download/store helpers |
archivr-cli |
Command-line interface, argument parsing, terminal behavior |
archivr-server |
Web server, API routes, mounted archive registry, static UI |
Archive Model
Each archive is still self-contained:
some-archive/
.archivr/
archivr.sqlite
name
store_path
store/
raw/
raw_tweets/
structured/
temp/
The web server can mount many independent archives through its own TOML registry. That registry is separate from the archives themselves.
Example:
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/path/to/archive/.archivr"
How To Run It
There are two user-facing binaries:
| Binary | Purpose |
|---|---|
archivr |
CLI for initializing archives and capturing material into one archive |
archivr-server |
Web server for browsing one or more existing archives |
The CLI writes archive data:
nix run .#archivr -- init ./my-archive --name "My Archive"
nix run .#archivr -- archive file:///absolute/path/to/file.pdf
The server reads archive data:
nix run .#archivr-server -- ./archivr-server.toml
If no config path is passed, the server reads ./archivr-server.toml.
The config is a server registry, not archive data:
[[archives]]
id = "personal"
label = "Personal"
archive_path = "/absolute/path/to/my-archive/.archivr"
The packaged Nix server wrapper sets ARCHIVR_STATIC_DIR so the server can find the installed web UI assets. Source-tree runs do not need that variable because they fall back to crates/archivr-server/static.
Write Data Flow
When archiving something through the CLI:
sequenceDiagram
participant User
participant CLI
participant Core
participant Store
participant DB
User->>CLI: archivr archive path-or-url
CLI->>Core: classify source and call downloader/store helpers
Core->>Store: save raw/structured artifacts
Core->>DB: insert run, source identity, entry, artifacts
CLI->>User: terminal result
Web Capture Pipeline
Web pages (Source::WebPage) take a longer path than yt-dlp or tweets:
- Fetch URL selection. If
via_freediumis set and the locator isn't already a Freedium URL, the downloader fetches the page throughfreedium-mirror.cfdwith no forwarded cookies. The canonical DB URL stays the original locator. - Browser capture.
downloader/singlefile.rsshells out tosingle-file-clidriving headless Chromium. Extensions (uBlock, cookie-consent) and injected browser scripts (modal closer, reader mode) attach perCaptureConfig. - Reader mode (optional) concatenates Mozilla's
Readability.jsfromvendor/readability/into the SingleFile browser script and stamps absolute URLs on lazy images so the serialised DOM points to fetchable sources. - Freedium cleanup strips mirror UI (nav, footer, toaster, author header, download control) using multi-signal selectors so article-authored controls aren't hit.
- Rust post-processing. After SingleFile writes the HTML, a Rust pass fetches any images the browser couldn't inline (bounded reads, same-origin cookie forwarding) and embeds them as data URIs. Title extraction runs after embedded font blocks are stripped so large fonts don't push
<title>beyond the read window.
Read Data Flow
When opening the web UI:
sequenceDiagram
participant Browser
participant Server
participant Core
participant DB
Browser->>Server: GET /api/archives
Browser->>Server: GET /api/archives/:id/entries
Server->>Core: list_root_entries(conn)
Core->>DB: query archive SQLite
DB-->>Core: rows
Core-->>Server: summaries
Server-->>Browser: JSON
Where To Edit
| Feature kind | Edit here |
|---|---|
| DB schema, inserts, archive runs, entries, tags | crates/archivr-core/src/database.rs |
Capture orchestration, Source routing, CaptureConfig |
crates/archivr-core/src/capture.rs |
| Archive opening, listing entries, entry detail, runs | crates/archivr-core/src/archive.rs |
| Download/save behavior | crates/archivr-core/src/downloader/ |
| CLI commands, argument parsing, terminal output | crates/archivr-cli/src/main.rs |
| Server API routes | crates/archivr-server/src/routes.rs |
| Auth model (users, sessions, tokens, roles) | crates/archivr-server/src/auth.rs |
| Mounted archive config model | crates/archivr-server/src/registry.rs |
| Frontend root state + routing | frontend/src/App.jsx |
| Frontend API client | frontend/src/api.js |
| Frontend components | frontend/src/components/ |
| Frontend styling | frontend/src/styles.css |
Practical Feature Rule
If a feature affects archive truth, start in archivr-core.
If a feature is only how the terminal behaves, edit archivr-cli.
If a feature is only how the browser sees or calls things, edit archivr-server and the static UI.
If a browser feature needs new data, the usual order is:
- Add or query the data in
archivr-core. - Expose it in
archivr-server. - Render it in the static UI.
Server Capabilities
The server both reads and writes archive data. Capture jobs are asynchronous: POST /api/archives/:id/captures inserts a job row, spawns a blocking task, and returns immediately; the frontend polls until the job completes or fails. Heavy work stays synchronous inside archivr-core.
Auth model. A separate archivr-auth.sqlite (path derived from the server config directory) holds users, sessions, and API tokens. Role bits are u32 flags (GUEST, USER, ADMIN, OWNER) so a single bitmask value covers assignment, checks, and visibility. The middleware stack is setup_guard → login_rate_limit → security_headers; route families are classified READ / ADMIN / WRITE / STATIC in routes.rs.
Search is client-side filtering over entries the frontend has already fetched.
Admin view covers mounted archives, users, sessions, and API tokens.