- Capture YouTube subtitles by default (opt-out in UI, API, CLI --no-subtitles) - Summarize YouTube videos from subtitles; fetch on demand, then local transcription, then error - Local transcription fallback: Whisper, Parakeet, Phonon-2 (English only) - Runtime-resolved, self-updating yt-dlp and Deno JS runtime (fixes YouTube 403s) - Settings > Instance > yt-dlp: status and in-app update without restart - X Article titles from article.title, with idempotent startup backfill - Thread title generation (single and bulk) with per-provider cheap models - Per-provider title model settings in Settings > Instance - Docs, mental model, AGENTS.md and transcription spec updated
78 KiB
Spec: local transcription fallback for YouTube summaries
- Status: Implemented (2026-10-05). See "Implementation deviations" at the end for where the code differs from this text.
- Date: 2026-10-05
- Depends on: the YouTube subtitle capture and subtitle-based summarization feature (subtitle artifacts,
crates/archivr-core/src/subtitles.rs, fetching subtitles on demand at summary time, and theNoSubtitlesAvailablefailure). The symbols named below come from that feature. - Audience: a model or engineer who implements this without any other context. Read
AGENTS.mdandARCHIVR-MENTAL-MODEL.mdfirst. The repo rules apply throughout:- core stays synchronous
- errors are
anyhow - logging uses
eprintln!withinfo:/warn:prefixes - external tools are configured by
ARCHIVR_*env vars, never TOML - yt-dlp processes are only built with
yt_dlp_command()(resolver-chosen binary plus--js-runtimes) - frontend API calls only go through
frontend/src/api.js - CSS is plain
- tests are in-file
#[cfg(test)]modules
Markers used in this document:
- [INFERENCE]: a claim about third-party software that was not run while writing this spec. The implementer must check it against the installed version before relying on it.
- [UNKNOWN]: information the sources did not provide.
1. Goal and non-goals
Goal
A summary is requested for a youtube/video entry. No subtitle artifact is archived, and fetching subtitles on demand from the original video adds none. Today that ends in NO_SUBTITLES_SUMMARY_MESSAGE. With this feature, Archivr can optionally transcribe the audio locally with an engine the user picks:
| Engine kind | Label | Languages |
|---|---|---|
whisper |
Whisper (whisper.cpp natively, or faster-whisper or any other Whisper runtime through a wrapper script) | multilingual |
parakeet |
NVIDIA Parakeet (parakeet-tdt-0.6b-v2/-v3 through a wrapper script around NeMo, parakeet-mlx, sherpa-onnx, …) |
English (v2) or 25 European languages (v3) |
phonon2 |
Fermion Research Phonon-2 | English only |
The transcript is stored as a normal subtitle artifact with kind: "transcribed", so the existing ranking, reduction and digest code turns it into summary input without any special cases. Later summaries of the same entry reuse it and do not transcribe again.
Non-goals
- No cloud ASR. Everything runs as a local subprocess on the server host. HTTP transcription APIs (including Phonon's own
fermion serveOpenAI-compatible endpoint) are out of scope; see §12. - No automatic transcription at capture time. It only happens when a user asks for a summary and picks an engine in that request. A capture-time option is a possible later extension (§12).
- No transcription for non-YouTube media. Other video and audio entries still get
UNSUPPORTED_SUMMARY_CONTENT_MESSAGE. Extending to them is §12. - Archivr does not bundle models or Python engine runtimes. Models are large and licensed separately. Users install engines and point env vars at them. The deployment changes (§10) only wire up
ffmpegand pass env vars through. - No re-transcription UI or engine switching for an entry that already has a transcript (§12).
- No word-level timestamps, diarization or translation. The output is a plain cue-level VTT.
- No new TOML config.
2. Background: what exists after the subtitle feature
These symbols are the shared contracts of the subtitle feature. Reuse them; do not reimplement them.
| Area | Symbol | Behaviour relevant here |
|---|---|---|
downloader/ytdlp.rs |
SubtitleKind { Manual, Auto, Unknown }, as_str, parse |
The kind is persisted in artifact metadata_json.kind. |
StagedSubtitle { path, language, kind, format, original_language } |
A staged sidecar file in store/temp/<key>/. |
|
plan_subtitle_request(metadata_json) -> Option<SubtitleRequest> |
Derives original_language from --dump-json (language field, or else an auto key ending -orig). |
|
pub(crate) language_base(code) |
Lowercases, strips -orig, keeps the first - segment ("de-orig" → "de", "en-GB" → "en"). |
|
private is_safe_language_code(code) |
^[A-Za-z0-9][A-Za-z0-9-]*$. |
|
fetch_metadata(url, cookies) -> Option<String>, fetch_metadata_with_timeout(url, cookies, timeout) |
--dump-json; None when the video can't be reached or the bound expires. Summary-time calls (and download_subtitles) are bounded by ARCHIVR_SUMMARY_CLI_TIMEOUT. |
|
resolve_yt_dlp(), yt_dlp_command(&ytdlp) |
resolve_yt_dlp() picks the binary; yt_dlp_command() is the only way to build a yt-dlp Command (adds the resolved --js-runtimes args). |
|
downloader/store.rs |
archive_staged_file(file, store_path) -> Result<PathBuf> |
SHA3 content-addressed move into raw/. |
subtitles.rs |
SUBTITLE_ARTIFACT_ROLE = "subtitle", SUBTITLE_ORIGIN_CAPTURE, SUBTITLE_ORIGIN_SUMMARY_FETCH |
Role and origin strings. |
SubtitleFormat { Vtt, Srt } with detect, mime, extension |
||
ArchivedSubtitle { raw_relpath, language, kind, format, original_language } |
||
archive_staged_subtitles(store_path, staged) -> Vec<ArchivedSubtitle> |
Per-file errors are logged and skipped. | |
register_subtitle_artifacts(conn, store_path, entry_id, subs, origin) -> Result<usize> |
IMMEDIATE transaction; skips an existing (entry, "subtitle", blob) and logs/skips a file it can't stat. |
|
fetch_subtitles_for_entry(paths, entry_uid, cookie_rules) -> Result<usize> |
Fetches on demand. Returns early with Ok(0) for non-YouTube or non-http(s) entries, and with Ok(n) when a usable (non-empty) subtitle track already exists. |
|
subtitle_to_transcript, parse_subtitle_metadata, subtitle_track_rank |
Reducer, metadata parse, ranking. | |
database.rs |
entry_source_info, list_entry_artifacts_by_role, entry_has_artifact_blob, update_entry_summary_input_sha256 |
|
summarizer.rs |
NO_SUBTITLES_SUMMARY_MESSAGE, SUBTITLE_FETCH_PENDING_INPUT_SHA256, is_no_subtitles_error, build_summary_input_with_subtitle_fetch(paths, entry_uid, options, cookie_rules) |
Background fetch-then-build entry point. |
private run_cli(executable, args, prompt, timeout_secs) |
Thread + channel watchdog: the child is killed when recv_timeout expires. |
|
required_env, env_or, optional_env, env_timeout, resolve_cli (private) |
Env-resolution helpers. | |
routes.rs |
request_entry_summary_handler with PreflightOutcome::{Cached, Pending, FetchSubtitles}; summary_failure_error_text; record_background_summary_failure |
Server flow. |
ContextRail.jsx |
SUMMARY_PROVIDERS, SUMMARY_PROVIDER_KEY sessionStorage, the .rail-summary-controls block, the "Generating…" spinner, 1500 ms polling, the failed-attempt <p className="form-msg form-msg--err rail-summary-error"> |
UI. |
The ranking table from the subtitle feature (D4), which this spec extends in §7:
| Rank | Track |
|---|---|
| 0 | manual + en |
| 1 | manual + orig |
| 2 | other manual |
| 3 | auto/unknown + orig |
| 4 | auto/unknown + en |
| 5 | anything else |
3. Where it plugs in
Transcription runs inside the background summary worker, between fetching subtitles on demand and the final NoSubtitlesAvailable error. During that time the summary row stays pending, holding the placeholder input_sha256 = SUBTITLE_FETCH_PENDING_INPUT_SHA256. That is the same row lifecycle the subtitle fetch already uses, so the UI shows the existing "Generating…" spinner and polls every 1500 ms.
Transcription runs only when all of these hold:
- The POST body names an engine in
transcribe_engine. That engine is listed inARCHIVR_TRANSCRIBE_ENGINESand fully configured. Both are checked synchronously at preflight, before any row is created; a failure is a 400. - The entry is
youtube/video.build_summary_inputreturnedNoSubtitlesAvailableat preflight, so the handler took thePreflightOutcome::FetchSubtitlesbranch. - After
fetch_subtitles_for_entry,build_summary_inputstill returnsNoSubtitlesAvailable. Rebuilding is the ground truth: the fetch can add nothing, or add only tracks that reduce to empty text, and both cases must lead to transcription. - The engine accepts the video's language (§6.5).
flowchart TD
POST["POST /summary {provider, transcribe_engine?}"] --> CFG{"provider_from_env + transcriber::request_from_env (if engine given)"}
CFG -- "error" --> E400["400 naming the env var"]
CFG -- ok --> PRE["build_summary_input (preflight)"]
PRE -- "Ok(input)" --> SYNC["existing path: cache lookup / pending row / provider"]
PRE -- "NoSubtitlesAvailable" --> ROW["pending row, input_sha256 = pending-subtitle-fetch, 202"]
ROW --> BG["spawn_blocking: build_summary_input_with_subtitle_fetch(.., transcription)"]
BG --> FETCH["subtitles::fetch_subtitles_for_entry -> SubtitleFetchOutcome"]
FETCH --> RB1{"build_summary_input"}
RB1 -- "Ok" --> SUM["update_entry_summary_input_sha256 + summarize_prebuilt_entry"]
RB1 -- "NoSubtitlesAvailable, no engine" --> FAIL1["failed row: NO_SUBTITLES_SUMMARY_MESSAGE"]
RB1 -- "NoSubtitlesAvailable, engine requested" --> GATE{"engine supports original language?"}
GATE -- no --> FAIL2["failed row: language-unsupported copy"]
GATE -- yes --> TR["transcriber::transcribe_entry: audio -> ffmpeg 16 kHz mono WAV -> engine -> VTT -> register subtitle artifact (kind transcribed)"]
TR -- "error / timeout" --> FAIL3["failed row: sanitized transcription copy"]
TR -- ok --> RB2{"build_summary_input"}
RB2 -- "Ok" --> SUM
RB2 -- "NoSubtitlesAvailable" --> FAIL4["failed row: NO_SUBTITLES_AFTER_TRANSCRIPTION_MESSAGE"]
No provider/LLM call happens on any failure path.
3.1 Changed entry point
Change the signature of summarizer::build_summary_input_with_subtitle_fetch and migrate its single caller in routes.rs. This is a clean cutover: no second function and no shim.
pub fn build_summary_input_with_subtitle_fetch(
paths: &ArchivePaths,
entry_uid: &str,
options: SummaryBuildOptions, // Copy
cookie_rules: &[database::CookieRule],
transcription: Option<&transcriber::TranscriptionRequest>,
) -> Result<SummaryInput>
Body, in order:
- Call
subtitles::fetch_subtitles_for_entry(paths, entry_uid, cookie_rules)?. It now returns aSubtitleFetchOutcome(§6.6). Logeprintln!("info: summary {entry_uid}: subtitle fetch added {} artifact(s)", outcome.added). - Match
build_summary_input(paths, entry_uid, options):Ok(input): return it.Err(e) if is_no_subtitles_error(&e) && transcription.is_some(): continue to step 3.Err(e): returnErr(e). This is the old behaviour when no engine was requested.
- Call
transcriber::transcribe_entry(paths, entry_uid, request, outcome.original_language.as_deref(), cookie_rules)?. It returns the number of artifact rows it inserted. Language refusal and failures come back as errors carrying aTranscriptionUserMessage(§8). - Match
build_summary_input(paths, entry_uid, options)once more:Ok(input): return it.Err(e) if is_no_subtitles_error(&e): returnErr(e.context(TranscriptionUserMessage(NO_SUBTITLES_AFTER_TRANSCRIPTION_MESSAGE.into()))). In practice this only happens if a race left an unusable track, becausetranscribe_entryalready rejects empty transcripts.Err(e): returnErr(e).
The server flow after this call is unchanged: update_entry_summary_input_sha256 with the real digest, then summarize_prebuilt_entry. On error, record_background_summary_failure.
4. Engines
4.1 Comparison
| Whisper (whisper.cpp / faster-whisper) | NVIDIA Parakeet TDT 0.6B (v2 / v3) | Fermion Research Phonon-2 | |
|---|---|---|---|
| Languages | ~99 languages with multilingual models; *.en models are English-only |
v2: English only. v3: 25 European languages with automatic language detection [INFERENCE: check the Hugging Face model card] | English only (vendor docs: "All of them transcribe English from 16 kHz audio") |
| Model size | tiny ≈75 MB … large-v3 ≈3 GB; large-v3-turbo 1,618 MB (figure from Fermion's comparison table) | 2,508 MB at full precision (Fermion's table, v3); int8 ONNX builds are smaller | 164 MB (≈2.1 bits per encoder weight) |
| Accuracy (Open ASR Leaderboard, 7 English sets, avg WER, Fermion's table) | large-v3-turbo 6.58 % | v3: 4.96 % | 5.21 % |
| Hardware | whisper.cpp: CPU (AVX/NEON), Apple Metal, CUDA/Vulkan builds. faster-whisper: CPU int8 or CUDA (CTranslate2) | NeMo: PyTorch, NVIDIA GPU recommended (CPU works but slowly) [INFERENCE]. parakeet-mlx: Apple silicon. sherpa-onnx: CPU int8 | Apple silicon GPU through MLX (174× realtime on an M5 MacBook Air); x86-64/Arm CPU engine in C (AVX-512 VNNI / AVX2 / NEON; 142.8× realtime on 8 Zen 5 cores); NVIDIA GPU (CUDA graphs); Windows CPU |
| Runtime install | whisper.cpp: one native binary whisper-cli plus a ggml model file (nixpkgs whisper-cpp [INFERENCE: attribute and binary name in the pinned nixpkgs]). faster-whisper: pip install faster-whisper |
pip install nemo_toolkit[asr] (heavy), pip install parakeet-mlx, or sherpa-onnx binaries [INFERENCE] |
pip install fermion-research plus a platform runtime: on Apple silicon pip install mlx mlx-audio mlx-lm soundfile scipy zstandard; on Linux/Windows CPU pip install fermion-research torch safetensors soundfile scipy zstandard (CPU torch wheel). Containers: ghcr.io/fermionresearch/phonon-cpu:2.0.6, ghcr.io/fermionresearch/phonon-cuda:1.0.5 |
| Native output | whisper.cpp writes .vtt/.srt/.json directly. faster-whisper: Python API only (segments with start, end, text) |
NeMo/parakeet-mlx: Python APIs with segment timestamps. No archivr-compatible CLI, so a wrapper script is needed | fermion transcribe <model> <file>: transcript-only stdout; --json gives text, model id, timings, per-segment start/end, a words list and a truncated flag. No VTT output |
| Licence | whisper.cpp MIT; OpenAI Whisper weights MIT; faster-whisper and CTranslate2 MIT | Weights CC-BY-4.0 (attribution required on redistribution); NeMo Apache-2.0 [INFERENCE] | Weights CC-BY-4.0 ("the licence of NVIDIA's Parakeet TDT 0.6B v3, from which they derive"). Phonon-1 models are Apache-2.0. The licence of the fermion-research CLI package is [UNKNOWN] (not stated on the pages read) |
| Long audio | Handled internally (30 s windows) | NeMo full attention has a maximum single-pass length, roughly 24 min, so the wrapper must chunk or switch to local attention [INFERENCE] | Files longer than 35 s are decoded in 25–35 s windows cut at pauses and joined with single spaces |
| Load cost per call | Model load is a few seconds | NeMo cold start is tens of seconds [INFERENCE] | "The first command in a session loads the engine (10 to 40 s)". Archivr starts one process per job, so every job pays this |
Sources: https://www.fermionresearch.com/research/phonon-2/ and https://www.fermionresearch.com/docs/speech/, fetched 2026-10-05. Weights: https://huggingface.co/FermionResearch/Phonon-2. Everything not marked as coming from those pages is general knowledge and carries [INFERENCE] where it matters.
Archivr never redistributes weights, so attribution under CC-BY-4.0 is the duty of whoever installs or redistributes the model. If a future Docker image bundles Parakeet or Phonon-2 weights, it must carry attribution (§10).
4.2 The single output contract
Every engine adapter must end with a VTT file inside the job's temp directory, written by a subprocess or derived from one. Archivr then reads that file. This is the same rule as the codex provider: parse a file, never a free-form stdout. The one controlled exception is Phonon-2's --json stdout, which the vendor documents as clean ("Standard output carries only the transcript … Progress, warnings, and timings go to standard error"). The adapter parses it strictly and writes the VTT itself, so everything downstream still sees a file (§4.5).
4.3 Whisper (whisper)
Two backends, chosen by ARCHIVR_WHISPER_BACKEND:
whisper_cpp (default). Invoke whisper.cpp's CLI directly:
<ARCHIVR_WHISPER_CLI> -m <ARCHIVR_WHISPER_MODEL> -f <job>/audio.wav -l <hint|auto> -ovtt -oj -of <job>/transcript -np
- Outputs:
<job>/transcript.vttand<job>/transcript.json.-oftakes the path without an extension.-npsuppresses everything except results. - Language: use the language detected by Whisper from the
-ojJSON, atresult.language[INFERENCE: check this key against the pinned whisper.cpp], falling back to the hint, falling back tound. - These flags were stable in whisper.cpp for a long time. Pin them with an argument-builder test, and verify them against the installed version during the manual smoke test [INFERENCE].
- Older builds name the binary
mainorwhisper-cpp.ARCHIVR_WHISPER_CLIhandles that.
script. ARCHIVR_WHISPER_CLI is a user-supplied wrapper that follows the script contract (§4.6), for example around faster-whisper (reference script in Appendix A.1).
Language hint (both backends): h = language_base(original_language). Pass h only if it matches ^[a-z]{2}$ (Whisper's codes are mostly ISO 639-1). Otherwise pass auto (whisper.cpp) or omit --language (script). Never pass an unvalidated string.
4.4 NVIDIA Parakeet (parakeet)
Always uses the script contract (§4.6). Parakeet has no CLI that writes VTT and takes archivr's arguments. ARCHIVR_PARAKEET_MODEL (default nvidia/parakeet-tdt-0.6b-v3) is passed to the script as --model; the script decides how to load it (Hugging Face id, .nemo path, MLX repo, ONNX dir). Reference wrappers: Appendix A.2 (NeMo) and A.3 (parakeet-mlx).
Language support depends on the model, which archivr can't introspect. ARCHIVR_PARAKEET_LANGUAGES is an optional allowlist of base codes (§5); recommend en for v2. --language is passed when the hint is known, and the script may ignore it (v3 auto-detects).
4.5 Fermion Research Phonon-2 (phonon2)
English only. This is hard-coded, not configurable.
Invocation (from the vendor docs; the --json position follows their example):
<ARCHIVR_PHONON2_CLI> transcribe <ARCHIVR_PHONON2_MODEL> <job>/audio.wav --json
Defaults: CLI fermion, model phonon-2. Documented aliases are phonon-2, phonon2, phonon, speech, stt, asr. phonon-1 and phonon-1-micro also work but are not the recommended model. The CLI also exposes phonon transcribe <file>; archivr uses the fermion transcribe <model> <file> form so the model is explicit.
-
Input: the CLI reads anything libsndfile decodes (wav/flac/ogg/aiff) and refuses mp3 and m4a, printing an ffmpeg command. Archivr always passes the 16 kHz mono PCM WAV from §6.3, so this never happens.
-
Output: stdout is a single JSON object. The vendor describes these fields: the text; the model id; decode-only and wall-clock seconds; a start and end time per decoded segment; a
wordslist with per-word start/end (Phonon-2 only); and atruncatedflag. The exact key names of the segment list and its members are [UNKNOWN]. Implementation steps:- Run
fermion transcribe phonon-2 sample.wav --jsononce on a real install. - Paste the output (trimmed) as a test fixture const in
transcriber.rs. - Write
phonon_json_to_vttagainst those real keys.
Until a real sample confirms the shape, the parser should accept, in this order:
- a top-level array of segment objects, each with numeric start/end seconds and a text string, under whichever key the sample shows (expected something like
segments); - otherwise, the
wordslist grouped into cues of at most 7 s or 84 characters, split at word boundaries; - otherwise, the top-level text as a single cue from
00:00:00.000to the WAV duration. The WAV duration is(file_len - 44) / 32000seconds for 16 kHz mono s16le; §6.3 guarantees that format.
- Run
-
If
truncatedistrue:eprintln!("warn: phonon2 reported truncated segments for {entry_uid}")and still accept the output. -
Write the VTT to
<job>/transcript.vtt. Cue timestamps are formattedHH:MM:SS.mmm. Cue text gets&,<,>escaped (&,<,>); the reducer decodes them again. -
Language stored on the artifact: always
en. -
The CLI is a Python program. On first use it may download weights into
~/.cache(the container examples mount/home/phonon/.cache), so the server user needs a writableHOMEor cache directory (§10). -
The licence of the CLI package is [UNKNOWN]. The weights are CC-BY-4.0.
4.6 Script contract (Whisper script backend, Parakeet)
Archivr runs:
<executable> --input <job>/audio.wav --output <job>/transcript.vtt --model <model> [--language <xx>]
The script must:
- Exit 0 only after writing a WebVTT file to
--output. That means aWEBVTTheader, then cuesHH:MM:SS.mmm --> HH:MM:SS.mmmfollowed by text lines, with blocks separated by blank lines. - Optionally write
<output>.langnext to it, containing a single language code it detected (e.g.de). Archivr uses it only if it passesis_safe_language_code. - Treat
--languageas a hint it may ignore. - Send anything it prints to stdout or stderr. Archivr ignores stdout and keeps the last 4 KiB of stderr for its logs.
- Write nothing outside
--output's directory except model caches. - Accept being killed with SIGKILL when the timeout expires.
The audio is already 16 kHz mono PCM WAV, so scripts never resample.
5. Configuration (env vars only, never TOML)
Resolution follows provider_from_env:
- A missing required var produces an error naming that exact var (
required_env). - Optional values use
env_or/optional_env. - Timeouts use
env_timeout. - CLIs that have a conventional install use
resolve_cli: env override → well-known absolute paths →$HOME/.local/bin/<bare>→ bare name onPATH.
Move these four private helpers from summarizer.rs into a new crates/archivr-core/src/env_config.rs as pub(crate) and update summarizer.rs to import them. That gives one convention with two users, not a copy.
| Variable | Default | Required when | Meaning |
|---|---|---|---|
ARCHIVR_TRANSCRIBE_ENGINES |
(unset: feature off) | always, to enable the feature | Comma-separated list of enabled engine kinds: whisper, parakeet, phonon2. Entries are trimmed and lowercased, empty entries are dropped, and duplicates are removed keeping the first. Unknown names get one eprintln!("warn: …") per call and are otherwise ignored. |
ARCHIVR_WHISPER_CLI |
resolved: /opt/homebrew/bin/whisper-cli, /usr/local/bin/whisper-cli, $HOME/.local/bin/whisper-cli, whisper-cli |
whisper enabled |
whisper.cpp binary, or the wrapper script when the backend is script. With the script backend this var is required (required_env): auto-discovery would find whisper-cli, which does not follow the script contract. |
ARCHIVR_WHISPER_MODEL |
— | whisper enabled |
whisper.cpp: path to a ggml model file. Script: passed through as --model (e.g. large-v3-turbo). |
ARCHIVR_WHISPER_BACKEND |
whisper_cpp |
— | whisper_cpp or script. Any other value is an error naming the var and the allowed values. |
ARCHIVR_WHISPER_LANGUAGES |
(unset: any) | — | Optional allowlist of base language codes (e.g. en for a *.en model). |
ARCHIVR_PARAKEET_CLI |
— | parakeet enabled |
Wrapper script following §4.6 (required_env). |
ARCHIVR_PARAKEET_MODEL |
nvidia/parakeet-tdt-0.6b-v3 |
— | Passed as --model. |
ARCHIVR_PARAKEET_LANGUAGES |
(unset: any) | — | Optional allowlist; recommend en for v2. |
ARCHIVR_PHONON2_CLI |
resolved: /opt/homebrew/bin/fermion, /usr/local/bin/fermion, $HOME/.local/bin/fermion, fermion |
— | The fermion CLI from pip install fermion-research. |
ARCHIVR_PHONON2_MODEL |
phonon-2 |
— | Model name or alias passed to fermion transcribe. |
ARCHIVR_TRANSCRIBE_TIMEOUT |
3600 |
— | Seconds of wall-clock budget for one transcription job (audio acquisition, ffmpeg and engine together; §8.2). |
ARCHIVR_FFMPEG |
ffmpeg |
— | ffmpeg binary (env_or). Set by the Nix wrappers and the Dockerfile (§10). |
Notes:
ARCHIVR_TRANSCRIBE_ENGINESis the gate. An engine whose vars are complete but which is not listed there is not offered and is rejected at POST. This keeps a half-configured host from accidentally exposing a CPU-heavy feature.- Language allowlists hold base codes compared with
language_base. Parsing is the same as for the engine list. An allowlist that ends up empty is the same as unset. - No secrets are involved. Engine vars are paths and model names, so nothing needs a secret file. NixOS users can still use
environmentFile(§10).
6. Core design
6.1 New module crates/archivr-core/src/transcriber.rs
Add pub mod transcriber; (and pub(crate) mod env_config;, pub(crate) mod process;) to lib.rs. Everything is synchronous.
pub const TRANSCRIBE_ENGINE_KINDS: [&str; 3] = ["whisper", "parakeet", "phonon2"];
pub const DEFAULT_TRANSCRIBE_TIMEOUT_SECS: u64 = 3600;
pub const TRANSCRIBE_SAMPLE_RATE_HZ: u32 = 16_000;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum WhisperBackend { WhisperCpp, Script }
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TranscriberConfig {
pub kind: &'static str, // one of TRANSCRIBE_ENGINE_KINDS
pub executable: PathBuf,
pub model: String,
pub whisper_backend: WhisperBackend, // ignored unless kind == "whisper"
pub languages: Option<Vec<String>>, // base codes; None = any. phonon2: always Some(["en"])
pub timeout_secs: u64,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TranscriptionSettings {
pub ffmpeg: PathBuf,
}
#[derive(Debug, Clone, PartialEq, Eq, serde::Serialize)]
pub struct TranscriberInfo {
pub kind: &'static str,
pub label: &'static str, // "Whisper" | "NVIDIA Parakeet" | "Phonon-2"
pub english_only: bool, // languages == Some(["en"])
pub languages: Option<Vec<String>>,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TranscriptOutput {
pub vtt_path: PathBuf, // inside out_dir
pub language: Option<String>, // detected/assumed; validated with is_safe_language_code
}
/// `Send + Sync` so a boxed transcriber can cross into the server's `spawn_blocking` worker.
pub trait Transcriber: Send + Sync {
fn kind(&self) -> &'static str;
fn label(&self) -> &'static str;
fn model(&self) -> &str;
fn timeout_secs(&self) -> u64;
/// `None` = language unknown → always true (see §6.5).
fn supports_language(&self, original_language: Option<&str>) -> bool;
fn supported_languages(&self) -> Option<&[String]>;
fn transcribe(&self, audio_wav: &Path, lang_hint: Option<&str>, out_dir: &Path,
deadline: std::time::Instant) -> Result<TranscriptOutput>;
}
pub struct TranscriptionRequest {
pub transcriber: Box<dyn Transcriber>,
pub settings: TranscriptionSettings,
}
pub fn enabled_engine_kinds() -> Vec<&'static str>; // parses ARCHIVR_TRANSCRIBE_ENGINES
pub fn transcriber_from_env(kind: &str) -> Result<TranscriberConfig>; // unknown kind → "unknown transcription engine: x (expected one of whisper, parakeet, phonon2)"
pub fn transcriber_from_config(cfg: TranscriberConfig) -> Box<dyn Transcriber>;
pub fn transcription_settings_from_env() -> TranscriptionSettings;
/// Enabled AND configured engines, in TRANSCRIBE_ENGINE_KINDS order. A configuration error is logged once per call and the engine is left out.
pub fn available_transcribers() -> Vec<TranscriberInfo>;
/// Server entry point: kind must be enabled (else error naming ARCHIVR_TRANSCRIBE_ENGINES) and configured (else the transcriber_from_env error).
pub fn request_from_env(kind: &str) -> Result<TranscriptionRequest>;
pub fn transcribe_entry(paths: &ArchivePaths, entry_uid: &str, request: &TranscriptionRequest,
original_language: Option<&str>, cookie_rules: &[database::CookieRule]) -> Result<usize>;
Notes on the plan's outline:
- The plan sketched
transcribe(..) -> Result<PathBuf>. This spec returnsTranscriptOutputso the detected language can be stored, and adds an explicitdeadlineso one budget covers all steps. - Implement one private struct per engine:
WhisperCppTranscriber,ScriptTranscriber(for both Whisperscriptand Parakeet; it carrieskind/label), andPhonon2Transcriber.transcriber_from_configboxes the matching one. - The trait makes in-process fakes possible in tests (§11).
Pure helpers. All are private unless noted, and all are unit-tested:
fn whisper_cpp_args(model: &str, wav: &Path, lang_hint: Option<&str>, out_prefix: &Path) -> Vec<OsString>;
fn script_args(wav: &Path, out_vtt: &Path, model: &str, lang_hint: Option<&str>) -> Vec<OsString>;
fn phonon2_args(model: &str, wav: &Path) -> Vec<OsString>;
fn ffmpeg_resample_args(input: &Path, out_wav: &Path) -> Vec<OsString>;
fn whisper_language_hint(original_language: Option<&str>) -> Option<String>; // ^[a-z]{2}$ after language_base
fn phonon_json_to_vtt(json: &str, wav_duration_secs: f64) -> Result<String>;
fn format_vtt_timestamp(seconds: f64) -> String; // "HH:MM:SS.mmm", clamps negatives to 0
fn wav_duration_secs(byte_len: u64) -> f64; // (len - 44).max(0) / 32000
fn parse_language_list(raw: Option<&str>) -> Option<Vec<String>>;
fn select_audio_source(store_path: &Path, primary: &[database::RoleArtifact]) -> Option<PathBuf>;
6.2 transcribe_entry, step by step
- Lock (§6.7). Acquire the process-wide transcription slot, waiting at most
timeout_secs. After the wait, start the job clock withdeadline = Instant::now() + timeout_secs. Waiting in the queue does not use up the job budget. - Re-check (concurrency, mirrors D6). Open the DB and call
entry_source_info(conn, entry_uid); a missing entry isbail!("entry not found: {entry_uid}"). If it is notyoutube/video,bail!(defensive; the caller already gated). Iflist_entry_artifacts_by_role(conn, entry_id, SUBTITLE_ARTIFACT_ROLE)now holds any artifact whosesubtitle_to_transcriptis non-empty, another request has already produced a track. ReturnOk(0)without transcribing. - Language gate (§6.5): if
!transcriber.supports_language(original_language), return the language-unsupported error (§8.1). - Job dir:
job = store_path/temp/transcribe-<Uuid::new_v4().simple()>. Wrap it in a privateTempDirGuard(PathBuf)whoseDroprunslet _ = fs::remove_dir_all(..), so cleanup also happens on?and on panics. - Audio source (§6.3): an archived media file, or failing that a yt-dlp audio-only download into
job. - Resample with ffmpeg into
job/audio.wav(§6.3). - Transcribe:
transcriber.transcribe(&job.join("audio.wav"), hint, &job, deadline). The hint isoriginal_language(each adapter derives its own form). - Validate:
vtt_pathmust exist and be non-empty.subtitle_to_transcript(read_to_string(vtt_path)?)must be non-empty; otherwise the job ends with the no-speech error (§8.1). - Stage and archive:
- Build
StagedSubtitle { path: vtt_path, language, kind: SubtitleKind::Transcribed, format: "vtt".into(), original_language: original_language.map(str::to_string) }.languageis the first ofoutput.language(if safe),original_language(if safe), then"und". Forphonon2it is always"en". - Call
subtitles::archive_staged_subtitles(store_path, vec![staged]). An empty result means the move failed, which is an error.
- Build
- Register:
subtitles::register_transcript_artifact(conn, store_path, entry_id, &archived, transcriber.kind(), transcriber.model())(§7). Logeprintln!("info: transcribed {entry_uid} with {kind} ({model}) in {secs:.1}s")and return the inserted count. - The guard drops and removes
job. The audio WAV and any yt-dlp audio are never archived.
6.3 Audio acquisition and resampling
Source selection (select_audio_source). Go through list_entry_artifacts_by_role(conn, entry_id, "primary_media") in id order. Take the first artifact where both hold:
- the extension (lowercased, from
relpath) is in{mp4, m4a, webm, mkv, mov, mp3, opus, ogg, oga, flac, wav, aac}, or the MIME type starts withaudio/orvideo/; store_path.join(relpath).is_file().
YouTube captures always archive media this way: an mp4 for video qualities, or the extracted audio file for the audio quality. So the archived file is the normal source, and it costs no network and no YouTube request.
Fallback: yt-dlp audio-only download. Used only when there is no usable archived file (pruned store, legacy entry) and canonical_url is http(s)://, the same gate as fetch_subtitles_for_entry. That gate also keeps tests from spawning yt-dlp. Add this to downloader/ytdlp.rs:
pub fn download_audio_for_transcription(url: &str, store_path: &Path, stage_key: &str,
cookies: &HashMap<String, String>, timeout_secs: u64) -> Result<PathBuf>;
fn audio_only_args(url: &str, cookie_file: Option<&Path>, out_template: &Path) -> Vec<OsString>;
// url, -f, bestaudio/best, --no-playlist, [--cookies f], -o temp/<key>/<key>.audio.%(ext)s
- Build it with
yt_dlp_command(&resolve_yt_dlp())and spawn it throughprocess::run_with_timeout(§6.4), with the remaining budget. - Use the same UUID-named cookie-file pattern and cleanup as
download(capture::resolve_cookies_for_url(cookie_rules, url)). - No
-x: archivr's ffmpeg step converts, and-xwould make yt-dlp call ffmpeg a second time. - The stage key is the job dir's name, so the guard cleans it up.
- Return the single non-
.part/.ytdlfile matching<key>.audio.*, or bail. - The fetched audio is transient and not archived: the entry already has its own media record, and a second media artifact would distort
cached_bytesand the entry view.
If there is no source at all, fail with the no-audio copy (§8.1).
Resampling. Always run ffmpeg, even when the input is already WAV, so every engine sees exactly one format:
<ARCHIVR_FFMPEG> -nostdin -hide_banner -loglevel error -y -i <input> -map 0:a:0 -vn -sn -dn -ac 1 -ar 16000 -c:a pcm_s16le <job>/audio.wav
-map 0:a:0makes ffmpeg fail when there is no audio stream. Map that failure to the audio-extraction copy (§8.1).- The output is 16 kHz mono signed 16-bit PCM WAV, 32,000 bytes per second: about 115 MB per hour of audio in
store/temp/. Document the disk requirement in the README. - Run through
process::run_with_timeoutwith the remaining budget. On non-zero exit, log the stderr tail (eprintln!) and return the sanitized copy.
6.4 Subprocess runner with timeout: crates/archivr-core/src/process.rs
The repo deliberately has no wait_timeout dependency. The existing summarizer::run_cli enforces timeouts with a thread, a channel and recv_timeout. It has a latent flaw for long jobs: it reads stderr only after the child exits. whisper.cpp, NeMo and yt-dlp write a lot to stderr. Once the pipe buffer (~64 KiB) fills, the child blocks, and the run ends in a spurious timeout.
Fix it once and share it:
pub(crate) struct ProcessOutput { pub stdout: String, pub stderr_tail: String /* last 4 KiB, lossy UTF-8 */ }
/// Spawns `executable args…`, optionally writes `stdin`, drains stdout and stderr on their own threads,
/// and kills the child if it is still running at `timeout`. Non-zero exit → Err("{exe} exited with {status}: {truncated stderr}").
/// Timeout → Err("{exe} timed out after {secs}s").
pub(crate) fn run_with_timeout(executable: &Path, args: &[OsString], stdin: Option<&str>, timeout: Duration) -> Result<ProcessOutput>;
- Move the body of
run_clihere and add a stderr-draining thread that keeps a bounded tail. summarizer::run_clibecomes a thin wrapper (argsmapped toOsString,Some(prompt), timeout), so provider behaviour is unchanged. Keep its existing tests (run_cli_round_trips_stdin_to_stdout,run_cli_kills_a_child_that_overruns_its_timeout,run_cli_reports_a_nonzero_exit,codex_positional_fallback_honors_cli_timeout).- Timeout errors must be recognisable without string matching. Attach a
ProcessTimedOut { secs }sentinel (Display"timed out after {secs}s") with.context(..), and detect it viachain().any(downcast_ref), the patternUnsupportedSummaryContentalready uses. - Timeout for each step:
deadline.saturating_duration_since(Instant::now()). If that is zero, fail with the timeout copy before spawning anything. - Deviation (implementation): process-group kill. "Kills the child" alone leaves grandchildren alive: a
scriptbackend that runspython …ornemowithoutexec, or yt-dlp's ffmpeg, kept running after a timeout and held the output pipes open. On unix the runner now spawns the child in its own process group (CommandExt::process_group(0)) and on timeout SIGKILLs the whole group vialibc::kill(-pgid, SIGKILL)(unix-onlylibcdependency; never for pid ≤ 1;ESRCHignored) before reaping. Shelling out tokill -KILL -- -<pid>was dropped because the Debian slim runtime image ships nokillbinary, so the group kill silently did nothing there. After the direct child exits normally, the pipe readers get a 2 s grace (capped by the remaining budget); if a grandchild still holds the pipes, the group is killed and the readers get one more grace. yt-dlp's privaterun_with_timeoutindownloader/ytdlp.rsfollows the same rules. This also tightens §4.6: script backends' subprocesses are killed with them.
6.5 Language gating (Phonon-2 English-only, optional allowlists)
supports_language(original_language):
languages == None:true.original_language == None(unknown):true. Many YouTube videos have nolanguagefield in yt-dlp metadata [INFERENCE]. A user who picks an English-only engine for such a video has made an explicit choice, and refusing would make Phonon-2 unusable in that case. Logeprintln!("warn: {kind}: original language unknown for {entry_uid}; assuming it is supported").- Otherwise:
languages.contains(&language_base(original_language)).en,en-US,en-GBanden-origall pass for Phonon-2;de,de-origandpt-BRare refused.
phonon2 is constructed with languages: Some(vec!["en".into()]) regardless of env, so it can't be configured away.
Promote ytdlp::language_base and ytdlp::is_safe_language_code to pub and reuse them. Do not copy them into transcriber.rs.
6.6 Knowing the original language at summary time
The --dump-json metadata is not persisted on the entry (capture only derives the title from it). The fallback therefore takes the language from the probe that fetch_subtitles_for_entry already makes. Change its return type (clean cutover; its only caller is build_summary_input_with_subtitle_fetch):
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct SubtitleFetchOutcome {
pub added: usize,
pub original_language: Option<String>,
}
pub fn fetch_subtitles_for_entry(paths: &ArchivePaths, entry_uid: &str,
cookie_rules: &[database::CookieRule]) -> Result<SubtitleFetchOutcome>;
- Factor the original-language derivation out of
plan_subtitle_requestintopub fn original_language_from_metadata(value: &serde_json::Value) -> Option<String>inytdlp.rs: thelanguagefield (trimmed, non-empty), otherwise the first sorted safeautomatic_captionskey ending in-origwith the suffix removed.plan_subtitle_requestcalls it. - In
fetch_subtitles_for_entry, setoriginal_languagefromoriginal_language_from_metadataright afterfetch_metadatasucceeds, before theplan_subtitle_request(..) == Noneearly return. A video with no captions is exactly the case where planning returnsNone, and the language must survive it. - Right after
entry_source_info, computeexisting_original_language: the first existingsubtitleartifact (id order) whoseparse_subtitle_metadata(..).original_languageisSome. This is one cheap query. - Every early return carries it: entry not
youtube/video, non-http(s)URL, existing subtitle artifacts (the S2 step-3 re-check), unreachable video, and no plannable tracks. A value from a successful probe overrides it. addedcounts only rows inserted by this call. The re-check early return therefore reportsadded: 0, not the number of existing artifacts as the S2 algorithm'sOk(len)did. The only caller just logs the count.
6.7 Concurrency
- One transcription at a time per server process. Engines saturate the CPU or GPU, and two parallel Whisper large runs can exhaust memory. Implement this as a private
static SLOT: (Mutex<bool>, Condvar)intranscriber.rs, and acquire it withCondvar::wait_timeout_while(guard, timeout, |busy| *busy).- If the wait times out, fail with the busy copy (§8.1).
- A small RAII guard releases the slot and calls
notify_oneon drop. - The CLI never transcribes (summaries are server-only), so per-process is per-server.
- Same entry, concurrent requests. The second request waits for the slot. Its re-check (§6.2 step 2) then finds the first request's transcribed track and returns
Ok(0), so the rebuild succeeds without a second transcription. - Dedup.
register_transcript_artifactuses the same IMMEDIATE transaction and theentry_has_artifact_blobcheck asregister_subtitle_artifacts. Identical VTT bytes never create two rows. Two runs that produce different bytes cannot happen, because the re-check runs under the slot. - Server restart mid-job.
fail_stalled_entry_summariesalready fails the pending row at startup. The job dirtemp/transcribe-*is left behind, which matches how interrupted captures leavetemp/<timestamp>today (§12). - Tokio's blocking pool holds one thread per waiting or running job. Each job is a
spawn_blocking, as provider calls already are.
7. Artifact storage
-
Role:
subtitle, not a newtranscriptrole. Withsubtitle,youtube_transcript_contentalready lists, ranks, reduces and labels the track. Registration, dedup,cached_bytesrefresh and the fetch re-check all work unchanged. A separate role would need a second candidate list in the summarizer and a second "has text" check infetch_subtitles_for_entry. The trade-off is accepted: once a transcribed track exists, the re-check infetch_subtitles_for_entrytreats the entry as having subtitles and no longer contacts YouTube (§12). -
New kind. Add
SubtitleKind::Transcribedinytdlp.rs:as_str() == "transcribed", andparse("transcribed") == Transcribed. Update every exhaustivematchonSubtitleKind. yt-dlp staging never produces this kind. -
New origin.
pub const SUBTITLE_ORIGIN_TRANSCRIPTION: &str = "transcription";insubtitles.rs. -
Storage.
storage_area = "raw",blob_id = Some,logical_path = None, blob MIMEtext/vtt, extensionvtt. All of this comes for free fromarchive_staged_subtitlesand the existingBlobRecordconstruction. -
metadata_json:
{"language":"en","kind":"transcribed","format":"vtt","original_language":"en"|null, "origin":"transcription","engine":"phonon2","model":"phonon-2"}modelis the configured value. If it is a filesystem path (it contains/or\), store only the file name, so no host path is persisted in the archive. -
Registration API. Refactor the insertion loop of
register_subtitle_artifactsinto a privateinsert_subtitle_rows(conn, store_path, entry_id, rows: &[(&ArchivedSubtitle, serde_json::Value)]) -> Result<usize>that owns the transaction, the dedup check, the commit andrefresh_entry_cached_bytes. Then:register_subtitle_artifacts(.., origin)builds the five-key metadata and calls it (behaviour unchanged).- New
pub fn register_transcript_artifact(conn, store_path, entry_id, sub: &ArchivedSubtitle, engine: &str, model: &str) -> Result<usize>addsengine/modelwith originSUBTITLE_ORIGIN_TRANSCRIPTION.
-
Ranking. Transcribed tracks go below every manual track and above auto captions.
subtitle_track_rankbecomes:Rank Track 0 manual + en 1 manual + orig 2 other manual 3 transcribed (any language) 4 auto/unknown + orig 5 auto/unknown + en 6 anything else Update the existing rank test's expected numbers. A transcribed track normally only exists when nothing else did, so the position only matters if subtitles appear later.
-
Summary label. No special case:
Transcript ({language}, transcribed subtitles):, from the existing format string withkind.as_str(). -
Digest. The content changes, so
input_sha256changes and no cached subtitle-less row is ever reused.PROMPT_VERSIONis not bumped.
8. Error handling, timeouts and copy
8.1 User-visible copy
Add a sentinel to transcriber.rs, following the UnsupportedSummaryContent pattern:
#[derive(Debug)]
pub struct TranscriptionUserMessage(pub String); // Display = the String
pub fn transcription_user_message(error: &anyhow::Error) -> Option<String>; // first in chain()
Every failure in transcribe_entry is built as Err(anyhow!(<detailed diagnostic>).context(TranscriptionUserMessage(<copy>))):
- The diagnostic (paths, exit status, stderr tail) goes to
eprintln!("warn: transcription {entry_uid}: {e:#}"). - Only the copy reaches the row.
Error text is visible to authenticated users only; public readers never see diagnostics. Even so, host paths and engine stderr do not belong in the archive DB.
In routes.rs, summary_failure_error_text checks in this order:
transcriber::transcription_user_message(error)is_no_subtitles_error→NO_SUBTITLES_SUMMARY_MESSAGEis_unsupported_summary_content_error→UNSUPPORTED_SUMMARY_CONTENT_MESSAGE- otherwise
format!("{error:#}")
Copy uses one line, curly apostrophes like the existing constants, and {label} from Transcriber::label():
| Case | Copy |
|---|---|
| Language refused | This video can’t be transcribed with {label} because it only supports {supported}, and the video’s original language is “{lang}”. Choose a different transcription engine. Here {supported} is English for ["en"], otherwise these languages: en, de, …. Put this in a pub fn transcription_language_unsupported_message(label, lang, supported: &[String]) -> String. |
| No audio source | Local transcription with {label} couldn’t start: the archived media file is missing and the original video couldn’t be downloaded. |
| ffmpeg failed / no audio track | Local transcription with {label} failed: the audio couldn’t be extracted from this video (it may have no audio track). |
| Engine non-zero exit | Local transcription with {label} failed: the transcription engine exited with an error. Check the server log for details. |
| Engine wrote no or empty VTT, or unparsable Phonon JSON | Local transcription with {label} failed: the transcription engine produced no subtitle file. Check the server log for details. |
Budget exceeded (ProcessTimedOut in the chain, or a zero remaining budget) |
Local transcription with {label} timed out after {timeout_secs} seconds. Raise ARCHIVR_TRANSCRIBE_TIMEOUT or choose a faster engine. |
| Slot wait timed out | Local transcription with {label} didn’t start because another transcription was still running. Try again later. |
| Transcript empty after reduction (silence or music) | const NO_SUBTITLES_AFTER_TRANSCRIPTION_MESSAGE: &str = "This video can’t be summarized because no subtitles are available and local transcription found no speech in its audio."; (in summarizer.rs) |
NO_SUBTITLES_SUMMARY_MESSAGE stays exactly as it is for requests that did not ask for transcription. When an engine was tried, every outcome uses one of the transcription-specific messages above, so the original copy never wrongly claims that nothing else was attempted.
8.2 Timeouts
- One budget per job:
ARCHIVR_TRANSCRIBE_TIMEOUT(default 3600 s), stored asTranscriberConfig::timeout_secs. It covers the yt-dlp audio fallback, ffmpeg and the engine. Every subprocess gets the remaining time; on expiry the child is killed (process::run_with_timeout). - Waiting for the slot has its own cap of the same value, and that wait is not counted in the job budget.
- The summary provider's timeout (
ARCHIVR_SUMMARY_*_TIMEOUT) applies only afterwards, to the LLM call. The two are independent. - A full hour is a realistic need on CPU-only hosts for Whisper large (whisper.cpp large-v3-turbo with Metal measured 17× realtime in Fermion's table; CPU-only runs are much slower [INFERENCE]). Phonon-2 on CPU handles an hour of audio in tens of seconds, plus 10–40 s of engine load.
8.3 Preflight (synchronous 400s, no row created)
In request_entry_summary_handler, right after provider_from_env succeeds and before the preflight spawn_blocking:
transcribe_enginethat isSomewith trimmed non-emptyk→transcriber::request_from_env(k). OnErr, returnApiError::bad_request(&format!("{e:#}")). The message names the missing var or saystranscription engine 'k' is not enabled (add it to ARCHIVR_TRANSCRIBE_ENGINES).- An empty string is treated as absent.
- Validation happens even if the entry later turns out to have subtitles. A configuration error is surfaced consistently and is cheap; the request is then simply never used.
8.4 Cleanup guarantees
TempDirGuard removes temp/transcribe-<uuid> (WAV, engine outputs, yt-dlp audio, cookie file) on success, error and panic. The archived VTT has already been moved to raw/ before the guard drops.
9. API and UI
9.1 Server (crates/archivr-server/src/routes.rs)
SummaryRequestBodygains#[serde(default)] transcribe_engine: Option<String>.- Update the handler doc comment: the body may carry
transcribe_engine, which is used only for YouTube videos without subtitles and runs in the background. - After the preflight validation (§8.3), hold
transcription: Option<transcriber::TranscriptionRequest>. - Move it into the background closure. Only the
FetchSubtitlesbranch uses it: it callssummarizer::build_summary_input_with_subtitle_fetch(&paths, &entry_uid, summary_options, &cookie_rules, transcription.as_ref()). - The
Pending(input already built) branch drops it unused. The 202 body is unchanged. - New route
.route("/api/summary/transcription-engines", get(transcription_engines_handler)):auth_user.require_role(ROLE_USER)?; guests and public readers get the existing 401/403 behaviour.- Returns
Json(transcriber::available_transcribers()), e.g.[{"kind":"phonon2","label":"Phonon-2","english_only":true,"languages":["en"]}]. - It reads env only and spawns nothing, so it is cheap enough to call once per ContextRail mount.
- An empty array means the feature is off.
9.2 Frontend
frontend/src/api.js:export async function fetchTranscriptionEngines({ signal } = {})returnsgetJson('/api/summary/transcription-engines', { signal }). On a 401/403 the caller treats the result as[].requestEntrySummary(archiveId, entryUid, { provider, force, includeImages, transcribeEngine, signal })addstranscribe_engine: transcribeEngineto the JSON body only when it is a non-empty string. Extend the comment.
frontend/src/components/ContextRail.jsx:- State:
transcriptionEngines(default[]), loaded once on mount when!isPublicSession; errors become[].transcribeEngineis initialised fromsessionStorage['archivr:summary:transcribe-engine'](constSUMMARY_TRANSCRIBE_ENGINE_KEY), default'', in the same try/catch style asSUMMARY_PROVIDER_KEY. Once engines load, reset to''if the stored value is not among them. - Render inside
.rail-summary-controls, after the provider<select>, only whentranscriptionEngines.length > 0 && detail.summary.source_kind === 'youtube' && detail.summary.entity_kind === 'video':<select className="rail-summary-select" value={transcribeEngine} onChange={e => handleTranscribeEngineChange(e.target.value)} aria-label="Local transcription if no subtitles"> <option value="">No local transcription</option> {transcriptionEngines.map(t => ( <option key={t.kind} value={t.kind}>{t.label}{t.english_only ? ' (English only)' : ''}</option> ))} </select> <p className="rail-summary-transcribe-note">Used only if this video has no subtitles. Transcription runs on this server and can take several minutes.</p> handleTranscribeEngineChangesets state and persists to sessionStorage (try/catch for private mode).- Pass
transcribeEnginetorequestEntrySummary. Keep the polling and generate callbacks scoped to the selected entry, asAGENTS.mdrequires. - Progress: the existing
running/ "Generating…" spinner and 1500 ms polling cover the whole fetch, transcribe and summarize sequence, because the row stayspendingthroughout. No new status values. - Errors: transcription copy arrives as the failed attempt's
error_textand renders in the existingrail-summary-errorparagraph. No logic change.
- State:
frontend/src/styles.css:.rail-summary-transcribe-note, styled like.rail-summary-image-option__note(small muted text). Use existing custom properties only.- CaptureDialog: no change. Capture-time transcription is a non-goal (§1). If it is added later, the control belongs next to the "Download subtitles" toggle as a
transcribe_enginecapture extension, with core support inCaptureConfig.
10. Deployment
Nix (flake.nix):
- Add
--set ARCHIVR_FFMPEG ${pkgs.ffmpeg}/bin/ffmpegto both thearchivrandarchivr-servermakeWrappercalls, following the existing--setpattern. Today ffmpeg is only on the PATH of theytDlpwrapper, not of the server. - Add
pkgs.ffmpegto the dev shellbuildInputs. - Optionally add
pkgs.whisper-cppto the dev shell for local testing [INFERENCE: check the attribute name and that it shipswhisper-cliin the pinnednixos-unstable]. - Do not wrap any engine or model into the packages. Engines and models stay user-supplied.
NixOS module (modules/nixos/archivr-server.nix). The module has no way to pass env vars today. Add:
environment = lib.mkOption { type = lib.types.attrsOf lib.types.str; default = { }; description = "Extra environment variables (e.g. ARCHIVR_TRANSCRIBE_ENGINES, ARCHIVR_PHONON2_CLI, LLM provider settings)."; }, mapped tosystemd.services.archivr-server.environment.environmentFile = lib.mkOption { type = lib.types.nullOr lib.types.path; default = null; }, mapped toserviceConfig.EnvironmentFilewhen non-null. This is useful for the LLM API keys too.- Hardening impact:
ProtectSystem = "strict"keeps model files readable but read-only, which is fine.- Python engines that download weights on first use need a writable cache. Set
HOME=/var/lib/archivr-server, already the user's home and insideStateDirectory, plusXDG_CACHE_HOME/HF_HOMEunder it in the module's defaultenvironment. Uselib.mkDefaultso users can override. - GPU engines need
/dev/nvidia*. The module sets noPrivateDevicesorDeviceAllow, so that access works today. Document that adding such hardening would break CUDA engines.
- Document an example using
pkgs.whisper-cppand a model fetched withpkgs.fetchurl, or a path under/var/lib/archivr-server/models.
Docker (Dockerfile, docker-compose.yml):
Dockerfile: addARCHIVR_FFMPEG=/usr/bin/ffmpegto the existingENVblock (ffmpeg is already apt-installed). Ship no engines; the image is CPU-only Debian bookworm.docker-compose.yml: add commented examples forARCHIVR_TRANSCRIBE_ENGINES,ARCHIVR_PHONON2_CLI,ARCHIVR_WHISPER_CLI/ARCHIVR_WHISPER_MODEL, and a commented read-only volume./models:/models:ro.- README: a derived-image example for Phonon-2 on CPU, using the vendor's documented commands. The image already has
python3/venv:
Note the cache volume (FROM archivr:latest RUN python3 -m venv /opt/transcribe && \ /opt/transcribe/bin/pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu && \ /opt/transcribe/bin/pip install fermion-research torch safetensors soundfile scipy zstandard ENV ARCHIVR_TRANSCRIBE_ENGINES=phonon2 ARCHIVR_PHONON2_CLI=/opt/transcribe/bin/fermionphonon-cachein the vendor examples), so weights survive container recreation. GPU containers need the NVIDIA container toolkit and a CUDA base image, which is out of scope.
Docs to update in the implementation change (repo rule: behaviour docs move with the code):
docs/README.md:- a "Local transcription (optional)" subsection under Supported Inputs → YouTube subtitles: the engine table, a setup example per engine, the disk note (~115 MB of temp WAV per hour), the English-only note for Phonon-2;
- a new
#### Local transcriptionenv table next to#### LLM providers; - NixOS
environment/environmentFileoptions; - the Docker example.
ARCHIVR-MENTAL-MODEL.md: a node for the transcription step in the LLM-summary mermaid diagram, thetranscribedkind andtranscriptionorigin, and a "Where To Edit" row fortranscriber.rs.AGENTS.md: add the new env vars to the "External tools by env var" bullet,transcriber.rs/process.rs/env_config.rsto Important Files, and the script contract under the "CLI providers parse a file" convention.
11. Tests
No real models, no network and no GPU in tests. Engines are faked either in-process (the Transcriber trait) or with executable #!/bin/sh stubs written into a tempfile dir, following the fake_yt_dlp pattern in ytdlp.rs (std::fs::write then set_permissions(0o755) under #[cfg(unix)]). Tests that mutate env vars must hold a module-level lock, like ENV_LOCK in summarizer.rs. Prefer the config-based constructors (transcriber_from_config) so most tests don't touch env at all.
process.rs
run_with_timeout_drains_large_stderr_without_deadlock:sh -c 'head -c 1000000 /dev/zero | tr "\0" x >&2; echo ok'with a 30 s timeout returnsstdout == "ok\n".run_with_timeout_kills_overrunning_child_and_marks_timeout:sleep 30, 1 s timeout; the error chain containsProcessTimedOutand it returns in under 5 s.run_with_timeout_reports_nonzero_exit_with_stderr_tail.- The existing
run_cli_*tests insummarizer.rsstay green unchanged.
env_config.rs: a resolve_cli priority test, if moving it leaves the existing summarizer tests without coverage. Otherwise, move the existing tests along with the helpers.
transcriber.rs
whisper_cpp_args_with_language_hint: exact vector[-m, M, -f, W, -l, de, -ovtt, -oj, -of, P, -np].whisper_cpp_args_without_hint_uses_auto.whisper_language_hint_only_for_two_letter_bases:de-orig→de,en-GB→en,yue→None,zh-Hans→zh, None→None,x;rm→None.script_args_follow_contract:--input,--output,--model, plus--languageonly when a hint is given.phonon2_args_are_transcribe_model_wav_json.ffmpeg_resample_args_are_16k_mono_pcm: contains-map 0:a:0,-ac 1,-ar 16000,-c:a pcm_s16le,-nostdin; the output path is last.phonon2_refuses_known_non_english_language:de,de-origandpt-BR→ false.phonon2_accepts_english_variants_and_unknown:en,en-US,en-origandNone→ true.allowlist_gating_for_whisper_and_parakeet:Some(["en","de"])acceptsde-origand refusesfr;Noneaccepts all.parse_language_list_trims_lowercases_dedups;enabled_engine_kinds_ignores_unknown_and_duplicates(env-locked).transcriber_from_env_whisper_missing_model_names_variable;transcriber_from_env_whisper_script_backend_requires_cli;transcriber_from_env_rejects_bad_backend;transcriber_from_env_parakeet_missing_cli_names_variable;transcriber_from_env_phonon2_defaults_to_fermion_and_phonon_2;transcriber_from_env_rejects_unknown_kind;transcriber_from_env_timeout_default_and_override. All env-locked; clear everyARCHIVR_*TRANSCRIBE*/engine var before and after.request_from_env_rejects_configured_but_not_enabled_engine: the message containsARCHIVR_TRANSCRIBE_ENGINES.available_transcribers_lists_enabled_and_configured_only: whisper enabled without a model is left out; phonon2 is listed withenglish_only == true.format_vtt_timestamp_formats_and_clamps:3661.5→01:01:01.500,-1.0→00:00:00.000.phonon_json_to_vtt_from_segments_fixture. Use the real captured sample (§4.5); output reduced withsubtitle_to_transcriptequals the joined segment texts.phonon_json_to_vtt_falls_back_to_single_cue_from_text.phonon_json_to_vtt_rejects_non_json.select_audio_source_prefers_existing_archived_media: an mp4 artifact whose file exists → Some. A missing file → None. Anhtmlprimary → None.transcribe_entry_with_stub_engine_registers_transcribed_artifact:- Scratch archive with a youtube/video entry and an existing
raw/…mp4primary. - Stub ffmpeg writes a 44-byte header plus zeros to its last argument.
- Stub
whisper-cliparses-of <prefix>and writes<prefix>.vtt(WEBVTT\n\n00:00:00.000 --> 00:00:02.000\nhello world\n) and<prefix>.json({"result":{"language":"en"}}). - Assert: one
subtitleartifact;metadata_jsonhaskind:"transcribed",origin:"transcription",engine:"whisper", a file-name-onlymodel,language:"en"; blob MIMEtext/vtt;store/temp/transcribe-*is gone.
- Scratch archive with a youtube/video entry and an existing
transcribe_entry_timeout_cleans_temp_and_returns_timeout_copy: the stub engine runssleep 30,timeout_secs = 1;transcription_user_messagecontains "timed out"; the temp dir is gone.transcribe_entry_engine_failure_message_is_sanitized: the stub writes/secret/pathto stderr and exits 1; the user message does not contain/secret/path.transcribe_entry_empty_transcript_is_no_speech: the stub writes onlyWEBVTT\n; the result carriesNO_SUBTITLES_AFTER_TRANSCRIPTION_MESSAGE.transcribe_entry_skips_when_usable_subtitle_already_exists: returns 0 and never spawns the engine (point the stub path at a non-existent file; spawning it would error).transcribe_entry_without_audio_and_non_http_url_fails_with_no_audio_copy.transcription_slot_serializes_jobs: two threads using an in-process fake that records overlap; overlap is never observed.
downloader/ytdlp.rs
original_language_from_metadata_prefers_language_field_then_orig_key.audio_only_args_bestaudio_without_extract: contains-f bestaudio/bestand--no-playlist; never-x.subtitle_kind_transcribed_round_trips.
subtitles.rs
track_rank_places_transcribed_below_manual_above_auto(and update the existing rank test).register_transcript_artifact_writes_engine_metadata_and_dedups(run twice → one row).fetch_outcome_reports_original_language_from_existing_artifacts: an entry with a non-HTTP canonical URL and an existing artifact whose metadata hasoriginal_language:"de"givesOk(SubtitleFetchOutcome { added: 0, original_language: Some("de") })without yt-dlp.- Update
fetch_subtitles_for_entry_skips_non_youtube_and_non_http_entriesfor the new return type.
summarizer.rs (in-process fake: struct FakeTranscriber { vtt: &'static str, calls: AtomicUsize } implementing Transcriber; it writes out_dir/transcript.vtt). Reuse youtube_summary_fixture, with the canonical URL non-HTTP so the subtitle fetch returns zero without yt-dlp, and a stub ffmpeg via TranscriptionSettings:
subtitle_fetch_with_transcriber_uses_transcribed_track: content starts withTranscript (en, transcribed subtitles):; the fake was called once.subtitle_fetch_without_transcriber_keeps_no_subtitles_error(regression).transcriber_not_called_when_usable_subtitles_exist.youtube_summary_digest_changes_when_transcript_added.phonon2_non_english_original_language_fails_before_audio_work. Seed an unusable subtitle artifact withoriginal_language:"de"so the fetch outcome carries it; the fake records zero calls and the error carries the language copy.
routes.rs
transcription_engines_endpoint_requires_user: a guest gets 401.transcription_engines_endpoint_lists_enabled_engines: env-locked,ARCHIVR_TRANSCRIBE_ENGINES=phonon2,ARCHIVR_PHONON2_CLI=/usr/bin/false→ one item withenglish_only: true.summary_post_rejects_unconfigured_transcribe_engine_with_400:{"provider":"codex_cli","transcribe_engine":"whisper"}with whisper enabled but no model. Expect 400, the body namesARCHIVR_WHISPER_MODEL, and no summary row exists.summary_post_rejects_engine_not_enabled.youtube_summary_with_stub_transcription_reaches_provider:- Env-locked.
make_test_youtube_entrywith theyoutube-test:offlineURL and a primary mp4 artifact whose file exists. - Stub ffmpeg and a stub whisper-cli (as above) through env vars;
ARCHIVR_CODEX_CLI=/usr/bin/false. - POST with
transcribe_engine:"whisper"→ 202. Poll untilfailed. - Assert
error_textis neitherNO_SUBTITLES_SUMMARY_MESSAGEnor a transcription copy. That proves transcription succeeded and the provider ran. Also assert atranscribedsubtitle artifact exists and the row'sinput_sha256is no longer the placeholder.
- Env-locked.
summary_failure_error_text_prefers_transcription_copy.- Existing tests (
youtube_summary_without_subtitles_fails_row_with_clear_message,summary_preflight_returns_safe_message_for_unsupported_video_content) stay green unchanged.
Frontend: no ContextRail component test exists, so none is added; bun test must stay green. If one has been added by then, extend it: the select is hidden when the engines list is empty and for non-YouTube entries, and transcribe_engine is sent only when one is selected.
Manual smoke test (documented, not CI). On a host with an engine installed:
- Capture a YouTube video with
--no-subtitlesthat has no captions, or delete its subtitle artifacts in a scratch archive. - Enable the engine. Request a summary with that engine selected.
- Check the
subtitlerow'smetadata_json, the VTT inraw/, the summary content label, and thatstore/temp/is empty. - Repeat with Phonon-2 on a non-English video and expect the language copy.
- Verify the whisper.cpp flags and JSON
result.languagekey, the Phonon--jsonkeys, and the Parakeet wrapper against the installed versions.
12. Open questions
- GPU vs CPU defaults. Should
available_transcribersshow a hardware hint, such as an "(slow on CPU)" label? That needs probingnvidia-smior Metal, which is out of scope for now. - Long-video chunking. whisper.cpp and Phonon-2 chunk internally. Parakeet wrappers must chunk themselves (Appendix A.2 notes this). Should archivr pre-split the WAV with ffmpeg (
-f segment -segment_time 600) and stitch the cue offsets, so wrappers can stay naive? This adds complexity and is deferred. - Concurrency cap. One job per process is fixed here. Should it be configurable (
ARCHIVR_TRANSCRIBE_CONCURRENCY) for multi-GPU hosts? Should waiting jobs be visible ("queued") in the UI? That needs a status that does not exist today. - Persistent engine servers. Phonon-2's
fermion serve(OpenAI-compatible/v1/audio/transcriptions, 32 MB request cap,verbose_jsonwith segments) and whisper.cpp's server would avoid the 10–40 s load per job. They need an HTTP client path and chunking under 32 MB (~17 min of 16 kHz mono s16 WAV). That falls under the "no cloud/HTTP ASR" non-goal for now, even when the server is localhost. - Re-transcription and engine switching. There is no UI to drop a transcribed track or prefer another engine. This could become a "Re-transcribe with…" action that deletes the
transcription-origin artifact (needs an artifact-delete path that keeps blob refcounts correct). - Fetching subtitles after a transcript exists. The
fetch_subtitles_for_entryre-check sees the transcribed track and never contacts YouTube again, even if creator captions are added later. One option is to count only non-transcribedartifacts in that re-check. This is deferred; it trades extra yt-dlp calls for freshness. - Capture-time transcription and non-YouTube audio/video (the generic
primary_mediapath): natural extensions once this path is proven. - Leftover
temp/transcribe-*after a crash. A startup sweep of staletemp/children (older than 24 h) would cover captures too. This is a separate change. - Whisper language detection with no hint. whisper.cpp detects the language from the first 30 s; a wrong detection hurts mixed-language videos. Should archivr pass
-l enwhen the title is ASCII-only? No for now; it's a heuristic.
Appendix A: reference wrapper scripts (script contract §4.6)
These are reference sketches and have not been run [INFERENCE: check the APIs against the installed library versions]. Users install them anywhere and point ARCHIVR_WHISPER_CLI (with ARCHIVR_WHISPER_BACKEND=script) or ARCHIVR_PARAKEET_CLI at them. Archivr does not ship them.
Shared helpers used by all three scripts:
def ts(s):
s = max(0.0, float(s)); h = int(s // 3600); m = int(s % 3600 // 60)
return f"{h:02d}:{m:02d}:{s % 60:06.3f}"
def write_vtt(path, cues): # cues: iterable of (start, end, text)
with open(path, "w", encoding="utf-8") as f:
f.write("WEBVTT\n\n")
for start, end, text in cues:
text = " ".join(text.split())
if text:
f.write(f"{ts(start)} --> {ts(end)}\n{text}\n\n")
A.1 faster-whisper
#!/usr/bin/env python3
import argparse
from faster_whisper import WhisperModel
# + ts/write_vtt from above
p = argparse.ArgumentParser()
p.add_argument("--input", required=True); p.add_argument("--output", required=True)
p.add_argument("--model", required=True); p.add_argument("--language")
a = p.parse_args()
model = WhisperModel(a.model, device="auto", compute_type="default")
segments, info = model.transcribe(a.input, language=a.language, vad_filter=True)
write_vtt(a.output, ((s.start, s.end, s.text) for s in segments))
open(a.output + ".lang", "w").write(info.language)
A.2 Parakeet via NeMo
#!/usr/bin/env python3
import argparse
import nemo.collections.asr as nemo_asr
# + ts/write_vtt from above
p = argparse.ArgumentParser()
p.add_argument("--input", required=True); p.add_argument("--output", required=True)
p.add_argument("--model", required=True); p.add_argument("--language") # ignored; v3 auto-detects
a = p.parse_args()
model = nemo_asr.models.ASRModel.from_pretrained(model_name=a.model)
# Long audio: full attention has a maximum single-pass length (~24 min). For longer files either
# switch to local attention (model.change_attention_model("rel_pos_local_attn", [256, 256])) or
# split the WAV into chunks and offset the timestamps. [INFERENCE: verify for the chosen model]
out = model.transcribe([a.input], timestamps=True)
segs = out[0].timestamp["segment"]
write_vtt(a.output, ((s["start"], s["end"], s["segment"]) for s in segs))
A.3 Parakeet via parakeet-mlx (Apple silicon)
#!/usr/bin/env python3
import argparse
from parakeet_mlx import from_pretrained
# + ts/write_vtt from above
p = argparse.ArgumentParser()
p.add_argument("--input", required=True); p.add_argument("--output", required=True)
p.add_argument("--model", required=True); p.add_argument("--language")
a = p.parse_args()
model = from_pretrained(a.model) # e.g. mlx-community/parakeet-tdt-0.6b-v3
result = model.transcribe(a.input)
write_vtt(a.output, ((s.start, s.end, s.text) for s in result.sentences))
Appendix B: file-by-file change list (implementation order)
crates/archivr-core/src/env_config.rs(new): moverequired_env,env_or,optional_env,env_timeout,resolve_clifromsummarizer.rsaspub(crate); updatesummarizer.rs.crates/archivr-core/src/process.rs(new):run_with_timeout,ProcessOutput,ProcessTimedOut;summarizer::run_clidelegates to it.crates/archivr-core/src/downloader/ytdlp.rs:SubtitleKind::Transcribed;publanguage_baseandis_safe_language_code;original_language_from_metadata;download_audio_for_transcriptionandaudio_only_args.crates/archivr-core/src/subtitles.rs:SUBTITLE_ORIGIN_TRANSCRIPTION;SubtitleFetchOutcomeand the newfetch_subtitles_for_entryreturn type;insert_subtitle_rowsrefactor;register_transcript_artifact; ranking table update.crates/archivr-core/src/transcriber.rs(new): everything in §6;lib.rsmodule declarations.crates/archivr-core/src/summarizer.rs: thebuild_summary_input_with_subtitle_fetchsignature and body (§3.1);NO_SUBTITLES_AFTER_TRANSCRIPTION_MESSAGE.crates/archivr-server/src/routes.rs:SummaryRequestBody.transcribe_engine, preflight validation, background wiring,transcription_engines_handlerplus route,summary_failure_error_textorder.frontend/src/api.js,frontend/src/components/ContextRail.jsx,frontend/src/styles.css.flake.nix,modules/nixos/archivr-server.nix,Dockerfile,docker-compose.yml.docs/README.md,ARCHIVR-MENTAL-MODEL.md,AGENTS.md.- Run
cargo build,cargo test,cd frontend && bun test, then the manual smoke test (§11).
Implementation deviations
The implementation follows this spec except where listed. Order of the summary path, as implemented: archived subtitles (preflight) → subtitles fetched from the original video → local transcription (only if both give nothing and an engine was requested) → error.
- D1. yt-dlp audio fallback uses
ytdlp.rs's privaterun_with_timeout(Command, Option<Duration>), notprocess::run_with_timeout. yt-dlp must be built with the privateyt_dlp_command(&resolve_yt_dlp()), which returns aCommand. So the signature isdownload_audio_for_transcription(.., timeout: Duration), and a timeout there is recognised by checking the job deadline after the error rather than byProcessTimedOut. - D2. The downloaded audio file is found with the existing
collect_staged_outputs(temp_dir, "<key>.audio", None).media, which already skips.part,.ytdl,.tempandcookies.txt. - D3. Sentinel detection.
ProcessTimedOutis the root error with a message on top (anyhow::Error::new(ProcessTimedOut{secs}).context("{exe} timed out after {secs}s"));TranscriptionUserMessageis attached as context. Both detectors useerror.chain().find_map(downcast_ref).or_else(|| error.downcast_ref()), because context layers are only reachable throughanyhow::Error::downcast_ref. Unit tests pin this. - D4. Stored
modelis reduced to its file name only if it contains\, is absolute, or exists on disk, instead of "contains/", so Hugging Face ids such asnvidia/parakeet-tdt-0.6b-v3are kept. A relative path that doesn't exist from the server's working directory is stored as-is. - D5.
is_safe_language_codebecamepub(crate), notpub;language_basewas alreadypub(crate). Both are only used inside the crate. - D6. The phonon2 truncation warning is
warn: phonon2 reported truncated segments, without the entry uid (the engine adapter doesn't have it). Theinfo: transcribed {uid} …andwarn: transcription {uid}: …lines name the entry. - D7. Phonon JSON. A real sample was captured (see "Verified facts"), pasted as
PHONON2_SAMPLE_JSONin thetranscriber.rstests, andphonon_json_to_vttis tested against it. The parser keeps the tolerant order (segments →wordsgrouped into cues of ≤7 s / ≤84 chars →textas one cue) and acceptstext/wordfor word text andstart/end(plusstart_s/start_timevariants) for times. - D8. ContextRail sends
transcribe_engineonly while the selector is visible (engines non-empty and the entry is a YouTube video), so a stale session choice can't cause a 400 on other entries. - D9. The dev shell adds
pkgs.ffmpegonly;whisper-cppis left tonix shell nixpkgs#whisper-cpp. - D10. Non-zero exit message from
process::run_with_timeoutis"{exe} exited with {status}: …{last ≤400 chars of the stderr tail}". The oldrun_cliquoted the first 400 chars; the tail holds the useful error. - D11. Reader threads after exit. Once the child exits, the runner waits for the stdout/stderr reader channels for at most the remaining budget, then reports a timeout, so a grandchild that keeps a pipe open can't hang the job.
- D12. Error copy when an engine was requested. The user's step (4) "error" is refined per §8.1: when an engine was tried, the transcription-specific copies replace
NO_SUBTITLES_SUMMARY_MESSAGE. The plain no-subtitles copy is still used whenever no engine was requested or the feature is off. - D13. The "original language unknown" warning is logged by
transcribe_entry(warn: {kind}: original language unknown for {entry_uid}; assuming it is supported) rather than insidesupports_language, which stays pure and has no uid. - D14. Script engines (Whisper
script, Parakeet) get the same validated two-letter hint as whisper.cpp (whisper_language_hint), or no--languageat all. - D15. If moving the transcript into
raw/fails, the job fails with the "produced no subtitle file" copy (§8.1 has no dedicated row for it). DB errors while registering propagate without a user copy, like other DB failures. - D16.
run_cliis a thin adapter overprocess::run_with_timeout;summarizer.rsno longer importsio::Write,process::{Command, Stdio},sync::mpscorthread.
Verified facts
- Phonon-2 (
fermion-research0.2.9, MLX backend on Apple silicon, 2026-10-05).fermion transcribe phonon-2 <wav> --jsonprints one JSON object with the keystext,model("FermionResearch/Phonon-2"),profile,backend,engine,duration_seconds,decode_seconds,wall_seconds,segment_count,segments([{id, start, end, text}]),words([{text, start, end}]) andtruncated(bool). The first run downloaded and verified the weights into~/.cache/fermion/speech/…. The package metadata declares no licence, so the CLI licence is still [UNKNOWN]. - whisper.cpp flags and the
result.languageJSON key are pinned by unit tests and checked during the orchestrator's live smoke run against nixpkgswhisper-cpp1.8.3.