Merge curation-v2-step3 into curation-v2 (steps 2 and 3 united)

Union merge of the Claude editor (step 2) and the embeddings/signals/telemetry
work (step 3): both config trees, per-provider LLM budgets plus the Voyage
meter, one startup provider log, runs.config_json with ranking and voyage
settings, and selected picks' why lines landing in candidate_runs.editor_why.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1rCLQeKBgnBo3oTgHuTMe
This commit is contained in:
2026-09-02 04:16:12 +00:00
co-authored by Claude Fable 5.1
13 changed files with 4947 additions and 37 deletions
+75 -12
View File
@@ -15,10 +15,10 @@ publishes the lot over its own OPDS catalog — which doubles as a
Each article chapter ends with Loved it / Good / Not for me links that feed back into tomorrow's curation. Each article chapter ends with Loved it / Good / Not for me links that feed back into tomorrow's curation.
Steady-state cost is roughly **$1/day**: $0.050.30 in DeepSeek tokens plus Steady-state cost is roughly **$1/day**: $0.050.30 in DeepSeek tokens plus
~$0.500.80 for the Claude editor, each with its own per-UTC-day ceiling ~$0.500.80 for the Claude editor and a few cents of Voyage AI embeddings, each
(`max_daily_usd` and `anthropic.max_daily_usd`). Those ceilings are runaway with its own per-UTC-day ceiling (`max_daily_usd`, `anthropic.max_daily_usd` and
guards, not accounting — set hard spend limits in both providers' dashboards as `voyage.max_daily_usd`). Those ceilings are runaway guards, not accounting — set
the real backstop. hard spend limits in the providers' dashboards as the real backstop.
- Full design: [`docs/plans/2026-08-15-the-daily-epub.md`](docs/plans/2026-08-15-the-daily-epub.md) - Full design: [`docs/plans/2026-08-15-the-daily-epub.md`](docs/plans/2026-08-15-the-daily-epub.md)
- Implementation decisions: [`docs/plans/2026-08-15-implementation-notes.md`](docs/plans/2026-08-15-implementation-notes.md) - Implementation decisions: [`docs/plans/2026-08-15-implementation-notes.md`](docs/plans/2026-08-15-implementation-notes.md)
@@ -29,7 +29,8 @@ the real backstop.
``` ```
Miniflux ingest ─▶ dedupe ─▶ extraction ─▶ persist ─▶ social enrichment Miniflux ingest ─▶ dedupe ─▶ extraction ─▶ persist ─▶ social enrichment
─▶ pre-filter ─▶ scoring (DeepSeek) ─▶ editor (Claude) ─▶ comments ─▶ editorial (Claude) ─▶ hygiene ─▶ embeddings (Voyage) + cheap signals ─▶ pre-filter ─▶ scoring (DeepSeek)
─▶ editor (Claude) ─▶ comments ─▶ editorial (Claude)
─▶ world briefing ─▶ EPUB (standard + X4) ─▶ XTC ─▶ publish ─▶ report ─▶ world briefing ─▶ EPUB (standard + X4) ─▶ XTC ─▶ publish ─▶ report
``` ```
@@ -57,6 +58,7 @@ fallback (`fallbacks = "default"`) is enabled on every editor request.
| **Miniflux** with an API key | the only content source | Settings → API Keys. The client is read-only and never mutates read state. | | **Miniflux** with an API key | the only content source | Settings → API Keys. The client is read-only and never mutates read state. |
| **DeepSeek API key** | scoring, and the fallback for every editor call | <https://platform.deepseek.com>. Optional: `--skip-llm` runs the whole pipeline without it. | | **DeepSeek API key** | scoring, and the fallback for every editor call | <https://platform.deepseek.com>. Optional: `--skip-llm` runs the whole pipeline without it. |
| **Anthropic API key** | the editor: selection, summaries, The Brief, the weekly profile rebuild | <https://console.anthropic.com>. Optional: without it every editor call runs on DeepSeek. Set a dashboard spend limit; `anthropic.max_daily_usd` is only a runaway guard. | | **Anthropic API key** | the editor: selection, summaries, The Brief, the weekly profile rebuild | <https://console.anthropic.com>. Optional: without it every editor call runs on DeepSeek. Set a dashboard spend limit; `anthropic.max_daily_usd` is only a runaway guard. |
| **Voyage AI API key** | article and interest embeddings behind the learned ranking signals | <https://www.voyageai.com>. Optional: without it (or with `--skip-embeddings`) the run uses cached vectors only and the learned signals are absent, never a penalty. |
| A 32+ byte random secret | signs the article rating links | `openssl rand -hex 32` | | A 32+ byte random secret | signs the article rating links | `openssl rand -hex 32` |
| **BookOrbit** library + watched folder | *optional* — a richer library UI on top of the same folder | Delivery does not need it: `daily-epub serve` has its own OPDS catalog over `publish.epub_dir`. If you do run it, create a dedicated "The Daily EPUB" library, enable *Watch folders*, and point `publish.epub_dir` at it. | | **BookOrbit** library + watched folder | *optional* — a richer library UI on top of the same folder | Delivery does not need it: `daily-epub serve` has its own OPDS catalog over `publish.epub_dir`. If you do run it, create a dedicated "The Daily EPUB" library, enable *Watch folders*, and point `publish.epub_dir` at it. |
| **Node.js 18+** and a clone of [`epub-to-xtc-converter`](https://github.com/bigbag/epub-to-xtc-converter) | XTC/XTCH output for the Xteink X4 | Optional (`xtc.enabled = false` turns it off). Needs `npm install` **inside `cli/`**, and a settings JSON naming a real TTF/OTF — see below. It has **no global npm bin** — it is invoked as `node <repo>/cli/index.js convert …`, which is why `xtc.command`/`xtc.args` are fully general. | | **Node.js 18+** and a clone of [`epub-to-xtc-converter`](https://github.com/bigbag/epub-to-xtc-converter) | XTC/XTCH output for the Xteink X4 | Optional (`xtc.enabled = false` turns it off). Needs `npm install` **inside `cli/`**, and a settings JSON naming a real TTF/OTF — see below. It has **no global npm bin** — it is invoked as `node <repo>/cli/index.js convert …`, which is why `xtc.command`/`xtc.args` are fully general. |
@@ -79,12 +81,16 @@ sudo install -m0755 target/release/daily-epub /usr/local/bin/
### Commands ### Commands
``` ```
daily-epub generate [--date YYYY-MM-DD] [--dry-run] [--out DIR] [--max-articles N] [--skip-llm] daily-epub generate [--date YYYY-MM-DD] [--dry-run] [--out DIR] [--max-articles N] [--skip-llm] [--skip-embeddings]
daily-epub serve # rating endpoints + OPDS catalog + downloads daily-epub serve # rating endpoints + OPDS catalog + downloads
daily-epub profile rebuild # regenerate learned profile adjustments daily-epub profile rebuild # regenerate learned profile adjustments
daily-epub ratings list --days 90 daily-epub ratings list --days 90
daily-epub ratings set --article 42 --label loved --note "excellent" daily-epub ratings set --article 42 --label loved --note "excellent"
daily-epub ratings clear --url https://example.com/article daily-epub ratings clear --url https://example.com/article
daily-epub explain --date YYYY-MM-DD (--article ID | --url URL) [--run-id N]
daily-epub explain --date YYYY-MM-DD --near-misses [N]
daily-epub features backfill [--days 30] [--rated-only] [--all] [--yes]
daily-epub features prune # stale embeddings + old candidate telemetry
daily-epub backfill-social # re-poll social scores for recent articles daily-epub backfill-social # re-poll social scores for recent articles
daily-epub db migrate # run migrations (also automatic on every start) daily-epub db migrate # run migrations (also automatic on every start)
``` ```
@@ -94,6 +100,25 @@ articles, curates and **builds both EPUBs into `--out`**, but it does not copy t
BookOrbit, does not run the retention sweep, does not write the `issues` row and BookOrbit, does not run the retention sweep, does not write the `issues` row and
does not advance the ingest watermark. It prints the lineup and the cost report. does not advance the ingest watermark. It prints the lineup and the cost report.
`--skip-embeddings` reads the embedding cache but makes zero Voyage calls.
`explain` answers "why was this (not) in the paper" from the `candidate_runs`
row the run persisted for every considered article: the stage it reached and the
reason it stopped, every raw and normalized signal with its presence and
effective weight, the top interests, the nearest rated neighbours, any cached
LLM assessments, and the editor's reason for a pick. `--url` canonicalizes the
address; an article that is not in the database at all is reported as never
ingested (a feed problem, not a ranking one). `--near-misses` lists the highest
ranked articles that were not selected.
`features backfill` embeds the rated and published articles first (the learned
set), then the standing interests, then — only with `--all` — every other
article first seen in the window. It prints an estimate and asks before spending
more than 5M tokens unless `--yes`; a warm cache makes zero calls. `features
prune` drops embeddings of articles neither rated nor published that are older
than `curation.ranking.embedding_retention_days`, and `candidate_runs` rows of
runs older than `curation.ranking.telemetry_retention_days`.
--- ---
## Configuration ## Configuration
@@ -152,6 +177,15 @@ Secrets belong in the environment file, never in the TOML.
| `anthropic.price_output_per_mtok` | `25.0` | USD per 1M output tokens. | | `anthropic.price_output_per_mtok` | `25.0` | USD per 1M output tokens. |
| `anthropic.max_daily_usd` | `3.0` | Claude ceiling per UTC day; tripping it moves the remaining editor work to DeepSeek. | | `anthropic.max_daily_usd` | `3.0` | Claude ceiling per UTC day; tripping it moves the remaining editor work to DeepSeek. |
| `anthropic.max_concurrent_requests` | `4` | Reserved for the parallel editor stages. | | `anthropic.max_concurrent_requests` | `4` | Reserved for the parallel editor stages. |
| `voyage.enabled` | `true` | Embed articles and interests with Voyage AI. `false` ⇒ cached vectors only. |
| `voyage.base_url` | `https://api.voyageai.com/v1` | `POST {base_url}/embeddings`. |
| `voyage.model` | `voyage-4-lite` | Embedding model; changing it invalidates the cache. |
| `voyage.api_key` | — | **`DAILY_EPUB_VOYAGE__API_KEY`**. Absent ⇒ cached vectors only. |
| `voyage.output_dimension` | `512` | One of 256, 512, 1024, 2048. |
| `voyage.batch_size` | `32` | Texts per request. |
| `voyage.max_concurrent_requests` | `4` | Requests in flight. |
| `voyage.max_input_chars` | `60000` | Per-article cut, on a char boundary. |
| `voyage.max_daily_usd` | `0.50` | Runaway guard at $0.02/M tokens. |
| `curation.max_article_count` | `28` | Hard ceiling on issue size. `--max-articles N` lowers it to `min(28, N)` and drags the soft target down with it. Must be ≥ `target_article_count`. | | `curation.max_article_count` | `28` | Hard ceiling on issue size. `--max-articles N` lowers it to `min(28, N)` and drags the soft target down with it. Must be ≥ `target_article_count`. |
| `curation.always_include_feeds` | `[]` | Miniflux feed ids or URL substrings that can never be dropped. | | `curation.always_include_feeds` | `[]` | Miniflux feed ids or URL substrings that can never be dropped. |
| `curation.blocked_domains` | `[]` | Hosts excluded outright. | | `curation.blocked_domains` | `[]` | Hosts excluded outright. |
@@ -161,6 +195,7 @@ Secrets belong in the environment file, never in the TOML.
| `curation.feedback.good_value` | `0.35` | Weight for a Good verdict. | | `curation.feedback.good_value` | `0.35` | Weight for a Good verdict. |
| `curation.feedback.not_for_me_value` | `-1.0` | Weight for a Not for me verdict. | | `curation.feedback.not_for_me_value` | `-1.0` | Weight for a Not for me verdict. |
| `curation.feedback.verdicts_in_prompt` | `60` | Recent explicit verdicts included in the system prompt. | | `curation.feedback.verdicts_in_prompt` | `60` | Recent explicit verdicts included in the system prompt. |
| `curation.ranking.*` | see below | Every weight, quota, gate and threshold of the personalized ranker. |
| `editorial.summary_model` | `editor` | `editor` (Claude) or `bulk` (DeepSeek) for the per-article summaries. | | `editorial.summary_model` | `editor` | `editor` (Claude) or `bulk` (DeepSeek) for the per-article summaries. |
| `editorial.summary_input_tokens` | `3000` | Article text offered to the summary prompt. | | `editorial.summary_input_tokens` | `3000` | Article text offered to the summary prompt. |
| `publish.epub_dir` | `/srv/bookorbit/libraries/daily-epub` | Both EPUB editions land here by atomic copy, and this is the directory the OPDS feed lists. The editions are distinguished by a `(X4)` tag in **both** the filename and `dc:title` — libraries and OPDS clients list books by title, so the filename alone would make them look identical. Point a BookOrbit watched folder at it if you want its UI too. **Renamed from `bookorbit_dir`**; the old key is a hard config error. | | `publish.epub_dir` | `/srv/bookorbit/libraries/daily-epub` | Both EPUB editions land here by atomic copy, and this is the directory the OPDS feed lists. The editions are distinguished by a `(X4)` tag in **both** the filename and `dc:title` — libraries and OPDS clients list books by title, so the filename alone would make them look identical. Point a BookOrbit watched folder at it if you want its UI too. **Renamed from `bookorbit_dir`**; the old key is a hard config error. |
@@ -175,6 +210,27 @@ Secrets belong in the environment file, never in the TOML.
| `server.hmac_secret` | — | **`DAILY_EPUB_SERVER__HMAC_SECRET`** (or `DAILY_EPUB_SECRET`). Without it, generated links are rejected with 403. | | `server.hmac_secret` | — | **`DAILY_EPUB_SERVER__HMAC_SECRET`** (or `DAILY_EPUB_SECRET`). Without it, generated links are rejected with 403. |
| `server.basic_auth_user` / `_pass` | unset | Optional Basic auth for `/opds/*` and `/files/*`. | | `server.basic_auth_user` / `_pass` | unset | Optional Basic auth for `/opds/*` and `/files/*`. |
`[curation.ranking]` holds the ranker's tunables. The learned signals are
gated: `knn` (rated-neighbour preference) ramps from `knn_floor` (8) to
`knn_full` (25) rated articles with embeddings, `feed` (feed affinity) from
`feed_floor` (15) to `feed_full` (40) attributable ratings; below the floor the
signal is absent. Ratings decay with `rating_half_life_days` (60) over
`rating_lookback_days` (180); `neighbour_k` (5) neighbours per side and
`negative_coefficient` (0.75) shape the signal. `triage_max` (800),
`deep_keep` (120), `shortlist_keep` (60), `assessment_reuse_days` (3),
`semantic_min_words` (300), `exploration_slots` (5), `[curation.ranking.quotas]`
(`triage` 60 · `interest` 20 · `knn` 20), `[curation.ranking.weights.utility]`
and `[curation.ranking.diversity]` (`cluster_threshold` 0.85, `per_cluster_cap`
2, `utility_protected` 10) are validated now and drive the LLM triage, deep
assessment and diversification stages as they land.
`[curation.ranking.weights.preliminary]` (`interest` 0.35 · `knn` 0.25 ·
`heuristic` 0.20 · `feed` 0.10 · `social` 0.10) blends the cheap signals; weights
are renormalized over the signals present for each article, so they need not sum
to 1. `embedding_retention_days` (120) and `telemetry_retention_days` (180) are
what `features prune` enforces. Validation: weights non-negative; `deep_keep ≥
shortlist_keep ≥ target_article_count`; `*_full > *_floor`; `0 ≤
cluster_threshold ≤ 1`; `per_cluster_cap ≥ 1`; Voyage batch size and concurrency ≥ 1.
--- ---
## Deployment (systemd) ## Deployment (systemd)
@@ -199,6 +255,7 @@ sudo tee /etc/daily-epub/env >/dev/null <<EOF
DAILY_EPUB_MINIFLUX__API_KEY=… DAILY_EPUB_MINIFLUX__API_KEY=…
DAILY_EPUB_DEEPSEEK__API_KEY=… DAILY_EPUB_DEEPSEEK__API_KEY=…
DAILY_EPUB_ANTHROPIC__API_KEY=… DAILY_EPUB_ANTHROPIC__API_KEY=…
DAILY_EPUB_VOYAGE__API_KEY=…
DAILY_EPUB_SERVER__HMAC_SECRET=$(openssl rand -hex 32) DAILY_EPUB_SERVER__HMAC_SECRET=$(openssl rand -hex 32)
EOF EOF
sudo chown daily-epub:daily-epub /etc/daily-epub/env && sudo chmod 0600 /etc/daily-epub/env sudo chown daily-epub:daily-epub /etc/daily-epub/env && sudo chmod 0600 /etc/daily-epub/env
@@ -440,7 +497,8 @@ server. The stages themselves:
```text ```text
miniflux.rs ingest curate/ scoring and selection miniflux.rs ingest curate/ scoring and selection
dedupe.rs clustering prefilter, llm, score, select, editorial dedupe.rs clustering prefilter, llm, score, select, editorial,
embedding, signals, telemetry
extract.rs body text profile/ the reader's taste profile extract.rs body text profile/ the reader's taste profile
images/ article images comments.rs discussion chapters images/ article images comments.rs discussion chapters
normalize usable <img> world.rs the world briefing normalize usable <img> world.rs the world briefing
@@ -511,10 +569,15 @@ From spec §7, plus what implementation turned up:
stack, so `curate/llm.rs` speaks the OpenAI-compatible wire protocol over the stack, so `curate/llm.rs` speaks the OpenAI-compatible wire protocol over the
shared `reqwest` client instead, behind a `ChatBackend` trait. The dependency shared `reqwest` client instead, behind a `ChatBackend` trait. The dependency
was removed. was removed.
- **Two providers are wired**, DeepSeek (bulk) and Anthropic (editor), each a - **Two chat providers are wired**, DeepSeek (bulk) and Anthropic (editor),
`ChatBackend` impl with its own `UsageMeter` and price table. A third means each a `ChatBackend` impl with its own `UsageMeter` and price table. A third
another impl. means another impl. Voyage AI embeddings sit behind the analogous
- **No embedding-based personal ranker yet** (spec §3.9 future work); the schema `EmbeddingBackend` trait in `curate/embedding.rs`.
is ready for it once ~200 ratings exist. - **The learned signals are computed but do not yet gate selection.** Every
eligible article gets interest, rated-neighbour, feed-affinity, social and
heuristic signals persisted to `candidate_runs.signals_json` (read them with
`explain`), while the heuristic pre-filter still decides what the LLM sees.
The rated-neighbour and feed signals stay absent until their gates open
(8 and 15 ratings respectively).
- **One reader, one issue per day.** There is no multi-user support and no - **One reader, one issue per day.** There is no multi-user support and no
weekly/retrospective edition (spec §6). weekly/retrospective edition (spec §6).
+65 -1
View File
@@ -5,6 +5,7 @@
# DAILY_EPUB_MINIFLUX__API_KEY=... # DAILY_EPUB_MINIFLUX__API_KEY=...
# DAILY_EPUB_DEEPSEEK__API_KEY=... # DAILY_EPUB_DEEPSEEK__API_KEY=...
# DAILY_EPUB_ANTHROPIC__API_KEY=... # DAILY_EPUB_ANTHROPIC__API_KEY=...
# DAILY_EPUB_VOYAGE__API_KEY=...
# DAILY_EPUB_SERVER__HMAC_SECRET=... # DAILY_EPUB_SERVER__HMAC_SECRET=...
# DAILY_EPUB_LOOKBACK_HOURS=30 # DAILY_EPUB_LOOKBACK_HOURS=30
@@ -14,7 +15,7 @@ target_article_count = 20
prefilter_keep = 120 prefilter_keep = 120
retention_days = 21 # EPUBs, by age retention_days = 21 # EPUBs, by age
xtc_retention_count = 5 # XTC issues, by count (~80-100 MB each) xtc_retention_count = 5 # XTC issues, by count (~80-100 MB each)
max_daily_usd = 2.0 # DeepSeek ceiling per UTC day; [anthropic] has its own max_daily_usd = 2.0 # DeepSeek ceiling per UTC day; [anthropic] and [voyage] have their own
world_briefing = true world_briefing = true
# SQLite database file. Parent directories are created on demand. # SQLite database file. Parent directories are created on demand.
@@ -63,6 +64,20 @@ price_output_per_mtok = 25.0
max_daily_usd = 3.0 max_daily_usd = 3.0
max_concurrent_requests = 4 max_concurrent_requests = 4
# Voyage AI embeddings behind the interest and rated-neighbour signals. Set
# `enabled = false` (or leave the key unset) and the paper still builds: the
# learned signals are simply absent, never a penalty.
[voyage]
enabled = true
base_url = "https://api.voyageai.com/v1"
model = "voyage-4-lite"
# api_key via DAILY_EPUB_VOYAGE__API_KEY env
output_dimension = 512 # 256 | 512 | 1024 | 2048
batch_size = 32
max_concurrent_requests = 4
max_input_chars = 60000 # per article, cut on a char boundary
max_daily_usd = 0.50 # runaway guard ($0.02 / M tokens)
[curation] [curation]
max_article_count = 28 # hard ceiling; there is no minimum (§13) max_article_count = 28 # hard ceiling; there is no minimum (§13)
always_include_feeds = [] # miniflux feed ids or site urls always_include_feeds = [] # miniflux feed ids or site urls
@@ -87,6 +102,55 @@ good_value = 0.35
not_for_me_value = -1.0 not_for_me_value = -1.0
verdicts_in_prompt = 60 verdicts_in_prompt = 60
# Every weight, quota, gate and threshold of the personalized ranker. The
# learned signals (`knn`, `feed`) contribute nothing until their gates open:
# the weight ramps linearly from `*_floor` to `*_full` rated articles.
[curation.ranking]
triage_max = 800 # eligible articles the triage LLM reads
deep_keep = 120 # deep-assessment set
shortlist_keep = 60 # what the editor sees
assessment_reuse_days = 3
rating_lookback_days = 180
rating_half_life_days = 60
neighbour_k = 5
negative_coefficient = 0.75
knn_floor = 8
knn_full = 25
feed_floor = 15
feed_full = 40
semantic_min_words = 300
exploration_slots = 5
embedding_retention_days = 120 # `features prune`: unrated, unpublished vectors
telemetry_retention_days = 180 # `features prune`: candidate_runs rows
[curation.ranking.quotas]
triage = 60
interest = 20
knn = 20
# Weights need not sum to 1; they are renormalized over the present signals.
[curation.ranking.weights.preliminary]
interest = 0.35
knn = 0.25
heuristic = 0.20
feed = 0.10
social = 0.10
[curation.ranking.weights.utility]
quality = 0.40
fit = 0.20
knn = 0.15
interest = 0.10
feed = 0.05
triage = 0.05
social = 0.03
heuristic = 0.02
[curation.ranking.diversity]
cluster_threshold = 0.85
per_cluster_cap = 2
utility_protected = 10
[editorial] [editorial]
summary_model = "editor" # editor (Claude) | bulk (DeepSeek) summary_model = "editor" # editor (Claude) | bulk (DeepSeek)
summary_input_tokens = 3000 # article text offered per summary summary_input_tokens = 3000 # article text offered per summary
+315
View File
@@ -75,6 +75,7 @@ pub struct Config {
pub miniflux: MinifluxConfig, pub miniflux: MinifluxConfig,
pub deepseek: DeepseekConfig, pub deepseek: DeepseekConfig,
pub anthropic: AnthropicConfig, pub anthropic: AnthropicConfig,
pub voyage: VoyageConfig,
pub curation: CurationConfig, pub curation: CurationConfig,
pub editorial: EditorialConfig, pub editorial: EditorialConfig,
pub publish: PublishConfig, pub publish: PublishConfig,
@@ -100,6 +101,7 @@ impl Default for Config {
miniflux: MinifluxConfig::default(), miniflux: MinifluxConfig::default(),
deepseek: DeepseekConfig::default(), deepseek: DeepseekConfig::default(),
anthropic: AnthropicConfig::default(), anthropic: AnthropicConfig::default(),
voyage: VoyageConfig::default(),
curation: CurationConfig::default(), curation: CurationConfig::default(),
editorial: EditorialConfig::default(), editorial: EditorialConfig::default(),
publish: PublishConfig::default(), publish: PublishConfig::default(),
@@ -229,6 +231,38 @@ impl Default for EditorialConfig {
} }
} }
/// `[voyage]` — embedding endpoint and cache shape (§4.3).
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct VoyageConfig {
pub enabled: bool,
pub base_url: String,
pub model: String,
/// Supply via `DAILY_EPUB_VOYAGE__API_KEY`; never put it in the TOML.
pub api_key: Option<String>,
pub output_dimension: usize,
pub batch_size: usize,
pub max_concurrent_requests: usize,
pub max_input_chars: usize,
pub max_daily_usd: f64,
}
impl Default for VoyageConfig {
fn default() -> Self {
Self {
enabled: true,
base_url: "https://api.voyageai.com/v1".into(),
model: "voyage-4-lite".into(),
api_key: None,
output_dimension: 512,
batch_size: 32,
max_concurrent_requests: 4,
max_input_chars: 60_000,
max_daily_usd: 0.50,
}
}
}
/// `[curation]` — pre-filter and section palette (§3.5, §3.6). /// `[curation]` — pre-filter and section palette (§3.5, §3.6).
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)] #[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)] #[serde(deny_unknown_fields, default)]
@@ -245,6 +279,7 @@ pub struct CurationConfig {
/// The only section names the LLM may use (§3.6 stage B). /// The only section names the LLM may use (§3.6 stage B).
pub sections: Vec<String>, pub sections: Vec<String>,
pub feedback: FeedbackConfig, pub feedback: FeedbackConfig,
pub ranking: RankingConfig,
} }
impl Default for CurationConfig { impl Default for CurationConfig {
@@ -268,6 +303,154 @@ impl Default for CurationConfig {
.map(|s| s.to_string()) .map(|s| s.to_string())
.collect(), .collect(),
feedback: FeedbackConfig::default(), feedback: FeedbackConfig::default(),
ranking: RankingConfig::default(),
}
}
}
/// `[curation.ranking]` — every weight, quota, gate and threshold of the
/// personalized ranker (plan §19). Steps 45 consume most of these; step 3
/// uses the learned-signal gates, the preliminary weights and the retention
/// windows.
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct RankingConfig {
pub triage_max: usize,
pub deep_keep: usize,
pub shortlist_keep: usize,
pub assessment_reuse_days: i64,
pub rating_lookback_days: i64,
pub rating_half_life_days: f64,
pub neighbour_k: usize,
pub negative_coefficient: f64,
pub knn_floor: usize,
pub knn_full: usize,
pub feed_floor: usize,
pub feed_full: usize,
pub semantic_min_words: i64,
pub exploration_slots: usize,
pub embedding_retention_days: i64,
pub telemetry_retention_days: i64,
pub quotas: RankingQuotas,
pub weights: RankingWeights,
pub diversity: DiversityConfig,
}
impl Default for RankingConfig {
fn default() -> Self {
Self {
triage_max: 800,
deep_keep: 120,
shortlist_keep: 60,
assessment_reuse_days: 3,
rating_lookback_days: 180,
rating_half_life_days: 60.0,
neighbour_k: 5,
negative_coefficient: 0.75,
knn_floor: 8,
knn_full: 25,
feed_floor: 15,
feed_full: 40,
semantic_min_words: 300,
exploration_slots: 5,
embedding_retention_days: 120,
telemetry_retention_days: 180,
quotas: RankingQuotas::default(),
weights: RankingWeights::default(),
diversity: DiversityConfig::default(),
}
}
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct RankingQuotas {
pub triage: usize,
pub interest: usize,
pub knn: usize,
}
impl Default for RankingQuotas {
fn default() -> Self {
Self {
triage: 60,
interest: 20,
knn: 20,
}
}
}
#[derive(Debug, Clone, Default, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct RankingWeights {
pub preliminary: PreliminaryWeights,
pub utility: UtilityWeights,
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct PreliminaryWeights {
pub interest: f64,
pub knn: f64,
pub heuristic: f64,
pub feed: f64,
pub social: f64,
}
impl Default for PreliminaryWeights {
fn default() -> Self {
Self {
interest: 0.35,
knn: 0.25,
heuristic: 0.20,
feed: 0.10,
social: 0.10,
}
}
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct UtilityWeights {
pub quality: f64,
pub fit: f64,
pub knn: f64,
pub interest: f64,
pub feed: f64,
pub triage: f64,
pub social: f64,
pub heuristic: f64,
}
impl Default for UtilityWeights {
fn default() -> Self {
Self {
quality: 0.40,
fit: 0.20,
knn: 0.15,
interest: 0.10,
feed: 0.05,
triage: 0.05,
social: 0.03,
heuristic: 0.02,
}
}
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct DiversityConfig {
pub cluster_threshold: f64,
pub per_cluster_cap: usize,
pub utility_protected: usize,
}
impl Default for DiversityConfig {
fn default() -> Self {
Self {
cluster_threshold: 0.85,
per_cluster_cap: 2,
utility_protected: 10,
} }
} }
} }
@@ -479,6 +662,70 @@ impl Config {
"anthropic.effort must be one of low, medium, high, xhigh, max".into(), "anthropic.effort must be one of low, medium, high, xhigh, max".into(),
)); ));
} }
let ranking = &self.curation.ranking;
if ranking.deep_keep < ranking.shortlist_keep
|| ranking.shortlist_keep < self.target_article_count
{
return Err(ConfigError::Invalid(
"curation.ranking must satisfy deep_keep >= shortlist_keep >= target_article_count"
.into(),
));
}
if ranking.knn_full <= ranking.knn_floor || ranking.feed_full <= ranking.feed_floor {
return Err(ConfigError::Invalid(
"curation.ranking *_full must be > *_floor >= 0".into(),
));
}
if !(0.0..=1.0).contains(&ranking.diversity.cluster_threshold) {
return Err(ConfigError::Invalid(
"curation.ranking.diversity.cluster_threshold must be between 0 and 1".into(),
));
}
if ranking.diversity.per_cluster_cap == 0 {
return Err(ConfigError::Invalid(
"curation.ranking.diversity.per_cluster_cap must be >= 1".into(),
));
}
let preliminary = &ranking.weights.preliminary;
let utility = &ranking.weights.utility;
let weights = [
preliminary.interest,
preliminary.knn,
preliminary.heuristic,
preliminary.feed,
preliminary.social,
utility.quality,
utility.fit,
utility.knn,
utility.interest,
utility.feed,
utility.triage,
utility.social,
utility.heuristic,
];
if weights
.iter()
.any(|weight| !weight.is_finite() || *weight < 0.0)
{
return Err(ConfigError::Invalid(
"curation.ranking weights must be finite and non-negative".into(),
));
}
if self.voyage.batch_size == 0 || self.voyage.max_concurrent_requests == 0 {
return Err(ConfigError::Invalid(
"voyage.batch_size and voyage.max_concurrent_requests must be >= 1".into(),
));
}
if ![256, 512, 1024, 2048].contains(&self.voyage.output_dimension) {
return Err(ConfigError::Invalid(
"voyage.output_dimension must be one of 256, 512, 1024, 2048".into(),
));
}
if ranking.rating_half_life_days <= 0.0 || !ranking.rating_half_life_days.is_finite() {
return Err(ConfigError::Invalid(
"curation.ranking.rating_half_life_days must be > 0".into(),
));
}
if self.curation.sections.is_empty() { if self.curation.sections.is_empty() {
return Err(ConfigError::Invalid( return Err(ConfigError::Invalid(
"curation.sections must not be empty".into(), "curation.sections must not be empty".into(),
@@ -545,8 +792,12 @@ mod tests {
jail.set_env("DAILY_EPUB_MINIFLUX__API_KEY", "secret-token"); jail.set_env("DAILY_EPUB_MINIFLUX__API_KEY", "secret-token");
jail.set_env("DAILY_EPUB_TARGET_ARTICLE_COUNT", "12"); jail.set_env("DAILY_EPUB_TARGET_ARTICLE_COUNT", "12");
jail.set_env("DAILY_EPUB_SERVER__HMAC_SECRET", "hunter2"); jail.set_env("DAILY_EPUB_SERVER__HMAC_SECRET", "hunter2");
jail.set_env("DAILY_EPUB_VOYAGE__API_KEY", "voyage-key");
jail.set_env("DAILY_EPUB_VOYAGE__ENABLED", "false");
let c = Config::load(None).map_err(|e| figment::Error::from(e.to_string()))?; let c = Config::load(None).map_err(|e| figment::Error::from(e.to_string()))?;
assert_eq!(c.voyage.api_key.as_deref(), Some("voyage-key"));
assert!(!c.voyage.enabled);
// from file // from file
assert_eq!(c.lookback_hours, 30); assert_eq!(c.lookback_hours, 30);
assert!(!c.world_briefing); assert!(!c.world_briefing);
@@ -639,6 +890,70 @@ mod tests {
assert!(c.validate().is_err()); assert!(c.validate().is_err());
} }
#[test]
fn voyage_and_ranking_defaults_and_validation() {
let cfg = Config::default();
assert!(cfg.voyage.enabled);
assert_eq!(cfg.voyage.base_url, "https://api.voyageai.com/v1");
assert_eq!(cfg.voyage.model, "voyage-4-lite");
assert_eq!(cfg.voyage.output_dimension, 512);
assert_eq!(cfg.voyage.batch_size, 32);
assert_eq!(cfg.voyage.max_concurrent_requests, 4);
assert_eq!(cfg.voyage.max_input_chars, 60_000);
assert_eq!(cfg.voyage.max_daily_usd, 0.50);
let ranking = &cfg.curation.ranking;
assert_eq!(
(
ranking.triage_max,
ranking.deep_keep,
ranking.shortlist_keep
),
(800, 120, 60)
);
assert_eq!((ranking.knn_floor, ranking.knn_full), (8, 25));
assert_eq!((ranking.feed_floor, ranking.feed_full), (15, 40));
assert_eq!(ranking.rating_half_life_days, 60.0);
assert_eq!(ranking.negative_coefficient, 0.75);
assert_eq!(ranking.weights.preliminary.interest, 0.35);
assert_eq!(ranking.weights.utility.quality, 0.40);
assert_eq!(ranking.diversity.per_cluster_cap, 2);
assert_eq!(ranking.embedding_retention_days, 120);
assert_eq!(ranking.telemetry_retention_days, 180);
cfg.validate().unwrap();
let mut bad = Config::default();
bad.voyage.output_dimension = 300;
assert!(bad.validate().is_err(), "dimension must be a Voyage size");
let mut bad = Config::default();
bad.voyage.batch_size = 0;
assert!(bad.validate().is_err());
let mut bad = Config::default();
bad.curation.ranking.weights.preliminary.knn = -0.1;
assert!(bad.validate().is_err(), "weights are non-negative");
let mut bad = Config::default();
bad.curation.ranking.knn_full = bad.curation.ranking.knn_floor;
assert!(bad.validate().is_err(), "*_full must exceed *_floor");
let mut bad = Config::default();
bad.curation.ranking.shortlist_keep = bad.curation.ranking.deep_keep + 1;
assert!(bad.validate().is_err(), "deep_keep >= shortlist_keep");
let mut bad = Config::default();
bad.curation.ranking.shortlist_keep = bad.target_article_count - 1;
assert!(bad.validate().is_err(), "shortlist_keep >= target");
let mut bad = Config::default();
bad.curation.ranking.diversity.cluster_threshold = 1.5;
assert!(bad.validate().is_err());
let mut bad = Config::default();
bad.curation.ranking.diversity.per_cluster_cap = 0;
assert!(bad.validate().is_err());
// Unknown keys inside a known section fail loudly.
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("config.toml");
std::fs::write(&path, "[voyage]\nenabled = true\nnot_a_key = 1\n").unwrap();
let err = Config::load(Some(&path)).expect_err("unknown voyage key must be rejected");
assert!(err.to_string().contains("not_a_key"), "{err}");
}
#[test] #[test]
fn validation_rejects_nonsense() { fn validation_rejects_nonsense() {
assert!( assert!(
File diff suppressed because it is too large Load Diff
+3
View File
@@ -11,11 +11,14 @@
//! client; selection and editorial on the editor with per-call bulk fallback. //! client; selection and editorial on the editor with per-call bulk fallback.
pub mod editorial; pub mod editorial;
pub mod embedding;
pub mod llm; pub mod llm;
pub mod prefilter; pub mod prefilter;
pub mod profile; pub mod profile;
pub mod score; pub mod score;
pub mod select; pub mod select;
pub mod signals;
pub mod telemetry;
use jiff::civil::Date; use jiff::civil::Date;
+24 -1
View File
@@ -90,7 +90,7 @@ impl PrefilterContext {
let since = today let since = today
.checked_sub(jiff::Span::new().days(STALE_LOOKBACK_DAYS)) .checked_sub(jiff::Span::new().days(STALE_LOOKBACK_DAYS))
.unwrap_or(today); .unwrap_or(today);
let already_published = db.previously_published_ids().await?; let already_published = db.previously_published_ids_before(today).await?;
let recently_rejected = db.recently_low_scored_ids(STALE_LOW_SCORE, since).await?; let recently_rejected = db.recently_low_scored_ids(STALE_LOW_SCORE, since).await?;
tracing::debug!( tracing::debug!(
published = already_published.len(), published = already_published.len(),
@@ -197,6 +197,29 @@ pub fn social_points(social_score: f64) -> f64 {
MAX_SOCIAL_POINTS * (social_score / SOCIAL_SATURATION).min(1.0).sqrt() MAX_SOCIAL_POINTS * (social_score / SOCIAL_SATURATION).min(1.0).sqrt()
} }
/// Text-only heuristic used by personalized ranking (§9.3).
pub fn text_heuristic(article: &Article) -> f64 {
longform_points(article.word_count)
- excerpt_only_penalty(article)
- roundup_penalty(&article.title)
}
pub fn excerpt_only_penalty(article: &Article) -> f64 {
if article.excerpt_only {
EXCERPT_ONLY_PENALTY
} else {
0.0
}
}
pub fn roundup_penalty(title: &str) -> f64 {
if looks_like_roundup(title) {
ROUNDUP_TITLE_PENALTY
} else {
0.0
}
}
/// Score one article 0100 from word count, social proof, source signals, /// Score one article 0100 from word count, social proof, source signals,
/// and the excerpt/roundup/blocklist penalties (§3.5). /// and the excerpt/roundup/blocklist penalties (§3.5).
pub fn score_article(article: &Article, _ctx: &PrefilterContext, cfg: &Config) -> f64 { pub fn score_article(article: &Article, _ctx: &PrefilterContext, cfg: &Config) -> f64 {
+10
View File
@@ -125,6 +125,16 @@ pub fn load_profile(path: &Path) -> anyhow::Result<ProfileFile> {
} }
} }
/// Load the exact standing-interest union used in the system prompt.
pub fn load_standing_interests(
opml_path: &Path,
profile_path: &Path,
) -> anyhow::Result<Vec<String>> {
let opml = parse_interests(opml_path)?;
let profile = load_profile(profile_path)?;
Ok(union_interests(opml, profile.interests))
}
fn union_interests(opml: Vec<String>, profile: Vec<String>) -> Vec<String> { fn union_interests(opml: Vec<String>, profile: Vec<String>) -> Vec<String> {
let mut seen = BTreeSet::new(); let mut seen = BTreeSet::new();
let mut out = Vec::new(); let mut out = Vec::new();
+981
View File
@@ -0,0 +1,981 @@
//! Cheap per-article ranking signals, the mid-rank percentile normalizer and the
//! preliminary blend (plan §9, §12.2, §12.4).
//!
//! Every signal is an `Option<f64>`: `None` means *absent*, which is never a
//! numeric zero. Absent signals are left out of the percentile computation and
//! of the blend, whose remaining weights are renormalized (§12.2, §12.4).
use std::collections::{BTreeMap, HashMap, HashSet};
use jiff::Timestamp;
use serde::{Deserialize, Serialize};
use crate::config::{PreliminaryWeights, RankingConfig, VoyageConfig};
use crate::curate::embedding::{dot, load_article_embeddings};
use crate::curate::prefilter;
use crate::db::Db;
use crate::types::{Article, ArticleId, FeedId, SourceKind};
/// Below this many embedded eligible articles the z-score is too noisy, so the
/// interest signal falls back to the raw top-1 cosine (§9.1).
pub const INTEREST_ZSCORE_MIN_ARTICLES: usize = 30;
/// Standard-deviation floor for the per-interest z-score (§9.1).
const ZSCORE_STD_FLOOR: f64 = 1e-3;
/// How many interests and rated neighbours `signals_json` records (§7.5).
const RECORDED_TOP: usize = 3;
/// The signal names that go through the percentile normalizer, in the order
/// they are rendered (§12.2). LLM scores (`triage`, `quality`, `fit`) are
/// absolute and arrive in steps 45.
pub const PERCENTILE_SIGNALS: [&str; 5] = ["interest", "knn", "feed", "social", "heuristic"];
#[derive(Debug, Clone, Default, PartialEq, Serialize, Deserialize)]
pub struct TopInterest {
pub name: String,
pub z: f64,
pub cos: f64,
}
#[derive(Debug, Clone, Default, PartialEq, Serialize, Deserialize)]
pub struct Neighbour {
pub article_id: ArticleId,
pub label: String,
pub cos: f64,
pub title: String,
}
/// Every cheap signal for one article, plus what the normalizer and the blend
/// derived from them (§9, §12.2, §12.4).
#[derive(Debug, Clone, Default, PartialEq, Serialize, Deserialize)]
pub struct Signals {
pub interest: Option<f64>,
/// Raw top-1 cosine behind `interest`, recorded for `explain` (§7.5).
pub interest_top1_cos: Option<f64>,
pub knn: Option<f64>,
pub feed: Option<f64>,
pub social: Option<f64>,
pub heuristic: Option<f64>,
/// Mid-rank percentiles of the present signals (§12.2).
#[serde(default)]
pub norm: BTreeMap<String, f64>,
/// Effective preliminary weights after gating and renormalization (§12.4).
#[serde(default)]
pub weights: BTreeMap<String, f64>,
#[serde(default)]
pub top_interests: Vec<TopInterest>,
#[serde(default)]
pub neighbours: Vec<Neighbour>,
#[serde(default)]
pub notes: Vec<String>,
/// Preliminary blend on a 0100 scale; `None` when nothing is present.
pub preliminary: Option<f64>,
/// Gate ramps applied to the learned signals' weights (§9.2, §9.3).
#[serde(skip)]
pub knn_gate: f64,
#[serde(skip)]
pub feed_gate: f64,
}
impl Signals {
/// The raw value of a named signal, `None` when absent or unknown.
pub fn raw(&self, name: &str) -> Option<f64> {
match name {
"interest" => self.interest,
"interest_top1_cos" => self.interest_top1_cos,
"knn" => self.knn,
"feed" => self.feed,
"social" => self.social,
"heuristic" => self.heuristic,
_ => None,
}
}
pub fn present(&self, name: &str) -> bool {
self.raw(name).is_some()
}
/// The signals every eligible article gets without embeddings or ratings.
pub fn baseline(article: &Article) -> Self {
Self {
social: (!article.social.is_empty()).then(|| article.social_score()),
heuristic: Some(prefilter::text_heuristic(article)),
..Self::default()
}
}
}
/// What the run log and the report say about the learned signals (§9.2, §15.4).
#[derive(Debug, Clone, Copy, Default, PartialEq)]
pub struct PreferenceSummary {
pub rated_with_embeddings: usize,
pub attributable_feed_ratings: usize,
pub knn_gate: f64,
pub feed_gate: f64,
}
/// One rated article with an embedding: the unit of the preference state (§9.2).
#[derive(Debug, Clone, PartialEq)]
pub struct RatedExample {
pub article_id: ArticleId,
pub label: String,
pub title: String,
/// The vote's value (`loved` 1.0, `good` 0.35, `not_for_me` 1.0).
pub value: f64,
/// `0.5 ^ (age_days / half_life_days)` at the time of the run.
pub decay: f64,
pub embedding: Vec<f32>,
/// Distinct direct feeds that carried the rated article (§9.3).
pub feeds: Vec<FeedId>,
}
impl RatedExample {
/// `weight_i = value_i × decay_i` (§9.2).
pub fn weight(&self) -> f64 {
self.value * self.decay
}
}
#[derive(Debug, Clone, Copy, Default, PartialEq)]
struct FeedRate {
up: f64,
down: f64,
}
impl FeedRate {
/// Beta-smoothed rate `(up + 1) / (up + down + 2)` (§9.3).
fn rate(self) -> f64 {
(self.up + 1.0) / (self.up + self.down + 2.0)
}
}
/// Rated-neighbour and feed-affinity state, built once per run (§9.2, §9.3).
#[derive(Debug, Clone, Default)]
pub struct PreferenceState {
pub examples: Vec<RatedExample>,
feed_rates: HashMap<FeedId, FeedRate>,
pub attributable_feed_ratings: usize,
pub knn_gate: f64,
pub feed_gate: f64,
}
impl PreferenceState {
/// Build the state from already-loaded examples (pure; tests use this).
pub fn build(examples: Vec<RatedExample>, ranking: &RankingConfig) -> Self {
let (feed_rates, attributable_feed_ratings) = feed_rates(&examples);
Self {
knn_gate: gate(examples.len(), ranking.knn_floor, ranking.knn_full),
feed_gate: gate(
attributable_feed_ratings,
ranking.feed_floor,
ranking.feed_full,
),
examples,
feed_rates,
attributable_feed_ratings,
}
}
/// Load `db::current_ratings(rating_lookback_days)` joined to
/// `article_embeddings`; ratings without an embedding are skipped (§9.2).
pub async fn load(
db: &Db,
voyage: &VoyageConfig,
ranking: &RankingConfig,
now: Timestamp,
) -> anyhow::Result<Self> {
let ratings = db.current_ratings(ranking.rating_lookback_days).await?;
let ids = ratings
.iter()
.map(|rating| rating.article_id)
.collect::<Vec<_>>();
let embeddings = load_article_embeddings(db, voyage, &ids).await?;
let mut examples = Vec::new();
for rating in ratings {
let Some(embedding) = embeddings.get(&rating.article_id).cloned() else {
continue;
};
let feeds = db
.get_article(rating.article_id)
.await?
.as_ref()
.map(direct_feeds)
.unwrap_or_default();
let age_days = (now.as_second() - rating.event_at.as_second()).max(0) as f64 / 86_400.0;
examples.push(RatedExample {
article_id: rating.article_id,
label: rating.label,
title: rating.title,
value: rating.value,
decay: decay(age_days, ranking.rating_half_life_days),
embedding,
feeds,
});
}
Ok(Self::build(examples, ranking))
}
pub fn summary(&self) -> PreferenceSummary {
PreferenceSummary {
rated_with_embeddings: self.examples.len(),
attributable_feed_ratings: self.attributable_feed_ratings,
knn_gate: self.knn_gate,
feed_gate: self.feed_gate,
}
}
/// The once-per-run log line of §9.2.
pub fn log(&self, ranking: &RankingConfig) {
let feed_detail = if self.feed_gate > 0.0 {
format!("(n={})", self.attributable_feed_ratings)
} else {
format!(
"(n={} < {})",
self.attributable_feed_ratings, ranking.feed_floor
)
};
tracing::info!(
rated_with_embeddings = self.examples.len(),
knn_gate = self.knn_gate,
feed_gate = self.feed_gate,
"preference: {} rated articles with embeddings → knn gate {:.2}; feed gate {:.1} {}",
self.examples.len(),
self.knn_gate,
self.feed_gate,
feed_detail
);
}
/// Signed rated-neighbour preference and the three nearest rated articles
/// (§9.2). Absent when the gate is closed or there are no examples.
pub fn knn(&self, candidate: &[f32], ranking: &RankingConfig) -> (Option<f64>, Vec<Neighbour>) {
if self.knn_gate <= 0.0 || self.examples.is_empty() {
return (None, Vec::new());
}
let mut scored = self
.examples
.iter()
.filter_map(|example| {
dot(candidate, &example.embedding)
.ok()
.map(|s| (s, example))
})
.collect::<Vec<_>>();
scored.sort_by(|left, right| right.0.total_cmp(&left.0));
let side = |positive: bool| -> Option<f64> {
let chosen = scored
.iter()
.filter(|(_, example)| (example.weight() > 0.0) == positive)
.take(ranking.neighbour_k.max(1))
.collect::<Vec<_>>();
let denominator = chosen
.iter()
.map(|(_, example)| example.weight().abs())
.sum::<f64>();
(denominator > 0.0).then(|| {
chosen
.iter()
.map(|(similarity, example)| example.weight().abs() * similarity)
.sum::<f64>()
/ denominator
})
};
let positive = side(true);
let negative = side(false);
let knn = (positive.is_some() || negative.is_some()).then(|| {
positive.unwrap_or(0.0) - ranking.negative_coefficient * negative.unwrap_or(0.0)
});
let neighbours = scored
.iter()
.take(RECORDED_TOP)
.map(|(cos, example)| Neighbour {
article_id: example.article_id,
label: example.label.clone(),
cos: *cos,
title: example.title.clone(),
})
.collect();
(knn, neighbours)
}
/// Mean Beta-smoothed rate over the article's rated direct feeds (§9.3).
pub fn feed(&self, article: &Article) -> Option<f64> {
if self.feed_gate <= 0.0 {
return None;
}
let rates = direct_feeds(article)
.into_iter()
.filter_map(|feed| self.feed_rates.get(&feed))
.map(|rate| rate.rate())
.collect::<Vec<_>>();
(!rates.is_empty()).then(|| rates.iter().sum::<f64>() / rates.len() as f64)
}
/// Per-feed `(up, down)` credit, exposed for tests of §9.3.
pub fn feed_credit(&self, feed: FeedId) -> Option<(f64, f64)> {
self.feed_rates.get(&feed).map(|rate| (rate.up, rate.down))
}
}
/// `0.5 ^ (age_days / half_life_days)` (§9.2).
pub fn decay(age_days: f64, half_life_days: f64) -> f64 {
if half_life_days <= 0.0 {
return 1.0;
}
0.5f64.powf(age_days.max(0.0) / half_life_days)
}
/// `clamp((n floor) / (full floor), 0, 1)` (§9.2).
pub fn gate(n: usize, floor: usize, full: usize) -> f64 {
if n <= floor {
0.0
} else if n >= full || full <= floor {
1.0
} else {
(n - floor) as f64 / (full - floor) as f64
}
}
/// Distinct `SourceKind::Feed` feeds that carried the article; the best entry's
/// feed when there are none (§9.3).
pub fn direct_feeds(article: &Article) -> Vec<FeedId> {
let mut feeds = article
.sources
.iter()
.filter(|source| source.kind == SourceKind::Feed)
.map(|source| source.feed_id)
.collect::<HashSet<_>>()
.into_iter()
.collect::<Vec<_>>();
if feeds.is_empty() && article.feed_id != 0 {
feeds.push(article.feed_id);
}
feeds.sort_unstable();
feeds
}
fn feed_rates(examples: &[RatedExample]) -> (HashMap<FeedId, FeedRate>, usize) {
let mut rates: HashMap<FeedId, FeedRate> = HashMap::new();
let mut attributable = 0;
for example in examples {
if example.feeds.is_empty() {
continue;
}
attributable += 1;
let credit = example.weight() / example.feeds.len() as f64;
for feed in &example.feeds {
let rate = rates.entry(*feed).or_default();
rate.up += credit.max(0.0);
rate.down += (-credit).max(0.0);
}
}
(rates, attributable)
}
/// The interest match of §9.1 for one article.
#[derive(Debug, Clone, Default, PartialEq)]
pub struct InterestMatch {
pub score: f64,
pub top1_cos: f64,
pub top_interests: Vec<TopInterest>,
}
/// Z-scored standing-interest match for every embedded article (§9.1).
///
/// Below [`INTEREST_ZSCORE_MIN_ARTICLES`] embedded articles the score is the raw
/// top-1 cosine instead, and that is logged.
pub fn interest_matches(
articles: &HashMap<ArticleId, Vec<f32>>,
interests: &HashMap<String, Vec<f32>>,
) -> HashMap<ArticleId, InterestMatch> {
if articles.is_empty() || interests.is_empty() {
return HashMap::new();
}
let fallback = articles.len() < INTEREST_ZSCORE_MIN_ARTICLES;
if fallback {
tracing::info!(
embedded = articles.len(),
"fewer than {INTEREST_ZSCORE_MIN_ARTICLES} embedded articles; interest uses the raw top-1 cosine"
);
}
let mut matches: HashMap<ArticleId, Vec<TopInterest>> = HashMap::new();
for (name, interest) in interests {
let similarities = articles
.iter()
.filter_map(|(article_id, article)| {
dot(interest, article).ok().map(|cos| (*article_id, cos))
})
.collect::<Vec<_>>();
if similarities.is_empty() {
continue;
}
let n = similarities.len() as f64;
let mean = similarities.iter().map(|(_, cos)| cos).sum::<f64>() / n;
let variance = similarities
.iter()
.map(|(_, cos)| (cos - mean).powi(2))
.sum::<f64>()
/ n;
let std = variance.sqrt().max(ZSCORE_STD_FLOOR);
for (article_id, cos) in similarities {
matches.entry(article_id).or_default().push(TopInterest {
name: name.clone(),
z: (cos - mean) / std,
cos,
});
}
}
matches
.into_iter()
.map(|(article_id, mut all)| {
all.sort_by(|left, right| {
right
.z
.total_cmp(&left.z)
.then_with(|| left.name.cmp(&right.name))
});
let top1_cos = all
.iter()
.map(|item| item.cos)
.fold(f64::NEG_INFINITY, f64::max);
all.truncate(RECORDED_TOP);
let score = if fallback {
top1_cos
} else {
let top_mean = all.iter().map(|item| item.z).sum::<f64>() / all.len() as f64;
0.7 * all[0].z + 0.3 * top_mean
};
(
article_id,
InterestMatch {
score,
top1_cos,
top_interests: all,
},
)
})
.collect()
}
/// Every cheap signal for the eligible set, normalized and blended (§9, §12.2,
/// §12.4). Pure: the preference state is already loaded.
pub fn compute(
articles: &[Article],
article_embeddings: &HashMap<ArticleId, Vec<f32>>,
interest_embeddings: &HashMap<String, Vec<f32>>,
preference: &PreferenceState,
ranking: &RankingConfig,
) -> HashMap<ArticleId, Signals> {
let interests = interest_matches(article_embeddings, interest_embeddings);
let mut all = articles
.iter()
.map(|article| {
let mut signals = Signals::baseline(article);
signals.knn_gate = preference.knn_gate;
signals.feed_gate = preference.feed_gate;
if let Some(matched) = interests.get(&article.id) {
signals.interest = Some(matched.score);
signals.interest_top1_cos = Some(matched.top1_cos);
signals.top_interests = matched.top_interests.clone();
}
if let Some(embedding) = article_embeddings.get(&article.id) {
let (knn, neighbours) = preference.knn(embedding, ranking);
signals.knn = knn;
signals.neighbours = neighbours;
}
signals.feed = preference.feed(article);
if preference.knn_gate > 0.0 {
signals.notes.push(format!(
"knn gate {:.2} (n={} rated with embeddings)",
preference.knn_gate,
preference.examples.len()
));
}
signals
})
.collect::<Vec<_>>();
normalize(&mut all.iter_mut().collect::<Vec<_>>());
for signals in &mut all {
preliminary_blend(signals, &ranking.weights.preliminary);
}
articles.iter().map(|article| article.id).zip(all).collect()
}
/// [`compute`] with the preference state loaded from the database.
pub async fn compute_all(
db: &Db,
articles: &[Article],
article_embeddings: &HashMap<ArticleId, Vec<f32>>,
interest_embeddings: &HashMap<String, Vec<f32>>,
voyage: &VoyageConfig,
ranking: &RankingConfig,
now: Timestamp,
) -> anyhow::Result<(HashMap<ArticleId, Signals>, PreferenceSummary)> {
let preference = PreferenceState::load(db, voyage, ranking, now).await?;
preference.log(ranking);
Ok((
compute(
articles,
article_embeddings,
interest_embeddings,
&preference,
ranking,
),
preference.summary(),
))
}
/// Mid-rank percentiles over the present values of each signal (§12.2).
///
/// `p(x) = (count_below + (count_equal + 1) / 2) / n_present`; fewer than two
/// present values or all-equal values give 0.5. Article id never breaks ties.
pub fn normalize(signals: &mut [&mut Signals]) {
for name in PERCENTILE_SIGNALS {
let mut values = signals
.iter()
.filter_map(|signal| signal.raw(name))
.collect::<Vec<_>>();
if values.is_empty() {
continue;
}
values.sort_by(f64::total_cmp);
let n = values.len() as f64;
let constant = values.len() < 2 || values.first() == values.last();
for signal in signals.iter_mut() {
let Some(value) = signal.raw(name) else {
continue;
};
let percentile = if constant {
0.5
} else {
let below = values.partition_point(|other| *other < value);
let equal = values.partition_point(|other| *other <= value) - below;
(below as f64 + (equal as f64 + 1.0) / 2.0) / n
};
signal.norm.insert(name.to_string(), percentile);
}
}
}
/// The preliminary blend of §12.4 on a 0100 scale: present-and-active
/// signals only, learned weights multiplied by their gate, renormalized to 1.
pub fn preliminary_blend(signals: &mut Signals, configured: &PreliminaryWeights) -> Option<f64> {
let candidates = [
("interest", configured.interest, 1.0),
("knn", configured.knn, signals.knn_gate),
("heuristic", configured.heuristic, 1.0),
("feed", configured.feed, signals.feed_gate),
("social", configured.social, 1.0),
];
let active = candidates
.into_iter()
.filter_map(|(name, weight, gate)| {
let norm = *signals.norm.get(name)?;
let effective = weight * gate;
(effective > 0.0).then_some((name, effective, norm))
})
.collect::<Vec<_>>();
let total = active.iter().map(|(_, weight, _)| weight).sum::<f64>();
if total <= 0.0 {
signals.weights.clear();
signals.preliminary = None;
return None;
}
signals.weights = active
.iter()
.map(|(name, weight, _)| ((*name).to_string(), weight / total))
.collect();
let blend = active
.iter()
.map(|(_, weight, norm)| weight / total * norm)
.sum::<f64>()
* 100.0;
signals.preliminary = Some(blend);
Some(blend)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::types::{ExtractMethod, SourceRef};
fn ranking() -> RankingConfig {
RankingConfig::default()
}
fn unit(values: &[f32]) -> Vec<f32> {
let norm = values.iter().map(|v| v * v).sum::<f32>().sqrt();
values.iter().map(|v| v / norm).collect()
}
fn example(id: ArticleId, label: &str, value: f64, embedding: &[f32]) -> RatedExample {
RatedExample {
article_id: id,
label: label.into(),
title: format!("rated {id}"),
value,
decay: 1.0,
embedding: unit(embedding),
feeds: vec![id],
}
}
fn article(id: ArticleId, feeds: &[FeedId]) -> Article {
Article {
id,
canonical_url: format!("https://example.com/{id}"),
title: format!("Article {id}"),
best_entry_id: id,
content_html: String::new(),
word_count: 1000,
excerpt_only: false,
image_count: 0,
sources: feeds
.iter()
.map(|feed| SourceRef {
entry_id: id,
feed_id: *feed,
feed_title: format!("feed {feed}"),
category: None,
kind: SourceKind::Feed,
})
.collect(),
first_seen: "2026-08-15T00:00:00Z".parse().unwrap(),
url: format!("https://example.com/{id}"),
author: None,
feed_id: feeds.first().copied().unwrap_or(0),
feed_title: String::new(),
category: None,
published_at: None,
comments_url: None,
image_urls: vec![],
social: vec![],
extract_method: ExtractMethod::Miniflux,
}
}
fn with_heuristic(value: Option<f64>) -> Signals {
Signals {
heuristic: value,
..Signals::default()
}
}
// --- §9.1 interest z-scores ---
fn interest_fixture(n: usize) -> (HashMap<ArticleId, Vec<f32>>, HashMap<String, Vec<f32>>) {
// Article 1 sits on axis x; the rest sit near axis y with a tiny spread.
let mut articles = HashMap::new();
articles.insert(1, unit(&[1.0, 0.0, 0.0]));
for id in 2..=n as ArticleId {
articles.insert(id, unit(&[0.0, 1.0, 0.001 * id as f32]));
}
// "Broad" is about equally close to everything; "Specific" matches only article 1.
let mut interests = HashMap::new();
interests.insert("Broad".to_string(), unit(&[1.0, 1.0, 0.0]));
interests.insert("Specific".to_string(), unit(&[1.0, 0.0, 0.0]));
(articles, interests)
}
#[test]
fn specific_interest_with_one_strong_match_beats_a_broad_one() {
let (articles, interests) = interest_fixture(40);
let matched = interest_matches(&articles, &interests);
let strong = &matched[&1];
assert_eq!(strong.top_interests[0].name, "Specific");
assert!(
strong.top_interests[0].z > 3.0,
"z = {}",
strong.top_interests[0].z
);
let others = (2..=40)
.map(|id| matched[&id].score)
.fold(f64::NEG_INFINITY, f64::max);
assert!(strong.score > others + 2.0, "{} vs {others}", strong.score);
// Raw cosine would have called Broad a near-tie everywhere (≈0.707).
assert!((matched[&2].top1_cos - 0.707).abs() < 0.01);
}
#[test]
fn interest_falls_back_to_raw_cosine_under_thirty_articles() {
let (articles, interests) = interest_fixture(10);
let matched = interest_matches(&articles, &interests);
for (id, m) in &matched {
assert!(
(m.score - m.top1_cos).abs() < 1e-9,
"article {id} should use raw top-1"
);
}
assert!((matched[&2].score - 0.707).abs() < 0.01);
}
// --- §9.2 preference ---
#[test]
fn one_loved_article_gives_a_positive_knn_to_a_near_neighbour() {
let mut ranking = ranking();
ranking.knn_floor = 0;
ranking.knn_full = 1;
let state = PreferenceState::build(vec![example(1, "loved", 1.0, &[1.0, 0.0])], &ranking);
let (knn, neighbours) = state.knn(&unit(&[0.9, 0.1]), &ranking);
assert!(knn.unwrap() > 0.9);
assert_eq!(neighbours.len(), 1);
assert_eq!(neighbours[0].label, "loved");
let (far, _) = state.knn(&unit(&[0.0, 1.0]), &ranking);
assert!(far.unwrap().abs() < 1e-6);
}
#[test]
fn two_unrelated_loved_clusters_both_score_high() {
let mut ranking = ranking();
ranking.knn_floor = 0;
ranking.knn_full = 1;
ranking.neighbour_k = 2;
let state = PreferenceState::build(
vec![
example(1, "loved", 1.0, &[1.0, 0.0, 0.0]),
example(2, "loved", 1.0, &[0.98, 0.02, 0.0]),
example(3, "loved", 1.0, &[0.0, 1.0, 0.0]),
example(4, "loved", 1.0, &[0.0, 0.98, 0.02]),
],
&ranking,
);
let (near_a, _) = state.knn(&unit(&[1.0, 0.0, 0.0]), &ranking);
let (near_b, _) = state.knn(&unit(&[0.0, 1.0, 0.0]), &ranking);
assert!(near_a.unwrap() > 0.95, "{near_a:?}");
assert!(near_b.unwrap() > 0.95, "{near_b:?}");
// A centroid would have put both at ~0.7.
}
#[test]
fn good_carries_a_third_of_loved() {
let loved = example(1, "loved", 1.0, &[1.0, 0.0]);
let good = example(2, "good", 0.35, &[1.0, 0.0]);
assert!((good.weight() / loved.weight() - 0.35).abs() < 1e-9);
// A mixed neighbourhood: the far example pulls the mean down by 0.35× as
// much weight when it is merely "good" as when it is "loved".
let mut ranking = ranking();
ranking.knn_floor = 0;
ranking.knn_full = 1;
let near = example(1, "loved", 1.0, &[1.0, 0.0]);
let candidate = unit(&[1.0, 0.0]);
let both_loved = PreferenceState::build(
vec![near.clone(), example(2, "loved", 1.0, &[0.0, 1.0])],
&ranking,
);
let one_good =
PreferenceState::build(vec![near, example(2, "good", 0.35, &[0.0, 1.0])], &ranking);
let pull_loved = 1.0 - both_loved.knn(&candidate, &ranking).0.unwrap();
let pull_good = 1.0 - one_good.knn(&candidate, &ranking).0.unwrap();
assert!(pull_good < pull_loved);
// Weighted means: 0.5 vs 1/1.35 → pulls 0.5 vs 0.35/1.35.
assert!((pull_good / pull_loved - 0.35 / 1.35 / 0.5).abs() < 1e-9);
}
#[test]
fn negatives_subtract_with_the_negative_coefficient() {
let mut ranking = ranking();
ranking.knn_floor = 0;
ranking.knn_full = 1;
let state =
PreferenceState::build(vec![example(1, "not_for_me", -1.0, &[1.0, 0.0])], &ranking);
let (knn, _) = state.knn(&unit(&[1.0, 0.0]), &ranking);
assert!((knn.unwrap() + ranking.negative_coefficient).abs() < 1e-9);
}
#[test]
fn decay_halves_at_the_half_life() {
assert!((decay(60.0, 60.0) - 0.5).abs() < 1e-12);
assert!((decay(0.0, 60.0) - 1.0).abs() < 1e-12);
assert!((decay(120.0, 60.0) - 0.25).abs() < 1e-12);
}
#[test]
fn gate_is_zero_below_floor_one_at_full_and_linear_between() {
assert_eq!(gate(0, 8, 25), 0.0);
assert_eq!(gate(8, 8, 25), 0.0);
assert_eq!(gate(25, 8, 25), 1.0);
assert_eq!(gate(100, 8, 25), 1.0);
assert!((gate(16, 8, 24) - 0.5).abs() < 1e-9);
assert!((gate(9, 8, 25) - 1.0 / 17.0).abs() < 1e-9);
}
#[test]
fn knn_is_absent_when_the_gate_is_closed() {
let ranking = ranking(); // knn_floor 8
let state = PreferenceState::build(vec![example(1, "loved", 1.0, &[1.0, 0.0])], &ranking);
assert_eq!(state.knn_gate, 0.0);
assert_eq!(state.knn(&unit(&[1.0, 0.0]), &ranking), (None, Vec::new()));
}
// --- §9.3 feed affinity ---
#[test]
fn feed_credit_sums_to_one_across_direct_feeds() {
let mut rated = example(1, "loved", 1.0, &[1.0, 0.0]);
rated.feeds = vec![10, 20, 30];
let state = PreferenceState::build(vec![rated], &ranking());
let total: f64 = [10, 20, 30]
.iter()
.map(|feed| state.feed_credit(*feed).unwrap().0)
.sum();
assert!((total - 1.0).abs() < 1e-9);
assert!((state.feed_credit(10).unwrap().0 - 1.0 / 3.0).abs() < 1e-9);
}
#[test]
fn feed_affinity_is_the_mean_over_rated_feeds() {
let mut ranking = ranking();
ranking.feed_floor = 0;
ranking.feed_full = 1;
let mut loved = example(1, "loved", 1.0, &[1.0, 0.0]);
loved.feeds = vec![10];
let mut down = example(2, "not_for_me", -1.0, &[1.0, 0.0]);
down.feeds = vec![20];
let state = PreferenceState::build(vec![loved, down], &ranking);
// feed 10: (1+1)/(1+0+2) = 2/3; feed 20: (0+1)/(0+1+2) = 1/3; unrated 99 ignored.
let both = state.feed(&article(7, &[10, 20, 99])).unwrap();
assert!((both - 0.5).abs() < 1e-9, "{both}");
let best = state.feed(&article(8, &[10])).unwrap();
assert!((best - 2.0 / 3.0).abs() < 1e-9);
assert_eq!(state.feed(&article(9, &[99])), None);
}
#[test]
fn feed_is_absent_when_the_gate_is_closed() {
let ranking = ranking(); // feed_floor 15
let mut loved = example(1, "loved", 1.0, &[1.0, 0.0]);
loved.feeds = vec![10];
let state = PreferenceState::build(vec![loved], &ranking);
assert_eq!(state.feed_gate, 0.0);
assert_eq!(state.feed(&article(7, &[10])), None);
}
#[test]
fn direct_feeds_fall_back_to_the_best_entry_feed() {
let mut a = article(1, &[]);
a.feed_id = 42;
assert_eq!(direct_feeds(&a), vec![42]);
assert_eq!(direct_feeds(&article(2, &[5, 3, 5])), vec![3, 5]);
}
// --- §12.2 normalization, §12.4 blend ---
#[test]
fn constant_signal_normalizes_to_half_for_everyone() {
let mut values = [with_heuristic(Some(7.0)), with_heuristic(Some(7.0))];
let mut refs = values.iter_mut().collect::<Vec<_>>();
normalize(&mut refs);
assert!(values.iter().all(|v| v.norm["heuristic"] == 0.5));
let mut single = [with_heuristic(Some(3.0))];
let mut refs = single.iter_mut().collect::<Vec<_>>();
normalize(&mut refs);
assert_eq!(single[0].norm["heuristic"], 0.5);
}
#[test]
fn ties_get_equal_percentiles_without_an_id_ramp() {
let mut values = (0..400)
.map(|_| with_heuristic(Some(0.0)))
.collect::<Vec<_>>();
let mut refs = values.iter_mut().collect::<Vec<_>>();
normalize(&mut refs);
assert!(values.iter().all(|v| v.norm["heuristic"] == 0.5));
let mut mixed = [
with_heuristic(Some(1.0)),
with_heuristic(Some(2.0)),
with_heuristic(Some(2.0)),
with_heuristic(Some(3.0)),
];
let mut refs = mixed.iter_mut().collect::<Vec<_>>();
normalize(&mut refs);
assert_eq!(mixed[0].norm["heuristic"], 0.25);
assert_eq!(mixed[1].norm["heuristic"], 0.625);
assert_eq!(mixed[2].norm["heuristic"], 0.625);
assert_eq!(mixed[3].norm["heuristic"], 1.0);
}
#[test]
fn absent_values_do_not_shift_present_values() {
let mut values = [
with_heuristic(Some(1.0)),
with_heuristic(Some(2.0)),
with_heuristic(None),
];
let mut refs = values.iter_mut().collect::<Vec<_>>();
normalize(&mut refs);
assert!(!values[2].norm.contains_key("heuristic"));
// n_present = 2: the absent third value does not widen the scale.
assert_eq!(values[0].norm["heuristic"], 0.5);
assert_eq!(values[1].norm["heuristic"], 1.0);
}
#[test]
fn effective_weights_sum_to_one_and_missing_signals_are_skipped() {
let mut signals = Signals {
interest: Some(1.0),
heuristic: Some(2.0),
norm: BTreeMap::from([("interest".into(), 0.8), ("heuristic".into(), 0.4)]),
..Signals::default()
};
let blend = preliminary_blend(&mut signals, &PreliminaryWeights::default()).unwrap();
assert!((signals.weights.values().sum::<f64>() - 1.0).abs() < 1e-9);
assert!(!signals.weights.contains_key("knn"));
assert!(!signals.weights.contains_key("social"));
// 0.35/0.55 × 0.8 + 0.20/0.55 × 0.4 = 0.6545…
assert!((blend - 65.4545).abs() < 0.01, "{blend}");
let mut only_heuristic = Signals {
heuristic: Some(2.0),
norm: BTreeMap::from([("heuristic".into(), 0.4)]),
..Signals::default()
};
let blend = preliminary_blend(&mut only_heuristic, &PreliminaryWeights::default());
assert!((blend.unwrap() - 40.0).abs() < 1e-9);
assert_eq!(only_heuristic.weights["heuristic"], 1.0);
}
#[test]
fn learned_weights_are_multiplied_by_their_gate() {
let mut signals = Signals {
knn: Some(0.5),
heuristic: Some(2.0),
knn_gate: 0.5,
norm: BTreeMap::from([("knn".into(), 1.0), ("heuristic".into(), 0.0)]),
..Signals::default()
};
preliminary_blend(&mut signals, &PreliminaryWeights::default());
// knn 0.25 × 0.5 = 0.125 against heuristic 0.20.
assert!((signals.weights["knn"] - 0.125 / 0.325).abs() < 1e-9);
}
#[test]
fn compute_scores_every_article_and_leaves_ungated_signals_absent() {
let articles = vec![article(1, &[10]), article(2, &[20]), article(3, &[30])];
let mut embeddings = HashMap::new();
embeddings.insert(1, unit(&[1.0, 0.0]));
embeddings.insert(2, unit(&[0.0, 1.0]));
let mut interests = HashMap::new();
interests.insert("Axis".to_string(), unit(&[1.0, 0.0]));
let state = PreferenceState::build(vec![example(9, "loved", 1.0, &[1.0, 0.0])], &ranking());
let signals = compute(&articles, &embeddings, &interests, &state, &ranking());
assert_eq!(signals.len(), 3);
assert!(signals[&1].interest.is_some());
assert!(signals[&3].interest.is_none(), "no embedding → absent");
assert!(
signals
.values()
.all(|s| s.knn.is_none() && s.feed.is_none())
);
assert!(
signals
.values()
.all(|s| s.heuristic.is_some() && s.preliminary.is_some())
);
}
}
+962
View File
@@ -0,0 +1,962 @@
//! Per-run candidate telemetry: the `candidate_runs` writer, `signals_json`,
//! the `explain` command and feature retention (plan §7.47.5, §15.2, §16).
//!
//! One row per considered article per run says where it stopped and why. Rows
//! are upserted on every stage transition with every column set (never
//! `COALESCE`), so the last write for a run is the whole truth.
use std::collections::BTreeMap;
use std::fmt::Write as _;
use jiff::Timestamp;
use jiff::civil::Date;
use serde::{Deserialize, Serialize};
use sqlx::Row as _;
use crate::curate::signals::{Neighbour, Signals, TopInterest};
use crate::db::{Db, fmt_ts};
use crate::types::ArticleId;
/// The stage vocabulary of §7.4, in pipeline order.
pub const STAGES: [&str; 7] = [
"excluded",
"eligible",
"triaged",
"admitted",
"assessed",
"shortlisted",
"selected",
];
/// Signal names rendered by `explain`, including the LLM ones steps 45 add.
const RENDERED_SIGNALS: [&str; 8] = [
"interest",
"knn",
"feed",
"social",
"heuristic",
"triage",
"quality",
"fit",
];
/// One `candidate_runs` row (§7.4).
#[derive(Debug, Clone)]
pub struct CandidateRun<'a> {
pub run_id: i64,
pub article_id: ArticleId,
pub stage: &'a str,
pub excluded_reason: Option<&'a str>,
/// JSON array of retriever names, first = the one that admitted it.
pub admitted_by: Option<&'a str>,
pub signals_json: &'a str,
pub utility: Option<f64>,
pub rank_utility: Option<i64>,
pub cluster_id: Option<i64>,
pub cluster_rank: Option<i64>,
pub editor_why: Option<&'a str>,
}
/// Upsert one row, setting every column (§7.4).
pub async fn write(db: &Db, row: &CandidateRun<'_>) -> Result<(), sqlx::Error> {
sqlx::query(
"INSERT INTO candidate_runs
(run_id, article_id, stage, excluded_reason, admitted_by, signals_json,
utility, rank_utility, cluster_id, cluster_rank, editor_why)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(run_id, article_id) DO UPDATE SET
stage = excluded.stage,
excluded_reason = excluded.excluded_reason,
admitted_by = excluded.admitted_by,
signals_json = excluded.signals_json,
utility = excluded.utility,
rank_utility = excluded.rank_utility,
cluster_id = excluded.cluster_id,
cluster_rank = excluded.cluster_rank,
editor_why = excluded.editor_why",
)
.bind(row.run_id)
.bind(row.article_id)
.bind(row.stage)
.bind(row.excluded_reason)
.bind(row.admitted_by)
.bind(row.signals_json)
.bind(row.utility)
.bind(row.rank_utility)
.bind(row.cluster_id)
.bind(row.cluster_rank)
.bind(row.editor_why)
.execute(db.pool())
.await?;
Ok(())
}
/// The thin row of a hygiene exclusion: keys, `stage = 'excluded'`, the reason
/// and `signals_json = '{}'` (§8.1).
pub async fn thin_excluded(
db: &Db,
run_id: i64,
article_id: ArticleId,
reason: &str,
) -> Result<(), sqlx::Error> {
write(
db,
&CandidateRun {
run_id,
article_id,
stage: "excluded",
excluded_reason: Some(reason),
admitted_by: None,
signals_json: "{}",
utility: None,
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: None,
},
)
.await
}
/// `signals_json` (§7.5). Missing signals are absent from `raw`/`norm` and
/// `false` in `present`; `weights` are the effective weights.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct SignalsJson {
pub v: i64,
#[serde(default)]
pub raw: BTreeMap<String, f64>,
#[serde(default)]
pub norm: BTreeMap<String, f64>,
#[serde(default)]
pub present: BTreeMap<String, bool>,
#[serde(default)]
pub weights: BTreeMap<String, f64>,
#[serde(default)]
pub top_interests: Vec<TopInterest>,
#[serde(default)]
pub neighbours: Vec<Neighbour>,
#[serde(default)]
pub exploration: bool,
#[serde(default)]
pub auto_include: bool,
#[serde(default)]
pub notes: Vec<String>,
}
impl SignalsJson {
/// The blend implied by the stored effective weights, 0100.
pub fn blend(&self) -> Option<f64> {
let mut score = 0.0;
let mut any = false;
for (name, weight) in &self.weights {
if let Some(value) = self.norm.get(name) {
score += weight * value;
any = true;
}
}
any.then_some(score * 100.0)
}
}
/// Serialize the signals of §7.5 for one article.
pub fn serialize_signals(signals: &Signals, auto_include: bool) -> String {
let mut raw = BTreeMap::new();
for name in [
"interest",
"interest_top1_cos",
"knn",
"feed",
"social",
"heuristic",
] {
if let Some(value) = signals.raw(name) {
raw.insert(name.to_string(), value);
}
}
let present = RENDERED_SIGNALS
.into_iter()
.map(|name| (name.to_string(), signals.present(name)))
.collect();
serde_json::to_string(&SignalsJson {
v: 1,
raw,
norm: signals.norm.clone(),
present,
weights: signals.weights.clone(),
top_interests: signals.top_interests.clone(),
neighbours: signals.neighbours.clone(),
exploration: false,
auto_include,
notes: signals.notes.clone(),
})
.unwrap_or_else(|_| "{}".into())
}
// ---------------------------------------------------------------------------
// `explain` (§15.2)
// ---------------------------------------------------------------------------
/// A `candidate_runs` row joined to its article title.
#[derive(Debug, Clone)]
pub struct ExplainRow {
pub run_id: i64,
pub article_id: ArticleId,
pub title: String,
pub stage: String,
pub excluded_reason: Option<String>,
pub admitted_by: Option<String>,
pub signals_json: String,
pub utility: Option<f64>,
pub rank_utility: Option<i64>,
pub cluster_id: Option<i64>,
pub cluster_rank: Option<i64>,
pub editor_why: Option<String>,
}
impl ExplainRow {
fn from_row(row: &sqlx::sqlite::SqliteRow) -> Self {
Self {
run_id: row.get("run_id"),
article_id: row.get("article_id"),
title: row.get("title"),
stage: row.get("stage"),
excluded_reason: row.get("excluded_reason"),
admitted_by: row.get("admitted_by"),
signals_json: row.get("signals_json"),
utility: row.get("utility"),
rank_utility: row.get("rank_utility"),
cluster_id: row.get("cluster_id"),
cluster_rank: row.get("cluster_rank"),
editor_why: row.get("editor_why"),
}
}
pub fn signals(&self) -> Option<SignalsJson> {
serde_json::from_str(&self.signals_json).ok()
}
/// Utility when step 5 has written it, else the preliminary blend.
pub fn score(&self) -> Option<f64> {
self.utility
.or_else(|| self.signals().and_then(|signals| signals.blend()))
}
}
/// The run `explain` reads: `--run-id` when given (and of that date), else the
/// latest non-dry run of the date.
pub async fn resolve_run(
db: &Db,
date: Date,
requested: Option<i64>,
) -> Result<Option<i64>, sqlx::Error> {
let row = match requested {
Some(run_id) => {
sqlx::query("SELECT id FROM runs WHERE id = ? AND date = ?")
.bind(run_id)
.bind(date.to_string())
.fetch_optional(db.pool())
.await?
}
None => {
sqlx::query(
"SELECT id FROM runs WHERE date = ? AND status != 'dry_run'
ORDER BY id DESC LIMIT 1",
)
.bind(date.to_string())
.fetch_optional(db.pool())
.await?
}
};
Ok(row.map(|row| row.get("id")))
}
pub async fn explain_row(
db: &Db,
run_id: i64,
article_id: ArticleId,
) -> Result<Option<ExplainRow>, sqlx::Error> {
let row = sqlx::query(
"SELECT cr.run_id, cr.article_id, COALESCE(a.title, '') AS title,
cr.stage, cr.excluded_reason, cr.admitted_by, cr.signals_json,
cr.utility, cr.rank_utility, cr.cluster_id, cr.cluster_rank, cr.editor_why
FROM candidate_runs cr JOIN articles a ON a.id = cr.article_id
WHERE cr.run_id = ? AND cr.article_id = ?",
)
.bind(run_id)
.bind(article_id)
.fetch_optional(db.pool())
.await?;
Ok(row.as_ref().map(ExplainRow::from_row))
}
/// The top `limit` rows by utility-or-blend that were not selected (§15.2).
pub async fn near_misses(
db: &Db,
run_id: i64,
limit: usize,
) -> Result<Vec<ExplainRow>, sqlx::Error> {
let rows = sqlx::query(
"SELECT cr.run_id, cr.article_id, COALESCE(a.title, '') AS title,
cr.stage, cr.excluded_reason, cr.admitted_by, cr.signals_json,
cr.utility, cr.rank_utility, cr.cluster_id, cr.cluster_rank, cr.editor_why
FROM candidate_runs cr JOIN articles a ON a.id = cr.article_id
WHERE cr.run_id = ? AND cr.stage != 'selected' AND cr.stage != 'excluded'",
)
.bind(run_id)
.fetch_all(db.pool())
.await?;
let mut output = rows.iter().map(ExplainRow::from_row).collect::<Vec<_>>();
output.sort_by(|left, right| {
right
.score()
.unwrap_or(f64::NEG_INFINITY)
.total_cmp(&left.score().unwrap_or(f64::NEG_INFINITY))
.then_with(|| left.article_id.cmp(&right.article_id))
});
output.truncate(limit);
Ok(output)
}
fn fmt_opt(value: Option<f64>) -> String {
value
.map(|v| format!("{v:.3}"))
.unwrap_or_else(|| "".into())
}
/// Render one persisted row the way §15.2 lists it.
pub async fn render_explain(db: &Db, row: &ExplainRow) -> Result<String, sqlx::Error> {
let mut out = String::new();
let _ = writeln!(out, "article {}: {}", row.article_id, row.title);
let _ = write!(out, "run {} · stage: {}", row.run_id, row.stage);
if let Some(reason) = &row.excluded_reason {
let _ = write!(out, " · reason: {reason}");
}
let _ = writeln!(out);
if let Some(signals) = row.signals() {
let _ = writeln!(out, "signals (raw · norm · weight):");
for name in RENDERED_SIGNALS {
let present = signals.present.get(name).copied().unwrap_or(false);
if present {
let _ = writeln!(
out,
" {name:<10} {:>8} · {:>6} · {:>6}",
fmt_opt(signals.raw.get(name).copied()),
fmt_opt(signals.norm.get(name).copied()),
fmt_opt(signals.weights.get(name).copied()),
);
} else {
let _ = writeln!(out, " {name:<10} absent");
}
}
if let Some(blend) = signals.blend() {
let _ = writeln!(out, "preliminary blend: {blend:.1}");
}
if let Some(cos) = signals.raw.get("interest_top1_cos") {
let _ = writeln!(out, "interest top-1 cosine: {cos:.3}");
}
if !signals.top_interests.is_empty() {
let _ = writeln!(out, "top interests:");
for interest in &signals.top_interests {
let _ = writeln!(
out,
" {} · z {:.2} · cos {:.3}",
interest.name, interest.z, interest.cos
);
}
}
if !signals.neighbours.is_empty() {
let _ = writeln!(out, "nearest rated neighbours:");
for neighbour in &signals.neighbours {
let _ = writeln!(
out,
" {} · cos {:.3} · article {} · {}",
neighbour.label, neighbour.cos, neighbour.article_id, neighbour.title
);
}
}
if signals.exploration || signals.auto_include {
let _ = writeln!(
out,
"flags: exploration={} auto_include={}",
signals.exploration, signals.auto_include
);
}
for note in &signals.notes {
let _ = writeln!(out, "note: {note}");
}
}
let assessments = sqlx::query(
"SELECT stage, model, score, fit, kind, facets_json, rationale, category,
paywalled_guess, assessed_at
FROM article_assessments WHERE article_id = ? ORDER BY stage",
)
.bind(row.article_id)
.fetch_all(db.pool())
.await?;
if !assessments.is_empty() {
let _ = writeln!(out, "assessments:");
for assessment in assessments {
let _ = writeln!(
out,
" {} · {} · score {} · fit {} · kind {} · category {} · paywalled={} · {}",
assessment.get::<String, _>("stage"),
assessment.get::<String, _>("model"),
fmt_opt(assessment.get::<Option<f64>, _>("score")),
fmt_opt(assessment.get::<Option<f64>, _>("fit")),
assessment
.get::<Option<String>, _>("kind")
.unwrap_or_else(|| "".into()),
assessment
.get::<Option<String>, _>("category")
.unwrap_or_else(|| "".into()),
assessment.get::<i64, _>("paywalled_guess") != 0,
assessment
.get::<Option<String>, _>("rationale")
.unwrap_or_default(),
);
if let Some(facets) = assessment.get::<Option<String>, _>("facets_json") {
let _ = writeln!(out, " facets: {facets}");
}
}
}
if row.utility.is_some() || row.rank_utility.is_some() {
let _ = writeln!(
out,
"utility: {} · rank {}",
fmt_opt(row.utility),
row.rank_utility
.map(|r| r.to_string())
.unwrap_or_else(|| "".into())
);
}
if let Some(cluster) = row.cluster_id {
let _ = writeln!(
out,
"cluster: {cluster} · rank {}",
row.cluster_rank
.map(|r| r.to_string())
.unwrap_or_else(|| "".into())
);
}
if let Some(admitted_by) = &row.admitted_by {
let _ = writeln!(out, "admitted by: {admitted_by}");
}
if let Some(why) = &row.editor_why {
let _ = writeln!(out, "editor: {why}");
}
Ok(out)
}
/// What `explain` was asked about.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ExplainTarget {
Article(ArticleId),
Url(String),
}
/// `explain --date D (--article ID | --url URL) [--run-id N]` as text (§15.2).
pub async fn explain(
db: &Db,
date: Date,
run_id: Option<i64>,
target: &ExplainTarget,
) -> anyhow::Result<String> {
let article_id = match target {
ExplainTarget::Article(id) => *id,
ExplainTarget::Url(url) => {
let canonical = crate::dedupe::canonical_url(url)
.ok_or_else(|| anyhow::anyhow!("invalid article URL {url:?}"))?;
match db.article_id_for_url(&canonical).await? {
Some(id) => id,
None => {
return Ok(format!(
"{canonical} was never ingested: it is not in `articles`, so this is a feed problem, not a ranking problem."
));
}
}
}
};
if db.get_article(article_id).await?.is_none() {
return Ok(format!(
"article {article_id} was never ingested: it is not in `articles`."
));
}
let Some(run_id) = resolve_run(db, date, run_id).await? else {
return Ok(match run_id {
Some(id) => format!("run {id} is not a run for {date}"),
None => format!("no non-dry run recorded for {date}"),
});
};
match explain_row(db, run_id, article_id).await? {
Some(row) => Ok(render_explain(db, &row).await?),
None => Ok(format!(
"article {article_id} was not considered by run {run_id} for {date} (outside its ingest window, or telemetry pruned)."
)),
}
}
/// `explain --date D --near-misses [N]` as text (§15.2).
pub async fn explain_near_misses(
db: &Db,
date: Date,
run_id: Option<i64>,
limit: usize,
) -> anyhow::Result<String> {
let Some(run_id) = resolve_run(db, date, run_id).await? else {
return Ok(format!("no non-dry run recorded for {date}"));
};
let rows = near_misses(db, run_id, limit).await?;
let mut out = String::new();
let _ = writeln!(
out,
"run {run_id} · {date} · top {} not selected, by {}:",
rows.len(),
if rows.iter().any(|row| row.utility.is_some()) {
"utility"
} else {
"preliminary blend"
}
);
for (index, row) in rows.iter().enumerate() {
let reason = row
.excluded_reason
.as_deref()
.map(|reason| format!(", {reason}"))
.unwrap_or_default();
let _ = writeln!(
out,
"{:>3}. {:>6} · {} · {}{} · article {}",
index + 1,
row.score()
.map(|score| format!("{score:.1}"))
.unwrap_or_else(|| "".into()),
row.title,
row.stage,
reason,
row.article_id
);
}
Ok(out)
}
// ---------------------------------------------------------------------------
// `features prune` (§7.1, §7.4)
// ---------------------------------------------------------------------------
/// Delete `article_embeddings` for articles neither rated nor published that
/// are older than `embedding_retention_days`, and `candidate_runs` rows whose
/// run started more than `telemetry_retention_days` ago. Returns the counts.
pub async fn prune(
db: &Db,
embedding_retention_days: i64,
telemetry_retention_days: i64,
now: Timestamp,
) -> Result<(u64, u64), sqlx::Error> {
let cutoff = |days: i64| {
now.checked_sub(jiff::Span::new().hours(days.max(0).saturating_mul(24)))
.unwrap_or(Timestamp::UNIX_EPOCH)
};
let embeddings = sqlx::query(
"DELETE FROM article_embeddings
WHERE article_id IN (
SELECT ae.article_id
FROM article_embeddings ae JOIN articles a ON a.id = ae.article_id
WHERE a.first_seen < ?
AND NOT EXISTS (SELECT 1 FROM rating_events re WHERE re.article_id = ae.article_id)
AND NOT EXISTS (SELECT 1 FROM issue_articles ia WHERE ia.article_id = ae.article_id)
)",
)
.bind(fmt_ts(cutoff(embedding_retention_days)))
.execute(db.pool())
.await?
.rows_affected();
let telemetry = sqlx::query(
"DELETE FROM candidate_runs
WHERE run_id IN (SELECT id FROM runs WHERE started_at < ?)",
)
.bind(fmt_ts(cutoff(telemetry_retention_days)))
.execute(db.pool())
.await?
.rows_affected();
Ok((embeddings, telemetry))
}
#[cfg(test)]
mod tests {
use super::*;
use crate::curate::embedding::encode_blob;
async fn db_with_articles(ids: &[ArticleId]) -> (tempfile::TempDir, Db) {
let dir = tempfile::tempdir().unwrap();
let db = Db::open_and_migrate(&dir.path().join("telemetry.db"))
.await
.unwrap();
for id in ids {
sqlx::query(
"INSERT INTO articles (id, canonical_url, title, first_seen)
VALUES (?, ?, ?, '2026-08-15T00:00:00Z')",
)
.bind(id)
.bind(format!("https://example.com/{id}"))
.bind(format!("Article {id}"))
.execute(db.pool())
.await
.unwrap();
}
(dir, db)
}
fn date() -> Date {
"2026-09-02".parse().unwrap()
}
fn signals(heuristic: f64, norm: f64) -> Signals {
Signals {
interest: Some(1.2),
interest_top1_cos: Some(0.61),
heuristic: Some(heuristic),
norm: BTreeMap::from([("heuristic".into(), norm), ("interest".into(), 0.9)]),
weights: BTreeMap::from([
("heuristic".into(), 0.2 / 0.55),
("interest".into(), 0.35 / 0.55),
]),
top_interests: vec![TopInterest {
name: "Gaussian Splatting".into(),
z: 3.4,
cos: 0.61,
}],
neighbours: vec![Neighbour {
article_id: 812,
label: "loved".into(),
cos: 0.71,
title: "A rated piece".into(),
}],
notes: vec!["knn gate 0.60 (n=14 rated with embeddings)".into()],
..Signals::default()
}
}
#[test]
fn signals_json_follows_the_plan_shape() {
let json = serialize_signals(&signals(41.0, 0.55), false);
let parsed: serde_json::Value = serde_json::from_str(&json).unwrap();
assert_eq!(parsed["v"], 1);
assert_eq!(parsed["raw"]["heuristic"], 41.0);
assert_eq!(parsed["raw"]["interest_top1_cos"], 0.61);
assert_eq!(parsed["present"]["heuristic"], true);
assert_eq!(parsed["present"]["knn"], false);
assert_eq!(parsed["present"]["quality"], false);
assert!(parsed["raw"].get("knn").is_none());
assert!(parsed["norm"].get("knn").is_none());
assert_eq!(parsed["exploration"], false);
assert_eq!(parsed["auto_include"], false);
assert_eq!(parsed["top_interests"][0]["name"], "Gaussian Splatting");
assert_eq!(parsed["neighbours"][0]["article_id"], 812);
assert_eq!(
parsed["notes"][0],
"knn gate 0.60 (n=14 rated with embeddings)"
);
let typed: SignalsJson = serde_json::from_str(&json).unwrap();
let weights: f64 = typed.weights.values().sum();
assert!((weights - 1.0).abs() < 1e-9);
assert!(typed.blend().is_some());
}
#[tokio::test]
async fn rows_are_upserted_with_every_column_replaced() {
let (_dir, db) = db_with_articles(&[1]).await;
let run_id = db.start_run(date(), Timestamp::now()).await.unwrap();
write(
&db,
&CandidateRun {
run_id,
article_id: 1,
stage: "eligible",
excluded_reason: Some("not_admitted"),
admitted_by: None,
signals_json: "{}",
utility: Some(1.0),
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: None,
},
)
.await
.unwrap();
write(
&db,
&CandidateRun {
run_id,
article_id: 1,
stage: "selected",
excluded_reason: None,
admitted_by: Some("[\"prefilter\"]"),
signals_json: "{\"v\":1}",
utility: None,
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: Some("because"),
},
)
.await
.unwrap();
let row = explain_row(&db, run_id, 1).await.unwrap().unwrap();
assert_eq!(row.stage, "selected");
assert_eq!(row.excluded_reason, None, "no COALESCE");
assert_eq!(row.utility, None);
assert_eq!(row.admitted_by.as_deref(), Some("[\"prefilter\"]"));
assert_eq!(row.editor_why.as_deref(), Some("because"));
let count: i64 = sqlx::query_scalar("SELECT COUNT(*) FROM candidate_runs")
.fetch_one(db.pool())
.await
.unwrap();
assert_eq!(count, 1);
}
#[tokio::test]
async fn explain_renders_persisted_rows_and_reports_never_ingested() {
let (_dir, db) = db_with_articles(&[1, 2]).await;
let run_id = db.start_run(date(), Timestamp::now()).await.unwrap();
thin_excluded(&db, run_id, 2, "blocked").await.unwrap();
let json = serialize_signals(&signals(41.0, 0.55), true);
write(
&db,
&CandidateRun {
run_id,
article_id: 1,
stage: "shortlisted",
excluded_reason: Some("not_selected"),
admitted_by: Some("[\"prefilter\"]"),
signals_json: &json,
utility: None,
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: None,
},
)
.await
.unwrap();
// A later dry run must not shadow the real one.
let dry = db.start_run(date(), Timestamp::now()).await.unwrap();
sqlx::query("UPDATE runs SET status = 'dry_run' WHERE id = ?")
.bind(dry)
.execute(db.pool())
.await
.unwrap();
assert_eq!(resolve_run(&db, date(), None).await.unwrap(), Some(run_id));
assert_eq!(
resolve_run(&db, date(), Some(dry)).await.unwrap(),
Some(dry)
);
assert_eq!(
resolve_run(&db, "2026-01-01".parse().unwrap(), Some(dry))
.await
.unwrap(),
None
);
let text = explain(&db, date(), None, &ExplainTarget::Article(1))
.await
.unwrap();
assert!(text.contains("article 1: Article 1"), "{text}");
assert!(
text.contains("stage: shortlisted · reason: not_selected"),
"{text}"
);
let squashed = text.split_whitespace().collect::<Vec<_>>().join(" ");
assert!(
squashed.contains("heuristic 41.000 · 0.550 · 0.364"),
"{text}"
);
assert!(squashed.contains("knn absent"), "{text}");
assert!(squashed.contains("quality absent"), "{text}");
assert!(
text.contains("Gaussian Splatting · z 3.40 · cos 0.610"),
"{text}"
);
assert!(
text.contains("loved · cos 0.710 · article 812 · A rated piece"),
"{text}"
);
assert!(text.contains("admitted by: [\"prefilter\"]"), "{text}");
assert!(text.contains("auto_include=true"), "{text}");
assert!(text.contains("preliminary blend:"), "{text}");
assert!(text.contains("note: knn gate"), "{text}");
let by_url = explain(
&db,
date(),
None,
&ExplainTarget::Url("https://example.com/1?utm_source=x".into()),
)
.await
.unwrap();
assert_eq!(by_url, text, "--url canonicalizes and finds the same row");
let thin = explain(&db, date(), None, &ExplainTarget::Article(2))
.await
.unwrap();
assert!(thin.contains("stage: excluded · reason: blocked"), "{thin}");
let missing = explain(
&db,
date(),
None,
&ExplainTarget::Url("https://nowhere.example/post".into()),
)
.await
.unwrap();
assert!(missing.contains("never ingested"), "{missing}");
let missing_id = explain(&db, date(), None, &ExplainTarget::Article(99))
.await
.unwrap();
assert!(missing_id.contains("never ingested"), "{missing_id}");
let no_run = explain(
&db,
"2026-01-01".parse().unwrap(),
None,
&ExplainTarget::Article(1),
)
.await
.unwrap();
assert!(no_run.contains("no non-dry run"), "{no_run}");
let not_considered = explain(&db, date(), Some(dry), &ExplainTarget::Article(1))
.await
.unwrap();
assert!(
not_considered.contains("was not considered by run"),
"{not_considered}"
);
}
#[tokio::test]
async fn near_misses_rank_by_blend_and_skip_selected_and_excluded() {
let (_dir, db) = db_with_articles(&[1, 2, 3, 4, 5]).await;
let run_id = db.start_run(date(), Timestamp::now()).await.unwrap();
let rows = [
(1, "selected", None, 0.9),
(2, "shortlisted", Some("not_selected"), 0.7),
(3, "eligible", Some("not_admitted"), 0.95),
(4, "shortlisted", Some("not_selected"), 0.1),
];
for (id, stage, reason, norm) in rows {
let json = serialize_signals(&signals(10.0, norm), false);
write(
&db,
&CandidateRun {
run_id,
article_id: id,
stage,
excluded_reason: reason,
admitted_by: None,
signals_json: &json,
utility: None,
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: None,
},
)
.await
.unwrap();
}
thin_excluded(&db, run_id, 5, "published_before")
.await
.unwrap();
let misses = near_misses(&db, run_id, 10).await.unwrap();
assert_eq!(
misses.iter().map(|row| row.article_id).collect::<Vec<_>>(),
vec![3, 2, 4]
);
let text = explain_near_misses(&db, date(), None, 2).await.unwrap();
assert!(
text.contains("top 2 not selected, by preliminary blend"),
"{text}"
);
assert!(
text.contains("Article 3 · eligible, not_admitted"),
"{text}"
);
assert!(!text.contains("Article 4"), "{text}");
}
#[tokio::test]
async fn prune_respects_rated_and_published() {
let (_dir, db) = db_with_articles(&[1, 2, 3, 4]).await;
let now = Timestamp::now();
let old = fmt_ts(now - jiff::Span::new().hours(200 * 24));
sqlx::query("UPDATE articles SET first_seen = ? WHERE id IN (1, 2, 3)")
.bind(&old)
.execute(db.pool())
.await
.unwrap();
let blob = encode_blob(&[0.5, 0.5]).unwrap();
for id in 1..=4 {
sqlx::query(
"INSERT INTO article_embeddings
(article_id, model, dimension, input_hash, embedding, created_at)
VALUES (?, 'voyage-4-lite', 2, 'h', ?, ?)",
)
.bind(id)
.bind(&blob)
.bind(&old)
.execute(db.pool())
.await
.unwrap();
}
sqlx::query(
"INSERT INTO rating_events (article_id, kind, source, label, value, event_at)
VALUES (1, 'explicit', 'cli', 'loved', 1.0, ?)",
)
.bind(&old)
.execute(db.pool())
.await
.unwrap();
sqlx::query(
"INSERT INTO issues (date, issue_number, generated_at) VALUES ('2026-02-01', 1, ?);
INSERT INTO issue_articles (issue_date, article_id, section) VALUES ('2026-02-01', 2, 'Top Stories');",
)
.bind(&old)
.execute(db.pool())
.await
.unwrap();
let old_run = db
.start_run("2026-02-01".parse().unwrap(), now)
.await
.unwrap();
sqlx::query("UPDATE runs SET started_at = ? WHERE id = ?")
.bind(&old)
.bind(old_run)
.execute(db.pool())
.await
.unwrap();
let new_run = db.start_run(date(), now).await.unwrap();
thin_excluded(&db, old_run, 1, "blocked").await.unwrap();
thin_excluded(&db, new_run, 1, "blocked").await.unwrap();
let (embeddings, telemetry) = prune(&db, 120, 180, now).await.unwrap();
assert_eq!(
embeddings, 1,
"only the old, unrated, unpublished article 3"
);
assert_eq!(telemetry, 1, "only the old run's rows");
let remaining: Vec<i64> =
sqlx::query_scalar("SELECT article_id FROM article_embeddings ORDER BY article_id")
.fetch_all(db.pool())
.await
.unwrap();
assert_eq!(remaining, vec![1, 2, 4]);
let runs: Vec<i64> = sqlx::query_scalar("SELECT run_id FROM candidate_runs")
.fetch_all(db.pool())
.await
.unwrap();
assert_eq!(runs, vec![new_run]);
}
}
+39
View File
@@ -328,6 +328,45 @@ impl Db {
Ok(rows.iter().map(|r| r.get::<i64, _>("article_id")).collect()) Ok(rows.iter().map(|r| r.get::<i64, _>("article_id")).collect())
} }
/// Article ids published before this issue date; same-date regeneration is allowed (§8.1).
pub async fn previously_published_ids_before(&self, date: Date) -> Result<Vec<ArticleId>> {
let rows =
sqlx::query("SELECT DISTINCT article_id FROM issue_articles WHERE issue_date < ?")
.bind(date.to_string())
.fetch_all(&self.pool)
.await?;
Ok(rows
.iter()
.map(|row| row.get::<i64, _>("article_id"))
.collect())
}
/// Published article ids first seen at or after `since` (`features backfill`).
pub async fn published_article_ids_since(&self, since: Timestamp) -> Result<Vec<ArticleId>> {
let rows = sqlx::query(
"SELECT DISTINCT ia.article_id FROM issue_articles ia
JOIN articles a ON a.id = ia.article_id
WHERE a.first_seen >= ?
ORDER BY ia.article_id",
)
.bind(fmt_ts(since))
.fetch_all(&self.pool)
.await?;
Ok(rows
.iter()
.map(|row| row.get::<i64, _>("article_id"))
.collect())
}
/// Every article id first seen at or after `since` (`features backfill --all`).
pub async fn article_ids_since(&self, since: Timestamp) -> Result<Vec<ArticleId>> {
let rows = sqlx::query("SELECT id FROM articles WHERE first_seen >= ? ORDER BY id")
.bind(fmt_ts(since))
.fetch_all(&self.pool)
.await?;
Ok(rows.iter().map(|row| row.get::<i64, _>("id")).collect())
}
/// Articles the LLM scored below `threshold` within the last `days` (§3.5). /// Articles the LLM scored below `threshold` within the last `days` (§3.5).
pub async fn recently_low_scored_ids( pub async fn recently_low_scored_ids(
&self, &self,
+308 -6
View File
@@ -3,6 +3,7 @@
//! Everything of substance lives in the library (`src/lib.rs`); this binary only //! Everything of substance lives in the library (`src/lib.rs`); this binary only
//! parses flags, loads config, opens the database and dispatches. //! parses flags, loads config, opens the database and dispatches.
use std::io::Write as _;
use std::path::PathBuf; use std::path::PathBuf;
use anyhow::{Context, Result}; use anyhow::{Context, Result};
@@ -10,6 +11,8 @@ use clap::{Parser, Subcommand, ValueEnum};
use tracing_subscriber::EnvFilter; use tracing_subscriber::EnvFilter;
use daily_epub::config::Config; use daily_epub::config::Config;
use daily_epub::curate::embedding::{self, BACKFILL_CONFIRM_TOKENS};
use daily_epub::curate::telemetry;
use daily_epub::db::Db; use daily_epub::db::Db;
use daily_epub::pipeline::{self, GenerateOptions, GenerateOutcome}; use daily_epub::pipeline::{self, GenerateOptions, GenerateOutcome};
use daily_epub::report::RunReport; use daily_epub::report::RunReport;
@@ -40,6 +43,11 @@ enum Command {
/// Inspect and edit explicit article verdicts. /// Inspect and edit explicit article verdicts.
#[command(subcommand)] #[command(subcommand)]
Ratings(RatingsCommand), Ratings(RatingsCommand),
/// Why an article was (not) in the paper, from persisted run telemetry.
Explain(ExplainArgs),
/// Embedding cache and telemetry maintenance.
#[command(subcommand)]
Features(FeaturesCommand),
/// Re-poll social scores for recent entries. /// Re-poll social scores for recent entries.
BackfillSocial(BackfillSocialArgs), BackfillSocial(BackfillSocialArgs),
/// Database maintenance. /// Database maintenance.
@@ -64,6 +72,9 @@ struct GenerateArgs {
/// Skip every LLM call: prefilter order selects, excerpts stand in for summaries. /// Skip every LLM call: prefilter order selects, excerpts stand in for summaries.
#[arg(long)] #[arg(long)]
skip_llm: bool, skip_llm: bool,
/// Use cached embeddings only: zero Voyage calls.
#[arg(long)]
skip_embeddings: bool,
} }
impl From<&GenerateArgs> for GenerateOptions { impl From<&GenerateArgs> for GenerateOptions {
@@ -74,6 +85,7 @@ impl From<&GenerateArgs> for GenerateOptions {
out: args.out.clone(), out: args.out.clone(),
max_articles: args.max_articles, max_articles: args.max_articles,
skip_llm: args.skip_llm, skip_llm: args.skip_llm,
skip_embeddings: args.skip_embeddings,
} }
} }
} }
@@ -158,6 +170,55 @@ struct RatingsClearArgs {
url: Option<String>, url: Option<String>,
} }
/// `explain --date D (--article ID | --url URL) [--run-id N]` or
/// `explain --date D --near-misses [N]` (plan §15.2).
#[derive(Debug, clap::Args)]
struct ExplainArgs {
/// Issue date whose run to read.
#[arg(long, value_name = "YYYY-MM-DD")]
date: String,
/// Article id, as printed by `ratings list` or `explain --near-misses`.
#[arg(
long,
required_unless_present_any = ["url", "near_misses"],
conflicts_with_all = ["url", "near_misses"]
)]
article: Option<ArticleId>,
/// Article URL; canonicalized before lookup.
#[arg(long, conflicts_with = "near_misses")]
url: Option<String>,
/// A specific run of that date instead of the latest non-dry one.
#[arg(long, value_name = "N")]
run_id: Option<i64>,
/// The top N articles that were considered but not selected (default 10).
#[arg(long, value_name = "N", num_args = 0..=1, default_missing_value = "10")]
near_misses: Option<usize>,
}
#[derive(Debug, Subcommand)]
enum FeaturesCommand {
/// Embed rated and published articles, then interests, into the cache.
Backfill(BackfillArgs),
/// Drop stale embeddings and old candidate telemetry per the retention config.
Prune,
}
#[derive(Debug, clap::Args)]
struct BackfillArgs {
/// Window for published (and, with --all, other) articles.
#[arg(long, default_value_t = 30)]
days: i64,
/// Only the rated set.
#[arg(long, conflicts_with = "all")]
rated_only: bool,
/// Also every other article first seen inside the window.
#[arg(long)]
all: bool,
/// Skip the confirmation prompt above the token threshold.
#[arg(long)]
yes: bool,
}
#[derive(Debug, clap::Args)] #[derive(Debug, clap::Args)]
struct BackfillSocialArgs { struct BackfillSocialArgs {
/// How many days back to re-poll. /// How many days back to re-poll.
@@ -196,6 +257,14 @@ async fn main() -> Result<()> {
let db = Db::open_and_migrate(&config.database_path).await?; let db = Db::open_and_migrate(&config.database_path).await?;
cmd_ratings(&config, &db, command).await?; cmd_ratings(&config, &db, command).await?;
} }
Command::Explain(args) => {
let db = Db::open_and_migrate(&config.database_path).await?;
cmd_explain(&db, args).await?;
}
Command::Features(command) => {
let db = Db::open_and_migrate(&config.database_path).await?;
cmd_features(&config, &db, command).await?;
}
Command::BackfillSocial(args) => { Command::BackfillSocial(args) => {
let db = Db::open_and_migrate(&config.database_path).await?; let db = Db::open_and_migrate(&config.database_path).await?;
cmd_backfill_social(&db, args.days).await?; cmd_backfill_social(&db, args.days).await?;
@@ -270,18 +339,27 @@ fn print_report(report: &RunReport) {
report.counts.duplicates_merged, report.counts.duplicates_merged,
report.counts.entries_dropped, report.counts.entries_dropped,
); );
if report.counts.llm_unscored > 0 { let unscored = if report.counts.llm_unscored > 0 {
format!(" ({} unscored)", report.counts.llm_unscored)
} else {
String::new()
};
println!( println!(
"curation: {} scored · {} unscored · {} selected", "curation: {} eligible · {} embedded · {} rated w/ embeddings → {} candidates → {} scored{unscored} → {} selected",
report.counts.llm_scored, report.counts.llm_unscored, report.counts.selected, report.counts.eligible,
report.counts.embedded,
report.counts.rated_with_embeddings,
report.counts.candidates,
report.counts.llm_scored,
report.counts.selected,
); );
}
println!( println!(
"tokens: {} input · {} cache read · {} cache write · {} output = ${:.4}", "tokens: {} input · {} cache read · {} cache write · {} output · {} voyage = ${:.4}",
report.usage.input_tokens, report.usage.input_tokens,
report.usage.cached_tokens, report.usage.cached_tokens,
report.usage.cache_write_tokens, report.usage.cache_write_tokens,
report.usage.output_tokens, report.usage.output_tokens,
report.voyage_tokens,
report.cost_usd, report.cost_usd,
); );
for (provider, usage) in &report.provider_costs { for (provider, usage) in &report.provider_costs {
@@ -294,6 +372,10 @@ fn print_report(report: &RunReport) {
usage.cost_usd, usage.cost_usd,
); );
} }
println!(
" voyage: {} tokens = ${:.4}",
report.voyage_tokens, report.voyage_cost_usd
);
for warning in &report.warnings { for warning in &report.warnings {
println!("warning: {warning}"); println!("warning: {warning}");
} }
@@ -478,6 +560,101 @@ async fn cmd_ratings(config: &Config, db: &Db, command: RatingsCommand) -> Resul
Ok(()) Ok(())
} }
async fn cmd_explain(db: &Db, args: ExplainArgs) -> Result<()> {
let date: jiff::civil::Date = args
.date
.parse()
.with_context(|| format!("invalid --date {:?}, expected YYYY-MM-DD", args.date))?;
let text = if let Some(limit) = args.near_misses {
telemetry::explain_near_misses(db, date, args.run_id, limit).await?
} else {
let target = match (args.article, args.url) {
(Some(id), _) => telemetry::ExplainTarget::Article(id),
(None, Some(url)) => telemetry::ExplainTarget::Url(url),
(None, None) => anyhow::bail!("provide --article, --url or --near-misses"),
};
telemetry::explain(db, date, args.run_id, &target).await?
};
print!("{text}");
Ok(())
}
async fn cmd_features(config: &Config, db: &Db, command: FeaturesCommand) -> Result<()> {
match command {
FeaturesCommand::Backfill(args) => {
if !config.voyage.enabled {
anyhow::bail!("voyage.enabled is false; nothing to backfill");
}
let service = embedding::EmbeddingService::real(db.clone(), config.voyage.clone())
.context("building the Voyage client")?;
let opts = embedding::BackfillOptions {
days: args.days,
rated_only: args.rated_only,
all: args.all,
};
let plan = embedding::plan_backfill(db, config, &service, &opts).await?;
println!(
"backfill: {} articles ({} learned, {} other) + {} interests to embed, {} already cached",
plan.article_count(),
plan.learned.len(),
plan.others.len(),
plan.interests.len(),
plan.cached
);
if plan.is_empty() {
println!("cache is warm; nothing to do");
return Ok(());
}
println!(
"estimate: ~{} tokens ≈ ${:.4} with {} at ${:.2}/M",
plan.estimated_tokens,
plan.estimated_cost_usd(),
config.voyage.model,
embedding::VOYAGE_PRICE_PER_MTOK
);
if plan.estimated_tokens > BACKFILL_CONFIRM_TOKENS
&& !args.yes
&& !confirm("continue?")?
{
println!("aborted");
return Ok(());
}
let outcome = embedding::run_backfill(&service, &plan).await?;
println!(
"embedded {} articles and {} interests · {} tokens · ${:.4}",
outcome.articles_embedded,
outcome.interests_embedded,
outcome.tokens,
outcome.cost_usd
);
}
FeaturesCommand::Prune => {
let ranking = &config.curation.ranking;
let (embeddings, rows) = telemetry::prune(
db,
ranking.embedding_retention_days,
ranking.telemetry_retention_days,
jiff::Timestamp::now(),
)
.await?;
println!(
"pruned {embeddings} embeddings older than {} days and {rows} candidate rows older than {} days",
ranking.embedding_retention_days, ranking.telemetry_retention_days
);
}
}
Ok(())
}
/// A y/N question on stdin; anything but a leading `y` is a no.
fn confirm(question: &str) -> Result<bool> {
print!("{question} [y/N] ");
std::io::stdout().flush()?;
let mut answer = String::new();
std::io::stdin().read_line(&mut answer)?;
Ok(answer.trim().to_lowercase().starts_with('y'))
}
async fn cmd_backfill_social(db: &Db, days: u32) -> Result<()> { async fn cmd_backfill_social(db: &Db, days: u32) -> Result<()> {
let http = http::build_client(http::DEFAULT_TIMEOUT)?; let http = http::build_client(http::DEFAULT_TIMEOUT)?;
let enricher = social::SocialEnricher::new(http, db.clone()); let enricher = social::SocialEnricher::new(http, db.clone());
@@ -509,6 +686,7 @@ mod tests {
"--max-articles", "--max-articles",
"6", "6",
"--skip-llm", "--skip-llm",
"--skip-embeddings",
]) ])
.unwrap(); .unwrap();
match cli.command { match cli.command {
@@ -518,10 +696,11 @@ mod tests {
assert_eq!(a.out, Some(PathBuf::from("./out"))); assert_eq!(a.out, Some(PathBuf::from("./out")));
assert_eq!(a.max_articles, Some(6)); assert_eq!(a.max_articles, Some(6));
assert!(a.skip_llm); assert!(a.skip_llm);
assert!(a.skip_embeddings);
let opts = GenerateOptions::from(&a); let opts = GenerateOptions::from(&a);
assert_eq!(opts.date.as_deref(), Some("2026-08-15")); assert_eq!(opts.date.as_deref(), Some("2026-08-15"));
assert!(opts.dry_run && opts.skip_llm); assert!(opts.dry_run && opts.skip_llm && opts.skip_embeddings);
assert_eq!(opts.max_articles, Some(6)); assert_eq!(opts.max_articles, Some(6));
} }
other => panic!("expected generate, got {other:?}"), other => panic!("expected generate, got {other:?}"),
@@ -583,6 +762,129 @@ mod tests {
assert_eq!(cli.config, Some(PathBuf::from("/tmp/x.toml"))); assert_eq!(cli.config, Some(PathBuf::from("/tmp/x.toml")));
} }
#[test]
fn parses_explain_and_features() {
match Cli::try_parse_from([
"daily-epub",
"explain",
"--date",
"2026-09-02",
"--article",
"42",
"--run-id",
"7",
])
.unwrap()
.command
{
Command::Explain(args) => {
assert_eq!(args.date, "2026-09-02");
assert_eq!(args.article, Some(42));
assert_eq!(args.run_id, Some(7));
assert_eq!(args.near_misses, None);
}
other => panic!("expected explain, got {other:?}"),
}
match Cli::try_parse_from([
"daily-epub",
"explain",
"--date",
"2026-09-02",
"--url",
"https://example.com/post",
])
.unwrap()
.command
{
Command::Explain(args) => {
assert_eq!(args.url.as_deref(), Some("https://example.com/post"))
}
other => panic!("expected explain, got {other:?}"),
}
match Cli::try_parse_from([
"daily-epub",
"explain",
"--date",
"2026-09-02",
"--near-misses",
])
.unwrap()
.command
{
Command::Explain(args) => assert_eq!(args.near_misses, Some(10)),
other => panic!("expected explain, got {other:?}"),
}
match Cli::try_parse_from([
"daily-epub",
"explain",
"--date",
"2026-09-02",
"--near-misses",
"3",
])
.unwrap()
.command
{
Command::Explain(args) => assert_eq!(args.near_misses, Some(3)),
other => panic!("expected explain, got {other:?}"),
}
assert!(Cli::try_parse_from(["daily-epub", "explain", "--date", "2026-09-02"]).is_err());
assert!(
Cli::try_parse_from([
"daily-epub",
"explain",
"--date",
"2026-09-02",
"--article",
"1",
"--near-misses"
])
.is_err()
);
match Cli::try_parse_from([
"daily-epub",
"features",
"backfill",
"--days",
"60",
"--all",
"--yes",
])
.unwrap()
.command
{
Command::Features(FeaturesCommand::Backfill(args)) => {
assert_eq!(args.days, 60);
assert!(args.all && args.yes && !args.rated_only);
}
other => panic!("expected features backfill, got {other:?}"),
}
match Cli::try_parse_from(["daily-epub", "features", "backfill"])
.unwrap()
.command
{
Command::Features(FeaturesCommand::Backfill(args)) => assert_eq!(args.days, 30),
other => panic!("expected features backfill, got {other:?}"),
}
assert!(
Cli::try_parse_from([
"daily-epub",
"features",
"backfill",
"--rated-only",
"--all"
])
.is_err()
);
assert!(matches!(
Cli::try_parse_from(["daily-epub", "features", "prune"])
.unwrap()
.command,
Command::Features(FeaturesCommand::Prune)
));
}
#[tokio::test] #[tokio::test]
async fn cli_set_and_clear_append_cli_events_with_latest_issue_date() { async fn cli_set_and_clear_append_cli_events_with_latest_issue_date() {
use sqlx::Row as _; use sqlx::Row as _;
+696 -10
View File
@@ -23,7 +23,7 @@
//! issue itself are upserted, `issue_articles` is replaced wholesale, and the //! issue itself are upserted, `issue_articles` is replaced wholesale, and the
//! published filenames are derived from the date. //! published filenames are derived from the date.
use std::collections::{BTreeMap, BTreeSet}; use std::collections::{BTreeMap, BTreeSet, HashMap, HashSet};
use std::path::PathBuf; use std::path::PathBuf;
use anyhow::{Context, Result}; use anyhow::{Context, Result};
@@ -32,14 +32,15 @@ use jiff::{Timestamp, Zoned};
use crate::config::Config; use crate::config::Config;
use crate::curate::llm::{Llms, PriceTable, UsageMeter}; use crate::curate::llm::{Llms, PriceTable, UsageMeter};
use crate::curate::{Curator, editorial, profile}; use crate::curate::{Curator, editorial, embedding, prefilter, profile, signals, telemetry};
use crate::db::Db; use crate::db::Db;
use crate::extract::Extractor; use crate::extract::Extractor;
use crate::miniflux::MinifluxClient; use crate::miniflux::MinifluxClient;
use crate::publish::Published; use crate::publish::Published;
use crate::report::{ProviderUsage, RunReport, RunStatus}; use crate::report::{ProviderUsage, RunReport, RunStatus};
use crate::types::{ use crate::types::{
Article, Artifact, Colophon, Edition, Issue, IssueMeta, Lineup, Models, reading_minutes, Article, ArticleId, Artifact, Colophon, Edition, Issue, IssueMeta, Lineup, Models,
reading_minutes,
}; };
use crate::{comments, dedupe, epub, http, miniflux, publish, social, world}; use crate::{comments, dedupe, epub, http, miniflux, publish, social, world};
@@ -56,6 +57,8 @@ pub struct GenerateOptions {
pub max_articles: Option<usize>, pub max_articles: Option<usize>,
/// `--skip-llm`: no DeepSeek call at all. /// `--skip-llm`: no DeepSeek call at all.
pub skip_llm: bool, pub skip_llm: bool,
/// `--skip-embeddings`: read the cache but make zero Voyage calls.
pub skip_embeddings: bool,
} }
/// What one run produced, for the caller to print (§3.13). /// What one run produced, for the caller to print (§3.13).
@@ -196,10 +199,12 @@ pub async fn generate(config: &Config, db: &Db, opts: &GenerateOptions) -> Resul
soft_target, soft_target,
hard_max, hard_max,
skip_llm = opts.skip_llm, skip_llm = opts.skip_llm,
skip_embeddings = opts.skip_embeddings,
voyage_enabled = config.voyage.enabled,
out = %out_dir.display(), out = %out_dir.display(),
"starting run" "starting run"
); );
log_resolved_providers(config, opts.skip_llm); log_resolved_providers(config, opts.skip_llm, opts.skip_embeddings);
let run_id = db.start_run(date, started_at).await?; let run_id = db.start_run(date, started_at).await?;
let mut report = RunReport::new(date, started_at); let mut report = RunReport::new(date, started_at);
@@ -213,6 +218,7 @@ pub async fn generate(config: &Config, db: &Db, opts: &GenerateOptions) -> Resul
let ctx = StageContext { let ctx = StageContext {
config, config,
db, db,
run_id,
date, date,
soft_target, soft_target,
hard_max, hard_max,
@@ -220,6 +226,7 @@ pub async fn generate(config: &Config, db: &Db, opts: &GenerateOptions) -> Resul
out_dir, out_dir,
dry_run: opts.dry_run, dry_run: opts.dry_run,
skip_llm: opts.skip_llm, skip_llm: opts.skip_llm,
skip_embeddings: opts.skip_embeddings,
}; };
let stages = match run_stages(&ctx, window_start, window_end, &mut report).await { let stages = match run_stages(&ctx, window_start, window_end, &mut report).await {
Ok(stages) => { Ok(stages) => {
@@ -275,6 +282,7 @@ struct StageOutput {
struct StageContext<'a> { struct StageContext<'a> {
config: &'a Config, config: &'a Config,
db: &'a Db, db: &'a Db,
run_id: i64,
date: Date, date: Date,
soft_target: usize, soft_target: usize,
hard_max: usize, hard_max: usize,
@@ -282,6 +290,7 @@ struct StageContext<'a> {
out_dir: PathBuf, out_dir: PathBuf,
dry_run: bool, dry_run: bool,
skip_llm: bool, skip_llm: bool,
skip_embeddings: bool,
} }
async fn run_stages( async fn run_stages(
@@ -369,7 +378,11 @@ async fn run_stages(
report.counts.social_hits = enricher.enrich_all(&mut articles).await as i64; report.counts.social_hits = enricher.enrich_all(&mut articles).await as i64;
report.timings.record("social", elapsed_ms(stage)); report.timings.record("social", elapsed_ms(stage));
// --- Stage 6: heuristic pre-filter (§3.5) --- // --- Stage 6: hygiene, embeddings, and cheap signals (§8.1, §9) ---
let embeddings = build_embedding_service(ctx, report);
let feature_signals = prepare_features(ctx, &articles, &embeddings, report).await;
// --- Stage 6b: the old heuristic pre-filter still gates in this step (§21) ---
let stage = Timestamp::now(); let stage = Timestamp::now();
let bulk_meter = let bulk_meter =
UsageMeter::with_prices(PriceTable::deepseek(&config.deepseek), config.max_daily_usd); UsageMeter::with_prices(PriceTable::deepseek(&config.deepseek), config.max_daily_usd);
@@ -399,6 +412,28 @@ async fn run_stages(
.await .await
.context("running the heuristic pre-filter")?; .context("running the heuristic pre-filter")?;
report.counts.candidates = candidates.len() as i64; report.counts.candidates = candidates.len() as i64;
let admitted = candidates
.iter()
.map(|candidate| candidate.article.id)
.collect::<Vec<_>>();
let admitted_set = admitted.iter().copied().collect::<HashSet<_>>();
let not_admitted = feature_signals
.keys()
.copied()
.filter(|id| !admitted_set.contains(id))
.collect::<Vec<_>>();
record_stage(
ctx,
&feature_signals,
&not_admitted,
"eligible",
Some("not_admitted"),
)
.await
.context("recording prefilter telemetry")?;
record_stage(ctx, &feature_signals, &admitted, "admitted", None)
.await
.context("recording prefilter telemetry")?;
report.timings.record("prefilter", elapsed_ms(stage)); report.timings.record("prefilter", elapsed_ms(stage));
// --- Stage 7: LLM scoring, then selection (§3.6 A + B) --- // --- Stage 7: LLM scoring, then selection (§3.6 A + B) ---
@@ -410,12 +445,54 @@ async fn run_stages(
} }
report.counts.llm_scored = candidates.iter().filter(|c| c.llm.is_some()).count() as i64; report.counts.llm_scored = candidates.iter().filter(|c| c.llm.is_some()).count() as i64;
report.counts.llm_unscored = report.counts.candidates - report.counts.llm_scored; report.counts.llm_unscored = report.counts.candidates - report.counts.llm_scored;
let assessed = candidates
.iter()
.filter(|candidate| candidate.llm.is_some())
.map(|candidate| candidate.article.id)
.collect::<Vec<_>>();
record_stage(ctx, &feature_signals, &assessed, "assessed", None)
.await
.context("recording assessment telemetry")?;
// Every prefilter survivor goes to the old selector, scored or not.
record_stage(ctx, &feature_signals, &admitted, "shortlisted", None)
.await
.context("recording shortlist telemetry")?;
let mut lineup = curator let mut lineup = curator
.select(candidates, date) .select(candidates, date)
.await .await
.context("selecting the lineup")?; .context("selecting the lineup")?;
report.counts.selected = lineup.picks.len() as i64; report.counts.selected = lineup.picks.len() as i64;
let selected = lineup
.picks
.iter()
.map(|pick| pick.article.id)
.collect::<Vec<_>>();
let selected_set = selected.iter().copied().collect::<HashSet<_>>();
let not_selected = admitted
.iter()
.copied()
.filter(|id| !selected_set.contains(id))
.collect::<Vec<_>>();
// The editor's one-line `why` (§13) lands in `candidate_runs.editor_why` so
// `explain` can quote it; heuristic picks leave it NULL.
let selected_with_why = lineup
.picks
.iter()
.map(|pick| (pick.article.id, pick.why.as_deref()))
.collect::<Vec<_>>();
record_stage_with_why(ctx, &feature_signals, &selected_with_why, "selected", None)
.await
.context("recording selection telemetry")?;
record_stage(
ctx,
&feature_signals,
&not_selected,
"shortlisted",
Some("not_selected"),
)
.await
.context("recording selection telemetry")?;
if lineup.picks.is_empty() { if lineup.picks.is_empty() {
report.warn("the lineup is empty — check the lookback window and pre-filter"); report.warn("the lineup is empty — check the lookback window and pre-filter");
} }
@@ -577,6 +654,253 @@ async fn run_stages(
}) })
} }
/// The cheap signals and hygiene outcome for one eligible article (§9).
#[derive(Debug, Clone)]
struct FeatureSignals {
signals: signals::Signals,
auto_include: bool,
}
/// The embedding cache with a Voyage client behind it, or cache-only under
/// `--skip-embeddings`, `voyage.enabled = false` or a missing key (§16, §17).
fn build_embedding_service(
ctx: &StageContext<'_>,
report: &mut RunReport,
) -> embedding::EmbeddingService {
let (db, voyage) = (ctx.db.clone(), ctx.config.voyage.clone());
if ctx.skip_embeddings {
tracing::info!("--skip-embeddings: using cached vectors only, no Voyage calls");
return embedding::EmbeddingService::cached_only(db, voyage);
}
if !voyage.enabled {
tracing::info!("voyage disabled: using cached embeddings only");
return embedding::EmbeddingService::cached_only(db, voyage);
}
match embedding::EmbeddingService::real(db.clone(), voyage.clone()) {
Ok(service) => service,
Err(embedding::EmbeddingError::MissingApiKey) => {
tracing::warn!(
"voyage enabled but {} is unset; using cached embeddings only",
embedding::VOYAGE_API_KEY_ENV
);
embedding::EmbeddingService::cached_only(db, voyage)
}
Err(error) => {
report.warn(format!(
"Voyage unavailable; using cached embeddings only: {error}"
));
embedding::EmbeddingService::cached_only(db, voyage)
}
}
}
/// Hygiene, embeddings and cheap signals for every article (§8.1, §9).
///
/// Hygiene-excluded articles get thin `candidate_runs` rows; every other
/// article gets an `eligible` row with its `signals_json`. Nothing here can
/// fail the run: embeddings and the learned signals degrade to absent (§17).
async fn prepare_features(
ctx: &StageContext<'_>,
articles: &[Article],
service: &embedding::EmbeddingService,
report: &mut RunReport,
) -> HashMap<ArticleId, FeatureSignals> {
let (config, db) = (ctx.config, ctx.db);
let hygiene = match prefilter::PrefilterContext::load(db, ctx.date).await {
Ok(context) => context,
Err(error) => {
report.warn(format!(
"could not load hygiene history; signals skipped: {error}"
));
return HashMap::new();
}
};
let published = hygiene
.already_published
.iter()
.copied()
.collect::<HashSet<_>>();
let rejected = hygiene
.recently_rejected
.iter()
.copied()
.collect::<HashSet<_>>();
let mut eligible = Vec::new();
for article in articles {
let auto_include = prefilter::is_auto_include(article, &config.curation);
let reason = if published.contains(&article.id) {
Some("published_before")
} else if !auto_include && prefilter::is_blocked(article, &config.curation) {
Some("blocked")
} else if !auto_include && rejected.contains(&article.id) {
Some("recently_rejected")
} else {
None
};
match reason {
Some(reason) => {
if let Err(error) =
telemetry::thin_excluded(db, ctx.run_id, article.id, reason).await
{
report.warn(format!(
"could not record excluded candidate {}: {error}",
article.id
));
}
}
None => eligible.push(article.clone()),
}
}
report.counts.eligible = eligible.len() as i64;
// --- embed (§7.1, §7.2) ---
let stage = Timestamp::now();
let article_embeddings = match service.articles(&eligible).await {
Ok(embeddings) => embeddings,
Err(error) => {
report.warn(format!("article embedding stage degraded: {error}"));
HashMap::new()
}
};
report.counts.embedded = article_embeddings.len() as i64;
let interests =
match profile::load_standing_interests(&config.interests_opml, &config.profile_path) {
Ok(interests) => interests,
Err(error) => {
tracing::warn!(%error, "could not load standing interests for embeddings");
Vec::new()
}
};
let interest_embeddings = match service.interests(&interests).await {
Ok(embeddings) => embeddings,
Err(error) => {
report.warn(format!("interest embedding stage degraded: {error}"));
HashMap::new()
}
};
if let Some(meter) = service.meter() {
report.voyage_tokens = meter.total_tokens();
report.voyage_cost_usd = meter.cost_usd();
}
tracing::info!(
eligible = eligible.len(),
embedded = article_embeddings.len(),
interests = interest_embeddings.len(),
voyage_tokens = report.voyage_tokens,
"embeddings ready"
);
report.timings.record("embed", elapsed_ms(stage));
// --- signals (§9, §12.2, §12.4) ---
let stage = Timestamp::now();
let ranking = &config.curation.ranking;
let (mut computed, preference) = match signals::compute_all(
db,
&eligible,
&article_embeddings,
&interest_embeddings,
&config.voyage,
ranking,
Timestamp::now(),
)
.await
{
Ok(result) => result,
Err(error) => {
report.warn(format!("signal computation degraded: {error:#}"));
let state = signals::PreferenceState::default();
(
signals::compute(
&eligible,
&article_embeddings,
&interest_embeddings,
&state,
ranking,
),
state.summary(),
)
}
};
report.counts.rated_with_embeddings = preference.rated_with_embeddings as i64;
let mut output = HashMap::new();
for article in &eligible {
let auto_include = prefilter::is_auto_include(article, &config.curation);
let signals = computed
.remove(&article.id)
.unwrap_or_else(|| signals::Signals::baseline(article));
output.insert(
article.id,
FeatureSignals {
signals,
auto_include,
},
);
}
let eligible_ids = eligible
.iter()
.map(|article| article.id)
.collect::<Vec<_>>();
if let Err(error) = record_stage(ctx, &output, &eligible_ids, "eligible", None).await {
report.warn(format!("could not record eligible candidates: {error}"));
}
report.timings.record("signals", elapsed_ms(stage));
output
}
/// Upsert the `candidate_runs` row of every listed article at a new stage
/// (§7.4). Articles without signals (hygiene-excluded) are left alone.
async fn record_stage(
ctx: &StageContext<'_>,
features: &HashMap<ArticleId, FeatureSignals>,
ids: &[ArticleId],
stage: &str,
excluded_reason: Option<&str>,
) -> Result<()> {
let rows = ids.iter().map(|id| (*id, None)).collect::<Vec<_>>();
record_stage_with_why(ctx, features, &rows, stage, excluded_reason).await
}
/// [`record_stage`] with the editor's `why` per article (§13, §7.4).
async fn record_stage_with_why(
ctx: &StageContext<'_>,
features: &HashMap<ArticleId, FeatureSignals>,
rows: &[(ArticleId, Option<&str>)],
stage: &str,
excluded_reason: Option<&str>,
) -> Result<()> {
let admitted = matches!(stage, "admitted" | "assessed" | "shortlisted" | "selected");
for (id, editor_why) in rows {
let Some(feature) = features.get(id) else {
continue;
};
let json = telemetry::serialize_signals(&feature.signals, feature.auto_include);
let admitted_by = admitted.then_some(if feature.auto_include {
"[\"auto\"]"
} else {
"[\"prefilter\"]"
});
telemetry::write(
ctx.db,
&telemetry::CandidateRun {
run_id: ctx.run_id,
article_id: *id,
stage,
excluded_reason,
admitted_by,
signals_json: &json,
utility: None,
rank_utility: None,
cluster_id: None,
cluster_rank: None,
editor_why: *editor_why,
},
)
.await
.with_context(|| format!("recording candidate {id} at stage {stage}"))?;
}
Ok(())
}
/// Insert/refresh the `articles` rows and stamp the returned ids back on (§3.13). /// Insert/refresh the `articles` rows and stamp the returned ids back on (§3.13).
async fn persist_articles(db: &Db, articles: &mut [Article]) -> Result<()> { async fn persist_articles(db: &Db, articles: &mut [Article]) -> Result<()> {
for article in articles.iter_mut() { for article in articles.iter_mut() {
@@ -701,9 +1025,10 @@ pub fn issue_size_bounds(config: &Config, max_articles: Option<usize>) -> (usize
} }
/// Startup line naming the resolved models and whether each provider is on /// Startup line naming the resolved models and whether each provider is on
/// (§19): the root config ignores unknown sections, so an `[anthropics]` typo /// (§19): the root config ignores unknown sections, so an `[anthropics]` or
/// would otherwise be silent. Keys are never logged, only their presence. /// `[voyages]` typo would otherwise be silent. Keys are never logged, only
fn log_resolved_providers(config: &Config, skip_llm: bool) { /// their presence.
fn log_resolved_providers(config: &Config, skip_llm: bool, skip_embeddings: bool) {
let has_key = |key: Option<&str>| key.is_some_and(|k| !k.trim().is_empty()); let has_key = |key: Option<&str>| key.is_some_and(|k| !k.trim().is_empty());
tracing::info!( tracing::info!(
bulk_model = %config.deepseek.model, bulk_model = %config.deepseek.model,
@@ -716,6 +1041,12 @@ fn log_resolved_providers(config: &Config, skip_llm: bool) {
editor_effort = %config.anthropic.effort, editor_effort = %config.anthropic.effort,
editor_max_daily_usd = config.anthropic.max_daily_usd, editor_max_daily_usd = config.anthropic.max_daily_usd,
summary_model = ?config.editorial.summary_model, summary_model = ?config.editorial.summary_model,
embedding_model = %config.voyage.model,
embedding_enabled = !skip_embeddings
&& config.voyage.enabled
&& has_key(config.voyage.api_key.as_deref()),
embedding_dimension = config.voyage.output_dimension,
embedding_max_daily_usd = config.voyage.max_daily_usd,
"resolved providers" "resolved providers"
); );
} }
@@ -730,20 +1061,25 @@ const PROMPT_VERSIONS: &[(&str, u32)] = &[
("profile", 2), ("profile", 2),
]; ];
/// The resolved `[curation]`, `[editorial]`, model names and prompt versions /// The resolved `[curation]` (ranking included), `[editorial]`, `[voyage]`,
/// written to `runs.config_json` (§7.6, §19). Never includes keys. /// model names and prompt versions written to `runs.config_json` (§7.6, §19).
/// Never includes keys.
fn resolved_run_config(config: &Config, soft_target: usize, hard_max: usize) -> serde_json::Value { fn resolved_run_config(config: &Config, soft_target: usize, hard_max: usize) -> serde_json::Value {
let mut curation = config.curation.clone(); let mut curation = config.curation.clone();
curation.max_article_count = hard_max; curation.max_article_count = hard_max;
let mut voyage = config.voyage.clone();
voyage.api_key = None;
serde_json::json!({ serde_json::json!({
"target_article_count": soft_target, "target_article_count": soft_target,
"prefilter_keep": config.prefilter_keep, "prefilter_keep": config.prefilter_keep,
"curation": curation, "curation": curation,
"editorial": config.editorial, "editorial": config.editorial,
"voyage": voyage,
"models": { "models": {
"bulk": config.deepseek.model, "bulk": config.deepseek.model,
"editor": if config.anthropic.enabled { config.anthropic.model.as_str() } else { "disabled" }, "editor": if config.anthropic.enabled { config.anthropic.model.as_str() } else { "disabled" },
"editor_effort": config.anthropic.effort, "editor_effort": config.anthropic.effort,
"embedding": if config.voyage.enabled { config.voyage.model.as_str() } else { "disabled" },
}, },
"prompt_versions": PROMPT_VERSIONS "prompt_versions": PROMPT_VERSIONS
.iter() .iter()
@@ -815,9 +1151,19 @@ mod tests {
let mut config = Config::default(); let mut config = Config::default();
config.anthropic.api_key = Some("sk-secret".into()); config.anthropic.api_key = Some("sk-secret".into());
config.deepseek.api_key = Some("ds-secret".into()); config.deepseek.api_key = Some("ds-secret".into());
config.voyage.api_key = Some("pa-secret".into());
let value = resolved_run_config(&config, 6, 6); let value = resolved_run_config(&config, 6, 6);
assert_eq!(value["target_article_count"], 6); assert_eq!(value["target_article_count"], 6);
assert_eq!(value["curation"]["max_article_count"], 6); assert_eq!(value["curation"]["max_article_count"], 6);
assert_eq!(value["curation"]["ranking"]["deep_keep"], 120);
assert_eq!(
value["curation"]["ranking"]["weights"]["utility"]["quality"],
0.40
);
assert_eq!(value["voyage"]["model"], "voyage-4-lite");
assert_eq!(value["voyage"]["output_dimension"], 512);
assert!(value["voyage"]["api_key"].is_null());
assert_eq!(value["models"]["embedding"], "voyage-4-lite");
assert_eq!(value["editorial"]["summary_model"], "editor"); assert_eq!(value["editorial"]["summary_model"], "editor");
assert_eq!(value["editorial"]["summary_input_tokens"], 3000); assert_eq!(value["editorial"]["summary_input_tokens"], 3000);
assert_eq!(value["models"]["bulk"], "deepseek-v4-flash"); assert_eq!(value["models"]["bulk"], "deepseek-v4-flash");
@@ -865,4 +1211,344 @@ mod tests {
assert_eq!(lineup.picks[0].summary.as_deref(), Some("An abstract.")); assert_eq!(lineup.picks[0].summary.as_deref(), Some("An abstract."));
assert!(lineup.picks[1..].iter().all(|p| p.summary.is_none())); assert!(lineup.picks[1..].iter().all(|p| p.summary.is_none()));
} }
use std::sync::Arc;
use crate::curate::embedding::{EmbeddingClient, EmbeddingService, MockBackend};
use crate::types::{Entry, ExtractMethod, SourceKind, SourceRef};
use sqlx::Row as _;
fn now() -> Timestamp {
"2026-09-02T09:00:00Z".parse().unwrap()
}
fn run_date() -> Date {
"2026-09-02".parse().unwrap()
}
fn fixture_article(entry_id: i64, host: &str, words: usize) -> Article {
let url = format!("https://{host}/post-{entry_id}");
let body = (0..words)
.map(|i| format!("word{i}"))
.collect::<Vec<_>>()
.join(" ");
Article {
id: 0,
canonical_url: url.clone(),
title: format!("Post {entry_id}"),
best_entry_id: entry_id,
content_html: format!("<p>{body}</p>"),
word_count: words as i64,
excerpt_only: false,
image_count: 0,
sources: vec![SourceRef {
entry_id,
feed_id: 100 + entry_id,
feed_title: format!("Feed {entry_id}"),
category: None,
kind: SourceKind::Feed,
}],
first_seen: now(),
url,
author: None,
feed_id: 100 + entry_id,
feed_title: format!("Feed {entry_id}"),
category: None,
published_at: None,
comments_url: None,
image_urls: vec![],
social: vec![],
extract_method: ExtractMethod::Miniflux,
}
}
fn entry_for(article: &Article) -> Entry {
Entry {
id: article.best_entry_id,
feed_id: article.feed_id,
feed_title: Some(article.feed_title.clone()),
category: None,
title: article.title.clone(),
url: article.url.clone(),
canonical_url: Some(article.canonical_url.clone()),
author: None,
published_at: None,
comments_url: None,
raw_content: article.content_html.clone(),
fetched_at: now(),
}
}
struct Harness {
_dir: tempfile::TempDir,
db: Db,
config: Config,
articles: Vec<Article>,
run_id: i64,
}
/// Four articles: two ordinary, one on a blocked host, one published yesterday.
async fn harness() -> Harness {
let dir = tempfile::tempdir().unwrap();
let db = Db::open_and_migrate(&dir.path().join("run.db"))
.await
.unwrap();
let mut config = Config::default();
config.curation.blocked_domains = vec!["blocked.example".into()];
config.voyage.output_dimension = 4;
config.target_article_count = 1;
config.interests_opml = dir.path().join("interests.opml");
std::fs::write(
&config.interests_opml,
"<opml><body><outline text=\"Writerdeck\"/></body></opml>",
)
.unwrap();
config.profile_path = dir.path().join("profile.md");
std::fs::write(&config.profile_path, "# Reader profile\n").unwrap();
let mut articles = vec![
fixture_article(1, "a.example", 1200),
fixture_article(2, "b.example", 900),
fixture_article(3, "blocked.example", 1500),
fixture_article(4, "d.example", 1400),
];
let entries = articles.iter().map(entry_for).collect::<Vec<_>>();
db.upsert_entries(&entries).await.unwrap();
persist_articles(&db, &mut articles).await.unwrap();
sqlx::query(
"INSERT INTO issues (date, issue_number, generated_at)
VALUES ('2026-09-01', 1, '2026-09-01T12:00:00Z')",
)
.execute(db.pool())
.await
.unwrap();
sqlx::query(
"INSERT INTO issue_articles (issue_date, article_id, section)
VALUES ('2026-09-01', ?, 'Top Stories')",
)
.bind(articles[3].id)
.execute(db.pool())
.await
.unwrap();
let run_id = db.start_run(run_date(), now()).await.unwrap();
Harness {
_dir: dir,
db,
config,
articles,
run_id,
}
}
fn context<'a>(h: &'a Harness, skip_embeddings: bool) -> StageContext<'a> {
StageContext {
config: &h.config,
db: &h.db,
run_id: h.run_id,
date: run_date(),
soft_target: h.config.target_article_count,
hard_max: h.config.curation.max_article_count,
started_at: now(),
out_dir: PathBuf::from("."),
dry_run: true,
skip_llm: true,
skip_embeddings,
}
}
fn mock_service(h: &Harness, backend: Arc<MockBackend>) -> EmbeddingService {
let client = EmbeddingClient::with_backend(h.config.voyage.clone(), backend);
EmbeddingService::with_client(h.db.clone(), h.config.voyage.clone(), client)
}
async fn stage_rows(
db: &Db,
run_id: i64,
) -> BTreeMap<i64, (String, Option<String>, Option<String>)> {
sqlx::query(
"SELECT article_id, stage, excluded_reason, admitted_by FROM candidate_runs
WHERE run_id = ? ORDER BY article_id",
)
.bind(run_id)
.fetch_all(db.pool())
.await
.unwrap()
.iter()
.map(|row| {
(
row.get::<i64, _>("article_id"),
(
row.get::<String, _>("stage"),
row.get::<Option<String>, _>("excluded_reason"),
row.get::<Option<String>, _>("admitted_by"),
),
)
})
.collect()
}
#[tokio::test]
async fn mocked_run_writes_a_candidate_runs_row_for_every_considered_article() {
let h = harness().await;
let ctx = context(&h, false);
let backend = Arc::new(MockBackend::auto(4));
let service = mock_service(&h, backend.clone());
let mut report = RunReport::new(run_date(), now());
let features = prepare_features(&ctx, &h.articles, &service, &mut report).await;
let [a, b, blocked, published] = [
h.articles[0].id,
h.articles[1].id,
h.articles[2].id,
h.articles[3].id,
];
assert_eq!(
features.keys().copied().collect::<BTreeSet<_>>(),
BTreeSet::from([a, b])
);
assert_eq!(report.counts.eligible, 2);
assert_eq!(report.counts.embedded, 2);
assert_eq!(report.counts.rated_with_embeddings, 0);
assert!(report.timings.0.contains_key("embed") && report.timings.0.contains_key("signals"));
assert!(report.voyage_tokens > 0);
// One batch for the two articles, one for the interest.
assert_eq!(backend.calls(), 2);
let signals = &features[&a].signals;
assert!(signals.heuristic.is_some());
assert!(
signals.interest.is_some(),
"interest present under the raw fallback"
);
assert!(
signals.knn.is_none() && signals.feed.is_none(),
"gates closed"
);
assert!(signals.preliminary.is_some());
let rows = stage_rows(&h.db, h.run_id).await;
assert_eq!(rows.len(), 4, "one row per considered article");
assert_eq!(rows[&blocked].0, "excluded");
assert_eq!(rows[&blocked].1.as_deref(), Some("blocked"));
assert_eq!(rows[&published].0, "excluded");
assert_eq!(rows[&published].1.as_deref(), Some("published_before"));
assert_eq!(rows[&a].0, "eligible");
assert_eq!(rows[&a].1, None);
let thin: String =
sqlx::query_scalar("SELECT signals_json FROM candidate_runs WHERE article_id = ?")
.bind(blocked)
.fetch_one(h.db.pool())
.await
.unwrap();
assert_eq!(thin, "{}");
// The old prefilter and selector, with the stage transitions of step 3.
let curator = Curator::new(h.config.clone(), h.db.clone(), Llms::default());
let candidates = curator
.prefilter(h.articles.clone(), run_date())
.await
.unwrap();
let admitted = candidates.iter().map(|c| c.article.id).collect::<Vec<_>>();
assert_eq!(
admitted.iter().copied().collect::<BTreeSet<_>>(),
BTreeSet::from([a, b])
);
record_stage(&ctx, &features, &admitted, "admitted", None)
.await
.unwrap();
record_stage(&ctx, &features, &admitted, "shortlisted", None)
.await
.unwrap();
let lineup = curator.select(candidates, run_date()).await.unwrap();
let selected = lineup
.picks
.iter()
.map(|p| p.article.id)
.collect::<Vec<_>>();
assert_eq!(selected.len(), 1);
let not_selected = admitted
.iter()
.copied()
.filter(|id| !selected.contains(id))
.collect::<Vec<_>>();
record_stage(&ctx, &features, &selected, "selected", None)
.await
.unwrap();
record_stage(
&ctx,
&features,
&not_selected,
"shortlisted",
Some("not_selected"),
)
.await
.unwrap();
let rows = stage_rows(&h.db, h.run_id).await;
assert_eq!(rows.len(), 4);
let (winner, loser) = (selected[0], not_selected[0]);
assert_eq!(
rows[&winner],
("selected".into(), None, Some("[\"prefilter\"]".into()))
);
assert_eq!(
rows[&loser],
(
"shortlisted".into(),
Some("not_selected".into()),
Some("[\"prefilter\"]".into())
)
);
let text = telemetry::explain(
&h.db,
run_date(),
Some(h.run_id),
&telemetry::ExplainTarget::Article(loser),
)
.await
.unwrap();
assert!(
text.contains("stage: shortlisted · reason: not_selected"),
"{text}"
);
}
#[tokio::test]
async fn skip_embeddings_makes_zero_voyage_calls_and_uses_the_cache() {
let h = harness().await;
let ctx = context(&h, true);
let mut report = RunReport::new(run_date(), now());
let service = build_embedding_service(&ctx, &mut report);
assert!(!service.has_client(), "--skip-embeddings is cache-only");
assert!(service.meter().is_none());
let features = prepare_features(&ctx, &h.articles, &service, &mut report).await;
assert_eq!(features.len(), 2);
assert_eq!(report.counts.embedded, 0, "nothing cached yet");
assert!(features.values().all(|f| f.signals.interest.is_none()));
assert!(features.values().all(|f| f.signals.heuristic.is_some()));
assert_eq!(report.voyage_tokens, 0);
}
#[tokio::test]
async fn a_voyage_failure_degrades_to_absent_signals_and_the_run_continues() {
let h = harness().await;
let ctx = context(&h, false);
let backend = Arc::new(MockBackend::new()); // nothing scripted: every call fails
let service = mock_service(&h, backend.clone());
let mut report = RunReport::new(run_date(), now());
let features = prepare_features(&ctx, &h.articles, &service, &mut report).await;
assert!(backend.calls() >= 1);
assert_eq!(features.len(), 2);
assert_eq!(report.counts.eligible, 2);
assert_eq!(report.counts.embedded, 0);
assert!(report.error.is_none());
for feature in features.values() {
assert!(feature.signals.interest.is_none() && feature.signals.knn.is_none());
assert!(feature.signals.heuristic.is_some());
assert!(
feature.signals.preliminary.is_some(),
"scored on what is present"
);
}
assert_eq!(stage_rows(&h.db, h.run_id).await.len(), 4);
}
} }
+21 -5
View File
@@ -67,6 +67,12 @@ pub struct StageCounts {
pub excerpt_only: i64, pub excerpt_only: i64,
/// Social lookups that returned a hit (§3.4). /// Social lookups that returned a hit (§3.4).
pub social_hits: i64, pub social_hits: i64,
/// Articles passing hygiene and eligible for personalized signals.
pub eligible: i64,
/// Eligible articles with a valid embedding.
pub embedded: i64,
/// Current rated articles with a valid embedding.
pub rated_with_embeddings: i64,
/// Articles surviving the heuristic pre-filter (§3.5). /// Articles surviving the heuristic pre-filter (§3.5).
pub candidates: i64, pub candidates: i64,
/// Articles scored by the LLM (§3.6 stage A). /// Articles scored by the LLM (§3.6 stage A).
@@ -113,10 +119,13 @@ pub struct RunReport {
pub counts: StageCounts, pub counts: StageCounts,
/// Aggregate usage retained for the legacy `runs` columns. /// Aggregate usage retained for the legacy `runs` columns.
pub usage: TokenUsage, pub usage: TokenUsage,
/// Provider-keyed usage and cost written to `runs.provider_costs_json`. /// LLM provider-keyed usage and cost written to `runs.provider_costs_json`.
pub provider_costs: BTreeMap<String, ProviderUsage>, pub provider_costs: BTreeMap<String, ProviderUsage>,
/// Resolved curation/editorial/model settings for this run. /// Resolved curation/editorial/model settings for this run.
pub config_json: serde_json::Value, pub config_json: serde_json::Value,
/// Voyage document/query tokens and cost for this run.
pub voyage_tokens: i64,
pub voyage_cost_usd: f64,
pub cost_usd: f64, pub cost_usd: f64,
pub timings: StageTimings, pub timings: StageTimings,
/// Ingest window actually used, RFC3339 (§3.1). /// Ingest window actually used, RFC3339 (§3.1).
@@ -141,6 +150,8 @@ impl RunReport {
usage: TokenUsage::default(), usage: TokenUsage::default(),
provider_costs: BTreeMap::new(), provider_costs: BTreeMap::new(),
config_json: serde_json::Value::Null, config_json: serde_json::Value::Null,
voyage_tokens: 0,
voyage_cost_usd: 0.0,
cost_usd: 0.0, cost_usd: 0.0,
timings: StageTimings::default(), timings: StageTimings::default(),
window_start: None, window_start: None,
@@ -163,11 +174,12 @@ impl RunReport {
self.error = Some(err.to_string()); self.error = Some(err.to_string());
} }
/// Stamp the end time, total provider costs and settle the status. /// Stamp the end time, total provider costs (LLM providers plus Voyage) and
/// settle the status.
pub fn finish(&mut self, finished_at: Timestamp) { pub fn finish(&mut self, finished_at: Timestamp) {
self.finished_at = Some(finished_at); self.finished_at = Some(finished_at);
self.usage = TokenUsage::default(); self.usage = TokenUsage::default();
self.cost_usd = 0.0; self.cost_usd = self.voyage_cost_usd;
for provider in self.provider_costs.values() { for provider in self.provider_costs.values() {
self.usage.add(provider.usage); self.usage.add(provider.usage);
self.cost_usd += provider.cost_usd; self.cost_usd += provider.cost_usd;
@@ -198,11 +210,12 @@ impl RunReport {
/// Compact human-readable summary printed at the end of `generate`. /// Compact human-readable summary printed at the end of `generate`.
pub fn summary_line(&self) -> String { pub fn summary_line(&self) -> String {
format!( format!(
"{} [{}] {} entries → {} articles → {} candidates → {} selected · ${:.4} · {}s", "{} [{}] {} entries → {} articles → {} eligible → {} candidates → {} selected · ${:.4} · {}s",
self.date, self.date,
self.status, self.status,
self.counts.entries_fetched, self.counts.entries_fetched,
self.counts.articles, self.counts.articles,
self.counts.eligible,
self.counts.candidates, self.counts.candidates,
self.counts.selected, self.counts.selected,
self.cost_usd, self.cost_usd,
@@ -257,9 +270,12 @@ mod tests {
cost_usd: 0.05, cost_usd: 0.05,
}, },
); );
r.voyage_tokens = 250_000;
r.voyage_cost_usd = 0.005;
r.finish(ts("2026-08-15T05:36:00Z")); r.finish(ts("2026-08-15T05:36:00Z"));
assert_eq!(r.status, RunStatus::Ok); assert_eq!(r.status, RunStatus::Ok);
assert!((r.cost_usd - 0.4728).abs() < 1e-9); // LLM providers plus Voyage; Voyage tokens stay out of the LLM aggregate.
assert!((r.cost_usd - 0.4778).abs() < 1e-9);
// The legacy aggregate columns are the sum across providers. // The legacy aggregate columns are the sum across providers.
assert_eq!(r.usage, usage(1_000_100, 1_003_000, 2_000, 1_000_800)); assert_eq!(r.usage, usage(1_000_100, 1_003_000, 2_000, 1_000_800));
assert_eq!(r.duration_secs(), Some(360)); assert_eq!(r.duration_secs(), Some(360));