Curation v2: session 2 handoff (steps 4-7 done, briefs preamble updated)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A1rCLQeKBgnBo3oTgHuTMe
This commit is contained in:
@@ -1,6 +1,6 @@
|
|||||||
# Personalized Curation v2 — implementation progress and handoff
|
# Personalized Curation v2 — implementation progress and handoff
|
||||||
|
|
||||||
**Updated:** 2026-09-02 (end of session 1)
|
**Updated:** 2026-09-02 (end of session 2)
|
||||||
**Plan:** `docs/plans/2026-09-02-personalized-curation-v2.md` (§21 is the step sequence)
|
**Plan:** `docs/plans/2026-09-02-personalized-curation-v2.md` (§21 is the step sequence)
|
||||||
**Branch:** `curation-v2` (branched from `main` at `a599745`; not merged, not pushed)
|
**Branch:** `curation-v2` (branched from `main` at `a599745`; not merged, not pushed)
|
||||||
**Briefs:** `docs/plans/curation-v2-briefs/` — `00-preamble.md` + one `stepN.md` per step. Each
|
**Briefs:** `docs/plans/curation-v2-briefs/` — `00-preamble.md` + one `stepN.md` per step. Each
|
||||||
@@ -13,43 +13,81 @@ brief was handed to an implementation agent as `cat 00-preamble.md stepN.md`.
|
|||||||
| 1. Feedback and profile | **done**, reviewed (codex review: no actionable issues) | `3a9f4b9` |
|
| 1. Feedback and profile | **done**, reviewed (codex review: no actionable issues) | `3a9f4b9` |
|
||||||
| 2. Claude editor and editorial | **done** | `57efbb4` |
|
| 2. Claude editor and editorial | **done** | `57efbb4` |
|
||||||
| 3. Embeddings, signals, telemetry | **done** (branch `curation-v2-step3`, merged) | `ea3b141` + merge commit |
|
| 3. Embeddings, signals, telemetry | **done** (branch `curation-v2-step3`, merged) | `ea3b141` + merge commit |
|
||||||
| 4. Triage replaces the gate | not started — brief ready: `step4.md` | |
|
| 4. Triage replaces the gate | **done** (Codex plugin, high effort), reviewed | `10f4afd` |
|
||||||
| 5. Deep assessment, utility, diversity | not started — brief ready: `step5.md` | |
|
| 5. Deep assessment, utility, diversity | **done** (Codex started, cut off by quota; Claude agent finished), reviewed | `05a74a0` |
|
||||||
| 6. Paper telemetry, stats, lock | not started — brief ready: `step6.md` | |
|
| 6. Paper telemetry, stats, lock | **done** (Claude agent), reviewed | `d261cd4` |
|
||||||
| 7. Cleanup + implementation notes | not started — brief ready: `step7.md` | |
|
| 7. Cleanup + implementation notes | **done** (Claude agent), reviewed | `d403c51` |
|
||||||
|
|
||||||
`cargo fmt --check`, `cargo clippy --all-targets` (only the three pre-existing `src/world.rs`
|
All seven plan steps are implemented. `cargo fmt --check`, `cargo clippy --all-targets`
|
||||||
`needless_borrow` warnings) and `cargo test` (294 lib tests + 6 integration suites) are green at HEAD.
|
(including `-W dead_code`) and `cargo test` (329 lib tests + 6 bin tests + the 7 integration suites) are green at HEAD.
|
||||||
|
|
||||||
Nothing has been run against the production database or the real providers. No local
|
Nothing has been run against the production database or the real providers. No local
|
||||||
`config.toml`, database, or API keys exist on the dev box, so verification so far is tests only.
|
`config.toml`, database, or API keys exist on the dev box, so verification so far is tests only.
|
||||||
|
|
||||||
## Decisions and deviations made while implementing (read before step 4)
|
## Decisions and deviations made while implementing
|
||||||
|
|
||||||
|
Steps 1–3 (session 1):
|
||||||
|
|
||||||
- **`scores` table not yet dropped.** Migration `0002_curation_v2.sql` created every new table, copied
|
|
||||||
`ratings` → `rating_events`, and dropped `ratings` and `feed_priors`, but kept `scores` because the
|
|
||||||
old Stage A scoring (`src/curate/score.rs`, `db::upsert_score`, `recently_low_scored_ids`) still
|
|
||||||
uses it. Step 4 adds `migrations/0003_drop_scores.sql` and moves the churn rule to
|
|
||||||
`article_assessments` (already in `step4.md`).
|
|
||||||
- **Old prefilter still gates.** Hygiene → embeddings → signals → preliminary blend now run for every
|
|
||||||
article, and `candidate_runs` rows are written with the stage vocabulary mapped onto the old flow
|
|
||||||
(`admitted_by` is `["prefilter"]`/`["auto"]` for now), but `prefilter::run` + `prefilter_keep`
|
|
||||||
still decide the deep set until step 4.
|
|
||||||
- **Legacy `up` links.** `Vote::parse("up")` → `Loved`, and `auth::verify_token` also accepts tokens
|
- **Legacy `up` links.** `Vote::parse("up")` → `Loved`, and `auth::verify_token` also accepts tokens
|
||||||
signed over the literal `up` segment so already-published issues keep working.
|
signed over the literal `up` segment so already-published issues keep working. Keep both.
|
||||||
- **Same-date regeneration** no longer excludes its own picks (`published_before` uses issue dates
|
- **Same-date regeneration** no longer excludes its own picks (`published_before` uses issue dates
|
||||||
strictly before the run date), per plan §8.1.
|
strictly before the run date), per plan §8.1.
|
||||||
- **`VoyageConfig.api_key`** exists as a field (figment maps `DAILY_EPUB_VOYAGE__API_KEY` into it;
|
- **`VoyageConfig.api_key`** exists as a field (figment maps `DAILY_EPUB_VOYAGE__API_KEY` into it;
|
||||||
`deny_unknown_fields` would otherwise reject the env var). Never document it in TOML.
|
`deny_unknown_fields` would otherwise reject the env var). Never document it in TOML.
|
||||||
- **`anthropic.max_concurrent_requests`** is validated but not consumed yet; summary concurrency is
|
- **`anthropic.max_concurrent_requests`** is validated but not consumed; summary concurrency is
|
||||||
the constant `SUMMARY_CONCURRENCY = 4` in `editorial.rs`.
|
the constant `SUMMARY_CONCURRENCY = 4` in `editorial.rs`.
|
||||||
- `UsageMeter::new(&DeepseekConfig, ..)` survives as a compat constructor over `with_prices`.
|
|
||||||
- `tests/fixtures/deepseek_front_page.json` was replaced by `tests/fixtures/claude_brief.json`.
|
|
||||||
- `RATINGS_LOOKBACK_DAYS` in `profile/mod.rs` is effectively unbounded (36,500) for the prompt
|
- `RATINGS_LOOKBACK_DAYS` in `profile/mod.rs` is effectively unbounded (36,500) for the prompt
|
||||||
verdict block and the weekly rebuild; the knn/feed preference state uses
|
verdict block and the weekly rebuild; the knn/feed preference state uses
|
||||||
`curation.ranking.rating_lookback_days` (180) as the plan says.
|
`curation.ranking.rating_lookback_days` (180) as the plan says.
|
||||||
- Footer CSS: `.rating` has no `white-space: nowrap` (it would clip on narrow e-ink screens).
|
- Footer CSS: `.rating` has no `white-space: nowrap` (it would clip on narrow e-ink screens).
|
||||||
|
|
||||||
|
Step 4:
|
||||||
|
|
||||||
|
- Hygiene order is blocked → published_before → recently_rejected, all skipped for auto-includes.
|
||||||
|
- `admitted_by` records every retriever whose own top-N would have taken the article (first entry
|
||||||
|
is the admitting one); exploration is flagged only when `exploration` admitted it.
|
||||||
|
- Under the pool cap (`triage_max`) the surplus gets `stage='eligible'`, `excluded_reason='not_admitted'`.
|
||||||
|
- Migration `0003_drop_scores.sql` drops `scores`; the churn rule reads `article_assessments`.
|
||||||
|
|
||||||
|
Step 5:
|
||||||
|
|
||||||
|
- `Candidate` is the only flow type; `ScoredArticle`, `combined_score`, `score.rs`, `select.rs` are gone.
|
||||||
|
- The editor omits the `facets:` line when every facet is unknown (plan renders it unconditionally).
|
||||||
|
- Articles cut at the shortlist keep the stage they reached (`admitted` if never deep-assessed,
|
||||||
|
else `assessed`) with `excluded_reason` `cluster_suppressed` (hit the base cap) or `shortlist_cap`.
|
||||||
|
- `Curator::assess` is a no-op under `--skip-llm`; cached deep rows are still reused when DeepSeek
|
||||||
|
is merely down.
|
||||||
|
- `--near-misses` and the Behind-the-paper list order by utility, falling back to the preliminary
|
||||||
|
blend per row.
|
||||||
|
- Fixtures: `tests/fixtures/deepseek_deep_batch{,_messy}.json` replace the old score fixtures;
|
||||||
|
`deepseek_triage_batch.json` added in step 4.
|
||||||
|
|
||||||
|
Step 6:
|
||||||
|
|
||||||
|
- `BehindThePaper` hangs off `Issue.behind` (filled after `build_issue`), not a `build_issue` argument.
|
||||||
|
Generation time and cost shown in the chapter are measured at issue assembly (they exclude the
|
||||||
|
EPUB build and publish, a few seconds); the cost includes Voyage.
|
||||||
|
- Voyage spend is recorded under `provider_costs["voyage"]`; `finish()` counts it once and keeps its
|
||||||
|
tokens out of the LLM aggregate. Voyage spend is not preloaded into its meter for the UTC day.
|
||||||
|
- `src/lock.rs` uses `libc::flock` (`libc` was already in `Cargo.lock`; now a direct dependency).
|
||||||
|
The holder's command name is written into the lock file for the "X is already running" message.
|
||||||
|
- The mid-run `admission:` line is `debug`; the four-line §15.4 block is logged once after
|
||||||
|
`finish()` and printed by `print_report`. Timings `summaries` and `brief` replace `editorial`.
|
||||||
|
|
||||||
|
Step 7:
|
||||||
|
|
||||||
|
- `features prune` sweeps `article_embeddings`, `candidate_runs` and now `article_assessments`;
|
||||||
|
`generate` runs it once after a published (non-dry) issue, best effort.
|
||||||
|
- The Brief chapter's TOC/title string is now "The Brief" (was still "From the Editor").
|
||||||
|
- Kept on purpose: `UsageMeter::new(&DeepseekConfig, ..)` (a convenience over `with_prices`, 12
|
||||||
|
call sites); `StageCounts.candidates` (feeds `runs.candidates` and the colophon); the
|
||||||
|
`prefilter_keep`/`score_batch_size` removed-key startup errors (operator guards); the SQL comment
|
||||||
|
"filled in by step 4/5" inside the applied migration `0002` (sqlx checksums); six v1-era
|
||||||
|
unreferenced `pub` items (`enrich_one`, `issues_before`, `today_in_tz`, `to_json_pretty`,
|
||||||
|
`with_retry_policy`, `db::issues_before`) as pre-v2 API surface.
|
||||||
|
- `config::tests::shipped_example_config_matches_the_defaults_key_for_key` compares the example
|
||||||
|
against `Config::default()` in both directions (numbers within 1e-6).
|
||||||
|
|
||||||
## Operator to-dos before the first real run
|
## Operator to-dos before the first real run
|
||||||
|
|
||||||
1. Set `DAILY_EPUB_ANTHROPIC__API_KEY` and `DAILY_EPUB_VOYAGE__API_KEY` in the systemd env file.
|
1. Set `DAILY_EPUB_ANTHROPIC__API_KEY` and `DAILY_EPUB_VOYAGE__API_KEY` in the systemd env file.
|
||||||
@@ -57,34 +95,37 @@ Nothing has been run against the production database or the real providers. No l
|
|||||||
guards, not accounting).
|
guards, not accounting).
|
||||||
3. Copy `data/profile.md` to wherever `profile_path` points on the server (default is relative to
|
3. Copy `data/profile.md` to wherever `profile_path` points on the server (default is relative to
|
||||||
`WorkingDirectory=/var/lib/daily-epub`, like `data/scour-interests.opml`).
|
`WorkingDirectory=/var/lib/daily-epub`, like `data/scour-interests.opml`).
|
||||||
4. Run `daily-epub db migrate` (0002 drops `ratings`/`feed_priors`; back up the DB first).
|
4. Back up the DB, then `daily-epub db migrate` (0002 drops `ratings`/`feed_priors`, 0003 drops `scores`).
|
||||||
5. `daily-epub features backfill --rated-only` then `--days 30` to warm the embedding cache.
|
5. `daily-epub features backfill --rated-only` then `--days 30` to warm the embedding cache.
|
||||||
6. A `generate --dry-run` and read the paper; `explain --near-misses` once step 4 lands.
|
6. `daily-epub generate --dry-run`, read the paper (including the new Behind-the-paper chapter),
|
||||||
|
then `explain --date … --near-misses` and `stats`.
|
||||||
|
7. Watch the first few real runs' `providers:` log line against the ~$0.80/day estimate (plan §3).
|
||||||
|
|
||||||
## How the work was run (so the next session can repeat it)
|
## How the work was run (so the next session can repeat it)
|
||||||
|
|
||||||
- Orchestrator: Claude Code (this repo), one implementation agent per step, review + commit by the
|
- Orchestrator: Claude Code (this repo), one implementation agent per step, review + commit by the
|
||||||
orchestrator after independent `cargo fmt/clippy/test`.
|
orchestrator after independent `cargo fmt/clippy/test`.
|
||||||
- Codex: `codex exec -C <repo> --sandbox workspace-write --add-dir ~/.cargo -c
|
- **Codex via the `openai-codex` Claude Code plugin (v1.0.6) works on this host as of session 2**
|
||||||
sandbox_workspace_write.network_access=true -c model_reasoning_effort=high -o <last.md> - < <brief>`,
|
(the operator fixed bubblewrap). Launch:
|
||||||
detached with `setsid nohup`, exit code written to a file and watched with a monitor. The
|
`node ~/.claude/plugins/cache/openai-codex/codex/1.0.6/scripts/codex-companion.mjs task --background --write --effort high "$(cat brief.md)"`,
|
||||||
`openai-codex` Claude Code plugin's `task` runs fail on this host (bubblewrap cannot create user
|
poll `… status <job-id> --json` (`.job.status`), read `… result <job-id>`. Inside its sandbox
|
||||||
namespaces: `kernel.apparmor_restrict_unprivileged_userns = 1`); the CLI works if the brief tells
|
`bind()` on 127.0.0.1 is forbidden, so exactly ten pre-existing loopback tests fail there
|
||||||
the agent to edit files via shell commands instead of the `apply_patch` tool (see the preamble).
|
(`curate::llm::tests::anthropic_*` ×4, `extract::tests::relative_urls_resolve_against_the_url_we_landed_on`,
|
||||||
The plugin's read-only `review --background --scope working-tree` does work and was used on step 1.
|
`server::tests::*` ×5, plus the two `m7_server` integration tests); the brief must tell the agent
|
||||||
Fix for the sandbox: `sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0`.
|
to ignore those and the orchestrator runs the full suite outside. The old "host quirk" paragraph
|
||||||
- Codex ran out of ChatGPT usage after ~560k tokens (step 1 complete, steps 2 and 3 cut off
|
in `00-preamble.md` is obsolete.
|
||||||
mid-way); Claude `general-purpose` sub-agents finished steps 2 and 3 from the partial trees and
|
- Codex usage budget is roughly one large step per ChatGPT usage window: in session 2 it finished
|
||||||
resolved the step 2/3 merge. Budget roughly one large step per Codex usage window.
|
step 4 (~26 min) and was cut off 21 minutes into step 5 ("try again at 7:44 PM"). The job status
|
||||||
- Steps 2 and 3 were run in parallel (step 3 in a git worktree); the merge cost ~19 conflict hunks
|
becomes `failed` with the partial work left in the tree; a Claude `general-purpose` agent with a
|
||||||
in config/pipeline/report/main/README/config.example. Steps 4 → 5 → 6 → 7 are sequential.
|
"finish, don't restart" brief (known failures, scope audit checklist) completed it. Steps 6 and 7
|
||||||
|
ran on Claude agents directly.
|
||||||
|
- Steps 4 → 5 → 6 → 7 were sequential, each reviewed against the plan sections named in its brief
|
||||||
|
before committing.
|
||||||
|
|
||||||
## Next session: exact starting instructions
|
## Next session: what remains
|
||||||
|
|
||||||
1. `git checkout curation-v2 && cargo test` (expect green).
|
1. `git checkout curation-v2 && cargo test` (expect green).
|
||||||
2. Read this file, then `docs/plans/curation-v2-briefs/00-preamble.md` and `step4.md`.
|
2. Optional: a Codex `review --background --scope branch --base main` pass over the whole branch.
|
||||||
3. Launch the step 4 agent with `cat 00-preamble.md step4.md` as the prompt (Codex CLI as above, or a
|
3. `git merge --no-ff curation-v2` into `main`, build, deploy, and do the operator to-dos above.
|
||||||
Claude general-purpose agent — tell it to ignore the "host quirk" paragraph in that case).
|
4. After a week of real runs: read `stats`, tune `[curation.ranking]` from what `explain` shows,
|
||||||
4. Review the diff against plan §10–§11, run the checks, commit as "Curation v2 step 4: …".
|
and revisit the deferred items in plan §23.
|
||||||
5. Repeat for steps 5, 6, 7. After step 7: `git merge --no-ff curation-v2` into `main`, deploy, and
|
|
||||||
do the operator to-dos above.
|
|
||||||
|
|||||||
@@ -44,9 +44,15 @@ End with a concise report (this is what the orchestrator reads): what you implem
|
|||||||
deviation from the plan and why, anything from the step you could not finish, and the exact
|
deviation from the plan and why, anything from the step you could not finish, and the exact
|
||||||
`cargo test` summary line(s). Keep it under 60 lines.
|
`cargo test` summary line(s). Keep it under 60 lines.
|
||||||
|
|
||||||
## Host quirk (important)
|
## Host notes
|
||||||
|
|
||||||
On this machine the built-in `apply_patch` tool's filesystem helper fails with
|
- The `openai-codex` Claude Code plugin sandbox works on this host (bubblewrap fixed 2026-09-02):
|
||||||
`bwrap: ... Operation not permitted`. Do not keep retrying it. Edit files through shell commands
|
`apply_patch`, the shell, `cargo build/test/clippy` and the warm `~/.cargo/registry` all work.
|
||||||
instead (the `apply_patch` CLI invoked from bash works, as do `python3 - <<'EOF'` rewrite scripts,
|
- Inside that sandbox `bind()` on 127.0.0.1 is forbidden, so ten pre-existing tests fail there for
|
||||||
`sed -i`, and heredocs). Shell commands, `cargo build`, `cargo test` and `cargo clippy` all work.
|
environmental reasons: the four `curate::llm::tests::anthropic_*` tests,
|
||||||
|
`extract::tests::relative_urls_resolve_against_the_url_we_landed_on`, and the five
|
||||||
|
`server::tests::*` (the two `tests/m7_server.rs` tests likewise). They are out of scope; do not
|
||||||
|
touch them. Every other test must pass; the orchestrator runs the full suite outside the sandbox.
|
||||||
|
- Cargo's registry is already warm; do not add dependencies that would need a network fetch.
|
||||||
|
- Work autonomously to completion. Do not stop to ask questions; make the closest-to-plan choice
|
||||||
|
and record it in the final report.
|
||||||
|
|||||||
Reference in New Issue
Block a user