@@ -91,6 +91,7 @@ If this returns article data, setup is complete.
...
@@ -91,6 +91,7 @@ If this returns article data, setup is complete.
| `chromium.userDataDir` | Optional | Browser profile dir for login state. |
| `chromium.userDataDir` | Optional | Browser profile dir for login state. |
| `chromium.maxConcurrentTabs` | Optional | Global max concurrent browser tabs across all harvest processes (default `5`). Enforced via CDP `Target.getTargets` browser self-check. |
| `chromium.maxConcurrentTabs` | Optional | Global max concurrent browser tabs across all harvest processes (default `5`). Enforced via CDP `Target.getTargets` browser self-check. |
| `sources[].paginate` | Optional | Multi-page list harvesting switch (default `false`). When `true` and the recipe configures `nextSelector` or `pageUrlTemplate`, the harvester follows article-list pages to top up candidates. See [Source Management](references/source-management.md#pagination). |
#### Dedicated profile (login state)
#### Dedicated profile (login state)
...
@@ -125,7 +126,7 @@ Two flows: **archive** (fetch + save to Obsidian) or **preview** (read-only).
...
@@ -125,7 +126,7 @@ Two flows: **archive** (fetch + save to Obsidian) or **preview** (read-only).
### How to work (read once)
### How to work (read once)
- **Scripts do all deterministic I/O** (fetch, parse, download images, write 原文.md / registry / index). The agent does only the intelligent part — Chinese translation + 4-section summaries — and hands them back as a JSON file for `finalize.js` to stitch in. The full article body is written to disk by the script and **never enters the agent's context**.
- **Scripts do all deterministic I/O** (fetch, parse, download images, write 原文.md / registry / index). The agent does only the intelligent part — Chinese translation + 4-section summaries — and hands them back as a JSON file for `finalize.js` to stitch in. The full article body is written to disk by the script and **never enters the agent's context**.
- **`count` is the target number of NEW articles**, not raw links. If a run returns fewer than `count`, the stderr prints `⚠️ … stream may be exhausted` — that's a genuine source limit, not a bug. Don't re-run for the same source/day just to "get more".
- **`count` is the target number of NEW articles**, not raw links. If a run returns fewer than `count`, the stderr prints `⚠️ … stream may be exhausted` — that's a genuine source limit, not a bug. Don't re-run for the same source/day just to "get more". For sources with pagination enabled (`paginate: true` + recipe pagination fields), this means the fetcher exhausted all `maxPages` and still couldn't find enough fresh articles.
- **Never `Read` the 原文.md files** to write summaries — the `bodyPreview` in the manifest is sufficient.
- **Never `Read` the 原文.md files** to write summaries — the `bodyPreview` in the manifest is sufficient.
- **Extract `tags` from the summary/body** (≤3 semantic keywords), not from the category/source.
- **Extract `tags` from the summary/body** (≤3 semantic keywords), not from the category/source.
- **One `finalize.js` call** after writing `.summaries.json`.
- **One `finalize.js` call** after writing `.summaries.json`.
Writes each `<中文标题>.md`, patches sentinel links in 原文.md files, updates `registry.json` + `index.md`, and removes `.manifest.json` + `.summaries.json`.
Writes each `<中文标题>.md`, patches sentinel links in 原文.md files, updates `registry.json` + `index.md`, and removes `.manifest.json` + `.summaries.json`.
> **Auto-push (optional).** If `config.push.enabled` is `true`, finalize auto-chains into `push.js`. A push failure never aborts finalize — re-push with `node scripts/push.js "<dayDir>"` once the backend is reachable.
> **Incomplete coverage.** If some articles still lack summaries, finalize stays non-fatal (finalized subset is written) but **keeps** `.manifest.json` + `.summaries.json` for resume and **skips auto-push** — 补齐缺失摘要 (`summary-add.js --status` → next batch → re-run finalize). If the temp files were already deleted by an earlier finalize, finalize rebuilds the pending list from `registry.json` (sentinel-titled entries), so writing `.summaries.json` + re-running finalize always recovers.
> **Auto-push (optional).** If `config.push.enabled` is `true` (and coverage is complete), finalize auto-chains into `push.js`. A push failure never aborts finalize — re-push with `node scripts/push.js "<dayDir>"` once the backend is reachable.
The global tab limit (`chromium.maxConcurrentTabs`, default `5`) is enforced by the browser itself via CDP `Target.getTargets` — even when multiple harvest processes run simultaneously, the total open tabs never stays above the limit. A source's `concurrency` is silently capped to the global limit.
The global tab limit (`chromium.maxConcurrentTabs`, default `5`) is enforced by the browser itself via CDP `Target.getTargets` — even when multiple harvest processes run simultaneously, the total open tabs never stays above the limit. A source's `concurrency` is silently capped to the global limit.
### Pagination
Each source has a `paginate` field (default `false`). When `true`**and** the source's recipe (`helpers/<id>.md`) configures pagination fields (`nextSelector` or `pageUrlTemplate`), the harvester follows multi-page article lists to top up candidates when a single page doesn't yield enough. This is a **two-level switch**:
-**`paginate` in config.json** — gates the feature per source (operator's call).
-**`nextSelector` / `pageUrlTemplate` in the recipe** — describes *how* to paginate (authored by the analyzer).
| Candidates are nav/section pages | Tighten `urlPattern`, or set `validateArticle: true`. |
| Candidates are nav/section pages | Tighten `urlPattern`, or set `validateArticle: true`. |
| Empty body | Fix `bodySelectors`. |
| Empty body | Fix `bodySelectors`. |
| 0 candidates, all nav | `listSelector: null` is wrong — scope to card containers. |
| 0 candidates, all nav | `listSelector: null` is wrong — scope to card containers. |
| List spans multiple pages | Configure pagination (`nextSelector` or `pageUrlTemplate`) so the harvester can follow pages. See `recipe-guide.md` → *Pagination*. Note: pagination only takes effect when the harvester's `config.json` also sets `paginate: true` for the source. |
Re-run preview after each recipe edit. Deep troubleshooting: `references/recipe-guide.md` → *Troubleshooting recipes*.
Re-run preview after each recipe edit. Deep troubleshooting: `references/recipe-guide.md` → *Troubleshooting recipes*.
| `nextSelector` | `null` | CSS selector for the "next page" element; click it to advance to the next page of the article list. `null` = pagination off. See [Pagination](#pagination-multi-page-list-harvesting) below. |
| `pageUrlTemplate` | `null` | URL template with a `{page}` placeholder, e.g. `https://site.com/news/page/{page}`. `null` = off. Mutually exclusive with `nextSelector`. |
| `pageStart` | `2` | First page number substituted into `pageUrlTemplate` (ignored when using `nextSelector`). Must be an integer ≥ 2. |
| `maxPages` | `5` | Hard cap on pages fetched beyond the first (cap: 20). |
| `pageWaitMs` | `null` | Wait after each page transition; falls back to `scrollWaitMs` when `null`. |
**Extraction** (single article body + metadata):
**Extraction** (single article body + metadata):
...
@@ -107,6 +112,69 @@ Set `validateArticle: true` when article and section/category URLs on the site *
...
@@ -107,6 +112,69 @@ Set `validateArticle: true` when article and section/category URLs on the site *
Don't set it when the URL alone cleanly separates articles from sections (e.g. Reuters'`/world/<slug>/<date>/` vs. `/world/`) — a tight `urlPattern` is cheaper than a per-candidate page fetch. `validateArticle` defaults to `false`, so existing recipes are unaffected.
Don't set it when the URL alone cleanly separates articles from sections (e.g. Reuters'`/world/<slug>/<date>/` vs. `/world/`) — a tight `urlPattern` is cheaper than a per-candidate page fetch. `validateArticle` defaults to `false`, so existing recipes are unaffected.
## Pagination (multi-page list harvesting)
Some sites split their article list across multiple pages (e.g. `/news/page/2`, `?page=3`, or a "下一页" / "Next" button). By default the harvester only collects from the **first page**(plus whatever lazy-load scrolling reveals). When a harvest requests more articles than one page yields, it prints `⚠️ stream may be exhausted` and stops short. Pagination lets the fetcher follow those pages to top up candidates.
### Two-level switch
Pagination requires **both** levels to be on:
1. **Recipe level**(this file): configure `nextSelector` or `pageUrlTemplate` to describe *how* to paginate.
2. **Config level**(`config.json`source entry): set`paginate: true` to enable the feature for that source.
The two modes are **mutually exclusive** — configuring both throws an error at load time.
| Mode | When to use | How it works |
|------|-------------|--------------|
| `nextSelector` | The page has a clickable "Next" / "下一页" / page-number element | Clicks the element, waits for navigation or SPA render, then collects. Stops when the selector is absent (last page). |
| `pageUrlTemplate` | Page URLs follow a numeric pattern (`/news/page/2`, `?page=2`) | Constructs the next URL by substituting `{page}`, navigates directly. Stops after `maxPages` or when consecutive pages yield no new candidates. |
**Pagination vs. scroll**: use `scrollSteps` / `maxScrollSteps` for **SPA infinite-scroll** lists (same URL, cards lazy-load on scroll). Use pagination for **traditional multi-page** lists (different URLs or explicit "Next" buttons). They compose: each page is still scrolled with `scrollSteps` to reveal its lazy-loaded cards before advancing.
### Recipe examples
**`nextSelector` mode** (click "Next" button):
```yaml
listSelector: 'div.article-list'
linkSelector: 'a[href]'
urlPattern: 'example\\.com/article/.+'
nextSelector: 'a.next-page' # clicked to advance; absence = last page
pageStart: 2 # first fetch beyond the homepage is /industry/all/2
maxPages: 10
```
### Notes
- **Preview never paginates** — `--preview` stays single-page for speed. Only archive harvest (where `desiredCount > 0`) follows pagination.
- **Cross-page dedup**: candidates are deduped by URL across all pages. `buildDiscoveryExpr` strips query/hash, so the same article appearing on two pages won't be double-counted.
- **Anti-bot**: each page transition re-runs the CAPTCHA/login-wall probe. A challenge mid-pagination exits with code 70 (same as a first-page block).
- **Stale exit**: if 2 consecutive pages add zero new candidates, pagination stops early (the list is exhausted). `maxPages` is a hard ceiling, not a target.