Commit 0e739208 authored by 谢宇轩's avatar 谢宇轩

feat: add page found and update push rule'

parent 2daf76f2
...@@ -91,6 +91,7 @@ If this returns article data, setup is complete. ...@@ -91,6 +91,7 @@ If this returns article data, setup is complete.
| `chromium.userDataDir` | Optional | Browser profile dir for login state. | | `chromium.userDataDir` | Optional | Browser profile dir for login state. |
| `chromium.maxConcurrentTabs` | Optional | Global max concurrent browser tabs across all harvest processes (default `5`). Enforced via CDP `Target.getTargets` browser self-check. | | `chromium.maxConcurrentTabs` | Optional | Global max concurrent browser tabs across all harvest processes (default `5`). Enforced via CDP `Target.getTargets` browser self-check. |
| `sources[].concurrency` | Optional | Per-source concurrent article fetch count (default `1` = serial). Capped by `chromium.maxConcurrentTabs`. See [Source Management](references/source-management.md#concurrency). | | `sources[].concurrency` | Optional | Per-source concurrent article fetch count (default `1` = serial). Capped by `chromium.maxConcurrentTabs`. See [Source Management](references/source-management.md#concurrency). |
| `sources[].paginate` | Optional | Multi-page list harvesting switch (default `false`). When `true` and the recipe configures `nextSelector` or `pageUrlTemplate`, the harvester follows article-list pages to top up candidates. See [Source Management](references/source-management.md#pagination). |
#### Dedicated profile (login state) #### Dedicated profile (login state)
...@@ -125,7 +126,7 @@ Two flows: **archive** (fetch + save to Obsidian) or **preview** (read-only). ...@@ -125,7 +126,7 @@ Two flows: **archive** (fetch + save to Obsidian) or **preview** (read-only).
### How to work (read once) ### How to work (read once)
- **Scripts do all deterministic I/O** (fetch, parse, download images, write 原文.md / registry / index). The agent does only the intelligent part — Chinese translation + 4-section summaries — and hands them back as a JSON file for `finalize.js` to stitch in. The full article body is written to disk by the script and **never enters the agent's context**. - **Scripts do all deterministic I/O** (fetch, parse, download images, write 原文.md / registry / index). The agent does only the intelligent part — Chinese translation + 4-section summaries — and hands them back as a JSON file for `finalize.js` to stitch in. The full article body is written to disk by the script and **never enters the agent's context**.
- **`count` is the target number of NEW articles**, not raw links. If a run returns fewer than `count`, the stderr prints `⚠️ … stream may be exhausted` — that's a genuine source limit, not a bug. Don't re-run for the same source/day just to "get more". - **`count` is the target number of NEW articles**, not raw links. If a run returns fewer than `count`, the stderr prints `⚠️ … stream may be exhausted` — that's a genuine source limit, not a bug. Don't re-run for the same source/day just to "get more". For sources with pagination enabled (`paginate: true` + recipe pagination fields), this means the fetcher exhausted all `maxPages` and still couldn't find enough fresh articles.
- **Never `Read` the 原文.md files** to write summaries — the `bodyPreview` in the manifest is sufficient. - **Never `Read` the 原文.md files** to write summaries — the `bodyPreview` in the manifest is sufficient.
- **Extract `tags` from the summary/body** (≤3 semantic keywords), not from the category/source. - **Extract `tags` from the summary/body** (≤3 semantic keywords), not from the category/source.
- **One `finalize.js` call** after writing `.summaries.json`. - **One `finalize.js` call** after writing `.summaries.json`.
...@@ -178,7 +179,9 @@ node scripts/finalize.js "<dayDir>" ...@@ -178,7 +179,9 @@ node scripts/finalize.js "<dayDir>"
``` ```
Writes each `<中文标题>.md`, patches sentinel links in 原文.md files, updates `registry.json` + `index.md`, and removes `.manifest.json` + `.summaries.json`. Writes each `<中文标题>.md`, patches sentinel links in 原文.md files, updates `registry.json` + `index.md`, and removes `.manifest.json` + `.summaries.json`.
> **Auto-push (optional).** If `config.push.enabled` is `true`, finalize auto-chains into `push.js`. A push failure never aborts finalize — re-push with `node scripts/push.js "<dayDir>"` once the backend is reachable. > **Incomplete coverage.** If some articles still lack summaries, finalize stays non-fatal (finalized subset is written) but **keeps** `.manifest.json` + `.summaries.json` for resume and **skips auto-push** — 补齐缺失摘要 (`summary-add.js --status` → next batch → re-run finalize). If the temp files were already deleted by an earlier finalize, finalize rebuilds the pending list from `registry.json` (sentinel-titled entries), so writing `.summaries.json` + re-running finalize always recovers.
> **Auto-push (optional).** If `config.push.enabled` is `true` (and coverage is complete), finalize auto-chains into `push.js`. A push failure never aborts finalize — re-push with `node scripts/push.js "<dayDir>"` once the backend is reachable.
### Flow B — Preview (read-only, no archive) ### Flow B — Preview (read-only, no archive)
...@@ -197,7 +200,9 @@ node scripts/push.js "<dayDir>" --status # read-only sync status ...@@ -197,7 +200,9 @@ node scripts/push.js "<dayDir>" --status # read-only sync status
node scripts/push.js "<dayDir>" --force # re-push all articles node scripts/push.js "<dayDir>" --force # re-push all articles
``` ```
`push.js` reads `registry.json`, parses each 原文.md / 摘要.md, uploads images to MinIO (SHA-256 deduped), and batch-pushes articles (≤10/batch, idempotent). Exit code `2` = partial failure. See `references/push-api-spec.md`. `push.js` reads `registry.json`, parses each 原文.md / 摘要.md, uploads images to MinIO (SHA-256 deduped), and batch-pushes articles (≤10/batch, idempotent).
**Readiness gate** — before uploading anything, every pending article must be summary-complete (real `titleZh`, not the `__SUMMARY_*__` sentinel; summary doc present with non-empty sections). Articles missing summaries block the whole push (exit `1`, nothing uploaded) with a 补充摘要 recovery guide: `summary-add.js --status` → write missing summaries → `finalize.js` → re-push. Exit code `2` = partial failure. See `references/push-api-spec.md`.
--- ---
......
{
"format": "news-harvester-source",
"version": 1,
"exportedAt": "2026-08-18T00:00:00.000Z",
"source": {
"id": "reuters-technology",
"name": "路透社科技",
"homepageUrl": "https://www.reuters.com/technology/",
"method": "browser",
"vaultFolder": "路透社/科技"
},
"files": {
"helpers/reuters-technology.md": "---\nsource: reuters-technology\nlistSelector: 'div[data-testid=\"Title\"]'\nlinkSelector: 'a[href]'\nurlPattern: 'reuters\\.com/.+-\\d{4}-\\d{2}-\\d{2}/?$'\nexcludeUrlPattern: 'reuters\\.com/(podcasts|newsletter)/'\nscrollSteps: 3\nscrollWaitMs: 2000\nwaitMs: 15000\nvalidateArticle: true\nbodyMinChars: 500\nbodySelectors:\n - '[data-testid^=\"paragraph\"]'\n - '[class*=\"articleBodyContent\"] p'\n - '[class*=\"article-body\"] p'\n - 'article p'\nbodyMinParagraphs: 3\ndateSelector: 'time[datetime]'\nauthorSelector: '[rel=\"author\"], [class*=\"author\"] a'\nimageSelector: 'img'\nimageMinWidth: 200\nexcludeImage:\n - logo\n - icon\n - avatar\n - profile\ntitleStripSuffix: ' \\| Reuters$'\nmaxBodyChars: 15000\n---\n# 路透社 Technology 抓取备忘\n\n## 列表页链接收集(`/technology/`)\n\n- **listSelector: `div[data-testid=\"Title\"]`**。实测全页文章链接全部在此容器内(单一容器覆盖全部,无分散问题)。该 data-testid 比 CSS Modules 类名稳定。\n- **urlPattern: `reuters\\.com/.+-\\d{4}-\\d{2}-\\d{2}/?$`**。Reuters 文章 URL 以 `-YYYY-MM-DD/` 结尾,日期后缀是极强的文章标识——分类页(`/technology/`、`/technology/artificial-intelligence/`)无日期后缀。放宽为全站匹配(不限 `/technology/`),因为板块页会推荐其他分区的文章。\n- **excludeUrlPattern** 排除 `/podcasts/`、`/newsletter/`(它们也带日期但非正文报道)。\n- `scrollSteps: 3` 触发懒加载拿全当天列表。\n\n## 详情页正文验证\n\n`validateArticle: true` 让 fetchPage 用 `og:type` + JSON-LD `@type` + 正文长度剔除分类页。实测信号:\n- 文章:`og:type=article`、`@type=NewsArticle`、正文 6000+ 字符\n- 分类页 `/technology/`:`og:type=website`、正文 0\n\nReuters 的日期后缀 URL 已能区分大部分文章/分类页,validateArticle 作为兜底防止非标准 URL 混入。\n\n## 其他\n\n- 反爬:headless 会被拦,需 `config.chromium.headless: false`。\n- 正文优先 `[data-testid^=\"paragraph\"]`,旧 `article p` 会被\"相关阅读\"\"订阅推广\"段落污染。\n"
}
}
{
"format": "news-harvester-source",
"version": 1,
"exportedAt": "2026-08-18T07:11:41.049Z",
"source": {
"id": "yna-industry",
"name": "연합뉴스 산업",
"homepageUrl": "https://www.yna.co.kr/industry/all/1",
"method": "browser",
"vaultFolder": "industry/yna"
},
"files": {
"helpers/yna-industry.md": "---\nsource: yna-industry\nlistSelector: 'figure.img-con01'\nlinkSelector: 'a[href]'\nurlPattern: 'yna\\.co\\.kr/view/AKR'\nexcludeUrlPattern: null\nscrollSteps: 3\nscrollWaitMs: 2000\nwaitMs: 15000\nmaxLinks: 30\n# ── 翻页 ──\n# 列表按页码分页:/industry/all/1(首页)→ /industry/all/2..11。\n# 首页已由 harvest 加载,翻页从 page=2 开始构造 URL。\npageUrlTemplate: 'https://www.yna.co.kr/industry/all/{page}'\npageStart: 2\nmaxPages: 10\npageWaitMs: 2000\nbodySelectors:\n - '.story-news article p'\n - '.story-news p'\n - 'article p'\nbodyMinParagraphs: 3\n# No <time> on Yonhap; JSON-LD datePublished is the primary source.\n# dateSelector set to a non-matching selector so the JSON-LD fallback fires.\ndateSelector: 'time[datetime]'\n# No visible author name on Yonhap (only \"기자\" generic label); JSON-LD\n# author.name is the real source — keep selector non-matching so it fires.\nauthorSelector: 'span.nonexistent-author'\n# Article body photos live inside .image-zone01 (within .story-news). Scoping\n# here excludes the bottom \"recommended articles\" strip (.img-con11 /\n# .item-box01) and reporter headshots by structure, not by fragile URL paths.\nimageSelector: '.image-zone01 img'\nimageMinWidth: 200\nexcludeImage:\n - logo\n - icon\n - avatar\n - profile\n - contentsfeed\n - ad.yna.co.kr\n - /reporter/\nexcludeImageSelector: null\nnoisePrefixes:\n - '제보는 카카오톡'\n - '<저작권자'\n - '무단 전재'\ntitleStripSuffix: ' \\| 연합뉴스$'\nmaxBodyChars: 15000\n---\n# 연합뉴스 산업 (Yonhap / industry) 抓取备忘\n\n- 列表页 `https://www.yna.co.kr/industry/all/1`,文章卡片容器 `figure.img-con01`(实测 30 篇/页)。\n- inspect 自动选中的 `li.swiper-slide`(38 链接)是顶部导航横滑菜单,非文章流——已手动改为 `figure.img-con01`。\n- 文章 URL 形如 `/view/AKR20260818084800017?section=industry/all`,urlPattern 用 `yna\\.co\\.kr/view/AKR` 精准命中正文页,排除 section/导航页。\n- 翻页:列表按页码分页 `/industry/all/1..11`(每页约 30 篇)。recipe 已配 `pageUrlTemplate`,从第 2 页起构造 URL 翻页采集;需在 config.json 设 `paginate: true` 启用。\n- 正文容器实测为 `.story-news article`(24 段);inspect 自动模板里的 `.article-let-box` 在本站不存在。\n- 元数据全靠 JSON-LD:文章页有 3 个 `application/ld+json` 块,`datePublished`(ISO 8601 带时区)和 `author.name` 在第 3 个 `NewsArticle` 块。fetcher-helper 现已扫描全部 JSON-LD 块(此前只读第 1 个 `NewsMediaOrganization` 块,拿不到日期/作者)。\n- 图片:正文图位于 `.image-zone01`(`.story-news` 内),通过结构化选择器精准定位,排除底部推荐/视频缩略图(`.img-con11`)和记者头像(`/reporter/`)。广告 CDN `contentsfeed` 也排除。\n- 正文噪声:`제보는 카카오톡`(举报引导)、`<저작권자` 及 `무단 전재`(版权声明)用 noisePrefixes 丢弃。\n- 验证: node scripts/preview.js --url https://www.yna.co.kr/industry/all/1 --recipe helpers/yna-industry.md --count 3\n",
"helpers/yna-industry.fixtures.json": "{\n \"articles\": [\n \"https://www.yna.co.kr/view/AKR20260818080451017\",\n \"https://www.yna.co.kr/view/AKR20260818049552017\",\n \"https://www.yna.co.kr/view/AKR20260818046751004\"\n ],\n \"sections\": [\n \"https://www.yna.co.kr/industry/all/1\",\n \"https://www.yna.co.kr/industry/index\"\n ],\n \"minLinks\": 10\n}\n"
}
}
\ No newline at end of file
...@@ -383,12 +383,13 @@ node scripts/push.js -h | --help ...@@ -383,12 +383,13 @@ node scripts/push.js -h | --help
**工作流**: **工作流**:
1. 校验 day 目录已最终化(无 `.manifest.json`/`.summaries.json`)。 1. 校验 day 目录已最终化(无 `.manifest.json`/`.summaries.json`)。
2. 读 `registry.json` → 文章列表;读 `.syncstate.json` → 过滤已 synced。 2. 就绪检查:每篇待推送文章必须有真实 `titleZh`(非 `__SUMMARY_*__` sentinel)且摘要文档四段非空 —— 缺摘要的文章**阻止整个推送**(exit 1,不上传任何内容),输出补充摘要的恢复指引(`summary-add.js --status` → 补摘要 → `finalize.js` → 重推)。
3. 对每篇文章:读 `原文.md` + `summary/<titleZh>.md` + `assets/*`。 3. 读 `registry.json` → 文章列表;读 `.syncstate.json` → 过滤已 synced。
4. 阶段一:上传未同步的图片(按 sha256 去重)→ 收集 `assetId`。 4. 对每篇文章:读 `原文.md` + `summary/<titleZh>.md` + `assets/*`。
5. 阶段二:按 ≤10/批切分 → `POST /articles/batch`(带 `Idempotency-Key`)。 5. 阶段一:上传未同步的图片(按 sha256 去重)→ 收集 `assetId`。
6. 按 `results` 更新 `.syncstate.json`;失败项保留待重试。 6. 阶段二:按 ≤10/批切分 → `POST /articles/batch`(带 `Idempotency-Key`)。
7. 输出摘要到 stderr,stdout 输出 JSON 结果(便于 agent/脚本消费)。 7. 按 `results` 更新 `.syncstate.json`;失败项保留待重试。
8. 输出摘要到 stderr,stdout 输出 JSON 结果(便于 agent/脚本消费)。
### 9.1 config.json 扩展 ### 9.1 config.json 扩展
......
...@@ -20,7 +20,7 @@ RECIPE ✅ = `helpers/<id>.md` exists (spec-driven fetcher-helper); — = no rec ...@@ -20,7 +20,7 @@ RECIPE ✅ = `helpers/<id>.md` exists (spec-driven fetcher-helper); — = no rec
node scripts/source-manage.js edit <id> --name "New Name" --url https://... --method browser node scripts/source-manage.js edit <id> --name "New Name" --url https://... --method browser
``` ```
Editable fields: `name`, `homepageUrl`, `method`, `articleUrlPattern`, `vaultFolder`, `enabled`, `concurrency`. Editable fields: `name`, `homepageUrl`, `method`, `articleUrlPattern`, `vaultFolder`, `enabled`, `concurrency`, `paginate`.
### Concurrency ### Concurrency
...@@ -32,6 +32,26 @@ node scripts/source-manage.js edit <id> --concurrency 3 ...@@ -32,6 +32,26 @@ node scripts/source-manage.js edit <id> --concurrency 3
The global tab limit (`chromium.maxConcurrentTabs`, default `5`) is enforced by the browser itself via CDP `Target.getTargets` — even when multiple harvest processes run simultaneously, the total open tabs never stays above the limit. A source's `concurrency` is silently capped to the global limit. The global tab limit (`chromium.maxConcurrentTabs`, default `5`) is enforced by the browser itself via CDP `Target.getTargets` — even when multiple harvest processes run simultaneously, the total open tabs never stays above the limit. A source's `concurrency` is silently capped to the global limit.
### Pagination
Each source has a `paginate` field (default `false`). When `true` **and** the source's recipe (`helpers/<id>.md`) configures pagination fields (`nextSelector` or `pageUrlTemplate`), the harvester follows multi-page article lists to top up candidates when a single page doesn't yield enough. This is a **two-level switch**:
- **`paginate` in config.json** — gates the feature per source (operator's call).
- **`nextSelector` / `pageUrlTemplate` in the recipe** — describes *how* to paginate (authored by the analyzer).
```bash
node scripts/source-manage.js edit <id> --paginate true
```
| `paginate` | recipe has pagination | Behaviour |
|------------|----------------------|-----------|
| `true` | yes | ✅ Multi-page harvesting |
| `true` | no | ⚠️ Warning printed, falls back to single-page (no error) |
| `false` / unset | yes | ℹ️ Info printed, single-page (fields ignored) |
| `false` / unset | no | ✅ Legacy single-page behaviour |
Install sets `paginate: false`; it is preserved across re-installs. See `recipe-guide.md` → *Pagination* for how to author the recipe side.
## Enable / Disable Source ## Enable / Disable Source
```bash ```bash
......
...@@ -403,6 +403,55 @@ async function scrollAndCollect(ws, { ...@@ -403,6 +403,55 @@ async function scrollAndCollect(ws, {
return merged; return merged;
} }
// ── Click + wait (pagination) ────────────────────────────────────────────────
// Click the element matching `selector` and wait for the resulting navigation
// or SPA render. Used by fetchHomeLinks to follow "next page" links across
// multi-page article lists.
//
// Returns true if the element was found and clicked (caller should continue
// paginating), false if the selector matched nothing (treat as "last page").
//
// The wait mirrors navigateAndWait: a full-page navigation fires
// Page.loadEventFired (raced against a 25s timeout), while an SPA transition
// that swaps content without a navigation falls through to `sleep(waitMs)` so
// the new cards have time to render before the next collect pass.
async function clickAndWait(ws, selector, waitMs = 2000) {
await cdpSend(ws, 'Page.enable');
// Attach the load listener BEFORE clicking, since loadEventFired can fire
// before the cdpEval response arrives. The handler is hoisted so we can
// detach it early if the selector matches nothing (no click → no navigation).
let loaded = false;
let resolveLoad;
const handler = (raw) => {
const msg = JSON.parse(raw.toString());
if (msg.method === 'Page.loadEventFired' && !loaded) {
loaded = true;
ws.removeListener('message', handler);
resolveLoad();
}
};
ws.on('message', handler);
const loadPromise = new Promise(resolve => { resolveLoad = resolve; });
const expr = `(function(){
var el = document.querySelector(${JSON.stringify(selector)});
if (!el) return false;
el.click();
return true;
})()`;
const clicked = await cdpEval(ws, expr);
if (!clicked) {
ws.removeListener('message', handler);
return false;
}
await Promise.race([loadPromise, sleep(25000)]);
await sleep(waitMs);
return true;
}
module.exports = { module.exports = {
sleep, sleep,
resolveCdpPort, resolveCdpPort,
...@@ -412,6 +461,7 @@ module.exports = { ...@@ -412,6 +461,7 @@ module.exports = {
cdpEval, cdpEval,
navigateAndWait, navigateAndWait,
scrollAndCollect, scrollAndCollect,
clickAndWait,
// Browser-level target management (multi-tab concurrency) // Browser-level target management (multi-tab concurrency)
getBrowserWs, getBrowserWs,
countPageTabs, countPageTabs,
......
...@@ -21,7 +21,7 @@ const fs = require('fs'); ...@@ -21,7 +21,7 @@ const fs = require('fs');
const path = require('path'); const path = require('path');
const yaml = require('js-yaml'); const yaml = require('js-yaml');
const { openSession, navigateAndWait, scrollAndCollect, cdpEval } = require('./cdp-client'); const { openSession, navigateAndWait, scrollAndCollect, clickAndWait, cdpEval } = require('./cdp-client');
const { evaluateBrowserBlock } = require('./block-check'); const { evaluateBrowserBlock } = require('./block-check');
// ── Spec defaults ──────────────────────────────────────────────────────────── // ── Spec defaults ────────────────────────────────────────────────────────────
...@@ -41,6 +41,19 @@ const DEFAULTS = { ...@@ -41,6 +41,19 @@ const DEFAULTS = {
waitMs: 12000, // initial render wait after load waitMs: 12000, // initial render wait after load
maxLinks: 30, maxLinks: 30,
// ── Pagination (multi-page list harvesting) ──
// Two mutually exclusive modes: nextSelector (click an element to advance)
// or pageUrlTemplate (construct the next-page URL by substituting {page}).
// Both are null by default → legacy single-page behaviour. Even when
// configured here, pagination only takes effect if the source's config.json
// entry has paginate: true (a two-level switch so operators can enable per
// source without touching the recipe).
nextSelector: null, // CSS selector for the "next page" element; click it to advance. null = off
pageUrlTemplate: null, // URL template with {page} placeholder, e.g. "https://site.com/news/page/{page}". null = off
pageStart: 2, // first page number substituted into pageUrlTemplate (ignored for nextSelector)
maxPages: 5, // hard cap on pages fetched beyond the first
pageWaitMs: null, // wait after each page transition; falls back to scrollWaitMs when null
// ── Extraction ── // ── Extraction ──
bodySelectors: [ bodySelectors: [
'[data-testid^="paragraph"]', '[data-testid^="paragraph"]',
...@@ -89,6 +102,16 @@ function loadSpec(helperPath) { ...@@ -89,6 +102,16 @@ function loadSpec(helperPath) {
if (!spec.urlPattern) { if (!spec.urlPattern) {
throw new Error(`helper.md ${path.basename(helperPath)} missing required 'urlPattern'`); throw new Error(`helper.md ${path.basename(helperPath)} missing required 'urlPattern'`);
} }
// Pagination: nextSelector and pageUrlTemplate are mutually exclusive.
if (spec.nextSelector && spec.pageUrlTemplate) {
throw new Error(`helper.md ${path.basename(helperPath)}: nextSelector and pageUrlTemplate are mutually exclusive — configure at most one`);
}
if (spec.pageUrlTemplate && (!Number.isInteger(spec.pageStart) || spec.pageStart < 2)) {
throw new Error(`helper.md ${path.basename(helperPath)}: pageStart must be an integer ≥ 2 when pageUrlTemplate is set`);
}
if (spec.maxPages > 20) {
throw new Error(`helper.md ${path.basename(helperPath)}: maxPages cap is 20 (got ${spec.maxPages})`);
}
return spec; return spec;
} }
...@@ -157,13 +180,19 @@ function buildExtractExpr(spec) { ...@@ -157,13 +180,19 @@ function buildExtractExpr(spec) {
const bodyMinChars = spec.bodyMinChars || 500; const bodyMinChars = spec.bodyMinChars || 500;
return `(function(){ return `(function(){
var title = document.title.replace(new RegExp(${titleStripRe}, 'g'), '').trim(); var title = document.title.replace(new RegExp(${titleStripRe}, 'g'), '').trim();
// JSON-LD fallback: many sites (Nikkei, Reuters) embed datePublished/author // JSON-LD fallback: many sites (Nikkei, Reuters, Yonhap) embed
// only in <script type="application/ld+json">, not in <time> or visible // datePublished/author only in <script type="application/ld+json">, not in
// author links. Parse it once up front so the CSS-selector results can // <time> or visible author links. Sites with multiple blocks (e.g. a
// fall back to it below. // NewsMediaOrganization block before the NewsArticle block) need all blocks
// scanned — ldGet already recurses into arrays. Parse every block up front
// so the CSS-selector results can fall back to it below.
var ld = null; var ld = null;
var ldEl = document.querySelector('script[type="application/ld+json"]'); var ldEls = document.querySelectorAll('script[type="application/ld+json"]');
if (ldEl) { try { ld = JSON.parse(ldEl.textContent); } catch (e) {} } if (ldEls.length) {
var blocks = [];
ldEls.forEach(function(el){ try { blocks.push(JSON.parse(el.textContent)); } catch(e){} });
ld = blocks.length === 1 ? blocks[0] : blocks;
}
function ldGet(obj, key) { function ldGet(obj, key) {
if (!obj) return null; if (!obj) return null;
if (Array.isArray(obj)) { for (var i=0;i<obj.length;i++){ var v=ldGet(obj[i],key); if(v) return v; } return null; } if (Array.isArray(obj)) { for (var i=0;i<obj.length;i++){ var v=ldGet(obj[i],key); if(v) return v; } return null; }
...@@ -266,28 +295,113 @@ function buildExtractExpr(spec) { ...@@ -266,28 +295,113 @@ function buildExtractExpr(spec) {
// avoids tab churn in serial workflows like preview.js / recipe-test.js where // avoids tab churn in serial workflows like preview.js / recipe-test.js where
// many pages are fetched one after another. When omitted, a new session is // many pages are fetched one after another. When omitted, a new session is
// created and closed in the finally block (harvester behavior). // created and closed in the finally block (harvester behavior).
async function fetchHomeLinks(url, spec, desiredCount, session) { //
// Optional `paginate` (harvest.js): when true AND desiredCount > 0 AND the
// recipe configures a pagination mode (nextSelector or pageUrlTemplate),
// follow multi-page article lists — scrolling each page with fixed scrollSteps,
// then advancing to the next page, cross-page deduping until desiredCount is
// met, maxPages is exhausted, or consecutive pages yield no new candidates.
// Two-level switch: config.json `paginate` gates the feature per source; the
// recipe fields describe *how* to paginate. A warning is printed when one
// level is on but the other is missing.
async function fetchHomeLinks(url, spec, desiredCount, session, paginate) {
const ownSession = !session; const ownSession = !session;
const { ws, close } = session || await openSession(); const { ws, close } = session || await openSession();
try { try {
await navigateAndWait(ws, url, spec.waitMs); await navigateAndWait(ws, url, spec.waitMs);
await evaluateBrowserBlock(ws); await evaluateBrowserBlock(ws);
const opts = {
waitMs: spec.scrollWaitMs, const recipePaginate = !!(spec.nextSelector || spec.pageUrlTemplate);
collectExpr: buildDiscoveryExpr(spec), const archiveMode = desiredCount && desiredCount > 0;
dedupKey: item => item.url const enablePaginate = archiveMode && paginate && recipePaginate;
};
if (desiredCount && desiredCount > 0) { // Two-level switch diagnostics: warn when only one level is configured.
opts.targetCount = desiredCount; if (archiveMode && paginate && !recipePaginate) {
opts.maxSteps = spec.maxScrollSteps; console.error(' ⚠️ config 已启用 paginate 但 recipe 未配置翻页字段(nextSelector / pageUrlTemplate),按单页采集。');
} else { }
opts.steps = spec.scrollSteps; if (archiveMode && !paginate && recipePaginate) {
console.error(' ℹ️ recipe 支持翻页但 config.paginate 未启用;如需跨页采集,在 config.json 该 source 下设 paginate: true。');
}
if (!enablePaginate) {
// Legacy single-page path (scroll lazy-load only).
const opts = {
waitMs: spec.scrollWaitMs,
collectExpr: buildDiscoveryExpr(spec),
dedupKey: item => item.url
};
if (archiveMode) {
opts.targetCount = desiredCount;
opts.maxSteps = spec.maxScrollSteps;
} else {
opts.steps = spec.scrollSteps;
}
const items = await scrollAndCollect(ws, opts);
const cap = archiveMode
? Math.max(spec.maxLinks || 30, desiredCount)
: (spec.maxLinks || 30);
return items.map(i => i.url).slice(0, cap);
} }
const items = await scrollAndCollect(ws, opts);
const cap = (desiredCount && desiredCount > 0) // ── Pagination mode ──
? Math.max(spec.maxLinks || 30, desiredCount) // Each page is scrolled with fixed scrollSteps (not target-count mode —
: (spec.maxLinks || 30); // the outer desiredCount drives the cross-page loop, not per-page scroll).
return items.map(i => i.url).slice(0, cap); // Candidates are deduped across pages via the seen Set; buildDiscoveryExpr
// already strips query/hash so the same article on different pages dedupes.
const seen = new Set();
const merged = [];
let stalePages = 0;
const pageWaitMs = spec.pageWaitMs || spec.scrollWaitMs;
const maxPages = Math.min(spec.maxPages || 5, 20);
for (let pageIdx = 0; pageIdx <= maxPages; pageIdx++) {
// pageIdx 0 is the already-loaded first page; subsequent iterations
// navigate/click to the next page before collecting.
if (pageIdx > 0) {
if (spec.nextSelector) {
const advanced = await clickAndWait(ws, spec.nextSelector, pageWaitMs);
if (!advanced) {
console.error(` 📄 已到最后一页(第 ${pageIdx + 1} 页无"下一页"元素),停止翻页。`);
break;
}
await evaluateBrowserBlock(ws);
} else {
// pageUrlTemplate: pageIdx=1 → pageStart, pageIdx=2 → pageStart+1, …
const nextPageNum = spec.pageStart + (pageIdx - 1);
const nextUrl = spec.pageUrlTemplate.replace('{page}', String(nextPageNum));
await navigateAndWait(ws, nextUrl, spec.waitMs);
await evaluateBrowserBlock(ws);
}
}
const pageOpts = {
steps: spec.scrollSteps,
waitMs: spec.scrollWaitMs,
collectExpr: buildDiscoveryExpr(spec),
dedupKey: item => item.url
};
const pageItems = await scrollAndCollect(ws, pageOpts);
const before = merged.length;
for (const item of pageItems) {
if (!seen.has(item.url)) {
seen.add(item.url);
merged.push(item);
}
}
const added = merged.length - before;
console.error(` 📄 第 ${pageIdx + 1} 页:+${added} 条候选(累计 ${merged.length})`);
if (merged.length >= desiredCount) break;
stalePages = (added === 0) ? stalePages + 1 : 0;
if (stalePages >= 2) {
console.error(' 📄 连续 2 页无新增候选,流可能已干涸,停止翻页。');
break;
}
}
const cap = Math.max(spec.maxLinks || 30, desiredCount);
return merged.map(i => i.url).slice(0, cap);
} finally { } finally {
if (ownSession) await close(); if (ownSession) await close();
} }
......
...@@ -13,8 +13,15 @@ ...@@ -13,8 +13,15 @@
* 3. Update the registry entry (titleZh / category / tags) * 3. Update the registry entry (titleZh / category / tags)
* Finally rewrites index.md with real titles and removes the two temp files. * Finally rewrites index.md with real titles and removes the two temp files.
* *
* Missing summary for an article → warn and leave its sentinel in place (non-fatal). * Missing summary for an article → warn and leave its sentinel in place (non-fatal),
* Missing .summaries.json → error and exit. * BUT with two guards for incomplete coverage:
* - .manifest.json / .summaries.json are KEPT (so summary-add.js resume works)
* - auto-push into push.js is skipped (push.js would block anyway: it refuses
* articles whose titleZh is still the __SUMMARY_*__ sentinel)
* If .manifest.json is absent (an older finalize already deleted it while
* coverage was incomplete), the pending list is rebuilt from registry.json —
* entries still carrying the sentinel titleZh — so summaries can be补齐 and
* finalize re-run. Missing .summaries.json → error and exit.
*/ */
const fs = require('fs'); const fs = require('fs');
...@@ -40,6 +47,54 @@ function loadJson(p, label) { ...@@ -40,6 +47,54 @@ function loadJson(p, label) {
} }
} }
const SENTINEL_RE = /^__SUMMARY_.*__$/;
// Best-effort source display name for the rebuilt manifest / index header.
function sourceNameFor(sourceId) {
try {
const cfg = JSON.parse(fs.readFileSync(path.join(__dirname, '..', 'config.json'), 'utf8'));
const s = (cfg.sources || []).find(x => x.id === sourceId);
if (s && s.name) return s.name;
} catch (e) { /* no config / no match — fall through */ }
return 'Unknown';
}
// Read `authors: "..."` back from a 原文.md frontmatter. The registry does not
// store authors, so this is needed when rebuilding the pending list from it.
function readAuthorsFromOriginal(origPath) {
try {
if (!origPath || !fs.existsSync(origPath)) return '';
const md = fs.readFileSync(origPath, 'utf8');
const m = md.match(/^authors: "(.*)"$/m);
return m ? m[1] : '';
} catch (e) { return ''; }
}
// Rebuild a pseudo-manifest from registry.json for a re-finalize after a
// previous finalize already removed .manifest.json while some articles never
// got summaries. Only sentinel-titled entries (never summarized) are included.
// Returns null when there is nothing left to finalize this way.
function rebuildManifestFromRegistry(dayDir) {
const regPath = path.join(dayDir, 'registry.json');
if (!fs.existsSync(regPath)) return null;
let registry;
try { registry = JSON.parse(fs.readFileSync(regPath, 'utf8')); } catch (e) { return null; }
const articles = (Array.isArray(registry.articles) ? registry.articles : [])
.filter(a => a && a.url && SENTINEL_RE.test((a.titleZh || '').trim()))
.map(a => ({
url: a.url,
slug: a.slug,
titleEn: a.titleEn || '',
originalFilename: a.originalFilename,
date: a.date,
authors: readAuthorsFromOriginal(path.join(dayDir, a.originalFilename || '')),
category: a.category,
tags: a.tags,
}));
if (articles.length === 0) return null;
return { sourceId: registry.source, sourceName: sourceNameFor(registry.source), articles };
}
function printHelp() { function printHelp() {
console.log(`finalize.js — Patch Chinese summaries into an archived day directory. console.log(`finalize.js — Patch Chinese summaries into an archived day directory.
...@@ -89,7 +144,21 @@ function main() { ...@@ -89,7 +144,21 @@ function main() {
const manifestPath = path.join(dayDir, '.manifest.json'); const manifestPath = path.join(dayDir, '.manifest.json');
const summariesPath = path.join(dayDir, '.summaries.json'); const summariesPath = path.join(dayDir, '.summaries.json');
const manifest = loadJson(manifestPath, '.manifest.json'); let manifest;
if (fs.existsSync(manifestPath)) {
manifest = loadJson(manifestPath, '.manifest.json');
} else {
// A previous finalize already ran but some articles never got summaries —
// recover the pending list from registry.json (sentinel-titled entries).
manifest = rebuildManifestFromRegistry(dayDir);
if (!manifest) {
console.error(`✗ .manifest.json not found in ${dayDir}`);
console.error(` Has harvest.js run for this day? (registry.json is missing or fully finalized)`);
process.exit(1);
}
console.error(`ℹ️ .manifest.json absent — rebuilt pending list from registry.json`);
console.error(` ${manifest.articles.length} article(s) still missing summaries`);
}
const summaries = loadJson(summariesPath, '.summaries.json'); const summaries = loadJson(summariesPath, '.summaries.json');
const sourceName = manifest.sourceName || 'Unknown'; const sourceName = manifest.sourceName || 'Unknown';
...@@ -196,17 +265,32 @@ function main() { ...@@ -196,17 +265,32 @@ function main() {
})); }));
updateIndex(dayDir, indexArticles, sourceName); updateIndex(dayDir, indexArticles, sourceName);
// 5. Cleanup temp files // 5. Cleanup temp files — ONLY on full coverage. When articles are still
try { fs.unlinkSync(manifestPath); } catch (e) {} // missing summaries, keep .manifest.json/.summaries.json so the batch
try { fs.unlinkSync(summariesPath); } catch (e) {} // resume flow (summary-add.js --status → next batch → re-run finalize)
// keeps working; push.js also refuses a day dir that still has them.
console.error(`\n✅ Done: ${patched} finalized, ${skipped} skipped`); const complete = skipped === 0;
console.error(` Registry + index updated, temp files removed`); if (complete) {
try { fs.unlinkSync(manifestPath); } catch (e) {}
try { fs.unlinkSync(summariesPath); } catch (e) {}
}
// 6. Auto-push to News Hub if enabled and configured. Push failures do NOT if (complete) {
// abort finalize (the archive is already complete and re-pushable), so console.error(`\n✅ Done: ${patched} finalized, 0 skipped`);
// a backend hiccup never leaves the local archive in a bad state. console.error(` Registry + index updated, temp files removed`);
maybeAutoPush(dayDir);
// 6. Auto-push to News Hub if enabled and configured. Push failures do NOT
// abort finalize (the archive is already complete and re-pushable), so
// a backend hiccup never leaves the local archive in a bad state.
maybeAutoPush(dayDir);
} else {
console.error(`\n⚠️ Finalize 部分完成: ${patched} finalized, ${skipped} 篇仍缺中文摘要`);
console.error(` 已完成文章的摘要文档 / registry / index 已更新;临时文件保留以便续补。`);
console.error(` 补齐缺失摘要:`);
console.error(` 1. node scripts/summary-add.js "${dayDir}" --status # 查看缺失清单`);
console.error(` 2. 分批写入缺失摘要并合并后,重新运行: node scripts/finalize.js "${dayDir}"`);
console.error(` 在全部补齐之前,push.js 会拒绝推送该目录(避免把无摘要的文章推上去)。`);
}
} }
// Chain into push.js when config.push.enabled && autoPushAfterFinalize. // Chain into push.js when config.push.enabled && autoPushAfterFinalize.
......
...@@ -548,9 +548,11 @@ Examples: ...@@ -548,9 +548,11 @@ Examples:
// NB: spec must NOT be passed to legacy fetchPage (it would be coerced to // NB: spec must NOT be passed to legacy fetchPage (it would be coerced to
// waitMs=NaN). desiredCount is only threaded into the spec path — legacy // waitMs=NaN). desiredCount is only threaded into the spec path — legacy
// sources do a single-page fetch with no scroll/pagination, so they can't // sources do a single-page fetch with no scroll/pagination, so they can't
// top up candidates (that's a separate, larger change). // top up candidates (that's a separate, larger change). `paginate` (from
// config.json source entry) is threaded into the spec path to enable
// multi-page list harvesting when the recipe also configures pagination.
const getLinks = (desiredCount) => spec const getLinks = (desiredCount) => spec
? fetcher.fetchHomeLinks(source.homepageUrl, spec, desiredCount) ? fetcher.fetchHomeLinks(source.homepageUrl, spec, desiredCount, undefined, source.paginate)
: fetcher.fetchHomeLinks(source.homepageUrl, source.articleUrlPattern); : fetcher.fetchHomeLinks(source.homepageUrl, source.articleUrlPattern);
const getPage = (url) => spec const getPage = (url) => spec
? fetcher.fetchPage(url, spec) ? fetcher.fetchPage(url, spec)
......
...@@ -16,9 +16,14 @@ ...@@ -16,9 +16,14 @@
* Prerequisites: * Prerequisites:
* - <dayDir>/registry.json exists (written by harvest.js, patched by finalize.js) * - <dayDir>/registry.json exists (written by harvest.js, patched by finalize.js)
* - <dayDir>/.manifest.json and .summaries.json are ABSENT (= finalize done) * - <dayDir>/.manifest.json and .summaries.json are ABSENT (= finalize done)
* - every pending article is summary-complete: titleZh is a real Chinese
* title (not the __SUMMARY_*__ sentinel) and its summary doc exists with
* non-empty sections — otherwise the push is blocked (exit 1) with
* instructions to补充摘要 and re-finalize
* - config.json has a `push` section (endpoint, appId, appKeyId, appSecret) * - config.json has a `push` section (endpoint, appId, appKeyId, appSecret)
* *
* Exit codes: 0 success (or all-pending-resolved), 1 usage/IO error, * Exit codes: 0 success (or all-pending-resolved), 1 usage/IO error
* (incl. readiness gate: articles missing summaries),
* 2 partial failure (some articles still error after push). * 2 partial failure (some articles still error after push).
*/ */
...@@ -143,7 +148,7 @@ function fileSha256(filePath) { ...@@ -143,7 +148,7 @@ function fileSha256(filePath) {
// Locate the Chinese summary doc for an article. finalize.js (current) writes // Locate the Chinese summary doc for an article. finalize.js (current) writes
// `summary/<titleZh>.md`; older layouts placed `<titleZh>.md` directly in the // `summary/<titleZh>.md`; older layouts placed `<titleZh>.md` directly in the
// day dir. Try the subdir first, then fall back to the root. Returns '' if // day dir. Try the subdir first, then fall back to the root. Returns '' if
// neither exists (the article is still pushed, just with an empty summary). // neither exists — readiness gating refuses to push such an article.
function readSummaryDoc(dayDir, titleZh) { function readSummaryDoc(dayDir, titleZh) {
if (!titleZh) return ''; if (!titleZh) return '';
const sub = path.join(dayDir, 'summary', `${titleZh}.md`); const sub = path.join(dayDir, 'summary', `${titleZh}.md`);
...@@ -153,6 +158,36 @@ function readSummaryDoc(dayDir, titleZh) { ...@@ -153,6 +158,36 @@ function readSummaryDoc(dayDir, titleZh) {
return ''; return '';
} }
// harvest.js writes `__SUMMARY_<slug>__` as the placeholder titleZh; finalize.js
// replaces it only when a summary exists. A registry entry still carrying it
// was never summarized and must not be pushed.
const SENTINEL_RE = /^__SUMMARY_.*__$/;
// Validate that an article is actually ready to be pushed: finalized (titleZh
// is a real Chinese title, not the harvest sentinel) and its summary doc
// exists with at least one non-empty section. Returns a list of problem
// strings (empty = ready).
function articleReadinessProblems(dayDir, a) {
const problems = [];
const titleZh = (a.titleZh || '').trim();
if (!titleZh) {
problems.push('titleZh 为空(finalize 未完成)');
} else if (SENTINEL_RE.test(titleZh)) {
problems.push('titleZh 仍为 __SUMMARY_*__ 占位符(未生成中文摘要,finalize 未覆盖该文章)');
}
const summaryMd = readSummaryDoc(dayDir, a.titleZh);
if (!summaryMd) {
problems.push('摘要文档缺失(summary/<titleZh>.md 不存在)');
} else {
const { body } = parseFrontmatter(summaryMd);
const s = parseSummarySections(body);
if (!s.overview && !s.background && !s.keyQuote && !s.impact) {
problems.push('摘要文档四段内容全为空');
}
}
return problems;
}
// Convert the 原文.md body (markdown AFTER frontmatter) into clean HTML for // Convert the 原文.md body (markdown AFTER frontmatter) into clean HTML for
// `originalDoc.body`. Strips Obsidian-specific syntax (callouts, wikilinks, // `originalDoc.body`. Strips Obsidian-specific syntax (callouts, wikilinks,
// `![[assets/...]]` embeds) and produces semantic HTML: <h2> for the English // `![[assets/...]]` embeds) and produces semantic HTML: <h2> for the English
...@@ -457,6 +492,31 @@ async function pushDayDir(dayDir, cfg, opts) { ...@@ -457,6 +492,31 @@ async function pushDayDir(dayDir, cfg, opts) {
} }
console.error(` Pending: ${pending.length} (already synced: ${articles.length - pending.length})`); console.error(` Pending: ${pending.length} (already synced: ${articles.length - pending.length})`);
// 2b. Readiness gate — refuse to push articles that were never summarized.
// finalize.js is non-fatal about missing summaries (it leaves the sentinel
// titleZh in place), so an incomplete summaries pass must be caught here:
// pushing would upload sentinel titles with empty summary docs. Block the
// whole run (nothing is uploaded) and tell the agent to补齐摘要 first.
const unready = pending
.map(a => ({ a, problems: articleReadinessProblems(dayDir, a) }))
.filter(x => x.problems.length > 0);
if (unready.length > 0) {
console.error(`\n🚫 推送已阻止 — ${unready.length}/${pending.length} 篇待推送文章尚未生成中文摘要(未完成 finalize)`);
for (const { a, problems } of unready) {
console.error(` ✗ ${a.titleEn || a.url}`);
for (const p of problems) console.error(` · ${p}`);
console.error(` ${a.url}`);
}
console.error(`\n 未推送任何文章(含配图上传)。请先为上述文章补充中文摘要:`);
console.error(` 1. 查看缺失清单: node scripts/summary-add.js "${dayDir}" --status`);
console.error(` (若 .manifest.json 已被上一次 finalize 删除,该命令同样能从 registry.json 列出缺失文章)`);
console.error(` 2. 将缺失文章的摘要写入 ${path.join(dayDir, '.summaries.json')}`);
console.error(` 格式: JSON 数组, 每条 { url, titleZh, category, tags[], summaryZh }; 多于 ~8 篇时分批走 summary-add.js`);
console.error(` 3. 重新 finalize: node scripts/finalize.js "${dayDir}"`);
console.error(` 最后重新推送: node scripts/push.js "${dayDir}"`);
process.exit(1);
}
// 3. Phase 1 — upload images for pending articles (local sha256 dedup). // 3. Phase 1 — upload images for pending articles (local sha256 dedup).
const shaCache = new Map(); // sha256 -> assetId (in-process dedup) const shaCache = new Map(); // sha256 -> assetId (in-process dedup)
// Seed cache from existing synced entries? Not needed: asset ownership is // Seed cache from existing synced entries? Not needed: asset ownership is
...@@ -652,16 +712,21 @@ Arguments: ...@@ -652,16 +712,21 @@ Arguments:
Prerequisites: Prerequisites:
- <dayDir>/registry.json exists (written by harvest.js, patched by finalize.js) - <dayDir>/registry.json exists (written by harvest.js, patched by finalize.js)
- <dayDir>/.manifest.json and .summaries.json are ABSENT (= finalize done) - <dayDir>/.manifest.json and .summaries.json are ABSENT (= finalize done)
- every pending article is summary-complete (real titleZh + non-empty summary
doc) — articles still missing summaries block the push with a "补充摘要"
recovery guide
- config.json has a \`push\` section with endpoint + appId/appKeyId/appSecret - config.json has a \`push\` section with endpoint + appId/appKeyId/appSecret
What it does: What it does:
1. Phase 1: upload each article's images to MinIO (POST /assets, sha256 deduped) 1. Readiness gate: refuse to push (exit 1, nothing uploaded) if any pending
2. Phase 2: batch-push articles (POST /articles, ≤10 per batch, idempotent) article still has a __SUMMARY_*__ sentinel titleZh or no summary content
3. Writes <dayDir>/.syncstate.json ledger — re-running resumes from failures 2. Phase 1: upload each article's images to MinIO (POST /assets, sha256 deduped)
3. Phase 2: batch-push articles (POST /articles, ≤10 per batch, idempotent)
4. Writes <dayDir>/.syncstate.json ledger — re-running resumes from failures
Exit codes: Exit codes:
0 success (all articles synced, or --status printed) 0 success (all articles synced, or --status printed)
1 usage/IO error (missing files, not finalized, config missing) 1 usage/IO error (missing files, not finalized, articles missing summaries, config missing)
2 partial/failed push (some articles still in error state) 2 partial/failed push (some articles still in error state)
Example: Example:
......
...@@ -80,7 +80,7 @@ switch (command) { ...@@ -80,7 +80,7 @@ switch (command) {
const source = config.sources.find(s => s.id === id); const source = config.sources.find(s => s.id === id);
if (!source) { console.error(`Source "${id}" not found`); process.exit(1); } if (!source) { console.error(`Source "${id}" not found`); process.exit(1); }
const editableFields = ['name', 'homepageUrl', 'method', 'articleUrlPattern', 'vaultFolder', 'enabled']; const editableFields = ['name', 'homepageUrl', 'method', 'articleUrlPattern', 'vaultFolder', 'enabled', 'paginate'];
let changed = false; let changed = false;
for (const field of editableFields) { for (const field of editableFields) {
...@@ -161,7 +161,9 @@ switch (command) { ...@@ -161,7 +161,9 @@ switch (command) {
} }
// Upsert the config entry. recipe sources don't carry articleUrlPattern // Upsert the config entry. recipe sources don't carry articleUrlPattern
// (it's legacy-direct-only and ignored when a recipe exists). // (it's legacy-direct-only and ignored when a recipe exists). `paginate`
// defaults to false on install but is intentionally NOT overwritten on
// re-install of an existing source (preserves the operator's setting).
const entry = { const entry = {
id, id,
name: src.name || id, name: src.name || id,
...@@ -169,10 +171,14 @@ switch (command) { ...@@ -169,10 +171,14 @@ switch (command) {
method: src.method || 'browser', method: src.method || 'browser',
vaultFolder: src.vaultFolder || src.name || id, vaultFolder: src.vaultFolder || src.name || id,
enabled: true, enabled: true,
concurrency: 1 // serial by default; user can raise to enable multi-tab fetching concurrency: 1, // serial by default; user can raise to enable multi-tab fetching
paginate: false // pagination switch — set true in config.json to enable multi-page harvesting
}; };
if (existing) { if (existing) {
// Preserve user-set paginate across re-installs; entry.paginate stays as-is for new sources.
const savedPaginate = existing.paginate;
Object.assign(existing, entry); Object.assign(existing, entry);
if (savedPaginate === true) existing.paginate = true;
} else { } else {
config.sources.push(entry); config.sources.push(entry);
} }
...@@ -201,6 +207,7 @@ switch (command) { ...@@ -201,6 +207,7 @@ switch (command) {
Usage: Usage:
node source-manage.js list node source-manage.js list
node source-manage.js edit <id> --name "New Name" --url https://... --method browser node source-manage.js edit <id> --name "New Name" --url https://... --method browser
node source-manage.js edit <id> --paginate true # enable multi-page list harvesting
node source-manage.js enable <id> node source-manage.js enable <id>
node source-manage.js disable <id> node source-manage.js disable <id>
node source-manage.js install <pkg.nhsource.json> [--force] # install a shared bundle (recipe + config entry) node source-manage.js install <pkg.nhsource.json> [--force] # install a shared bundle (recipe + config entry)
......
...@@ -82,19 +82,36 @@ function isValidEntry(s) { ...@@ -82,19 +82,36 @@ function isValidEntry(s) {
} }
// ── --status: diff .manifest.json urls against .summaries.json urls ────────── // ── --status: diff .manifest.json urls against .summaries.json urls ──────────
const SENTINEL_RE = /^__SUMMARY_.*__$/;
function status(dayDir) { function status(dayDir) {
const manifestPath = path.join(dayDir, MANIFEST); const manifestPath = path.join(dayDir, MANIFEST);
const summariesPath = path.join(dayDir, SUMMARIES); const summariesPath = path.join(dayDir, SUMMARIES);
if (!fs.existsSync(manifestPath)) { let manifestArticles;
console.error(`✗ ${MANIFEST} not found in ${dayDir}`); if (fs.existsSync(manifestPath)) {
console.error(` (already finalized? manifest is removed by finalize.js)`); const manifest = loadJson(manifestPath, MANIFEST);
process.exit(1); manifestArticles = Array.isArray(manifest.articles) ? manifest.articles : [];
} else {
// A previous finalize already removed .manifest.json while coverage was
// incomplete — derive the pending list from registry.json's sentinel
// entries (same fallback as finalize.js) so the resume loop keeps a
// "what's missing" view: write summaries → re-run finalize.js.
const regPath = path.join(dayDir, 'registry.json');
const noManifest = () => {
console.error(`✗ ${MANIFEST} not found in ${dayDir}`);
console.error(` (already finalized? manifest is removed by finalize.js)`);
process.exit(1);
};
if (!fs.existsSync(regPath)) noManifest();
const registry = loadJson(regPath, 'registry.json');
manifestArticles = (Array.isArray(registry.articles) ? registry.articles : [])
.filter(a => a && a.url && SENTINEL_RE.test((a.titleZh || '').trim()));
if (manifestArticles.length === 0) noManifest();
console.error(`ℹ️ .manifest.json absent — listing registry.json entries still missing summaries`);
} }
const manifest = loadJson(manifestPath, MANIFEST);
const manifestArticles = Array.isArray(manifest.articles) ? manifest.articles : [];
let summaries = []; let summaries = [];
if (fs.existsSync(summariesPath)) { if (fs.existsSync(summariesPath)) {
try { summaries = JSON.parse(fs.readFileSync(summariesPath, 'utf8')); } try { summaries = JSON.parse(fs.readFileSync(summariesPath, 'utf8')); }
......
...@@ -104,6 +104,7 @@ node scripts/preview.js --url <homepageUrl> --recipe helpers/<id>.md --count 3 ...@@ -104,6 +104,7 @@ node scripts/preview.js --url <homepageUrl> --recipe helpers/<id>.md --count 3
| Candidates are nav/section pages | Tighten `urlPattern`, or set `validateArticle: true`. | | Candidates are nav/section pages | Tighten `urlPattern`, or set `validateArticle: true`. |
| Empty body | Fix `bodySelectors`. | | Empty body | Fix `bodySelectors`. |
| 0 candidates, all nav | `listSelector: null` is wrong — scope to card containers. | | 0 candidates, all nav | `listSelector: null` is wrong — scope to card containers. |
| List spans multiple pages | Configure pagination (`nextSelector` or `pageUrlTemplate`) so the harvester can follow pages. See `recipe-guide.md` → *Pagination*. Note: pagination only takes effect when the harvester's `config.json` also sets `paginate: true` for the source. |
Re-run preview after each recipe edit. Deep troubleshooting: `references/recipe-guide.md` → *Troubleshooting recipes*. Re-run preview after each recipe edit. Deep troubleshooting: `references/recipe-guide.md` → *Troubleshooting recipes*.
......
{
"format": "news-harvester-source",
"version": 1,
"exportedAt": "2026-08-14T02:47:27.378Z",
"source": {
"id": "bloomberg-technology",
"name": "彭博社科技",
"homepageUrl": "https://www.bloomberg.com/technology",
"method": "browser",
"vaultFolder": "Bloomberg/Technology"
},
"files": {
"helpers/bloomberg-technology.md": "---\nsource: bloomberg-technology\n\n# ── 发现阶段 ──\n# Bloomberg Technology 首页文章分散在多个 LineupContent* 卡片容器 +\n# styles_itemTextContainer 列表中,用复合 listSelector 一次性圈中。\nlistSelector: '[class*=\"LineupContent\"], [class*=\"itemTextContainer\"]'\nlinkSelector: 'a[href]'\n\n# 文章 URL 形如 /news/articles/YYYY-MM-DD/<slug> 或 /news/features/YYYY-MM-DD/<slug>\n# 视频 /news/videos/ 用 excludeUrlPattern 排除。\nurlPattern: 'bloomberg\\.com/news/(articles|features)/.+'\nexcludeUrlPattern: 'bloomberg\\.com/news/videos/'\n\nscrollSteps: 3\nscrollWaitMs: 2000\nwaitMs: 20000\nmaxLinks: 30\n\n# ── 提取阶段 ──\n# Bloomberg 正文段落 class 为 ArticleBodyText_articleBodyContent__xxxxx,\n# 直接命中正文 <p>,避免误抓 newsletter / MostRead 等侧边栏内容。\n# body-content p 作为 fallback(整个正文容器内的所有 <p>)。\nbodySelectors:\n - '[class*=\"ArticleBodyText\"]'\n - '[data-testid=\"paragraph\"]'\n - '[class*=\"body-content\"] p'\n - '[class*=\"articleBodyContent\"] p'\n - '[class*=\"ArticleBody\"] p'\n - '[class*=\"article-body\"] p'\n - 'article p'\nbodyMinParagraphs: 3\n\ndateSelector: 'time[datetime]'\n# Bloomberg 作者署名在 byline 容器中,JSON-LD 作为最终 fallback。\nauthorSelector: '[rel=\"author\"], [class*=\"author\"] a, [class*=\"byline\"] a, [class*=\"Byline\"] a'\n\nimageSelector: 'img'\nimageMinWidth: 200\nexcludeImage:\n - logo\n - icon\n - avatar\n - profile\n\n# Bloomberg 文末常见署名/订阅推广噪声 + 推荐文章标题噪声。\nnoisePrefixes:\n - 'Sign up for'\n - 'Subscribe to'\n - 'Read more:'\n - 'Read More:'\n - 'Most Read from Bloomberg'\n - 'This article was produced'\n - '©2026 Bloomberg'\n - 'Bloomberg may send me'\n - 'By submitting my information'\n - 'By continuing, I agree'\n\ntitleStripSuffix: ' - Bloomberg$'\nmaxBodyChars: 15000\n---\n# bloomberg-technology 抓取备忘\n\n- 首页 `https://www.bloomberg.com/technology` 文章卡片分布在多个 `LineupContent*`\n 容器(4Up/3Up/2Up/Basic)以及 `styles_itemTextContainer` 列表中,inspect 实测\n 首屏约 22 篇文章链接 + 4 个视频。用复合 `listSelector` 圈中所有卡片。\n- `urlPattern` 限定 `/news/(articles|features)/.+`,天然排除 `/technology/ai` 等\n section 页与 `/markets` 等导航;`excludeUrlPattern` 排除 `/news/videos/`。\n- 反爬/付费墙:Bloomberg 对 headless 有反爬,正文可能因未登录被截断。\n 如 preview 出现 charCount<500 或 CAPTCHA(exit 70),需\n `node scripts/ensure-chromium.js --login --url https://www.bloomberg.com/technology`\n 在窗口浏览器登录后退出,持久化 cookie 到 userDataDir。\n- 正文段落 class 为 `ArticleBodyText_articleBodyContent__xxxxx`(哈希后缀会变),\n 用 `[class*=\"ArticleBodyText\"]` 精确命中,避免误抓 newsletter / MostRead 推荐区块。\n waitMs=20000 给足页面渲染时间,防止正文容器加载不全时 fallback 兜底抓到推荐标题。\n- 作者署名在 byline 容器中,authorSelector 已覆盖 `[class*=\"byline\"] a`;\n 若 CSS 选择器全部未命中,fetcher-helper 会自动 fallback 到 JSON-LD `author` 字段。\n- 验证: node scripts/preview.js --url https://www.bloomberg.com/technology \\\n --recipe helpers/bloomberg-technology.md --count 3\n",
"helpers/bloomberg-technology.fixtures.json": "{\n \"articles\": [\n \"https://www.bloomberg.com/news/articles/2026-08-13/trump-enlists-private-sector-to-boost-cyber-offensive-arsenal\",\n \"https://www.bloomberg.com/news/articles/2026-08-13/anthropic-said-in-talks-to-buy-ai-startup-decart-for-6-billion\",\n \"https://www.bloomberg.com/news/features/2026-08-12/thousands-of-india-workers-are-helping-ai-firms-train-robots-to-replace-them\"\n ],\n \"sections\": [\n \"https://www.bloomberg.com/technology/ai\",\n \"https://www.bloomberg.com/technology/big-tech\"\n ],\n \"minLinks\": 10\n}\n"
}
}
\ No newline at end of file
...@@ -27,6 +27,17 @@ scrollWaitMs: 2000 # 每次滚动后等待毫秒(等新内容渲染) ...@@ -27,6 +27,17 @@ scrollWaitMs: 2000 # 每次滚动后等待毫秒(等新内容渲染)
waitMs: 15000 # 页面初始加载后等待毫秒(等 JS/React 渲染完) waitMs: 15000 # 页面初始加载后等待毫秒(等 JS/React 渲染完)
maxLinks: 30 # 链接上限 maxLinks: 30 # 链接上限
# ── 翻页(多页列表采集,可选)──
# 当站点把文章列表分成多页时(如 /news/page/2 或"下一页"按钮),配置翻页让
# harvest 在单页候选不足 count 时自动跨页采集。两种模式互斥,最多配一个。
# 注意:recipe 配了翻页字段后,还需在 config.json 该 source 下设 paginate: true
# 才会生效(两级开关)。详见 recipe-guide.md → Pagination 小节。
# nextSelector: 'a.next-page' # 点击"下一页"元素翻页;元素不存在 = 最后一页
# pageUrlTemplate: 'https://site.com/news/page/{page}' # URL 模板翻页,{page} 替换为页码
# pageStart: 2 # pageUrlTemplate 起始页码(≥2,nextSelector 模式忽略)
# maxPages: 5 # 最多翻页数(不含首页,上限 20)
# pageWaitMs: 2000 # 翻页后等待毫秒(null 时回退到 scrollWaitMs)
# ── 提取阶段:在文章页定位正文/元数据 ── # ── 提取阶段:在文章页定位正文/元数据 ──
# bodySelectors: 正文容器选择器,按序尝试,首个满足 ≥ bodyMinParagraphs 的即用。 # bodySelectors: 正文容器选择器,按序尝试,首个满足 ≥ bodyMinParagraphs 的即用。
bodySelectors: bodySelectors:
...@@ -81,6 +92,7 @@ maxBodyChars: 15000 ...@@ -81,6 +92,7 @@ maxBodyChars: 15000
# #
# 发现阶段: listSelector / linkSelector / urlPattern(必填) / excludeUrlPattern # 发现阶段: listSelector / linkSelector / urlPattern(必填) / excludeUrlPattern
# scrollSteps / scrollWaitMs / waitMs / maxLinks # scrollSteps / scrollWaitMs / waitMs / maxLinks
# nextSelector / pageUrlTemplate / pageStart / maxPages / pageWaitMs (翻页,可选)
# 提取阶段: bodySelectors / bodyMinParagraphs / dateSelector / authorSelector # 提取阶段: bodySelectors / bodyMinParagraphs / dateSelector / authorSelector
# imageSelector / imageMinWidth / excludeImage / noisePrefixes # imageSelector / imageMinWidth / excludeImage / noisePrefixes
# titleStripSuffix / maxBodyChars # titleStripSuffix / maxBodyChars
......
...@@ -51,6 +51,11 @@ All fields optional except `urlPattern`; defaults reproduce legacy behaviour. ...@@ -51,6 +51,11 @@ All fields optional except `urlPattern`; defaults reproduce legacy behaviour.
| `scrollWaitMs` | `1500` | Pause after each scroll | | `scrollWaitMs` | `1500` | Pause after each scroll |
| `waitMs` | `12000` | Initial render wait after load | | `waitMs` | `12000` | Initial render wait after load |
| `maxLinks` | `30` | Link cap | | `maxLinks` | `30` | Link cap |
| `nextSelector` | `null` | CSS selector for the "next page" element; click it to advance to the next page of the article list. `null` = pagination off. See [Pagination](#pagination-multi-page-list-harvesting) below. |
| `pageUrlTemplate` | `null` | URL template with a `{page}` placeholder, e.g. `https://site.com/news/page/{page}`. `null` = off. Mutually exclusive with `nextSelector`. |
| `pageStart` | `2` | First page number substituted into `pageUrlTemplate` (ignored when using `nextSelector`). Must be an integer ≥ 2. |
| `maxPages` | `5` | Hard cap on pages fetched beyond the first (cap: 20). |
| `pageWaitMs` | `null` | Wait after each page transition; falls back to `scrollWaitMs` when `null`. |
**Extraction** (single article body + metadata): **Extraction** (single article body + metadata):
...@@ -107,6 +112,69 @@ Set `validateArticle: true` when article and section/category URLs on the site * ...@@ -107,6 +112,69 @@ Set `validateArticle: true` when article and section/category URLs on the site *
Don't set it when the URL alone cleanly separates articles from sections (e.g. Reuters' `/world/<slug>/<date>/` vs. `/world/`) — a tight `urlPattern` is cheaper than a per-candidate page fetch. `validateArticle` defaults to `false`, so existing recipes are unaffected. Don't set it when the URL alone cleanly separates articles from sections (e.g. Reuters' `/world/<slug>/<date>/` vs. `/world/`) — a tight `urlPattern` is cheaper than a per-candidate page fetch. `validateArticle` defaults to `false`, so existing recipes are unaffected.
## Pagination (multi-page list harvesting)
Some sites split their article list across multiple pages (e.g. `/news/page/2`, `?page=3`, or a "下一页" / "Next" button). By default the harvester only collects from the **first page** (plus whatever lazy-load scrolling reveals). When a harvest requests more articles than one page yields, it prints `⚠️ stream may be exhausted` and stops short. Pagination lets the fetcher follow those pages to top up candidates.
### Two-level switch
Pagination requires **both** levels to be on:
1. **Recipe level** (this file): configure `nextSelector` or `pageUrlTemplate` to describe *how* to paginate.
2. **Config level** (`config.json` source entry): set `paginate: true` to enable the feature for that source.
| config `paginate` | recipe has pagination fields | Behaviour |
|-------------------|------------------------------|----------|
| `true` | yes | ✅ Multi-page harvesting |
| `true` | no | ⚠️ Warning: "已启用 paginate 但 recipe 未配置翻页字段" — falls back to single-page, no error |
| `false` / unset | yes | ℹ️ Info: "recipe 支持翻页但 config.paginate 未启用" — single-page, fields ignored |
| `false` / unset | no | ✅ Legacy single-page behaviour |
Install sets `paginate: false` by default. Enable it per source:
```bash
node scripts/source-manage.js edit <id> --paginate true
```
### Choosing a mode
The two modes are **mutually exclusive** — configuring both throws an error at load time.
| Mode | When to use | How it works |
|------|-------------|--------------|
| `nextSelector` | The page has a clickable "Next" / "下一页" / page-number element | Clicks the element, waits for navigation or SPA render, then collects. Stops when the selector is absent (last page). |
| `pageUrlTemplate` | Page URLs follow a numeric pattern (`/news/page/2`, `?page=2`) | Constructs the next URL by substituting `{page}`, navigates directly. Stops after `maxPages` or when consecutive pages yield no new candidates. |
**Pagination vs. scroll**: use `scrollSteps` / `maxScrollSteps` for **SPA infinite-scroll** lists (same URL, cards lazy-load on scroll). Use pagination for **traditional multi-page** lists (different URLs or explicit "Next" buttons). They compose: each page is still scrolled with `scrollSteps` to reveal its lazy-loaded cards before advancing.
### Recipe examples
**`nextSelector` mode** (click "Next" button):
```yaml
listSelector: 'div.article-list'
linkSelector: 'a[href]'
urlPattern: 'example\\.com/article/.+'
nextSelector: 'a.next-page' # clicked to advance; absence = last page
maxPages: 5
pageWaitMs: 2000
```
**`pageUrlTemplate` mode** (URL pattern):
```yaml
listSelector: 'ul.news-items'
linkSelector: 'a[href]'
urlPattern: 'yna\\.co\\.kr/view/.+'
pageUrlTemplate: 'https://en.yna.co.kr/industry/all/{page}'
pageStart: 2 # first fetch beyond the homepage is /industry/all/2
maxPages: 10
```
### Notes
- **Preview never paginates** — `--preview` stays single-page for speed. Only archive harvest (where `desiredCount > 0`) follows pagination.
- **Cross-page dedup**: candidates are deduped by URL across all pages. `buildDiscoveryExpr` strips query/hash, so the same article appearing on two pages won't be double-counted.
- **Anti-bot**: each page transition re-runs the CAPTCHA/login-wall probe. A challenge mid-pagination exits with code 70 (same as a first-page block).
- **Stale exit**: if 2 consecutive pages add zero new candidates, pagination stops early (the list is exhausted). `maxPages` is a hard ceiling, not a target.
## Fixtures format ## Fixtures format
`recipe-test.js` requires `helpers/<id>.fixtures.json`: `recipe-test.js` requires `helpers/<id>.fixtures.json`:
......
...@@ -403,6 +403,55 @@ async function scrollAndCollect(ws, { ...@@ -403,6 +403,55 @@ async function scrollAndCollect(ws, {
return merged; return merged;
} }
// ── Click + wait (pagination) ────────────────────────────────────────────────
// Click the element matching `selector` and wait for the resulting navigation
// or SPA render. Used by fetchHomeLinks to follow "next page" links across
// multi-page article lists.
//
// Returns true if the element was found and clicked (caller should continue
// paginating), false if the selector matched nothing (treat as "last page").
//
// The wait mirrors navigateAndWait: a full-page navigation fires
// Page.loadEventFired (raced against a 25s timeout), while an SPA transition
// that swaps content without a navigation falls through to `sleep(waitMs)` so
// the new cards have time to render before the next collect pass.
async function clickAndWait(ws, selector, waitMs = 2000) {
await cdpSend(ws, 'Page.enable');
// Attach the load listener BEFORE clicking, since loadEventFired can fire
// before the cdpEval response arrives. The handler is hoisted so we can
// detach it early if the selector matches nothing (no click → no navigation).
let loaded = false;
let resolveLoad;
const handler = (raw) => {
const msg = JSON.parse(raw.toString());
if (msg.method === 'Page.loadEventFired' && !loaded) {
loaded = true;
ws.removeListener('message', handler);
resolveLoad();
}
};
ws.on('message', handler);
const loadPromise = new Promise(resolve => { resolveLoad = resolve; });
const expr = `(function(){
var el = document.querySelector(${JSON.stringify(selector)});
if (!el) return false;
el.click();
return true;
})()`;
const clicked = await cdpEval(ws, expr);
if (!clicked) {
ws.removeListener('message', handler);
return false;
}
await Promise.race([loadPromise, sleep(25000)]);
await sleep(waitMs);
return true;
}
module.exports = { module.exports = {
sleep, sleep,
resolveCdpPort, resolveCdpPort,
...@@ -412,6 +461,7 @@ module.exports = { ...@@ -412,6 +461,7 @@ module.exports = {
cdpEval, cdpEval,
navigateAndWait, navigateAndWait,
scrollAndCollect, scrollAndCollect,
clickAndWait,
// Browser-level target management (multi-tab concurrency) // Browser-level target management (multi-tab concurrency)
getBrowserWs, getBrowserWs,
countPageTabs, countPageTabs,
......
...@@ -21,7 +21,7 @@ const fs = require('fs'); ...@@ -21,7 +21,7 @@ const fs = require('fs');
const path = require('path'); const path = require('path');
const yaml = require('js-yaml'); const yaml = require('js-yaml');
const { openSession, navigateAndWait, scrollAndCollect, cdpEval } = require('./cdp-client'); const { openSession, navigateAndWait, scrollAndCollect, clickAndWait, cdpEval } = require('./cdp-client');
const { evaluateBrowserBlock } = require('./block-check'); const { evaluateBrowserBlock } = require('./block-check');
// ── Spec defaults ──────────────────────────────────────────────────────────── // ── Spec defaults ────────────────────────────────────────────────────────────
...@@ -41,6 +41,19 @@ const DEFAULTS = { ...@@ -41,6 +41,19 @@ const DEFAULTS = {
waitMs: 12000, // initial render wait after load waitMs: 12000, // initial render wait after load
maxLinks: 30, maxLinks: 30,
// ── Pagination (multi-page list harvesting) ──
// Two mutually exclusive modes: nextSelector (click an element to advance)
// or pageUrlTemplate (construct the next-page URL by substituting {page}).
// Both are null by default → legacy single-page behaviour. Even when
// configured here, pagination only takes effect if the source's config.json
// entry has paginate: true (a two-level switch so operators can enable per
// source without touching the recipe).
nextSelector: null, // CSS selector for the "next page" element; click it to advance. null = off
pageUrlTemplate: null, // URL template with {page} placeholder, e.g. "https://site.com/news/page/{page}". null = off
pageStart: 2, // first page number substituted into pageUrlTemplate (ignored for nextSelector)
maxPages: 5, // hard cap on pages fetched beyond the first
pageWaitMs: null, // wait after each page transition; falls back to scrollWaitMs when null
// ── Extraction ── // ── Extraction ──
bodySelectors: [ bodySelectors: [
'[data-testid^="paragraph"]', '[data-testid^="paragraph"]',
...@@ -89,6 +102,16 @@ function loadSpec(helperPath) { ...@@ -89,6 +102,16 @@ function loadSpec(helperPath) {
if (!spec.urlPattern) { if (!spec.urlPattern) {
throw new Error(`helper.md ${path.basename(helperPath)} missing required 'urlPattern'`); throw new Error(`helper.md ${path.basename(helperPath)} missing required 'urlPattern'`);
} }
// Pagination: nextSelector and pageUrlTemplate are mutually exclusive.
if (spec.nextSelector && spec.pageUrlTemplate) {
throw new Error(`helper.md ${path.basename(helperPath)}: nextSelector and pageUrlTemplate are mutually exclusive — configure at most one`);
}
if (spec.pageUrlTemplate && (!Number.isInteger(spec.pageStart) || spec.pageStart < 2)) {
throw new Error(`helper.md ${path.basename(helperPath)}: pageStart must be an integer ≥ 2 when pageUrlTemplate is set`);
}
if (spec.maxPages > 20) {
throw new Error(`helper.md ${path.basename(helperPath)}: maxPages cap is 20 (got ${spec.maxPages})`);
}
return spec; return spec;
} }
...@@ -157,13 +180,19 @@ function buildExtractExpr(spec) { ...@@ -157,13 +180,19 @@ function buildExtractExpr(spec) {
const bodyMinChars = spec.bodyMinChars || 500; const bodyMinChars = spec.bodyMinChars || 500;
return `(function(){ return `(function(){
var title = document.title.replace(new RegExp(${titleStripRe}, 'g'), '').trim(); var title = document.title.replace(new RegExp(${titleStripRe}, 'g'), '').trim();
// JSON-LD fallback: many sites (Nikkei, Reuters) embed datePublished/author // JSON-LD fallback: many sites (Nikkei, Reuters, Yonhap) embed
// only in <script type="application/ld+json">, not in <time> or visible // datePublished/author only in <script type="application/ld+json">, not in
// author links. Parse it once up front so the CSS-selector results can // <time> or visible author links. Sites with multiple blocks (e.g. a
// fall back to it below. // NewsMediaOrganization block before the NewsArticle block) need all blocks
// scanned — ldGet already recurses into arrays. Parse every block up front
// so the CSS-selector results can fall back to it below.
var ld = null; var ld = null;
var ldEl = document.querySelector('script[type="application/ld+json"]'); var ldEls = document.querySelectorAll('script[type="application/ld+json"]');
if (ldEl) { try { ld = JSON.parse(ldEl.textContent); } catch (e) {} } if (ldEls.length) {
var blocks = [];
ldEls.forEach(function(el){ try { blocks.push(JSON.parse(el.textContent)); } catch(e){} });
ld = blocks.length === 1 ? blocks[0] : blocks;
}
function ldGet(obj, key) { function ldGet(obj, key) {
if (!obj) return null; if (!obj) return null;
if (Array.isArray(obj)) { for (var i=0;i<obj.length;i++){ var v=ldGet(obj[i],key); if(v) return v; } return null; } if (Array.isArray(obj)) { for (var i=0;i<obj.length;i++){ var v=ldGet(obj[i],key); if(v) return v; } return null; }
...@@ -266,28 +295,113 @@ function buildExtractExpr(spec) { ...@@ -266,28 +295,113 @@ function buildExtractExpr(spec) {
// avoids tab churn in serial workflows like preview.js / recipe-test.js where // avoids tab churn in serial workflows like preview.js / recipe-test.js where
// many pages are fetched one after another. When omitted, a new session is // many pages are fetched one after another. When omitted, a new session is
// created and closed in the finally block (harvester behavior). // created and closed in the finally block (harvester behavior).
async function fetchHomeLinks(url, spec, desiredCount, session) { //
// Optional `paginate` (harvest.js): when true AND desiredCount > 0 AND the
// recipe configures a pagination mode (nextSelector or pageUrlTemplate),
// follow multi-page article lists — scrolling each page with fixed scrollSteps,
// then advancing to the next page, cross-page deduping until desiredCount is
// met, maxPages is exhausted, or consecutive pages yield no new candidates.
// Two-level switch: config.json `paginate` gates the feature per source; the
// recipe fields describe *how* to paginate. A warning is printed when one
// level is on but the other is missing.
async function fetchHomeLinks(url, spec, desiredCount, session, paginate) {
const ownSession = !session; const ownSession = !session;
const { ws, close } = session || await openSession(); const { ws, close } = session || await openSession();
try { try {
await navigateAndWait(ws, url, spec.waitMs); await navigateAndWait(ws, url, spec.waitMs);
await evaluateBrowserBlock(ws); await evaluateBrowserBlock(ws);
const opts = {
waitMs: spec.scrollWaitMs, const recipePaginate = !!(spec.nextSelector || spec.pageUrlTemplate);
collectExpr: buildDiscoveryExpr(spec), const archiveMode = desiredCount && desiredCount > 0;
dedupKey: item => item.url const enablePaginate = archiveMode && paginate && recipePaginate;
};
if (desiredCount && desiredCount > 0) { // Two-level switch diagnostics: warn when only one level is configured.
opts.targetCount = desiredCount; if (archiveMode && paginate && !recipePaginate) {
opts.maxSteps = spec.maxScrollSteps; console.error(' ⚠️ config 已启用 paginate 但 recipe 未配置翻页字段(nextSelector / pageUrlTemplate),按单页采集。');
} else { }
opts.steps = spec.scrollSteps; if (archiveMode && !paginate && recipePaginate) {
console.error(' ℹ️ recipe 支持翻页但 config.paginate 未启用;如需跨页采集,在 config.json 该 source 下设 paginate: true。');
}
if (!enablePaginate) {
// Legacy single-page path (scroll lazy-load only).
const opts = {
waitMs: spec.scrollWaitMs,
collectExpr: buildDiscoveryExpr(spec),
dedupKey: item => item.url
};
if (archiveMode) {
opts.targetCount = desiredCount;
opts.maxSteps = spec.maxScrollSteps;
} else {
opts.steps = spec.scrollSteps;
}
const items = await scrollAndCollect(ws, opts);
const cap = archiveMode
? Math.max(spec.maxLinks || 30, desiredCount)
: (spec.maxLinks || 30);
return items.map(i => i.url).slice(0, cap);
} }
const items = await scrollAndCollect(ws, opts);
const cap = (desiredCount && desiredCount > 0) // ── Pagination mode ──
? Math.max(spec.maxLinks || 30, desiredCount) // Each page is scrolled with fixed scrollSteps (not target-count mode —
: (spec.maxLinks || 30); // the outer desiredCount drives the cross-page loop, not per-page scroll).
return items.map(i => i.url).slice(0, cap); // Candidates are deduped across pages via the seen Set; buildDiscoveryExpr
// already strips query/hash so the same article on different pages dedupes.
const seen = new Set();
const merged = [];
let stalePages = 0;
const pageWaitMs = spec.pageWaitMs || spec.scrollWaitMs;
const maxPages = Math.min(spec.maxPages || 5, 20);
for (let pageIdx = 0; pageIdx <= maxPages; pageIdx++) {
// pageIdx 0 is the already-loaded first page; subsequent iterations
// navigate/click to the next page before collecting.
if (pageIdx > 0) {
if (spec.nextSelector) {
const advanced = await clickAndWait(ws, spec.nextSelector, pageWaitMs);
if (!advanced) {
console.error(` 📄 已到最后一页(第 ${pageIdx + 1} 页无"下一页"元素),停止翻页。`);
break;
}
await evaluateBrowserBlock(ws);
} else {
// pageUrlTemplate: pageIdx=1 → pageStart, pageIdx=2 → pageStart+1, …
const nextPageNum = spec.pageStart + (pageIdx - 1);
const nextUrl = spec.pageUrlTemplate.replace('{page}', String(nextPageNum));
await navigateAndWait(ws, nextUrl, spec.waitMs);
await evaluateBrowserBlock(ws);
}
}
const pageOpts = {
steps: spec.scrollSteps,
waitMs: spec.scrollWaitMs,
collectExpr: buildDiscoveryExpr(spec),
dedupKey: item => item.url
};
const pageItems = await scrollAndCollect(ws, pageOpts);
const before = merged.length;
for (const item of pageItems) {
if (!seen.has(item.url)) {
seen.add(item.url);
merged.push(item);
}
}
const added = merged.length - before;
console.error(` 📄 第 ${pageIdx + 1} 页:+${added} 条候选(累计 ${merged.length})`);
if (merged.length >= desiredCount) break;
stalePages = (added === 0) ? stalePages + 1 : 0;
if (stalePages >= 2) {
console.error(' 📄 连续 2 页无新增候选,流可能已干涸,停止翻页。');
break;
}
}
const cap = Math.max(spec.maxLinks || 30, desiredCount);
return merged.map(i => i.url).slice(0, cap);
} finally { } finally {
if (ownSession) await close(); if (ownSession) await close();
} }
......
...@@ -216,6 +216,13 @@ excludeUrlPattern: null ...@@ -216,6 +216,13 @@ excludeUrlPattern: null
scrollSteps: ${patternStr ? 0 : 3} scrollSteps: ${patternStr ? 0 : 3}
scrollWaitMs: 2000 scrollWaitMs: 2000
waitMs: 15000 waitMs: 15000
# ── 翻页(可选)── 站点把文章列表分多页时启用。两种模式互斥,最多配一个。
# 配了翻页字段后还需在 config.json 该 source 下设 paginate: true 才生效。
# nextSelector: 'a.next' # 点击"下一页"元素翻页
# pageUrlTemplate: null # 或用 URL 模板 'https://site.com/news/page/{page}'
# pageStart: 2 # pageUrlTemplate 起始页码(≥2)
# maxPages: 5 # 最多翻页数(不含首页)
# pageWaitMs: null # 翻页后等待(null 回退到 scrollWaitMs)
bodySelectors: bodySelectors:
- '[data-testid^="paragraph"]' - '[data-testid^="paragraph"]'
- '[class*="articleBodyContent"] p' - '[class*="articleBodyContent"] p'
......
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment