# DailyPaper-CCUS **Repository Path**: bwang31/daily-paper-ccus ## Basic Information - **Project Name**: DailyPaper-CCUS - **Description**: 基于https://github.com/huangkiki/dailypaper-skills定制的个人paper抓取 - **Primary Language**: Python - **License**: Apache-2.0 - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-06-06 - **Last Updated**: 2026-07-12 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # daily-paper-ccus `daily-paper-ccus` 是给 Codex 使用的文献工作流技能包。日常使用时,用户主要通过自然语言 prompt 触发 skill,而不是手动记命令或跑脚本。 它面向 CCUS / 地质能源 / 地下工程方向,覆盖从“今日论文推荐”到“进 Zotero、补 PDF、精读论文、生成 Markdown 笔记、同步回 Zotero”的一整条链路。 现在的推荐用法是 **Zotero-first**:Zotero 是主沉淀中心,Markdown 笔记只是便于 Codex 精读和版本管理的工作副本。 ## How To Use This Repo 安装完成后,你在 Codex 里直接说需求即可: ```text 今日论文推荐 读一下这篇论文 10.1016/j.physrep.2025.03.001 帮我下载这篇 ScienceDirect 论文 PDF,如果需要机构验证就让我操作 从 Zotero 的 PaperClaw 抓取文献列表 批量读一下 Zotero 里 PaperClaw 分类下的论文 把这篇精读笔记同步到 Zotero ``` Codex 会根据 prompt 自动调用对应 skill。命令行脚本主要用于首次安装、环境配置、调试和维护。 ## Skill Trigger Map | 你对 Codex 说 | 触发能力 | 结果 | | --- | --- | --- | | `今日论文推荐` | `daily-papers` 总入口 | 搜索、评分、去重、写今日推荐 | | `过去3天论文推荐` | `daily-papers` 总入口 | 扩大时间窗做推荐 | | `过去一周论文推荐` | `daily-papers` 总入口 | 最近一周推荐 | | `用 geologic_hydrogen 做今日论文推荐` | `daily-papers` + `geologic_hydrogen` profile | 地质氢方向推荐 | | `用 particle_transport 做今日论文推荐` | `daily-papers` + `fractured_reservoir_particle_transport` profile | 裂隙颗粒/支撑剂运移方向推荐 | | `论文点评` / `跑一下论文点评` | `daily-papers-review` | 读取富化结果,写今日推荐工作副本,并写入 Zotero PaperClaw | | `批量笔记` / `跑一下论文笔记` | `daily-papers-notes` | 从今日推荐中批量生成论文笔记,并尝试同步到 Zotero | | `读一下这篇论文 ` | `paper-reader` | 读取论文,生成结构化 Markdown 笔记,并同步到 Zotero child note | | `快速看一下这篇论文 ` | `paper-reader` | 快速摘要、贡献、方法和结论 | | `批判性分析这篇论文 ` | `paper-reader` | 更偏审稿视角的优缺点评估 | | `读一下 Zotero 里的 XXX` | `paper-reader` + Zotero 查询 | 从 Zotero 找条目和 PDF 后读论文 | | `从 Zotero 的 PaperClaw 抓取文献列表` | Zotero 查询 | 列出分类下文献、PDF 状态和已同步笔记状态 | | `批量读一下 Zotero 里 PaperClaw 分类下的论文` | `paper-reader` 批量模式 | 递归读取 Zotero 分类论文,跳过已有合格笔记 | | `把这篇精读笔记同步到 Zotero` | Zotero Web API note sync | 将本地 Markdown 精读笔记写成 Zotero 条目的 child note | | `帮我下载这篇论文 PDF` | shared PDF resolver | 依次尝试 OA、Elsevier API、InstSci 机构浏览器、HTML 重组 PDF | | `View PDF 被验证挡住,用网页正文重组 PDF` | `authorized_article_pdf.py` fallback | 用已授权 HTML 页面打印成 PDF | | `帮我生成 PubMed / Web of Science 检索式` | `daily-papers-fetch` SCI planner | 生成 broad / moderate / precise 检索式 | | `帮我检查这些论文元数据完整性` | `daily-papers-fetch` metadata audit | 审计 DOI、标题、摘要、作者、PMID/PMCID 等 | `particle_transport` 是 `fractured_reservoir_particle_transport` 的常用短别名。 ## Prompt Recipes ### Daily Recommendation ```text 今日论文推荐 ``` ```text 过去3天论文推荐,优先看高质量综述和实验/数值结合的文章 ``` ```text 用 geologic_hydrogen 做今日论文推荐,并把值得跟进的论文写入 Zotero PaperClaw ``` ```text 用 particle_transport 做今日论文推荐,重点关注裂隙粗糙度、支撑剂运移、颗粒堵塞和传热耦合 ``` ```text 今日论文推荐,推荐结果和必读论文的精读笔记都同步到 Zotero ``` ### Paper Reading ```text 读一下这篇论文 10.1016/j.physrep.2025.03.001 ``` ```text 快速看一下这篇论文 C:\Users\pc\Downloads\paper.pdf ``` ```text 批判性分析这篇论文 https://arxiv.org/abs/2509.24527,重点看方法假设和可复现性 ``` ```text 读一下 Zotero 里的 Diffusion Policy ``` ```text 读一下 Zotero 里的 A comprehensive review of particle-laden flows modeling,生成 md 精读笔记并同步回 Zotero ``` ```text 批量读一下 Zotero 里 PaperClaw 分类下的论文,已有合格 Zotero 笔记的跳过 ``` ### PDF Acquisition ```text 帮我下载这篇论文 PDF:10.1016/j.physrep.2025.03.001 ``` ```text 这篇 ScienceDirect 我有机构权限,如果需要浏览器人工验证就让我操作 ``` ```text View PDF 总是被机器验证挡住,请用已登录浏览器里的网页正文重组一个 PDF ``` ```text 这篇只有 Zotero 条目没有附件,帮我补 PDF 并更新本地路径 ``` PDF 获取策略按顺序尝试: 1. OA/direct:arXiv、Unpaywall、OpenAlex OA、Semantic Scholar、DOAJ、EuropePMC、PMC、Crossref 页面。 2. Elsevier API:使用本地 Elsevier API key 和 InstSci 兼容路线。 3. InstSci visible browser:复用你的机构登录,遇到 SSO/2FA/CAPTCHA 时由你手动完成。 4. Authorized HTML-to-PDF:如果 HTML 全文已经能读,就把网页正文、图、表、公式打印成 PDF,避开 `View PDF` 按钮触发的额外验证。 5. Zotero/manual:如果仍然需要人工操作,保存条目或手动拖入附件。 ### Search Planning And Audit ```text 帮我围绕 geologic_hydrogen 生成 PubMed 和 Web of Science 检索式 ``` ```text 帮我为裂隙颗粒运移方向做 broad / moderate / precise 三档 WOS 检索式 ``` ```text 检查今天推荐论文的 DOI、摘要、作者和期刊信息是否完整 ``` ### Zotero As The Main Memory ```text 从 Zotero 的 PaperClaw 抓取文献列表,告诉我哪些有 PDF、哪些已经有 Codex 精读笔记 ``` ```text 把今天新生成的 md 精读笔记同步到 Zotero 对应条目 ``` ```text 把这篇 PDF 精读成 Markdown,并把笔记作为 Zotero child note 保存 ``` ```text 把今日推荐写入 Zotero PaperClaw;如果 Web API 不可写,就尝试 Zotero Connector ``` ## What This Repo Contains - `daily-papers`: 每日论文推荐总入口。 - `daily-papers-fetch`: 搜索、SCI 检索式规划、元数据审计。 - `daily-papers-review`: 推荐页写作、Zotero PaperClaw 写入。 - `daily-papers-notes`: 从推荐列表批量生成论文笔记,并同步到 Zotero。 - `paper-reader`: 单篇/批量读论文,支持 PDF、DOI、URL、arXiv、Zotero,并把精读笔记写回 Zotero。 - shared PDF resolver: 集成 `scansci-pdf`、InstSci、Elsevier API、HTML-to-PDF fallback。 Bundled profiles: - `geologic_hydrogen`: stimulated geologic hydrogen, natural/white hydrogen, serpentinization, ultramafic reservoirs, stimulation, reactive transport, THMC risk, monitoring, commercialization readiness. - `fractured_reservoir_particle_transport`: particle/proppant/fines transport in fractured geo-energy reservoirs, turbulent fracture flow, rough fractures, slurry flow, fracture conductivity. ## First-Time Setup Prompts 推荐让 Codex 帮你完成安装,而不是自己记命令: ```text 请帮我把 daily-paper-ccus 安装到当前 Codex,Zotero 在 D:\Research\Zotero,本地 Markdown 工作目录用 D:\Research\DailyPaperNotes ``` 首次部署时请先让用户输入或粘贴必要密钥,不要把密钥写入仓库。通常需要: - Elsevier API key:用于 ScienceDirect / InstSci 路线。 - Zotero Web API key:用于写入 PaperClaw 和同步 child note。 - Unpaywall/OA email:默认可用 `bin.wang@cup.edu.cn`。 ```text 请帮我配置 PDF 获取工具。Elsevier API key 我稍后粘贴;Unpaywall email 用 bin.wang@cup.edu.cn;机构是 China University of Petroleum Beijing / 中国石油大学(北京) ``` ```text 请检查 scansci-pdf、InstSci、Playwright、Zotero Connector 是否都能用 ``` Codex 背后会使用这些脚本: ```powershell python scripts/install_pack.py --force ` --workspace "D:\Research\DailyPaperNotes" ` --zotero-db "D:\Research\Zotero\zotero.sqlite" ` --zotero-storage "D:\Research\Zotero\storage" python scripts/setup_scansci_pdf.py ``` 这些命令是维护参考;日常使用仍然以 prompt 为主。 ## Local Configuration Machine-specific paths belong in: ```text %USERPROFILE%\.codex\skills\_shared\user-config.local.json ``` PDF and institution-access settings live locally in: ```text %USERPROFILE%\.scansci-pdf\config.json %USERPROFILE%\.instsci\config.json ``` Default Unpaywall/OA metadata email: ```text bin.wang@cup.edu.cn ``` Do not commit API keys, institution passwords, cookies, login tokens, or browser profiles. ## API Readiness First ### Local-Code-Only Policy 论文检索、推荐、精读、批次分类和 Zotero 归档默认只从本仓库已有 skill 与脚本触发。不要临时改用网页搜索、Computer Use、一次性脚本或另一套自创流程。若本地代码尚未覆盖需求,先向用户说明功能缺口;未经明确同意不新增替代实现。若 readiness check 显示必需 API key、邮箱、collection key 或写权限缺失,立即停止对应步骤并请用户提供,不做静默降级或绕过认证。 On a new computer, or whenever the Python/Zotero/PDF environment changes, run the readiness check before starting a prompt-driven workflow: ```powershell python scripts\check_api_readiness.py --mode full ``` For narrower checks: ```powershell python scripts\check_api_readiness.py --mode pdf python scripts\check_api_readiness.py --mode zotero-write python scripts\check_api_readiness.py --mode notes ``` If a required API key is missing, Codex should pause and ask the user to provide it before continuing. API keys stay in environment variables or local user config files; they are never committed to this repository. ## Zotero PaperClaw Packaged Zotero target: - Zotero group: `TopJournalPapers` - Group id: `6384936` - Collection: `PaperClaw` - Collection key: `AWSWCB9V` Zotero Web API key stays in the environment: ```powershell $env:ZOTERO_API_KEY = "..." ``` If Web API write access is unavailable, the review workflow can use Zotero Desktop Connector when Zotero is running and the correct collection is selected. Zotero browser Connector can save items and sometimes PDFs from an already-authorized browser session, but it cannot bypass publisher verification. ## Zotero Note Sync 精读笔记会先生成本地 Markdown 工作副本,然后通过 Zotero Web API 写入对应 Zotero 条目的 child note。后续再次同步会更新同一条 Codex note,不会无限重复创建。 Batch sync for a prepared reading set: ```powershell $env:ZOTERO_API_KEY = "..." python skills\daily-papers\zotero_web_api.py sync-batch ` "runs\paperclaw-priority\ultramafic-h2-co2-2026\zotero-batch.json" Remove-Item Env:ZOTERO_API_KEY ``` `sync-batch` creates or reuses a PaperClaw child collection, creates/updates Zotero items, uploads local PDF attachments, and syncs Markdown reading notes as child notes. DailyPaper 默认使用这条批次归档路径。每次推荐或精读触发都会在 PaperClaw 下创建或复用一个子分类,默认名称为 `YYYY-MM-DD {topic} 精读`: ```powershell python skills\daily-papers\prepare_zotero_batch.py ` runs\my-topic\papers.resolved.json ` runs\my-topic\zotero-batch.json ` --date 2026-07-12 --topic "particle laden flow" ` --notes-dir runs\my-topic\notes python skills\daily-papers\zotero_web_api.py sync-batch ` runs\my-topic\zotero-batch.json ``` Zotero Desktop Connector 只保留为“用户已经明确选中某个现有分类”时的题录应急写入方式。它不负责创建每次运行的归档子分类,主流程也不通过 UI 自动化切换 Zotero 分类。 Zotero item tags are intentionally short Chinese workflow labels, normally one tag per paper, such as `评分5蛇纹石化`, `评分5碳矿产氢`, or `评分4PDF待补`. Older workflow tags like `/unread`, `daily-paper`, `PaperClaw`, and `daily-papers` are cleaned during sync while user-created tags are preserved. Zotero 条目的“其他”(`extra`)字段只保存简洁来源值,例如 `codex_search`。评分、PDF 路径、解析器状态和 `/unread` 等工作流信息不再写入该字段;同步旧条目时会自动清理这些历史噪声行。 Prompt: ```text 把 C:\Users\pc\Downloads\paper-note.md 同步到 Zotero,DOI 是 10.xxxx/xxxxx ``` Behind the scenes, Codex may run: ```powershell python skills\daily-papers\zotero_web_api.py sync-note ` "C:\Users\pc\Downloads\paper-note.md" ` --doi "10.xxxx/xxxxx" ``` 也可以用 Zotero item key 精确指定: ```powershell python skills\daily-papers\zotero_web_api.py sync-note ` "C:\Users\pc\Downloads\paper-note.md" ` --item-key "PXW99EKT" ``` 如果没有 `ZOTERO_API_KEY` 或 key 没有 group 写权限,Codex 会保留本地 Markdown,并在最终汇报里明确说明 Zotero note sync 未执行。 ## Authorized HTML-to-PDF Fallback This fallback exists for the case you described: you have legitimate institutional access and the article HTML is readable, but publisher `View PDF` triggers extra machine verification. Prompt: ```text View PDF 被验证挡住,但网页全文能看。请用网页正文重组一个 PDF。 ``` Behind the scenes, Codex may run: ```powershell python scripts/authorized_article_pdf.py ` "10.1016/j.physrep.2025.03.001" ` --output "runs/html-pdf/particle_laden_flows_review.pdf" ` --headed-auth ` --auth-wait 300 ``` Important boundaries: - It does not solve CAPTCHA. - It does not bypass paywalls. - It only exports pages already accessible in your browser session. - If the browser page is blocked by verification/CPE but the Elsevier API returns sufficient XML, the fallback rebuilds an HTML article from API full text, figures, tables, formulas, and references, then prints it to PDF. - If both browser HTML and API XML are incomplete, the tool fails clearly instead of producing a fake complete paper. - The reconstructed PDF may not match publisher pagination, but it preserves rendered text, figures, tables, and equations when those assets are available through authorized HTML or API XML. ## Quality Rules MDPI and Frontiers articles are blocked by default: - Domains: `mdpi.com`, `frontiersin.org`, `frontiersin.org.cn` - Publishers: `MDPI`, `Frontiers Media`, `Frontiers` - DOI prefixes: `10.3390`, `10.3389` These filters apply across structured sources and Codex Search. ## Repository Layout ```text daily-paper-ccus/ skills/ # installable Codex skills user-data/ # profiles, seed papers, example config scripts/ # install, export, validate, PDF helpers ARCHITECTURE.md README.md ``` ## Maintenance Commands Use these only when maintaining the repo itself. Export installed profiles and seed data back into the repo: ```powershell python scripts/export_user_data.py --from-skills "%USERPROFILE%\.codex\skills" --output user-data ``` Validate before publishing: ```powershell python scripts/validate_pack.py --quick-validate "%USERPROFILE%\.codex\skills\.system\skill-creator\scripts\quick_validate.py" ``` Publish target: ```text https://gitee.com/bwang31/daily-paper-ccus ``` Use credentials only for the push operation. Do not save credentials in scripts, configs, or notes. ## License Apache-2.0. See `LICENSE`.