# book-convert **Repository Path**: kpret/book-convert ## Basic Information - **Project Name**: book-convert - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-05-29 - **Last Updated**: 2026-05-29 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Book Convert PDF 转 Markdown 工具,使用 OpenAI VLM(如 GPT-4o)进行内容识别。 ## 功能特性 - 逐页读取 PDF 文档 - 使用 OpenAI VLM 识别页面内容 - 提取页眉、页脚、正文、水印等结构化信息 - 跨页上下文关联(带上一页末尾100字符) - 输出结构化 Markdown 格式 ## 技术栈 - TypeScript + Node.js - pdf-lib - PDF 解析 - canvas - 页面渲染 - axios - HTTP 请求 - OpenAI API - VLM 内容识别 ## 安装 ```bash npm install ``` ## 配置 ### 环境变量 | 变量名 | 说明 | 默认值 | | :--- | :--- | :--- | | VLM_API_KEY | OpenAI API 密钥 | - | | VLM_BASE_URL | API 地址 | https://api.openai.com | | VLM_MODEL | 使用的模型 | gpt-4o | ### 配置文件 编辑 `config/config.yaml`: ```yaml vlm: apiKey: ${VLM_API_KEY} baseUrl: https://api.openai.com timeout: 60000 maxTokens: 4096 model: gpt-4o conversion: contextLength: 100 skipWatermark: true preserveHeaderFooter: false outputFormat: markdown logging: level: info file: logs/conversion.log ``` ## 使用方法 ### 命令行接口 ```bash # 基本用法 npx book-convert convert input.pdf output.md # 带 API 密钥 npx book-convert convert input.pdf output.md --api-key your-openai-api-key # 详细模式 npx book-convert convert input.pdf output.md --verbose ``` ### 选项 | 选项 | 说明 | 默认值 | | :--- | :--- | :--- | | --api-key | OpenAI API 密钥 | - | | --context-length | 上下文长度 | 100 | | --verbose | 详细日志模式 | false | ## API 使用 ```typescript import { PDFConverter } from './src/services/PDFConverter'; const converter = new PDFConverter(); const result = await converter.convert('input.pdf', 'output.md'); if (result.success) { console.log(`转换成功!共处理 ${result.pageCount} 页`); } else { console.error(`转换失败: ${result.message}`); } ``` ## 项目结构 ``` book-convert/ ├── config/ │ └── config.yaml # 配置文件 ├── src/ │ ├── components/ # 核心组件 │ │ ├── PDFParser.ts # PDF 解析器 │ │ ├── PageRenderer.ts # 页面渲染器 │ │ ├── VLMClient.ts # VLM 客户端 │ │ ├── ContentAnalyzer.ts # 内容分析器 │ │ ├── MarkdownGenerator.ts # Markdown 生成器 │ │ └── ContextManager.ts # 上下文管理器 │ ├── services/ │ │ └── PDFConverter.ts # PDF 转换服务 │ ├── types/ │ │ └── index.ts # 类型定义 │ ├── utils/ │ │ ├── logger.ts # 日志工具 │ │ └── config.ts # 配置管理 │ └── index.ts # CLI 入口 ├── package.json ├── tsconfig.json └── README.md ``` ## 工作流程 1. **加载 PDF**: 使用 pdf-lib 读取 PDF 文件 2. **逐页处理**: 循环处理每一页 3. **页面渲染**: 将页面转换为 PNG 图像 4. **VLM 识别**: 调用 OpenAI API 分析图像内容 5. **上下文关联**: 传递上一页末尾内容作为上下文 6. **结构提取**: 识别页眉、页脚、正文、水印 7. **Markdown 生成**: 合并所有页面内容 ## 开发 ```bash # 编译项目 npm run build # 开发模式 npm run dev # 运行测试 npm run test # 代码检查 npm run lint ``` ## 许可证 MIT