# docmore_sdk **Repository Path**: fireae/docmore_sdk ## Basic Information - **Project Name**: docmore_sdk - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-11 - **Last Updated**: 2026-09-11 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # OCR Parse SDK (Stainless) 使用 [Stainless](https://www.stainless.com/) 生成的 OCR 统一识别接口 SDK 项目配置与完整实现。 项目基于 OpenAPI 3.1 规范 + stainless.yml 配置,可生成 TypeScript 和 Python 两种语言的 SDK。 ## 项目结构 ``` sdk/ ├── .stainless/ │ ├── openapi.yaml # OpenAPI 3.1 规范(/v1/parse 接口) │ ├── stainless.yml # Stainless SDK 生成配置 │ └── workspace.json # 工作区配置 ├── ocr-parse-typescript/ # TypeScript SDK 实现 │ ├── src/ │ │ ├── index.ts # 主入口 (OCRParse 类) │ │ ├── types.ts # 类型定义 │ │ ├── core.ts # HTTP 客户端 + 错误处理 + 重试 │ │ ├── streaming.ts # SSE 流式处理 │ │ └── resources/ │ │ ├── parse.ts # /parse 资源(create / stream) │ │ └── models.ts # /models 资源(list / retrieve) │ ├── package.json │ └── tsconfig.json └── ocr-parse-python/ # Python SDK 实现 ├── ocr_parse/ │ ├── __init__.py │ ├── client.py # 主入口 (OCRParse 类) │ ├── types.py # 类型定义(dataclass) │ ├── _client.py # HTTP 客户端 + 重试 │ ├── _exceptions.py # 错误类层级 │ ├── _streaming.py # SSE 流式处理 │ └── resources/ │ ├── parse.py # /parse 资源 │ └── models.py # /models 资源 └── pyproject.toml ``` ## Stainless 配置说明 ### 核心文件 - **`openapi.yaml`** — OpenAPI 3.1 规范,完整定义了 `/v1/parse` 和 `/v1/models` 两个端点,包括: - 三种输入格式(OpenAI messages / Mistral document / 原生 content 数组) - 请求选项(output_format、bbox、pages、表格格式等) - 结构化输出(response_schema) - 流式输出(SSE) - 完整的响应结构(pages、bounding_boxes、images、tables、usage 等) - **`stainless.yml`** — Stainless SDK 生成配置,包括: - 资源层级(`parse` / `models`) - 流式配置(SSE events) - 错误类层级(9 种错误类型) - 重试配置(可重试状态码) - 多目标输出(TypeScript + Python) - 环境配置(production / staging) - **`workspace.json`** — Stainless CLI 工作区配置 ### 生成 SDK 安装 Stainless CLI 后,在 `.stainless` 目录运行: ```bash # 登录 stl auth login # 初始化项目 stl init # 预览构建 stl preview # 监听模式 stl preview --watch ``` 生成的 SDK 将输出到 `ocr-parse-typescript/` 和 `ocr-parse-python/` 目录。 > 注:本项目已提供了手动编写的完整 SDK 实现,代码风格与 Stainless 生成的一致,可直接使用。 ## TypeScript SDK 使用 ### 安装 ```bash npm install @ocr-parse/sdk # 或 yarn add @ocr-parse/sdk ``` ### 基本使用 ```typescript import OCRParse from '@ocr-parse/sdk'; const client = new OCRParse({ apiKey: 'your-api-key', // baseURL: 'https://staging.example.com/v1', // 可选,自定义端点 // maxRetries: 2, // 可选,重试次数 // timeout: 600000, // 可选,超时(毫秒) }); // 解析图片 const result = await client.parse.create({ model: 'deepseek-ocr-2', content: [ { type: 'text', text: '提取这张发票的全部内容' }, { type: 'image_url', image_url: { url: 'https://example.com/invoice.png', detail: 'high' } }, ], options: { output_format: 'markdown', include_bounding_boxes: true, }, }); console.log(result.content); console.log(`共 ${result.pages.length} 页`); ``` ### 解析 PDF ```typescript const result = await client.parsePdf('https://example.com/report.pdf', { pages: '0-4', output_format: 'markdown', include_bounding_boxes: true, }); for (const page of result.pages) { console.log(`=== Page ${page.index} ===`); console.log(page.markdown); if (page.bounding_boxes) { for (const bb of page.bounding_boxes) { console.log(` [${bb.type}] ${bb.content}`); } } } ``` ### 流式处理 ```typescript const stream = await client.parse.stream({ model: 'deepseek-ocr-2', content: [ { type: 'text', text: '解析这份 PDF' }, { type: 'file_url', file_url: { url: 'https://example.com/report.pdf' } }, ], }); for await (const event of stream) { if (event.event === 'page.complete') { console.log(`第 ${event.data.index} 页解析完成`); console.log(event.data.markdown); } else if (event.event === 'done') { console.log('全部完成!'); } } ``` ### OpenAI SDK 兼容调用 ```typescript // 使用 OpenAI 格式的 messages 数组 const result = await client.parse.create({ model: 'deepseek-ocr-2', messages: [ { role: 'user', content: [ { type: 'text', text: '识别这份文档' }, { type: 'image_url', image_url: { url: 'https://example.com/doc.png' } }, ]}, ], }); ``` ### Mistral SDK 兼容调用 ```typescript // 使用 Mistral 格式的 document 对象 const result = await client.parse.create({ model: 'deepseek-ocr-2', document: { type: 'document_url', document_url: 'https://example.com/report.pdf', }, options: { pages: '0-4', output_format: 'markdown', }, }); ``` ### 结构化提取 ```typescript const result = await client.parse.create({ model: 'deepseek-ocr-2', content: [ { type: 'text', text: '提取发票信息' }, { type: 'image_url', image_url: { url: 'https://example.com/invoice.png' } }, ], response_schema: { name: 'invoice_extraction', schema: { type: 'object', properties: { invoice_number: { type: 'string' }, total: { type: 'number' }, date: { type: 'string' }, items: { type: 'array', items: { type: 'object', properties: { name: { type: 'string' }, quantity: { type: 'number' }, price: { type: 'number' }, }, }, }, }, required: ['invoice_number', 'total'], }, strict: true, }, }); console.log(result.annotation); // { invoice_number: 'INV-2026-001', total: 2500, ... } ``` ### 获取可用模型 ```typescript const models = await client.models.list(); for (const model of models.data) { console.log(`${model.id}: ${model.description}`); console.log(` 支持语言: ${model.capabilities.languages.join(', ')}`); } ``` ## Python SDK 使用 ### 安装 ```bash pip install ocr-parse ``` ### 基本使用 ```python from ocr_parse import OCRParse client = OCRParse(api_key="your-api-key") # 解析图片 result = client.parse.create( model="deepseek-ocr-2", content=[ {"type": "text", "text": "提取这张发票的全部内容"}, {"type": "image_url", "image_url": {"url": "https://example.com/invoice.png"}}, ], options={ "output_format": "markdown", "include_bounding_boxes": True, }, ) print(result.content) print(f"共 {len(result.pages)} 页") ``` ### 便捷方法 ```python # 解析图片 result = client.parse_image("https://example.com/receipt.jpg") # 解析 PDF(指定页码范围) result = client.parse_pdf( "https://example.com/report.pdf", pages="0-4", output_format="markdown", include_bounding_boxes=True, ) for page in result.pages: print(f"=== Page {page.index} ===") print(page.markdown) ``` ### 流式处理 ```python for event in client.parse.stream( model="deepseek-ocr-2", content=[ {"type": "text", "text": "解析这份 PDF"}, {"type": "file_url", "file_url": {"url": "https://example.com/report.pdf"}}, ], ): if event["event"] == "page.complete": print(f"第 {event['data']['index']} 页解析完成") print(event["data"]["markdown"]) elif event["event"] == "done": print("全部完成!") ``` ### 上下文管理器 ```python with OCRParse(api_key="your-api-key") as client: result = client.parse_image("https://example.com/doc.png") print(result.content) ``` ### 错误处理 ```python from ocr_parse import ( BadRequestError, AuthenticationError, RateLimitError, APIError, ) try: result = client.parse.create(model="invalid-model", content=[...]) except BadRequestError as e: print(f"请求错误: {e}") print(f"参数: {e.param}") except AuthenticationError as e: print(f"认证失败: {e}") except RateLimitError as e: print(f"频率超限,请稍后重试: {e}") except APIError as e: print(f"API 错误: {e.status_code} - {e}") ``` ## API 参考 ### Endpoints | 端点 | 方法 | SDK 方法 | 说明 | |------|------|----------|------| | `/v1/parse` | POST | `client.parse.create()` | 解析文档 | | `/v1/parse` | POST (stream) | `client.parse.stream()` | 流式解析 | | `/v1/models` | GET | `client.models.list()` | 列出可用模型 | ### 支持的 OCR 模型 | 模型 ID | 别名 | |---------|------| | `paddleocr-vl` | `paddleocr-vl-latest` | | `olmocr` | `olmocr-latest` | | `chandra2-ocr` | `chandra-ocr-latest` | | `deepseek-ocr-2` | `deepseek-ocr-latest` | | `dots.ocr` | `dots-ocr-latest` | ## License MIT