# Parse PDF
`POST /v1/parse` turns one uploaded PDF into a structured document tree plus optional Markdown, HTML, plain text, and annotated PDF artifacts. For up to 10 files in one request, use [Parse Batch](/parse-batch).
Source: https://docs.docushell.com/parse-pdf
Category: Reference
## Related
- [Getting started](/getting-started.md): Review auth, idempotency, and the shared jobs flow first.
- [Parse Batch](/parse-batch.md): Submit up to 10 PDFs in one request with shared parse options.
- [RAG ingestion](/rag-ingestion.md): Chunking strategy, citation metadata, and tagged-PDF structure trees.
- [Try Parse live](/playgrounds/parse): Open the parser playground with annotated PDF overlays, extracted blocks, JSON, and live API execution.
## Artifacts
Select artifacts with `formats`. Structured JSON is always produced; only one markdown-style format is allowed per job.
| `formats` value | What it is | Use it for |
| --- | --- | --- |
| `json` | Hierarchical document tree â headings, paragraphs, lists, tables, cells, captions, images, with bounding boxes where available | Automation, indexing, citations, structured QA |
| `markdown` | Flattened text in reading order | Previews, search indexing, LLM ingestion |
| `markdown_with_html` | Markdown that retains richer inline and table markup | Content that loses meaning as plain Markdown |
| `markdown_with_images` | Image-capable Markdown; a zip when `image_output=external` | Documents where figures carry the meaning |
| `html` | Styled companion document | Rendering and human review |
| `text` | Plain text | Lightweight search and simple ingestion |
| `annotated_pdf` | Source pages overlaid with detected structure | Debugging extraction before you trust it |
| `tagged_pdf` | Generated structure tree â inference, not a PDF/UA guarantee | Accessibility review workflows |
## Extraction Controls
Tuning knobs that change *how* text is extracted. Artifact selection is separate â see the table above. The parse status payload also includes `result.metadata` so you can inspect which extraction-tuning options were applied to a completed job.
| Control | Default | Effect |
| --- | --- | --- |
| `reading_order=xycut\|off` | Layout-aware | Force or disable reading-order reconstruction. |
| `table_method=default\|cluster` | `default` | Use `cluster` when the default path misses cell groupings. |
| `use_struct_tree=true` | `false` | Prefer the PDF's native structure tags when it has usable ones. |
| `keep_line_breaks=true` | `false` | Preserve source line breaks in text-oriented output. |
| `include_header_footer=true` | `false` | Keep repeated page furniture instead of dropping it. |
| `sanitize=true` | `false` | Mask visible emails, URLs, and phone numbers in extracted output. |
| `image_output=off\|embedded\|external` | `embedded` | Image handling for image-capable formats. |
| `page_range` | All pages | Select pages, e.g. `1-3,5,9-11`. Selected pages set the credit cost. |
| `hybrid_mode=auto\|full` | Backend-gated | Requires the hybrid backend to be enabled by operations. |
| OCR for scanned PDFs | Backend-gated | Without the hybrid OCR profile, image-only PDFs return `ocr_required`. |
| Formula and chart enrichment | Backend-gated | Only when the active backend profile includes those enrichments. |
## Single-File Endpoint
- Method: `POST`
- Path: `/v1/parse`
- Auth: Bearer token required on submit, status, and artifact download requests.
- Idempotency: Server-minted `job_id` values with optional `Idempotency-Key` replay support.
- Content type: `multipart/form-data`
Submit a PDF for queued parsing and receive structured JSON plus Markdown, HTML, plain text, and annotated PDF debug output.
### Headers
| Name | Type | Required | Location | Description |
| --- | --- | --- | --- | --- |
| Authorization | Bearer | Yes | header | User-owned API key created in the DocuShell dashboard. |
| Idempotency-Key | string | No | header | Recommended for safely retrying submit requests without creating duplicate jobs. |
### Request Fields
| Name | Type | Required | Location | Description |
| --- | --- | --- | --- | --- |
| file | file | Yes | multipart | PDF upload. The gateway validates PDF magic bytes before forwarding the file. |
| file_name | string | No | multipart | Optional file name override used for storage metadata and downstream artifact names. |
| page_range | string | No | multipart | Comma-separated pages or ranges such as `1-3,5,9-11`. |
| include_header_footer | boolean | No | multipart | Set to `true` to keep header and footer content in the extracted output. Default: false |
| use_struct_tree | boolean | No | multipart | Set to `true` to prefer native tagged-PDF structure when the source document includes a usable structure tree. Default: false |
| sanitize | boolean | No | multipart | Set to `true` to mask email addresses, URLs, and phone numbers in extracted output. Default: false |
| reading_order | `xycut` \| `off` | No | multipart | Optional reading-order strategy. Omit it to keep the current default extraction behavior. |
| table_method | `default` \| `cluster` | No | multipart | Optional table-detection strategy. Omit it to keep the current default extraction behavior. |
| keep_line_breaks | boolean | No | multipart | Set to `true` to preserve source line breaks more aggressively in text-oriented output. Default: false |
| output_mode | `json` \| `both` \| `html` \| `all` | No | multipart | Backward-compatible artifact bundle selector. `json` keeps only structured JSON, `both` adds Markdown, `html` adds HTML, and `all` returns the common legacy bundle: JSON, Markdown, HTML, text, and annotated PDF. Default: both |
| formats | `json` \| `markdown` \| `html` \| `text` \| `annotated_pdf` \| `markdown_with_html` \| `markdown_with_images` \| `tagged_pdf` | No | multipart | Optional explicit artifact list. Send as repeated fields or a comma-separated value, such as `formats=json,text`. |
| hybrid_mode | `auto` \| `full` | No | multipart | Optional per-job hybrid triage override. Requires the hybrid backend to be enabled by operations. |
| image_output | `off` \| `embedded` \| `external` | No | multipart | Controls image handling for image-capable outputs. `markdown_with_images` defaults to embedded images unless `external` is requested. |
### Request Notes
- Plan limits are enforced before the job is queued. Starter keeps the 50 MB per-file cap; Pro, Growth, and Scale raise upload size, per PDF/job page limits, and concurrency as monthly credits grow.
- Set `use_struct_tree=true` when tagged PDFs should favor their native structure tree. Leave it off for the default reading-order-oriented extraction path.
- Structured JSON remains the canonical parse result and is always generated for successful jobs so status responses can keep returning `result.document`.
- `sanitize`, `reading_order`, `table_method`, `keep_line_breaks`, `hybrid_mode`, and `image_output` are extraction-tuning knobs. `output_mode` and `formats` control which companion artifacts are emitted.
- Request newer artifact types such as `markdown_with_html`, `markdown_with_images`, and `tagged_pdf` with `formats`; only one markdown-style format (`markdown`, `markdown_with_html`, or `markdown_with_images`) can be requested per job because the parse engine emits one Markdown-family file per run.
- DocuShell keeps rendering-mismatch safety filters enabled for Parse PDF output. `sanitize=true` is a separate optional control for masking visible sensitive data.
- OCR, formula extraction, and chart/image descriptions follow the active DocuShell backend profile. They are not per-request fields on the shared public API.
- Status polling stays on `/v1/jobs/:jobId`. Artifact streaming happens through the shared download route with `format=json|markdown|html|text|annotated_pdf|markdown_with_html|markdown_with_images|tagged_pdf`.
### Sample Requests
```bash
curl -X POST "https://api.docushell.com/api/v1/parse" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Idempotency-Key: parse-demo-001" \
-F "file=@./quarterly-report.pdf;type=application/pdf" \
-F "file_name=quarterly-report.pdf" \
-F "page_range=1-3" \
-F "include_header_footer=true" \
-F "use_struct_tree=true" \
-F "sanitize=true" \
-F "reading_order=xycut" \
-F "table_method=cluster" \
-F "keep_line_breaks=true" \
-F "formats=json,markdown_with_images" \
-F "image_output=embedded"
```
### Queued Response
```json
{
"job_id": "job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT",
"status": "queued",
"cost": 24,
"service": "parse-pdf",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E",
"links": {
"status": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT"
}
}
```
### Status Response
```json
{
"job_id": "job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT",
"status": "done",
"service": "parse-pdf",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E",
"result": {
"document": {
"file name": "quarterly-report.pdf",
"number of pages": 2,
"kids": [
{
"type": "heading",
"heading level": 1,
"page number": 1,
"bounding box": [72.0, 708.2, 238.6, 726.1],
"content": "Executive summary"
},
{
"type": "paragraph",
"page number": 1,
"bounding box": [72.0, 668.4, 540.0, 694.9],
"content": "Revenue rose 18% year over year across the managed-services portfolio."
},
{
"type": "list",
"page number": 1,
"bounding box": [90.0, 616.3, 380.4, 654.8],
"list items": [
{ "type": "list item", "page number": 1, "bounding box": [90.0, 638.9, 262.7, 654.8], "content": "Renewals remained above 92%." },
{ "type": "list item", "page number": 1, "bounding box": [90.0, 616.3, 380.4, 632.2], "content": "Average contract value increased in EMEA." }
]
},
{
"type": "table",
"page number": 1,
"bounding box": [72.0, 534.6, 340.0, 596.8],
"number of rows": 2,
"number of columns": 2,
"rows": [
{
"type": "table row",
"row number": 1,
"cells": [
{ "type": "table cell", "row number": 1, "column number": 1, "page number": 1, "bounding box": [72.0, 566.0, 206.0, 596.8],
"kids": [{ "type": "paragraph", "page number": 1, "bounding box": [76.0, 574.1, 112.9, 588.0], "content": "Region" }] },
{ "type": "table cell", "row number": 1, "column number": 2, "page number": 1, "bounding box": [206.0, 566.0, 340.0, 596.8],
"kids": [{ "type": "paragraph", "page number": 1, "bounding box": [210.0, 574.1, 250.4, 588.0], "content": "Growth" }] }
]
},
{
"type": "table row",
"row number": 2,
"cells": [
{ "type": "table cell", "row number": 2, "column number": 1, "page number": 1, "bounding box": [72.0, 534.6, 206.0, 566.0],
"kids": [{ "type": "paragraph", "page number": 1, "bounding box": [76.0, 542.7, 160.3, 556.6], "content": "North America" }] },
{ "type": "table cell", "row number": 2, "column number": 2, "page number": 1, "bounding box": [206.0, 534.6, 340.0, 566.0],
"kids": [{ "type": "paragraph", "page number": 1, "bounding box": [210.0, 542.7, 232.4, 556.6], "content": "21%" }] }
]
}
]
},
{
"type": "caption",
"page number": 1,
"bounding box": [72.0, 512.3, 286.5, 526.4],
"content": "Table 1. Regional growth by quarter."
}
]
},
"artifacts": {
"markdown_download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=markdown",
"json_download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=json",
"html_download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=html",
"text_download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=text",
"annotated_pdf_download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=annotated_pdf"
},
"metadata": {
"engine": "docushell_parse",
"output_mode": "all",
"include_header_footer": true,
"use_struct_tree": true,
"sanitize": true,
"reading_order": "xycut",
"table_method": "cluster",
"keep_line_breaks": true
}
},
"metrics": {
"queue_wait_ms": 214,
"duration_ms": 1789
},
"links": {
"status": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT",
"download": "/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download"
}
}
```
The JSON artifact preserves reading order and exposes the structured document tree through `number of pages` and `kids`. Each `bounding box` is `[left, bottom, right, top]` in PDF points (1/72 inch), measured from the bottom-left corner of the page.
### Artifacts
- Structured JSON is always generated. It carries the document root, `numberOfPages`, and the hierarchical `kids` array with bounding boxes where the layout provides them.
- Only one markdown-style artifact can be requested per job: `markdown`, `markdown_with_html`, or `markdown_with_images`.
- `markdown_with_images` returns a self-contained file when `image_output=embedded`, and a zip when `image_output=external`.
- `tagged_pdf` is automated structure inference for accessibility review, not a PDF/UA compliance guarantee.
### Poll And Download
- Poll `GET /v1/jobs/:jobId` until `status` becomes `done` or `failed`.
- Completed jobs expose artifact links under `result.artifacts`.
- Stream one artifact at a time from `GET /v1/jobs/:jobId/download?format=`. See [Artifact Downloads](#artifact-downloads).
### Failure Notes
- `invalid_pdf` covers invalid file types and malformed uploads rejected before the worker starts.
- `corrupt_pdf` is reserved for damaged PDFs that fail deeper validation or parser execution.
- `password_protected` is returned when the document requires a password.
- `ocr_required` is returned for scans or image-only PDFs when hybrid OCR is disabled, unavailable, or still produces too little extractable text.
- `invalid_page_range` is returned when the submitted page selector is malformed or selects no valid pages.
- `page_limit_exceeded` is returned when the requested page set is larger than the plan-specific parse cap.
- `server_busy` or `backend_unavailable` indicate temporary capacity problems. Retry with the same Idempotency-Key when safe.
- Fixes for each code are on the [Troubleshooting](/troubleshooting) page.
### Error Examples
#### Password-protected PDF â `400 password_protected`
The document cannot be parsed until it is decrypted outside the public API lane.
```json
{
"error": {
"code": "password_protected",
"message": "This PDF is password-protected and cannot be parsed without a password.",
"type": "invalid_request_error",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E"
}
}
```
#### OCR required â `400 ocr_required`
The parser could not extract text from a scan or image-only file.
```json
{
"error": {
"code": "ocr_required",
"message": "This PDF appears to require OCR before it can be parsed.",
"type": "invalid_request_error",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E"
}
}
```
#### Invalid page range â `400 invalid_page_range`
The submitted selector is malformed or does not resolve to valid pages.
```json
{
"error": {
"code": "invalid_page_range",
"message": "The requested page_range is invalid for this PDF.",
"type": "invalid_request_error",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E"
}
}
```
#### Plan page limit exceeded â `400 page_limit_exceeded`
The requested document or selected page range is larger than the active plan allows.
```json
{
"error": {
"code": "page_limit_exceeded",
"message": "Requested page range exceeds your plan limit.",
"type": "invalid_request_error",
"request_id": "req_01JX8Y62XCDNZ2BM7TBM2M9Q8E"
}
}
```
## Completion Webhooks
Use `webhook_url` and `webhook_secret` for per-request completion callbacks, or send `x-docushell-webhook-url` and its matching secret header on batch parse requests.
Receivers must validate `x-docushell-signature`, deduplicate by `x-docushell-delivery`, and finish within 10-second request timeouts. Use public HTTPS staging endpoints or approved tunnels for receiver tests.
Terminal event names include `pdf.parse.completed`, `pdf.parse.failed`, `pdf.parse.batch.completed`, `pdf.parse.batch.completed_with_failures`, `pdf.parse.batch.failed`, `resume.parse.completed`, `resume.parse.failed`, `resume.batch.completed`, `resume.batch.completed_with_failures`, and `resume.batch.failed`.
## Artifact Downloads
One artifact per request. Swap `format` for any value the job produced; binary artifacts need `--output`.
```bash
curl "https://api.docushell.com/api/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=json" \
-H "Authorization: Bearer YOUR_API_KEY"
# binary artifacts: annotated_pdf, tagged_pdf, and external-image markdown (zip)
curl "https://api.docushell.com/api/v1/jobs/job_01JX8Y5YJ2M2D8N1AQ5F7Q3KVT/download?format=annotated_pdf" \
-H "Authorization: Bearer YOUR_API_KEY" \
--output document.annotated.pdf
```
Requesting a format the job did not produce returns `400`.