# OCR models

> Choose the OCR model per request with ocr_model; auto tries a chain of models and falls back on failure.

The `ocr_model` field of [`/v1/convert`](https://md.tlelabs.com/docs/convert/) and [`/v1/quote`](https://md.tlelabs.com/docs/quote/) picks the model that reads pictures and rendered PDF pages. Ids are `provider/model` and match exactly (case-sensitive).

## auto (default)

Sending nothing is the same as `ocr_model=auto`. Each kind of work has a chain of models; when a model fails (HTTP error, rate limit, timeout, malformed answer) the next one is tried. Every model in the chains costs 3 credits per OCR page.

- Pictures: `novita/deepseek-ocr-2 → openparser/paddleocr-vl-1.6 → paddle/serving`
- Rendered PDF pages: `openparser/paddleocr-vl-1.6 → paddle/serving`

Models the server has not configured are skipped. [`GET /v1/health`](https://md.tlelabs.com/docs/health-and-status/) lists the chains and the models available right now.

## A specific model

A specific id uses only that model, with no fallback, and OCR pages cost that model’s price (reported in `usage.ocr_page_credits`). A model that cannot do a kind of work is treated as if no model were available for it: with a picture-only model, PDF pages that need rendering keep their text layer or stay empty with a warning.

An unknown id, or one whose provider is not configured on the server, gets `400 invalid_request` with the accepted ids in `error.details.allowed`, before anything is reserved. `ocr=off` still validates `ocr_model` but runs no OCR.

| Model                                   | Reads               | Credits per OCR page |
| --------------------------------------- | ------------------- | -------------------- |
| `novita/deepseek-ocr-2`                 | pictures            | 3                    |
| `novita/paddleocr-vl`                   | pictures            | 3                    |
| `openparser/paddleocr-vl-1.6`           | pictures, PDF pages | 3                    |
| `openparser/mistral-ocr-3`              | pictures, PDF pages | 3                    |
| `openparser/azure-di-read`              | pictures, PDF pages | 3                    |
| `openparser/google-docai-ocr`           | pictures, PDF pages | 3                    |
| `openparser/aws-textract-detect`        | pictures, PDF pages | 3                    |
| `openparser/mistral-ocr-4`              | pictures, PDF pages | 6                    |
| `openparser/aws-textract-layout`        | pictures, PDF pages | 6                    |
| `openparser/azure-di-layout`            | pictures, PDF pages | 15                   |
| `openparser/aws-textract-tables-layout` | pictures, PDF pages | 23                   |
| `paddle/serving`                        | pictures, PDF pages | 3                    |

A page read without OCR always costs 1 credit, whichever model you choose. See [Formats, pages and pricing](https://md.tlelabs.com/docs/formats-and-pricing/).

## What is sent to the model

Only what the converter cannot read from the file goes to an OCR provider: single pictures, and PDFs of at most 10 pages that contain only the pages needing a rendered read. When a PDF cannot be split (encrypted, unreadable or rejected by the splitter), the original file is sent once instead and only the pages that need a rendered read are used from the result. The OCR text is put back in place in the Markdown. All OCR for one conversion shares a 150-second budget; work that cannot start in time becomes a warning.
