Skip to content

The ocr_model field of /v1/convert and /v1/quote picks the model that reads pictures and rendered PDF pages. Ids are provider/model and match exactly (case-sensitive).

Sending nothing is the same as ocr_model=auto. Each kind of work has a chain of models; when a model fails (HTTP error, rate limit, timeout, malformed answer) the next one is tried. Every model in the chains costs 3 credits per OCR page.

  • Pictures: novita/deepseek-ocr-2 → openparser/paddleocr-vl-1.6 → paddle/serving
  • Rendered PDF pages: openparser/paddleocr-vl-1.6 → paddle/serving

Models the server has not configured are skipped. GET /v1/health lists the chains and the models available right now.

A specific id uses only that model, with no fallback, and OCR pages cost that model’s price (reported in usage.ocr_page_credits). A model that cannot do a kind of work is treated as if no model were available for it: with a picture-only model, PDF pages that need rendering keep their text layer or stay empty with a warning.

An unknown id, or one whose provider is not configured on the server, gets 400 invalid_request with the accepted ids in error.details.allowed, before anything is reserved. ocr=off still validates ocr_model but runs no OCR.

Every ocr_model id, what it can read and its price
ModelReadsCredits per OCR page
novita/deepseek-ocr-2pictures3
novita/paddleocr-vlpictures3
openparser/paddleocr-vl-1.6pictures, PDF pages3
openparser/mistral-ocr-3pictures, PDF pages3
openparser/azure-di-readpictures, PDF pages3
openparser/google-docai-ocrpictures, PDF pages3
openparser/aws-textract-detectpictures, PDF pages3
openparser/mistral-ocr-4pictures, PDF pages6
openparser/aws-textract-layoutpictures, PDF pages6
openparser/azure-di-layoutpictures, PDF pages15
openparser/aws-textract-tables-layoutpictures, PDF pages23
paddle/servingpictures, PDF pages3

A page read without OCR always costs 1 credit, whichever model you choose. See Formats, pages and pricing.

Only what the converter cannot read from the file goes to an OCR provider: single pictures, and PDFs of at most 10 pages that contain only the pages needing a rendered read. When a PDF cannot be split (encrypted, unreadable or rejected by the splitter), the original file is sent once instead and only the pages that need a rendered read are used from the result. The OCR text is put back in place in the Markdown. All OCR for one conversion shares a 150-second budget; work that cannot start in time becomes a warning.