Formats, pages and pricing
Formats
Section titled “Formats”| Format | Extensions | What you get |
|---|---|---|
| Text layer, form fields and FreeText annotations; scanned or unreadable pages are OCR'd. | ||
| DOCX | .docx | Headings, lists, tables, links, charts and SmartArt; pictures are OCR'd. |
| PPTX | .pptx | Slides in order with titles, text, tables, charts, SmartArt and speaker notes; pictures are OCR'd. |
| XLSX | .xlsx | Every sheet as a Markdown table, plus charts, shapes and pictures. |
| HTML | .html, .htm | Sanitized headings, lists, tables and links, text inside SVG; embedded data: pictures are OCR'd. |
The converter works cheapest first: (1) what the file contains as structure — text, tables, charts, SmartArt, text inside SVG and EMF/WMF pictures — with no OCR; (2) OCR of single pictures; (3) OCR of rendered PDF pages, only for pages that cannot be read otherwise.
| Format | One billed page |
|---|---|
| One PDF page | |
| DOCX | 3,000 characters of text |
| PPTX | One slide |
| XLSX | 3,000 characters of cell text per sheet |
| HTML | 3,000 characters of visible text |
Counts are rounded up, with a minimum of one page (one per sheet for XLSX). HTML counts 3,000 characters of visible text after whitespace is collapsed; DOCX counts 3,000 characters of the document’s text (the body of word/document.xml; headers, footers, footnotes and pictures excluded), so the count is known before converting.
Pricing
Section titled “Pricing”- 1 credit per page read from the file.
- 3 credits per OCR page with
ocr_model=auto; a specific model has its own price (OCR models). - A page counts as an OCR page only when OCR succeeded on it. For DOCX and HTML, a page chunk with at least one picture read by OCR is an OCR page; for PPTX a slide, and for XLSX a sheet (charged once per sheet).
- Nothing is charged when a conversion fails.
Example: a DOCX with 4,500 characters of text and 2 pictures in the first 3,000 characters is 2 pages, 1 of them OCR: 1 × 1 + 1 × 3 = 4 credits.
The most a file can cost is reserved before converting (you need that much available); afterwards you are charged only for the pages actually converted. Use /v1/quote to see both numbers in advance. Credits are added by the operator: md@tlelabs.com.
Free trial
Section titled “Free trial”The home page converts files for free without signing in: 10 conversions per IP address per UTC day, files up to 10 MB and 50 pages. It never uses OCR: everything read from the file itself is included, pictures keep their placeholder and scanned pages are not read. The free converter is only on the home page, not in the API.
Limits
Section titled “Limits”| Limit | Value | Over the limit |
|---|---|---|
| File size | 50 MB | 413 file_too_large |
| PDF pages | 500 | 422 page_limit_exceeded |
| PPTX slides | 200 | 422 page_limit_exceeded |
| XLSX sheets | 50 | 422 page_limit_exceeded |
Pictures to OCR per HTML, DOCX, PPTX or XLSX file (when ocr is not off) | 50 | 422 image_limit_exceeded |
| PDF that needs OCR | 20 MB | 413 file_too_large |
Page, slide, sheet and picture limits are checked before anything is reserved. Pictures on PDF pages with text are limited to 50 per document: the rest are skipped with a warning instead of an error. Very large XML parts or Markdown output inside DOCX, PPTX and XLSX files also get 413 file_too_large, without a charge.