Skip to content
FormatExtensionsWhat you get
PDF.pdfText layer, form fields and FreeText annotations; scanned or unreadable pages are OCR'd.
DOCX.docxHeadings, lists, tables, links, charts and SmartArt; pictures are OCR'd.
PPTX.pptxSlides in order with titles, text, tables, charts, SmartArt and speaker notes; pictures are OCR'd.
XLSX.xlsxEvery sheet as a Markdown table, plus charts, shapes and pictures.
HTML.html, .htmSanitized headings, lists, tables and links, text inside SVG; embedded data: pictures are OCR'd.

The converter works cheapest first: (1) what the file contains as structure — text, tables, charts, SmartArt, text inside SVG and EMF/WMF pictures — with no OCR; (2) OCR of single pictures; (3) OCR of rendered PDF pages, only for pages that cannot be read otherwise.

FormatOne billed page
PDFOne PDF page
DOCX3,000 characters of text
PPTXOne slide
XLSX3,000 characters of cell text per sheet
HTML3,000 characters of visible text

Counts are rounded up, with a minimum of one page (one per sheet for XLSX). HTML counts 3,000 characters of visible text after whitespace is collapsed; DOCX counts 3,000 characters of the document’s text (the body of word/document.xml; headers, footers, footnotes and pictures excluded), so the count is known before converting.

  • 1 credit per page read from the file.
  • 3 credits per OCR page with ocr_model=auto; a specific model has its own price (OCR models).
  • A page counts as an OCR page only when OCR succeeded on it. For DOCX and HTML, a page chunk with at least one picture read by OCR is an OCR page; for PPTX a slide, and for XLSX a sheet (charged once per sheet).
  • Nothing is charged when a conversion fails.

Example: a DOCX with 4,500 characters of text and 2 pictures in the first 3,000 characters is 2 pages, 1 of them OCR: 1 × 1 + 1 × 3 = 4 credits.

The most a file can cost is reserved before converting (you need that much available); afterwards you are charged only for the pages actually converted. Use /v1/quote to see both numbers in advance. Credits are added by the operator: md@tlelabs.com.

The home page converts files for free without signing in: 10 conversions per IP address per UTC day, files up to 10 MB and 50 pages. It never uses OCR: everything read from the file itself is included, pictures keep their placeholder and scanned pages are not read. The free converter is only on the home page, not in the API.

LimitValueOver the limit
File size50 MB413 file_too_large
PDF pages500422 page_limit_exceeded
PPTX slides200422 page_limit_exceeded
XLSX sheets50422 page_limit_exceeded
Pictures to OCR per HTML, DOCX, PPTX or XLSX file (when ocr is not off)50422 image_limit_exceeded
PDF that needs OCR20 MB413 file_too_large

Page, slide, sheet and picture limits are checked before anything is reserved. Pictures on PDF pages with text are limited to 50 per document: the rest are skipped with a warning instead of an error. Very large XML parts or Markdown output inside DOCX, PPTX and XLSX files also get 413 file_too_large, without a charge.