# mdworker documentation > mdworker is an HTTP API from TLE Labs that converts HTML, DOCX, PDF, PPTX and XLSX files to clean Markdown, with OCR for scans and pictures, billed in credits per page. --- # Docs > mdworker converts HTML, DOCX, PDF, PPTX and XLSX files to clean Markdown over one HTTP request, with OCR for scans and pictures, billed in credits per page. mdworker is an HTTP API that turns documents into Markdown. Send a file to `POST /v1/convert` with an API key and get Markdown back: headings, lists, tables and links come from the file itself, and an OCR model reads scans and pictures. You pay credits for the pages converted and nothing when a conversion fails. - Base URL: `https://api.md.tlelabs.com` - Formats: PDF, DOCX, PPTX, XLSX, HTML, detected from the file’s bytes. - Price: 1 credit per page read from the file, 3 credits per OCR page with `ocr_model=auto`. - Files up to 50 MB. Service health: [status.md.tlelabs.com](https://status.md.tlelabs.com/). ## Start here - [Quickstart](https://md.tlelabs.com/docs/quickstart/): create a key and convert your first file. - [Authentication](https://md.tlelabs.com/docs/authentication/): API keys, rotation and the one 401. ## API reference - [Convert a file](https://md.tlelabs.com/docs/convert/): `POST /v1/convert`, its options and both response types. - [Quote before converting](https://md.tlelabs.com/docs/quote/): `POST /v1/quote`, the price of up to 10 files without converting them. - [Wallet](https://md.tlelabs.com/docs/wallet/): `GET /v1/wallet` and `GET /v1/wallet/transactions`. - [Health and status](https://md.tlelabs.com/docs/health-and-status/): `GET /v1/health` and the public status page. ## Concepts - [OCR models](https://md.tlelabs.com/docs/ocr-models/): `ocr_model`, the `auto` chains and the price of each model. - [Formats, pages and pricing](https://md.tlelabs.com/docs/formats-and-pricing/): what is extracted, how pages are counted, limits. - [Errors](https://md.tlelabs.com/docs/errors/): the error format and every error code. - [Rate limits](https://md.tlelabs.com/docs/rate-limits/): concurrent conversions and retries. These pages are also available as Markdown for tools and language models: add `.md` to a page URL (for example [`/docs/convert.md`](https://md.tlelabs.com/docs/convert.md)), or read [`/llms.txt`](https://md.tlelabs.com/llms.txt) and [`/llms-full.txt`](https://md.tlelabs.com/llms-full.txt). Source: https://md.tlelabs.com/docs/ --- # Quickstart > Create an API key in the dashboard and convert your first document with curl, JavaScript or Python. ## 1. Create an API key Sign in at [md.tlelabs.com](https://md.tlelabs.com/login) with Google, open [API keys](https://md.tlelabs.com/keys) and create a key. The key (`mdw_live_…`) is shown once: store it in a secret manager or an environment variable. New accounts start with 0 credits. Credits are added by the operator: write to . ## 2. Convert a file `POST /v1/convert` with `multipart/form-data`; the file goes in the `file` field. Replace `` with your key. **curl** ```sh curl -X POST 'https://api.md.tlelabs.com/v1/convert' \ -H 'Authorization: Bearer ' \ -F 'file=@document.pdf' \ -F 'ocr=auto' ``` **JavaScript** ```js import { readFile } from "node:fs/promises"; const form = new FormData(); form.append("file", new Blob([await readFile("document.pdf")]), "document.pdf"); form.append("ocr", "auto"); const response = await fetch("https://api.md.tlelabs.com/v1/convert", { method: "POST", headers: { Authorization: "Bearer " }, body: form, }); const result = await response.json(); console.log(result.markdown); ``` **Python** ```py import requests with open("document.pdf", "rb") as file: response = requests.post( "https://api.md.tlelabs.com/v1/convert", headers={"Authorization": "Bearer "}, files={"file": file}, data={"ocr": "auto"}, ) print(response.json()["markdown"]) ``` ## 3. Read the result The JSON response holds the Markdown and what the conversion cost: ```json { "markdown": "# Quarterly report\n\nRevenue grew 12%…", "format": "pdf", "usage": { "pages": 10, "ocr_pages": 3, "credits_charged": 16, "credits_remaining": 9984, "ocr_page_credits": 3 }, "metadata": { "byte_size": 524288, "duration_ms": 1842, "ocr_model": "auto", "ocr_models_used": ["openparser/paddleocr-vl-1.6"] }, "warnings": [] } ``` ## Next steps - [Convert a file](https://md.tlelabs.com/docs/convert/): OCR options and the plain Markdown response. - [Quote before converting](https://md.tlelabs.com/docs/quote/): know the price before you spend credits. - [Errors](https://md.tlelabs.com/docs/errors/): what to retry and what to fix. Source: https://md.tlelabs.com/docs/quickstart/ --- # Authentication > Every request except GET /v1/health sends an API key as a Bearer token; keys belong to an account and share its wallet. Send your API key in the `Authorization` header of every request: ```http Authorization: Bearer mdw_live_ ``` - A key is `mdw_live_` followed by 32 lowercase characters (`a-z`, `2-7`). Only a hash is stored: the key is shown once, when you create it. - Keys belong to your account, and credits belong to the account, not to a key. Revoking or replacing a key never changes your balance. - An account can have up to 10 active keys. Create, rename and revoke them in the dashboard under [API keys](https://md.tlelabs.com/keys). - `GET /v1/health` is the only endpoint without authentication. ## Rotating a key Create a new key, switch your clients to it, then revoke the old one. A revoked key stops working on the next request. There are no refresh tokens. ## Failed authentication Every authentication failure is status `401` with the code `invalid_api_key`. A missing or malformed header says so in the message; an unknown key, a revoked key and a suspended account all get the same response, so the API never reveals which case it was: ```json { "error": { "code": "invalid_api_key", "message": "Invalid API key." } } ``` The `/v1/*` endpoints accept only API keys: the dashboard’s sign-in session does not work there. The API is meant to be called from servers. It sends no CORS headers, so browsers block calls from web pages, and a key in front-end code would be visible to anyone. Source: https://md.tlelabs.com/docs/authentication/ --- # Convert a file > POST /v1/convert takes one document as multipart/form-data and returns its Markdown as JSON or as plain text/markdown. `POST /v1/convert` converts one document to Markdown. The format is detected from the file’s bytes, never from its name or `Content-Type`. ```sh curl -X POST https://api.md.tlelabs.com/v1/convert \ -H "Authorization: Bearer " \ -F "file=@report.pdf" \ -F "ocr=auto" \ -F "ocr_model=auto" ``` ## Request `Content-Type: multipart/form-data` with these fields: | Field | Required | Value | | ----------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------- | | `file` | yes | The document. Up to 50 MB. | | `ocr` | no | `auto` (default), `force` or `off`. See [OCR modes](#ocr-modes). | | `ocr_model` | no | `auto` (default) or a model id such as `openparser/mistral-ocr-4`. See [OCR models](https://md.tlelabs.com/docs/ocr-models/). | | `preserve_images` | no | `true` (default) or `false`. See [Picture placeholders](#picture-placeholders). | Supported formats: PDF (.pdf), DOCX (.docx), PPTX (.pptx), XLSX (.xlsx), HTML (.html, .htm). Anything else gets `415 unsupported_format`. A field with an invalid value gets `400 invalid_request` with `error.details.field` naming it; for `ocr_model` the allowed ids are in `error.details.allowed`. Headers: `Authorization` (required) and `Accept` — `application/json` (default) or `text/markdown`. ## OCR modes | `ocr` | Behavior | | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `auto` | OCR only what the file cannot give as text: pictures inside DOCX, PPTX, XLSX and HTML (`data:` pictures), scanned PDF pages, PDF pages whose text layer is missing or garbled, and large pictures on PDF pages with text. | | `force` | PDF only: every non-blank page is read by the OCR model even when it has a text layer (useful for handwriting on typed pages). Other formats behave like `auto`. | | `off` | No OCR. Pictures keep only their placeholder. A PDF with no text at all fails with `422 parse_failed`. | Content the converter reads from the file itself — text, tables, charts, SmartArt, text inside SVG and EMF/WMF pictures — is always included and costs 1 credit per page, even with `ocr=off`. See [Formats, pages and pricing](https://md.tlelabs.com/docs/formats-and-pricing/). An OCR failure does not fail the request: the affected picture or page is left without OCR text, you are charged a text page for it, and `warnings` says what happened. Only a document that ends up with no text at all because OCR failed gets `502 ocr_provider_error`, and nothing is charged. ## Picture placeholders With `preserve_images=true` each picture leaves a placeholder such as `![image](image-1.png)` where it was, followed by its OCR text as a quote: ```markdown Revenue grew ![Scanned chart](image-1.png) this quarter. > Q1 120 · Q2 134 · Q3 151 ``` With `preserve_images=false` only the text remains. Picture bytes are never stored or returned: the placeholder only marks the position. ## Response: JSON `200 OK`, `Content-Type: application/json`: ```json { "markdown": "# Quarterly report\n\n…", "format": "pdf", "usage": { "pages": 10, "ocr_pages": 3, "credits_charged": 16, "credits_remaining": 9984, "ocr_page_credits": 3 }, "metadata": { "byte_size": 524288, "duration_ms": 1842, "ocr_model": "auto", "ocr_models_used": ["openparser/paddleocr-vl-1.6"] }, "warnings": [] } ``` | Field | Meaning | | -------------------------- | --------------------------------------------------------------------------------------------------------------- | | `usage.pages` | Billed pages ([how pages are counted](https://md.tlelabs.com/docs/formats-and-pricing/#pages)). | | `usage.ocr_pages` | Pages charged at the OCR price (OCR succeeded on them). | | `usage.credits_charged` | `(pages − ocr_pages) × 1 + ocr_pages × ocr_page_credits`. | | `usage.credits_remaining` | Wallet balance after this charge. | | `usage.ocr_page_credits` | Price of one OCR page with the chosen model. | | `metadata.ocr_model` | The `ocr_model` you sent, `auto` when omitted. | | `metadata.ocr_models_used` | Models that returned at least one OCR result, in order of first use (shows `auto` fallbacks); `[]` without OCR. | | `warnings` | Non-fatal problems, for example a picture whose OCR timed out. | ## Response: Markdown With `Accept: text/markdown` the body is the Markdown itself (`Content-Type: text/markdown; charset=utf-8`) and the usage comes in headers: ```http HTTP/1.1 200 OK Content-Type: text/markdown; charset=utf-8 X-Pages: 10 X-OCR-Pages: 3 X-Credits-Charged: 16 X-Credits-Remaining: 9984 X-Format: pdf X-Duration-Ms: 1842 X-OCR-Model: auto X-OCR-Models-Used: openparser/paddleocr-vl-1.6 X-OCR-Page-Credits: 3 ``` Warnings are only in the JSON response. ## Charging Before converting, the most the file can cost is reserved from your available credits (`402 insufficient_credits` when there are not enough). After the conversion you are charged only for the pages actually converted and the rest of the reservation is released. Any error releases the whole reservation. See [Errors](https://md.tlelabs.com/docs/errors/) for the one exception (`503` with a `request_id`). Limits per request — file size, pages, slides, sheets and pictures to OCR — are listed in [Formats, pages and pricing](https://md.tlelabs.com/docs/formats-and-pricing/#limits). Base URL: `https://api.md.tlelabs.com`; files up to 50 MB. Source: https://md.tlelabs.com/docs/convert/ --- # Quote before converting > POST /v1/quote prices several files at once without converting them, calling OCR, reserving credits or charging anything. `POST /v1/quote` tells you what converting each file would cost. It is a dry run: nothing is converted, no OCR model is called, no credits are reserved or charged, and it does not count toward the [concurrent conversion limit](https://md.tlelabs.com/docs/rate-limits/). Authentication is the same as for `/v1/convert`. ```sh curl -X POST https://api.md.tlelabs.com/v1/quote \ -H "Authorization: Bearer " \ -F "file=@a.pdf" -F "file=@b.docx" \ -F "ocr_model=openparser/mistral-ocr-4" ``` ## Request `Content-Type: multipart/form-data`: | Field | Required | Value | | ----------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------- | | `file` | yes | 1 to 10 files (repeat the field), priced in the order sent. | | `ocr` | no | `auto` (default), `force` or `off`, as for [`/v1/convert`](https://md.tlelabs.com/docs/convert/#ocr-modes); applies to every file. | | `ocr_model` | no | `auto` (default) or a model id, as for `/v1/convert`; applies to every file. | The whole request may be up to 50 MB (`413 file_too_large`). No file, more than 10 files (`details: {"field": "file", "max": 10}`) or an invalid `ocr` / `ocr_model` gets `400 invalid_request`. ## Response `200 OK`. Example: `a.pdf` has 5 pages, 4 with a text layer and 1 scan; the second file is not a supported format. ```json { "ocr": "auto", "ocr_model": "auto", "ocr_page_credits": 3, "files": [ { "index": 0, "filename": "a.pdf", "size_bytes": 81234, "status": "ok", "format": "pdf", "pages": 5, "text_pages": 4, "ocr_pages": 1, "details": { "estimate": "exact_if_ocr_succeeds", "pdf_pages": { "blank": 0, "text": 4, "scan": 1, "text_render": 0, "render": 0 }, "images_to_ocr": 1 }, "ocr_models": { "image": "novita/deepseek-ocr-2", "pdf": "openparser/paddleocr-vl-1.6" }, "credits": { "estimated": 7, "max": 15 }, "warnings": [] }, { "index": 1, "filename": "x.zip", "size_bytes": 120, "status": "error", "error": { "code": "unsupported_format", "message": "File format not recognized. Supported: HTML, DOCX, PDF, PPTX, XLSX." } } ], "total": { "files": 2, "ok": 1, "estimated_credits": 7, "max_credits": 15 } } ``` | Field | Meaning | | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `pages`, `text_pages`, `ocr_pages` | Billed pages ([how pages are counted](https://md.tlelabs.com/docs/formats-and-pricing/#pages)); `text_pages = pages − ocr_pages`. | | `credits.estimated` | `text_pages × 1 + ocr_pages × ocr_page_credits`: the price if every expected OCR succeeds. | | `credits.max` | Exactly what `/v1/convert` would reserve for this file; you need this much available. | | `details.estimate` | `exact_if_ocr_succeeds`: `/v1/convert` charges `estimated` when every OCR succeeds (less when some fail). `upper_bound`: the file could not be analyzed in detail, `estimated = max`. | | `details.pdf_pages` | PDF: pages by kind — `blank`, `text`, `scan`, `text_render` (text layer plus a rendered read), `render`; `null` when not analyzed. | | `details.images_to_ocr` | Pictures expected to be sent to OCR. | | `details.basis` | HTML: `"bytes"` — priced by file size, like the reservation. | | `ocr_models` | The model tried first for pictures and for PDF pages; `null` when there is none, with `ocr=off`, and `pdf` for non-PDF files. | | `total` | Number of files, files `ok`, and the sums of `estimated` and `max` over the `ok` files. | ## How each format is estimated - **PDF up to 20 MB** (in a request up to 20 MB): pages are classified with the same logic as `/v1/convert`, so the estimate is exact when OCR succeeds. - **PDF up to 20 MB in a request over 20 MB**, with OCR on: not analyzed (`upper_bound`, `estimated = max`, warning `page analysis skipped: request larger than 20 MB`); quote that file alone for a detailed price. With `ocr=off` it is still counted exactly. - **PDF over 20 MB**: too large for OCR, so it is priced from its text layer (`ocr_pages` 0). With `ocr=auto` a page without text, and with `ocr=force` any file, gets `file_too_large`, like `/v1/convert`; with `ocr=off` it is priced from the text layer whatever it contains. - **DOCX, PPTX, XLSX**: `ocr_pages` counts the chunks, slides or sheets with a picture to OCR (`estimated = max` when an image OCR model is available). Without one, `ocr_pages` is 0, `estimated` can be lower than `max` and a warning says pictures will not be read. - **HTML**: `estimated = max` (`upper_bound`, `basis: "bytes"`). ## Errors per file A file that cannot be priced gets `status: "error"` with the same code and message `/v1/convert` would return — `unsupported_format`, `file_too_large`, `page_limit_exceeded`, `image_limit_exceeded`, `parse_failed` — or `internal_error`. It does not fail the request (still `200`) and is not counted in `total`. Source: https://md.tlelabs.com/docs/quote/ --- # Wallet > GET /v1/wallet returns the account's balance, reserved and available credits; GET /v1/wallet/transactions lists charges and top-ups. Credits belong to your account; every key of the account sees the same wallet. ## GET /v1/wallet ```sh curl https://api.md.tlelabs.com/v1/wallet -H "Authorization: Bearer " ``` ```json { "account_id": "acc_xxxxxxxxxxxxxxxxxxxxxxxxxx", "balance": 9984, "reserved": 30, "available": 9954, "lifetime_charged": 1247, "key_prefix": "mdw_live_xxxx", "created_at": "2026-05-01T03:14:15Z" } ``` | Field | Meaning | | ------------------ | ------------------------------------------------- | | `balance` | Credits in the wallet. | | `reserved` | Credits held by conversions in progress. | | `available` | `balance − reserved`: what a new request can use. | | `lifetime_charged` | Total credits charged so far. | | `key_prefix` | Prefix of the key that made the request. | | `created_at` | When the wallet was created. | ## GET /v1/wallet/transactions Charges and top-ups, newest first. ```sh curl "https://api.md.tlelabs.com/v1/wallet/transactions?limit=50" -H "Authorization: Bearer " ``` Query parameters: - `limit`: 1–200, default 50. - `cursor`: the `next_cursor` of the previous page. It is opaque: pass it back unchanged. A value out of range, not an integer, or a damaged cursor gets `400 invalid_request`. ```json { "transactions": [ { "id": "txn_01H…", "type": "convert", "amount": -16, "balance_after": 9984, "created_at": "2026-06-16T10:23:11Z", "request_id": "req_01H…", "details": { "format": "pdf", "pages": 10, "ocr_pages": 3 } }, { "id": "txn_01G…", "type": "topup", "amount": 10000, "balance_after": 10000, "created_at": "2026-05-01T03:14:15Z" } ], "next_cursor": "eyJ…" } ``` - `type` is `convert` (negative `amount`) or `topup` (positive). `details` exists only on `convert`. - `next_cursor` is `null` on the last page. - Failed conversions create no transaction. Use `request_id` to check whether a request was charged (see the `503` case in [Errors](https://md.tlelabs.com/docs/errors/)). ## Adding credits Credits are added by the operator; there is no self-serve payment. Write to . The dashboard shows the same wallet and transactions under [Billing](https://md.tlelabs.com/billing). Source: https://md.tlelabs.com/docs/wallet/ --- # Health and status > GET /v1/health reports the API and OCR provider status without authentication; status.md.tlelabs.com shows the last hour, day and month. ## GET /v1/health No authentication. It never calls an OCR provider: it reports the snapshot that a check every 5 minutes records, cached for 60 seconds, so it can be up to about 6 minutes old. ```json { "status": "ok", "version": "0.1.0", "ocr_provider": "ok", "ocr_providers": { "image": "novita", "pdf": "openparser" }, "ocr_models": { "available": ["novita/deepseek-ocr-2", "novita/paddleocr-vl", "openparser/paddleocr-vl-1.6", "…"], "auto": { "image": ["novita/deepseek-ocr-2", "openparser/paddleocr-vl-1.6"], "pdf": ["openparser/paddleocr-vl-1.6"] } }, "checked_at": "2026-10-09T10:05:00.000Z", "ocr_provider_checks": { "novita": { "status": "operational", "latency_ms": 640 }, "openparser": { "status": "operational", "latency_ms": 340 } } } ``` | Field | Meaning | | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ocr_provider` | `ok` when every configured provider is operational; `degraded` when one is not (conversions still run, OCR failures become warnings) or none is configured; `unknown` when there is no recent check. | | `ocr_providers` | Provider tried first for pictures and for PDF pages with `ocr_model=auto`. | | `ocr_models.available` | The `ocr_model` values this server accepts. | | `ocr_models.auto` | The `auto` fallback chains for pictures and PDF pages. | | `checked_at` | When the last check ran; `null` before the first one. | | `ocr_provider_checks` | Per provider: `operational`, `degraded`, `major_outage` or `unknown`, and the last latency. | The response has `Cache-Control: public, max-age=60` and is not rate limited, so monitoring can poll it. ## Status page [status.md.tlelabs.com](https://status.md.tlelabs.com/) shows the API and each OCR provider over the last hour, 24 hours and 30 days. The same data is JSON at (CORS allowed). A provider is `degraded` after one failed check or a check slower than 10 seconds, and `major_outage` after two failures in a row. Source: https://md.tlelabs.com/docs/health-and-status/ --- # OCR models > Choose the OCR model per request with ocr_model; auto tries a chain of models and falls back on failure. The `ocr_model` field of [`/v1/convert`](https://md.tlelabs.com/docs/convert/) and [`/v1/quote`](https://md.tlelabs.com/docs/quote/) picks the model that reads pictures and rendered PDF pages. Ids are `provider/model` and match exactly (case-sensitive). ## auto (default) Sending nothing is the same as `ocr_model=auto`. Each kind of work has a chain of models; when a model fails (HTTP error, rate limit, timeout, malformed answer) the next one is tried. Every model in the chains costs 3 credits per OCR page. - Pictures: `novita/deepseek-ocr-2 → openparser/paddleocr-vl-1.6 → paddle/serving` - Rendered PDF pages: `openparser/paddleocr-vl-1.6 → paddle/serving` Models the server has not configured are skipped. [`GET /v1/health`](https://md.tlelabs.com/docs/health-and-status/) lists the chains and the models available right now. ## A specific model A specific id uses only that model, with no fallback, and OCR pages cost that model’s price (reported in `usage.ocr_page_credits`). A model that cannot do a kind of work is treated as if no model were available for it: with a picture-only model, PDF pages that need rendering keep their text layer or stay empty with a warning. An unknown id, or one whose provider is not configured on the server, gets `400 invalid_request` with the accepted ids in `error.details.allowed`, before anything is reserved. `ocr=off` still validates `ocr_model` but runs no OCR. | Model | Reads | Credits per OCR page | | --------------------------------------- | ------------------- | -------------------- | | `novita/deepseek-ocr-2` | pictures | 3 | | `novita/paddleocr-vl` | pictures | 3 | | `openparser/paddleocr-vl-1.6` | pictures, PDF pages | 3 | | `openparser/mistral-ocr-3` | pictures, PDF pages | 3 | | `openparser/azure-di-read` | pictures, PDF pages | 3 | | `openparser/google-docai-ocr` | pictures, PDF pages | 3 | | `openparser/aws-textract-detect` | pictures, PDF pages | 3 | | `openparser/mistral-ocr-4` | pictures, PDF pages | 6 | | `openparser/aws-textract-layout` | pictures, PDF pages | 6 | | `openparser/azure-di-layout` | pictures, PDF pages | 15 | | `openparser/aws-textract-tables-layout` | pictures, PDF pages | 23 | | `paddle/serving` | pictures, PDF pages | 3 | A page read without OCR always costs 1 credit, whichever model you choose. See [Formats, pages and pricing](https://md.tlelabs.com/docs/formats-and-pricing/). ## What is sent to the model Only what the converter cannot read from the file goes to an OCR provider: single pictures, and PDFs of at most 10 pages that contain only the pages needing a rendered read. When a PDF cannot be split (encrypted, unreadable or rejected by the splitter), the original file is sent once instead and only the pages that need a rendered read are used from the result. The OCR text is put back in place in the Markdown. All OCR for one conversion shares a 150-second budget; work that cannot start in time becomes a warning. Source: https://md.tlelabs.com/docs/ocr-models/ --- # Formats, pages and pricing > What mdworker extracts from each format, how billed pages are counted, what a conversion costs and the limits per request. ## Formats | Format | Extensions | What you get | | ------ | ----------- | -------------------------------------------------------------------------------------------------- | | PDF | .pdf | Text layer, form fields and FreeText annotations; scanned or unreadable pages are OCR'd. | | DOCX | .docx | Headings, lists, tables, links, charts and SmartArt; pictures are OCR'd. | | PPTX | .pptx | Slides in order with titles, text, tables, charts, SmartArt and speaker notes; pictures are OCR'd. | | XLSX | .xlsx | Every sheet as a Markdown table, plus charts, shapes and pictures. | | HTML | .html, .htm | Sanitized headings, lists, tables and links, text inside SVG; embedded data: pictures are OCR'd. | The converter works cheapest first: (1) what the file contains as structure — text, tables, charts, SmartArt, text inside SVG and EMF/WMF pictures — with no OCR; (2) OCR of single pictures; (3) OCR of rendered PDF pages, only for pages that cannot be read otherwise. ## Pages | Format | One billed page | | ------ | --------------------------------------- | | PDF | One PDF page | | DOCX | 3,000 characters of text | | PPTX | One slide | | XLSX | 3,000 characters of cell text per sheet | | HTML | 3,000 characters of visible text | Counts are rounded up, with a minimum of one page (one per sheet for XLSX). HTML counts 3,000 characters of visible text after whitespace is collapsed; DOCX counts 3,000 characters of the document’s text (the body of `word/document.xml`; headers, footers, footnotes and pictures excluded), so the count is known before converting. ## Pricing - 1 credit per page read from the file. - 3 credits per OCR page with `ocr_model=auto`; a specific model has its own price ([OCR models](https://md.tlelabs.com/docs/ocr-models/)). - A page counts as an OCR page only when OCR succeeded on it. For DOCX and HTML, a page chunk with at least one picture read by OCR is an OCR page; for PPTX a slide, and for XLSX a sheet (charged once per sheet). - Nothing is charged when a conversion fails. Example: a DOCX with 4,500 characters of text and 2 pictures in the first 3,000 characters is 2 pages, 1 of them OCR: 1 × 1 + 1 × 3 = 4 credits. The most a file can cost is reserved before converting (you need that much available); afterwards you are charged only for the pages actually converted. Use [`/v1/quote`](https://md.tlelabs.com/docs/quote/) to see both numbers in advance. Credits are added by the operator: . ## Free trial The [home page](https://md.tlelabs.com/#try) converts files for free without signing in: 10 conversions per IP address per UTC day, files up to 10 MB and 50 pages. It never uses OCR: everything read from the file itself is included, pictures keep their placeholder and scanned pages are not read. The free converter is only on the home page, not in the API. ## Limits | Limit | Value | Over the limit | | --------------------------------------------------------------------------- | ----- | -------------------------- | | File size | 50 MB | `413 file_too_large` | | PDF pages | 500 | `422 page_limit_exceeded` | | PPTX slides | 200 | `422 page_limit_exceeded` | | XLSX sheets | 50 | `422 page_limit_exceeded` | | Pictures to OCR per HTML, DOCX, PPTX or XLSX file (when `ocr` is not `off`) | 50 | `422 image_limit_exceeded` | | PDF that needs OCR | 20 MB | `413 file_too_large` | Page, slide, sheet and picture limits are checked before anything is reserved. Pictures on PDF pages with text are limited to 50 per document: the rest are skipped with a warning instead of an error. Very large XML parts or Markdown output inside DOCX, PPTX and XLSX files also get `413 file_too_large`, without a charge. Source: https://md.tlelabs.com/docs/formats-and-pricing/ --- # Errors > Every error is JSON with a code and a message; no credits are charged for a failed request, with one documented 503 exception. An error response has a JSON body: ```json { "error": { "code": "insufficient_credits", "message": "Account has 5 credits available, this request requires at least 30 reserved.", "details": { "balance": 5, "required_reserve": 30 } } } ``` `code` is stable and meant for programs; `message` is for people and may change. `details` appears only for some codes. | Status | Code | Meaning | What to do | | ------ | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | | 400 | `invalid_request` | Missing `file`, wrong `Content-Type`, or an invalid field value; `details.field` names it (and `details.allowed` lists valid `ocr_model` ids). | Fix the request. | | 401 | `invalid_api_key` | Missing, malformed, revoked or unknown key, or a suspended account. | Check the key ([Authentication](https://md.tlelabs.com/docs/authentication/)). | | 402 | `insufficient_credits` | Not enough available credits for the reservation; `details` has `balance` and `required_reserve`. | Add credits, or wait for conversions in progress to finish. | | 404 | `not_found` | No such endpoint. | Check the path. | | 405 | `method_not_allowed` | Wrong HTTP method for the endpoint. | Check the method. | | 413 | `file_too_large` | File over 50 MB, a PDF over 20 MB that needs OCR, or a part inside the file too large. | Split or shrink the file. | | 415 | `unsupported_format` | Not HTML, DOCX, PDF, PPTX or XLSX. | Convert the file to a supported format first. | | 422 | `page_limit_exceeded` | Over 500 PDF pages, 200 slides or 50 sheets. | Split the document. | | 422 | `image_limit_exceeded` | More than 50 pictures to OCR in an HTML, DOCX, PPTX or XLSX file. | Split the file or send `ocr=off`. | | 422 | `parse_failed` | The file is damaged or password-protected, has no slides, or contains no text that could be extracted. | Check the file; do not retry unchanged. | | 429 | `too_many_concurrent` | Too many conversions in progress for this key. | Retry after the `Retry-After` header ([Rate limits](https://md.tlelabs.com/docs/rate-limits/)). | | 500 | `internal_error` | Unexpected failure. | Retry later. | | 502 | `ocr_provider_error` | The document has no text left because OCR failed or was unavailable. | Retry later, or choose another `ocr_model`. | | 503 | `service_unavailable` | Temporary failure. | Retry with backoff; see below when `details.request_id` is present. | ## Charging on errors No credits are charged for a request that returns an error: the reservation is released immediately. The one exception is a `503 service_unavailable` with `details.request_id`. It means the charge could not be confirmed and may have been recorded. Before retrying, look for that `request_id` in [`GET /v1/wallet/transactions`](https://md.tlelabs.com/docs/wallet/#get-v1wallettransactions): if a `convert` transaction has it, the conversion was charged (the Markdown is lost; converting again charges again). A response that is not JSON (for example an HTML error page from the network edge) should be treated like a `5xx` and retried later. A reservation left by an interrupted request expires after 5 minutes and is then released automatically. Source: https://md.tlelabs.com/docs/errors/ --- # Rate limits > Each API key may run a few conversions at the same time; extra requests get 429 too_many_concurrent with Retry-After. ## Concurrent conversions Each API key may have **5 conversions in progress** at the same time. The next one gets: ```http HTTP/1.1 429 Too Many Requests Retry-After: 5 Content-Type: application/json; charset=utf-8 {"error":{"code":"too_many_concurrent","message":"Too many conversions in progress for this API key (max 5)."}} ``` Nothing is reserved or charged for it. Wait the number of seconds in `Retry-After`, then retry. To convert more files in parallel, queue them on your side, or use separate keys for separate workloads (all keys of an account share one wallet). [`POST /v1/quote`](https://md.tlelabs.com/docs/quote/) reserves nothing, so it does not count toward this limit. [`GET /v1/health`](https://md.tlelabs.com/docs/health-and-status/) is never rate limited. ## Retrying - `429`: wait for `Retry-After`, then retry. - `500`, `502`, `503`: retry with exponential backoff. For a `503` with `details.request_id`, check whether you were charged first ([Errors](https://md.tlelabs.com/docs/errors/#charging-on-errors)). - `4xx` other than `429`: fix the request; retrying unchanged gets the same answer. The API may also limit request rates at the network edge. Treat any `429` the same way: back off and honor `Retry-After` when it is present. Source: https://md.tlelabs.com/docs/rate-limits/