Parsing PDF, DOCX, XLSX and EPUB for agents
An agent extracts text, Markdown or JSON by sending a public document URL to the endpoint for its format: PDF, DOCX, PPTX, XLSX, EPUB or CSV. Use image OCR for scans, screenshots and photos, then inspect extraction limits and warnings before passing the result into the next stage.
Updated
01Start with the output your agent needs
Call POST https://grist.tools/v1/<name> with a JSON body and content-type: application/json. Most document endpoints fetch the public URL in url; fetch the free input schema at GET /v1/<name>/schema before constructing the request. /openapi.json also publishes the schemas. No account or API key is required.
Choose text for a reading task, Markdown when the supported conversion structure is useful, and JSON rows for tabular work. The table below covers the documents family, including extraction, PDF utilities and output rendering. Use its endpoint documentation links for the full input contract and its generated limits when planning each call.
02PDF text layers and image OCR are separate paths
Use pdf-to-text to extract the PDF text layer page by page in reading order. Choose pdf-to-markdown for Markdown with headings inferred from font size. That inference is best effort and does not reproduce the layout exactly. Both accept max_pages for leading pages, and neither performs OCR. The example below demonstrates pdf-to-markdown and its returned Markdown.
Both PDF extractors operate without fetching external CMap or standard-font data. Self-contained character mappings can yield text, including CJK, but a document needing unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction.
For an image, choose ocr-image-to-text. It accepts PNG, JPEG, TIFF, WebP and GIF identified by their bytes, and returns text with mean confidence. Its accepted lang is eng. The input is an image URL, so do not send it a PDF URL expecting automatic page conversion.
03Choose the reader for DOCX, PPTX or EPUB
Use docx-to-markdown for a DOCX document and retain its warnings alongside markdown. Conversion limitations surface as warnings rather than a promise that every element translated. The fetched body must be a ZIP containing a DOCX document; an arbitrary archive is insufficient.
Use pptx-to-text for ordered slide body text. Its slides entries preserve slide numbers and text, while text combines the extraction. This is a body-text reader, so do not plan on speaker notes or slide rendering. The archive must contain the expected slide XML parts.
Use epub-to-text for chapter text in reading order and Dublin Core metadata. Its archive container must point to an OPF package. max_chapters selects leading XHTML or HTML spine chapters; a longer book is cut and reported as truncated. Keep chapters when chapter boundaries matter instead of using only the combined text.
04Extract tabular data without assuming completeness
Use xlsx-to-json for XLSX rows, with dates returned as ISO strings. Set sheet to an exact worksheet name, or omit it to return every worksheet. header_row selects the column-name row; setting it to 0 returns positional arrays and columns set to null. max_rows limits each worksheet separately, with truncation reported per sheet.
Use csv-to-json for UTF-8 CSV or TSV, including quoted fields, embedded newlines and escaped quotes. The delimiter is detected from the first rows unless supplied. header chooses keyed objects or arrays; header names are trimmed, blank names replaced and duplicates made unique. A UTF-8 BOM is accepted, but unsupported charset declarations, NUL bytes and malformed UTF-8 are rejected. A truncated success does not validate the entire CSV structure.
05Services covered
| Endpoint | Summary | Price | Timeout | max_timeout_seconds | Size cap | Max duration |
|---|---|---|---|---|---|---|
| csv-to-json | Parses a CSV/TSV document into JSON rows, keyed by a header row or as arrays; the delimiter is auto-detected. | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - |
| docx-to-markdown | Converts a Word .docx to Markdown, surfacing anything that did not translate as warnings. | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - |
| epub-to-text | Extracts an EPUB’s reading-order text and Dublin Core metadata, chapter by chapter. | $0.005 | 45 s | 66 | 25 MB (26,214,400 bytes) | - |
| html-to-pdf | Renders a web page or an inline HTML string to a PDF, returned base64-encoded with its page count. | $0.005 | 12 s | 33 | 5 MB (5,242,880 bytes) | - |
| json-to-csv | Renders JSON objects as raw CSV with no spreadsheet formula protection; columns are inferred or given, nested values are JSON-encoded. | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - |
| markdown-to-pdf | Renders a Markdown document to a PDF, returned base64-encoded with its page count. | $0.005 | 12 s | 33 | 2 MB (2,097,152 bytes) | - |
| ocr-image-to-text | Reads the text out of an image (scan, screenshot, photo) with Tesseract, returned with a mean confidence. | $0.005 | 45 s | 66 | 10 MB (10,485,760 bytes) | - |
| pdf-merge | Concatenates 2–8 source PDFs, in order, into one PDF returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding. | $0.005 | 45 s | 66 | 8 MB (8,388,608 bytes) | - |
| pdf-metadata | Reads a PDF's page count, document-info fields, version and encryption flag without rendering it. | $0.005 | 25 s | 46 | 25 MB (26,214,400 bytes) | - |
| pdf-split | Extracts a 1-indexed page selection (e.g. "1,3-4") from a PDF into one new PDF, returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding. | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - |
| pdf-to-markdown | Converts a PDF to Markdown, inferring headings from font size (best-effort, not layout-perfect; no OCR). | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - |
| pdf-to-text | Extracts the text layer of a PDF, page by page, in reading order (no OCR). | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - |
| pptx-to-text | Extracts the slide text of a PowerPoint .pptx, slide by slide, in order (body text only, up to 1,000 slides). | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - |
| xlsx-to-json | Reads an Excel .xlsx into JSON rows, keyed by a header row or as arrays; dates become ISO strings. Archives are limited to 64 worksheets. | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - |
06Limits
- Largest response cap: 25 MB (26,214,400 bytes) on csv-to-json.
- Smallest response cap: 2 MB (2,097,152 bytes) on markdown-to-pdf.
- Longest timeout: 45 s on epub-to-text.
- Shortest timeout: 12 s on html-to-pdf.
- max_timeout_seconds: from 33 to 66.
07End-to-end example
A real recorded example of pdf-to-markdown, the same one its docs page publishes. Strings in the response longer than 400 characters are cut and end with ....
POST /v1/pdf-to-markdown
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/documents/markdown-guide.pdf",
"max_pages": 100
}200 OK
payment-response: <base64 settlement receipt>
{
"source_url": "https://grist.tools/samples/documents/markdown-guide.pdf",
"total_pages": 1,
"extracted_pages": 1,
"truncated": false,
"chars": 249,
"markdown": "# Field Notes\n\n## A short guide in three parts\n\nThis sample shows how headings are inferred.\n\nLarger text becomes a heading, body text stays plain.\n\n### Second section\n\nEach line of the page becomes its own block.\n\nThe document ends after this line."
}08Inspect and select PDFs before extracting
Use pdf-metadata when the agent needs page count, document-info fields, version or the encryption flag. It reads these without rendering and can describe an encrypted PDF without decrypting it. Treat the metadata as inspection data, not extracted document text.
Use pdf-split to make a new PDF from ranges. Page numbering starts at the first page; comma-separated selections and inclusive ranges are emitted in the written order, including repeats. Descending ranges and pages beyond the end are errors. Use pdf-merge to concatenate source PDFs in the order of urls; sources are fetched sequentially. Both refuse encrypted sources and return pdf_base64 with page and byte counts. Neither call replaces a text-extraction step.
09Render a deliverable or export rows
Choose html-to-pdf for an HTML or XHTML URL, or provide self-contained html. Supply exactly one source. JavaScript is disabled, external images, stylesheets and fonts are blocked, and inline PNG data images are supported. Build the input around those restrictions instead of expecting a page dependent on external resources to render fully.
Choose markdown-to-pdf for a UTF-8 Markdown URL or inline markdown, again supplying exactly one source. Raw HTML is escaped. PNG, JPEG, GIF and WebP data-image syntax is accepted; SVG data images and external image loading are blocked. Both renderers accept format and landscape and return pdf_base64 with a page count.
Choose json-to-csv for a public UTF-8 JSON document containing an object or an array of objects. columns fixes column order; otherwise columns are inferred from rows retained by max_rows. Nested values are JSON-encoded. The output uses CSV quoting and CRLF record separators, but it does not neutralize spreadsheet formulas. Headers and strings preserve formula prefixes, tabs and carriage returns; use spreadsheet-specific import controls for untrusted data.
10Chain calls through their actual input contracts
For a PDF reading pipeline, inspect with pdf-metadata if page count or encryption affects your decision, then call pdf-to-markdown on the source URL. Compare total_pages, extracted_pages and truncated, and inspect markdown itself. Here truncated reports only the max_pages limit: a false value does not certify complete text extraction.
When you need selected pages first, insert pdf-split. Its pdf_base64 is output data, while the extractor expects a public URL. Your pipeline must decode and make the resulting PDF available at a public URL before submitting the extraction request. The same handoff applies after pdf-merge; do not treat base64 output as an automatically hosted document.
A DOCX conversion can feed its returned markdown directly into the inline input of markdown-to-pdf. For an XLSX export, select the returned worksheet rows and make that JSON available at a public URL for json-to-csv. Preserve warnings, chapter or slide boundaries, and truncation indicators with the extracted content so the agent can decide whether its next task has enough input.
11Budget each successful call and its deadline
A PDF inspection and Markdown extraction costs $0.005 plus $0.005 when both calls succeed. Adding page selection adds $0.005. A DOCX extraction followed by PDF rendering costs $0.005 plus $0.005; budget image OCR separately at $0.005 per successful call. Use the table for other branches of your pipeline.
Payment uses x402 in USDC on Base. Without payment-signature, a call answers HTTP 402 with the price (amount, in atomic units of the asset) and maxTimeoutSeconds before payment; the free /schema route reports it as max_timeout_seconds. For pdf-to-markdown, the handler timeout is 40 s, the input cap is 25 MB, and the whole-call allowance in seconds is 61. Size the client deadline around the whole-call allowance, which includes verification and settlement.
Payment is verified before the handler and settled only after success. Gate refusals and handler failures are never settled. A settlement call that fails without answering or an answer that fails validation (that case is HTTP 500 internal, with a message starting "Payment settlement returned") leaves an uncertainty identified in the error message; inspect that message before deciding to repeat a paid request.
12Retry according to the error
Correct invalid_input against the schema before retrying. blocked_target means the network guard refused the destination, such as a disallowed address, unsupported URL scheme or port, or credentials in the URL. Repeating the same destination does not address that cause. For unsupported_content_type, supply a supported document type; for unprocessable, check details: document endpoints also return it when the target answered with an HTTP error status (details.status), otherwise it covers bytes that arrived but could not be processed, such as a malformed or encrypted document.
too_large means the response exceeded its streaming size cap or decompression ratio, or the document exceeded a processing limit of the endpoint, such as pixel count, slide or worksheet count, the CSV parsing budget, the PDF text extraction limit or the output PDF size. Change the source accordingly. unreachable_target means resolution failed, or the connection failed or closed before the body arrived; upstream_timeout means the destination did not answer within the declared timeout. Check source availability before scheduling another attempt.
rate_limited is HTTP 429 for excessive calls from a payer or towards a destination domain, with nothing charged. Reduce that traffic. Retryable HTTP 503 errors include concurrency_limit, registry_at_capacity and upstream_unavailable; wait the seconds specified by retry-after. When upstream_unavailable comes from the front proxy, the call may already have run and settled; check authorizationState before paying again. An internal error is a service fault, with nothing charged in almost every case; preserve its message, especially any settlement uncertainty.
13Questions
- Does PDF text extraction perform OCR?
- Neither
pdf-to-textnorpdf-to-markdownperforms OCR; both extract the existing text layer.ocr-image-to-textaccepts supported image URLs and returns text with mean confidence, usingengas its accepted language. - Does a successful PDF response mean all text was extracted?
- No: unavailable predefined character mappings can produce empty or partial text without an error.
truncatedreports only themax_pageslimit, so inspect the returned content even when it is false. - How do I select a worksheet from an XLSX file?
- Pass its exact name in
sheettoxlsx-to-json; omitting that field returns every worksheet. An unknown name produces HTTP 422 listing available names.max_rowsapplies separately to each returned worksheet. - Can an agent pass extracted Markdown directly to PDF rendering?
- Yes, pass the extracted string as
markdowntomarkdown-to-pdf, without also supplyingurl. Raw HTML is escaped, and external images and SVG data images are blocked. - Can these PDF utilities decrypt an encrypted document?
pdf-metadatadescribes an encrypted PDF without decrypting it.pdf-splitandpdf-mergerefuse encrypted sources, so metadata inspection does not make those files eligible for transformation.- Does CSV export protect against spreadsheet formulas?
json-to-csvreturns raw CSV without spreadsheet formula protection. Quoting handles delimiters, quotes and newlines while formula prefixes remain intact; use spreadsheet-specific import controls when handling untrusted data.- Are failed document calls charged?
- Gate refusals and failures inside the handler are never settled. Payment settlement follows handler success, but if a settlement call fails without answering or its answer fails validation (that case is HTTP 500
internal, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance. The error message identifies that uncertainty.
14Related pages
- CSV to JSON API docs
- DOCX to Markdown API docs
- EPUB to Text API docs
- HTML to PDF API docs
- JSON to CSV API docs
- Markdown to PDF API docs
- OCR Image to Text API docs
- PDF Merge API docs
- PDF Metadata API docs
- PDF Split API docs
- PDF to Markdown API docs
- PDF to Text API docs
- PPTX to Text API docs
- XLSX to JSON API docs
- CSV to JSON API
- DOCX to Markdown API
- EPUB to Text API
- HTML to PDF API
- JSON to CSV API
- Markdown to PDF API
- OCR Image to Text API
- PDF Merge API
- PDF Metadata API
- PDF Split API
- PDF to Markdown API
- PDF to Text API
- PPTX to Text API
- XLSX to JSON API
- Documents
- API documentation
- Guides
- How AI agents pay for APIs with x402
- Turning web pages into LLM-ready text
- Audio and video conversion API without ffmpeg servers