catalogue / tools / Documents

Documents tools

Document APIs extract text or tables, inspect and rearrange PDFs, and render documents to PDF. Agents and scripts can use them when document parsing, OCR or layout requires a dedicated engine and guarded fetching. csv-to-json accepts delimited text and auto-detects the delimiter; xlsx-to-json reads a spreadsheet workbook and converts dates to ISO strings. Both return header-keyed rows or arrays. json-to-csv takes JSON objects, infers or accepts columns, and JSON-encodes nested values.

docx-to-markdown converts a word-processing document with warnings for content that did not translate. pptx-to-text extracts slide body text in order; epub-to-text returns chapter text in reading order and Dublin Core metadata. pdf-to-text reads the text layer page by page; pdf-to-markdown infers headings from font size. Neither performs OCR or guarantees complete extraction. Unavailable character mappings can leave empty or partial text without an error. ocr-image-to-text reads images with Tesseract and returns text with mean confidence.

pdf-metadata returns page count, document-info fields, version and encryption status without rendering. pdf-split takes a page selection; pdf-merge concatenates source PDFs in order. Both return a PDF encoded as base64. html-to-pdf renders a web page or inline HTML; markdown-to-pdf renders Markdown. Both return a PDF encoded as base64 with its page count. PDF Markdown conversion is best-effort, without exact layout preservation. CSV output does not neutralize spreadsheet formulas: quoting handles delimiters, quotes and newlines while preserving formula prefixes.

01APIs

APIWhat it doesPriceMax inputMax durationTimeout
CSV to JSON APIParses a CSV/TSV document into JSON rows, keyed by a header row or as arrays; the delimiter is auto-detected.$0.005 per call25 MBnone30 s
DOCX to Markdown APIConverts a Word .docx to Markdown, surfacing anything that did not translate as warnings.$0.005 per call25 MBnone40 s
EPUB to Text APIExtracts an EPUB’s reading-order text and Dublin Core metadata, chapter by chapter.$0.005 per call25 MBnone45 s
HTML to PDF APIRenders a web page or an inline HTML string to a PDF, returned base64-encoded with its page count.$0.005 per call5 MBnone12 s
JSON to CSV APIRenders JSON objects as raw CSV with no spreadsheet formula protection; columns are inferred or given, nested values are JSON-encoded.$0.005 per call25 MBnone30 s
Markdown to PDF APIRenders a Markdown document to a PDF, returned base64-encoded with its page count.$0.005 per call2 MBnone12 s
OCR Image to Text APIReads the text out of an image (scan, screenshot, photo) with Tesseract, returned with a mean confidence.$0.005 per call10 MBnone45 s
PDF Merge APIConcatenates 2–8 source PDFs, in order, into one PDF returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding.$0.005 per call8 MBnone45 s
PDF Metadata APIReads a PDF's page count, document-info fields, version and encryption flag without rendering it.$0.005 per call25 MBnone25 s
PDF Split APIExtracts a 1-indexed page selection (e.g. "1,3-4") from a PDF into one new PDF, returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding.$0.005 per call25 MBnone30 s
PDF to Markdown APIConverts a PDF to Markdown, inferring headings from font size (best-effort, not layout-perfect; no OCR).$0.005 per call25 MBnone40 s
PDF to Text APIExtracts the text layer of a PDF, page by page, in reading order (no OCR).$0.005 per call25 MBnone40 s
PPTX to Text APIExtracts the slide text of a PowerPoint .pptx, slide by slide, in order (body text only, up to 1,000 slides).$0.005 per call25 MBnone40 s
XLSX to JSON APIReads an Excel .xlsx into JSON rows, keyed by a header row or as arrays; dates become ISO strings. Archives are limited to 64 worksheets.$0.005 per call25 MBnone40 s

02Conversions

03Questions

Do the PDF text extractors perform OCR or guarantee complete text?
No. pdf-to-text and pdf-to-markdown extract the text layer without OCR. They do not fetch external CMap or standard-font data. A PDF that needs unavailable predefined CMaps can yield empty or partial text without an error. A successful response does not certify completeness, and truncated reports only the max_pages limit.
Does json-to-csv protect against spreadsheet formulas?
No. It preserves formula prefixes in headers and string values without adding apostrophes or trimming tabs and carriage returns. CSV quoting escapes delimiters, quotes and newlines but does not neutralize formulas. Records are separated by CRLF.
How do PDF splitting and merging differ?
pdf-split extracts a page selection from a source PDF into a new PDF. pdf-merge concatenates source PDFs in order. Both return the resulting PDF encoded as base64.

04Related