catalogue / tools / Documents / PDF to Markdown

Convert PDF to Markdown

To convert PDF to Markdown, send one POST to /v1/pdf-to-markdown with the PDF source. Each call costs $0.005 in USDC on Base via x402, with no account and no API key. The source file can be up to 25 MB.

Endpoint
POST /v1/pdf-to-markdown
Price
$0.005 per call

01Limits

LimitValue
Largest source file25 MB (26,214,400 bytes)
Timeout40 s
Source formatsPDF
Output formatsMarkdown
Option max_pagesdefault 100

02Example

A real PDF to Markdown run of pdf-to-markdown.

Request
POST /v1/pdf-to-markdown
content-type: application/json
payment-signature: <base64 x402 payload>

{
  "url": "https://grist.tools/samples/documents/markdown-guide.pdf",
  "max_pages": 100
}
Response
200 OK
payment-response: <base64 settlement receipt>

{
  "source_url": "https://grist.tools/samples/documents/markdown-guide.pdf",
  "total_pages": 1,
  "extracted_pages": 1,
  "truncated": false,
  "chars": 249,
  "markdown": "# Field Notes\n\n## A short guide in three parts\n\nThis sample shows how headings are inferred.\n\nLarger text becomes a heading, body text stays plain.\n\n### Second section\n\nEach line of the page becomes its own block.\n\nThe document ends after this line."
}

Source file: documents/markdown-guide.pdf, 1.1 KB (1,128 bytes). One-page PDF with lines drawn at 24, 18, 14 and 12 pt Helvetica, so a 12 pt body sits under three heading sizes; built with pdf-lib with pinned metadata. The request sets max_pages to 100. The recorded run returned 249 bytes of Markdown, 22% of the source size, with 1 page in the source, 1 extracted and 249 characters.

In the sample itself, the PDF 1.7 file has 3 top-level objects, 1 object stream among them, and a cross-reference stream; inflated, its streams declare 1 page with MediaBox 0 0 420 320, base font Helvetica and 7 text runs at 24, 18, 12 and 14 pt, the first reading "Field Notes". The returned Markdown splits into 7 blocks separated by blank lines, with 3 headings (1 at level 1, 1 at level 2, 1 at level 3) reading "Field Notes", "A short guide in three parts" and "Second section". Plain paragraphs make up 4 blocks, the first reading "This sample shows how headings are inferred." and the last "The document ends after this line." Its headings total 10 words, its paragraphs 31 words.

03Format notes

PDF as a source

PDF input must begin with the %PDF- signature. pdfjs, opened without network access, reads the text layer that already exists in the content stream; pages are not rasterised and no OCR runs, so a scanned, image-only PDF yields little or no text. Output is a pure function of the input, with no clock or randomness involved. For Markdown, the body font size is taken as the most common rounded line height, with ties going to the smaller size, and lines sufficiently larger than that body size are promoted to heading levels. Tables, columns, figures and inline styling are not reconstructed, and Markdown metacharacters in the text are not escaped. The max_pages option bounds how many leading pages are read; a longer document is cut and reported as truncated.

Markdown as a target

Markdown output comes from two kinds of source. Where the input already has document structure, docx-to-markdown and html-clean-text both hand HTML to turndown, which writes ATX headings, fenced code blocks and hyphen bullets. PDF offers only positioned text, so pdf-to-markdown infers headings from font size relative to the body size, and it does not rebuild tables, columns, figures or inline styles; its output is a transcription with no escaping of Markdown characters. For HTML, Readability keeps the main article by default; setting readability to false, or a page where no article is found, converts the whole body including navigation and footer. html-clean-text cuts both Markdown and text at max_chars, counted in Unicode code points, and flags a shortened result as truncated. The same call also returns plain text alongside the Markdown.

Note: It reads the text layer of the PDF and runs no OCR: a scanned page without one comes back with little or no text.

04Errors

CodeHTTP
invalid_input400
blocked_target403
unreachable_target424
upstream_timeout424
unsupported_content_type415
too_large413
unprocessable422
internal500

When each is raised, and what it means for payment: /docs/pdf-to-markdown.

05Questions

What does it cost to convert PDF to Markdown?
$0.005 per call, paid in USDC on Base via x402.
How large can the PDF file be?
Up to 25 MB (26,214,400 bytes). Past a limit the call answers too_large (413).
What happens if the conversion fails?
The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.

06Related