catalogue / tools / Documents / PDF to plain text

Convert PDF to plain text

To convert PDF to plain text, send one POST to /v1/pdf-to-text with the PDF source. Each call costs $0.005 in USDC on Base via x402, with no account and no API key. The source file can be up to 25 MB.

Endpoint
POST /v1/pdf-to-text
Price
$0.005 per call

01Limits

LimitValue
Largest source file25 MB (26,214,400 bytes)
Timeout40 s
Source formatsPDF
Output formatsplain text
Option max_pagesdefault 100

02Example

A real PDF to plain text run of pdf-to-text.

Request
POST /v1/pdf-to-text
content-type: application/json
payment-signature: <base64 x402 payload>

{
  "url": "https://grist.tools/samples/documents/report.pdf",
  "max_pages": 100
}
Response
200 OK
payment-response: <base64 settlement receipt>

{
  "source_url": "https://grist.tools/samples/documents/report.pdf",
  "total_pages": 2,
  "extracted_pages": 2,
  "truncated": false,
  "chars": 122,
  "text": "Sample Report\nThis is a public test document.\nIt has two pages of plain text.\n\nPage Two\nThe second page closes the report."
}

Source file: documents/report.pdf, 1.2 KB (1,185 bytes). Two-page PDF with a Helvetica text layer: five short lines of neutral sample prose, built with pdf-lib with pinned metadata. The request sets max_pages to 100. The recorded run returned 122 bytes of plain text, 10% of the source size, with 2 pages in the source, 2 extracted and 122 characters.

In the sample itself, the PDF 1.7 file has 4 top-level objects, 1 object stream among them, and a cross-reference stream; inflated, its streams declare 2 pages with MediaBox 0 0 420 300, base font Helvetica and 5 text runs at 14 pt, the first reading "Sample Report". The returned text runs to 6 lines, 1 blank one among them, and its longest line, "The second page closes the report.", holds 34 characters. Its first line reads "Sample Report" and its last "The second page closes the report." Split on whitespace, it yields 23 tokens over 5 non-blank lines.

03Format notes

PDF as a source

PDF input must begin with the %PDF- signature. pdfjs, opened without network access, reads the text layer that already exists in the content stream; pages are not rasterised and no OCR runs, so a scanned, image-only PDF yields little or no text. Output is a pure function of the input, with no clock or randomness involved. For Markdown, the body font size is taken as the most common rounded line height, with ties going to the smaller size, and lines sufficiently larger than that body size are promoted to heading levels. Tables, columns, figures and inline styling are not reconstructed, and Markdown metacharacters in the text are not escaped. The max_pages option bounds how many leading pages are read; a longer document is cut and reported as truncated.

plain text as a target

Plain text is recovered in several ways, each tied to one reader. ocr-image-to-text reads pixels with Tesseract running offline from local files, returns a mean confidence with the text, and offers only the languages whose data ships with the service. pdf-to-text returns the existing text layer page by page and performs no OCR. epub-to-text follows the spine's reading order and adds title, creator and language metadata. pptx-to-text returns slide body text in slide order, leaving out speaker notes. html-clean-text extracts article text with Readability by default and returns Markdown in the same response. Because each reader works differently, the text reflects its source: recognised characters with a confidence score from an image, but characters already present in the file from documents and pages.

Note: It reads the text layer of the PDF and runs no OCR: a scanned page without one comes back with little or no text.

04Errors

CodeHTTP
invalid_input400
blocked_target403
unreachable_target424
upstream_timeout424
unsupported_content_type415
too_large413
unprocessable422
internal500

When each is raised, and what it means for payment: /docs/pdf-to-text.

05Questions

What does it cost to convert PDF to plain text?
$0.005 per call, paid in USDC on Base via x402.
How large can the PDF file be?
Up to 25 MB (26,214,400 bytes). Past a limit the call answers too_large (413).
What happens if the conversion fails?
The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.

06Related