catalogue / tools / Documents / PDF to Text API

PDF to Text API

Extracts the text layer of a PDF, page by page, in reading order (no OCR).

Endpoint
POST /v1/pdf-to-text
Price
$0.005 per call
Family
Documents

01When to use it

  • Pulling a PDF text layer needs a real PDF engine and a hardened fetch: an agent that fetches the URL itself walks past every SSRF control, and carrying pdfjs per call is not something an agent does inline. Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit.
  • Your source is PDF and you need plain text.

When not to use it

  • The input is over 25 MB (26,214,400 bytes): pdf-to-text answers too_large (413).
  • The destination needs more than 40 s to answer: the call ends with upstream_timeout (424).
  • The destination is on a private or local network: the network guard refuses it with blocked_target (403).

02Parameters

FieldTypeRequiredDescription
urlstringyesAbsolute http(s) URL of the PDF, up to 2048 characters; redirects are followed and the body must start with the %PDF- signature.
max_pagesintegernoHow many leading pages to extract, 1 to 1000; a shorter document is not an error, and a longer one is cut and reported as truncated. Default: 100.

03Limits

LimitValue
Largest source file25 MB (26,214,400 bytes)
Timeout40 s
Source formatsPDF
Output formatsplain text

04Example

The published example of pdf-to-text, verbatim: the request body and the response it returns.

Request
POST /v1/pdf-to-text
content-type: application/json
payment-signature: <base64 x402 payload>

{
  "url": "https://grist.tools/samples/documents/report.pdf",
  "max_pages": 100
}
Response
200 OK
payment-response: <base64 settlement receipt>

{
  "source_url": "https://grist.tools/samples/documents/report.pdf",
  "total_pages": 2,
  "extracted_pages": 2,
  "truncated": false,
  "chars": 122,
  "text": "Sample Report\nThis is a public test document.\nIt has two pages of plain text.\n\nPage Two\nThe second page closes the report."
}

Sample file: documents/report.pdf. Two-page PDF with a Helvetica text layer: five short lines of neutral sample prose, built with pdf-lib with pinned metadata.

In the sample itself, the PDF 1.7 file has 4 top-level objects, 1 object stream among them, and a cross-reference stream; inflated, its streams declare 2 pages with MediaBox 0 0 420 300, base font Helvetica and 5 text runs at 14 pt, the first reading "Sample Report". The returned text runs to 6 lines, 1 blank one among them, and its longest line, "The second page closes the report.", holds 34 characters. Its first line reads "Sample Report" and its last "The second page closes the report." Split on whitespace, it yields 23 tokens over 5 non-blank lines.

Response fieldTypeIn the example
source_urlstring"https://grist.tools/samples/documents/report.pdf"
total_pagesnumber2
extracted_pagesnumber2
truncatedbooleanfalse
charsnumber122
textstring"Sample Report This is a public test document. It has two pages of plain text. Page Two The second page closes the report."

05Call it from code

JavaScript
// pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
// Install: npm install @x402/fetch@2 @x402/evm@2 viem@2
// Save as client.mjs (ES module, Node 18+), export EVM_PRIVATE_KEY with the paying wallet's key in your shell, then run: node client.mjs
// The wrapper reads the 402 (payment-required), signs and retries with payment-signature.
import { wrapFetchWithPayment, x402Client, decodePaymentResponseHeader } from '@x402/fetch'
import { ExactEvmScheme } from '@x402/evm/exact/client'
import { privateKeyToAccount } from 'viem/accounts'

const account = privateKeyToAccount(process.env.EVM_PRIVATE_KEY)
const client = new x402Client().register('eip155:8453', new ExactEvmScheme(account))
const fetchWithPayment = wrapFetchWithPayment(fetch, client)

const body = {
  "url": "https://grist.tools/samples/documents/report.pdf",
  "max_pages": 100
}

const res = await fetchWithPayment('https://grist.tools/v1/pdf-to-text', {
  method: 'POST',
  headers: { 'content-type': 'application/json' },
  body: JSON.stringify(body),
})
console.log(res.status, await res.json())
const settle = res.headers.get('payment-response')
if (settle) console.log(decodePaymentResponseHeader(settle).transaction)
Python
# pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
# Install: pip install "x402[requests,evm]>=2.24,<3"
# The session reads the 402 (payment-required), signs and retries with payment-signature.
import os
from eth_account import Account
from x402 import x402ClientSync
from x402.http import decode_payment_response_header
from x402.http.clients.requests import x402_requests
from x402.mechanisms.evm.exact import register_exact_evm_client
from x402.mechanisms.evm.signers import EthAccountSigner

account = Account.from_key(os.environ["EVM_PRIVATE_KEY"])
client = x402ClientSync()
register_exact_evm_client(client, EthAccountSigner(account), networks="eip155:8453")
session = x402_requests(client)

payload = {
    "url": "https://grist.tools/samples/documents/report.pdf",
    "max_pages": 100
}

res = session.post("https://grist.tools/v1/pdf-to-text", json=payload)
print(res.status_code, res.json())
settle = res.headers.get("payment-response")
if settle:
    print(decode_payment_response_header(settle).transaction)
curl
# pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
# 1. Unpaid call: HTTP 402, the requirements in the payment-required header (base64 JSON) and in the body.
curl -i -X POST 'https://grist.tools/v1/pdf-to-text' -H 'content-type: application/json' -d '{"url":"https://grist.tools/samples/documents/report.pdf","max_pages":100}'

# 2. Same call with the signed payment (an EIP-3009 authorization, EIP-712 signed: it cannot be
#    typed by hand). x-payment is accepted as the v1 alternative. HTTP 200 carries payment-response.
curl -i -X POST 'https://grist.tools/v1/pdf-to-text' -H 'content-type: application/json' -H 'payment-signature: <base64 x402 payload>' -d '{"url":"https://grist.tools/samples/documents/report.pdf","max_pages":100}'

06Errors

CodeHTTPFor this service
invalid_input400The body does not match the schema; its fields are url and max_pages.
blocked_target403
unreachable_target424
upstream_timeout424
unsupported_content_type415The source is not a type pdf-to-text reads; it reads PDF.
too_large413
unprocessable422
internal500

Input that is rejected before the tool runs (malformed JSON, schema mismatch, body over the size limit) is never charged.

When each code is raised, and what it means for payment: /docs/pdf-to-text.

07Conversions

ConversionRecorded sampleSource sizeOutput size
PDF to plain textdocuments/report.pdf1.2 KB122 bytes

08Questions

How much does a pdf-to-text call cost?
$0.005 per call, paid in USDC on Base via x402, with no account and no API key.
What does pdf-to-text return?
A JSON object with 6 top-level fields: source_url, total_pages, extracted_pages, truncated, chars and text. The example on this page is a real call over documents/report.pdf.
What happens when a pdf-to-text call fails?
The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled. A 503 upstream_unavailable can follow a settled call; the x402 payments guide covers it, 429, 402 and an uncertain 500.

09Related