PDF to Text API
Extracts the text layer of a PDF, page by page, in reading order (no OCR).
- Endpoint
POST /v1/pdf-to-text- Price
- $0.005 per call
- Family
- Documents
01When to use it
- Pulling a PDF text layer needs a real PDF engine and a hardened fetch: an agent that fetches the URL itself walks past every SSRF control, and carrying pdfjs per call is not something an agent does inline. Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit.
- Your source is PDF and you need plain text.
When not to use it
- The input is over 25 MB (26,214,400 bytes): pdf-to-text answers too_large (413).
- The destination needs more than 40 s to answer: the call ends with upstream_timeout (424).
- The destination is on a private or local network: the network guard refuses it with blocked_target (403).
02Parameters
| Field | Type | Required | Description |
|---|---|---|---|
| url | string | yes | Absolute http(s) URL of the PDF, up to 2048 characters; redirects are followed and the body must start with the %PDF- signature. |
| max_pages | integer | no | How many leading pages to extract, 1 to 1000; a shorter document is not an error, and a longer one is cut and reported as truncated. Default: 100. |
03Limits
| Limit | Value |
|---|---|
| Largest source file | 25 MB (26,214,400 bytes) |
| Timeout | 40 s |
| Source formats | |
| Output formats | plain text |
04Example
The published example of pdf-to-text, verbatim: the request body and the response it returns.
POST /v1/pdf-to-text
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/documents/report.pdf",
"max_pages": 100
}200 OK
payment-response: <base64 settlement receipt>
{
"source_url": "https://grist.tools/samples/documents/report.pdf",
"total_pages": 2,
"extracted_pages": 2,
"truncated": false,
"chars": 122,
"text": "Sample Report\nThis is a public test document.\nIt has two pages of plain text.\n\nPage Two\nThe second page closes the report."
}Sample file: documents/report.pdf. Two-page PDF with a Helvetica text layer: five short lines of neutral sample prose, built with pdf-lib with pinned metadata.
In the sample itself, the PDF 1.7 file has 4 top-level objects, 1 object stream among them, and a cross-reference stream; inflated, its streams declare 2 pages with MediaBox 0 0 420 300, base font Helvetica and 5 text runs at 14 pt, the first reading "Sample Report". The returned text runs to 6 lines, 1 blank one among them, and its longest line, "The second page closes the report.", holds 34 characters. Its first line reads "Sample Report" and its last "The second page closes the report." Split on whitespace, it yields 23 tokens over 5 non-blank lines.
| Response field | Type | In the example |
|---|---|---|
| source_url | string | "https://grist.tools/samples/documents/report.pdf" |
| total_pages | number | 2 |
| extracted_pages | number | 2 |
| truncated | boolean | false |
| chars | number | 122 |
| text | string | "Sample Report This is a public test document. It has two pages of plain text. Page Two The second page closes the report." |
05Call it from code
// pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
// Install: npm install @x402/fetch@2 @x402/evm@2 viem@2
// Save as client.mjs (ES module, Node 18+), export EVM_PRIVATE_KEY with the paying wallet's key in your shell, then run: node client.mjs
// The wrapper reads the 402 (payment-required), signs and retries with payment-signature.
import { wrapFetchWithPayment, x402Client, decodePaymentResponseHeader } from '@x402/fetch'
import { ExactEvmScheme } from '@x402/evm/exact/client'
import { privateKeyToAccount } from 'viem/accounts'
const account = privateKeyToAccount(process.env.EVM_PRIVATE_KEY)
const client = new x402Client().register('eip155:8453', new ExactEvmScheme(account))
const fetchWithPayment = wrapFetchWithPayment(fetch, client)
const body = {
"url": "https://grist.tools/samples/documents/report.pdf",
"max_pages": 100
}
const res = await fetchWithPayment('https://grist.tools/v1/pdf-to-text', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
})
console.log(res.status, await res.json())
const settle = res.headers.get('payment-response')
if (settle) console.log(decodePaymentResponseHeader(settle).transaction)# pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
# Install: pip install "x402[requests,evm]>=2.24,<3"
# The session reads the 402 (payment-required), signs and retries with payment-signature.
import os
from eth_account import Account
from x402 import x402ClientSync
from x402.http import decode_payment_response_header
from x402.http.clients.requests import x402_requests
from x402.mechanisms.evm.exact import register_exact_evm_client
from x402.mechanisms.evm.signers import EthAccountSigner
account = Account.from_key(os.environ["EVM_PRIVATE_KEY"])
client = x402ClientSync()
register_exact_evm_client(client, EthAccountSigner(account), networks="eip155:8453")
session = x402_requests(client)
payload = {
"url": "https://grist.tools/samples/documents/report.pdf",
"max_pages": 100
}
res = session.post("https://grist.tools/v1/pdf-to-text", json=payload)
print(res.status_code, res.json())
settle = res.headers.get("payment-response")
if settle:
print(decode_payment_response_header(settle).transaction)# pdf-to-text: $0.005 per call, paid in USD Coin on eip155:8453 via x402
# 1. Unpaid call: HTTP 402, the requirements in the payment-required header (base64 JSON) and in the body.
curl -i -X POST 'https://grist.tools/v1/pdf-to-text' -H 'content-type: application/json' -d '{"url":"https://grist.tools/samples/documents/report.pdf","max_pages":100}'
# 2. Same call with the signed payment (an EIP-3009 authorization, EIP-712 signed: it cannot be
# typed by hand). x-payment is accepted as the v1 alternative. HTTP 200 carries payment-response.
curl -i -X POST 'https://grist.tools/v1/pdf-to-text' -H 'content-type: application/json' -H 'payment-signature: <base64 x402 payload>' -d '{"url":"https://grist.tools/samples/documents/report.pdf","max_pages":100}'06Errors
| Code | HTTP | For this service |
|---|---|---|
| invalid_input | 400 | The body does not match the schema; its fields are url and max_pages. |
| blocked_target | 403 | |
| unreachable_target | 424 | |
| upstream_timeout | 424 | |
| unsupported_content_type | 415 | The source is not a type pdf-to-text reads; it reads PDF. |
| too_large | 413 | |
| unprocessable | 422 | |
| internal | 500 |
Input that is rejected before the tool runs (malformed JSON, schema mismatch, body over the size limit) is never charged.
When each code is raised, and what it means for payment: /docs/pdf-to-text.
07Conversions
| Conversion | Recorded sample | Source size | Output size |
|---|---|---|---|
| PDF to plain text | documents/report.pdf | 1.2 KB | 122 bytes |
08Questions
- How much does a pdf-to-text call cost?
- $0.005 per call, paid in USDC on Base via x402, with no account and no API key.
- What does pdf-to-text return?
- A JSON object with 6 top-level fields: source_url, total_pages, extracted_pages, truncated, chars and text. The example on this page is a real call over documents/report.pdf.
- What happens when a pdf-to-text call fails?
- The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled. A 503 upstream_unavailable can follow a settled call; the x402 payments guide covers it, 429, 402 and an uncertain 500.