Convert EPUB to plain text
To convert EPUB to plain text, send one POST to /v1/epub-to-text with the EPUB source. Each call costs $0.005 in USDC on Base via x402, with no account and no API key. The source file can be up to 25 MB.
- Endpoint
POST /v1/epub-to-text- Price
- $0.005 per call
- Service
- epub-to-text
01Limits
| Limit | Value |
|---|---|
| Largest source file | 25 MB (26,214,400 bytes) |
| Timeout | 45 s |
| Source formats | EPUB |
| Output formats | plain text |
| Option max_chapters | default 500 |
02Example
A real EPUB to plain text run of epub-to-text.
POST /v1/epub-to-text
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/documents/seed-saving.epub",
"max_chapters": 500
}200 OK
payment-response: <base64 settlement receipt>
{
"source_url": "https://grist.tools/samples/documents/seed-saving.epub",
"metadata": {
"title": "A Short Guide to Seed Saving",
"creator": "Grist Samples",
"language": "en"
},
"chapter_count": 3,
"extracted_chapters": 3,
"truncated": false,
"chars": 441,
"chapters": [
{
"href": "OEBPS/ch1.xhtml",
"text": "Why Save Seeds\n\nSaving seeds keeps a variety that grows well in your own soil.\n\nIt also costs nothing and makes next spring easier to plan."
},
{
"href": "OEBPS/ch2.xhtml",
"text": "Choosing Plants\n\nPick the healthiest plants and let a few fruits ripen fully.\n\nBeans and peas are the easiest to start with.\n\nTomatoes need their seeds rinsed and dried."
},
{
"href": "OEBPS/ch3.xhtml",
"text": "Drying and Storing\n\nDry seeds on paper for a week, away from direct sun.\n\nStore them in labelled envelopes in a cool, dark place."
}
],
"text": "Why Save Seeds\n\nSaving seeds keeps a variety that grows well in your own soil.\n\nIt also costs nothing and makes next spring easier to plan.\n\nChoosing Plants\n\nPick the healthiest plants and let a few fruits ripen fully.\n\nBeans and peas are the easiest to start with.\n\nTomatoes need their seeds rinsed and dried.\n\nDrying and Storing\n\nDry seeds on paper for a week, away from direct sun.\n\nStore them in labelled envelopes in a cool, dark place."
}Source file: documents/seed-saving.epub, 2.4 KB (2,422 bytes). Three-chapter EPUB 3 book about seed saving, with a nav document outside the spine and Dublin Core title, creator and language, hand-written with pinned zip dates. The request sets max_chapters to 500. The recorded run returned 441 bytes of plain text, 18% of the source size, with 3 chapters and 441 characters.
The returned text runs to 19 lines, 9 blank ones among them, and its longest line, "Saving seeds keeps a variety that grows well in your own soil.", holds 62 characters. Its first line reads "Why Save Seeds" and its last "Store them in labelled envelopes in a cool, dark place." Split on whitespace, it yields 79 tokens over 10 non-blank lines.
03Format notes
EPUB as a source
An EPUB is a ZIP archive whose reading order comes from its OPF package. After the PK 03 04 signature check and a bound on inflated size, Grist follows META-INF/container.xml to the package document; an archive without that pointer is refused as not an EPUB. The spine lists content in reading order, and only items whose media type is XHTML or HTML contribute chapters, so items left out of the spine, such as a navigation document not listed there, are skipped. Chapter paths are resolved inside the archive only, with parent-directory steps clamped at the root. Script and style elements are removed, block elements get their own lines, and whitespace is collapsed. Dublin Core title, creator and language come back separately as metadata. The max_chapters option bounds how many spine chapters are read, and a longer book is cut and marked truncated.
plain text as a target
Plain text is recovered in several ways, each tied to one reader. ocr-image-to-text reads pixels with Tesseract running offline from local files, returns a mean confidence with the text, and offers only the languages whose data ships with the service. pdf-to-text returns the existing text layer page by page and performs no OCR. epub-to-text follows the spine's reading order and adds title, creator and language metadata. pptx-to-text returns slide body text in slide order, leaving out speaker notes. html-clean-text extracts article text with Readability by default and returns Markdown in the same response. Because each reader works differently, the text reflects its source: recognised characters with a confidence score from an image, but characters already present in the file from documents and pages.
04Errors
| Code | HTTP |
|---|---|
invalid_input | 400 |
blocked_target | 403 |
unreachable_target | 424 |
upstream_timeout | 424 |
unsupported_content_type | 415 |
too_large | 413 |
unprocessable | 422 |
internal | 500 |
When each is raised, and what it means for payment: /docs/epub-to-text.
05Questions
- What does it cost to convert EPUB to plain text?
- $0.005 per call, paid in USDC on Base via x402.
- How large can the EPUB file be?
- Up to 25 MB (26,214,400 bytes). Past a limit the call answers too_large (413).
- What happens if the conversion fails?
- The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.