Convert HTML to plain text
To convert HTML to plain text, send one POST to /v1/html-clean-text with the HTML source. Each call costs $0.002 in USDC on Base via x402, with no account and no API key. The source file can be up to 1 MB.
- Endpoint
POST /v1/html-clean-text- Price
- $0.002 per call
- Service
- html-clean-text
01Limits
| Limit | Value |
|---|---|
| Largest source file | 1 MB (1,048,576 bytes) |
| Timeout | 12 s |
| Source formats | HTML |
| Output formats | Markdown, plain text |
| Option include_links | default false |
| Option readability | default true |
| Option max_chars | default 50000 |
02Example
A real HTML to plain text run of html-clean-text.
POST /v1/html-clean-text
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/web/render-article.html",
"include_links": true,
"readability": true,
"max_chars": 2000
}200 OK
payment-response: <base64 settlement receipt>
{
"requested_url": "https://grist.tools/samples/web/render-article.html",
"source_url": "https://grist.tools/samples/web/render-article.html",
"final_url": "https://grist.tools/samples/web/render-article.html",
"content_type": "text/html",
"charset": "utf-8",
"http_status": 200,
"title": "Measuring Rainfall with a Simple Gauge",
"byline": "Sample Author",
"excerpt": "A short guide to reading a cylinder rain gauge at the same hour every day.",
"site_name": "Grist Samples",
"lang": "en",
"lang_source": "html_lang",
"text": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\nPlacing the gauge\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\nMount it level, so the scale reads true.\nKeep the funnel clear of leaves and insects.\nEmpty the tube right after each reading.\nReading the scale\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\nWrite the reading in a notebook with the date and time. A printable schedule helps when several people share the job.",
"markdown": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\n\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\n\n## Placing the gauge\n\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\n\n- Mount it level, so the scale reads true.\n- Keep the funnel clear of leaves and insects.\n- Empty the tube right after each reading.\n\n## Reading the scale\n\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\n\nWrite the reading in a notebook with the date and time. A [printable schedule](https://grist.tools/samples/web/render-print.html) helps when several people share the job.",
"word_count": 178,
"char_count": 933,
"truncated": false,
"readability_applied": true,
"links": [
{
"href": "https://grist.tools/samples/web/render-print.html",
"text": "printable schedule",
"rel": null,
"is_internal": true
}
],
"link_count": 1,
"fetched_at": "2026-09-01T12:00:00.000Z"
}Source file: web/render-article.html, 2.7 KB (2,752 bytes). Short English article page about reading a rain gauge, with a nav bar, footer boilerplate and a full head of description, canonical, hreflang, favicon, Open Graph and Twitter tags. The request sets include_links to true, readability to true and max_chars to 2000. The recorded run returned 933 bytes of plain text, 34% of the source size, with 933 characters, 178 words and language en.
In the sample itself, the document declares language en and the title "Measuring Rainfall with a Simple Gauge"; 1 style element carries 1 CSS rule; the body holds 6 p, 3 a, 3 li, 2 h2, 1 article, 1 footer, 1 h1, 1 header, 1 main, 1 nav and 1 ul elements; no script element; its headings read "Measuring Rainfall with a Simple Gauge", "Placing the gauge" and "Reading the scale"; its 3 links point to render-article.html, render-print.html and render-print.html. Across the whole file, the markup opens 44 tags, 20 of them inside the head. The returned text runs to 10 lines, none of them blank, and its longest line holds 202 characters. Split on whitespace, it yields 178 tokens over 10 non-blank lines.
03Format notes
HTML as a source
HTML can be supplied inline in the html field or fetched from a URL, and exactly one of the two is required. html-clean-text turns it into Markdown and text, and html-to-pdf renders it. When html-to-pdf fetches a page, it refuses a declared content type other than HTML or XHTML, looks past a byte order mark and leading whitespace, and rejects a body that starts with a recognised binary signature such as an archive, image or audio header. The character set is taken from a byte order mark, then the HTTP charset, then a meta tag near the start. The browser never navigates to the caller's URL and runs no page JavaScript. For inline HTML, html-clean-text accepts a base_url to resolve relative links, which is not allowed together with url, where the final fetched address serves as the base.
plain text as a target
Plain text is recovered in several ways, each tied to one reader. ocr-image-to-text reads pixels with Tesseract running offline from local files, returns a mean confidence with the text, and offers only the languages whose data ships with the service. pdf-to-text returns the existing text layer page by page and performs no OCR. epub-to-text follows the spine's reading order and adds title, creator and language metadata. pptx-to-text returns slide body text in slide order, leaving out speaker notes. html-clean-text extracts article text with Readability by default and returns Markdown in the same response. Because each reader works differently, the text reflects its source: recognised characters with a confidence score from an image, but characters already present in the file from documents and pages.
Note: By default (readability: true) only the main article is kept; send readability: false to convert the whole body.
04Errors
| Code | HTTP |
|---|---|
invalid_input | 400 |
blocked_target | 403 |
unreachable_target | 424 |
upstream_status | 424 |
upstream_timeout | 424 |
unsupported_content_type | 415 |
too_large | 413 |
unprocessable | 422 |
internal | 500 |
When each is raised, and what it means for payment: /docs/html-clean-text.
05Questions
- What does it cost to convert HTML to plain text?
- $0.002 per call, paid in USDC on Base via x402.
- How large can the HTML file be?
- Up to 1 MB (1,048,576 bytes). Past a limit the call answers too_large (413).
- What happens if the conversion fails?
- The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.