HTML Clean Text API
Fetches a web page and returns the article as clean text and markdown, with its declared metadata.
- Endpoint
POST /v1/html-clean-text- Price
- $0.002 per call
- Family
- Fetch & extract
01When to use it
- It needs the network and a real DOM: boilerplate removal, legacy charset decoding and HTML-to-markdown are three heavy libraries an agent does not carry, and feeding raw HTML to a model instead costs far more in tokens than this call costs.
- Your source is HTML and you need Markdown or plain text.
When not to use it
- The input is over 1 MB (1,048,576 bytes): html-clean-text answers too_large (413).
- The destination needs more than 12 s to answer: the call ends with upstream_timeout (424).
- The destination is on a private or local network: the network guard refuses it with blocked_target (403).
02Parameters
| Field | Type | Required | Description |
|---|---|---|---|
| url | string | no | URL of an HTML, XHTML or plain-text page to fetch, up to 2048 characters; redirects are followed, the body is capped at 1 MB, and exactly one of url or html is required. |
| html | string | no | Inline HTML source to clean instead of fetching, up to 262,144 characters; binary content such as PDF or JSON is refused. |
| base_url | string | no | Base URL for resolving relative links in the html input; not allowed together with url, where the final fetched URL is the base. |
| include_links | boolean | no | Also return the links inside the extracted content, each with its text, rel and same-host flag, resolved against the fetched URL or base_url when there is one, else kept as authored. Default: false. |
| readability | boolean | no | Run Readability to keep only the main article; false, or an article Readability cannot find, converts the whole body, navigation and footer included. Default: true. |
| max_chars | integer | no | Maximum length of text and markdown, 200 to 200,000 Unicode code points (default 50,000); a longer result is cut and flagged as truncated. Default: 50000. |
03Limits
| Limit | Value |
|---|---|
| Largest source file | 1 MB (1,048,576 bytes) |
| Timeout | 12 s |
| Source formats | HTML |
| Output formats | Markdown, plain text |
04Example
The published example of html-clean-text, verbatim: the request body and the response it returns.
POST /v1/html-clean-text
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/web/render-article.html",
"include_links": true,
"readability": true,
"max_chars": 2000
}200 OK
payment-response: <base64 settlement receipt>
{
"requested_url": "https://grist.tools/samples/web/render-article.html",
"source_url": "https://grist.tools/samples/web/render-article.html",
"final_url": "https://grist.tools/samples/web/render-article.html",
"content_type": "text/html",
"charset": "utf-8",
"http_status": 200,
"title": "Measuring Rainfall with a Simple Gauge",
"byline": "Sample Author",
"excerpt": "A short guide to reading a cylinder rain gauge at the same hour every day.",
"site_name": "Grist Samples",
"lang": "en",
"lang_source": "html_lang",
"text": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\nPlacing the gauge\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\nMount it level, so the scale reads true.\nKeep the funnel clear of leaves and insects.\nEmpty the tube right after each reading.\nReading the scale\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\nWrite the reading in a notebook with the date and time. A printable schedule helps when several people share the job.",
"markdown": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\n\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\n\n## Placing the gauge\n\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\n\n- Mount it level, so the scale reads true.\n- Keep the funnel clear of leaves and insects.\n- Empty the tube right after each reading.\n\n## Reading the scale\n\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\n\nWrite the reading in a notebook with the date and time. A [printable schedule](https://grist.tools/samples/web/render-print.html) helps when several people share the job.",
"word_count": 178,
"char_count": 933,
"truncated": false,
"readability_applied": true,
"links": [
{
"href": "https://grist.tools/samples/web/render-print.html",
"text": "printable schedule",
"rel": null,
"is_internal": true
}
],
"link_count": 1,
"fetched_at": "2026-09-01T12:00:00.000Z"
}Sample file: web/render-article.html. Short English article page about reading a rain gauge, with a nav bar, footer boilerplate and a full head of description, canonical, hreflang, favicon, Open Graph and Twitter tags. Sample file: web/render-print.html. Printable one-page HTML table with inline CSS: a weekly rain gauge reading schedule.
In web/render-article.html, the document declares language en and the title "Measuring Rainfall with a Simple Gauge"; 1 style element carries 1 CSS rule; the body holds 6 p, 3 a, 3 li, 2 h2, 1 article, 1 footer, 1 h1, 1 header, 1 main, 1 nav and 1 ul elements; no script element; its headings read "Measuring Rainfall with a Simple Gauge", "Placing the gauge" and "Reading the scale"; its 3 links point to render-article.html, render-print.html and render-print.html. Across the whole file, the markup opens 44 tags, 20 of them inside the head. In web/render-print.html, the document declares language en and the title "Rain Gauge Reading Schedule"; 1 style element carries 8 CSS rules; the body holds 28 td, 8 tr, 4 th, 1 a, 1 footer, 1 h1, 1 p, 1 table, 1 tbody and 1 thead elements; no script element; its heading reads "Rain Gauge Reading Schedule"; the table header cells are Day, Reader, Reading (mm) and Notes; its one link points to render-article.html. Across the whole file, the markup opens 53 tags, 3 of them inside the head. The returned Markdown splits into 8 blocks separated by blank lines, with 2 headings (2 at level 2) reading "Placing the gauge" and "Reading the scale" and 3 bullet list items. Plain paragraphs make up 5 blocks. Its headings total 6 words, its paragraphs 157 words.
| Response field | Type | In the example |
|---|---|---|
| requested_url | string | "https://grist.tools/samples/web/render-article.html" |
| source_url | string | "https://grist.tools/samples/web/render-article.html" |
| final_url | string | "https://grist.tools/samples/web/render-article.html" |
| content_type | string | "text/html" |
| charset | string | "utf-8" |
| http_status | number | 200 |
| title | string | "Measuring Rainfall with a Simple Gauge" |
| byline | string | "Sample Author" |
| excerpt | string | "A short guide to reading a cylinder rain gauge at the same hour every day." |
| site_name | string | "Grist Samples" |
| lang | string | "en" |
| lang_source | string | "html_lang" |
| text | string | a string of 933 characters |
| markdown | string | a string of 1011 characters |
| word_count | number | 178 |
| char_count | number | 933 |
| truncated | boolean | false |
| readability_applied | boolean | true |
| links | array | 1 item, each with href, text, rel and is_internal |
| link_count | number | 1 |
| fetched_at | string | "2026-09-01T12:00:00.000Z" |
05Call it from code
// html-clean-text: $0.002 per call, paid in USD Coin on eip155:8453 via x402
// Install: npm install @x402/fetch@2 @x402/evm@2 viem@2
// Save as client.mjs (ES module, Node 18+), export EVM_PRIVATE_KEY with the paying wallet's key in your shell, then run: node client.mjs
// The wrapper reads the 402 (payment-required), signs and retries with payment-signature.
import { wrapFetchWithPayment, x402Client, decodePaymentResponseHeader } from '@x402/fetch'
import { ExactEvmScheme } from '@x402/evm/exact/client'
import { privateKeyToAccount } from 'viem/accounts'
const account = privateKeyToAccount(process.env.EVM_PRIVATE_KEY)
const client = new x402Client().register('eip155:8453', new ExactEvmScheme(account))
const fetchWithPayment = wrapFetchWithPayment(fetch, client)
const body = {
"url": "https://grist.tools/samples/web/render-article.html",
"include_links": true,
"readability": true,
"max_chars": 2000
}
const res = await fetchWithPayment('https://grist.tools/v1/html-clean-text', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
})
console.log(res.status, await res.json())
const settle = res.headers.get('payment-response')
if (settle) console.log(decodePaymentResponseHeader(settle).transaction)# html-clean-text: $0.002 per call, paid in USD Coin on eip155:8453 via x402
# Install: pip install "x402[requests,evm]>=2.24,<3"
# The session reads the 402 (payment-required), signs and retries with payment-signature.
import os
from eth_account import Account
from x402 import x402ClientSync
from x402.http import decode_payment_response_header
from x402.http.clients.requests import x402_requests
from x402.mechanisms.evm.exact import register_exact_evm_client
from x402.mechanisms.evm.signers import EthAccountSigner
account = Account.from_key(os.environ["EVM_PRIVATE_KEY"])
client = x402ClientSync()
register_exact_evm_client(client, EthAccountSigner(account), networks="eip155:8453")
session = x402_requests(client)
payload = {
"url": "https://grist.tools/samples/web/render-article.html",
"include_links": True,
"readability": True,
"max_chars": 2000
}
res = session.post("https://grist.tools/v1/html-clean-text", json=payload)
print(res.status_code, res.json())
settle = res.headers.get("payment-response")
if settle:
print(decode_payment_response_header(settle).transaction)# html-clean-text: $0.002 per call, paid in USD Coin on eip155:8453 via x402
# 1. Unpaid call: HTTP 402, the requirements in the payment-required header (base64 JSON) and in the body.
curl -i -X POST 'https://grist.tools/v1/html-clean-text' -H 'content-type: application/json' -d '{"url":"https://grist.tools/samples/web/render-article.html","include_links":true,"readability":true,"max_chars":2000}'
# 2. Same call with the signed payment (an EIP-3009 authorization, EIP-712 signed: it cannot be
# typed by hand). x-payment is accepted as the v1 alternative. HTTP 200 carries payment-response.
curl -i -X POST 'https://grist.tools/v1/html-clean-text' -H 'content-type: application/json' -H 'payment-signature: <base64 x402 payload>' -d '{"url":"https://grist.tools/samples/web/render-article.html","include_links":true,"readability":true,"max_chars":2000}'06Errors
| Code | HTTP | For this service |
|---|---|---|
| invalid_input | 400 | The body does not match the schema; its fields are url, html, base_url, include_links, readability and max_chars. |
| blocked_target | 403 | |
| unreachable_target | 424 | |
| upstream_status | 424 | |
| upstream_timeout | 424 | |
| unsupported_content_type | 415 | The source is not a type html-clean-text reads; it reads HTML. |
| too_large | 413 | |
| unprocessable | 422 | |
| internal | 500 |
Input that is rejected before the tool runs (malformed JSON, schema mismatch, body over the size limit) is never charged.
When each code is raised, and what it means for payment: /docs/html-clean-text.
07Conversions
| Conversion | Recorded sample | Source size | Output size |
|---|---|---|---|
| HTML to Markdown | web/render-article.html | 2.7 KB | 1,011 bytes |
| HTML to plain text | web/render-article.html | 2.7 KB | 933 bytes |
08Questions
- How much does a html-clean-text call cost?
- $0.002 per call, paid in USDC on Base via x402, with no account and no API key.
- What does html-clean-text return?
- A JSON object with 21 top-level fields: requested_url, source_url, final_url, content_type, charset, http_status, title, byline, excerpt, site_name, lang, lang_source, text, markdown, word_count, char_count, truncated, readability_applied, links, link_count and fetched_at. The example on this page is a real call over web/render-article.html and web/render-print.html.
- What happens when a html-clean-text call fails?
- The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled. A 503 upstream_unavailable can follow a settled call; the x402 payments guide covers it, 429, 402 and an uncertain 500.