Turning web pages into LLM-ready text
Turn a public web page into model context with url-to-markdown for article Markdown or html-clean-text for cleaned text and Markdown bounded by max_chars. Add page-metadata, structured-data-extract or link-extract when your agent also needs head metadata, embedded data or page links.
Updated
01Choose the content your model needs
Raw HTML includes navigation, scripts and markup that consume model tokens and require removal. Start by deciding whether the task needs the article, its declared metadata, embedded structured data or links to further pages. Make that choice before assembling the model context, and request the output that serves the task.
Use url-to-markdown when the article is the input you want. It fetches a page and returns article Markdown with headings, lists and resolved links, together with its title, byline and final URL. Its max_chars option cuts the returned Markdown by Unicode code points. Keep the final URL alongside the article when preparing the context.
Calls use POST https://grist.tools/v1/<name> with a JSON body and content-type: application/json. Read the free input schema at GET /v1/<name>/schema before building a request; /openapi.json also publishes the schemas.
02Use cleaned text when you need extraction controls
html-clean-text returns both text and markdown, plus declared metadata. Choose it when you need plain text, control over article extraction, inline HTML input or links from the extracted content. For fetching, supply url for an HTML, XHTML or plain-text page. For source you already hold, supply html instead; exactly one of those inputs is required.
Article extraction is enabled by default through readability. Disabling it converts the whole body, including navigation and footer. The same whole-body fallback applies when article extraction cannot find an article. Decide whether that broader content suits the task before passing the result to the model.
Set max_chars to bound both text and Markdown by Unicode code points; a longer result is cut and flagged as truncated. The example below uses html-clean-text on a sample article about measuring rainfall. It demonstrates fetching with article extraction, a character limit and include_links enabled.
03Select head metadata or embedded structured data
Choose page-metadata for fields declared in the page head: title, description, canonical URL, language, hreflang, favicon and Open Graph metadata. Its canonical, favicon, hreflang and og:image URLs are resolved to absolute against the final redirect destination or the page base; og:url and twitter:image are returned as authored, and a URL that cannot be resolved is null. Use these fields when the agent needs page identity and declared descriptions alongside, or instead of, article content.
Choose structured-data-extract for JSON-LD, Open Graph, microdata and meta tags, plus a normalized merge of the requested JSON-LD, Open Graph and meta sources; microdata is returned but not merged. Set types to select from jsonld, opengraph, microdata and meta, or omit it to request all of them. A source you do not request returns null and contributes nothing to the merge.
Broken JSON-LD scripts are skipped rather than causing extraction to throw. Inspect jsonld_parse_errors when reviewing a result, and keep the extracted sources available when using the merged fields. Ask for article content separately when the model also needs the prose.
04Choose links according to the next action
Enable include_links on html-clean-text when you need links inside the extracted content. Each comes with link text, rel and a same-host flag. Links resolve against the fetched URL or the supplied base_url for inline HTML; without a base for inline input, they remain as authored.
Choose link-extract for page links, with absolute URLs, rel, nofollow and internal-link flags. Set internal_only when you want links whose host equals the host of the final fetched page. The totals then describe that filtered set. Use those results to select further URLs for your agent to process.
The table below compares the endpoints and their declared prices and limits. Use the generated documentation links to inspect the contract for each call you plan to make.
05Services covered
| Endpoint | Summary | Price | Timeout | max_timeout_seconds | Size cap | Max duration |
|---|---|---|---|---|---|---|
| url-to-markdown | Fetches a web page and returns the article as clean markdown, with its title, byline and final URL. | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - |
| html-clean-text | Fetches a web page and returns the article as clean text and markdown, with its declared metadata. | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - |
| page-metadata | Fetches a page and returns its <head> metadata -- title, description, canonical, language, hreflang, favicon, Open Graph and Twitter cards -- with every URL resolved to absolute. | $0.002 | 8 s | 29 | 3 MB (3,145,728 bytes) | - |
| structured-data-extract | Fetches a web page and returns its embedded machine-readable data -- JSON-LD, OpenGraph, microdata and meta tags -- plus a normalised merge across all four. | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - |
| link-extract | Fetches a web page and returns every link on it, resolved to absolute, each tagged with its rel, a nofollow flag and whether it stays on the same site. | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - |
06Limits
- Largest response cap: 5 MB (5,242,880 bytes) on structured-data-extract.
- Smallest response cap: 1 MB (1,048,576 bytes) on url-to-markdown.
- Longest timeout: 12 s on url-to-markdown.
- Shortest timeout: 8 s on page-metadata.
- max_timeout_seconds: from 29 to 33.
07End-to-end example
A real recorded example of html-clean-text, the same one its docs page publishes. Strings in the response longer than 400 characters are cut and end with ....
POST /v1/html-clean-text
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/web/render-article.html",
"include_links": true,
"readability": true,
"max_chars": 2000
}200 OK
payment-response: <base64 settlement receipt>
{
"requested_url": "https://grist.tools/samples/web/render-article.html",
"source_url": "https://grist.tools/samples/web/render-article.html",
"final_url": "https://grist.tools/samples/web/render-article.html",
"content_type": "text/html",
"charset": "utf-8",
"http_status": 200,
"title": "Measuring Rainfall with a Simple Gauge",
"byline": "Sample Author",
"excerpt": "A short guide to reading a cylinder rain gauge at the same hour every day.",
"site_name": "Grist Samples",
"lang": "en",
"lang_source": "html_lang",
"text": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\nPlacing the gauge\nPut the gauge in an open spot, aw ...",
"markdown": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\n\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\n\n## Placing the gauge\n\nPut the gauge in an open sp ...",
"word_count": 178,
"char_count": 933,
"truncated": false,
"readability_applied": true,
"links": [
{
"href": "https://grist.tools/samples/web/render-print.html",
"text": "printable schedule",
"rel": null,
"is_internal": true
}
],
"link_count": 1,
"fetched_at": "2026-09-01T12:00:00.000Z"
}08Chain extraction around the question
For a task that needs an article and its declared identity, start with page-metadata, then send the page URL to html-clean-text or url-to-markdown. Keep the returned title and URL next to the chosen text representation. Add structured-data-extract only when the question also needs embedded data, such as the author or publication fields shown in its normalized output.
For a task that starts from an index page, call link-extract, select relevant returned URLs, and submit those URLs to the article endpoint. Inspect links_truncated before treating the returned link list as complete. If the task only needs references within the article, use html-clean-text with include_links and evaluate whether that already supplies the links you need.
Build the final context from the fields the task needs. Choose text or markdown from the cleaning response, inspect the truncation flag, and attach any selected metadata or structured fields. Keep the endpoint choice explicit at each stage so article extraction, metadata extraction and link selection remain deliberate decisions.
09Budget calls and allow for the payment deadline
A metadata and cleaned-text pipeline costs $0.002 plus $0.002 when both calls succeed. Adding structured data adds $0.002. An article-only path through url-to-markdown costs $0.002, while a discovery step through link-extract adds $0.002. Budget each selected page extraction separately when following links.
Payment uses x402 in USDC on Base, with no account or API key. A request without payment-signature receives HTTP 402 with the requirements, including the price (amount, in atomic units of the asset) and the whole-call maxTimeoutSeconds, before payment; the free /schema route reports the same value as max_timeout_seconds. Use those requirements when deciding whether to proceed with the call.
For html-clean-text, the handler timeout is 12 s, while the whole-call allowance in max_timeout_seconds is 33 seconds. That envelope includes payment verification and settlement around execution. Size the client deadline around the whole call.
Verification happens before execution; settlement follows handler success. Gate refusals and handler failures are never settled. If a settlement call fails without answering or its answer fails validation (that case is HTTP 500 internal, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance, and the error message identifies that case. Preserve that message when deciding what to do next.
10Handle errors before choosing a retry
Use the typed code to decide what needs to change. invalid_input means the body does not match the schema, so correct the request. blocked_target means the network guard refused the destination. Check for a disallowed address, a scheme outside HTTP or HTTPS, a port outside the allowed web ports, or credentials in the URL. Repeating the unchanged request does not address those causes.
unreachable_target means the destination did not resolve, or the connection to it failed or closed before the body arrived. upstream_status means it answered with a status the endpoint cannot use; upstream_timeout means it did not answer within the declared timeout. Check the target and allow for a later retry when the underlying availability problem may have changed.
unsupported_content_type calls for a supported input type. too_large means the response exceeded the streaming size cap or the allowed decompression ratio, or the document exceeded the conversion limits of the endpoint on length, tag count or markup cost. unprocessable means the redirect chain could not be followed (too many hops or an unparseable Location), or bytes arrived but could not be processed, for example a page that yields no text or nests too deeply. Correct the source or endpoint choice where appropriate; reducing max_chars only changes returned text length.
rate_limited means too many calls from the same payer or toward the same destination domain, so reduce request pressure. Retryable HTTP 503 errors, including concurrency_limit, registry_at_capacity and upstream_unavailable, require waiting for the seconds specified by retry-after. When upstream_unavailable comes from the front proxy, the call may already have run and settled; check authorizationState before paying again. internal identifies a service fault; inspect the error message, including any settlement uncertainty, before retrying.
11Keep the extraction boundary visible
The fetched body cap for html-clean-text is 1 MB. Treat that input limit separately from max_chars, which bounds the returned text and Markdown. Check the generated limits for each endpoint you add to a chain; selecting another extraction purpose also means selecting its input contract.
html-clean-text refuses PDF and JSON content. Inline HTML uses base_url to resolve relative links, but base_url is not allowed together with url, where the final fetched URL supplies the base. Public URL fetching remains subject to the network guard, and url-to-markdown rechecks every redirect against private and reserved addresses. Keep unsupported content and refused destinations outside this extraction path.
12Questions
- Which endpoint should supply the text for my model context?
- Use
url-to-markdownfor article Markdown with headings, lists, resolved links, title and byline. Choosehtml-clean-textwhen you want both plain text and Markdown, inline HTML input or control over article extraction. Select the returned representation that the task needs and keep its source URL alongside it. - Can I clean HTML that my agent already has?
- Yes, send it in the
htmlfield tohtml-clean-text, withouturl. Supplybase_urlif relative links need a base for resolution, and enableinclude_linksif you need links from the extracted content. Without a base, those links remain as authored. - Does
max_charsset a model token budget? max_charscounts Unicode code points, so use it as a character limit rather than a model token count. Onhtml-clean-text, it bounds both text and Markdown, with longer output cut and flagged as truncated. Onurl-to-markdown, it cuts the returned Markdown.- Why does cleaned output sometimes include navigation?
html-clean-textconverts the whole body whenreadabilityis disabled or article extraction cannot find an article. That body includes navigation and footer content. Review whether the broader output suits your task before placing it in the model context.- When should I request structured data instead of page metadata?
- Choose
structured-data-extractwhen you need JSON-LD, microdata or the normalized merge of embedded sources. Usetypesto select the sources you need; unrequested sources returnnulland are excluded from the merge. Choosepage-metadatafor head fields such as canonical URL, language,hreflangand favicon. - Should an agent retry every failed extraction?
- No; inspect the typed error and correct invalid input, refused targets or unsuitable content before repeating the request. For retryable HTTP 503 errors, wait for the interval in
retry-after. An upstream availability failure may warrant a later retry, while a settlement call that failed without answering or an answer that failed validation (that case is HTTP 500internal, with a message starting "Payment settlement returned") requires attention to the uncertainty stated in the error message. - Do I need an account before checking the price and schema?
- No account or API key is required, and
GET /v1/<name>/schemais free. A call withoutpayment-signaturereturns HTTP 402 with the price inamountand the whole-call deadline inmaxTimeoutSecondsbefore you pay. Successful calls settle through x402 in USDC on Base after the handler completes.
13Related pages
- URL to Markdown API docs
- HTML Clean Text API docs
- Page Metadata API docs
- Structured Data Extract API docs
- Link Extract API docs
- URL to Markdown API
- HTML Clean Text API
- Page Metadata API
- Structured Data Extract API
- Link Extract API
- Fetch & extract
- API documentation
- Guides
- How AI agents pay for APIs with x402
- Parsing PDF, DOCX, XLSX and EPUB for agents
- Audio and video conversion API without ffmpeg servers