Fetch & extract tools
Fetch and extract APIs turn web URLs into article content, page declarations, link lists or HTTP response details. An agent or script can call them when it needs readable content from a page, metadata for that page, or the destination behind a shortened link. The extraction work uses network access and HTML parsing, including boilerplate removal, legacy charset decoding and URL resolution where described by the member.
html-clean-text returns article text and Markdown with declared metadata. url-to-markdown returns article Markdown with headings, lists and resolved links, plus the title, byline and final URL. link-extract returns absolute links with rel values, a nofollow flag and same-site classification. page-metadata reads head declarations such as title, description, canonical URL, language and social cards. structured-data-extract returns JSON-LD, Open Graph, microdata and meta tags, together with a normalized merge.
For HTTP inspection, http-headers takes a URL and returns headers, final status, the redirect chain and a security-header grade without downloading the body. url-unshorten follows a shortened or tracking URL and reports its destination and every hop. Article extraction removes navigation, scripts and other surrounding markup; url-to-markdown does not return the entire raw page. The header grade follows a deterministic, documented rubric, while the extraction members report content and declarations from the fetched page.
01APIs
| API | What it does | Price | Max input | Max duration | Timeout |
|---|---|---|---|---|---|
| HTML Clean Text API | Fetches a web page and returns the article as clean text and markdown, with its declared metadata. | $0.002 per call | 1 MB | none | 12 s |
| HTTP Headers API | Returns the response headers, final status, redirect chain and a security-header grade for a URL, without downloading the body. | $0.002 per call | 64 KB | none | 8 s |
| Link Extract API | Fetches a web page and returns every link on it, resolved to absolute, each tagged with its rel, a nofollow flag and whether it stays on the same site. | $0.002 per call | 5 MB | none | 8 s |
| Page Metadata API | Fetches a page and returns its <head> metadata -- title, description, canonical, language, hreflang, favicon, Open Graph and Twitter cards -- with every URL resolved to absolute. | $0.002 per call | 3 MB | none | 8 s |
| Structured Data Extract API | Fetches a web page and returns its embedded machine-readable data -- JSON-LD, OpenGraph, microdata and meta tags -- plus a normalised merge across all four. | $0.002 per call | 5 MB | none | 8 s |
| URL to Markdown API | Fetches a web page and returns the article as clean markdown, with its title, byline and final URL. | $0.002 per call | 1 MB | none | 12 s |
| URL Unshorten API | Follows a shortened or tracking URL to its real destination and returns every hop it went through. | $0.002 per call | 64 KB | none | 10 s |
02Conversions
03Questions
- How do html-clean-text and url-to-markdown differ?
- html-clean-text returns the article as clean text and Markdown with declared metadata. url-to-markdown returns article Markdown with its title, byline and final URL, preserving headings, lists and resolved links.
- Can response headers be inspected without downloading the body?
- http-headers returns response headers, final status, the redirect chain and a security-header grade without downloading the body. The grade uses a deterministic, documented rubric.
- What does structured-data-extract return?
- It returns embedded JSON-LD, Open Graph, microdata and meta tags, plus a normalized merge. A broken JSON-LD script is skipped instead of causing the extraction to throw.