catalogue / tools / Feeds & site structure

Feeds & site structure tools

Feeds and site structure APIs read published feed declarations, feed contents, sitemap URLs, crawl rules and domain files. An agent or script can use them to find a feed from a known page, normalize its items, enumerate a sitemap or inspect the robots.txt rules for a user agent and URL. These operations need network access and parsers for the different document formats.

feed-discover takes a web page and returns the feed links declared in its head, resolved to absolute URLs with their types and titles. rss-parse takes a feed document and returns channel metadata and normalized items, including content, authors, categories and dates in ISO 8601 UTC. sitemap-parse accepts a sitemap, sitemap index or bare origin and produces a flat, bounded URL list; lastmod, changefreq and priority are populated only when include_metadata is true, and are null when it is false (the default).

robots-check fetches robots.txt and reports the matching group, winning rule and competing rules. Conflict resolution uses the longest match, not the first match. wellknown-fetch parses the well-known files of a domain into typed fields, seller records and a cross-file summary. Feed discovery reports declared links rather than parsed feed items. Sitemap traversal is bounded, so it does not fetch an unbounded tree of documents; RSS parsing uses an XML parser protected against external entity attacks.

01APIs

APIWhat it doesPriceMax inputMax durationTimeout
Feed Discover APIFetches a web page and returns the RSS, Atom and JSON feeds it declares in its <head>, as absolute URLs with their type and title.$0.002 per call5 MBnone8 s
Robots Check APIFetches robots.txt and answers whether a user agent may crawl a URL, showing the group, the winning rule and every rule that competed with it.$0.002 per call500 KBnone8 s
RSS Parse APIFetches an RSS, Atom or RSS 1.0 feed and returns its channel metadata and a normalised list of items -- id, title, link, summary, content, author, dates and categories -- with every date in ISO 8601 UTC.$0.002 per call5 MBnone8 s
Sitemap Parse APIParses a sitemap, a sitemap index or a bare origin into a flat, bounded list of URLs with lastmod, changefreq and priority.$0.002 per call10 MBnone14 s
Well-Known Fetch APIFetches and parses a domain's well-known files -- security.txt (RFC 9116), ads.txt and humans.txt -- returning typed fields, seller records and a cross-file summary in one call.$0.002 per call2 MBnone12 s

02Questions

How does feed discovery differ from feed parsing?
feed-discover reads the feed declarations of a web page and returns absolute URLs, types and titles. rss-parse reads the contents of a feed and returns channel metadata and normalized items with dates in ISO 8601 UTC.
How does robots-check resolve competing rules?
It resolves conflicts by longest match, following RFC 9309. The result shows the matching group, the winning rule and every competing rule.
Which inputs does sitemap-parse accept?
It accepts a sitemap, sitemap index or bare origin and returns a flat, bounded list of URLs; lastmod, changefreq and priority are populated only when include_metadata is true, and are null when it is false (the default). It can follow an index across multiple sitemap documents.

03Related