# Grist > Deterministic HTTP micro-utilities for AI agents, paid per call via x402 (USD Coin on eip155:8453). > One endpoint does one thing, always the same way, for a fraction of a cent. ## How a call works 1. POST https://grist.tools/v1/ with a JSON body. 2. A paid endpoint answers 402 with the payment requirements, including the price and max_timeout_seconds for that endpoint, so the deadline is known before paying. 3. Repeat the call with a PAYMENT-SIGNATURE header (X-PAYMENT is also accepted). The payment is verified before the handler runs and settled only after it succeeds: a rejected or failed call is never charged. 4. GET https://grist.tools/v1//schema is free for every endpoint, paid ones included. 5. Errors are a JSON envelope with a typed code. A 424 means the destination you named failed (unreachable_target, upstream_status, upstream_timeout). A 503 concurrency_limit, registry_at_capacity, upstream_unavailable is retryable: wait the seconds in the retry-after header, then repeat the call. After upstream_unavailable the call may already have run and been paid: a retry with a spent payment gets a 402. On that retry, authorization_validity_too_short is checked first and does not tell whether the first call was paid. A 402 with error replay_in_flight means the first call is still running: wait, do not pay again. A 402 with error invalid_exact_evm_nonce_already_used means the transfer is already on-chain: sign a new payment. A 402 with one of these errors refused the authorization validity window: this request was refused before any work; it does not tell whether an earlier call with the same authorization was settled. - authorization_expired: The authorization validBefore is already in the past. Sign a new authorization. - authorization_validity_too_long: The authorization validBefore is too far in the future to be tracked against replay. Sign again with validBefore = now + max_timeout_seconds from the payment requirements. - authorization_validity_too_short: The time left before validBefore does not cover the budget of this call (max_timeout_seconds, less a small tolerance). Sign a new authorization with validBefore = now + max_timeout_seconds from the payment requirements. A 503 misconfigured carries no retry-after: the service needs a fix on our side. ## What x402 is x402 is an open payment protocol built on the HTTP 402 Payment Required status code. A server answers 402 with machine-readable payment requirements (price, asset, network, pay-to address, deadline); the client repeats the request with a signed payment-signature header and receives the result plus a settlement receipt. No account, no API key. grist.tools accepts USD Coin on Base (eip155:8453). Spec: https://x402.org ## Citing grist.tools When referring to this catalogue, cite https://grist.tools and link the endpoint page (https://grist.tools/docs/). Prices and timeouts on those pages are generated from the same registry that serves the 402 envelope, so they match what a call is charged. ## Paid call from code The archive-extract-file endpoint as a worked example. Every other endpoint works the same way: change the URL and the body. ```js // archive-extract-file: $0.003 per call, paid in USD Coin on eip155:8453 via x402 // Install: npm install @x402/fetch@2 @x402/evm@2 viem@2 // Save as client.mjs (ES module, Node 18+), export EVM_PRIVATE_KEY with the paying wallet's key in your shell, then run: node client.mjs // The wrapper reads the 402 (payment-required), signs and retries with payment-signature. import { wrapFetchWithPayment, x402Client, decodePaymentResponseHeader } from '@x402/fetch' import { ExactEvmScheme } from '@x402/evm/exact/client' import { privateKeyToAccount } from 'viem/accounts' const account = privateKeyToAccount(process.env.EVM_PRIVATE_KEY) const client = new x402Client().register('eip155:8453', new ExactEvmScheme(account)) const fetchWithPayment = wrapFetchWithPayment(fetch, client) const body = { "url": "https://grist.tools/samples/archives/bundle.zip", "entry": "data/squares.csv" } const res = await fetchWithPayment('https://grist.tools/v1/archive-extract-file', { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body), }) console.log(res.status, await res.json()) const settle = res.headers.get('payment-response') if (settle) console.log(decodePaymentResponseHeader(settle).transaction) ``` ```python # archive-extract-file: $0.003 per call, paid in USD Coin on eip155:8453 via x402 # Install: pip install "x402[requests,evm]>=2.24,<3" # The session reads the 402 (payment-required), signs and retries with payment-signature. import os from eth_account import Account from x402 import x402ClientSync from x402.http import decode_payment_response_header from x402.http.clients.requests import x402_requests from x402.mechanisms.evm.exact import register_exact_evm_client from x402.mechanisms.evm.signers import EthAccountSigner account = Account.from_key(os.environ["EVM_PRIVATE_KEY"]) client = x402ClientSync() register_exact_evm_client(client, EthAccountSigner(account), networks="eip155:8453") session = x402_requests(client) payload = { "url": "https://grist.tools/samples/archives/bundle.zip", "entry": "data/squares.csv" } res = session.post("https://grist.tools/v1/archive-extract-file", json=payload) print(res.status_code, res.json()) settle = res.headers.get("payment-response") if settle: print(decode_payment_response_header(settle).transaction) ``` ```sh # archive-extract-file: $0.003 per call, paid in USD Coin on eip155:8453 via x402 # 1. Unpaid call: HTTP 402, the requirements in the payment-required header (base64 JSON) and in the body. curl -i -X POST 'https://grist.tools/v1/archive-extract-file' -H 'content-type: application/json' -d '{"url":"https://grist.tools/samples/archives/bundle.zip","entry":"data/squares.csv"}' # 2. Same call with the signed payment (an EIP-3009 authorization, EIP-712 signed: it cannot be # typed by hand). x-payment is accepted as the v1 alternative. HTTP 200 carries payment-response. curl -i -X POST 'https://grist.tools/v1/archive-extract-file' -H 'content-type: application/json' -H 'payment-signature: ' -d '{"url":"https://grist.tools/samples/archives/bundle.zip","entry":"data/squares.csv"}' ``` ## The three rules 1. No third-party marginal cost. No paid upstream APIs, no model tokens, no billed external service. 2. Determinism. The same input returns the same output. Clock-derived fields are allowed only when the endpoint declares them in its output. 3. It must require the network or a heavy dependency. Agents do not pay for a hash or a uuid. ## Guides ### How AI agents pay for APIs with x402 URL: https://grist.tools/guides/x402-payments-for-ai-agents (updated 2026-09-28) An AI agent pays for a paid API call by reading the HTTP 402 requirements, signing a payment authorization and resending the request with `payment-signature`. On grist.tools, payment uses USDC on Base without an account or API key, with verification before execution and settlement only after the handler succeeds. #### Discover the endpoint before authorizing payment Start with `/.well-known/x402`, which lists endpoints with their payment requirements and metadata, including name, description, family, price, free status, deadline, schema URL and documentation URL. The catalogue covers 50 endpoints. You can also discover services through `/openapi.json`, `/llms.txt`, `/llms-full.txt` and `/index.json`. Read `GET /v1//schema` before constructing the body. This free response includes the input schema, input and output examples, errors, price and timeouts. Use it to select the endpoint and prepare valid JSON before signing. The table below compares declared service prices and limits and links to each endpoint contract. Free endpoints take a plain call without verification or settlement. For example, the price of `dns-lookup` is free, and its manifest entry has an empty `accepts` array. Do not require a payment challenge on that path. Budget each successful extraction separately when processing several pages through `url-to-markdown`. Its per-call price is $0.002; read each additional endpoint contract before adding another operation. #### Read the unpaid response as payment requirements Send `POST https://grist.tools/v1/url-to-markdown` with a JSON body and `content-type: application/json`. Without a payment header, a paid endpoint returns HTTP 402 after checking the request body size. This happens before schema validation, so the challenge does not confirm that your input is valid. The body contains `x402Version`, `error`, `resource` and `accepts`. The resource describes the requested URL, service, response type and family; `accepts` supplies payment requirements. The `payment-required` response header contains the same envelope as base64-encoded JSON. The generated example below uses `page-metadata` and shows its envelope alongside the utility request and response. The requirements use x402 version 2 and the `exact` scheme. Inspect `network`, `asset`, `amount`, `payTo` and `maxTimeoutSeconds`. Payment runs on Base; the price of `url-to-markdown` is $0.002. Read the atomic-unit `amount` from the requirements when preparing payment. #### Sign an authorization that matches the challenge The payment payload carries an EIP-3009 `transferWithAuthorization`, signed using EIP-712. Its fields include `from`, `to`, `value`, `nonce` and `validBefore`. The recipient and value are checked against `payTo` and `amount`. Use the advertised requirements to construct the authorization, and keep its nonce available for settlement checks. Encode the JSON `PaymentPayload` as base64 and send it in `payment-signature`. The older `x-payment` header remains accepted as an alias; if both headers are present, `payment-signature` takes precedence. That alias does not remove payload checks: malformed payloads, unsupported versions, mismatched requirements and invalid validity windows receive HTTP 402 with a reason in `error`. In the JavaScript documentation, `x402Client` registers the advertised network with `ExactEvmScheme(account)`, and `wrapFetchWithPayment` supplies the payment wrapper. Use the endpoint example when connecting these pieces. Keep the advertised recipient, amount and deadline visible in your integration so you can compare them with a refused authorization. #### Submit the paid request and retain the receipt Resend the JSON request with the payment header. The server checks the body against the schema before payment verification. After verification come the payer limit, replay reservation and destination-domain limit, followed by the handler. A facilitator verifies the payment before work and settles it after handler success. A verification result of `isValid: false` returns HTTP 402 with the reason, before the handler runs. A successful paid response returns HTTP 200, the utility output and a receipt in `payment-response`. The `x-payment-response` alias carries the same value. Retain the receipt with the result when recording a completed call. On `/v1/*`, CORS exposes `payment-required`, `payment-response` and `x-payment-response`, allowing browser clients to read those response headers. #### Services covered | Endpoint | Price | Timeout | max_timeout_seconds | Size cap | Max duration | | --- | --- | --- | --- | --- | --- | | [archive-extract-file](https://grist.tools/docs/archive-extract-file) | $0.003 | 90 s | 111 | 25 MB (26,214,400 bytes) | - | | [archive-inspect](https://grist.tools/docs/archive-inspect) | $0.003 | 60 s | 81 | 100 MB (104,857,600 bytes) | - | | [audio-convert](https://grist.tools/docs/audio-convert) | $0.003 | 94 s | 115 | 100 MB (104,857,600 bytes) | 60 min (3,600 s) | | [audio-extract](https://grist.tools/docs/audio-extract) | $0.003 | 94 s | 115 | 200 MB (209,715,200 bytes) | 60 min (3,600 s) | | [color-palette-extract](https://grist.tools/docs/color-palette-extract) | $0.003 | 60 s | 81 | 25 MB (26,214,400 bytes) | - | | [csv-to-json](https://grist.tools/docs/csv-to-json) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [dns-lookup](https://grist.tools/docs/dns-lookup) | free | 5 s | 6 | 2 MB (2,097,152 bytes) | - | | [docx-to-markdown](https://grist.tools/docs/docx-to-markdown) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [email-validate](https://grist.tools/docs/email-validate) | $0.002 | 5 s | 26 | 2 MB (2,097,152 bytes) | - | | [epub-to-text](https://grist.tools/docs/epub-to-text) | $0.005 | 45 s | 66 | 25 MB (26,214,400 bytes) | - | | [favicon-extract](https://grist.tools/docs/favicon-extract) | $0.003 | 45 s | 66 | 25 MB (26,214,400 bytes) | - | | [feed-discover](https://grist.tools/docs/feed-discover) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | | [file-type-detect](https://grist.tools/docs/file-type-detect) | $0.003 | 30 s | 51 | 32 MB (33,554,432 bytes) | - | | [html-clean-text](https://grist.tools/docs/html-clean-text) | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - | | [html-to-pdf](https://grist.tools/docs/html-to-pdf) | $0.005 | 12 s | 33 | 5 MB (5,242,880 bytes) | - | | [http-headers](https://grist.tools/docs/http-headers) | $0.002 | 8 s | 29 | 64 KB (65,536 bytes) | - | | [image-compress](https://grist.tools/docs/image-compress) | $0.003 | 60 s | 81 | 25 MB (26,214,400 bytes) | - | | [image-convert](https://grist.tools/docs/image-convert) | $0.003 | 60 s | 81 | 25 MB (26,214,400 bytes) | - | | [image-metadata](https://grist.tools/docs/image-metadata) | $0.003 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [image-probe](https://grist.tools/docs/image-probe) | $0.003 | 20 s | 41 | 25 MB (26,214,400 bytes) | - | | [image-resize](https://grist.tools/docs/image-resize) | $0.003 | 60 s | 81 | 25 MB (26,214,400 bytes) | - | | [ip-info](https://grist.tools/docs/ip-info) | $0.002 | 5 s | 26 | 2 MB (2,097,152 bytes) | - | | [json-to-csv](https://grist.tools/docs/json-to-csv) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [link-extract](https://grist.tools/docs/link-extract) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | | [markdown-to-pdf](https://grist.tools/docs/markdown-to-pdf) | $0.005 | 12 s | 33 | 2 MB (2,097,152 bytes) | - | | [media-probe](https://grist.tools/docs/media-probe) | $0.003 | 30 s | 51 | 100 MB (104,857,600 bytes) | - | | [ocr-image-to-text](https://grist.tools/docs/ocr-image-to-text) | $0.005 | 45 s | 66 | 10 MB (10,485,760 bytes) | - | | [page-metadata](https://grist.tools/docs/page-metadata) | $0.002 | 8 s | 29 | 3 MB (3,145,728 bytes) | - | | [pdf-merge](https://grist.tools/docs/pdf-merge) | $0.005 | 45 s | 66 | 8 MB (8,388,608 bytes) | - | | [pdf-metadata](https://grist.tools/docs/pdf-metadata) | $0.005 | 25 s | 46 | 25 MB (26,214,400 bytes) | - | | [pdf-split](https://grist.tools/docs/pdf-split) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [pdf-to-markdown](https://grist.tools/docs/pdf-to-markdown) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [pdf-to-text](https://grist.tools/docs/pdf-to-text) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [pptx-to-text](https://grist.tools/docs/pptx-to-text) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [remote-file-hash](https://grist.tools/docs/remote-file-hash) | $0.003 | 60 s | 81 | 100 MB (104,857,600 bytes) | - | | [robots-check](https://grist.tools/docs/robots-check) | $0.002 | 8 s | 29 | 500 KB (512,000 bytes) | - | | [rss-parse](https://grist.tools/docs/rss-parse) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | | [sitemap-parse](https://grist.tools/docs/sitemap-parse) | $0.002 | 14 s | 35 | 10 MB (10,485,760 bytes) | - | | [ssl-cert-info](https://grist.tools/docs/ssl-cert-info) | $0.002 | 10 s | 31 | 2 MB (2,097,152 bytes) | - | | [structured-data-extract](https://grist.tools/docs/structured-data-extract) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | | [subtitle-convert](https://grist.tools/docs/subtitle-convert) | $0.003 | 15 s | 36 | 5 MB (5,242,880 bytes) | - | | [subtitle-extract](https://grist.tools/docs/subtitle-extract) | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - | | [svg-to-png](https://grist.tools/docs/svg-to-png) | $0.003 | 60 s | 81 | 25 MB (26,214,400 bytes) | - | | [url-to-markdown](https://grist.tools/docs/url-to-markdown) | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - | | [url-unshorten](https://grist.tools/docs/url-unshorten) | $0.002 | 10 s | 31 | 64 KB (65,536 bytes) | - | | [video-thumbnail](https://grist.tools/docs/video-thumbnail) | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - | | [waveform-data](https://grist.tools/docs/waveform-data) | $0.003 | 90 s | 111 | 100 MB (104,857,600 bytes) | - | | [wellknown-fetch](https://grist.tools/docs/wellknown-fetch) | $0.002 | 12 s | 33 | 2 MB (2,097,152 bytes) | - | | [whois](https://grist.tools/docs/whois) | $0.002 | 12 s | 33 | 2 MB (2,097,152 bytes) | - | | [xlsx-to-json](https://grist.tools/docs/xlsx-to-json) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | #### Pay for an API call with x402 1. Discover the endpoint: Read `/.well-known/x402` and the free `GET /v1//schema` response. Prepare a JSON body matching the input schema. 2. Request the payment requirements: POST the JSON body to the paid endpoint with `content-type: application/json` and without a payment header. 3. Inspect the challenge: Read HTTP 402 and inspect `accepts` for `network`, `asset`, `amount`, `payTo` and `maxTimeoutSeconds` before authorizing payment. 4. Sign the authorization: Create the EIP-712 signed EIP-3009 authorization matching the requirements, with validity covering the whole-call allowance. Retain its nonce. 5. Resubmit with payment: Base64-encode the JSON `PaymentPayload` and resend the request with that value in `payment-signature`. 6. Record the result: On HTTP 200, retain the output and `payment-response` receipt. Otherwise, inspect the error and any settlement uncertainty before retrying. #### Use the whole-call deadline for payment validity The handler timeout for `url-to-markdown` is 12 s. Its whole-call allowance is 33 seconds, published as `max_timeout_seconds` in discovery metadata and `maxTimeoutSeconds` in the payment requirements. That allowance includes verification and settlement around the handler. Size the client deadline around the whole call. The authorization uses `validBefore` in Unix seconds. Its remaining validity must cover the whole-call allowance, subject to a small tolerance. Too little remaining time produces `authorization_validity_too_short`; a deadline in the past produces `authorization_expired`. An excessively distant deadline produces `authorization_validity_too_long`. When correcting the validity window, follow the instruction to sign again with `validBefore = now + max_timeout_seconds`. Check the requirements for the endpoint you are calling instead of carrying a handler timeout from another endpoint into the authorization. #### Distinguish execution failure from uncertain settlement Verification does not settle the payment. Gate refusals and handler failures are never settled, and a handler failure returns its typed error without a charge. If the client disconnects before settlement, settlement is skipped. These rules make the point of failure relevant when interpreting an unsuccessful call. If the handler succeeds but the settle call fails without an answer, or returns a `tx_timeout` that does not qualify for delivery (no usable transaction hash, a different network, or a payer or amount mismatch), the response is HTTP 500 `internal` with the message "The work was done but settlement could not be completed". A facilitator answer that fails validation also returns HTTP 500 `internal`, with a message starting "Payment settlement returned". In these cases the payment outcome is uncertain and the output is not delivered. Preserve the message and check the on-chain EIP-3009 `authorizationState` for the nonce before signing a new payment. A `tx_timeout` with a well-formed transaction hash on the charged network and no payer or amount mismatch returns HTTP 200 with the utility output and a `payment-response` receipt whose `success` is `false`; the transfer is in flight and may still revert. Any other failed settlement (a reason other than `tx_timeout`) returns HTTP 402 with the reason in the `error` field, the `payment-response` header and no utility output. Read the reason rather than treating every 402 as an initial request for payment. Keep the response status, error and nonce together when investigating a failed attempt. #### Handle replay responses before signing again Each authorization is tracked by its nonce, and the EIP-3009 nonce can be settled only once. Reusing an authorization can therefore produce a replay error. Use the reason to decide what happens next; automatically signing again on every payment error ignores whether earlier work is still running or may already have settled. `replay_in_flight` means the first call is still running: wait and do not pay again. `replay_consumed` means the authorization was already used and may have been settled; another call requires a new payment. `replay_exhausted` means the authorization has reached its attempt limit. To determine whether the authorization was settled, read the on-chain EIP-3009 `authorizationState` for its nonce. Use that check when a consumed authorization or uncertain settlement leaves the result unclear. Retaining the nonce gives your retry handling a reference to the payment whose outcome you need to establish. #### Correct request failures and respect retry signals Utility calls require POST; other methods receive HTTP 405 `method_not_allowed`. A content type other than `application/json` receives HTTP 415 before payment. An oversized request body receives HTTP 413 `too_large`, with the cap in `details.limit_bytes`, also before payment. Schema failure with a payment header returns HTTP 400 `invalid_input` before verification. HTTP 429 `rate_limited` indicates too many calls from a payer or toward a destination domain. Nothing is charged; use `retry-after` to schedule a later attempt. Retryable HTTP 503 errors, including `concurrency_limit`, `registry_at_capacity` and `upstream_unavailable`, also require waiting for the interval in that header. When `upstream_unavailable` comes from the front proxy, the call may already have run and settled; check `authorizationState` before paying again. For handler errors, use the typed code to identify the cause. `blocked_target` means the network guard refused the destination; `unreachable_target` means resolution or connection failed. `upstream_timeout` means the call exceeded the declared handler timeout of the endpoint, while `unprocessable` means the received bytes could not be processed. Correct the source or request where appropriate before retrying. `unsupported_content_type` means the destination returned an unsupported type. `too_large` can identify a fetched response exceeding the streaming size cap or allowed decompression ratio. `upstream_status` means the destination answered with a status the endpoint cannot use. #### Questions Q: Do I need an account or API key to pay? A: No account or API key is required. A paid endpoint advertises its requirements in HTTP 402 before payment, and `GET /v1//schema` is free. The agent supplies a signed payment authorization when submitting the paid call. Q: Which header carries the signed payment? A: Send base64-encoded JSON `PaymentPayload` in `payment-signature`. The older `x-payment` header is accepted, but `payment-signature` wins if both are present. Read `payment-required` for the challenge and `payment-response` for the successful payment receipt. Q: Does receiving HTTP 402 mean my input passed validation? A: No; an unpaid paid-endpoint request receives its challenge after the body size check and before schema validation. With a payment header, the body is validated before verification. Read the free schema and correct `invalid_input` errors before resubmitting. Q: Am I charged when the handler fails? A: No; handler failures are never settled, and payment verification happens before execution. A settlement request that fails without an answer or a settlement answer that fails validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned") is a separate case: the work completed, but the payment outcome is uncertain. Preserve that error and check the nonce through `authorizationState`. Q: Can I reuse a payment authorization for another call? A: An authorization nonce can be settled only once. `replay_in_flight` means wait without paying again, while `replay_consumed` means it was used and may have settled. Check its on-chain `authorizationState` when the outcome is unclear before signing a new payment. Q: Which timeout belongs in my authorization? A: Use the whole-call allowance advertised in the payment requirements as `maxTimeoutSeconds`, also published as `max_timeout_seconds` in metadata. It includes verification and settlement around the handler. The validity guidance is `validBefore = now + max_timeout_seconds`. Q: Can I use a client library for the payment flow? A: The endpoint documentation provides JavaScript examples using `wrapFetchWithPayment` and `ExactEvmScheme`, plus `decodePaymentResponseHeader` for the receipt. Python examples use `x402ClientSync`, `register_exact_evm_client`, `EthAccountSigner` and `x402_requests(client)`. The curl examples show the unpaid request followed by resubmission with `payment-signature`. ### Turning web pages into LLM-ready text URL: https://grist.tools/guides/web-to-markdown-for-llms (updated 2026-09-28) Turn a public web page into model context with `url-to-markdown` for article Markdown or `html-clean-text` for cleaned text and Markdown bounded by `max_chars`. Add `page-metadata`, `structured-data-extract` or `link-extract` when your agent also needs head metadata, embedded data or page links. #### Choose the content your model needs Raw HTML includes navigation, scripts and markup that consume model tokens and require removal. Start by deciding whether the task needs the article, its declared metadata, embedded structured data or links to further pages. Make that choice before assembling the model context, and request the output that serves the task. Use `url-to-markdown` when the article is the input you want. It fetches a page and returns article Markdown with headings, lists and resolved links, together with its title, byline and final URL. Its `max_chars` option cuts the returned Markdown by Unicode code points. Keep the final URL alongside the article when preparing the context. Calls use `POST https://grist.tools/v1/` with a JSON body and `content-type: application/json`. Read the free input schema at `GET /v1//schema` before building a request; `/openapi.json` also publishes the schemas. #### Use cleaned text when you need extraction controls `html-clean-text` returns both `text` and `markdown`, plus declared metadata. Choose it when you need plain text, control over article extraction, inline HTML input or links from the extracted content. For fetching, supply `url` for an HTML, XHTML or plain-text page. For source you already hold, supply `html` instead; exactly one of those inputs is required. Article extraction is enabled by default through `readability`. Disabling it converts the whole body, including navigation and footer. The same whole-body fallback applies when article extraction cannot find an article. Decide whether that broader content suits the task before passing the result to the model. Set `max_chars` to bound both text and Markdown by Unicode code points; a longer result is cut and flagged as truncated. The example below uses `html-clean-text` on a sample article about measuring rainfall. It demonstrates fetching with article extraction, a character limit and `include_links` enabled. #### Select head metadata or embedded structured data Choose `page-metadata` for fields declared in the page head: title, description, canonical URL, language, `hreflang`, favicon and Open Graph metadata. Its canonical, favicon, `hreflang` and `og:image` URLs are resolved to absolute against the final redirect destination or the page base; `og:url` and `twitter:image` are returned as authored, and a URL that cannot be resolved is `null`. Use these fields when the agent needs page identity and declared descriptions alongside, or instead of, article content. Choose `structured-data-extract` for JSON-LD, Open Graph, microdata and meta tags, plus a `normalized` merge of the requested JSON-LD, Open Graph and meta sources; microdata is returned but not merged. Set `types` to select from `jsonld`, `opengraph`, `microdata` and `meta`, or omit it to request all of them. A source you do not request returns `null` and contributes nothing to the merge. Broken JSON-LD scripts are skipped rather than causing extraction to throw. Inspect `jsonld_parse_errors` when reviewing a result, and keep the extracted sources available when using the merged fields. Ask for article content separately when the model also needs the prose. #### Choose links according to the next action Enable `include_links` on `html-clean-text` when you need links inside the extracted content. Each comes with link text, `rel` and a same-host flag. Links resolve against the fetched URL or the supplied `base_url` for inline HTML; without a base for inline input, they remain as authored. Choose `link-extract` for page links, with absolute URLs, `rel`, `nofollow` and internal-link flags. Set `internal_only` when you want links whose host equals the host of the final fetched page. The totals then describe that filtered set. Use those results to select further URLs for your agent to process. The table below compares the endpoints and their declared prices and limits. Use the generated documentation links to inspect the contract for each call you plan to make. #### Services covered | Endpoint | Price | Timeout | max_timeout_seconds | Size cap | Max duration | | --- | --- | --- | --- | --- | --- | | [url-to-markdown](https://grist.tools/docs/url-to-markdown) | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - | | [html-clean-text](https://grist.tools/docs/html-clean-text) | $0.002 | 12 s | 33 | 1 MB (1,048,576 bytes) | - | | [page-metadata](https://grist.tools/docs/page-metadata) | $0.002 | 8 s | 29 | 3 MB (3,145,728 bytes) | - | | [structured-data-extract](https://grist.tools/docs/structured-data-extract) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | | [link-extract](https://grist.tools/docs/link-extract) | $0.002 | 8 s | 29 | 5 MB (5,242,880 bytes) | - | #### Chain extraction around the question For a task that needs an article and its declared identity, start with `page-metadata`, then send the page URL to `html-clean-text` or `url-to-markdown`. Keep the returned title and URL next to the chosen text representation. Add `structured-data-extract` only when the question also needs embedded data, such as the author or publication fields shown in its normalized output. For a task that starts from an index page, call `link-extract`, select relevant returned URLs, and submit those URLs to the article endpoint. Inspect `links_truncated` before treating the returned link list as complete. If the task only needs references within the article, use `html-clean-text` with `include_links` and evaluate whether that already supplies the links you need. Build the final context from the fields the task needs. Choose `text` or `markdown` from the cleaning response, inspect the truncation flag, and attach any selected metadata or structured fields. Keep the endpoint choice explicit at each stage so article extraction, metadata extraction and link selection remain deliberate decisions. #### Budget calls and allow for the payment deadline A metadata and cleaned-text pipeline costs $0.002 plus $0.002 when both calls succeed. Adding structured data adds $0.002. An article-only path through `url-to-markdown` costs $0.002, while a discovery step through `link-extract` adds $0.002. Budget each selected page extraction separately when following links. Payment uses x402 in USDC on Base, with no account or API key. A request without `payment-signature` receives HTTP 402 with the requirements, including the price (`amount`, in atomic units of the asset) and the whole-call `maxTimeoutSeconds`, before payment; the free `/schema` route reports the same value as `max_timeout_seconds`. Use those requirements when deciding whether to proceed with the call. For `html-clean-text`, the handler timeout is 12 s, while the whole-call allowance in `max_timeout_seconds` is 33 seconds. That envelope includes payment verification and settlement around execution. Size the client deadline around the whole call. Verification happens before execution; settlement follows handler success. Gate refusals and handler failures are never settled. If a settlement call fails without answering or its answer fails validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance, and the error message identifies that case. Preserve that message when deciding what to do next. #### Handle errors before choosing a retry Use the typed code to decide what needs to change. `invalid_input` means the body does not match the schema, so correct the request. `blocked_target` means the network guard refused the destination. Check for a disallowed address, a scheme outside HTTP or HTTPS, a port outside the allowed web ports, or credentials in the URL. Repeating the unchanged request does not address those causes. `unreachable_target` means the destination did not resolve, or the connection to it failed or closed before the body arrived. `upstream_status` means it answered with a status the endpoint cannot use; `upstream_timeout` means it did not answer within the declared timeout. Check the target and allow for a later retry when the underlying availability problem may have changed. `unsupported_content_type` calls for a supported input type. `too_large` means the response exceeded the streaming size cap or the allowed decompression ratio, or the document exceeded the conversion limits of the endpoint on length, tag count or markup cost. `unprocessable` means the redirect chain could not be followed (too many hops or an unparseable `Location`), or bytes arrived but could not be processed, for example a page that yields no text or nests too deeply. Correct the source or endpoint choice where appropriate; reducing `max_chars` only changes returned text length. `rate_limited` means too many calls from the same payer or toward the same destination domain, so reduce request pressure. Retryable HTTP 503 errors, including `concurrency_limit`, `registry_at_capacity` and `upstream_unavailable`, require waiting for the seconds specified by `retry-after`. When `upstream_unavailable` comes from the front proxy, the call may already have run and settled; check `authorizationState` before paying again. `internal` identifies a service fault; inspect the error message, including any settlement uncertainty, before retrying. #### Keep the extraction boundary visible The fetched body cap for `html-clean-text` is 1 MB. Treat that input limit separately from `max_chars`, which bounds the returned text and Markdown. Check the generated limits for each endpoint you add to a chain; selecting another extraction purpose also means selecting its input contract. `html-clean-text` refuses PDF and JSON content. Inline HTML uses `base_url` to resolve relative links, but `base_url` is not allowed together with `url`, where the final fetched URL supplies the base. Public URL fetching remains subject to the network guard, and `url-to-markdown` rechecks every redirect against private and reserved addresses. Keep unsupported content and refused destinations outside this extraction path. #### Questions Q: Which endpoint should supply the text for my model context? A: Use `url-to-markdown` for article Markdown with headings, lists, resolved links, title and byline. Choose `html-clean-text` when you want both plain text and Markdown, inline HTML input or control over article extraction. Select the returned representation that the task needs and keep its source URL alongside it. Q: Can I clean HTML that my agent already has? A: Yes, send it in the `html` field to `html-clean-text`, without `url`. Supply `base_url` if relative links need a base for resolution, and enable `include_links` if you need links from the extracted content. Without a base, those links remain as authored. Q: Does `max_chars` set a model token budget? A: `max_chars` counts Unicode code points, so use it as a character limit rather than a model token count. On `html-clean-text`, it bounds both text and Markdown, with longer output cut and flagged as truncated. On `url-to-markdown`, it cuts the returned Markdown. Q: Why does cleaned output sometimes include navigation? A: `html-clean-text` converts the whole body when `readability` is disabled or article extraction cannot find an article. That body includes navigation and footer content. Review whether the broader output suits your task before placing it in the model context. Q: When should I request structured data instead of page metadata? A: Choose `structured-data-extract` when you need JSON-LD, microdata or the normalized merge of embedded sources. Use `types` to select the sources you need; unrequested sources return `null` and are excluded from the merge. Choose `page-metadata` for head fields such as canonical URL, language, `hreflang` and favicon. Q: Should an agent retry every failed extraction? A: No; inspect the typed error and correct invalid input, refused targets or unsuitable content before repeating the request. For retryable HTTP 503 errors, wait for the interval in `retry-after`. An upstream availability failure may warrant a later retry, while a settlement call that failed without answering or an answer that failed validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned") requires attention to the uncertainty stated in the error message. Q: Do I need an account before checking the price and schema? A: No account or API key is required, and `GET /v1//schema` is free. A call without `payment-signature` returns HTTP 402 with the price in `amount` and the whole-call deadline in `maxTimeoutSeconds` before you pay. Successful calls settle through x402 in USDC on Base after the handler completes. ### Parsing PDF, DOCX, XLSX and EPUB for agents URL: https://grist.tools/guides/document-parsing-api (updated 2026-09-28) An agent extracts text, Markdown or JSON by sending a public document URL to the endpoint for its format: PDF, DOCX, PPTX, XLSX, EPUB or CSV. Use image OCR for scans, screenshots and photos, then inspect extraction limits and warnings before passing the result into the next stage. #### Start with the output your agent needs Call `POST https://grist.tools/v1/` with a JSON body and `content-type: application/json`. Most document endpoints fetch the public URL in `url`; fetch the free input schema at `GET /v1//schema` before constructing the request. `/openapi.json` also publishes the schemas. No account or API key is required. Choose text for a reading task, Markdown when the supported conversion structure is useful, and JSON rows for tabular work. The table below covers the documents family, including extraction, PDF utilities and output rendering. Use its endpoint documentation links for the full input contract and its generated limits when planning each call. #### PDF text layers and image OCR are separate paths Use `pdf-to-text` to extract the PDF text layer page by page in reading order. Choose `pdf-to-markdown` for Markdown with headings inferred from font size. That inference is best effort and does not reproduce the layout exactly. Both accept `max_pages` for leading pages, and neither performs OCR. The example below demonstrates `pdf-to-markdown` and its returned Markdown. Both PDF extractors operate without fetching external CMap or standard-font data. Self-contained character mappings can yield text, including CJK, but a document needing unavailable predefined CMaps can return empty or partial text without an error. Embedded `ToUnicode` alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. For an image, choose `ocr-image-to-text`. It accepts PNG, JPEG, TIFF, WebP and GIF identified by their bytes, and returns `text` with mean `confidence`. Its accepted `lang` is `eng`. The input is an image URL, so do not send it a PDF URL expecting automatic page conversion. #### Choose the reader for DOCX, PPTX or EPUB Use `docx-to-markdown` for a DOCX document and retain its `warnings` alongside `markdown`. Conversion limitations surface as warnings rather than a promise that every element translated. The fetched body must be a ZIP containing a DOCX document; an arbitrary archive is insufficient. Use `pptx-to-text` for ordered slide body text. Its `slides` entries preserve slide numbers and text, while `text` combines the extraction. This is a body-text reader, so do not plan on speaker notes or slide rendering. The archive must contain the expected slide XML parts. Use `epub-to-text` for chapter text in reading order and Dublin Core metadata. Its archive container must point to an OPF package. `max_chapters` selects leading XHTML or HTML spine chapters; a longer book is cut and reported as `truncated`. Keep `chapters` when chapter boundaries matter instead of using only the combined `text`. #### Extract tabular data without assuming completeness Use `xlsx-to-json` for XLSX rows, with dates returned as ISO strings. Set `sheet` to an exact worksheet name, or omit it to return every worksheet. `header_row` selects the column-name row; setting it to `0` returns positional arrays and `columns` set to `null`. `max_rows` limits each worksheet separately, with truncation reported per sheet. Use `csv-to-json` for UTF-8 CSV or TSV, including quoted fields, embedded newlines and escaped quotes. The delimiter is detected from the first rows unless supplied. `header` chooses keyed objects or arrays; header names are trimmed, blank names replaced and duplicates made unique. A UTF-8 BOM is accepted, but unsupported charset declarations, NUL bytes and malformed UTF-8 are rejected. A truncated success does not validate the entire CSV structure. #### Services covered | Endpoint | Price | Timeout | max_timeout_seconds | Size cap | Max duration | | --- | --- | --- | --- | --- | --- | | [csv-to-json](https://grist.tools/docs/csv-to-json) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [docx-to-markdown](https://grist.tools/docs/docx-to-markdown) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [epub-to-text](https://grist.tools/docs/epub-to-text) | $0.005 | 45 s | 66 | 25 MB (26,214,400 bytes) | - | | [html-to-pdf](https://grist.tools/docs/html-to-pdf) | $0.005 | 12 s | 33 | 5 MB (5,242,880 bytes) | - | | [json-to-csv](https://grist.tools/docs/json-to-csv) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [markdown-to-pdf](https://grist.tools/docs/markdown-to-pdf) | $0.005 | 12 s | 33 | 2 MB (2,097,152 bytes) | - | | [ocr-image-to-text](https://grist.tools/docs/ocr-image-to-text) | $0.005 | 45 s | 66 | 10 MB (10,485,760 bytes) | - | | [pdf-merge](https://grist.tools/docs/pdf-merge) | $0.005 | 45 s | 66 | 8 MB (8,388,608 bytes) | - | | [pdf-metadata](https://grist.tools/docs/pdf-metadata) | $0.005 | 25 s | 46 | 25 MB (26,214,400 bytes) | - | | [pdf-split](https://grist.tools/docs/pdf-split) | $0.005 | 30 s | 51 | 25 MB (26,214,400 bytes) | - | | [pdf-to-markdown](https://grist.tools/docs/pdf-to-markdown) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [pdf-to-text](https://grist.tools/docs/pdf-to-text) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [pptx-to-text](https://grist.tools/docs/pptx-to-text) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | | [xlsx-to-json](https://grist.tools/docs/xlsx-to-json) | $0.005 | 40 s | 61 | 25 MB (26,214,400 bytes) | - | #### Inspect and select PDFs before extracting Use `pdf-metadata` when the agent needs page count, document-info fields, version or the encryption flag. It reads these without rendering and can describe an encrypted PDF without decrypting it. Treat the metadata as inspection data, not extracted document text. Use `pdf-split` to make a new PDF from `ranges`. Page numbering starts at the first page; comma-separated selections and inclusive ranges are emitted in the written order, including repeats. Descending ranges and pages beyond the end are errors. Use `pdf-merge` to concatenate source PDFs in the order of `urls`; sources are fetched sequentially. Both refuse encrypted sources and return `pdf_base64` with page and byte counts. Neither call replaces a text-extraction step. #### Render a deliverable or export rows Choose `html-to-pdf` for an HTML or XHTML URL, or provide self-contained `html`. Supply exactly one source. JavaScript is disabled, external images, stylesheets and fonts are blocked, and inline PNG data images are supported. Build the input around those restrictions instead of expecting a page dependent on external resources to render fully. Choose `markdown-to-pdf` for a UTF-8 Markdown URL or inline `markdown`, again supplying exactly one source. Raw HTML is escaped. PNG, JPEG, GIF and WebP data-image syntax is accepted; SVG data images and external image loading are blocked. Both renderers accept `format` and `landscape` and return `pdf_base64` with a page count. Choose `json-to-csv` for a public UTF-8 JSON document containing an object or an array of objects. `columns` fixes column order; otherwise columns are inferred from rows retained by `max_rows`. Nested values are JSON-encoded. The output uses CSV quoting and CRLF record separators, but it does not neutralize spreadsheet formulas. Headers and strings preserve formula prefixes, tabs and carriage returns; use spreadsheet-specific import controls for untrusted data. #### Chain calls through their actual input contracts For a PDF reading pipeline, inspect with `pdf-metadata` if page count or encryption affects your decision, then call `pdf-to-markdown` on the source URL. Compare `total_pages`, `extracted_pages` and `truncated`, and inspect `markdown` itself. Here `truncated` reports only the `max_pages` limit: a false value does not certify complete text extraction. When you need selected pages first, insert `pdf-split`. Its `pdf_base64` is output data, while the extractor expects a public URL. Your pipeline must decode and make the resulting PDF available at a public URL before submitting the extraction request. The same handoff applies after `pdf-merge`; do not treat base64 output as an automatically hosted document. A DOCX conversion can feed its returned `markdown` directly into the inline input of `markdown-to-pdf`. For an XLSX export, select the returned worksheet rows and make that JSON available at a public URL for `json-to-csv`. Preserve warnings, chapter or slide boundaries, and truncation indicators with the extracted content so the agent can decide whether its next task has enough input. #### Budget each successful call and its deadline A PDF inspection and Markdown extraction costs $0.005 plus $0.005 when both calls succeed. Adding page selection adds $0.005. A DOCX extraction followed by PDF rendering costs $0.005 plus $0.005; budget image OCR separately at $0.005 per successful call. Use the table for other branches of your pipeline. Payment uses x402 in USDC on Base. Without `payment-signature`, a call answers HTTP 402 with the price (`amount`, in atomic units of the asset) and `maxTimeoutSeconds` before payment; the free `/schema` route reports it as `max_timeout_seconds`. For `pdf-to-markdown`, the handler timeout is 40 s, the input cap is 25 MB, and the whole-call allowance in seconds is 61. Size the client deadline around the whole-call allowance, which includes verification and settlement. Payment is verified before the handler and settled only after success. Gate refusals and handler failures are never settled. A settlement call that fails without answering or an answer that fails validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned") leaves an uncertainty identified in the error message; inspect that message before deciding to repeat a paid request. #### Retry according to the error Correct `invalid_input` against the schema before retrying. `blocked_target` means the network guard refused the destination, such as a disallowed address, unsupported URL scheme or port, or credentials in the URL. Repeating the same destination does not address that cause. For `unsupported_content_type`, supply a supported document type; for `unprocessable`, check `details`: document endpoints also return it when the target answered with an HTTP error status (`details.status`), otherwise it covers bytes that arrived but could not be processed, such as a malformed or encrypted document. `too_large` means the response exceeded its streaming size cap or decompression ratio, or the document exceeded a processing limit of the endpoint, such as pixel count, slide or worksheet count, the CSV parsing budget, the PDF text extraction limit or the output PDF size. Change the source accordingly. `unreachable_target` means resolution failed, or the connection failed or closed before the body arrived; `upstream_timeout` means the destination did not answer within the declared timeout. Check source availability before scheduling another attempt. `rate_limited` is HTTP 429 for excessive calls from a payer or towards a destination domain, with nothing charged. Reduce that traffic. Retryable HTTP 503 errors include `concurrency_limit`, `registry_at_capacity` and `upstream_unavailable`; wait the seconds specified by `retry-after`. When `upstream_unavailable` comes from the front proxy, the call may already have run and settled; check `authorizationState` before paying again. An `internal` error is a service fault, with nothing charged in almost every case; preserve its message, especially any settlement uncertainty. #### Questions Q: Does PDF text extraction perform OCR? A: Neither `pdf-to-text` nor `pdf-to-markdown` performs OCR; both extract the existing text layer. `ocr-image-to-text` accepts supported image URLs and returns text with mean confidence, using `eng` as its accepted language. Q: Does a successful PDF response mean all text was extracted? A: No: unavailable predefined character mappings can produce empty or partial text without an error. `truncated` reports only the `max_pages` limit, so inspect the returned content even when it is false. Q: How do I select a worksheet from an XLSX file? A: Pass its exact name in `sheet` to `xlsx-to-json`; omitting that field returns every worksheet. An unknown name produces HTTP 422 listing available names. `max_rows` applies separately to each returned worksheet. Q: Can an agent pass extracted Markdown directly to PDF rendering? A: Yes, pass the extracted string as `markdown` to `markdown-to-pdf`, without also supplying `url`. Raw HTML is escaped, and external images and SVG data images are blocked. Q: Can these PDF utilities decrypt an encrypted document? A: `pdf-metadata` describes an encrypted PDF without decrypting it. `pdf-split` and `pdf-merge` refuse encrypted sources, so metadata inspection does not make those files eligible for transformation. Q: Does CSV export protect against spreadsheet formulas? A: `json-to-csv` returns raw CSV without spreadsheet formula protection. Quoting handles delimiters, quotes and newlines while formula prefixes remain intact; use spreadsheet-specific import controls when handling untrusted data. Q: Are failed document calls charged? A: Gate refusals and failures inside the handler are never settled. Payment settlement follows handler success, but if a settlement call fails without answering or its answer fails validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance. The error message identifies that uncertainty. ### Audio and video conversion API without ffmpeg servers URL: https://grist.tools/guides/audio-video-conversion-api (updated 2026-09-28) An agent can convert and inspect audio or video through grist.tools HTTP endpoints, leaving ffmpeg execution and guarded source fetching to the service. Choose an endpoint for audio output, metadata, captions, a frame or waveform peaks, and receive its result as JSON. #### Make media work an HTTP call Send a public source URL in `url` to `POST https://grist.tools/v1/`, using a JSON body and `content-type: application/json`. These media operations need guarded fetching, and every one except `subtitle-convert` also runs ffmpeg or ffprobe. Your agent selects the operation and its declared options without running its own media server. Read the free input schema at `GET /v1//schema` before constructing a request. The schemas also appear in `/openapi.json`; use the endpoint documentation links for the individual contracts. Start with the desired result: converting audio, choosing a track, inspecting a container, retrieving captions, grabbing a frame and drawing a waveform call for different endpoints. #### Convert audio or select an audio track Choose `audio-convert` to transcode the first audio stream. Its `to` targets are `wav`, `mp3`, `flac`, `aac` and `opus`, with `mp3` as the default. WAV output uses PCM, AAC output is an ADTS stream, and Opus output uses an Ogg container. Recognised source containers include MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF and MP3. Set `bitrate_kbps` for the lossy targets `mp3`, `aac` and `opus`. AAC has a stricter upper bound than the other lossy targets, and exceeding it returns `invalid_input`. Set `sample_rate` only to a value in the schema: it accepts a fixed set of rates. When omitted, ffmpeg chooses the rate, normally the source rate. Choose `audio-extract` when you need to select an audio track from a video or another recognised container. Its `stream` is a zero-based index among audio streams only, with the first selected by default. It offers the same `to` targets, also defaulting to `mp3`; a nonexistent track is `unprocessable`. The generated `audio-convert` example below converts the sample FLAC file to WAV and returns `audio_base64` with output metadata. #### Inspect media or request a visual representation Use `media-probe` for container, duration, bit rate and per-stream codec or resolution information. It reads the container index without decoding media. Its response exposes `streams` with fields such as `type`, `codec`, `width`, `height`, `sample_rate` and `channels`. Use that information to decide which operation and track the task needs. Use `video-thumbnail` for a single frame at `time_seconds`, returned in `image_base64`. Choose `png` or `jpeg` through `format`; PNG is the default. Optional `width` scales the image while preserving aspect ratio, with height rounded to an even number. Omitting it keeps native size. A timestamp past the end fails with `unprocessable`, with `details.reason` set to `no_frame`. Use `waveform-data` for normalised peaks suitable for drawing. It decodes the first audio stream to mono, then returns the largest absolute sample in each equal slice. Set `peaks` for the bucket count and `sample_rate` for the decode rate. The result includes `sample_count`, `duration_seconds`, `peak_count` and the `peaks` array, without shipping the decoded samples. #### Extract existing captions or convert a subtitle file Choose `subtitle-extract` for an existing text subtitle track inside a media container. Supported source containers include MKV/WebM, MP4/MOV, AVI and Ogg. Its zero-based `stream` counts subtitle streams only; a missing selection fails with `unprocessable`, with `details.reason` set to `no_subtitle_stream`. Set `to` to `srt` or `vtt`; extraction defaults to `srt` and returns caption text in `subtitle`. Choose `subtitle-convert` when the source URL already points to a SubRip or WebVTT file. Detection uses the content rather than the extension, and `to` defaults to `vtt`. Conversion re-serialises cues: SRT output is renumbered, while VTT cue identifiers, settings and NOTE blocks are dropped. Preserving those extras is outside this conversion contract. The table below compares the services, with their prices and declared limits. #### Services covered | Endpoint | Price | Timeout | max_timeout_seconds | Size cap | Max duration | | --- | --- | --- | --- | --- | --- | | [audio-convert](https://grist.tools/docs/audio-convert) | $0.003 | 94 s | 115 | 100 MB (104,857,600 bytes) | 60 min (3,600 s) | | [audio-extract](https://grist.tools/docs/audio-extract) | $0.003 | 94 s | 115 | 200 MB (209,715,200 bytes) | 60 min (3,600 s) | | [media-probe](https://grist.tools/docs/media-probe) | $0.003 | 30 s | 51 | 100 MB (104,857,600 bytes) | - | | [subtitle-convert](https://grist.tools/docs/subtitle-convert) | $0.003 | 15 s | 36 | 5 MB (5,242,880 bytes) | - | | [subtitle-extract](https://grist.tools/docs/subtitle-extract) | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - | | [video-thumbnail](https://grist.tools/docs/video-thumbnail) | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - | | [waveform-data](https://grist.tools/docs/waveform-data) | $0.003 | 90 s | 111 | 100 MB (104,857,600 bytes) | - | #### Chain operations around the source URL For a media inspection pipeline, start with `media-probe`, then pass the source URL to the operation the result calls for. Select audio with `audio-extract`, captions with `subtitle-extract`, or a frame with `video-thumbnail`. Match audio and subtitle selections within their respective stream types; do not copy an overall probe stream index into a type-specific selection without checking it. For first-track waveform data, send the original audio or video URL directly to `waveform-data`. For audio extraction, choose the required output format in `audio-extract` itself. Add `audio-convert` when its bit rate or sample rate controls are needed, rather than treating conversion as a mandatory step after every extraction. These inputs take URLs. If a later call must process returned `audio_base64` or `subtitle` content, your pipeline needs to make that content available at a public URL before passing it as the next `url`. Keep returned media and text distinct from source URLs. For captions, requesting the required format during extraction can remove the need for a subsequent `subtitle-convert` call. #### Check input caps before planning the chain `audio-convert` accepts source bytes up to 100 MB and audio duration up to 60 min (3,600 s). For `audio-extract`, the size cap is 200 MB and the selected audio stream must fit 60 min (3,600 s). Longer input fails with `too_large`; an input whose duration cannot be measured fails with `unprocessable`. Those duration decisions happen before encoding, and neither refusal is settled. Check each stage separately. In particular, `media-probe` has a 100 MB cap, so a source eligible for extraction may exceed the probe allowance. The Size cap column of the table lists the other caps, including those for caption files, thumbnails and waveform input. Do not transfer the audio conversion duration cap to an endpoint that does not declare it. Keep the requested operation within its contract. Probing reads metadata without decoding; a thumbnail request selects a single frame; waveform output contains peaks; subtitle extraction selects existing text captions. These results serve different purposes. Changing thumbnail width or waveform bucket count does not change the source file that the endpoint must fetch under its byte cap. #### Budget successful calls and the whole response deadline A probe followed by audio conversion costs $0.003 plus $0.003 when both calls succeed. Selecting a track instead costs $0.003 plus $0.003. A thumbnail adds $0.003, waveform data adds $0.003, and caption extraction adds $0.003. Converting a separate caption file costs $0.003. Budget the operations your pipeline actually selects. No account or API key is required. A call without `payment-signature` returns HTTP 402 with the price in `amount` and the whole-call deadline in `maxTimeoutSeconds`, which the free schema publishes as `max_timeout_seconds`. Payment uses x402 in USDC on Base. For `audio-convert`, the handler timeout is 94 s, while `max_timeout_seconds` is 115 seconds. Use the whole-call allowance for client deadlines because it includes payment verification and settlement around the handler. Payment is verified before execution and settled only after the handler succeeds. A gate refusal or handler failure is never settled. If a settlement call fails without answering or its answer fails validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance; the error message identifies that case. Preserve the message and account for that uncertainty before repeating a paid operation. #### Choose retries from the error meaning `invalid_input` means the body fails the schema, so correct fields or options before retrying. `blocked_target` means the network guard refuses the destination: check the address, HTTP or HTTPS scheme, permitted web port and absence of URL credentials. Repeating an unchanged request does not correct these causes. `unreachable_target` means resolution failed, or the connection failed or closed before the body arrived. `upstream_timeout` means it did not answer before the endpoint timeout. Check source availability and consider a later retry when that condition may have changed. `unsupported_content_type` requires a source type the endpoint handles; `unprocessable` means the source answered with an HTTP error status, or its bytes arrived but could not be processed within the endpoint deadline. `too_large` can indicate the streaming size cap, decompression ratio or a declared media duration cap. Correct the source or endpoint choice before resubmitting. For missing tracks, check the relevant stream selection; when `unprocessable` carries `details.reason` `no_frame`, choose a timestamp within the video. Preserve the error detail so the agent can distinguish these cases. `rate_limited` means excessive calls from the payer or toward the destination domain; reduce request pressure. For retryable HTTP 503 errors such as `concurrency_limit`, `registry_at_capacity` and `upstream_unavailable`, wait for the seconds in `retry-after`. When `upstream_unavailable` comes from the front proxy, the call may already have run and settled; check `authorizationState` before paying again. `internal` identifies a service fault; inspect the message, including any settlement uncertainty, before retrying. #### Questions Q: Which endpoint converts audio from a video? A: `audio-convert` converts the first audio stream in a recognised container to the selected audio format. Use `audio-extract` when you need to choose a track with `stream`. That index counts audio streams only, and extraction can return the chosen output format directly. Q: Can I control the output bit rate and sample rate? A: `audio-convert` exposes `bitrate_kbps` for `mp3`, `aac` and `opus`, with a stricter upper bound for AAC. Its `sample_rate` must come from the allowed set in the free schema. Omitting the sample rate lets ffmpeg choose it, normally from the source. Q: Does media inspection decode the audio or video? A: `media-probe` reads the container index without decoding media. It returns container, duration, bit rate and stream information, including codecs and applicable resolution fields. Choose `video-thumbnail` for a decoded frame or `waveform-data` for peaks derived from decoded audio. Q: Can I preserve every WebVTT detail during conversion? A: `subtitle-convert` re-serialises cues and drops VTT cue identifiers, settings and NOTE blocks. SRT output is renumbered. Choose this endpoint for normalised caption conversion when those extra details do not need to survive. Q: What happens when audio exceeds the duration cap? A: `audio-convert` caps audio duration at 60 min (3,600 s), and `audio-extract` caps the selected audio stream at 60 min (3,600 s). Longer input returns `too_large`, while unmeasurable duration returns `unprocessable`. Both decisions happen before encoding, and the call is never settled. Q: Can I pass an extraction response directly into another call? A: The media endpoints take a public source URL in `url`. Audio extraction returns `audio_base64`, and subtitle extraction returns text in `subtitle`. To process that output in another URL-based call, make the resulting content available at a public URL first. Q: Should my agent retry every failed media request? A: Inspect the typed error before retrying: invalid options, missing tracks and oversized inputs need a corrected request or source. For retryable HTTP 503 errors, wait for `retry-after`. If a settlement call failed without answering or its answer failed validation (that case is HTTP 500 `internal`, with a message starting "Payment settlement returned"), retain the stated uncertainty when deciding whether to repeat the operation. ## Web Requires the network. - [dns-lookup](https://grist.tools/docs/dns-lookup): DNS records for a domain (A, AAAA, MX, TXT, NS, CNAME, SOA, CAA) with their TTLs. (family domain, category Verification, free, timeout 5000ms) POST https://grist.tools/v1/dns-lookup Why this endpoint exists: A DNS query needs the network, so an agent cannot answer it on its own. It is what you run in a loop over a list of domains. Price: free · max_timeout_seconds: 6 · response cap: 2097152 bytes Typed errors: invalid_input (400), unreachable_target (424), upstream_timeout (424) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "domain": { "description": "Fully qualified domain name to resolve, up to 253 characters; it is trimmed, lowercased and stripped of trailing dots, and IP literals, single-label names and private suffixes such as .local or .internal are refused.", "type": "string", "minLength": 1, "maxLength": 253 }, "types": { "description": "Record types to query, from a, aaaa, mx, txt, ns, cname, soa and caa in either case; duplicates collapse to one query per type, answers come back in that fixed order, and the default is all eight.", "minItems": 1, "maxItems": 16, "type": "array", "items": { "type": "string", "enum": [ "a", "aaaa", "mx", "txt", "ns", "cname", "soa", "caa", "A", "AAAA", "MX", "TXT", "NS", "CNAME", "SOA", "CAA" ] } } }, "required": [ "domain" ], "additionalProperties": false } Example request body: { "domain": "example.com", "types": [ "a", "mx" ] } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "domain": "example.com", "types_requested": [ "a", "mx" ], "records": [ { "type": "a", "values": [ { "address": "93.184.215.14", "ttl": 3600 } ], "truncated": false, "resolution_error": null }, { "type": "mx", "values": [ { "exchange": "mail.example.com", "priority": 10 } ], "truncated": false, "resolution_error": null } ], "queried_at": "2026-08-26T18:00:00.000Z" } - [email-validate](https://grist.tools/docs/email-validate): Validates an email address: syntax, MX records for its domain, and whether it is a disposable or role account. (family domain, category Verification, $0.002, timeout 5000ms) POST https://grist.tools/v1/email-validate Why it is worth paying for: The MX check needs a resolver an agent may not have, and the disposable/role assessment is a curated dataset and rule set it would otherwise have to build and keep current. Price: $0.002 · max_timeout_seconds: 26 · response cap: 2097152 bytes Typed errors: invalid_input (400), unreachable_target (424), upstream_timeout (424), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "email": { "type": "string", "minLength": 1, "maxLength": 320, "description": "The address to assess, up to 320 characters; it is trimmed, and a malformed address is answered with syntax_valid false rather than rejected." }, "check_mx": { "default": true, "description": "Whether to resolve MX records for the domain; when true, a domain with no MX makes valid false, and false skips every network lookup.", "type": "boolean" }, "check_disposable": { "default": true, "description": "Whether to match the domain against the bundled disposable-provider list; when true, a disposable domain makes valid false.", "type": "boolean" } }, "required": [ "email" ], "additionalProperties": false } Example request body: { "email": "support@grist.tools", "check_mx": false } Example response: { "email": "support@grist.tools", "valid": true, "syntax_valid": true, "mx_found": false, "mx_records": [], "is_disposable": false, "is_role_account": true } - [feed-discover](https://grist.tools/docs/feed-discover): Fetches a web page and returns the RSS, Atom and JSON feeds it declares in its , as absolute URLs with their type and title. (family feeds, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/feed-discover Why it is worth paying for: It needs the network, and the feed an agent wants is declared in the page it already has as a -- one guarded fetch and a DOM query returns it, where feeding the whole page to a model costs far more in tokens and still guesses. Price: $0.002 · max_timeout_seconds: 29 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the HTML or XHTML page whose or rel=\"feed\" declarations with an RSS, Atom or JSON Feed media type are read, up to 2048 characters; up to 5 redirects are followed and no feed paths are probed." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/index.html" } Example response: { "url": "https://grist.tools/samples/web/fieldnotes/index.html", "source_url": "https://grist.tools/samples/web/fieldnotes/index.html", "final_url": "https://grist.tools/samples/web/fieldnotes/index.html", "http_status": 200, "feeds": [ { "url": "https://grist.tools/samples/web/fieldnotes/feed.xml", "type": "rss", "title": "Field Notes (RSS)", "rel": "alternate" }, { "url": "https://grist.tools/samples/web/fieldnotes/atom.xml", "type": "atom", "title": "Field Notes (Atom)", "rel": "alternate" } ], "feeds_total": 2, "feeds_truncated": false, "fetched_at": "2026-09-01T12:00:00.000Z" } - [html-clean-text](https://grist.tools/docs/html-clean-text): Fetches a web page and returns the article as clean text and markdown, with its declared metadata. (family fetch, category Data, $0.002, timeout 12000ms) POST https://grist.tools/v1/html-clean-text Why it is worth paying for: It needs the network and a real DOM: boilerplate removal, legacy charset decoding and HTML-to-markdown are three heavy libraries an agent does not carry, and feeding raw HTML to a model instead costs far more in tokens than this call costs. Price: $0.002 · max_timeout_seconds: 33 · response cap: 1048576 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "description": "URL of an HTML, XHTML or plain-text page to fetch, up to 2048 characters; redirects are followed, the body is capped at 1 MB, and exactly one of url or html is required.", "type": "string", "maxLength": 2048, "format": "uri" }, "html": { "description": "Inline HTML source to clean instead of fetching, up to 262,144 characters; binary content such as PDF or JSON is refused.", "type": "string", "minLength": 1, "maxLength": 262144 }, "base_url": { "description": "Base URL for resolving relative links in the html input; not allowed together with url, where the final fetched URL is the base.", "type": "string", "maxLength": 2048, "format": "uri" }, "include_links": { "default": false, "description": "Also return the links inside the extracted content, each with its text, rel and same-host flag, resolved against the fetched URL or base_url when there is one, else kept as authored.", "type": "boolean" }, "readability": { "default": true, "description": "Run Readability to keep only the main article; false, or an article Readability cannot find, converts the whole body, navigation and footer included.", "type": "boolean" }, "max_chars": { "default": 50000, "description": "Maximum length of text and markdown, 200 to 200,000 Unicode code points (default 50,000); a longer result is cut and flagged as truncated.", "type": "integer", "minimum": 200, "maximum": 200000 } }, "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/render-article.html", "include_links": true, "readability": true, "max_chars": 2000 } Example response: { "requested_url": "https://grist.tools/samples/web/render-article.html", "source_url": "https://grist.tools/samples/web/render-article.html", "final_url": "https://grist.tools/samples/web/render-article.html", "content_type": "text/html", "charset": "utf-8", "http_status": 200, "title": "Measuring Rainfall with a Simple Gauge", "byline": "Sample Author", "excerpt": "A short guide to reading a cylinder rain gauge at the same hour every day.", "site_name": "Grist Samples", "lang": "en", "lang_source": "html_lang", "text": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\nPlacing the gauge\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\nMount it level, so the scale reads true.\nKeep the funnel clear of leaves and insects.\nEmpty the tube right after each reading.\nReading the scale\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\nWrite the reading in a notebook with the date and time. A printable schedule helps when several people share the job.", "markdown": "A cylinder rain gauge is a clear tube with a scale printed on the side. Rain falls through a funnel at the top and collects in the tube, where the water level shows how much fell since the last reading.\n\nThe most useful habit is consistency. Read the gauge at the same hour every day, ideally in the morning, so that each number covers one full day.\n\n## Placing the gauge\n\nPut the gauge in an open spot, away from walls, trees and roofs. A good rule is to keep it at least twice as far from any obstacle as the obstacle is tall.\n\n- Mount it level, so the scale reads true.\n- Keep the funnel clear of leaves and insects.\n- Empty the tube right after each reading.\n\n## Reading the scale\n\nBend down until your eye is level with the water. The surface curves slightly where it meets the tube; read the number at the bottom of that curve.\n\nWrite the reading in a notebook with the date and time. A [printable schedule](https://grist.tools/samples/web/render-print.html) helps when several people share the job.", "word_count": 178, "char_count": 933, "truncated": false, "readability_applied": true, "links": [ { "href": "https://grist.tools/samples/web/render-print.html", "text": "printable schedule", "rel": null, "is_internal": true } ], "link_count": 1, "fetched_at": "2026-09-01T12:00:00.000Z" } - [http-headers](https://grist.tools/docs/http-headers): Returns the response headers, final status, redirect chain and a security-header grade for a URL, without downloading the body. (family fetch, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/http-headers Why it is worth paying for: It needs a hardened HTTP client -- an agent that fetches the URL itself can walk a redirect into a private address -- and the security grade is a deterministic, documented rubric it would otherwise have to build and maintain. Price: $0.002 · max_timeout_seconds: 29 · response cap: 65536 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL to request, up to 2048 characters; up to five redirects are followed and every hop goes through the SSRF guard." }, "method": { "default": "HEAD", "description": "HTTP method for the request, HEAD by default; GET is for servers that mishandle HEAD, and its body is still never read.", "type": "string", "enum": [ "HEAD", "GET" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://example.com/", "method": "HEAD" } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "url": "https://example.com/", "source_url": "https://example.com/", "final_url": "https://example.com/", "redirect_count": 0, "http_status": 200, "headers": { "content-security-policy": "default-src 'self'", "content-type": "text/html; charset=utf-8", "cross-origin-opener-policy": "same-origin", "cross-origin-resource-policy": "same-origin", "permissions-policy": "geolocation=()", "referrer-policy": "strict-origin-when-cross-origin", "strict-transport-security": "max-age=63072000; includeSubDomains; preload", "x-content-type-options": "nosniff", "x-frame-options": "DENY" }, "security": { "hsts": { "present": true, "max_age": 63072000, "include_subdomains": true, "preload": true }, "csp_present": true, "x_content_type_options": "nosniff", "x_frame_options": "DENY", "referrer_policy": "strict-origin-when-cross-origin", "permissions_policy": "geolocation=()", "cross_origin_opener_policy": "same-origin", "cross_origin_resource_policy": "same-origin" }, "security_grade": "A", "missing_security_headers": [], "fetched_at": "2026-08-26T18:00:00.000Z" } - [ip-info](https://grist.tools/docs/ip-info): Reverse DNS (PTR) and network ownership for an IP address or a domain. (family domain, category Verification, $0.002, timeout 5000ms) POST https://grist.tools/v1/ip-info Why it is worth paying for: A reverse lookup needs the network, and the ownership data is a versioned dataset an agent does not carry. It is what you run over a list of addresses from a log. Price: $0.002 · max_timeout_seconds: 26 · response cap: 2097152 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "ip": { "description": "A public IPv4 or IPv6 address to look up; private, loopback and reserved addresses are refused. Give exactly one of ip or domain.", "type": "string", "minLength": 1 }, "domain": { "description": "A public fully qualified domain name, resolved to its first address, which is then looked up; the call is refused if the name resolves to any private address. Give exactly one of ip or domain.", "type": "string", "minLength": 1 } }, "additionalProperties": false } Example request body: { "ip": "8.8.8.8" } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "query": "8.8.8.8", "ip": "8.8.8.8", "ptr": "dns.google", "asn": 15169, "asn_org": "GOOGLE", "country": "US" } - [link-extract](https://grist.tools/docs/link-extract): Fetches a web page and returns every link on it, resolved to absolute, each tagged with its rel, a nofollow flag and whether it stays on the same site. (family fetch, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/link-extract Why it is worth paying for: It needs the network and a real DOM: resolving relative and protocol-relative hrefs against a possible , classifying internal versus external, and reading rel/nofollow correctly is fiddly to get right, and feeding raw HTML to a model to do it costs far more in tokens than this call costs. Price: $0.002 · max_timeout_seconds: 29 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the HTML or XHTML page to read, up to 2048 characters; up to 5 redirects are followed and links resolve against the final URL or its ." }, "internal_only": { "default": false, "description": "When true, return only links whose host equals the host of the fetched page (the final redirect hop); the totals then count only those links.", "type": "boolean" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/index.html", "internal_only": false } Example response: { "url": "https://grist.tools/samples/web/fieldnotes/index.html", "source_url": "https://grist.tools/samples/web/fieldnotes/index.html", "final_url": "https://grist.tools/samples/web/fieldnotes/index.html", "http_status": 200, "base_url": "https://grist.tools/samples/web/fieldnotes/index.html", "links": [ { "href": "https://grist.tools/samples/web/fieldnotes/index.html", "text": "Home", "rel": null, "nofollow": false, "internal": true }, { "href": "https://grist.tools/samples/web/fieldnotes/about.html", "text": "About", "rel": null, "nofollow": false, "internal": true }, { "href": "https://grist.tools/samples/web/fieldnotes/feed.xml", "text": "Subscribe", "rel": "alternate", "nofollow": false, "internal": true }, { "href": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "text": "The first frost of the season", "rel": null, "nofollow": false, "internal": true }, { "href": "https://grist.tools/samples/web/fieldnotes/notes/rain-gauge.html", "text": "Reading a rain gauge by hand", "rel": null, "nofollow": false, "internal": true }, { "href": "https://example.org/glossary", "text": "an outside glossary", "rel": "nofollow noopener", "nofollow": true, "internal": false }, { "href": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "text": "Sitemap", "rel": null, "nofollow": false, "internal": true } ], "links_total": 7, "links_returned": 7, "links_truncated": false, "fetched_at": "2026-09-01T12:00:00.000Z" } - [page-metadata](https://grist.tools/docs/page-metadata): Fetches a page and returns its metadata -- title, description, canonical, language, hreflang, favicon, Open Graph and Twitter cards -- with every URL resolved to absolute. (family fetch, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/page-metadata Why it is worth paying for: It needs the network and three libraries an agent scraping many pages does not carry: an HTML parser, a legacy charset decoder, and a resolver that honours and the final redirect hop. Feeding raw HTML to a model to pull these fields costs far more in tokens than this call. Price: $0.002 · max_timeout_seconds: 29 · response cap: 3145728 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of an HTML or XHTML page, up to 2048 characters; up to 5 redirects are followed and relative head URLs resolve against the final hop or its ." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/render-article.html" } Example response: { "url": "https://grist.tools/samples/web/render-article.html", "source_url": "https://grist.tools/samples/web/render-article.html", "final_url": "https://grist.tools/samples/web/render-article.html", "http_status": 200, "title": "Measuring Rainfall with a Simple Gauge", "description": "A short guide to reading a cylinder rain gauge at the same hour every day.", "canonical": "https://grist.tools/samples/web/render-article.html", "lang": "en", "hreflang": [ { "lang": "en", "href": "https://grist.tools/samples/web/render-article.html" }, { "lang": "x-default", "href": "https://grist.tools/samples/web/render-article.html" } ], "hreflang_total": 2, "hreflang_truncated": false, "favicon": "https://grist.tools/favicon.png", "charset": "utf-8", "opengraph": { "title": "Measuring Rainfall with a Simple Gauge", "description": "Read the gauge at eye level, at the same hour, and write the number down.", "image": "https://grist.tools/og.png", "type": "article", "site_name": "Grist Samples", "url": "https://grist.tools/samples/web/render-article.html" }, "twitter": { "card": "summary_large_image", "title": "Measuring Rainfall with a Simple Gauge", "description": "Read the gauge at eye level, at the same hour, and write the number down.", "image": "https://grist.tools/og.png" }, "robots_meta": "index, follow", "fetched_at": "2026-09-01T12:00:00.000Z" } - [robots-check](https://grist.tools/docs/robots-check): Fetches robots.txt and answers whether a user agent may crawl a URL, showing the group, the winning rule and every rule that competed with it. (family feeds, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/robots-check Why it is worth paying for: It needs the network, and the part an agent gets wrong on its own is precedence: RFC 9309 resolves conflicts by longest match, not first match. One fetch answers up to 50 paths. Price: $0.002 · max_timeout_seconds: 29 · response cap: 512000 bytes Typed errors: invalid_input (400), blocked_target (403), upstream_timeout (424), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL to check, up to 2048 characters; robots.txt is fetched from the root of its origin and the verdict is for its path and query." }, "user_agent": { "default": "*", "description": "Crawler product token or full User-Agent string, printable ASCII up to 200 characters; the text before the first slash or space picks the group, and the default is *.", "type": "string", "minLength": 1, "maxLength": 200, "pattern": "^[\\x20-\\x7e]+$" }, "include_sitemaps": { "default": true, "description": "Whether to return the Sitemap URLs listed in robots.txt; false returns null for sitemaps and its counters.", "type": "boolean" }, "additional_paths": { "description": "Up to 50 more root-relative paths or absolute URLs on the same origin as url, each answered from the same robots.txt fetch in additional_results.", "maxItems": 50, "type": "array", "items": { "type": "string", "minLength": 1, "maxLength": 2048 } } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/v1/whois", "user_agent": "GPTBot", "include_sitemaps": true, "additional_paths": [ "/docs/whois", "/v1/whois/schema" ] } Example response: { "url": "https://grist.tools/v1/whois", "source_url": "https://grist.tools/robots.txt", "robots_url": "https://grist.tools/robots.txt", "robots_status": "found", "robots_http_status": 200, "user_agent": "GPTBot", "user_agent_token": "GPTBot", "allowed": false, "decision_reason": "rule", "matched_rule": { "type": "disallow", "path": "/v1/", "line": 3, "octet_length": 4 }, "matched_group": { "user_agents": [ "*" ], "is_wildcard": true, "line": 1 }, "competing_rules": [ { "type": "disallow", "path": "/v1/", "line": 3, "octet_length": 4 }, { "type": "allow", "path": "/", "line": 2, "octet_length": 1 } ], "competing_rules_total": 2, "competing_rules_truncated": false, "crawl_delay": null, "sitemaps": [ "https://grist.tools/sitemap.xml" ], "sitemaps_total": 1, "sitemaps_truncated": false, "groups_found": 1, "rules_parsed": 3, "unparsed_lines": 0, "robots_size_bytes": 141, "robots_truncated": false, "additional_results": [ { "path": "/docs/whois", "allowed": true, "decision_reason": "rule", "matched_rule": { "type": "allow", "path": "/", "line": 2, "octet_length": 1 } }, { "path": "/v1/whois/schema", "allowed": true, "decision_reason": "rule", "matched_rule": { "type": "allow", "path": "/v1/*/schema", "line": 4, "octet_length": 12 } } ], "fetched_at": "2026-09-01T12:00:00.000Z", "checked_at": "2026-09-01T12:00:00.000Z" } - [rss-parse](https://grist.tools/docs/rss-parse): Fetches an RSS, Atom or RSS 1.0 feed and returns its channel metadata and a normalised list of items -- id, title, link, summary, content, author, dates and categories -- with every date in ISO 8601 UTC. (family feeds, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/rss-parse Why it is worth paying for: It needs the network and normalises three feed formats that say the same things in different tags and two different date formats, behind an XXE-safe XML parser. One call replaces the three-parser, RFC-822-date, entity-decoding scraper an agent would otherwise carry. Price: $0.002 · max_timeout_seconds: 29 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the RSS 2.0, Atom or RSS 1.0 (RDF) feed, up to 2048 characters; up to 5 redirects are followed and a DTD is refused." }, "limit": { "default": 50, "description": "Maximum number of items returned, 1 to 1000 (default 50), counted after the since filter; a longer feed is cut and reported as truncated.", "type": "integer", "minimum": 1, "maximum": 1000 }, "since": { "description": "ISO 8601 timestamp with Z or an offset; when set, only items whose published (or else updated) date is at or after it are kept, and undated items are dropped.", "type": "string", "format": "date-time", "pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d(?::[0-5]\\d(?:\\.\\d+)?)?(?:Z|([+-](?:[01]\\d|2[0-3]):[0-5]\\d)))$" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/feed.xml", "limit": 10 } Example response: { "url": "https://grist.tools/samples/web/fieldnotes/feed.xml", "source_url": "https://grist.tools/samples/web/fieldnotes/feed.xml", "final_url": "https://grist.tools/samples/web/fieldnotes/feed.xml", "http_status": 200, "format": "rss", "feed": { "title": "Field Notes", "link": "https://grist.tools/samples/web/fieldnotes/index.html", "description": "Short notes on frost, rain and wind from a small test garden.", "language": "en", "updated": "2026-08-25T06:15:00Z", "generator": "Hand-written sample" }, "items": [ { "id": "fieldnotes-rain-gauge", "title": "Reading a rain gauge by hand", "link": "https://grist.tools/samples/web/fieldnotes/notes/rain-gauge.html", "summary": "How the daily rain reading is taken & written down.", "content": "

The gauge is read once a day at eight in the morning.

", "author": "Field Notes Editors", "published": "2026-08-25T06:15:00Z", "updated": null, "categories": [ "rain", "method" ] }, { "id": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "title": "The first frost of the season", "link": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "summary": "A light ground frost on the lower beds, measured at dawn.", "content": null, "author": "Field Notes Editors", "published": "2026-08-24T06:30:00Z", "updated": null, "categories": [ "frost" ] } ], "count": 2, "total": 2, "truncated": false, "fetched_at": "2026-09-01T12:00:00.000Z" } - [sitemap-parse](https://grist.tools/docs/sitemap-parse): Parses a sitemap, a sitemap index or a bare origin into a flat, bounded list of URLs with lastmod, changefreq and priority. (family feeds, category Data, $0.002, timeout 14000ms) POST https://grist.tools/v1/sitemap-parse Why it is worth paying for: Needs the network, follows a sitemap index across many documents and normalises four different legal sitemap formats. An agent that does it itself either writes an XML parser with the XXE and fan-out defences, or fetches an unbounded tree of documents a stranger chose. Price: $0.002 · max_timeout_seconds: 35 · response cap: 10485760 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "pattern": "^[hH][tT][tT][pP][sS]?:", "description": "Absolute http(s) URL, up to 2048 characters, of a sitemap, a sitemap index or a bare origin (path / and no query); a bare origin is resolved through robots.txt Sitemap: lines or /sitemap.xml." }, "limit": { "default": 5000, "description": "Maximum number of URLs returned across every document read, 1 to 50000 (default 5000); past it the list is cut and reported as truncated.", "type": "integer", "minimum": 1, "maximum": 50000 }, "include_metadata": { "default": false, "description": "When true each URL carries its lastmod, changefreq and priority; when false (default) those three fields are null.", "type": "boolean" }, "max_sitemaps": { "default": 10, "description": "Maximum number of sitemap documents fetched, entry document included, 1 to 50 (default 10); children past it are reported as skipped with reason sitemap_limit.", "type": "integer", "minimum": 1, "maximum": 50 }, "discover_from_robots": { "default": true, "description": "Only for a bare origin: when true (default) the Sitemap: lines of /robots.txt are read first with /sitemap.xml as the fallback, when false /sitemap.xml is used directly.", "type": "boolean" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "limit": 100, "include_metadata": true } Example response: { "requested_url": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "entry_url": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "source_url": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "discovery": "direct", "kind": "mixed", "format": "xml", "urls": [ { "loc": "https://grist.tools/samples/web/fieldnotes/index.html", "lastmod": "2026-08-25T00:00:00Z", "changefreq": "daily", "priority": 1 }, { "loc": "https://grist.tools/samples/web/fieldnotes/about.html", "lastmod": "2026-08-20T00:00:00Z", "changefreq": "yearly", "priority": 0.3 }, { "loc": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "lastmod": "2026-08-24T06:30:00Z", "changefreq": "monthly", "priority": 0.6 }, { "loc": "https://grist.tools/samples/web/fieldnotes/notes/rain-gauge.html", "lastmod": "2026-08-25T06:15:00Z", "changefreq": null, "priority": 0.6 } ], "count": 4, "total_seen": 4, "truncated": false, "stopped_reason": null, "sitemap_count": 3, "sitemaps_visited": [ { "url": "https://grist.tools/samples/web/fieldnotes/sitemap-index.xml", "kind": "sitemapindex", "url_count": 0, "http_status": 200, "bytes": 407, "gzip": false, "depth": 0 }, { "url": "https://grist.tools/samples/web/fieldnotes/sitemap-pages.xml", "kind": "urlset", "url_count": 2, "http_status": 200, "bytes": 479, "gzip": false, "depth": 1 }, { "url": "https://grist.tools/samples/web/fieldnotes/sitemap-notes.xml", "kind": "urlset", "url_count": 2, "http_status": 200, "bytes": 493, "gzip": false, "depth": 1 } ], "sitemaps_skipped": [], "sitemaps_skipped_total": 0, "sitemaps_skipped_by_reason": { "depth_limit": 0, "sitemap_limit": 0, "cross_host": 0, "too_large": 0, "malformed": 0, "duplicate": 0, "time_budget": 0, "not_fetchable": 0 }, "invalid_loc_count": 0, "invalid_lastmod_count": 0, "robots_sitemaps": null, "fetched_at": "2026-09-01T12:00:00.000Z" } - [ssl-cert-info](https://grist.tools/docs/ssl-cert-info): Fetches the TLS certificate a host presents: issuer, subject, SAN list, validity window, days to expiry and the full chain. (family domain, category Verification, $0.002, timeout 10000ms) POST https://grist.tools/v1/ssl-cert-info Why it is worth paying for: Reading a live certificate needs a TLS handshake against the host, which an agent may be unable to make; the days-to-expiry and chain view are what you monitor across a fleet of domains. Price: $0.002 · max_timeout_seconds: 31 · response cap: 2097152 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "domain": { "description": "Public fully qualified domain name to connect to and send as SNI. Give exactly one of domain or url.", "type": "string", "minLength": 1 }, "url": { "description": "An http(s) URL whose host is probed; an explicit port in the URL overrides the port field, and the path is ignored. Give exactly one of domain or url.", "type": "string", "minLength": 1 }, "port": { "default": 443, "description": "TLS port for the handshake, one of 443, 465, 636, 990, 993, 995, 8443; defaults to 443.", "type": "integer", "minimum": -9007199254740991, "maximum": 9007199254740991 } }, "additionalProperties": false } Example request body: { "domain": "example.com" } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "host": "example.com", "port": 443, "issuer": "CN=R3, O=Let's Encrypt, C=US", "subject": "CN=example.com, O=Example Inc, C=US", "san": [ "example.com", "www.example.com" ], "valid_from": "2026-08-01T00:00:00.000Z", "valid_to": "2026-10-30T23:59:59.000Z", "days_to_expiry": 65, "self_signed": false, "chain": [ { "issuer": "CN=R3, O=Let's Encrypt, C=US", "subject": "CN=example.com, O=Example Inc, C=US", "valid_from": "2026-08-01T00:00:00.000Z", "valid_to": "2026-10-30T23:59:59.000Z" }, { "issuer": "CN=ISRG Root X1, O=Internet Security Research Group, C=US", "subject": "CN=R3, O=Let's Encrypt, C=US", "valid_from": "2020-09-04T00:00:00.000Z", "valid_to": "2025-09-15T16:00:00.000Z" } ], "queried_at": "2026-08-26T18:00:00.000Z" } - [structured-data-extract](https://grist.tools/docs/structured-data-extract): Fetches a web page and returns its embedded machine-readable data -- JSON-LD, OpenGraph, microdata and meta tags -- plus a normalised merge across all four. (family fetch, category Data, $0.002, timeout 8000ms) POST https://grist.tools/v1/structured-data-extract Why it is worth paying for: It needs the network and a real DOM, and the extraction is the code an agent keeps re-implementing wrong: skipping a broken JSON-LD script instead of throwing, telling og: from name= apart, and walking microdata with its parent/child nesting rule. One call replaces four brittle scrapers. Price: $0.002 · max_timeout_seconds: 29 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the HTML or XHTML page to read, up to 2048 characters; up to 5 redirects are followed and relative URLs resolve against the final URL or its ." }, "types": { "description": "Which of jsonld, opengraph, microdata and meta to extract, at least one; omitted means all four, and a type not requested comes back null and is left out of the normalized merge.", "minItems": 1, "type": "array", "items": { "type": "string", "enum": [ "jsonld", "opengraph", "microdata", "meta" ] } } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html" } Example response: { "url": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "source_url": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "final_url": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "http_status": 200, "jsonld": [ { "@context": "https://schema.org", "@type": "Article", "headline": "The first frost of the season", "description": "A light ground frost on the lower beds, measured at dawn.", "image": "../icon-32.png", "datePublished": "2026-08-24T06:30:00Z", "author": { "@type": "Organization", "name": "Field Notes Editors" } } ], "jsonld_total": 1, "jsonld_truncated": false, "jsonld_parse_errors": 0, "opengraph": { "og:title": "The first frost of the season", "og:type": "article", "og:url": "https://grist.tools/samples/web/fieldnotes/notes/first-frost.html", "og:image": "https://grist.tools/samples/web/fieldnotes/icon-32.png" }, "microdata": [ { "type": "https://schema.org/Place", "properties": { "name": "Lower beds", "description": "The lowest corner of the garden, where cold air settles first." } } ], "microdata_total": 1, "microdata_truncated": false, "meta": { "description": "A light ground frost on the lower beds, measured at dawn.", "author": "Field Notes Editors", "article:published_time": "2026-08-24T06:30:00Z" }, "normalized": { "title": "The first frost of the season", "description": "A light ground frost on the lower beds, measured at dawn.", "image": "https://grist.tools/samples/web/fieldnotes/icon-32.png", "type": "Article", "author": "Field Notes Editors", "published": "2026-08-24T06:30:00Z" }, "fetched_at": "2026-09-01T12:00:00.000Z" } - [url-to-markdown](https://grist.tools/docs/url-to-markdown): Fetches a web page and returns the article as clean markdown, with its title, byline and final URL. (family fetch, category Data, $0.002, timeout 12000ms) POST https://grist.tools/v1/url-to-markdown Why it is worth paying for: Raw HTML is mostly navigation, scripts and markup an agent pays for in tokens and then has to strip itself; this returns only the article, as markdown with headings, lists and resolved links, fetched through an SSRF-safe client that re-checks every redirect. Price: $0.002 · max_timeout_seconds: 33 · response cap: 1048576 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_status (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Page to fetch, http or https. Redirects are followed and every hop is checked against private and reserved addresses." }, "max_chars": { "default": 50000, "description": "Cut applied to the returned markdown, counted in Unicode code points.", "type": "integer", "minimum": 200, "maximum": 200000 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://kitchen-ledger.blog/posts/weekday-starter", "max_chars": 2000 } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "title": "Feeding a Sourdough Starter on a Weekday Schedule", "byline": "Marta Rinaldi", "excerpt": "A feeding routine that keeps a rye starter active on a nine-to-five schedule, with ratios and timings.", "site_name": "Kitchen Ledger", "lang": "en", "markdown": "By [Marta Rinaldi](https://kitchen-ledger.blog/authors/marta-rinaldi/)\n\nMost starter guides assume you are home all day, ready to feed every twelve hours on the dot. A starter kept on a weekday routine can stay just as lively, provided the flour, the ratio and the temperature do the work that a flexible timetable cannot.\n\n## The ratio that buys you time\n\nA stiffer, less frequently fed starter ferments more slowly. Feeding one part starter to five parts flour and five parts water stretches the peak from four hours to roughly ten at a kitchen temperature of twenty-two degrees, which is long enough to cover a working day.\n\n- Morning, before work: discard all but 10 g and feed 50 g flour and 50 g water.\n- Evening: use the starter at its peak, or refrigerate it overnight.\n- Weekend: return to two feeds a day at one to one to one if you plan to bake.\n\n## Choosing the flour\n\nWhole rye ferments faster and more predictably than white wheat, so a small share of it keeps a slow schedule from stalling. The [guide to flour protein](https://kitchen-ledger.blog/guides/flour-protein/) explains why, and the [baker's percentage tables](https://www.bakerspercentage.net/tables) help when you scale a recipe up.\n\nIf the starter smells sharply of acetone by evening, it is hungry: shorten the interval or lower the temperature rather than adding more flour at once.", "length": 1368, "truncated": false, "readability_applied": true, "url_final": "https://kitchen-ledger.blog/posts/weekday-starter", "fetched_at": "2026-09-26T09:30:00.000Z" } - [url-unshorten](https://grist.tools/docs/url-unshorten): Follows a shortened or tracking URL to its real destination and returns every hop it went through. (family fetch, category Data, $0.002, timeout 10000ms) POST https://grist.tools/v1/url-unshorten Why it is worth paying for: Resolving a short link needs the network, and doing it safely needs a hardened client: an agent that follows the redirects itself will happily walk into a private address or a redirect loop. Price: $0.002 · max_timeout_seconds: 31 · response cap: 65536 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "The short or tracking URL to follow, up to 2048 characters; every hop, the first included, goes through the SSRF guard." }, "max_hops": { "default": 10, "description": "Most redirects to follow, 1 to 10, default 10; a chain that has not ended by then fails as unprocessable, and a meta refresh at the limit stops with stopped_reason hop_limit.", "type": "integer", "minimum": 1, "maximum": 10 }, "method": { "default": "auto", "description": "auto sends HEAD and retries the same hop with GET on a 400, 403, 405 or 501; head and get force one method for every hop.", "type": "string", "enum": [ "auto", "head", "get" ] }, "follow_meta_refresh": { "default": false, "description": "Whether to also follow an HTML meta refresh with a delay of 30 seconds or less, found in the first 64 KiB of a 2xx HTML or untyped page; method head turns it off, and JavaScript redirects are never followed.", "type": "boolean" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/unshorten-start.html", "max_hops": 10, "follow_meta_refresh": true } Example response: { "requested_url": "https://grist.tools/samples/web/unshorten-start.html", "source_url": "https://grist.tools/samples/web/unshorten-start.html", "final_url": "https://grist.tools/samples/web/unshorten-landing.html", "final_status": 200, "final_host": "grist.tools", "final_content_type": "text/html; charset=utf-8", "hop_count": 1, "reached_final": true, "stopped_reason": null, "cross_host": false, "hops": [ { "url": "https://grist.tools/samples/web/unshorten-start.html", "method": "GET", "status": 200, "location": "unshorten-landing.html", "resolved_location": "https://grist.tools/samples/web/unshorten-landing.html", "kind": "meta_refresh", "host": "grist.tools" }, { "url": "https://grist.tools/samples/web/unshorten-landing.html", "method": "GET", "status": 200, "location": null, "resolved_location": null, "kind": "final", "host": "grist.tools" } ], "initial_host_is_known_shortener": false, "shortener_list_version": "2026-08-26", "checked_at": "2026-09-01T12:00:00.000Z", "fetched_at": "2026-09-01T12:00:00.000Z" } - [wellknown-fetch](https://grist.tools/docs/wellknown-fetch): Fetches and parses a domain's well-known files -- security.txt (RFC 9116), ads.txt and humans.txt -- returning typed fields, seller records and a cross-file summary in one call. (family feeds, category Data, $0.002, timeout 12000ms) POST https://grist.tools/v1/wellknown-fetch Why it is worth paying for: It needs the network and a hardened client: the files are named by a stranger, so a naive fetch of /ads.txt is an SSRF probe, and the value is the parse -- PGP-signed security.txt fields, ads.txt split into DIRECT/RESELLER records -- an agent would otherwise build and maintain itself. Price: $0.002 · max_timeout_seconds: 33 · response cap: 2097152 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), too_large (413), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "domain": { "description": "Fully qualified domain name whose files are fetched over https, up to 253 characters; it is trimmed, lowercased and stripped of trailing dots, and IP literals, single-label names and private suffixes are refused.", "type": "string", "minLength": 1, "maxLength": 253 }, "files": { "description": "Which files to fetch, from security.txt (read at /.well-known/security.txt), ads.txt and humans.txt (read at the root); duplicates collapse, and the default is all three.", "minItems": 1, "maxItems": 3, "type": "array", "items": { "type": "string", "enum": [ "security.txt", "ads.txt", "humans.txt" ] } } }, "required": [ "domain" ], "additionalProperties": false } Example request body: { "domain": "example.com" } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "domain": "example.com", "files_requested": [ "security.txt", "ads.txt", "humans.txt" ], "security_txt": { "requested": true, "present": true, "status": 200, "url": "https://example.com/.well-known/security.txt", "contact": [ "mailto:security@example.com", "https://example.com/security-contact" ], "expires": "2027-01-01T00:00:00.000Z", "is_expired": false, "encryption": [ "https://example.com/pgp-key.txt" ], "acknowledgments": [ "https://example.com/hall-of-fame" ], "preferred_languages": "en, fr", "canonical": [ "https://example.com/.well-known/security.txt" ], "policy": [ "https://example.com/security-policy" ], "hiring": [ "https://example.com/jobs" ], "signed": false, "field_count": 9 }, "ads_txt": { "requested": true, "present": true, "status": 200, "url": "https://example.com/ads.txt", "records": [ { "advertising_system": "greenadexchange.com", "publisher_id": "12345", "relationship": "DIRECT", "certification_authority_id": "d75815a79", "line": 4 }, { "advertising_system": "blueadexchange.com", "publisher_id": "xf7", "relationship": "RESELLER", "certification_authority_id": null, "line": 5 }, { "advertising_system": "silverssp.com", "publisher_id": "9675", "relationship": "DIRECT", "certification_authority_id": null, "line": 6 } ], "records_total": 3, "records_truncated": false, "direct_count": 2, "reseller_count": 1, "variables": [ { "name": "CONTACT", "value": "adops@example.com" } ], "variables_truncated": false, "invalid_lines": 0 }, "humans_txt": { "requested": true, "present": true, "status": 200, "url": "https://example.com/humans.txt", "text": "/* TEAM */\nWebmaster: Jane Doe\nSite: example.com\n\n/* THANKS */\nEveryone.\n", "truncated": false, "bytes": 73 }, "normalised": { "files_present": [ "ads.txt", "humans.txt", "security.txt" ], "security_contacts": [ "mailto:security@example.com", "https://example.com/security-contact" ], "security_expires": "2027-01-01T00:00:00.000Z", "security_expired": false, "security_signed": false, "ads_sellers": 3, "ads_direct": 2, "ads_reseller": 1 }, "fetched_at": "2026-08-26T18:00:00.000Z" } - [whois](https://grist.tools/docs/whois): Registration record for a domain over whois: registrar, creation/update/expiry dates, status codes and nameservers, with the raw response. (family domain, category Verification, $0.002, timeout 12000ms) POST https://grist.tools/v1/whois Why it is worth paying for: whois is a port-43 protocol an agent is unlikely to speak, behind a per-TLD server map and an IANA referral step; the parsed fields plus a confidence score are what make the raw record usable. Price: $0.002 · max_timeout_seconds: 33 · response cap: 2097152 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "domain": { "description": "Public fully qualified domain name to look up, up to 253 characters; it is lowercased and stripped of trailing dots, and its TLD picks the registry whois server (built-in map, else the IANA referral).", "type": "string", "minLength": 1, "maxLength": 253 } }, "required": [ "domain" ], "additionalProperties": false } Example request body: { "domain": "example.com" } Example response. This example is illustrative: it shows the exact shape the handler returns, but it was not produced by a reproducible call, because the real answer depends on live network data that changes over time. { "domain": "example.com", "registrar": "Example Registrar, LLC", "created": "1995-08-14T04:00:00.000Z", "updated": "2025-08-14T07:01:44.000Z", "expires": "2026-08-13T04:00:00.000Z", "status": [ "clientTransferProhibited" ], "nameservers": [ "a.iana-servers.net", "b.iana-servers.net" ], "raw": "Domain Name: EXAMPLE.COM\nRegistrar: Example Registrar, LLC\nRegistrar WHOIS Server: whois.exampleregistrar.com\nUpdated Date: 2025-08-14T07:01:44Z\nCreation Date: 1995-08-14T04:00:00Z\nRegistry Expiry Date: 2026-08-13T04:00:00Z\nDomain Status: clientTransferProhibited https://icann.org/epp#clientTransferProhibited\nName Server: A.IANA-SERVERS.NET\nName Server: B.IANA-SERVERS.NET", "parsed_confidence": 1 } ## Documents Requires a heavy dependency. - [csv-to-json](https://grist.tools/docs/csv-to-json): Parses a CSV/TSV document into JSON rows, keyed by a header row or as arrays; the delimiter is auto-detected. (family documents, category Content, $0.005, timeout 30000ms) POST https://grist.tools/v1/csv-to-json Why it is worth paying for: A correct CSV parser is more than a split on commas (quoted fields, embedded newlines and escaped quotes all bite), and doing it over a URL an agent fetches itself walks past every SSRF control. Price: $0.005 · max_timeout_seconds: 51 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "UTF-8 CSV/TSV URL. Missing charset means UTF-8; only utf-8 and utf8 declarations are accepted (case-insensitive, optionally quoted). Unsupported or invalid declarations and NUL bytes return 415; malformed UTF-8 returns 422. A UTF-8 BOM is accepted. Encoding is checked across the whole body." }, "delimiter": { "description": "A single field delimiter character such as \",\" or a tab; omitted, it is auto-detected from the first rows among comma, tab, pipe, semicolon and the ASCII record and unit separators, and a line break, quote or BOM given here falls back to a comma.", "type": "string", "minLength": 1, "maxLength": 1 }, "header": { "default": true, "description": "Use the first nonempty record as object keys. Names are trimmed; blank names become column_. Names are made unique left to right using the next unused _2, _3 suffix, including collisions with literal suffixed names. Special names remain data keys. Short rows pad with empty strings; extra fields return 422 csv_extra_fields. With false, preserve actual row widths as string arrays.", "type": "boolean" }, "max_rows": { "default": 10000, "description": "Maximum returned nonempty data rows. One additional nonempty record sets truncated; its width is budgeted but not checked against the header. A truncated success does not validate the entire CSV structure.", "type": "integer", "minimum": 1, "maximum": 100000 } }, "required": [ "url" ], "additionalProperties": false, "description": "All values remain strings. Records whose parsed fields concatenate to whitespace only are skipped, including blank lines, delimiter-only and quoted-empty records. They do not establish headers, consume max_rows or alone set truncated; structural budgets still apply." } Example request body: { "url": "https://grist.tools/samples/documents/garden-plots.csv", "header": true } Example response: { "source_url": "https://grist.tools/samples/documents/garden-plots.csv", "delimiter": ",", "columns": [ "plot", "holder", "crop", "area_m2", "notes" ], "rows": [ { "plot": "A1", "holder": "Maple Group", "crop": "tomatoes", "area_m2": "12", "notes": "Stakes needed, north side" }, { "plot": "A2", "holder": "Oak Group", "crop": "beans", "area_m2": "08", "notes": "" }, { "plot": "B1", "holder": "Willow Group", "crop": "herbs, mixed", "area_m2": "6", "notes": "Shares water with \"B2\"" }, { "plot": "B2", "holder": "Birch Group", "crop": "squash", "area_m2": "10.50", "notes": "Compost added in April" } ], "row_count": 4, "truncated": false } - [docx-to-markdown](https://grist.tools/docs/docx-to-markdown): Converts a Word .docx to Markdown, surfacing anything that did not translate as warnings. (family documents, category Content, $0.005, timeout 40000ms) POST https://grist.tools/v1/docx-to-markdown Why it is worth paying for: A faithful docx→Markdown conversion needs a real OOXML engine plus a hardened fetch: an agent that fetches the URL itself walks past every SSRF control, and carrying mammoth per call is not something an agent does inline. Price: $0.005 · max_timeout_seconds: 61 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the .docx, up to 2048 characters; redirects are followed, the body must be a zip, and mammoth must find a Word document in it." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/garden-handbook.docx" } Example response: { "source_url": "https://grist.tools/samples/documents/garden-handbook.docx", "chars": 341, "warnings": [], "markdown": "# Community Garden Handbook\n\nPlots are assigned for **one season** and renewed each spring.\n\n## Watering\n\n- Water early in the morning.\n- Use the rain barrels first.\n- Log each watering on the shed board.\n\n## Tools\n\nReturn every tool to the shed _before sunset_.\n\nThe seasonal calendar lives on [the garden site](https://grist.tools/)." } - [epub-to-text](https://grist.tools/docs/epub-to-text): Extracts an EPUB’s reading-order text and Dublin Core metadata, chapter by chapter. (family documents, category Content, $0.005, timeout 45000ms) POST https://grist.tools/v1/epub-to-text Why it is worth paying for: Reading an EPUB means walking a zip, its container, its OPF spine and each XHTML chapter through a hardened fetch. An agent doing that inline walks past every SSRF control, and it is fiddly enough to be worth a paid endpoint. Price: $0.005 · max_timeout_seconds: 66 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the .epub, up to 2048 characters; redirects are followed and the body must be a zip whose META-INF/container.xml points at an OPF package." }, "max_chapters": { "default": 500, "description": "How many leading XHTML/HTML spine chapters to extract, 1 to 5000; a shorter book is not an error, and a longer one is cut and reported as truncated.", "type": "integer", "minimum": 1, "maximum": 5000 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/seed-saving.epub", "max_chapters": 500 } Example response: { "source_url": "https://grist.tools/samples/documents/seed-saving.epub", "metadata": { "title": "A Short Guide to Seed Saving", "creator": "Grist Samples", "language": "en" }, "chapter_count": 3, "extracted_chapters": 3, "truncated": false, "chars": 441, "chapters": [ { "href": "OEBPS/ch1.xhtml", "text": "Why Save Seeds\n\nSaving seeds keeps a variety that grows well in your own soil.\n\nIt also costs nothing and makes next spring easier to plan." }, { "href": "OEBPS/ch2.xhtml", "text": "Choosing Plants\n\nPick the healthiest plants and let a few fruits ripen fully.\n\nBeans and peas are the easiest to start with.\n\nTomatoes need their seeds rinsed and dried." }, { "href": "OEBPS/ch3.xhtml", "text": "Drying and Storing\n\nDry seeds on paper for a week, away from direct sun.\n\nStore them in labelled envelopes in a cool, dark place." } ], "text": "Why Save Seeds\n\nSaving seeds keeps a variety that grows well in your own soil.\n\nIt also costs nothing and makes next spring easier to plan.\n\nChoosing Plants\n\nPick the healthiest plants and let a few fruits ripen fully.\n\nBeans and peas are the easiest to start with.\n\nTomatoes need their seeds rinsed and dried.\n\nDrying and Storing\n\nDry seeds on paper for a week, away from direct sun.\n\nStore them in labelled envelopes in a cool, dark place." } - [html-to-pdf](https://grist.tools/docs/html-to-pdf): Renders a web page or an inline HTML string to a PDF, returned base64-encoded with its page count. (family documents, category Content, $0.005, timeout 12000ms) POST https://grist.tools/v1/html-to-pdf Why it is worth paying for: It needs a headless browser AND a hardened fetch: an agent that renders a URL itself walks the browser straight past every SSRF control, and carrying a Chromium is not something an agent does per call. Price: $0.005 · max_timeout_seconds: 33 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "description": "Absolute http(s) URL of an HTML or XHTML page to fetch and render, up to 2048 characters; redirects are followed, the body is capped at 5 MB, and exactly one of url or html is required.", "type": "string", "maxLength": 2048, "format": "uri" }, "html": { "type": "string", "minLength": 1, "maxLength": 262144, "description": "Self-contained HTML. Inline PNG data images are supported; external images, stylesheets and fonts are blocked. JavaScript is disabled." }, "format": { "default": "A4", "description": "Paper size of every page, A4 by default.", "type": "string", "enum": [ "A3", "A4", "A5", "Letter", "Legal", "Tabloid" ] }, "landscape": { "default": false, "description": "Print in landscape orientation instead of portrait.", "type": "boolean" } }, "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/render-print.html", "format": "A4", "landscape": false } Example response: { "source_url": "https://grist.tools/samples/web/render-print.html", "pages": 1, "bytes": 38404, "format": "A4", "landscape": false, "pdf_base64": "JVBERi0xLjcKJYGBgYEKCjYgMCBvYmoKPDwKL0ZpbHRlciAvRmxhdGVEZWNvZGUK...", "rendered_at": "2026-09-01T12:00:00.000Z" } - [json-to-csv](https://grist.tools/docs/json-to-csv): Renders JSON objects as raw CSV with no spreadsheet formula protection; columns are inferred or given, nested values are JSON-encoded. (family documents, category Content, $0.005, timeout 30000ms) POST https://grist.tools/v1/json-to-csv Why it is worth paying for: Correct CSV output is quoting and escaping to RFC-4180, not a join on commas, and the column set has to be reconciled across rows with different keys, fiddly enough that paying for it beats hand-rolling it per call. Headers and string values preserve formula prefixes (=, +, -, @), tabs and carriage returns without added apostrophes or trimming. CSV quoting escapes delimiters, quotes and newlines but does not neutralize spreadsheet formulas. Use spreadsheet-specific import controls for untrusted data. CRLF separates records. Price: $0.005 · max_timeout_seconds: 51 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of a UTF-8 JSON document (a leading BOM is allowed) that is an array of objects or a single object, up to 2048 characters; redirects are followed." }, "columns": { "description": "Selects columns in exactly the supplied order; a missing key produces an empty field, except an absent __proto__ column renders {}. When omitted, visits only rows retained by max_rows, in input order, appending each unseen key once. Within each row, JavaScript enumerates canonical array-index keys (0 through 4294967294) numerically ascending, then other string keys in insertion order. Names such as 01, -1 and 4294967295 are ordinary string keys. New keys from later rows are appended, not globally sorted; textual JSON key order is not preserved for array-index keys.", "minItems": 1, "maxItems": 1000, "type": "array", "items": { "type": "string", "minLength": 1, "maxLength": 255 } }, "max_rows": { "default": 10000, "description": "Maximum rows rendered, 1 to 100000, default 10000; later rows are dropped and truncated is set, and inferred columns come only from the kept rows.", "type": "integer", "minimum": 1, "maximum": 100000 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/garden-harvest.json" } Example response: { "source_url": "https://grist.tools/samples/documents/garden-harvest.json", "columns": [ "week", "crop", "kg", "plots", "note" ], "csv": "week,crop,kg,plots,note\r\n1,radish,2.5,\"[\"\"A1\"\",\"\"B2\"\"]\",\r\n1,lettuce,4,,\"first cut, outer leaves only\"\r\n2,peas,3.25,\"[\"\"A2\"\"]\",", "row_count": 3, "truncated": false } - [markdown-to-pdf](https://grist.tools/docs/markdown-to-pdf): Renders a Markdown document to a PDF, returned base64-encoded with its page count. (family documents, category Content, $0.005, timeout 12000ms) POST https://grist.tools/v1/markdown-to-pdf Why it is worth paying for: Turning styled Markdown into a paginated PDF needs a real layout engine AND a hardened fetch: an agent that renders a URL itself walks the browser straight past every SSRF control, and carrying a Chromium is not something an agent does per call. Price: $0.005 · max_timeout_seconds: 33 · response cap: 2097152 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "description": "Absolute http(s) URL of a UTF-8 Markdown text file to fetch and render, up to 2048 characters; redirects are followed, the body is capped at 2 MB, and exactly one of url or markdown is required.", "type": "string", "maxLength": 2048, "format": "uri" }, "markdown": { "type": "string", "minLength": 1, "maxLength": 262144, "description": "Markdown with escaped raw HTML. PNG/JPEG/GIF/WebP data-image syntax is accepted; SVG data images and external image loading are blocked." }, "format": { "default": "A4", "description": "Paper size of every page, A4 by default.", "type": "string", "enum": [ "A3", "A4", "A5", "Letter", "Legal", "Tabloid" ] }, "landscape": { "default": false, "description": "Print in landscape orientation instead of portrait.", "type": "boolean" } }, "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/render-notes.md", "format": "A4", "landscape": false } Example response: { "source_url": "https://grist.tools/samples/documents/render-notes.md", "pages": 1, "bytes": 69844, "format": "A4", "landscape": false, "pdf_base64": "JVBERi0xLjcKJYGBgYEKCjkgMCBvYmoKPDwKL0ZpbHRlciAvRmxhdGVEZWNvZGUK...", "rendered_at": "2026-09-01T12:00:00.000Z" } - [ocr-image-to-text](https://grist.tools/docs/ocr-image-to-text): Reads the text out of an image (scan, screenshot, photo) with Tesseract, returned with a mean confidence. (family documents, category Content, $0.005, timeout 45000ms) POST https://grist.tools/v1/ocr-image-to-text Why it is worth paying for: OCR needs a real engine and its language data AND a hardened fetch: an agent that pulls the image URL itself walks past every SSRF control, and carrying a WASM Tesseract plus a language model is not something an agent does per call. Price: $0.005 · max_timeout_seconds: 66 · response cap: 10485760 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the image, up to 2048 characters and 10 MB; redirects are followed, the bytes must be PNG, JPEG, TIFF, WebP or GIF by signature, and at most 40 megapixels." }, "lang": { "default": "eng", "description": "Tesseract language model to read the text with; only languages whose data ships with the service are accepted, today eng (the default).", "type": "string", "enum": [ "eng" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/ocr-scan.png", "lang": "eng" } Example response: { "source_url": "https://grist.tools/samples/images/ocr-scan.png", "lang": "eng", "image_format": "png", "text": "SAMPLE SCAN FOR OCR\nSECOND LINE 2026\n", "confidence": 86 } - [pdf-merge](https://grist.tools/docs/pdf-merge): Concatenates 2–8 source PDFs, in order, into one PDF returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding. (family documents, category Content, $0.005, timeout 45000ms) POST https://grist.tools/v1/pdf-merge Why it is worth paying for: Merging PDFs needs a real PDF library and a hardened fetch over several URLs at once: an agent doing it inline walks every one past the SSRF controls, and the joined output must be byte-reproducible. Price: $0.005 · max_timeout_seconds: 66 · response cap: 8388608 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "urls": { "minItems": 2, "maxItems": 8, "type": "array", "items": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of one source PDF, up to 2048 characters; redirects are followed and the body must start with the %PDF- signature." }, "description": "The source PDFs, 2 to 8 URLs fetched one at a time and concatenated in this order; each is read up to 8 MiB and an encrypted source is refused." } }, "required": [ "urls" ], "additionalProperties": false } Example request body: { "urls": [ "https://grist.tools/samples/documents/merge-intro.pdf", "https://grist.tools/samples/documents/merge-appendix.pdf" ] } Example response: { "sources": [ { "url": "https://grist.tools/samples/documents/merge-intro.pdf", "pages": 1 }, { "url": "https://grist.tools/samples/documents/merge-appendix.pdf", "pages": 2 } ], "pages": 3, "bytes": 1474, "pdf_base64": "JVBERi0xLjcKJYGBgYEKCjQgMCBvYmoKPDwKL0ZpbHRlciAvRmxhdGVEZWNvZGUK..." } - [pdf-metadata](https://grist.tools/docs/pdf-metadata): Reads a PDF's page count, document-info fields, version and encryption flag without rendering it. (family documents, category Content, $0.005, timeout 25000ms) POST https://grist.tools/v1/pdf-metadata Why it is worth paying for: Reading a PDF trailer needs a real PDF parser and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying a PDF library per call is not something an agent does inline. Price: $0.005 · max_timeout_seconds: 46 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the PDF, up to 2048 characters; redirects are followed, the body must start with %PDF-, and an encrypted file is still described without being decrypted." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/metadata-brochure.pdf" } Example response: { "source_url": "https://grist.tools/samples/documents/metadata-brochure.pdf", "bytes": 1010, "version": "1.7", "pages": 1, "encrypted": false, "title": "Sample Brochure", "author": "Grist Samples", "subject": "Public test document", "keywords": "sample metadata test", "creator": "grist.tools", "producer": "grist.tools", "creation_date": "2000-01-01T00:00:00.000Z", "modification_date": "2000-01-01T00:00:00.000Z" } - [pdf-split](https://grist.tools/docs/pdf-split): Extracts a 1-indexed page selection (e.g. "1,3-4") from a PDF into one new PDF, returned base64-encoded. Maximum output PDF size: 25 MiB before base64 encoding. (family documents, category Content, $0.005, timeout 30000ms) POST https://grist.tools/v1/pdf-split Why it is worth paying for: Splitting a PDF needs a real PDF library and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and the output must be byte-reproducible, which a naive save is not. Price: $0.005 · max_timeout_seconds: 51 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the source PDF, up to 2048 characters; redirects are followed, the body must start with %PDF-, and an encrypted file is refused." }, "ranges": { "type": "string", "minLength": 1, "maxLength": 512, "description": "Comma-separated 1-indexed pages or inclusive ranges such as \"1,3-4\", up to 512 characters; pages are emitted in the order written, repeats are kept, and a descending range or a page past the end is an error." } }, "required": [ "url", "ranges" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/split-five-pages.pdf", "ranges": "1,3-4" } Example response: { "source_url": "https://grist.tools/samples/documents/split-five-pages.pdf", "source_pages": 5, "pages": 3, "bytes": 1392, "pdf_base64": "JVBERi0xLjcKJYGBgYEKCjQgMCBvYmoKPDwKL0ZpbHRlciAvRmxhdGVEZWNvZGUK..." } - [pdf-to-markdown](https://grist.tools/docs/pdf-to-markdown): Converts a PDF to Markdown, inferring headings from font size (best-effort, not layout-perfect; no OCR). (family documents, category Content, $0.005, timeout 40000ms) POST https://grist.tools/v1/pdf-to-markdown Why it is worth paying for: A PDF has no heading structure, only sized glyphs; recovering Markdown needs a real PDF engine plus a hardened fetch, which an agent cannot do inline without walking past the SSRF controls. Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit. Price: $0.005 · max_timeout_seconds: 61 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the PDF, up to 2048 characters; redirects are followed and the body must start with the %PDF- signature." }, "max_pages": { "default": 100, "description": "How many leading pages to convert, 1 to 1000; a shorter document is not an error, and a longer one is cut and reported as truncated.", "type": "integer", "minimum": 1, "maximum": 1000 } }, "required": [ "url" ], "additionalProperties": false, "description": "Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit. PDF text extraction accepts at most 1000000 raw and separately 1000000 output UTF-16 code units, 10000 items per page and 100000 items across selected pages. Exceeding a limit rejects the call with 413 too_large. truncated refers only to max_pages." } Example request body: { "url": "https://grist.tools/samples/documents/markdown-guide.pdf", "max_pages": 100 } Example response: { "source_url": "https://grist.tools/samples/documents/markdown-guide.pdf", "total_pages": 1, "extracted_pages": 1, "truncated": false, "chars": 249, "markdown": "# Field Notes\n\n## A short guide in three parts\n\nThis sample shows how headings are inferred.\n\nLarger text becomes a heading, body text stays plain.\n\n### Second section\n\nEach line of the page becomes its own block.\n\nThe document ends after this line." } - [pdf-to-text](https://grist.tools/docs/pdf-to-text): Extracts the text layer of a PDF, page by page, in reading order (no OCR). (family documents, category Content, $0.005, timeout 40000ms) POST https://grist.tools/v1/pdf-to-text Why it is worth paying for: Pulling a PDF text layer needs a real PDF engine and a hardened fetch: an agent that fetches the URL itself walks past every SSRF control, and carrying pdfjs per call is not something an agent does inline. Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit. Price: $0.005 · max_timeout_seconds: 61 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the PDF, up to 2048 characters; redirects are followed and the body must start with the %PDF- signature." }, "max_pages": { "default": 100, "description": "How many leading pages to extract, 1 to 1000; a shorter document is not an error, and a longer one is cut and reported as truncated.", "type": "integer", "minimum": 1, "maximum": 1000 } }, "required": [ "url" ], "additionalProperties": false, "description": "Offline, best-effort text-layer extraction; no OCR. No external CMap or standard-font data is configured or fetched. PDFs with self-contained character mappings can yield text, including CJK; documents requiring unavailable predefined CMaps can return empty or partial text without an error. Embedded ToUnicode alone does not guarantee extraction when the encoding still needs an external CMap. Missing standard-font data does not necessarily prevent extraction. A successful response does not certify complete text; truncated reports only the max_pages limit. PDF text extraction accepts at most 1000000 raw and separately 1000000 output UTF-16 code units, 10000 items per page and 100000 items across selected pages. Exceeding a limit rejects the call with 413 too_large. truncated refers only to max_pages." } Example request body: { "url": "https://grist.tools/samples/documents/report.pdf", "max_pages": 100 } Example response: { "source_url": "https://grist.tools/samples/documents/report.pdf", "total_pages": 2, "extracted_pages": 2, "truncated": false, "chars": 122, "text": "Sample Report\nThis is a public test document.\nIt has two pages of plain text.\n\nPage Two\nThe second page closes the report." } - [pptx-to-text](https://grist.tools/docs/pptx-to-text): Extracts the slide text of a PowerPoint .pptx, slide by slide, in order (body text only, up to 1,000 slides). (family documents, category Content, $0.005, timeout 40000ms) POST https://grist.tools/v1/pptx-to-text Why it is worth paying for: Reading .pptx slide text needs to unpack an OOXML zip through a hardened fetch: an agent fetching the URL itself walks past every SSRF control, and the parsing is fiddly enough to be worth a paid endpoint. Price: $0.005 · max_timeout_seconds: 61 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the .pptx, up to 2048 characters; redirects are followed and the body must be a zip holding ppt/slides/slideN.xml parts, at most 1,000 of them." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/garden-plan.pptx" } Example response: { "source_url": "https://grist.tools/samples/documents/garden-plan.pptx", "slide_count": 3, "chars": 244, "slides": [ { "number": 1, "text": "Spring Planting Plan\nCommunity garden season overview" }, { "number": 2, "text": "Beds and Crops\nNorth beds: beans and peas\nSouth beds: tomatoes & peppers\nHerb spiral: basil, thyme, mint" }, { "number": 3, "text": "Next Steps\nOrder compost by March\nRepair the east fence\nSchedule the first work day" } ], "text": "Spring Planting Plan\nCommunity garden season overview\n\nBeds and Crops\nNorth beds: beans and peas\nSouth beds: tomatoes & peppers\nHerb spiral: basil, thyme, mint\n\nNext Steps\nOrder compost by March\nRepair the east fence\nSchedule the first work day" } - [xlsx-to-json](https://grist.tools/docs/xlsx-to-json): Reads an Excel .xlsx into JSON rows, keyed by a header row or as arrays; dates become ISO strings. Archives are limited to 64 worksheets. (family documents, category Content, $0.005, timeout 40000ms) POST https://grist.tools/v1/xlsx-to-json Why it is worth paying for: Parsing a real .xlsx needs a spreadsheet engine plus a hardened fetch: an agent fetching the URL itself walks past every SSRF control, and carrying exceljs per call is not something an agent does inline. Price: $0.005 · max_timeout_seconds: 61 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the .xlsx, up to 2048 characters; redirects are followed, the body must be a zip, and exceljs must load at least one worksheet from it (at most 64)." }, "sheet": { "description": "Exact name of the one worksheet to return, 1 to 255 characters; omitted, every worksheet is returned, and a name not in the workbook is a 422 listing the available names.", "type": "string", "minLength": 1, "maxLength": 255 }, "header_row": { "default": 1, "description": "The 1-based row holding the column names, 0 to 1000, default 1; data starts on the next row, and 0 returns every row as a positional array with columns set to null.", "type": "integer", "minimum": 0, "maximum": 1000 }, "max_rows": { "default": 1000, "description": "Maximum data rows returned per worksheet, 1 to 50000, default 1000; a sheet with more rows is cut and marked truncated.", "type": "integer", "minimum": 1, "maximum": 50000 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/documents/garden-inventory.xlsx", "header_row": 1, "max_rows": 1000 } Example response: { "source_url": "https://grist.tools/samples/documents/garden-inventory.xlsx", "sheet_count": 2, "sheets": [ { "name": "Seeds", "columns": [ "Variety", "Crop", "Packets", "Sown On", "Organic" ], "rows": [ { "Variety": "Early Scarlet", "Crop": "Radish", "Packets": 4, "Sown On": "2026-03-14T00:00:00.000Z", "Organic": true }, { "Variety": "Green Arrow", "Crop": "Pea", "Packets": 6, "Sown On": "2026-03-21T00:00:00.000Z", "Organic": true }, { "Variety": "Golden Cherry", "Crop": "Tomato", "Packets": 2, "Sown On": "2026-04-11T00:00:00.000Z", "Organic": false } ], "row_count": 3, "truncated": false }, { "name": "Tools", "columns": [ "Tool", "Count", "Location" ], "rows": [ { "Tool": "Spade", "Count": 3, "Location": "Shed" }, { "Tool": "Watering can", "Count": 5, "Location": "Rain barrels" } ], "row_count": 2, "truncated": false } ] } ## Media Requires a heavy dependency. - [archive-extract-file](https://grist.tools/docs/archive-extract-file): Extracts one named entry from a remote zip, tar or tar.gz and returns its bytes (base64) with a sha256. (family archives, category Compute, $0.003, timeout 90000ms) POST https://grist.tools/v1/archive-extract-file Why it is worth paying for: Pulling a single file out of an archive without unpacking the whole thing needs a real reader, a hardened fetch and the decompression-bomb care that caps a single entry twice, not something an agent should run inline. Price: $0.003 · max_timeout_seconds: 111 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the archive (at most 25 MB, also the cap on an inflated tar.gz), up to 2048 characters; redirects are followed and the format (zip, tar or gzip-compressed tar) is decided from the magic bytes, not the extension." }, "entry": { "type": "string", "minLength": 1, "maxLength": 4096, "description": "Exact entry path to extract, 1 to 4096 characters, matched case-sensitively against the names archive-inspect lists; directories, links, encrypted entries and entries over 25 MB are refused." } }, "required": [ "url", "entry" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/archives/bundle.zip", "entry": "data/squares.csv" } Example response: { "source_url": "https://grist.tools/samples/archives/bundle.zip", "kind": "zip", "entry_name": "data/squares.csv", "size": 31, "compression": "deflate", "crc32": "89761172", "sha256": "202663eb5661a93d557520ce47e40bb6ccdcd7fded543fca46592d2291e2c8cf", "content_base64": "bixzcXVhcmUKMSwxCjIsNAozLDkKNCwxNgo1LDI1Cg==" } - [archive-inspect](https://grist.tools/docs/archive-inspect): Lists the entries of a remote zip, tar or tar.gz (paths, sizes, compression and CRC) without extracting any of them. (family archives, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/archive-inspect Why it is worth paying for: Seeing inside an archive without extracting it needs a real zip/tar reader and a hardened fetch, plus the decompression-bomb care that reads declared sizes instead of inflating, not something an agent does inline. Price: $0.003 · max_timeout_seconds: 81 · response cap: 104857600 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the archive (at most 100 MB, also the cap on an inflated tar.gz), up to 2048 characters; redirects are followed and the format (zip, tar or gzip-compressed tar) is decided from the magic bytes, not the extension." }, "max_entries": { "default": 1000, "description": "How many entries to list, 1 to 10000 (default 1000), taken in archive order and then sorted by name; any beyond the limit still count in total_entries and set truncated.", "type": "integer", "minimum": 1, "maximum": 10000 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/archives/bundle.zip", "max_entries": 1000 } Example response: { "source_url": "https://grist.tools/samples/archives/bundle.zip", "kind": "zip", "compression": "none", "entry_count": 5, "total_entries": 5, "total_uncompressed_bytes": 159, "encrypted": false, "truncated": false, "entries": [ { "name": "data/", "is_directory": true, "entry_type": "directory", "size": 0, "compressed_size": 0, "compression": "none", "encrypted": false, "crc32": "00000000" }, { "name": "data/squares.csv", "is_directory": false, "entry_type": "file", "size": 31, "compressed_size": 33, "compression": "deflate", "encrypted": false, "crc32": "89761172" }, { "name": "docs/", "is_directory": true, "entry_type": "directory", "size": 0, "compressed_size": 0, "compression": "none", "encrypted": false, "crc32": "00000000" }, { "name": "docs/readme.md", "is_directory": false, "entry_type": "file", "size": 116, "compressed_size": 94, "compression": "deflate", "encrypted": false, "crc32": "b6cd27c3" }, { "name": "hello.txt", "is_directory": false, "entry_type": "file", "size": 12, "compressed_size": 12, "compression": "store", "encrypted": false, "crc32": "af083b2d" } ] } - [audio-convert](https://grist.tools/docs/audio-convert): Transcodes an audio file between WAV, MP3, FLAC, AAC and Opus, with an optional bit rate and sample rate. (family av, category Compute, $0.003, timeout 94000ms) POST https://grist.tools/v1/audio-convert Why it is worth paying for: Re-encoding audio needs ffmpeg AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 115 · response cap: 104857600 bytes · max input duration: 3600 s Input limits: Longer input is refused with too_large (413), and input whose duration cannot be measured with unprocessable (422). Both are decided before any encoding, so the call is never settled. With to set to aac, bitrate_kbps is at most 192; a higher value is refused with invalid_input (400). Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the source file, up to 2048 characters; the first audio stream of a recognised container (MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF or MP3) is converted, and it must last at most 3600 seconds." }, "to": { "default": "mp3", "description": "Output format: wav (16-bit PCM), mp3 (LAME), flac, aac (ADTS stream) or opus (in an Ogg container); defaults to mp3.", "type": "string", "enum": [ "wav", "mp3", "flac", "aac", "opus" ] }, "bitrate_kbps": { "description": "Target bit rate in kbps, lossy formats only (mp3, aac, opus): 8 to 512, and at most 192 for aac.", "type": "integer", "minimum": 8, "maximum": 512 }, "sample_rate": { "description": "Output sample rate in Hz, one of 8000, 11025, 16000, 22050, 32000, 44100 or 48000; when omitted ffmpeg chooses the rate, normally the source rate.", "type": "integer", "minimum": -9007199254740991, "maximum": 9007199254740991 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/convert-beeps.flac", "to": "wav" } Example response: { "source_url": "https://grist.tools/samples/av/convert-beeps.flac", "bytes": 32086, "source_container": "flac", "format": "wav", "bitrate_kbps": null, "sample_rate": null, "audio_bytes": 32044, "audio_base64": "UklGRiR9AABXQVZFZm10IBAAAAABAAEAQB8AAIA+AAACABAAZGF0YQB9AAAA4ADo..." } - [audio-extract](https://grist.tools/docs/audio-extract): Extracts the audio track from a video (or any container) and returns it in a chosen audio format. (family av, category Compute, $0.003, timeout 94000ms) POST https://grist.tools/v1/audio-extract Why it is worth paying for: Demuxing and re-encoding an audio track needs ffmpeg AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 115 · response cap: 209715200 bytes · max input duration: 3600 s Input limits: Longer input is refused with too_large (413), and input whose duration cannot be measured with unprocessable (422). Both are decided before any encoding, so the call is never settled. Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of a video or audio file, up to 2048 characters, in a recognised container (MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF or MP3); the selected audio stream must last at most 3600 seconds." }, "stream": { "default": 0, "description": "Zero-based index among the audio streams only (0 is the first audio track), 0 to 63; an index the file does not have is refused as unprocessable.", "type": "integer", "minimum": 0, "maximum": 63 }, "to": { "default": "mp3", "description": "Output format for the extracted track: wav (16-bit PCM), mp3 (LAME), flac, aac (ADTS stream) or opus (in an Ogg container); defaults to mp3.", "type": "string", "enum": [ "wav", "mp3", "flac", "aac", "opus" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/probe-clip.mov", "stream": 0, "to": "wav" } Example response: { "source_url": "https://grist.tools/samples/av/probe-clip.mov", "bytes": 33850, "container": "mp4", "stream_index": 0, "format": "wav", "audio_bytes": 32044, "audio_base64": "UklGRiR9AABXQVZFZm10IBAAAAABAAEAQB8AAIA+AAACABAAZGF0YQB9AAAA4ADo..." } - [color-palette-extract](https://grist.tools/docs/color-palette-extract): Extracts an image's dominant colours as a small palette (hex, RGB and coverage fraction). (family images, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/color-palette-extract Why it is worth paying for: Extracting a palette safely needs a real decoder and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and a byte-reproducible quantiser is not something an agent writes inline. Price: $0.003 · max_timeout_seconds: 81 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the image, up to 2048 characters; redirects are followed and the format is detected from the bytes, not the content-type." }, "colors": { "default": 5, "description": "Maximum number of palette colours to return, 1 to 16, most common first; fewer come back when the image has fewer distinct colour buckets.", "type": "integer", "minimum": 1, "maximum": 16 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/palette-bands.png", "colors": 5 } Example response: { "source_url": "https://grist.tools/samples/images/palette-bands.png", "source_format": "png", "sample_size": 64, "count": 5, "dominant": "#3b4a5a", "colors": [ { "hex": "#3b4a5a", "rgb": [ 59, 74, 90 ], "fraction": 0.375 }, { "hex": "#d8c9a3", "rgb": [ 216, 201, 163 ], "fraction": 0.25 }, { "hex": "#2f8f83", "rgb": [ 47, 143, 131 ], "fraction": 0.1875 }, { "hex": "#e0735a", "rgb": [ 224, 115, 90 ], "fraction": 0.125 }, { "hex": "#f4efe6", "rgb": [ 244, 239, 230 ], "fraction": 0.0625 } ] } - [favicon-extract](https://grist.tools/docs/favicon-extract): Discovers a page's favicons and returns the best one decoded to a normalised PNG with its size. (family images, category Compute, $0.003, timeout 45000ms) POST https://grist.tools/v1/favicon-extract Why it is worth paying for: It fans out over URLs a page author wrote and must decode several icon formats: an agent that does this itself walks past every SSRF control and becomes a request amplifier, and hand-rolling ICO/SVG decoding is not inline work. Price: $0.003 · max_timeout_seconds: 66 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the web page whose icon tags (rel icon, shortcut, apple-touch-icon or mask-icon) are read, up to 2048 characters; /favicon.ico on its origin is added as a last candidate and icons are only fetched from the same registrable domain." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/web/fieldnotes/index.html" } Example response: { "page_url": "https://grist.tools/samples/web/fieldnotes/index.html", "icons": [ { "url": "https://grist.tools/samples/web/fieldnotes/icon-32.png", "rel": "icon", "type": "image/png", "sizes": "32x32" }, { "url": "https://grist.tools/samples/web/fieldnotes/icon.svg", "rel": "icon", "type": "image/svg+xml", "sizes": null }, { "url": "https://grist.tools/favicon.ico", "rel": "icon", "type": null, "sizes": null } ], "selected": { "url": "https://grist.tools/samples/web/fieldnotes/icon-32.png", "source_format": "png", "width": 32, "height": 32, "bytes": 106, "image_base64": "iVBORw0KGgoAAAANSUhEUgAAACAAAAAgCAIAAAD8GO2jAAAACXBIWXMAAAsSAAALEgHS3X78AAAAMUlEQVRIx2PQjw2gKWIYtWDUgiFswaf3z0hCoxaMWjBqwagFoxYMTwtGq8xRC0aQBQBMN1tMlMrjbQAAAABJRU5ErkJggg==" } } - [file-type-detect](https://grist.tools/docs/file-type-detect): Detects a remote file's true type from its magic bytes, returning the canonical extension and media type. (family archives, category Compute, $0.003, timeout 30000ms) POST https://grist.tools/v1/file-type-detect Why it is worth paying for: Trusting a URL extension or a server content-type is exactly how an agent mislabels a payload; deciding the type from the bytes needs a signature database and a hardened fetch an agent should not run inline. Price: $0.003 · max_timeout_seconds: 51 · response cap: 33554432 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the file, up to 2048 characters; redirects are followed, the whole body (at most 32 MB) is downloaded, and the type comes from its magic bytes, ignoring the extension and content-type header." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/archives/bundle.zip" } Example response: { "source_url": "https://grist.tools/samples/archives/bundle.zip", "bytes": 639, "detected": true, "ext": "zip", "mime": "application/zip" } - [image-compress](https://grist.tools/docs/image-compress): Re-encodes an image in its own format at a lower quality to shrink the file, keeping its dimensions. (family images, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/image-compress Why it is worth paying for: Compressing safely needs a real codec and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and the output must be byte-reproducible, which an ad-hoc re-encode is not. Price: $0.003 · max_timeout_seconds: 81 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of a JPEG, PNG, WebP or GIF image, up to 2048 characters; redirects are followed and the output keeps the source format and dimensions." }, "quality": { "default": 75, "description": "Encoder quality from 1 to 100, lower gives a smaller file; drives JPEG and WebP re-compression and PNG palette quantisation, and is ignored for GIF.", "type": "integer", "minimum": 1, "maximum": 100 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/compress-scene.jpg", "quality": 60 } Example response: { "source_url": "https://grist.tools/samples/images/compress-scene.jpg", "format": "jpeg", "width": 96, "height": 64, "source_bytes": 1836, "bytes": 758, "saved_bytes": 1078, "ratio": 0.4129, "image_base64": "/9j/2wBDAA0JCgsKCA0LCgsODg0PEyAVExISEyccHhcgLikxMC4pLSwzOko+MzZG..." } - [image-convert](https://grist.tools/docs/image-convert): Transcodes an image to PNG, JPEG or WebP at a chosen quality, returned base64-encoded. (family images, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/image-convert Why it is worth paying for: Transcoding safely needs a real codec and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and the output must be byte-reproducible, which an ad-hoc re-encode is not. Price: $0.003 · max_timeout_seconds: 81 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the source image, up to 2048 characters; redirects are followed and the source format is detected from the bytes." }, "format": { "type": "string", "enum": [ "png", "jpeg", "webp" ], "description": "Target format to encode to, png, jpeg or webp; the pixel dimensions are kept and metadata is dropped." }, "quality": { "default": 80, "description": "Encoder quality from 1 to 100 for jpeg and webp output; ignored for png, which is always lossless.", "type": "integer", "minimum": 1, "maximum": 100 } }, "required": [ "url", "format" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/convert-gradient.png", "format": "webp", "quality": 80 } Example response: { "source_url": "https://grist.tools/samples/images/convert-gradient.png", "source_format": "png", "format": "webp", "mime_type": "image/webp", "width": 32, "height": 32, "bytes": 156, "image_base64": "UklGRpQAAABXRUJQVlA4IIgAAAAQBQCdASogACAALk02m02hJCQkBABMS2AE6ZiX..." } - [image-metadata](https://grist.tools/docs/image-metadata): Reads an image's EXIF/ICC/IPTC/XMP metadata and can return a copy with all metadata stripped. (family images, category Compute, $0.003, timeout 30000ms) POST https://grist.tools/v1/image-metadata Why it is worth paying for: Reading and stripping image metadata safely needs a real decoder and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and a strip must be byte-reproducible, which an ad-hoc re-encode is not. Price: $0.003 · max_timeout_seconds: 51 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the image to inspect, up to 2048 characters; redirects are followed and the format is detected from the bytes." }, "strip": { "default": false, "description": "When true, also return a same-format copy re-encoded with every metadata block removed (lossy sources at quality 80); allowed for JPEG, PNG, WebP, TIFF and GIF.", "type": "boolean" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/metadata-exif.jpg", "strip": true } Example response: { "source_url": "https://grist.tools/samples/images/metadata-exif.jpg", "format": "jpeg", "width": 32, "height": 24, "channels": 3, "space": "srgb", "depth": "uchar", "has_alpha": false, "has_profile": true, "orientation": 6, "density": 72, "is_progressive": false, "metadata": { "exif_bytes": 186, "icc_bytes": 480, "iptc_bytes": null, "xmp_bytes": null }, "bytes": 1024, "stripped": { "format": "jpeg", "bytes": 333, "image_base64": "/9j/2wBDAAYEBQYFBAYGBQYHBwYIChAKCgkJChQODwwQFxQYGBcUFhYaHSUfGhsj..." } } - [image-probe](https://grist.tools/docs/image-probe): Reports an image's format and pixel dimensions from its header, without decoding the pixels. (family images, category Compute, $0.003, timeout 20000ms) POST https://grist.tools/v1/image-probe Why it is worth paying for: Reading an image header safely needs a real decoder and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an image library per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 41 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the image, up to 2048 characters; redirects are followed, the bytes must be jpeg, png, gif, webp, tiff, avif, heic or svg by signature, and only the header is parsed." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/tiles-64x40.png" } Example response: { "source_url": "https://grist.tools/samples/images/tiles-64x40.png", "format": "png", "width": 64, "height": 40, "channels": 4, "space": "srgb", "has_alpha": true, "bytes": 10348 } - [image-resize](https://grist.tools/docs/image-resize): Resizes an image to a target width/height under a chosen fit, returned base64-encoded. (family images, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/image-resize Why it is worth paying for: Resizing safely needs a real decoder and a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and the output must be byte-reproducible, which an ad-hoc resample is not. Price: $0.003 · max_timeout_seconds: 81 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the source image, up to 2048 characters; redirects are followed, the bytes must be jpeg, png, gif, webp, tiff, avif, heic or svg by signature, and at most 40 megapixels." }, "width": { "description": "Target width in pixels, 1 to 20000; at least one of width or height is required, and a missing side follows the aspect ratio (except under fill).", "type": "integer", "minimum": 1, "maximum": 20000 }, "height": { "description": "Target height in pixels, 1 to 20000; at least one of width or height is required, and a missing side follows the aspect ratio (except under fill).", "type": "integer", "minimum": 1, "maximum": 20000 }, "fit": { "default": "inside", "description": "How the image fills the target box, as in sharp: inside fits within keeping the ratio (default), cover crops from the centre, contain pads to the full box, fill stretches, outside covers the box without cropping.", "type": "string", "enum": [ "cover", "contain", "fill", "inside", "outside" ] }, "format": { "description": "Output format; when omitted the source format is kept if it is png, jpeg or webp, otherwise the output is PNG.", "type": "string", "enum": [ "png", "jpeg", "webp" ] }, "quality": { "default": 80, "description": "Encoder quality from 1 to 100, default 80, used by the lossy jpeg and webp encoders and ignored for png.", "type": "integer", "minimum": 1, "maximum": 100 }, "allow_enlarge": { "default": false, "description": "When true the image may be scaled up past its native size; default false keeps each side at most its source size.", "type": "boolean" } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/tiles-64x40.png", "width": 32, "fit": "inside" } Example response: { "source_url": "https://grist.tools/samples/images/tiles-64x40.png", "source_format": "png", "source_width": 64, "source_height": 40, "format": "png", "mime_type": "image/png", "fit": "inside", "width": 32, "height": 20, "bytes": 779, "image_base64": "iVBORw0KGgoAAAANSUhEUgAAACAAAAAUCAYAAADskT9PAAAACXBIWXMAAAsSAAAL..." } - [media-probe](https://grist.tools/docs/media-probe): Reads a media file's container, duration, bit rate and per-stream codec/resolution without decoding it. (family av, category Compute, $0.003, timeout 30000ms) POST https://grist.tools/v1/media-probe Why it is worth paying for: Probing a media file needs the ffprobe binary AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 51 · response cap: 104857600 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the audio or video file, up to 2048 characters, in a recognised container (MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF or MP3); ffprobe reads its container index and nothing is decoded." } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/probe-clip.mov" } Example response: { "source_url": "https://grist.tools/samples/av/probe-clip.mov", "bytes": 33850, "container": "mp4", "format_name": "mov,mp4,m4a,3gp,3g2,mj2", "duration_seconds": 2, "bit_rate": 135400, "streams": [ { "index": 0, "type": "video", "codec": "png", "codec_long_name": "PNG (Portable Network Graphics) image", "width": 64, "height": 36, "frame_rate": 2, "sample_rate": null, "channels": null }, { "index": 1, "type": "audio", "codec": "pcm_s16le", "codec_long_name": "PCM signed 16-bit little-endian", "width": null, "height": null, "frame_rate": null, "sample_rate": 8000, "channels": 1 } ] } - [remote-file-hash](https://grist.tools/docs/remote-file-hash): Computes the sha256, sha1 or md5 digest of a remote file, buffered in memory under a size cap, for checksum verification. (family archives, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/remote-file-hash Why it is worth paying for: Verifying an artefact against a published checksum means pulling the whole file through a hardened fetch and hashing it. An agent that fetches the URL itself walks past every SSRF control and pays the bandwidth anyway. Price: $0.003 · max_timeout_seconds: 81 · response cap: 104857600 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the file, up to 2048 characters; redirects are followed and the digest is taken over the whole response body (after any HTTP content-encoding is decoded), which must be at most 100 MB." }, "algorithm": { "default": "sha256", "description": "Digest to compute over the bytes: sha256 (default), sha1 or md5, returned as lowercase hex.", "type": "string", "enum": [ "sha256", "sha1", "md5" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/archives/bundle.zip", "algorithm": "sha256" } Example response: { "source_url": "https://grist.tools/samples/archives/bundle.zip", "bytes": 639, "algorithm": "sha256", "hash": "07af68ae7f049cf267c89c4c89df15fd1f59153773b308ea6bb7bb0d0585e5c3" } - [subtitle-convert](https://grist.tools/docs/subtitle-convert): Converts a subtitle track between SubRip (.srt) and WebVTT (.vtt), normalising the cues. (family av, category Compute, $0.003, timeout 15000ms) POST https://grist.tools/v1/subtitle-convert Why it is worth paying for: It fetches through a hardened guard and re-serialises the cue model rather than string-patching, so positioning settings, cue ids and NOTE blocks convert cleanly instead of corrupting the output an agent would have to repair. Price: $0.003 · max_timeout_seconds: 36 · response cap: 5242880 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of a SubRip or WebVTT file, at most 2048 characters; the source format is detected from the content, not the extension." }, "to": { "default": "vtt", "description": "Target format, vtt (default) or srt; cues are re-serialised, so SRT output is renumbered and VTT cue ids, settings and NOTE blocks are dropped.", "type": "string", "enum": [ "srt", "vtt" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/subtitle-captions.srt", "to": "vtt" } Example response: { "source_url": "https://grist.tools/samples/av/subtitle-captions.srt", "source_format": "srt", "target_format": "vtt", "cue_count": 3, "subtitle": "WEBVTT\n\n00:00:01.000 --> 00:00:03.500\nThe tide goes out slowly.\n\n00:00:04.000 --> 00:00:06.250\nShells line the wet sand,\none row after another.\n\n00:00:07.000 --> 00:00:09.000\nBy noon the water returns.\n" } - [subtitle-extract](https://grist.tools/docs/subtitle-extract): Extracts a text subtitle track from a media container (MKV, MP4, WebM) as SubRip or WebVTT. (family av, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/subtitle-extract Why it is worth paying for: Demuxing a caption stream out of a container needs ffmpeg AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 81 · response cap: 209715200 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of a media container that can carry subtitles (MKV/WebM, MP4/MOV, AVI or Ogg), at most 2048 characters." }, "stream": { "default": 0, "description": "Zero-based index counted among the subtitle streams only (0 to 63, default 0); a missing stream fails with no_subtitle_stream.", "type": "integer", "minimum": 0, "maximum": 63 }, "to": { "default": "srt", "description": "Format of the returned captions, srt (default) or vtt, written by ffmpeg.", "type": "string", "enum": [ "srt", "vtt" ] } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/subtitle-movie.mkv", "stream": 0, "to": "srt" } Example response: { "source_url": "https://grist.tools/samples/av/subtitle-movie.mkv", "bytes": 1486, "container": "matroska", "stream_index": 0, "format": "srt", "cue_count": 2, "subtitle": "1\n00:00:00,500 --> 00:00:02,000\nThe kettle starts to hum.\n\n2\n00:00:02,500 --> 00:00:04,000\nSteam rises from the spout.\n\n" } - [svg-to-png](https://grist.tools/docs/svg-to-png): Rasterises an SVG to a PNG at a chosen resolution, returned base64-encoded. (family images, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/svg-to-png Why it is worth paying for: Rasterising an SVG safely needs a real renderer and a hardened fetch: an SVG can point at internal addresses, so an agent that renders one itself walks past every SSRF control, and the output must be byte-reproducible. Price: $0.003 · max_timeout_seconds: 81 · response cap: 26214400 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the SVG, up to 2048 characters; redirects are followed and the body must look like SVG markup; external references in it are never fetched." }, "width": { "description": "Target width in pixels, 1 to 20000; when width or height is set the rendered raster is resized to fit inside that box, keeping its aspect ratio, and may be scaled up.", "type": "integer", "minimum": 1, "maximum": 20000 }, "height": { "description": "Target height in pixels, 1 to 20000; when width or height is set the rendered raster is resized to fit inside that box, keeping its aspect ratio, and may be scaled up.", "type": "integer", "minimum": 1, "maximum": 20000 }, "density": { "default": 96, "description": "Render resolution in DPI, 1 to 2400, default 96; the natural raster size scales with density / 72, so a 64-unit-wide SVG renders 85 pixels wide at 96.", "type": "integer", "minimum": 1, "maximum": 2400 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/images/svg-badge.svg", "density": 192 } Example response: { "source_url": "https://grist.tools/samples/images/svg-badge.svg", "format": "png", "width": 171, "height": 128, "bytes": 2021, "image_base64": "iVBORw0KGgoAAAANSUhEUgAAAKsAAACACAYAAAB0g5nsAAAACXBIWXMAAB2HAAAd..." } - [video-thumbnail](https://grist.tools/docs/video-thumbnail): Grabs a single frame from a video at a chosen timestamp as a PNG or JPEG, optionally scaled to a width. (family av, category Compute, $0.003, timeout 60000ms) POST https://grist.tools/v1/video-thumbnail Why it is worth paying for: Seeking and decoding one frame needs ffmpeg AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 81 · response cap: 209715200 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of the video, at most 2048 characters; the container (MP4/MOV, MKV/WebM, AVI or Ogg) is detected from the bytes." }, "time_seconds": { "default": 0, "description": "Timestamp of the frame to grab, in seconds from the start (0 to 86400, default 0); a time past the end fails with no_frame.", "type": "number", "minimum": 0, "maximum": 86400 }, "format": { "default": "png", "description": "Encoding of the returned frame: png (default) or jpeg.", "type": "string", "enum": [ "png", "jpeg" ] }, "width": { "description": "Output width in pixels (16 to 4096); the height keeps the aspect ratio, rounded to an even number. Native size when omitted.", "type": "integer", "minimum": 16, "maximum": 4096 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/thumbnail-clip.mp4", "time_seconds": 1, "format": "png" } Example response: { "source_url": "https://grist.tools/samples/av/thumbnail-clip.mp4", "bytes": 2587, "container": "mp4", "time_seconds": 1, "format": "png", "width": 128, "height": 72, "image_bytes": 292, "image_base64": "iVBORw0KGgoAAAANSUhEUgAAAIAAAABICAIAAACx52pFAAAACXBIWXMAAAABAAAA..." } - [waveform-data](https://grist.tools/docs/waveform-data): Reduces an audio file to a fixed number of normalised waveform peaks for drawing, without shipping samples. (family av, category Compute, $0.003, timeout 90000ms) POST https://grist.tools/v1/waveform-data Why it is worth paying for: Decoding audio to peaks needs ffmpeg AND a hardened fetch: an agent that pulls the URL itself walks past every SSRF control, and carrying an ffmpeg per call is not something an agent does inline. Price: $0.003 · max_timeout_seconds: 111 · response cap: 104857600 bytes Typed errors: invalid_input (400), blocked_target (403), unreachable_target (424), upstream_timeout (424), unsupported_content_type (415), too_large (413), unprocessable (422), internal (500) Input schema: { "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { "url": { "type": "string", "maxLength": 2048, "format": "uri", "description": "Absolute http(s) URL of an audio or video file, up to 2048 characters, in a recognised container (MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF or MP3); its first audio stream is decoded to mono." }, "peaks": { "default": 200, "description": "Number of buckets to return, 1 to 4000, default 200; each is the largest absolute sample in an equal slice of the track, scaled to 0..1 against 16-bit full scale.", "type": "integer", "minimum": 1, "maximum": 4000 }, "sample_rate": { "default": 8000, "description": "Decode rate in Hz, one of 4000, 8000, 11025, 16000, 22050 or 44100, default 8000; it sets sample_count and how finely each bucket is scanned.", "type": "integer", "minimum": -9007199254740991, "maximum": 9007199254740991 } }, "required": [ "url" ], "additionalProperties": false } Example request body: { "url": "https://grist.tools/samples/av/waveform-beeps.wav", "peaks": 20, "sample_rate": 8000 } Example response: { "source_url": "https://grist.tools/samples/av/waveform-beeps.wav", "bytes": 32044, "container": "wav", "sample_rate": 8000, "sample_count": 16000, "duration_seconds": 2, "peak_count": 20, "peaks": [ 0.25, 0.25, 0.25, 0.25, 0, 0, 0, 0.5, 0.5, 0.5, 0.5, 0, 0, 0, 0.8999, 0.8999, 0.8999, 0.8999, 0, 0 ] } Endpoints published: 50.