Audio and video conversion API without ffmpeg servers
An agent can convert and inspect audio or video through grist.tools HTTP endpoints, leaving ffmpeg execution and guarded source fetching to the service. Choose an endpoint for audio output, metadata, captions, a frame or waveform peaks, and receive its result as JSON.
Updated
01Make media work an HTTP call
Send a public source URL in url to POST https://grist.tools/v1/<name>, using a JSON body and content-type: application/json. These media operations need guarded fetching, and every one except subtitle-convert also runs ffmpeg or ffprobe. Your agent selects the operation and its declared options without running its own media server.
Read the free input schema at GET /v1/<name>/schema before constructing a request. The schemas also appear in /openapi.json; use the endpoint documentation links for the individual contracts. Start with the desired result: converting audio, choosing a track, inspecting a container, retrieving captions, grabbing a frame and drawing a waveform call for different endpoints.
02Convert audio or select an audio track
Choose audio-convert to transcode the first audio stream. Its to targets are wav, mp3, flac, aac and opus, with mp3 as the default. WAV output uses PCM, AAC output is an ADTS stream, and Opus output uses an Ogg container. Recognised source containers include MP4/MOV, Matroska/WebM, AVI, Ogg, WAV, FLAC, AIFF and MP3.
Set bitrate_kbps for the lossy targets mp3, aac and opus. AAC has a stricter upper bound than the other lossy targets, and exceeding it returns invalid_input. Set sample_rate only to a value in the schema: it accepts a fixed set of rates. When omitted, ffmpeg chooses the rate, normally the source rate.
Choose audio-extract when you need to select an audio track from a video or another recognised container. Its stream is a zero-based index among audio streams only, with the first selected by default. It offers the same to targets, also defaulting to mp3; a nonexistent track is unprocessable. The generated audio-convert example below converts the sample FLAC file to WAV and returns audio_base64 with output metadata.
03Inspect media or request a visual representation
Use media-probe for container, duration, bit rate and per-stream codec or resolution information. It reads the container index without decoding media. Its response exposes streams with fields such as type, codec, width, height, sample_rate and channels. Use that information to decide which operation and track the task needs.
Use video-thumbnail for a single frame at time_seconds, returned in image_base64. Choose png or jpeg through format; PNG is the default. Optional width scales the image while preserving aspect ratio, with height rounded to an even number. Omitting it keeps native size. A timestamp past the end fails with unprocessable, with details.reason set to no_frame.
Use waveform-data for normalised peaks suitable for drawing. It decodes the first audio stream to mono, then returns the largest absolute sample in each equal slice. Set peaks for the bucket count and sample_rate for the decode rate. The result includes sample_count, duration_seconds, peak_count and the peaks array, without shipping the decoded samples.
04Extract existing captions or convert a subtitle file
Choose subtitle-extract for an existing text subtitle track inside a media container. Supported source containers include MKV/WebM, MP4/MOV, AVI and Ogg. Its zero-based stream counts subtitle streams only; a missing selection fails with unprocessable, with details.reason set to no_subtitle_stream. Set to to srt or vtt; extraction defaults to srt and returns caption text in subtitle.
Choose subtitle-convert when the source URL already points to a SubRip or WebVTT file. Detection uses the content rather than the extension, and to defaults to vtt. Conversion re-serialises cues: SRT output is renumbered, while VTT cue identifiers, settings and NOTE blocks are dropped. Preserving those extras is outside this conversion contract. The table below compares the services, with their prices and declared limits.
05Services covered
| Endpoint | Summary | Price | Timeout | max_timeout_seconds | Size cap | Max duration |
|---|---|---|---|---|---|---|
| audio-convert | Transcodes an audio file between WAV, MP3, FLAC, AAC and Opus, with an optional bit rate and sample rate. | $0.003 | 94 s | 115 | 100 MB (104,857,600 bytes) | 60 min (3,600 s) |
| audio-extract | Extracts the audio track from a video (or any container) and returns it in a chosen audio format. | $0.003 | 94 s | 115 | 200 MB (209,715,200 bytes) | 60 min (3,600 s) |
| media-probe | Reads a media file's container, duration, bit rate and per-stream codec/resolution without decoding it. | $0.003 | 30 s | 51 | 100 MB (104,857,600 bytes) | - |
| subtitle-convert | Converts a subtitle track between SubRip (.srt) and WebVTT (.vtt), normalising the cues. | $0.003 | 15 s | 36 | 5 MB (5,242,880 bytes) | - |
| subtitle-extract | Extracts a text subtitle track from a media container (MKV, MP4, WebM) as SubRip or WebVTT. | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - |
| video-thumbnail | Grabs a single frame from a video at a chosen timestamp as a PNG or JPEG, optionally scaled to a width. | $0.003 | 60 s | 81 | 200 MB (209,715,200 bytes) | - |
| waveform-data | Reduces an audio file to a fixed number of normalised waveform peaks for drawing, without shipping samples. | $0.003 | 90 s | 111 | 100 MB (104,857,600 bytes) | - |
06Limits
- Largest response cap: 200 MB (209,715,200 bytes) on audio-extract.
- Smallest response cap: 5 MB (5,242,880 bytes) on subtitle-convert.
- Longest timeout: 94 s on audio-convert.
- Shortest timeout: 15 s on subtitle-convert.
- max_timeout_seconds: from 36 to 115.
- Max input duration on audio-convert: 60 min (3,600 s).
- Max input duration on audio-extract: 60 min (3,600 s).
07End-to-end example
A real recorded example of audio-convert, the same one its docs page publishes. Strings in the response longer than 400 characters are cut and end with ....
POST /v1/audio-convert
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/av/convert-beeps.flac",
"to": "wav"
}200 OK
payment-response: <base64 settlement receipt>
{
"source_url": "https://grist.tools/samples/av/convert-beeps.flac",
"bytes": 32086,
"source_container": "flac",
"format": "wav",
"bitrate_kbps": null,
"sample_rate": null,
"audio_bytes": 32044,
"audio_base64": "UklGRiR9AABXQVZFZm10IBAAAAABAAEAQB8AAIA+AAACABAAZGF0YQB9AAAA4ADo..."
}08Chain operations around the source URL
For a media inspection pipeline, start with media-probe, then pass the source URL to the operation the result calls for. Select audio with audio-extract, captions with subtitle-extract, or a frame with video-thumbnail. Match audio and subtitle selections within their respective stream types; do not copy an overall probe stream index into a type-specific selection without checking it.
For first-track waveform data, send the original audio or video URL directly to waveform-data. For audio extraction, choose the required output format in audio-extract itself. Add audio-convert when its bit rate or sample rate controls are needed, rather than treating conversion as a mandatory step after every extraction.
These inputs take URLs. If a later call must process returned audio_base64 or subtitle content, your pipeline needs to make that content available at a public URL before passing it as the next url. Keep returned media and text distinct from source URLs. For captions, requesting the required format during extraction can remove the need for a subsequent subtitle-convert call.
09Check input caps before planning the chain
audio-convert accepts source bytes up to 100 MB and audio duration up to 60 min (3,600 s). For audio-extract, the size cap is 200 MB and the selected audio stream must fit 60 min (3,600 s). Longer input fails with too_large; an input whose duration cannot be measured fails with unprocessable. Those duration decisions happen before encoding, and neither refusal is settled.
Check each stage separately. In particular, media-probe has a 100 MB cap, so a source eligible for extraction may exceed the probe allowance. The Size cap column of the table lists the other caps, including those for caption files, thumbnails and waveform input. Do not transfer the audio conversion duration cap to an endpoint that does not declare it.
Keep the requested operation within its contract. Probing reads metadata without decoding; a thumbnail request selects a single frame; waveform output contains peaks; subtitle extraction selects existing text captions. These results serve different purposes. Changing thumbnail width or waveform bucket count does not change the source file that the endpoint must fetch under its byte cap.
10Budget successful calls and the whole response deadline
A probe followed by audio conversion costs $0.003 plus $0.003 when both calls succeed. Selecting a track instead costs $0.003 plus $0.003. A thumbnail adds $0.003, waveform data adds $0.003, and caption extraction adds $0.003. Converting a separate caption file costs $0.003. Budget the operations your pipeline actually selects.
No account or API key is required. A call without payment-signature returns HTTP 402 with the price in amount and the whole-call deadline in maxTimeoutSeconds, which the free schema publishes as max_timeout_seconds. Payment uses x402 in USDC on Base. For audio-convert, the handler timeout is 94 s, while max_timeout_seconds is 115 seconds. Use the whole-call allowance for client deadlines because it includes payment verification and settlement around the handler.
Payment is verified before execution and settled only after the handler succeeds. A gate refusal or handler failure is never settled. If a settlement call fails without answering or its answer fails validation (that case is HTTP 500 internal, with a message starting "Payment settlement returned"), its outcome cannot be stated in advance; the error message identifies that case. Preserve the message and account for that uncertainty before repeating a paid operation.
11Choose retries from the error meaning
invalid_input means the body fails the schema, so correct fields or options before retrying. blocked_target means the network guard refuses the destination: check the address, HTTP or HTTPS scheme, permitted web port and absence of URL credentials. Repeating an unchanged request does not correct these causes.
unreachable_target means resolution failed, or the connection failed or closed before the body arrived. upstream_timeout means it did not answer before the endpoint timeout. Check source availability and consider a later retry when that condition may have changed. unsupported_content_type requires a source type the endpoint handles; unprocessable means the source answered with an HTTP error status, or its bytes arrived but could not be processed within the endpoint deadline.
too_large can indicate the streaming size cap, decompression ratio or a declared media duration cap. Correct the source or endpoint choice before resubmitting. For missing tracks, check the relevant stream selection; when unprocessable carries details.reason no_frame, choose a timestamp within the video. Preserve the error detail so the agent can distinguish these cases.
rate_limited means excessive calls from the payer or toward the destination domain; reduce request pressure. For retryable HTTP 503 errors such as concurrency_limit, registry_at_capacity and upstream_unavailable, wait for the seconds in retry-after. When upstream_unavailable comes from the front proxy, the call may already have run and settled; check authorizationState before paying again. internal identifies a service fault; inspect the message, including any settlement uncertainty, before retrying.
12Questions
- Which endpoint converts audio from a video?
audio-convertconverts the first audio stream in a recognised container to the selected audio format. Useaudio-extractwhen you need to choose a track withstream. That index counts audio streams only, and extraction can return the chosen output format directly.- Can I control the output bit rate and sample rate?
audio-convertexposesbitrate_kbpsformp3,aacandopus, with a stricter upper bound for AAC. Itssample_ratemust come from the allowed set in the free schema. Omitting the sample rate lets ffmpeg choose it, normally from the source.- Does media inspection decode the audio or video?
media-probereads the container index without decoding media. It returns container, duration, bit rate and stream information, including codecs and applicable resolution fields. Choosevideo-thumbnailfor a decoded frame orwaveform-datafor peaks derived from decoded audio.- Can I preserve every WebVTT detail during conversion?
subtitle-convertre-serialises cues and drops VTT cue identifiers, settings and NOTE blocks. SRT output is renumbered. Choose this endpoint for normalised caption conversion when those extra details do not need to survive.- What happens when audio exceeds the duration cap?
audio-convertcaps audio duration at 60 min (3,600 s), andaudio-extractcaps the selected audio stream at 60 min (3,600 s). Longer input returnstoo_large, while unmeasurable duration returnsunprocessable. Both decisions happen before encoding, and the call is never settled.- Can I pass an extraction response directly into another call?
- The media endpoints take a public source URL in
url. Audio extraction returnsaudio_base64, and subtitle extraction returns text insubtitle. To process that output in another URL-based call, make the resulting content available at a public URL first. - Should my agent retry every failed media request?
- Inspect the typed error before retrying: invalid options, missing tracks and oversized inputs need a corrected request or source. For retryable HTTP 503 errors, wait for
retry-after. If a settlement call failed without answering or its answer failed validation (that case is HTTP 500internal, with a message starting "Payment settlement returned"), retain the stated uncertainty when deciding whether to repeat the operation.
13Related pages
- Audio Convert API docs
- Audio Extract API docs
- Media Probe API docs
- Subtitle Convert API docs
- Subtitle Extract API docs
- Video Thumbnail API docs
- Waveform Data API docs
- Audio Convert API
- Audio Extract API
- Media Probe API
- Subtitle Convert API
- Subtitle Extract API
- Video Thumbnail API
- Waveform Data API
- Audio & video
- API documentation
- Guides
- How AI agents pay for APIs with x402
- Turning web pages into LLM-ready text
- Parsing PDF, DOCX, XLSX and EPUB for agents