Convert Ogg to WebVTT (subtitle track)
To convert Ogg to WebVTT (subtitle track), send one POST to /v1/subtitle-extract with the Ogg source and "to": "vtt" in the body. Left out, "to" falls back to its default "srt", so a WebVTT request has to send it. Each call costs $0.003 in USDC on Base via x402, with no account and no API key. The source file can be up to 200 MB.
- Endpoint
POST /v1/subtitle-extract- Price
- $0.003 per call
- Service
- subtitle-extract
01Limits
| Limit | Value |
|---|---|
| Largest source file | 200 MB (209,715,200 bytes) |
| Timeout | 60 s |
| Source formats | AVI, MKV, MP4, Ogg |
| Output formats | SRT, WebVTT |
| Option stream | default 0 |
02Example
A real Ogg to WebVTT run of subtitle-extract.
POST /v1/subtitle-extract
content-type: application/json
payment-signature: <base64 x402 payload>
{
"url": "https://grist.tools/samples/av/pair-clip.ogg",
"stream": 0,
"to": "vtt"
}200 OK
payment-response: <base64 settlement receipt>
{
"source_url": "https://grist.tools/samples/av/pair-clip.ogg",
"bytes": 12553,
"container": "ogg",
"stream_index": 0,
"format": "vtt",
"cue_count": 2,
"subtitle": "WEBVTT\n\n00:00.000 --> 00:01.750\nRain taps on the window.\n\n00:02.250 --> 00:03.750\nA lamp glows in the hall.\n"
}Source file: av/pair-clip.ogg, 12.3 KB (12,553 bytes). 4 s Ogg file: a 64x36 Theora video stream at 1 fps (a slate field with a small amber square moving right) and a mono 8 kHz FLAC audio track of two 500 Hz triangle beeps, both muxed by ffmpeg, plus a two-cue OGM text subtitle stream added by the generator. The request sets stream to 0. The recorded run returned 108 bytes of WebVTT, 1% of the source size, with the source read as the ogg container, stream index 0 and 2 cues. Recorded with ffmpeg 8.1.2.
In the sample itself, 16 Ogg pages carry 3 logical streams: Theora 3.2.1 video of 64 by 36 pixels at 1 frame per second (serial 0), FLAC 1.0 mapping audio at 8,000 Hz with 1 channel (serial 1) and an OGM text stream (serial 2). The end-of-stream pages hold granule positions 2,250 for serial 2, 67 for serial 0 and 32,000 for serial 1. The returned text starts with the line WEBVTT. Its 2 timing lines are 00:00.000 --> 00:01.750 and 00:02.250 --> 00:03.750, giving minutes and seconds with no hour field, and a full stop before the milliseconds. No cue has an identifier line above its timing.
03Format notes
Ogg as a source
An Ogg file opens with the capture pattern OggS, and the container can carry Vorbis, Opus or Theora streams. Grist sniffs that pattern and hands Ogg to three of the conversion services on these pages. Audio conversion decodes the first audio stream, leaving any Theora picture aside. A thumbnail request seeks before opening the input and decodes a single frame from the first video stream, discarding sound. Caption extraction maps a subtitle stream counted from zero among the subtitle streams and serialises it through ffmpeg. Because audio-convert writes its Opus output inside Ogg, a file produced by that service comes back as an Ogg source, not as a separate Opus format. Theora is a lossy video codec, so a frame taken from it carries compression noise into the thumbnail, unlike a clip whose video stream is itself stored losslessly.
WebVTT as a target
subtitle-convert writes WebVTT as the WEBVTT header, a blank line, then cues with dotted timestamps in HH:MM:SS.mmm form. WebVTT is that service's default destination. Output is rebuilt from each cue's start, end and text, so SubRip index numbers disappear and no positioning settings are added. Text subtitle tracks inside media containers take a different writer: subtitle-extract passes them through ffmpeg's webvtt muxer, and in the recorded extraction results that muxer prints timestamps without an hour component. Metadata is stripped and bit-exact flags are set on that route, so the same stream yields the same text for a pinned ffmpeg. Both routes end in the same format, while the serialiser, and therefore the exact timestamp layout, depends on whether the captions started as a file or as a track. Nothing follows the end timestamp on a timing line, because subtitle-convert writes no cue settings, and no identifier precedes a cue.
Note: The container must carry a text subtitle track; image-based subtitles cannot be turned into text. An audio-only Ogg (Vorbis or Opus) has no subtitle track and fails with unprocessable (no_subtitle_stream).
04Errors
| Code | HTTP |
|---|---|
invalid_input | 400 |
blocked_target | 403 |
unreachable_target | 424 |
upstream_timeout | 424 |
unsupported_content_type | 415 |
too_large | 413 |
unprocessable | 422 |
internal | 500 |
When each is raised, and what it means for payment: /docs/subtitle-extract.
05Questions
- What does it cost to convert Ogg to WebVTT?
- $0.003 per call, paid in USDC on Base via x402.
- How large can the Ogg file be?
- Up to 200 MB (209,715,200 bytes). Past a limit the call answers too_large (413).
- What happens if the conversion fails?
- The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.