catalogue / tools / Audio & video / MKV to WebVTT

Convert MKV to WebVTT (subtitle track)

To convert MKV to WebVTT (subtitle track), send one POST to /v1/subtitle-extract with the MKV source and "to": "vtt" in the body. Left out, "to" falls back to its default "srt", so a WebVTT request has to send it. Each call costs $0.003 in USDC on Base via x402, with no account and no API key. The source file can be up to 200 MB.

Endpoint
POST /v1/subtitle-extract
Price
$0.003 per call

01Limits

LimitValue
Largest source file200 MB (209,715,200 bytes)
Timeout60 s
Source formatsAVI, MKV, MP4, Ogg
Output formatsSRT, WebVTT
Option streamdefault 0

02Example

A real MKV to WebVTT run of subtitle-extract.

Request
POST /v1/subtitle-extract
content-type: application/json
payment-signature: <base64 x402 payload>

{
  "url": "https://grist.tools/samples/av/pair-clip.mkv",
  "stream": 0,
  "to": "vtt"
}
Response
200 OK
payment-response: <base64 settlement receipt>

{
  "source_url": "https://grist.tools/samples/av/pair-clip.mkv",
  "bytes": 10531,
  "container": "matroska",
  "stream_index": 0,
  "format": "vtt",
  "cue_count": 2,
  "subtitle": "WEBVTT\n\n00:00.000 --> 00:01.750\nRain taps on the window.\n\n00:02.250 --> 00:03.750\nA lamp glows in the hall.\n"
}

Source file: av/pair-clip.mkv, 10.3 KB (10,531 bytes). 4 s Matroska file made by ffmpeg: a 64x36 PNG video stream at 1 fps (a slate field with a small amber square moving right), a mono 8 kHz FLAC audio track of two 500 Hz triangle beeps, and an English SubRip subtitle stream with two cues. The request sets stream to 0. The recorded run returned 108 bytes of WebVTT, 1% of the source size, with the source read as the matroska container, stream index 0 and 2 cues. Recorded with ffmpeg 8.1.2.

In the sample itself, the EBML header names DocType matroska, version 4, readable by a version 2 parser; Segment Info sets a timestamp scale of 1,000,000 ns and a duration of 4,000 ticks, 4 s, with Lavf as the muxing application; the Tracks element lists codec IDs V_MS/VFW/FOURCC (video, language und), A_FLAC (audio, language und) and S_TEXT/UTF8 (subtitle, language eng). The returned text starts with the line WEBVTT. Its 2 timing lines are 00:00.000 --> 00:01.750 and 00:02.250 --> 00:03.750, giving minutes and seconds with no hour field, and a full stop before the milliseconds. No cue has an identifier line above its timing.

03Format notes

MKV as a source

Matroska files begin with the EBML magic 1A 45 DF A3, and WebM uses the same signature, so the sniffer treats both as a single container and reports it as matroska in the container field of each response. Matroska is built from tracks that ffmpeg can address independently. When audio is the target, the first audio track is decoded and any picture is ignored. For a still frame, the requested timestamp is sought at the demuxer level before one picture is decoded, with sound excluded. A subtitle track is chosen by its index counted among subtitle tracks only, starting from zero, and an index that matches nothing is refused rather than answered with an empty file. SubRip or WebVTT text written by ffmpeg comes back together with a cue count taken from the timing arrows in that text.

WebVTT as a target

subtitle-convert writes WebVTT as the WEBVTT header, a blank line, then cues with dotted timestamps in HH:MM:SS.mmm form. WebVTT is that service's default destination. Output is rebuilt from each cue's start, end and text, so SubRip index numbers disappear and no positioning settings are added. Text subtitle tracks inside media containers take a different writer: subtitle-extract passes them through ffmpeg's webvtt muxer, and in the recorded extraction results that muxer prints timestamps without an hour component. Metadata is stripped and bit-exact flags are set on that route, so the same stream yields the same text for a pinned ffmpeg. Both routes end in the same format, while the serialiser, and therefore the exact timestamp layout, depends on whether the captions started as a file or as a track. Nothing follows the end timestamp on a timing line, because subtitle-convert writes no cue settings, and no identifier precedes a cue.

Note: The container must carry a text subtitle track; image-based subtitles cannot be turned into text.

04Errors

CodeHTTP
invalid_input400
blocked_target403
unreachable_target424
upstream_timeout424
unsupported_content_type415
too_large413
unprocessable422
internal500

When each is raised, and what it means for payment: /docs/subtitle-extract.

05Questions

What does it cost to convert MKV to WebVTT?
$0.003 per call, paid in USDC on Base via x402.
How large can the MKV file be?
Up to 200 MB (209,715,200 bytes). Past a limit the call answers too_large (413).
What happens if the conversion fails?
The tool answers with a typed JSON error from the errors table. Payment is settled only after the tool has produced its result; a call that fails inside the tool is never settled.

06Related