All articles
Captions

SRT vs VTT: Choosing The Right Caption Format

A practical comparison of SRT and VTT subtitles, when each one is the safer choice, and how to avoid the formatting mistakes that break players.

August 10, 20265 min read

Both formats carry the same essential payload: timed blocks of text. The difference is what else they can carry and which players accept them without complaint.

Pick SRT for maximum compatibility

SRT is the lowest common denominator and that is its strength. Practically every editor, social platform and TV-adjacent pipeline accepts it. If your caption file is going somewhere you do not control, send SRT.

Pick VTT for the web

WebVTT is the format the HTML video element expects. It supports styling cues, positioning, regions and metadata tracks, which matters when captions must sit clear of on-screen graphics or when you want chapter markers alongside them.

Formatting rules that prevent rejected files

  • Two lines per cue, around 42 characters per line — longer cues get truncated by players.
  • Minimum cue duration of roughly one second so captions do not flash.
  • Never let a cue overlap the next one; most strict parsers drop the whole file.
  • Keep speaker labels consistent, and prefix them the same way in every cue.

Export both. They are generated from the same transcript, cost nothing extra, and save a round trip when a platform turns out to be picky.

Get SRT And VTT From One Upload

Transcribe your video once and export captions in both formats, plus TXT, PDF, Markdown, CSV and JSON.

Caption A Video Free

Transcribe Your Next Recording Free

Speaker labels, timestamps and all 15 export formats on the free plan.

Transcribe Free