Video Editing

How to Make Subtitles With Whisper (On Your Own Computer)

Whisper will transcribe a video on your own computer, for free, and write the SRT file for you. What it will not do is make good subtitles: its cues are too long, too fast or split in the wrong place. We ran it, measured the output against published reading speed rules, and fixed it.

By Linkeddit·October 4, 2026·11 min read

Key takeaways

  • Four on-device routes: whisper.cpp (command line, any OS), OpenAI's Python whisper command, MacWhisper (Mac app) and Subtitle Edit (free editor with Whisper built in). All four write SRT and VTT.
  • On an M1 Pro, whisper.cpp with the base.en model transcribed 24.5 seconds of speech in 0.71 seconds by its own timer.
  • The raw output was not usable as subtitles: lines of 85 to 93 characters against Netflix's 42, and two cues over 20 characters per second.
  • Splitting at 42 characters made it worse in a new way: four one-word cues on screen for 0.31 to 0.6 seconds, below Netflix's 5/6 of a second minimum.
  • Regrouping by word timestamps, fixing one misheard word and trimming two cues took our file from 5 of 6 cues flagged to 0 of 8 under Netflix's rules.

1How do you make subtitles with Whisper?

Run Whisper on the video's audio, ask for SRT or VTT output, then fix line length and reading speed before you publish. You can run it from the command line with whisper.cpp or OpenAI's Python package, or through an app such as MacWhisper or Subtitle Edit. Everything here runs on your own computer, so the footage is never uploaded. The pages ranking for this query today show how to start Whisper, not what its output looks like as subtitles; sections 5 and 6 measure that on a real run.

RouteRuns onSubtitle outputBest for
whisper.cppMac, Windows, Linux (build from source)-osrt, -ovttFast local runs, scripts, Apple silicon
openai-whisper (Python)Anything with Python and ffmpeg--output_format srt, vtt or allLine width and line count controls
MacWhisperMacSRT and VTT export on the free tierNo terminal, drag and drop
Subtitle EditWindows, Mac, Linux (v5.2.0)SRT, VTT and 300+ other formatsTranscribe and then edit timing in one app

The app facts, checked October 2026 on each vendor's own pages: MacWhisper's site lists “Export subtitles to .srt & .vtt” in its free tier and says it processes sensitive content “locally without data ever leaving your Mac”. Pro is a one-time licence that adds batch transcription and speaker recognition; the site showed us its price in euros only, so we do not quote it. Subtitle Edit's site lists “Speech to text (speech recognition) via Whisper” and says it can “read, write, and convert between more than 300 subtitle formats”, and its source code carries engines for whisper.cpp, Purfview's Faster-Whisper XXL, CTranslate2, Const-me, OpenAI's Whisper and WhisperX. Its v5.2.0 release (September 10, 2026) ships Windows, macOS and Linux builds, so it is not only a Windows route.

2How do you use Whisper on a video?

Extract the audio as 16 kHz WAV with ffmpeg, then point Whisper at it and ask for SRT. whisper.cpp needs the WAV step; OpenAI's Python command reads the video file directly, using the ffmpeg its README lists as a requirement. The whisper.cpp README (checked October 2026) says whisper-cli “currently runs only with 16-bit WAV files” and gives the ffmpeg line to convert. These are the commands we ran on October 4, 2026, on whisper.cpp 1.9.4 (built from source as the README describes, under the MIT License; Homebrew also lists a whisper.cpp formula at 1.9.4):

git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh base.en
cmake -B build
cmake --build build -j --config Release

# whisper-cli reads 16-bit WAV, so pull the audio out of the video first
ffmpeg -i talk.mp4 -ar 16000 -ac 1 -c:a pcm_s16le talk.wav
./build/bin/whisper-cli -m models/ggml-base.en.bin -f talk.wav -osrt -ovtt -of talk

Our test video is 24.5 seconds of speech made with the macOS say command from a script we wrote, so we know every word. whisper.cpp's own timer reported 0.71 seconds on an Apple M1 Pro, about 34 times faster than real time on this clip. Here is part of the SRT it wrote:

1
00:00:00,000 --> 00:00:04,000
 Welcome back to the channel. Today we are testing whether an open-source speech model

2
00:00:04,000 --> 00:00:09,160
 can write subtitles for a video on this computer, without uploading anything. The transcript

4
00:00:13,720 --> 00:00:18,240
 passed in under a second is useless to someone reading it, so we will check every cue against

Two things to notice. First, the words are close but not right: the script said “flashes past” and Whisper wrote “flashes passed”, and it turned our full stop before “So we will check” into a comma. Second, every cue runs to the end of a sentence fragment, and each line starts with a space.

3Is the Python version of Whisper better for subtitles?

It was slower on our Mac, where it ran on the CPU, but it has the subtitle controls whisper.cpp lacks: a maximum line width and a maximum line count, which together produce two-line cues. The openai/whisper README (checked October 2026) installs it with pip and asks for ffmpeg on your system. We installed openai-whisper 20250625 in a scratch Python 3.12 environment and ran:

pip install -U openai-whisper
whisper talk.mp4 --model base --language English --output_format srt \
  --word_timestamps True --max_line_width 42 --max_line_count 2

The help text says --max_line_width and --max_line_count both require --word_timestamps True. The run took 7.4 seconds of wall time, including loading the model, on the CPU (it warned that FP16 is not supported there and fell back to FP32). Note the comparison is not like for like: we used the multilingual base model here and base.en in whisper.cpp. The output:

1
00:00:00,000 --> 00:00:03,600
Welcome back to the channel, today we are
testing whether an open source speech

3
00:00:08,460 --> 00:00:12,740
the transcript is usually good, the hard
part is timing and line length, because a

Every line is now within 42 characters and every cue has two lines. But look at the breaks: “the hard / part” and “because a” at the end of a line. Netflix's general requirements say to break after punctuation or before conjunctions and prepositions, and not to separate an article from its noun. A character count cannot see any of that. It also made the same “passed” mistake and wrote “channel, today” where we said “channel. Today”.

4Which Whisper model size should you use for subtitles?

Start small. For clear English, base.en or small.en is usually enough; move up a size only when the words are wrong, because a bigger model costs memory and time and does nothing for line length or timing. The sizes below are OpenAI's, from the openai/whisper README, and the memory column is whisper.cpp's, from its README (both checked October 2026).

SizeParametersVRAM (Python)Relative speedwhisper.cpp disk / memory
tiny39 M~1 GB~10x75 MiB / ~273 MB
base74 M~1 GB~7x142 MiB / ~388 MB
small244 M~2 GB~4x466 MiB / ~852 MB
medium769 M~5 GB~2x1.5 GiB / ~2.1 GB
large1550 M~10 GB1x2.9 GiB / ~3.9 GB
turbo809 M~6 GB~8xNot listed

OpenAI says the relative speeds were measured transcribing English on an A100 and that real-world speed “may vary significantly”. It also says the English-only .en models tend to do better, especially tiny.en and base.en, and that turbo is “not trained for translation tasks”. A commenter on the r/software thread for a free Whisper subtitle GUI gave the same advice for English-only work: go for small.en.

5What does raw Whisper output get wrong as subtitles?

Line length, orphan words and reading speed. We ran each output through the logic behind our caption reading speed checker using Netflix's adult English limits, and none of them passed as written. Netflix's English style guide allows 42 characters per line and up to 20 characters per second for adult programs; its general requirements set a minimum of 5/6 of a second and a maximum of 7 seconds per subtitle (both checked October 2026). We trimmed the leading space whisper.cpp writes before counting.

OutputCuesFlagged (Netflix adult)What failed
whisper.cpp default65Lines of 85 to 93 characters; 2 cues over 20 per second
whisper.cpp -ml 42 -sow1694 one-word cues of 0.31 to 0.6 s; 5 cues over 20 per second
Python, 42 x 2 lines633 cues at 20.5 to 21.7 per second; breaks inside phrases
Ours: regrouped and proofread82Cues at 21.2 and 20.2 per second
Ours: plus two trims80None (BBC rules still flag 5)

The -ml 42 row is the trap. whisper.cpp's README calls --max-len experimental, and on our clip it split each long segment at 42 characters and left the remainder as its own cue: “ model” for 0.31 seconds, “transcript” for 0.6, “flashes” and “captions” for 0.45 each. A one-word flash is unreadable. A poster in a r/VideoEditing subtitle QC thread put the rule in one line:

“break by meaning, not by character count. Avoid orphan words.”
r/VideoEditing

Reading speed is the harder problem, because no split fixes it. By our count Whisper heard 84 words in 24.28 seconds, about 208 words a minute. The BBC's subtitle guidelines (checked October 2026) recommend 160 to 180 words per minute, while noting that viewers tend to prefer verbatim subtitles. Across the whole clip the text averages 19.8 characters per second (our arithmetic: 481 characters over 24.28 seconds), so a verbatim file sits right at Netflix's ceiling and above the BBC's.

6How do you fix Whisper subtitles?

Get word-level timestamps, regroup the words into cues of at most two lines that end at punctuation, proofread the words, then trim text in any cue that still reads too fast. Check the file again after every change. This is the order that worked on our clip, and the rule we used is ours, not a standard: close a cue at a comma or full stop once it has 30 characters, never let it pass 84, then split it into two lines with the checker's line splitter at 42.

  1. Word timestamps. whisper.cpp gives one word per cue with -ml 1 -sow. One word, “for”, came back with the same start and end time (00:00:05,480). Our SRT to VTT converter reports it as line 74, “The end time must be later than the start time”, so keep zero-length words when you regroup or the sentence loses a word.
  2. Regroup. Eight cues instead of six or sixteen, every line within 42 characters, no orphan cues. Two cues were still fast, at 21.2 and 20.2 characters per second.
  3. Proofread. “passed” became “past”. No model size setting would have told us.
  4. Trim. Cue 1 had 78 characters in 3.68 seconds; at 20 per second it can hold 73 (our arithmetic: 20 x 3.68 = 73.6). Cutting “to the channel” brought it to 17.1. In cue 7 we dropped “then” and moved a stranded “a” to the next cue, giving 18.4.
1
00:00:00,010 --> 00:00:03,690
Welcome back. Today we are testing
whether an open-source speech

5
00:00:12,100 --> 00:00:15,910
because a caption that flashes past in
under a second is useless to someone

The result passes Netflix's adult rules on all 8 cues. The BBC rules still flag 5, because 37 characters per line and about 18 characters per second (the checker's conversion of 180 words a minute) are stricter than this speaker's pace allows without rewriting. Condense further or accept a verbatim file for a fast talker; either way, run the file through the reading speed checker before you publish.

7How do I add subtitles to a video without subtitles?

Make an SRT with Whisper, then choose one of three: upload the file next to the video where the platform accepts caption files, add it to the video file as a subtitle track viewers can switch on, or burn it into the picture so it shows everywhere. A subtitle track needs no re-encode. This is the command we ran to add our cleaned file to the MP4 as a soft track:

ffmpeg -i talk.mp4 -i talk.en.srt -map 0 -map 1 -c copy \
  -c:s mov_text -metadata:s:s:0 language=eng talk-subbed.mp4
# stream 0 h264 video, stream 1 aac audio, stream 2 mov_text subtitle (eng)

Burning in needs ffmpeg's subtitles filter, which draws the text with the libass library; the documentation (checked October 2026) says FFmpeg must be configured with --enable-libass. The ffmpeg 9.0.2 build on our test Mac has no libass, and -vf subtitles=final.srt failed with “Error parsing filterchain”, so we do not show a burn-in run here. Check yours with ffmpeg -filters first, or burn the captions in from an editor. More on ffmpeg pitfalls in the ffmpeg commands AI gets wrong.

8Whisper vs Descript for transcripts and captions?

Use Whisper when you need transcripts and caption files; pay for Descript when you want to edit by editing the transcript. Both write SRT and VTT. Descript meters how much media it processes each month; Whisper on your own machine does not. A podcaster on r/podcasting with 120 untranscribed back episodes asked exactly this, and listed what mattered:

“I care about how many minutes each plan covers, if it has speaker labels and SRT export for captions.”
r/podcasting

On caption files they are level. Descript's Export subtitles help page (checked October 2026) exports .srt or .vtt with a maximum characters per line and maximum lines per card, the same controls as the Python CLI. One reply in the thread said it plainly: “SRT export is free. Whisper writes it directly, no editor involved.” On speaker labels Descript is ahead out of the box: its pricing page lists speaker detection on every plan, while another reply pointed out that “plain Whisper doesn't label speakers”, which needs a separate diarization step.

On minutes, the same pricing page (checked October 2026) lists 60 media minutes a month on Free, 600 on Hobbyist, 1,800 on Creator and 2,400 on Business. Our arithmetic for that backlog: 120 episodes of 45 to 90 minutes is 5,400 to 10,800 minutes, so 9 to 18 months of a Hobbyist allowance, or 3 to 6 months of Creator. If whisper.cpp held the 34 times real time we measured with base.en (it will be slower with larger models and longer audio), the same backlog would be roughly 2.5 to 5 hours of computer time. If you already have an editor you like, run the backlog through Whisper and judge Descript only on its text-based editing. Our Descript alternatives guide covers that side.

One more tip from that thread: keep the JSON output, with word timestamps, next to the SRT, so captions for a clip cut months later need no new transcription.

9How does Linkeddit Studio make captions with Whisper?

Studio runs Whisper base on your computer and burns the captions into the video you export. Studio does not export SRT or VTT files. We make Studio, so read this as the vendor talking. If you need a caption file for YouTube, a client or a player, use one of the four routes above; Studio is for when the captions belong in the picture.

You download the speech model once in Settings, select a clip and press Transcribe, and the captions arrive as a styled text track. There are five styles (plain, highlight, pop, karaoke and box), 1 to 12 words per line, 1 to 3 lines, and top, center or bottom placement. Studio groups captions by words, not by characters, so its burned-in captions are not checked against a characters per second limit; pick fewer words per line for a fast speaker. Claude Code or Codex, running in Studio's agent panel, can add the captions as one undo step, restyle them, and correct a misheard word in a caption, which is where the “passed” in our test would get fixed.

Captions in the picture, transcribed on your computer

Linkeddit Studio is a desktop editor for Mac and Windows where Claude Code or Codex edits your footage. On-device Whisper captions, five burned-in styles, files on your disk. One-time $99, no credits.
See Linkeddit Studio

Frequently asked questions

Can Whisper make SRT files?+

Yes. whisper.cpp writes SRT with -osrt and VTT with -ovtt, and OpenAI's Python whisper command writes txt, vtt, srt, tsv and json, all of them by default, or one with --output_format srt. We ran both on October 4, 2026 and got valid SRT files from each.

Is Whisper free to use for subtitles?+

The model and OpenAI's code are released under the MIT License, according to the openai/whisper README, and whisper.cpp is also MIT licensed, according to the LICENSE file in its repository. Running either on your own computer costs nothing beyond the machine's time. Paid apps such as MacWhisper Pro wrap the same models in a graphical interface.

Which Whisper model is best for subtitles?+

For clear English speech, start with base.en or small.en and move up only if the words are wrong. OpenAI's README lists base at 74 million parameters and about 1 GB of video memory, against 1,550 million and about 10 GB for large. A bigger model fixes misheard words; it does not fix line length or timing.

Does Whisper label speakers?+

No. Plain Whisper returns text and timestamps without speaker names. Speaker labels need a separate diarization step, and some apps bundle one: MacWhisper lists automatic speaker recognition as a Pro feature, and Descript lists speaker detection on its plans.

Can Linkeddit Studio export an SRT file?+

No. Studio transcribes speech on your device with Whisper base and burns the captions into the exported video. If you need a separate SRT or VTT file, use whisper.cpp, the Python CLI, MacWhisper or Subtitle Edit.