Video Editing
How Fast Is Whisper on a Mac? We Timed 47 Runs on an M1 Pro
On an M1 Pro, whisper.cpp with Metal turned 30 minutes of real speech into text in 43 seconds with the base model. The same model size running in WebAssembly on one CPU thread took 11 minutes. Here are all 47 runs, the accuracy of each model, and the method to rerun it on your own Mac.
Key takeaways
- Every engine and model we tested ran faster than real time on an M1 Pro with 16 GB, from 2.6x to 56.1x.
- whisper.cpp base with Metal: 16.83 s for 10 minutes of speech, 42.72 s for 30 minutes.
- Metal is the biggest lever: the same base model on 4 CPU threads took 122.84 s for 10 minutes, about 7.3 times longer.
- Bigger models cost time steadily: on 10 minutes, word error rate fell from 4.25% (base) to 2.56% (small) for about 1.5 times the wall time.
- The engine matters as much as the model. Whisper base in WebAssembly on one thread, the engine Linkeddit Studio uses, took 237.99 s for 10 minutes.
1How fast is Whisper on a Mac?
Fast, if the app uses the GPU. On our M1 Pro, whisper.cpp with the base model and Metal transcribed 10 minutes of speech in 16.83 seconds (36.4 times real time) and 30 minutes in 42.72 seconds. The slowest setup we tested, Whisper base in WebAssembly on one CPU thread, needed about 4 minutes for the same 10 minutes. Every figure here is a median of runs we did on October 5, 2026, on one machine, with real read speech and a known transcript.
The pages that rank for this question mostly come from dictation app makers. They give per-chip tables built on 10 to 60 second clips, and one says outright that its numbers are “illustrative ranges drawn from public whisper.cpp benchmarks plus internal JustVoice testing” (checked October 2026). We publish every run instead, on 1, 10 and 30 minute files, with measured accuracy. We make Linkeddit Studio, and one engine we tested is the runtime and model Studio uses for on-device captions, so section 6 includes our own result, which is the slowest one here.
2What did each Whisper model and backend measure?
Seventeen cells, three runs each except two 30 minute cells that ran once. Wall time is the median run; real-time factor (RTF) is seconds of audio per second of work, so higher is faster. Word error rate (WER) was identical across runs of a cell, because decoding is deterministic.
| Engine and model | Backend | Clip | Runs | Median wall time | Median RTF | WER |
|---|---|---|---|---|---|---|
| whisper.cpp tiny | Metal | 1 min | 3 | 1.68 s | 41.6x | 8.33% |
| whisper.cpp tiny | Metal | 10 min | 3 | 10.93 s | 56.1x | 5.75% |
| whisper.cpp tiny | Metal | 30 min | 3 | 34.84 s | 51.8x | 5.38% |
| whisper.cpp base | Metal | 1 min | 3 | 2.41 s | 29.1x | 6.55% |
| whisper.cpp base | Metal | 10 min | 3 | 16.83 s | 36.4x | 4.25% |
| whisper.cpp base | Metal | 30 min | 3 | 42.72 s | 42.2x | 3.53% |
| whisper.cpp base | CPU, 4 threads | 1 min | 3 | 14.42 s | 4.9x | 5.95% |
| whisper.cpp base | CPU, 4 threads | 10 min | 3 | 122.84 s | 5.0x | 3.94% |
| whisper.cpp base | CPU, 4 threads | 30 min | 1 | 300.54 s | 6.0x | 3.29% |
| whisper.cpp small | Metal | 1 min | 3 | 3.53 s | 19.8x | 2.38% |
| whisper.cpp small | Metal | 10 min | 3 | 25.80 s | 23.8x | 2.56% |
| whisper.cpp small | Metal | 30 min | 3 | 91.39 s | 19.7x | 2.26% |
| whisper.cpp large-v3-turbo | Metal | 1 min | 3 | 6.11 s | 11.5x | 2.38% |
| whisper.cpp large-v3-turbo | Metal | 10 min | 3 | 38.81 s | 15.8x | 1.81% |
| transformers.js base q8 | WebAssembly, 1 thread | 1 min | 3 | 26.15 s | 2.7x | 5.95% |
| transformers.js base q8 | WebAssembly, 1 thread | 10 min | 3 | 237.99 s | 2.6x | 5.62% |
| transformers.js base q8 | WebAssembly, 1 thread | 30 min | 1 | 657.12 s | 2.7x | 5.31% |
The transformers.js rows are Linkeddit Studio's on-device runtime and model (transformers.js 4.3.0, whisper-base q8, WebAssembly, one thread), measured outside the app in Node. The whisper.cpp rows use the official unquantized ggml weights.
3What does real-time factor mean, and why does 8x realtime confuse people?
We define real-time factor as audio length divided by processing time. 36x means 36 seconds of speech per second of work, so a 10 minute file takes about 17 seconds. People use the phrase in both directions. A Hacker News commenter wrote that on an M1 Mac, with the large model:
“8 minutes of processing for every 1 minute of audio.”
They called that “8x realtime”, the opposite of how most benchmarks use it. Another commenter reported that “whisper-cpp can do about 6x-10x realtime” on a 2021 M1 MacBook, meaning faster than real time. Before you compare two numbers, check which way each one runs and what hardware produced it. As a reply on r/LocalLLaMA put it, after a speed claim for a different model:
“What hardware are you using?”
4Which Whisper model should you use on Apple silicon?
Small is the sweet spot for captions on this machine. On the 10 minute file with Metal it made 2.56% word errors at 23.8x real time, against 4.25% at 36.4x for base. Large-v3-turbo was the most accurate at 1.81% and still ran 15.8 times faster than real time.
| Model (whisper.cpp, Metal) | 10 min wall time | RTF | WER |
|---|---|---|---|
| tiny | 10.93 s | 56.1x | 5.75% |
| base | 16.83 s | 36.4x | 4.25% |
| small | 25.80 s | 23.8x | 2.56% |
| large-v3-turbo | 38.81 s | 15.8x | 1.81% |
Going from base to small cut errors by about 40 percent (4.25% to 2.56%) for about 1.5 times the wall time (16.83 s to 25.80 s). For captions, every wrong word is one a viewer reads, so on a Mac with a GPU the extra seconds are usually worth it. Tiny is fast, but its 8.33% WER on the 1 minute clip is roughly one wrong word in twelve.
Our WER reads higher than published LibriSpeech scores for a known reason. The reference transcripts use the books' British spellings (“counselled”, “ardour”) and Whisper writes American ones, and we did not normalize spelling or numbers. Compare these models with each other, not with paper figures. For choosing a size for subtitles specifically, our guide to making subtitles with Whisper covers the workflow.
5Does Metal make Whisper faster on Apple silicon?
Yes, by far more than most pages say. Same base model, same file, same machine: 10 minutes took 122.84 seconds on 4 CPU threads and 16.83 seconds with Metal, about 7.3 times faster. On 30 minutes the CPU run took 300.54 s (one run) against 42.72 s with Metal.
The dictation app page quoted above says enabling Metal “yields a 30 to 60% speedup” over CPU-only execution, with an M2 Pro table it calls representative (checked October 2026). Its setup differs from ours: quantized weights, whisper.cpp around 1.6, and 8 performance cores in its CPU column. Our CPU rows used whisper.cpp's default of 4 threads on a busy laptop. We did not test 8 threads or an idle machine, so we cannot say how far they would narrow the gap. Engine choice has mattered on Macs for years; a Hacker News commenter who had been running Whisper on a GTX 1070 wrote in 2022 that it:
“was terribly slow on M1 Mac. Whisper.cpp has comparable performance to the 1070 while running on M1 CPU.”
Accuracy did not move in one direction. CPU-only base scored 3.94% on 10 minutes against 4.25% with Metal. Different numeric paths give slightly different text; read that as noise between backends, not as the CPU being more accurate.
6Why do some Mac apps run Whisper slower than others?
Because they run it on a different engine. A native app can use whisper.cpp with Metal. An app built on web technology can run Whisper through WebAssembly on the CPU, as Linkeddit Studio does, on one thread. On our M1 Pro that made the same model size about 14 times slower.
This is the engine Linkeddit Studio uses for its on-device captions, and we make Studio, so here is our own number plainly. Studio runs whisper-base, quantized to 8 bits, through transformers.js in a WebAssembly worker on one CPU thread, with no GPU path. Measured outside the app with the same runtime and model files, it took 237.99 seconds for 10 minutes of speech (2.6x real time) and 657.12 seconds, about 11 minutes, for 30 minutes (one run). That is about 14 times slower than whisper.cpp base with Metal (16.83 s) and about 1.9 times slower than whisper.cpp base on 4 CPU threads (122.84 s).
For a user, that means a 10 minute clip takes about 4 minutes to caption on this Mac, still faster than playing it back, and the audio never leaves the computer. Accuracy was a little lower than whisper.cpp base: 5.62% against 4.25% on 10 minutes. Quantization and 30 second chunked decoding both play a part, and this test does not separate them. Electron and Node share V8 but are not identical, so treat the in-app time as approximately equal.
7Why do short benchmark clips understate Whisper's speed?
Because startup is a fixed cost. whisper.cpp base with Metal measured 29.1x on the 1 minute clip, 36.4x on 10 minutes and 42.2x on 30 minutes, the same engine and model. Our wall times include process start and model load, as a real transcription does, and on a 70 second file that overhead is a large share of the total.
The first run also paid a warm-up. On the 1 minute clip, the first base run took 3.92 s and the next two 2.41 s and 2.24 s; tiny went 2.91 s, then 1.68 s and 1.62 s. Once warm, long runs were tight: the three base Metal runs on 10 minutes spanned 16.79 to 16.91 s. A benchmark built on one short clip measures startup as much as transcription.
8How long will my recording take to transcribe?
Divide the length by the real-time factor. On this M1 Pro, an hour would take about 85 seconds with whisper.cpp base and Metal, and about 22 minutes in WebAssembly on one thread. The hour column is a projection from our measured RTF; we did not time a full hour.
| Setup | 10 min (measured) | 30 min (measured) | 1 hour (projected) |
|---|---|---|---|
| whisper.cpp tiny, Metal | 10.93 s | 34.84 s | about 70 s |
| whisper.cpp base, Metal | 16.83 s | 42.72 s | about 85 s |
| whisper.cpp small, Metal | 25.80 s | 91.39 s | about 3 min |
| whisper.cpp large-v3-turbo, Metal | 38.81 s | not run | about 4 min |
| whisper.cpp base, CPU | 122.84 s | 300.54 s | about 10 min |
| transformers.js base q8, WebAssembly | 237.99 s | 657.12 s | about 22 min |
Turbo's hour uses its 10 minute RTF, since we did not run it on 30 minutes. Transcription is only the first step for video. If you made subtitle files with whisper.cpp, check their pacing in the caption reading speed checker, and convert them for web players with the SRT to VTT converter. To cut the ums the transcript catches, see how to remove filler words without choppy cuts.
9How did we test, and how can you reproduce it?
One working MacBook, two engines, six model and backend setups, three clip lengths of real read speech with exact reference transcripts, and wall-clock timing of the whole process.
- Machine. Apple M1 Pro, 10-core CPU (8 performance, 2 efficiency), 16-core GPU, 16 GB memory, macOS 27.0 (26A428). pmset recorded no thermal warning.
- whisper.cpp. whisper.cpp 1.9.4 from Homebrew (ggml 0.25.3), 4 threads, flash attention and Metal on by default, with the official ggml models tiny, base, small and large-v3-turbo.
- WebAssembly. transformers.js 4.3.0 browser build with onnxruntime-web 1.31.0-dev, the onnx-community whisper-base q8 encoder and decoder, one thread, 30 second chunks with a 5 second stride, under Node 23.4.0. Model load took 1.09 to 1.36 s per run.
- Audio. LibriSpeech test-clean (CC BY 4.0, read from public-domain LibriVox audiobooks), utterances joined in id order into 16 kHz mono WAV clips of 70.00 s, 613.06 s and 1,804.64 s, with 168, 1,601 and 4,594 reference words.
- WER. Lowercase, hyphens to spaces, strip all but letters, digits and in-word apostrophes, then word-level edit distance divided by reference words. No spelling or number normalization.
brew install whisper-cpp # whisper.cpp 1.9.4, ggml 0.25.3 # Metal (the default) whisper-cli -m ggml-base.bin -f 10min.wav -l en -otxt -of out -np # CPU only, same model file whisper-cli -m ggml-base.bin -f 10min.wav -l en -otxt -of out -np -ng
The limits, stated plainly. This is one machine on one day, not a lab: Chrome, VS Code and other work were running, and the 1 minute load average at the start of runs ranged from about 6 to 23, so an idle Mac would likely be somewhat faster, most of all on the CPU rows. whisper.cpp wall times include model load, which we did not capture separately. The two slowest 30 minute cells ran once, not three times, and large-v3-turbo ran on the 1 and 10 minute clips only. Read speech is clean; noisy or overlapping voices will raise WER. We did not test Core ML, quantized ggml models, MLX, WhisperKit, faster-whisper, medium or large-v3. Newer chips with more GPU cores should be faster, so run the commands on your own Mac and compare real-time factors.
Captions on your own computer
Frequently asked questions
Is Whisper faster than real time on a Mac?+
On our M1 Pro, every engine we tested was, from 2.6x real time (transformers.js in WebAssembly on one thread, 10 minute file) to 56.1x (whisper.cpp tiny with Metal, 10 minute file).
How long does Whisper take to transcribe an hour of audio on a Mac?+
We did not time an hour. Projected from our 30 minute runs on the M1 Pro, whisper.cpp base with Metal would need about 85 seconds and WebAssembly on one thread about 22 minutes.
Does whisper.cpp use the GPU on Apple silicon?+
Yes. The Homebrew build uses Metal by default and -ng turns it off. With base on our 10 minute file, Metal took 16.83 s and the CPU on 4 threads took 122.84 s.
Which Whisper model is best on a Mac?+
On our 10 minute file with Metal: base for speed (4.25% word error rate at 36.4x), small for most captioning (2.56% at 23.8x), and large-v3-turbo when every word matters (1.81% at 15.8x).
Why is Whisper slow in some Mac apps?+
The engine, more than the model. Apps built on web technology can run Whisper in WebAssembly on the CPU, sometimes on one thread. The same base size ran about 14 times slower that way than in whisper.cpp with Metal.
Can I reproduce this benchmark?+
Yes. Install whisper.cpp 1.9.4 with Homebrew, download the official ggml models, build clips from LibriSpeech test-clean, and run the commands in the method section. Compare your real-time factors with ours.