Video Editing

How to Remove Filler Words From Video Without Choppy Cuts

Finding the ums is easy. Cutting them without taking a real word, leaving half a filler behind, or turning a talking head into a string of jumps is the hard part. Here is what each method detects, where cuts go wrong, and a script we ran to check it.

By Linkeddit·October 5, 2026·10 min read

Key takeaways

  • Four routes: by hand, a transcript editor (Descript, Premiere, CapCut), a Whisper and ffmpeg script, or a transcript cleanup like Linkeddit Studio's. All depend on the transcript and its word timing.
  • Um and uh are safe to cut. Like and you know are often real words, so take them only when commas or pauses set them off.
  • In our test, Whisper's end times for the fillers came back 0.17 to 0.19 s early, and cutting at them left every filler audible. Cutting from the filler's start to the next word's start removed all four.
  • On video, every removed filler is a jump cut. Plan the cover (punch-in, B-roll, second angle) or leave more fillers in.
  • A detector only removes what the transcript wrote down. In our test Whisper wrote uh as ah.

1How do you remove filler words from a video?

Transcribe with word timings, mark the fillers, review the list, cut the same stretch from picture and sound, and listen to every join. The pages that rank for this question are mostly tool landing pages with three steps: upload, detect, export. None says which words the detector matches, why a cut can leave half an um behind, or what to do about the picture jumping. This guide covers those.

MethodHow fillers are foundBest for
By handYour ears and the waveformShort videos, speakers who roll ums into words
Transcript editorThe tool's speech model and filler listWork already inside that editor
Whisper plus ffmpegYour own word list over Whisper's word timesFree, scriptable, batch jobs
Linkeddit StudioA fixed rule set over the transcriptTalking heads edited by their words

Pauses are a separate job. If dead air is the bigger problem, see how to remove silence from video, which measures what each silence setting cuts.

2Which words actually count as filler words?

Um and uh are filler in almost every use. Like, you know, so and actually are filler only sometimes, and every tool draws that line differently. The sociolinguist Valerie Fridland made the same split in a Stanford GSB podcast interview (checked October 2026), calling um and uh “filled pauses” and saying like and you know “don't actually pattern in the same way”. What each tool names, on its own pages:

ToolWords it names (checked October 2026)
Descriptum and uh (Help Center); its tool page FAQ adds you know, like, so, actually
Adobe Premiere"filler words like uh and umm"
CapCut desktopum, uh, like, you know, as examples
DaVinci ResolveNo filler detector in the Resolve 20 New Features Guide
Linkeddit Studioum, umm, uh, uhh, uhm, er, erm, hmm, mm; like and you know when set off; the first of a doubled word

3Will an automatic remover cut real words?

Yes, if it removes discourse words everywhere. “I like it”, “so what?” and “do you know him?” all contain a word from some filler list. The safer rule takes like and you know only when set off. Studio counts a comma after, or punctuation or a pause on both sides; our script uses a simpler version (a comma after and punctuation before, no pause check). “It was, like, great” loses its like; “I like it” keeps it.

Repeats have the same problem. “I I think” is a stumble, but “had had” can be correct English, and a rule that removes the first of any doubled word takes both. Read the proposed cuts before you apply them, and search for so and actually yourself.

4Why does it sound choppy after removing filler words?

Usually because the cut misses where the filler starts and ends. Transcript word times are estimates, and an early end time leaves the tail of the um in. We measured it. The macOS say voice read “So, um, I think the the plan is, uh, to ship it next week. Umm, we tested it, like, twice. I like it.” into an 8.4 second clip. We transcribed it with Whisper base.en and word timestamps, then compared Whisper's times with the real speech boundaries from ffmpeg's silencedetect:

# Whisper (no prompt)   the audio itself
#  0.62-0.78  um,      um    0.66-0.95
#  2.58-2.76  ah,      uh    2.63-2.95
#  4.62-4.78  um,      um    4.67-4.97
#  6.06-6.34  like,    like  6.16-6.52

Every filler's end came back 0.17 to 0.19 seconds early. Cut exactly at those times, the result still transcribed as “So, um, I think that the plan is, uh, to ship it”, with all four fillers audible. One synthetic voice and one model, so treat the numbers as an illustration, but people report the same with commercial tools. Under a popular Descript tutorial, a YouTube commenter said the tool:

“sometimes it leaves part of the word or the sound of my voice from saying the word.”
YouTube comment

The fix that worked for us matches what one r/podcasting commenter does:

“I cut from the beginning of the UM to the very start of the next word or phrase.”
r/podcasting

Cut that way, the pause before each filler stayed as a breath and all four fillers were gone. Two more rules help. Skip fillers that run straight into the next word; an editor in the same thread has a “soft rule” not to remove filler that rolls into the following words. And add a very short fade, 20 ms in our script, at each join so it does not click. Descript's Avoid harsh cuts option skips fillers it cannot remove without clipping nearby words.

5How do you hide the jump cuts?

Change what the viewer sees at the cut: punch in, cut to B-roll, switch angle, or do not cut there. Audio-first tools treat a filler as sound. As a reply in an r/davinciresolve thread about cutting ums put it:

“The person is still saying the words visually.”
r/davinciresolve

On one camera, each removed um is a jump. Under the same Descript tutorial, a commenter wrote that “the video portion would look like I'm having a seizure”. The covers, cheapest first:

  • Punch-in. Scale the clip after the cut (we start at 10 to 20 percent), then back out at a later cut. Best when you shot above your publish resolution.
  • B-roll or a screen recording over the join, so the sound cuts and the picture does not.
  • A second angle. A video podcast editor on r/podcasting says most of their cuts can be masked with angle changes or smoothing effects.

Tightening pauses as well adds more cuts; the silence guide counts the jump cuts a 10 minute talking head gets at each setting.

6Should you remove every filler word?

No. For podcasts and interviews, remove the fillers that break a sentence and keep the rest. For Shorts and ads, cut harder: pace matters more, and captions and reframing hide the cuts. That is our rule of thumb, and podcast editors agree. In an r/podcasting thread on sounding natural, one editor who removes only the easy ums explained:

“People say um. People breathe.”
r/podcasting

Another reply was blunter: “AI edits that remove 100% of filler don't sound good!” Our suggestion for a middle path: remove fillers that open a sentence or sit inside a clause, and keep the ones at the end of a thought, where a pause belongs anyway.

7Why doesn't my transcript show the ums?

Many speech models are trained to write clean text, so they drop or rewrite disfluencies, and a transcript editor cannot remove a filler it never wrote down. In a discussion on Whisper's GitHub (checked October 2026), the asker reports Whisper correcting these disfluencies out, and the suggested fix is an initial prompt written in disfluent speech. It is a mitigation, not a guarantee.

Our test showed a subtler version. Whisper kept the synthetic voice's ums, with and without the prompt, but wrote the uh as “ah” and the doubled “the the” as “that the”. A detector that only knows uh misses the first, and nothing catches the second, because the repeat is gone from the text. TechSmith's blog (checked October 2026) puts it well: “accuracy may vary across accents, speaking styles, and non-native speech, so review still matters”.

8Can you remove filler words with Whisper and ffmpeg?

Yes. Transcribe with word timestamps, list the fillers, build the stretches to keep, and let ffmpeg cut picture and sound together. We ran every command below on the test clip with openai-whisper and ffmpeg 9.0.2. The Whisper CLI source (checked October 2026) marks --word_timestamps experimental. The prompt is the disfluent one quoted in the GitHub discussion and in a Hugging Face Whisper space thread (both checked October 2026). It writes talk.json, with a words list in each segment:

pip install -U openai-whisper

whisper talk.mp4 --model base.en --word_timestamps True \
  --output_format json --fp16 False \
  --initial_prompt "Umm, let me think like, hmm... Okay, here's what I'm, like, thinking."

The script removes hesitations, removes like only between punctuation and a comma, cuts each filler to the next word's start, and writes a filtergraph that trims each kept stretch, fades its audio 20 ms at the joins, and concatenates them:

import json, re, sys

HESITATIONS = {"um", "umm", "uh", "uhh", "uhm", "er", "erm", "ah", "hmm", "mm"}
FADE = 0.02  # 20 ms audio fade on each side of every join

words = [w for s in json.load(open(sys.argv[1]))["segments"] for w in s["words"]]
bare = lambda w: re.sub(r"[^\w']", "", w["word"]).lower()

drop = []
for i, w in enumerate(words):
    prev = words[i - 1]["word"].strip() if i else "."
    # Cut to where the next word starts, not where Whisper says the filler ends.
    stop = words[i + 1]["start"] if i + 1 < len(words) else w["end"]
    if bare(w) in HESITATIONS:
        drop.append((w["start"], stop))
    elif bare(w) == "like" and w["word"].strip().endswith(",") and prev[-1] in ",.?!":
        drop.append((w["start"], stop))

keep, t = [], 0.0
for a, b in drop:
    if a > t:
        keep.append((t, a))
    t = b
keep.append((t, None))

parts, labels = [], ""
for n, (a, b) in enumerate(keep):
    end = f":end={b:.3f}" if b is not None else ""
    parts.append(f"[0:v]trim=start={a:.3f}{end},setpts=PTS-STARTPTS[v{n}]")
    fade_out = f",afade=t=out:st={b - a - FADE:.3f}:d={FADE}" if b is not None else ""
    parts.append(f"[0:a]atrim=start={a:.3f}{end},asetpts=PTS-STARTPTS,afade=t=in:d={FADE}{fade_out}[a{n}]")
    labels += f"[v{n}][a{n}]"
parts.append(f"{labels}concat=n={len(keep)}:v=1:a=1[v][a]")

print(f"{len(drop)} filler words, {sum(b - a for a, b in drop):.2f} s", file=sys.stderr)
print(";\n".join(parts))
python3 fillers.py talk.json > cuts.txt
# 4 filler words, 2.08 s

ffmpeg -i talk.mp4 -/filter_complex cuts.txt -map "[v]" -map "[a]" \
  -c:v libx264 -crf 18 -c:a aac clean.mp4
# video 8.37 s -> 6.27 s, audio 8.37 s -> 6.31 s

It read back as “So, I think that the plan is, to ship it next week, we tested it, twice, I like it.”: four fillers gone, the real like kept. ffmpeg 9 reads a filtergraph file with -/filter_complex; older builds used -filter_complex_script, which we did not test. The script skips the review step, so print its drop list and read it before trusting it on real footage. Our ffmpeg command generator writes simpler keep-segment cuts, and the loudness checker confirms the level after you re-encode.

9How do Descript, Premiere, DaVinci Resolve and CapCut compare?

Descript has the most controls, Premiere and CapCut detect and bulk delete from the transcript, and the Resolve 20 guide documents no filler detector. Only what each vendor documents, checked October 2026.

  • Descript. Its filler words help page says fillers are detected and underlined automatically (English transcripts only). Remove filler words lists each one with its timestamp; you choose Delete, Delete and replace with gap, Ignore, or Remove from transcript. It uses AI Credits on current plans.
  • Adobe Premiere. The Text-Based Editing help (updated January 7, 2026) has a Transcript panel filter where you choose Filler words and delete them in bulk.
  • DaVinci Resolve. The word filler does not appear in the Resolve 20 New Features Guide, which documents a Fairlight Remove Silence tool instead. We did not check Resolve 21.
  • CapCut. Its filler words page describes the desktop flow: Auto captions with the filler option on, then delete the highlighted fillers or adjust them by hand.

10How does Linkeddit Studio remove filler words?

Studio's transcript panel has a Remove filler words button. It strikes through each word it would cut, shows the count and seconds, such as “12 filler words, 4.1 s”, and cuts only when you press Remove. We make Studio, so read this as the vendor talking. The rules are fixed in code:

  • Hesitations are always taken: um, umm, uh, uhh, uhm, erm, er, hmm and mm, in any case. Ah is not on the list, so an uh written as ah stays.
  • Like and you know are taken only when set off: a comma after, or punctuation or a pause of 0.3 seconds or more on both sides.
  • The first of a word said twice is taken unless it ends a clause, so a correct “had had” would go too. So and actually are not detected.

There is no per-word toggle: to keep one, press Cancel, select the words you want gone and press Delete. Pressing Remove is one undo step and takes the same stretch from every track, picture, sound, music and captions, then closes the gap.

The limits, given section 4: Studio cuts each filler exactly from its transcript start to its end, with no padding and no fade at the join. With your own OpenAI, Groq or Google Gemini key it gets word-level times. The on-device Whisper model times whole phrases, so Studio spreads each phrase across its words by length and greys those words as estimated; cuts on estimated times are rougher. Listen to the joins and undo any that leave a tail.

See every filler before it goes

Linkeddit Studio is a desktop editor for Mac and Windows where Claude Code or Codex edits your footage. Remove filler words and tighten pauses from the transcript, see what will go first, and undo any cut. Files stay on your disk. One-time $99, no credits.
See Linkeddit Studio

Frequently asked questions

How do I remove um and uh from a video for free?+

Cut them by hand in any free editor, or use the Whisper and ffmpeg script in this guide, which is free. The FAQ on CapCut's filler words page says you can remove them "for free with CapCut's trial", so check its current terms first.

Does removing filler words make you sound robotic?+

It can if you take every one. Remove the fillers that interrupt a sentence, keep the ones at natural thinking points, and listen back at normal speed.

Should I remove like and you know too?+

Only when they are set off from the sentence, as in "it was, like, great". In "I like it" or "do you know him?" they are real words.

Why does my editor find no filler words when I can hear them?+

A transcript-based detector only finds fillers the transcript wrote down. Speech models often drop or rewrite them; in our test Whisper wrote uh as ah. Re-transcribe with another model or prompt, or cut those by hand.

Can Whisper detect filler words?+

It can transcribe them but often leaves them out. A disfluent initial prompt helps, and --word_timestamps True times each word. In our test the end times came back early, so cut to the next word's start.

Should I cut filler words from the audio or the video?+

Both at once. Deleting them from the audio alone pushes everything after the first cut out of sync with the lips.