Video Editing
How to Remove Silence From Video Without Killing the Pacing
Removing silence is easy. Removing the right silence is the hard part: the same tool that clears dead air will also eat a laugh, a breath or a pause that was doing work. Here are the four ways to do it, the three settings that decide what goes, and a measured test of what each setting actually cuts.
Key takeaways
- Four methods: cut by hand, detect and cut with ffmpeg, an editor's auto tool (Premiere's Text-Based Editing, Descript's Shorten word gaps, Resolve Studio's IntelliCut), or tighten pauses in a transcript.
- Three settings decide everything: the threshold (what counts as quiet), the minimum silence (which pauses are touched) and the padding (how long the pause you leave behind is).
- In our 25 second test, -40 dB kept a quiet laugh and removed 5.9 s; -30 dB removed 6.8 s and the laugh with it. The detector cannot tell a laugh from dead air.
- ffmpeg's silenceremove filter shortens only the audio. On our clip the audio came out at 21.39 s and the video stayed at 25 s.
- On a 10 minute talking head, our arithmetic gives 100 jump cuts at long-form settings and 250 at Shorts settings. The minimum silence length is the pacing control.
1How do you remove silence from a video?
Find the quiet stretches, decide which ones to cut, and cut video and audio together so they stay in sync. You can do that by hand, with ffmpeg, with an editor's automatic tool, or by tightening the pauses in a transcript. The search results for this query are an App Store listing, a vendor's guide, a Reddit thread, a help article and a GitHub project. None of them compares the methods or measures what a setting really removes, so that is what this guide does.
| Method | What it listens to | Best for |
|---|---|---|
| By hand | Your ears and the waveform | Short pieces where every pause is a choice |
| ffmpeg (silencedetect, then cut) | Loudness only | Free, scriptable, any length |
| Editor auto tools | Loudness or the transcript, by tool | Work already inside that editor |
| Transcript pause tightening | Gaps between transcribed words | Talking heads, podcasts, interviews |
Neither kind knows whether a pause mattered, which is why the settings in the next section carry the weight.
2What settings decide what gets cut?
Three settings: the threshold sets what counts as quiet, the minimum silence sets which pauses are long enough to touch, and the padding sets how much pause is left behind after each cut. ffmpeg's silencedetect documentation (checked October 2026) names the first two: a noise tolerance with a default of -60 dB, and a duration with a default of 2 seconds. Padding is not a silencedetect option; you apply it when you turn the silences into cuts.
The way we think about padding: it is not a safety margin, it is your new pause length. A cut with 150 ms of padding leaves 150 ms on each side, so every removed gap becomes a 0.3 second pause, no matter whether it started as 0.6 seconds or 2.5. We measured exactly that in the test below: after the cut, every pause that had been cut read 0.30 seconds, and the one short beat below the minimum stayed at 0.25. The minimum silence decides which pauses get that treatment and which stay as recorded.
The threshold depends on your room more than on any rule. Our silence detector has a worked example of picking one for a quiet room and a noisy one, and it recomputes as you drag, so you can see the effect on your own file before you cut anything.
3What does each setting actually remove? A measured test
On our 25 second test clip, -40 dB with a 0.5 second minimum and 150 ms padding removed 5.9 seconds and kept a quiet laugh. Moving only the threshold to -30 dB removed 6.8 seconds, and the laugh went with it. We built the clip with ffmpeg 9.0.2: a 720p test picture at 30 fps, with a 220 Hz tone standing in for speech (about -23 dB RMS), a 400 Hz tone standing in for a laugh (about -37 dB RMS) and faint room tone (about -65 dB RMS) in between. Synthetic on purpose: we know where every gap is. The gaps are 0.25, 0.6, 1.2, 2.5 (with the laugh 0.9 seconds in) and 1.0 seconds, with 1 second of lead-in and 2 seconds of tail.
Step one is detection. ffmpeg prints each silence it finds:
ffmpeg -hide_banner -i talk.mp4 -af silencedetect=n=-40dB:d=0.5 -f null - # silence_start: 0 silence_end: 1.000083 # silence_start: 6.999937 silence_end: 7.600083 # silence_start: 9.999958 silence_end: 11.200083 # silence_start: 13.999937 silence_end: 14.900229 # silence_start: 15.499792 silence_end: 16.500083 # silence_start: 18.999937 silence_end: 20.000083 # silence_start: 22.999937 silence_end: 25
The 2.5 second gap comes back as two silences, 14.0 to 14.9 and 15.5 to 16.5, because the laugh breaks it. The 0.25 second beat is shorter than the 0.5 second minimum, so it is not listed. Step two is turning silences into ranges to keep, with 150 ms of padding on each side that touches speech (our arithmetic: 7.0 + 0.15 = 7.15, 7.6 - 0.15 = 7.45, and so on). Then the keep-segments command, the same recipe our ffmpeg command generator writes, cuts picture and sound together:
ffmpeg -i talk.mp4 \ -vf 'select=between(t\,0.85\,7.15)+between(t\,7.45\,10.15)+between(t\,11.05\,14.15)+between(t\,14.75\,15.65)+between(t\,16.35\,19.15)+between(t\,19.85\,23.15),setpts=N/FRAME_RATE/TB' \ -af 'aselect=between(t\,0.85\,7.15)+between(t\,7.45\,10.15)+between(t\,11.05\,14.15)+between(t\,14.75\,15.65)+between(t\,16.35\,19.15)+between(t\,19.85\,23.15),asetpts=N/SR/TB' \ long.mp4
| Setting | Predicted kept | Measured video | Measured audio | Laugh |
|---|---|---|---|---|
| -40 dB, 0.5 s minimum, 150 ms padding | 19.10 s | 19.10 s (573 frames) | 19.11 s | Kept |
| -30 dB, 0.5 s minimum, 150 ms padding | 18.20 s | 18.20 s | 18.22 s | Cut |
| -40 dB, 0.2 s minimum, 80 ms padding | 18.17 s | 18.20 s | 18.18 s | Kept |
Our arithmetic for the first row: the cuts are 0.85 + 0.3 + 0.9 + 0.6 + 0.7 + 0.7 + 1.85 = 5.9 seconds, so 25 - 5.9 = 19.1 seconds stay. At -30 dB the laugh, which peaks around -34 dBFS, falls below the threshold, the whole 2.5 second gap becomes one silence, and the cut grows to 2.2 seconds. We measured the laugh's spot in each output: -37 dB RMS in the first file, -23 dB (speech) in the second, because the laugh was gone and the next sentence had moved up. The third row also removed the 0.25 second beat, the kind of short gap where a speaker might take a breath.
The third row's video is 0.03 seconds over the prediction because select keeps whole frames (33 ms each at 30 fps). The logic behind our silence detector gave identical keep ranges for all three settings.
4Why not just use ffmpeg's silenceremove filter?
Because silenceremove is an audio filter. On a video file it shortens the sound and leaves the picture at full length, so every cut pushes the audio further ahead of the lips. The silenceremove documentation (checked October 2026) describes it as removing silence from the beginning, middle or end of the audio. We ran it on the test clip with the same -40 dB and 0.5 second settings:
ffmpeg -i talk.mp4 \ -af "silenceremove=stop_periods=-1:stop_duration=0.5:stop_threshold=-40dB" \ -c:v copy sr.mp4 # video stream: 25.000000 s # audio stream: 21.393458 s
3.6 seconds of drift by the end of a 25 second clip. The same trap exists in editors. In a thread on r/editors about Avid, the poster said the “Strip Silence from Sequence” feature finds the gaps well but:
“it only affects the audio, not the picture.”
The keep-segments command in section 3 cuts both. Use silenceremove only on audio you will publish as audio.
5How do I keep laughter, breaths and natural pauses?
Lower the threshold so quiet sounds stay above it, raise the minimum silence so short beats are left alone, and listen to every cut near a laugh. No detector, whether it reads the waveform or the transcript, can tell a meaningful pause from dead air. Our test shows how small the margin is: the laugh survived at -40 dB and vanished at -30 dB, a 10 dB move.
Experienced editors treat silence as material. The poster of an r/editors thread on whether AI cutting tools work listed the ways they bring intent to their own cuts, including “holding silence when it improves impact”. And in a thread about jump cuts from auto silence removal, one reply suggested exactly the setting this guide keeps coming back to:
“I'd raise that duration to leave shorter natural pauses but chop out longer silences.”
Short beats are the other casualty: our 0.2 second minimum cut the 0.25 second beat in the test. If speech sounds hurried after a cut, raise the minimum before you touch anything else. One more point: cut pauses, do not mute them. A poster on r/podcasting noticed that a voice cleaned to the point of “digital silence between phrases” often ends up sounding cold or lifeless. Our reading: cutting keeps the room tone that is there, while muting replaces it with nothing.
The other question people ask about every silence tool is the one from an r/VideoEditing thread about Movavi: “Does it clip the beginnings or ends of words?” It can. The end of a word fades out, and with too little padding or too loose a threshold the cut lands inside the fade. More padding is the fix, and listening to the joins is the only real check.
6How tight should you cut for long-form vs Shorts?
Tight for Shorts, loose for long-form, and the setting that makes the difference is the minimum silence, because it decides how many cuts you make. Each cut on a single camera is a jump cut. The poster of an r/editors thread on auto silence removal found the cuts jarring “whenever the head moved or the eyeline shifted”. The top reply was blunt:
“That feature was designed specifically to create jump cuts in order to get runtime shorter.”
Here is a worked example with our own numbers, for a hypothetical 10 minute talking head with 250 pauses: 150 short beats of 0.4 seconds, 80 pauses of 0.9 seconds and 20 long gaps of 2.5 seconds. That is 60 + 72 + 50 = 182 seconds of pause in 600 seconds.
| Setting (ours, not a standard) | Pauses cut | Time removed | New length | Every cut pause becomes |
|---|---|---|---|---|
| Long-form: 0.7 s minimum, 200 ms padding | 100 | 82 s | 8:38 | 0.4 s |
| Shorts: 0.3 s minimum, 80 ms padding | 250 | 142 s | 7:38 | 0.16 s |
| Pause limit of 1 s (transcript method) | 20 | 30 s | 9:30 | 1.0 s |
The arithmetic: at long-form settings only the 0.9 and 2.5 second pauses qualify, and each loses its length minus 0.4 seconds, so 80 x 0.5 + 20 x 2.1 = 82 seconds. At Shorts settings every pause qualifies and loses its length minus 0.16, so 150 x 0.24 + 80 x 0.74 + 20 x 2.34 = 142 seconds. A pause limit of 1 second shortens only the 20 long gaps to 1 second each, 20 x 1.5 = 30 seconds.
Look at the cut count, not just the time. The Shorts settings save one more minute and make two and a half times as many jump cuts. For a vertical clip with captions, that can be right. For a 10 minute video it is 250 visible jumps, about one every 1.8 seconds of the finished 7:38 (458 / 250). Our suggestion: start long-form at a high minimum, cover the cuts you keep with a second angle, a punch-in or B-roll, and tighten further only if retention asks for it.
7Which editors remove silence automatically?
Premiere Pro, Descript and DaVinci Resolve Studio all document a way to remove pauses or silence. We list only what each vendor documents on its own pages, checked October 2026.
- Premiere Pro. Adobe's Detect and delete pauses in transcripts page (last updated January 7, 2026) says Text-Based Editing lets you detect pauses and bulk delete them: open the filter in the Transcript panel, choose Pauses, then Delete for one instance or Delete all. It works from the transcript.
- Descript. Shorten word gaps lets you define a gap as more than, or between, two durations and set a new target length, such as 200 ms, then Shorten one or Shorten all. Descript says it uses AI Credits on current plans, as does Remove filler words, which has an option to skip fillers that cannot be removed without a harsh cut.
- DaVinci Resolve. Blackmagic's own Mac App Store listing says the paid DaVinci Resolve Studio includes “IntelliCut to remove silence”, and its version notes list “AI IntelliCut to remove silences” under the Fairlight page. The 20.2.1 notes add an “Option to ripple delete silence for selected clip” on the Cut and Edit pages. The listing does not document the settings.
- CapCut. CapCut has a Silence Remover page, but it describes a general editing workflow and names no threshold, duration or padding control, and we found no help article on it. So we cannot tell you what it measures.
Two of these four work from the transcript, so the laughter problem is the same in a different form: a laugh that is not written down as a word sits inside a pause, and the pause is what gets shortened.
8How does Linkeddit Studio tighten pauses?
Studio tightens pauses from the transcript, not the waveform: you pick a pause limit, it shows how many pauses are longer and how many seconds would go, and you press Remove. We make Studio, so read this as the vendor talking, and note what it does not do: Studio has no waveform silence remover.
The pause limit is 0.5, 1, 1.5 or 2 seconds. Every gap between transcribed words longer than the limit is shortened to the limit, with half of it kept on each side, so a 2.5 second gap at a 1 second limit loses 1.5 seconds. That is the third row of the table in section 6: few cuts, long pauses kept. Remove filler words works the same way, with um and uh, the first of a word said twice in a row, and “like” or “you know” when set off by commas or pauses. Either cut is one undo step and takes every track with it, so picture and sound stay together.
The caveat from section 7 applies. If our test gap were in a Studio transcript, the laugh 0.9 seconds into it would sit in the part a 1 second limit removes, unless the transcript had written it as a word. Studio strikes through each pause it would cut before you press Remove, which is the moment to listen.
Tighten pauses by reading, not by guessing
Frequently asked questions
How do I remove silence from a video for free?+
Find the silences with ffmpeg's silencedetect filter (or our browser silence detector), turn them into a list of ranges to keep, and cut video and audio together with ffmpeg's select and aselect filters. Both tools are free. Do not use silenceremove on a video file: it shortens only the audio, so the picture drifts out of sync.
What is a good silence threshold in dB?+
There is no standard. ffmpeg's own default is -60 dB with a 2 second minimum. Our starting point for a close microphone in a quiet room is -40 dB, then we move it until only real pauses are marked. In our test, a quiet laugh at about -37 dB RMS survived at -40 dB and was cut at -30 dB.
How much padding should I leave around each cut?+
Padding is the length of the pause you leave behind: with 150 ms of padding, every removed gap becomes a 0.3 second pause. In our view 150 to 250 ms suits long-form talking heads and 80 to 120 ms suits Shorts. If the ends of words sound clipped, raise the padding first.
Why does automatic silence removal make jump cuts?+
Every removed pause is a cut, and on a single camera the head and eyeline move between the two sides. Raising the minimum silence length cuts fewer pauses and so makes fewer jump cuts; a second angle, a punch-in or B-roll can cover the ones you keep.
Does Premiere Pro have automatic silence removal?+
Yes, through Text-Based Editing. Adobe's help page (last updated January 7, 2026, checked October 2026) says you open the filter in the Transcript panel, choose Pauses, and select Delete or Delete all.