Free tool · Video

Silence Detector: Find and Remove Silence From Video

Open a video or audio file and this page marks every stretch of silence. You set three things: the threshold in dB, below which audio counts as quiet, the minimum length a quiet stretch must reach, and the padding kept around each cut. You get a cut list, a ready ffmpeg command and a trimmed WAV. The file is analysed in your browser and never uploaded.

Free, no account required.

Open an audio or video file

Your file never leaves your device. It is decoded and analysed in this browser tab.

How silence detection works

The detector looks at the audio in short slices of 10 milliseconds. For each slice it measures the loudness as an RMS level in decibels relative to full scale, where 0 dB is the loudest a file can be and quiet room tone sits far below it. A slice quieter than your threshold is silent. Neighbouring silent slices join into a run, and a run only counts as a silence once it is at least as long as the minimum silence you set.

That is all it does, and it is worth being clear about the limit. The detector hears loudness and nothing else. It cannot tell a pause that carries meaning from dead air, a laugh from a cough, or a whisper from room noise. It is a fast way to produce a first cut list, which you then check by eye on the timeline and by ear on the result.

Each silence is then shrunk by the padding on the sides that touch speech. What is left over is the keep list: the ranges to retain. The cut list on the page shows both, the silences found and the ranges kept, in seconds with three decimals. The ranges are written in the form ffmpeg needs, so there is no copying of timecodes by hand. A stretch of silence at the very start or end of the file is removed whole, because there is no speech on the far side to protect.

Stereo and multichannel files are averaged to one channel before the measurement, so a silence in the mix is silence in the result. The audio is decoded at 48 kHz, which is a resampling step for files recorded at another rate and does not change where the pauses are.

Picking a threshold, with a worked example

The threshold is the dB level below which audio is treated as quiet. A lower number, such as -50 dB, is stricter: only very quiet audio counts as silence. A higher number, such as -30 dB, is looser: more of the file counts as silence, including soft speech. Because the scale is negative, the direction trips people up. Moving the slider left makes the detector cut less.

Here is a worked example with our own numbers, not a standard. Say you record with a close microphone in a quiet room. Speech reads around -20 dB, the room tone between sentences reads around -55 dB, and a soft aside reads around -35 dB. At -40 dB, the room tone is cut and the soft aside survives, which is right. At -30 dB the aside is cut too, and you lose a line you meant to keep. At -60 dB nothing is cut, because even the room tone is louder than the threshold.

Now move to a noisy room, where the fan and traffic put the room tone around -35 dB. At -40 dB the detector finds no silence at all, because the pauses are louder than the line. The fix is to raise the threshold to about -30 dB, then check that speech still reads above it. The practical method is the same in both rooms: set the threshold, look at the timeline, and move it until the grey regions sit on the pauses you can see and nothing else. Our default of -40 dB is a starting point for a quiet room and nothing more.

If speech and noise are too close in level for any threshold to separate them, no setting will work and you are better off cleaning the audio first or cutting by hand. Detecting silence cannot fix a recording where the silence is not quiet.

Keep the laughter: why padding and minimum length matter

The most common complaint about automatic silence removal is that it cuts things people wanted. In a Descript thread on r/Descript, a podcaster listing why they were leaving the tool wrote that its gap shortening would also remove stuff like laughter, calling it a really important piece of content. In a thread on r/podcasting about the same tool, the complaint was that automatic gap removal sometimes cuts audio where someone has clearly spoken, and the poster asked for a sensitivity setting.

"Auto gap removal cuts spoken audio... Needs a sensitivity setting badly."A post on r/podcasting

The three sliders on this page are that sensitivity setting. A laugh is often quieter than speech, and it comes after a gap, so it is exactly what a loose threshold and a short minimum silence will eat. Three adjustments protect it. Lower the threshold so quiet sounds stay above the line. Raise the minimum silence so only long gaps count. Raise the padding so each cut leaves a cushion of sound on both sides.

Padding also covers a plain mechanical problem. The detector works on 10 millisecond slices, and the end of a word fades out. Without padding, a cut can land inside the fade and the word sounds chopped. A padding of 100 to 200 milliseconds is a good place to start, and it costs little: each cut is shortened by twice that amount, so a one second pause cut with 150 ms padding removes 0.7 seconds, not a full second. That is our own rule of thumb from testing, not a published standard.

No setting fully solves this. Because the tool sees only loudness, it will always be possible to find a laugh or a pause with meaning that it cuts. Treat the cut list as a draft, listen to the result, and put back anything that mattered by loosening the settings and running it again.

Long-form vs shorts: how tight to cut

How aggressively to cut depends on where the video will play. Short vertical video rewards tight pacing, and a pause can read as a reason to scroll. Long talking-head video is different. A creator on r/VideoEditing described auto cutting every breath and filler on YouTube vlogs, and wrote that people said a 12-minute video felt like a TikTok. They now leave some air in and cut only the true dead spots.

"people said a 12-minute video felt like a TikTok."A post on r/VideoEditing

This is one creator's report, not a measurement, but it matches a sound instinct: listeners use pauses to take in what was said. Our suggestion, again our own and not a rule, is to use two presets. For a short, try a minimum silence of 250 to 400 ms with 80 to 120 ms of padding, which tightens the delivery. For a long-form video or a podcast, try a minimum silence of 700 to 1,000 ms with 150 to 250 ms of padding, which removes dead air and leaves the breathing room.

Because this page recomputes as you move a slider, you can try both and compare the total time removed. If a long-form cut removes a surprisingly large share of the file, check whether the threshold is too loose before accepting it, and listen to the result.

How to apply the cut list

There are three ways to use the result. The first is the ffmpeg command. It uses the select and aselect filters to keep only frames and audio inside the listed ranges, then rewrites the timestamps so the gaps close. Run it in a terminal from the folder that holds your video. Because the video is filtered, it is re-encoded, which takes time on a long file and may change the file size. The command is built with the same recipe as the FFmpeg Command Generator, where you can change codec, quality and the output name.

The second way is the trimmed WAV. It contains only the kept audio, joined end to end. Use it when you only need the sound, for example a podcast that you will publish as audio, or to check by ear how the cut will sound before you commit the video to a long encode. It is 16-bit PCM at 48 kHz with the same channel count as the source.

The third way is the cut list itself. The silences are listed in seconds, so you can type them as markers into any editor that accepts timecodes and cut by hand. That is slower, but it lets you skip a cut you do not trust. After the cut, the Loudness Checker is a natural next step, because the trimmed file is a new file with its own integrated loudness and it is worth measuring again.

One caution about timing. The video filter selects whole frames, so each cut snaps to a frame boundary. At 30 frames per second that is about 33 milliseconds, well inside the padding. The audio can end up a few hundredths of a second shorter than the video because audio is stored in blocks. Neither is audible, but if you chain the result into a longer edit, check the sync on your own file.

How to remove silence from a video

Step 1

Choose a video or audio file

Press Choose file and pick an MP4, MOV, MP3, WAV or similar. The browser decodes the audio on your device. If it cannot decode the codec, the page says so.

Step 2

Set the threshold, minimum silence and padding

Start with -40 dB, 500 ms and 150 ms. Raise the minimum silence if breaths are being cut, and raise the padding if the starts and ends of words are clipped.

Step 3

Read the timeline and the total

Black on the timeline is kept and grey is removed. The headline shows how much time the cut removes from the whole file.

Step 4

Copy the ffmpeg command or download the WAV

Copy the command and run it in a terminal from the folder that holds the video. It keeps the listed ranges in video and audio together. The WAV button gives you only the trimmed audio.

When you outgrow this tool

In Linkeddit Studio you tighten pauses and remove filler words from the transcript, and every cut can be undone.

This detector is free and complete on its own, with no account. Linkeddit Studio is the desktop editor for the cut itself. In Studio you tighten pauses and remove filler words from the transcript, so you decide what goes by reading, and every cut can be undone. Studio is a one-time purchase and your media stays on your disk.

FAQ

Silence Detector (Find and Remove Silence From Video) questions

Direct answers on thresholds, padding and applying the cut list.

How do I remove silence from a video for free?

Open the video on this page, adjust the three sliders until the timeline looks right, then copy the ffmpeg command and run it. The command keeps the ranges that have speech and joins them, in video and audio together. ffmpeg is free. This page only finds the silence and writes the command, so nothing is uploaded and no account is needed.

What dB threshold should I use to detect silence?

Our starting point is -40 dB, and it suits a quiet room with a close microphone. In a noisy room the pauses sit higher than that, so the threshold has to be raised toward -30 dB or the room tone will never count as silence. Watch the timeline: if the grey regions sit only on real pauses, the threshold is right. This is our own advice, not a standard.

Why is the minimum silence length so important?

It decides which pauses are cut. A short minimum such as 200 ms removes breaths and the small beats between words, which makes speech sound hurried. A minimum near one second removes only real dead air. Raise it first if the result sounds rushed. This is advice from our own testing, not a rule.

What does padding do?

Padding keeps a little audio on each side of every cut, so a cut does not land on the last syllable of a word or the first breath of the next. At 150 ms the tail of a word survives. Padding is not applied to the very start or end of the file. If your cut list removes less than you expected, a long padding is the usual reason.

Does the video get re-encoded in my browser?

No. The browser only decodes the audio to find the silence. The video is never touched here. You run the generated ffmpeg command yourself, and that command does re-encode the video, because it drops frames and rebuilds the timeline. That takes time on a long file, and cuts snap to video frames, so a cut can land a few hundredths of a second away from the exact time.

Will it cut laughter or quiet speech?

It can. This tool only sees loudness, so it cannot tell a laugh, a soft aside or a meaningful pause from dead air. Quiet speech below the threshold will be cut. Lower the threshold, raise the minimum silence and check the timeline before you run the command. Always listen to the result.

What file types work?

Any audio or video file your browser can decode, which usually means MP4, MOV and WebM video and MP3, WAV, M4A and AAC audio. A file with no audio track or an unusual codec shows an error. The ffmpeg command works on any file ffmpeg can read, so a file the browser cannot decode may still work if you write the ranges by hand.

Is my file uploaded?

No. The browser reads the file from your disk, decodes the audio with the Web Audio API on your device, and plain JavaScript finds the silence. There is no account and no network request for the file, so unreleased footage stays private.