Free tool · Video
Open a video or audio file and this page marks every stretch of silence. You set three things: the threshold in dB, below which audio counts as quiet, the minimum length a quiet stretch must reach, and the padding kept around each cut. You get a cut list, a ready ffmpeg command and a trimmed WAV. The file is analysed in your browser and never uploaded.
Free, no account required.
Open an audio or video file
Your file never leaves your device. It is decoded and analysed in this browser tab.
The detector looks at the audio in short slices of 10 milliseconds. For each slice it measures the loudness as an RMS level in decibels relative to full scale, where 0 dB is the loudest a file can be and quiet room tone sits far below it. A slice quieter than your threshold is silent. Neighbouring silent slices join into a run, and a run only counts as a silence once it is at least as long as the minimum silence you set.
That is all it does, and it is worth being clear about the limit. The detector hears loudness and nothing else. It cannot tell a pause that carries meaning from dead air, a laugh from a cough, or a whisper from room noise. It is a fast way to produce a first cut list, which you then check by eye on the timeline and by ear on the result.
Each silence is then shrunk by the padding on the sides that touch speech. What is left over is the keep list: the ranges to retain. The cut list on the page shows both, the silences found and the ranges kept, in seconds with three decimals. The ranges are written in the form ffmpeg needs, so there is no copying of timecodes by hand. A stretch of silence at the very start or end of the file is removed whole, because there is no speech on the far side to protect.
Stereo and multichannel files are averaged to one channel before the measurement, so a silence in the mix is silence in the result. The audio is decoded at 48 kHz, which is a resampling step for files recorded at another rate and does not change where the pauses are.
The threshold is the dB level below which audio is treated as quiet. A lower number, such as -50 dB, is stricter: only very quiet audio counts as silence. A higher number, such as -30 dB, is looser: more of the file counts as silence, including soft speech. Because the scale is negative, the direction trips people up. Moving the slider left makes the detector cut less.
Here is a worked example with our own numbers, not a standard. Say you record with a close microphone in a quiet room. Speech reads around -20 dB, the room tone between sentences reads around -55 dB, and a soft aside reads around -35 dB. At -40 dB, the room tone is cut and the soft aside survives, which is right. At -30 dB the aside is cut too, and you lose a line you meant to keep. At -60 dB nothing is cut, because even the room tone is louder than the threshold.
Now move to a noisy room, where the fan and traffic put the room tone around -35 dB. At -40 dB the detector finds no silence at all, because the pauses are louder than the line. The fix is to raise the threshold to about -30 dB, then check that speech still reads above it. The practical method is the same in both rooms: set the threshold, look at the timeline, and move it until the grey regions sit on the pauses you can see and nothing else. Our default of -40 dB is a starting point for a quiet room and nothing more.
If speech and noise are too close in level for any threshold to separate them, no setting will work and you are better off cleaning the audio first or cutting by hand. Detecting silence cannot fix a recording where the silence is not quiet.
The most common complaint about automatic silence removal is that it cuts things people wanted. In a Descript thread on r/Descript, a podcaster listing why they were leaving the tool wrote that its gap shortening would also remove stuff like laughter, calling it a really important piece of content. In a thread on r/podcasting about the same tool, the complaint was that automatic gap removal sometimes cuts audio where someone has clearly spoken, and the poster asked for a sensitivity setting.
"Auto gap removal cuts spoken audio... Needs a sensitivity setting badly."A post on r/podcasting
The three sliders on this page are that sensitivity setting. A laugh is often quieter than speech, and it comes after a gap, so it is exactly what a loose threshold and a short minimum silence will eat. Three adjustments protect it. Lower the threshold so quiet sounds stay above the line. Raise the minimum silence so only long gaps count. Raise the padding so each cut leaves a cushion of sound on both sides.
Padding also covers a plain mechanical problem. The detector works on 10 millisecond slices, and the end of a word fades out. Without padding, a cut can land inside the fade and the word sounds chopped. A padding of 100 to 200 milliseconds is a good place to start, and it costs little: each cut is shortened by twice that amount, so a one second pause cut with 150 ms padding removes 0.7 seconds, not a full second. That is our own rule of thumb from testing, not a published standard.
No setting fully solves this. Because the tool sees only loudness, it will always be possible to find a laugh or a pause with meaning that it cuts. Treat the cut list as a draft, listen to the result, and put back anything that mattered by loosening the settings and running it again.
How aggressively to cut depends on where the video will play. Short vertical video rewards tight pacing, and a pause can read as a reason to scroll. Long talking-head video is different. A creator on r/VideoEditing described auto cutting every breath and filler on YouTube vlogs, and wrote that people said a 12-minute video felt like a TikTok. They now leave some air in and cut only the true dead spots.
"people said a 12-minute video felt like a TikTok."A post on r/VideoEditing
This is one creator's report, not a measurement, but it matches a sound instinct: listeners use pauses to take in what was said. Our suggestion, again our own and not a rule, is to use two presets. For a short, try a minimum silence of 250 to 400 ms with 80 to 120 ms of padding, which tightens the delivery. For a long-form video or a podcast, try a minimum silence of 700 to 1,000 ms with 150 to 250 ms of padding, which removes dead air and leaves the breathing room.
Because this page recomputes as you move a slider, you can try both and compare the total time removed. If a long-form cut removes a surprisingly large share of the file, check whether the threshold is too loose before accepting it, and listen to the result.
There are three ways to use the result. The first is the ffmpeg command. It uses the select and aselect filters to keep only frames and audio inside the listed ranges, then rewrites the timestamps so the gaps close. Run it in a terminal from the folder that holds your video. Because the video is filtered, it is re-encoded, which takes time on a long file and may change the file size. The command is built with the same recipe as the FFmpeg Command Generator, where you can change codec, quality and the output name.
The second way is the trimmed WAV. It contains only the kept audio, joined end to end. Use it when you only need the sound, for example a podcast that you will publish as audio, or to check by ear how the cut will sound before you commit the video to a long encode. It is 16-bit PCM at 48 kHz with the same channel count as the source.
The third way is the cut list itself. The silences are listed in seconds, so you can type them as markers into any editor that accepts timecodes and cut by hand. That is slower, but it lets you skip a cut you do not trust. After the cut, the Loudness Checker is a natural next step, because the trimmed file is a new file with its own integrated loudness and it is worth measuring again.
One caution about timing. The video filter selects whole frames, so each cut snaps to a frame boundary. At 30 frames per second that is about 33 milliseconds, well inside the padding. The audio can end up a few hundredths of a second shorter than the video because audio is stored in blocks. Neither is audible, but if you chain the result into a longer edit, check the sync on your own file.
Step 1
Press Choose file and pick an MP4, MOV, MP3, WAV or similar. The browser decodes the audio on your device. If it cannot decode the codec, the page says so.
Step 2
Start with -40 dB, 500 ms and 150 ms. Raise the minimum silence if breaths are being cut, and raise the padding if the starts and ends of words are clipped.
Step 3
Black on the timeline is kept and grey is removed. The headline shows how much time the cut removes from the whole file.
Step 4
Copy the command and run it in a terminal from the folder that holds the video. It keeps the listed ranges in video and audio together. The WAV button gives you only the trimmed audio.
When you outgrow this tool
This detector is free and complete on its own, with no account. Linkeddit Studio is the desktop editor for the cut itself. In Studio you tighten pauses and remove filler words from the transcript, so you decide what goes by reading, and every cut can be undone. Studio is a one-time purchase and your media stays on your disk.
FAQ
Direct answers on thresholds, padding and applying the cut list.
Open the video on this page, adjust the three sliders until the timeline looks right, then copy the ffmpeg command and run it. The command keeps the ranges that have speech and joins them, in video and audio together. ffmpeg is free. This page only finds the silence and writes the command, so nothing is uploaded and no account is needed.
Our starting point is -40 dB, and it suits a quiet room with a close microphone. In a noisy room the pauses sit higher than that, so the threshold has to be raised toward -30 dB or the room tone will never count as silence. Watch the timeline: if the grey regions sit only on real pauses, the threshold is right. This is our own advice, not a standard.
It decides which pauses are cut. A short minimum such as 200 ms removes breaths and the small beats between words, which makes speech sound hurried. A minimum near one second removes only real dead air. Raise it first if the result sounds rushed. This is advice from our own testing, not a rule.
Padding keeps a little audio on each side of every cut, so a cut does not land on the last syllable of a word or the first breath of the next. At 150 ms the tail of a word survives. Padding is not applied to the very start or end of the file. If your cut list removes less than you expected, a long padding is the usual reason.
No. The browser only decodes the audio to find the silence. The video is never touched here. You run the generated ffmpeg command yourself, and that command does re-encode the video, because it drops frames and rebuilds the timeline. That takes time on a long file, and cuts snap to video frames, so a cut can land a few hundredths of a second away from the exact time.
It can. This tool only sees loudness, so it cannot tell a laugh, a soft aside or a meaningful pause from dead air. Quiet speech below the threshold will be cut. Lower the threshold, raise the minimum silence and check the timeline before you run the command. Always listen to the result.
Any audio or video file your browser can decode, which usually means MP4, MOV and WebM video and MP3, WAV, M4A and AAC audio. A file with no audio track or an unusual codec shows an error. The ffmpeg command works on any file ffmpeg can read, so a file the browser cannot decode may still work if you write the ranges by hand.
No. The browser reads the file from your disk, decodes the audio with the Web Audio API on your device, and plain JavaScript finds the silence. There is no account and no network request for the file, so unreleased footage stays private.
Every tool on the shelf is free, with no card and no trial. Most need no account at all. See all free tools.
Work out file size from bitrate, the bitrate that fits a size limit, or how long a video can run, with audio included.
Build a correct ffmpeg command for trimming, compressing, converting, cropping and more, with every flag explained and linked to the ffmpeg docs.
Pick a resolution, frame rate and SDR or HDR and get YouTube's recommended container, codec, bitrate and audio settings, with the source linked.
Reduce a resolution to its ratio, scale to a width or height with even dimensions, and see the crop or padding needed to reframe, such as 16:9 to 9:16.