FFmpeg’s silencedetect filter finds the silent sections in an audio or video file and reports where each one starts and ends and how long it lasts. You decide what counts as silence with a volume threshold (noise) and a minimum length (duration). Use it to find the silence before you cut a podcast or recording automatically or split it into chapters.

Basic Command

ffmpeg -i input.mp3 -af silencedetect -f null /dev/null

This writes no output file. It prints the silent sections in the terminal (standard error).

How to Read the Output

[silencedetect @ 0x...] silence_start: 5.32
[silencedetect @ 0x...] silence_end: 8.15 | silence_duration: 2.83
[silencedetect @ 0x...] silence_start: 45.7
[silencedetect @ 0x...] silence_end: 48.2 | silence_duration: 2.5
Field Description
silence_start Start time of the silence (seconds)
silence_end End time of the silence (seconds)
silence_duration Length of the silence (seconds)

Tuning the Threshold (noise) and Minimum Duration (duration)

Default Values

Parameter Default Description
noise -60dB Anything at or below this volume is treated as “silence”
duration 2 Only report silence lasting at least this many seconds

Custom Example: Treat -40dB or lower for 0.5s or longer as silence

ffmpeg -i input.mp3 -af "silencedetect=noise=-40dB:duration=0.5" -f null /dev/null
  • Raise noise (e.g., -40dB): quiet background noise also counts as silence
  • Shorten duration (e.g., 0.1s): even short pauses for breath are detected
  • Lengthen duration (e.g., 3s): only long pauses are detected

Applying to Video Files

ffmpeg -i input.mp4 -af "silencedetect=noise=-50dB:duration=1" -f null /dev/null

For a video file, the filter analyzes its audio track.

Narrowing the Output (Combining with grep)

To show only the lines you need (Linux/macOS):

ffmpeg -i input.mp3 -af silencedetect -f null /dev/null 2>&1 | grep silence_

On the Windows Command Prompt, use findstr silence_ instead.

Application: Automatic Cutting Using Silent Sections

To cut with -ss/-to based on the silencedetect output:

  1. Get silence_start / silence_end with silencedetect
  2. Work out the time ranges between the silent sections
  3. Cut those ranges with an ffmpeg trim command

If you automate these steps with a shell script or Python, you can remove the silent parts of podcasts or lecture videos in bulk.

Common Configuration Patterns

Use Case noise duration
Removing silence from podcasts -40dB 0.3
Skipping long silences in recorded footage -50dB 3
Detecting track boundaries in music files -60dB 1
Detecting only the silent parts of a recording -30dB 0.5

Measured: time and size

The measured command:

ffmpeg -i input.mp4 -af silencedetect=noise=-40dB:duration=0.5 -f null -
Measured
Wall time 7.07 s
Speed 16.97x realtime
Output file none — analysis only

Test machine: Core i9-14900KF (32 threads), FFmpeg 8.1 (gyan.dev). Source: 1920x1080, 30 fps, 120 s, 351.4 MB of video (Big Buck Bunny, looped, CC BY 3.0) plus a 128 kb/s AAC sine-wave track, 353.29 MB in total. Measured 2026-09-05, one run at a time. The raw numbers are in the dataset.

Why it lands at 7.07 seconds

Nothing is written to disk. On the same source and machine, a plain re-encode with no filter (ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset medium -an output.mp4) takes 27.97 s and writes an 86.26 MB file. silencedetect only reads the audio and prints to the log, and -f null - takes the video frames and throws them away. The x264 encode is gone, and that accounts for the whole difference. Most of the remaining 7.07 s goes into decoding the 1080p video only to throw it away. If you only need the audio analysed, add -vn; the video is then not decoded at all.

It still takes a few seconds. In the same set of measurements, jobs that use -c copy were much faster: 0.41 s for stream mapping, 0.49 s for moving the moov atom and 1.03 s for an audio fade. All three copy the video stream instead of decoding it; the fade re-encodes only the audio. silencedetect decodes all 120 seconds to find every silent range. An analysis pass like this is much faster than a re-encode, but slower than a stream copy.

Detection is only the step before the edit. Seven seconds is short next to what you do with the result: in the same set of measurements, scaling to 720p took 18.43 s and a conversion to VP9 took 729.16 s. When you estimate the total time, plan around the step that re-encodes, not the analysis.

The threshold changes how many silent ranges you get. A lower threshold (such as -50dB) counts quieter sounds as sound, so you get fewer silent ranges. A higher one (such as -30dB) counts quiet room noise as silence, so you get more. Try two or three thresholds on your own material.

When you use the result for automatic cuts, leave 100–300 ms of padding around the speech instead of cutting each silent range right up to its edges. Otherwise the starts of words get cut off.

With several microphones, one channel can be silent while the conversation goes on in another. Decide first whether to detect on the mixed track or to check each channel separately.

Frequently Asked Questions

What threshold should I use for silence detection?

FFmpeg’s default is noise=0.001, which is about -60dB. As a guide, set it 5–10 dB above the noise floor (the level of the background noise) of your recording. Start around -50dB for studio recordings and around -25dB for outdoor recordings, then adjust.

I want to automatically cut out silent sections

Use the silenceremove filter. -af "silenceremove=start_periods=1:start_silence=0.5:start_threshold=-30dB" removes the silence at the beginning only; start_periods never touches the end. To shorten gaps in the middle, use stop_periods=-1:stop_duration=1. For silence at the end, reverse the audio with areverse, remove the silence at the start, and reverse it back. Remove silence automatically shows the measured output length for each setting.

How do I get the silence timestamps?

ffmpeg -i in.mp4 -af silencedetect=noise=-30dB:d=0.5 -f null - 2>&1 | grep silence prints the start and end times of the silent sections.

How should I set the duration?

d=0.5 (silence of 0.5 seconds or longer) is a common choice. If it is too short, breaths and short pauses between words are detected too. Adjust it to the use: 1 second or more for vlogs, 0.3 seconds for narration, and so on.

Can I detect silence even with background music?

Not if the music is louder than the threshold: it keeps the level up during pauses in the speech, so the pauses are not detected as silence. Music quieter than the threshold counts as silence. If the music is on its own audio track, select only the speech track with -map 0:a:0 (or whichever number it has) and run silencedetect on that. If speech and music are already mixed into one track, FFmpeg alone cannot separate them.