What is Audio Ducking?
Audio ducking lowers the background music (BGM) automatically while someone is speaking. It is common in podcasts and narrated videos. In FFmpeg, the sidechaincompress filter does it.
How It Works
Narration (sidechain signal)
↓ level rises
Compressor attenuates the BGM
↓
BGM level drops automatically
sidechaincompress takes two audio inputs:
- The main signal (the BGM): the track that gets turned down
- The sidechain signal (the narration): the track whose level decides when that happens
Basic Command
This command ducks the audio of a video under a separate narration file:
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked]" \
-map 0:v -map "[ducked]" -c:v copy output_ducked.mp4
Here the audio in input.mp4 is the BGM and input.mp3 is the narration. The BGM gets quieter whenever the narration is speaking. apad fills the time after the narration ends with silence. Without it, a narration shorter than the BGM cuts the BGM off where the narration ends. The narration itself is not in the output: it only controls the ducking. The amix example further down puts it back in.
Parameter Reference
| Parameter | Description | Default | Recommended range |
|---|---|---|---|
threshold |
Sidechain level above which compression starts | 0.125 | 0.01–0.05 |
ratio |
Compression ratio (e.g. 4 = 4:1) | 2 | 3–8 |
attack |
Time before the compressor kicks in (ms) | 20 | 100–300 |
release |
Time to return to the original level (ms) | 250 | 500–2000 |
makeup |
Gain applied after compression (a factor from 1 to 64, not dB) | 1 | — |
knee |
Knee width (softness of the threshold) | 2.82843 | — |
Tuning threshold
Set threshold to suit how loud your narration is. A smaller value makes the ducking react to quieter speech. This example raises it to 0.05, so the narration has to be louder before the BGM ducks:
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.05:ratio=4:attack=200:release=1000[ducked]" \
-map 0:v -map "[ducked]" -c:v copy output.mp4
Shaping a Natural Ducking Curve
attack sets how quickly the BGM drops when speech starts. release sets how quickly it comes back after speech stops.
Slow, natural ducking
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=6:attack=300:release=1500[ducked]" \
-map 0:v -map "[ducked]" -c:v copy output_natural.mp4
Fast, tight ducking
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=8:attack=50:release=500[ducked]" \
-map 0:v -map "[ducked]" -c:v copy output_tight.mp4
Pre-Attenuate the BGM Before Ducking
If the BGM is too loud even without ducking, lower it with volume first, then duck it:
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]volume=0.5,aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked]" \
-map 0:v -map "[ducked]" -c:v copy output.mp4
Mix Narration Back In After Ducking
For a final mix with both the ducked BGM and the narration, duck the BGM first and then mix the narration in with amix:
ffmpeg -i input.mp4 -i input.mp3 \
-filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,asplit=2[sc][voice2];[sc]apad[voice1];[bg][voice1]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked];[ducked][voice2]amix=inputs=2:duration=longest[out]" \
-map 0:v -map "[out]" -c:v copy output_mixed.mp4
By default amix turns each input down to half (1/N for N inputs), so the narration and the BGM come out about 6 dB quieter. Add normalize=0 to amix to keep their levels, and watch for clipping.
Why asplit is needed here
The narration is used twice: once as the sidechain that drives the compressor,
and once as an input to amix. A label in a filtergraph can be used as an
input only once, so asplit=2 makes two copies of it.
FFmpeg 8.x inserts that split for you, so the missing asplit goes unnoticed
there. FFmpeg 6.1, the version Ubuntu 24.04 ships, does not, and fails with:
Invalid stream specifier: voice.
Stream specifier 'voice' in filtergraph description ... matches no streams.
With asplit written out, the command works on every version.
Notes
The two inputs of sidechaincompress can have different sample rates and channel counts; FFmpeg converts them automatically. To choose the shared format yourself, add aformat:
Example: aformat=fltp:44100:stereo
Frequently Asked Questions
What is sidechain compression in plain English?
It is a compressor that turns one signal down whenever a different signal gets loud. The usual use is lowering BGM under a voice-over. In FFmpeg, that is sidechaincompress.
What threshold and ratio should I start with?
A threshold of 0.05 (about -26 dB) and a ratio of 8:1 give clearly audible ducking for podcasts. Set attack to 200 ms and release to 1000 ms, as in the basic command. The BGM then stays down between words and comes back up naturally between sentences.
Is ducking better than just lowering BGM volume?
For narrated content, yes. A fixed volume cut also makes the music quiet between sentences. Ducking lowers it only during speech, so the music stays full the rest of the time.
Can I duck two BGM tracks with one voice?
Yes. Run each BGM through its own sidechaincompress with the same voice as the sidechain input, then mix the two ducked outputs together.
Why does my output have audible pumping?
Pumping (the BGM audibly rising and falling) means attack and release are too fast. Lengthen them to the ranges in the parameter table: attack 100–300 ms, release 500–2000 ms. Lowering the ratio to 4:1 also softens the effect.
Related Articles
- Loudness Normalization (loudnorm / LUFS)
- Volume Detection and Adjustment (volumedetect / volume)
- Audio Format Conversion
Primary source: ffmpeg.org/ffmpeg-filters.html#sidechaincompress