What is Audio Ducking?

Audio ducking lowers the background music (BGM) automatically while someone is speaking. It is common in podcasts and narrated videos. In FFmpeg, the sidechaincompress filter does it.

How It Works

Narration (sidechain signal)
       ↓ level rises
Compressor attenuates the BGM
       ↓
BGM level drops automatically

sidechaincompress takes two audio inputs:

  1. The main signal (the BGM): the track that gets turned down
  2. The sidechain signal (the narration): the track whose level decides when that happens

Basic Command

This command ducks the audio of a video under a separate narration file:

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked]" \
  -map 0:v -map "[ducked]" -c:v copy output_ducked.mp4

Here the audio in input.mp4 is the BGM and input.mp3 is the narration. The BGM gets quieter whenever the narration is speaking. apad fills the time after the narration ends with silence. Without it, a narration shorter than the BGM cuts the BGM off where the narration ends. The narration itself is not in the output: it only controls the ducking. The amix example further down puts it back in.

Parameter Reference

Parameter Description Default Recommended range
threshold Sidechain level above which compression starts 0.125 0.01–0.05
ratio Compression ratio (e.g. 4 = 4:1) 2 3–8
attack Time before the compressor kicks in (ms) 20 100–300
release Time to return to the original level (ms) 250 500–2000
makeup Gain applied after compression (a factor from 1 to 64, not dB) 1 —
knee Knee width (softness of the threshold) 2.82843 —

Tuning threshold

Set threshold to suit how loud your narration is. A smaller value makes the ducking react to quieter speech. This example raises it to 0.05, so the narration has to be louder before the BGM ducks:

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.05:ratio=4:attack=200:release=1000[ducked]" \
  -map 0:v -map "[ducked]" -c:v copy output.mp4

Shaping a Natural Ducking Curve

attack sets how quickly the BGM drops when speech starts. release sets how quickly it comes back after speech stops.

Slow, natural ducking

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=6:attack=300:release=1500[ducked]" \
  -map 0:v -map "[ducked]" -c:v copy output_natural.mp4

Fast, tight ducking

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=8:attack=50:release=500[ducked]" \
  -map 0:v -map "[ducked]" -c:v copy output_tight.mp4

Pre-Attenuate the BGM Before Ducking

If the BGM is too loud even without ducking, lower it with volume first, then duck it:

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]volume=0.5,aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,apad[voice];[bg][voice]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked]" \
  -map 0:v -map "[ducked]" -c:v copy output.mp4

Mix Narration Back In After Ducking

For a final mix with both the ducked BGM and the narration, duck the BGM first and then mix the narration in with amix:

ffmpeg -i input.mp4 -i input.mp3 \
  -filter_complex "[0:a]aformat=fltp:44100:stereo[bg];[1:a]aformat=fltp:44100:stereo,asplit=2[sc][voice2];[sc]apad[voice1];[bg][voice1]sidechaincompress=threshold=0.02:ratio=4:attack=200:release=1000[ducked];[ducked][voice2]amix=inputs=2:duration=longest[out]" \
  -map 0:v -map "[out]" -c:v copy output_mixed.mp4

By default amix turns each input down to half (1/N for N inputs), so the narration and the BGM come out about 6 dB quieter. Add normalize=0 to amix to keep their levels, and watch for clipping.

Why asplit is needed here

The narration is used twice: once as the sidechain that drives the compressor, and once as an input to amix. A label in a filtergraph can be used as an input only once, so asplit=2 makes two copies of it.

FFmpeg 8.x inserts that split for you, so the missing asplit goes unnoticed there. FFmpeg 6.1, the version Ubuntu 24.04 ships, does not, and fails with:

Invalid stream specifier: voice.
Stream specifier 'voice' in filtergraph description ... matches no streams.

With asplit written out, the command works on every version.

Notes

The two inputs of sidechaincompress can have different sample rates and channel counts; FFmpeg converts them automatically. To choose the shared format yourself, add aformat:

Example: aformat=fltp:44100:stereo

Frequently Asked Questions

What is sidechain compression in plain English?

It is a compressor that turns one signal down whenever a different signal gets loud. The usual use is lowering BGM under a voice-over. In FFmpeg, that is sidechaincompress.

What threshold and ratio should I start with?

A threshold of 0.05 (about -26 dB) and a ratio of 8:1 give clearly audible ducking for podcasts. Set attack to 200 ms and release to 1000 ms, as in the basic command. The BGM then stays down between words and comes back up naturally between sentences.

Is ducking better than just lowering BGM volume?

For narrated content, yes. A fixed volume cut also makes the music quiet between sentences. Ducking lowers it only during speech, so the music stays full the rest of the time.

Can I duck two BGM tracks with one voice?

Yes. Run each BGM through its own sidechaincompress with the same voice as the sidechain input, then mix the two ducked outputs together.

Why does my output have audible pumping?

Pumping (the BGM audibly rising and falling) means attack and release are too fast. Lengthen them to the ranges in the parameter table: attack 100–300 ms, release 500–2000 ms. Lowering the ratio to 4:1 also softens the effect.


Primary source: ffmpeg.org/ffmpeg-filters.html#sidechaincompress