Before you raise, lower or normalise the volume of a file, check its current level. volumedetect reads the whole audio track and reports the peak level and the mean level in dBFS, plus a count of the loudest samples. The numbers show whether the recording is too quiet or close to clipping, and how much headroom is left (how far the peak is below the maximum level).

Basic command

ffmpeg -i input.mp3 -af volumedetect -f null /dev/null

-f null /dev/null throws the output away, so no file is written. The results are printed to stderr (the terminal) when the run ends. On Windows you can use NUL in place of /dev/null; -f null - works on every system.

Reading the output

[Parsed_volumedetect_0 @ 0x...] n_samples: 2646000
[Parsed_volumedetect_0 @ 0x...] mean_volume: -18.3 dB
[Parsed_volumedetect_0 @ 0x...] max_volume: -0.5 dB
[Parsed_volumedetect_0 @ 0x...] histogram_0db: 26
[Parsed_volumedetect_0 @ 0x...] histogram_1db: 98
[Parsed_volumedetect_0 @ 0x...] histogram_2db: 496
[Parsed_volumedetect_0 @ 0x...] histogram_3db: 1830
[Parsed_volumedetect_0 @ 0x...] histogram_4db: 4410
Field Meaning
n_samples Number of samples analysed (all channels)
mean_volume Average (RMS) level of the whole file, in dBFS
max_volume Level of the loudest single sample, in dBFS
histogram_Xdb Number of samples in the 1 dB band from −X dB to −(X+1) dB. The lines start at the loudest band and stop once the total so far reaches 1/1000 of all samples

What dBFS means

dBFS (decibels relative to full scale) is the level unit for digital audio.

  • 0 dBFS = the largest value the format can hold; anything louder is clipped
  • -6 dBFS ≈ half the amplitude of full scale
  • Closer to 0 is louder; every value is zero or negative

Checking a video’s audio

ffmpeg -i input.mp4 -af volumedetect -f null /dev/null

The same command works on video files. If the file has several audio tracks, choose one with -map 0:a:1 (the second track). Add -vn so the video is not decoded; on long files this makes the scan much faster:

ffmpeg -i input.mp4 -vn -af volumedetect -f null /dev/null

Using the numbers before normalising

If max_volume is −0.5 dB, you can raise the level by 0.5 dB before the peak reaches 0 dBFS. For peak normalisation, apply a gain equal to max_volume with the sign flipped:

1. Measure

ffmpeg -i input.mp3 -af volumedetect -f null /dev/null

2. Apply the gain with the volume filter

ffmpeg -i input.mp3 -af "volume=+0.5dB" output.mp3

Peak normalisation only makes sure the file doesn’t clip. It says nothing about how loud the file sounds. To match a loudness target such as −16 LUFS (podcasts) or −14 LUFS (streaming platforms), use loudnorm, which measures loudness as defined in ITU-R BS.1770. See Loudness normalisation.

Common checks

Is the recording too quiet?

If mean_volume is below about −30 dB, the recording level was probably low. To print only the two lines you need:

ffmpeg -i input.mp3 -af volumedetect -f null /dev/null 2>&1 | grep -E "mean_volume|max_volume"

2>&1 sends stderr into the pipe, so grep can filter it.

Is it clipping?

max_volume is rounded to 0.1 dB, so 0.0 dB or -0.0 dB means the peak is at or just below full scale. histogram_0db counts every sample within 1 dB of full scale, clipped or not. Neither number alone proves clipping or audible distortion: loud audio that does not clip also gives a large count. Clipped audio gives one too. For example, summing two identical channels with pan=mono|c0=c0+c1 raises the level by 6 dB. On a file that peaked at −0.1 dBFS, this put 310,400 of 441,000 samples (about 70%) into histogram_0db, and the result was badly distorted. Stereo to mono and channel mapping has the measurement and a safe alternative.

How much did a filter change the level?

Run volumedetect before and after the filter. This quickly shows whether a pan, amix or -ac 1 step kept the level, or how much amix lowered it with its default normalisation (about 6 dB for two inputs).

FAQ

mean_volume and loudnorm’s LUFS don’t agree

They measure different things. mean_volume is a plain RMS over all samples. LUFS weights the frequencies to roughly match human hearing (K-weighting) and ignores quiet passages (gating). Two files with the same mean_volume can differ by several LU. Use volumedetect for headroom and clipping, and loudnorm (or ebur128) for how loud the audio sounds.

Can I get the values per channel?

Not from volumedetect alone, because it reports all channels together. Split the channels with channelsplit in -filter_complex and run volumedetect on each branch, or use astats, which prints the peak and RMS of each channel.

Is there a faster way for very long files?

volumedetect has to decode all of the audio, so the decoder sets the speed. For video files, add -vn so the video is not decoded too.