Before you raise, lower or normalise the volume of a file, check its current level. volumedetect reads the whole audio track and reports the peak level and the mean level in dBFS, plus a count of the loudest samples. The numbers show whether the recording is too quiet or close to clipping, and how much headroom is left (how far the peak is below the maximum level).
Basic command
ffmpeg -i input.mp3 -af volumedetect -f null /dev/null
-f null /dev/null throws the output away, so no file is written. The results are printed to stderr (the terminal) when the run ends. On Windows you can use NUL in place of /dev/null; -f null - works on every system.
Reading the output
[Parsed_volumedetect_0 @ 0x...] n_samples: 2646000
[Parsed_volumedetect_0 @ 0x...] mean_volume: -18.3 dB
[Parsed_volumedetect_0 @ 0x...] max_volume: -0.5 dB
[Parsed_volumedetect_0 @ 0x...] histogram_0db: 26
[Parsed_volumedetect_0 @ 0x...] histogram_1db: 98
[Parsed_volumedetect_0 @ 0x...] histogram_2db: 496
[Parsed_volumedetect_0 @ 0x...] histogram_3db: 1830
[Parsed_volumedetect_0 @ 0x...] histogram_4db: 4410
| Field | Meaning |
|---|---|
n_samples |
Number of samples analysed (all channels) |
mean_volume |
Average (RMS) level of the whole file, in dBFS |
max_volume |
Level of the loudest single sample, in dBFS |
histogram_Xdb |
Number of samples in the 1 dB band from −X dB to −(X+1) dB. The lines start at the loudest band and stop once the total so far reaches 1/1000 of all samples |
What dBFS means
dBFS (decibels relative to full scale) is the level unit for digital audio.
0 dBFS= the largest value the format can hold; anything louder is clipped-6 dBFS≈ half the amplitude of full scale- Closer to 0 is louder; every value is zero or negative
Checking a video’s audio
ffmpeg -i input.mp4 -af volumedetect -f null /dev/null
The same command works on video files. If the file has several audio tracks, choose one with -map 0:a:1 (the second track). Add -vn so the video is not decoded; on long files this makes the scan much faster:
ffmpeg -i input.mp4 -vn -af volumedetect -f null /dev/null
Using the numbers before normalising
If max_volume is −0.5 dB, you can raise the level by 0.5 dB before the peak reaches 0 dBFS. For peak normalisation, apply a gain equal to max_volume with the sign flipped:
1. Measure
ffmpeg -i input.mp3 -af volumedetect -f null /dev/null
2. Apply the gain with the volume filter
ffmpeg -i input.mp3 -af "volume=+0.5dB" output.mp3
Peak normalisation only makes sure the file doesn’t clip. It says nothing about how loud the file sounds. To match a loudness target such as −16 LUFS (podcasts) or −14 LUFS (streaming platforms), use loudnorm, which measures loudness as defined in ITU-R BS.1770. See Loudness normalisation.
Common checks
Is the recording too quiet?
If mean_volume is below about −30 dB, the recording level was probably low. To print only the two lines you need:
ffmpeg -i input.mp3 -af volumedetect -f null /dev/null 2>&1 | grep -E "mean_volume|max_volume"
2>&1 sends stderr into the pipe, so grep can filter it.
Is it clipping?
max_volume is rounded to 0.1 dB, so 0.0 dB or -0.0 dB means the peak is at or just below full scale. histogram_0db counts every sample within 1 dB of full scale, clipped or not. Neither number alone proves clipping or audible distortion: loud audio that does not clip also gives a large count. Clipped audio gives one too. For example, summing two identical channels with pan=mono|c0=c0+c1 raises the level by 6 dB. On a file that peaked at −0.1 dBFS, this put 310,400 of 441,000 samples (about 70%) into histogram_0db, and the result was badly distorted. Stereo to mono and channel mapping has the measurement and a safe alternative.
How much did a filter change the level?
Run volumedetect before and after the filter. This quickly shows whether a pan, amix or -ac 1 step kept the level, or how much amix lowered it with its default normalisation (about 6 dB for two inputs).
FAQ
mean_volume and loudnorm’s LUFS don’t agree
They measure different things. mean_volume is a plain RMS over all samples. LUFS weights the frequencies to roughly match human hearing (K-weighting) and ignores quiet passages (gating). Two files with the same mean_volume can differ by several LU. Use volumedetect for headroom and clipping, and loudnorm (or ebur128) for how loud the audio sounds.
Can I get the values per channel?
Not from volumedetect alone, because it reports all channels together. Split the channels with channelsplit in -filter_complex and run volumedetect on each branch, or use astats, which prints the peak and RMS of each channel.
Is there a faster way for very long files?
volumedetect has to decode all of the audio, so the decoder sets the speed. For video files, add -vn so the video is not decoded too.
Related articles
- Loudness normalisation —
loudnormtwo-pass procedure - Stereo to mono and channel mapping — measured levels for
-ac 1vspan - Remove silence automatically — pick a threshold from the measured level
- Extract audio — pull the audio track out of a video first