Audio Compression

Perceptual coding, MP3 basics.

Darshan N
Updated: 19 March 2026
5 min read

Audio compression is the process of reducing the data rate or storage size of audio signals while maintaining acceptable perceptual quality for human listeners. Without compression, a CD-quality stereo audio signal requires 1.41 Mbps. Techniques such as MP3 reduce this to 128 kbps with minimal perceptual degradation. Understanding how perceptual coding exploits the limitations of human hearing is fundamental to DSP applications in multimedia and communication systems.

Audio Compression: Perceptual Coding and MP3 FrameworkPCM Audio44.1kHz 16-bitFilter Bank /MDCT Sub-bandsPsychoacousticModel (masking)Bit AllocationNoise shapingEntropy CodingHuffman + bitstreamPsychoacoustic Masking EffectsSimultaneous (Frequency) MaskingA loud tone masks nearby quiet tonesat same time. Masking threshold risesnear masker frequency. Quantizationnoise hidden below threshold.Temporal MaskingMasking occurs before (pre-masking ~2ms)and after (post-masking ~200ms) a loudsound. Quantization noise in thesewindows is inaudible to human ear.Bit Rate ComparisonUncompressed PCM stereo 44.1kHz 16-bit:1411 kbpsMP3 high quality (320 kbps):320 kbps (compression ratio ~4.4:1)MP3 standard (128 kbps):128 kbps (compression ratio ~11:1)AAC (96 kbps) perceptually similar to MP3 128 kbps due to better psychoacoustic model
Figure 1: MP3 encoding pipeline and the two psychoacoustic masking effects that enable high compression ratios.

Core Concept: Why Audio Can Be Compressed

The human auditory system does not perceive all frequencies or all time-domain events equally. This perceptual non-uniformity is what audio compression algorithms exploit. There are two primary masking effects. Simultaneous masking (frequency masking) occurs when a loud tone prevents the hearing of nearby quieter tones at the same time instant. The masking threshold rises in the neighborhood of the masker frequency, creating a region where quantization noise can be hidden.

The second effect is temporal masking. After a loud sound occurs, the auditory system remains partially desensitized for approximately 100-200 ms (post-masking). There is also a brief pre-masking effect lasting about 2-5 ms before the loud sound. Audio in these windows can be quantized more coarsely without the distortion being audible.

The absolute threshold of hearing is the minimum sound pressure level a human can detect as a function of frequency. The ear is most sensitive between 1 kHz and 5 kHz. Frequencies below 20 Hz and above 20 kHz are generally inaudible. Any signal component below the absolute threshold can be discarded entirely without perceptual loss.

MDCT and Sub-band Coding

MP3 uses a hybrid filter bank consisting of a 32-band polyphase filter bank followed by a Modified Discrete Cosine Transform (MDCT) applied to each sub-band. The MDCT provides better frequency resolution and uses overlapping windows to avoid blocking artifacts. The MDCT of a signal of length 2N gives N real-valued coefficients, providing 2:1 overlap between consecutive frames while maintaining perfect reconstruction.

The MDCT basis functions are cosines: MDCT[k] = sum_{n=0}^{2N-1} x(n) * cos( pi/N * (n + 1/2 + N/2) * (k + 1/2) ) for k = 0 to N-1. The overlapping windows with 50 percent overlap are enabled by the principle of time-domain aliasing cancellation (TDAC), which guarantees that the aliasing introduced by overlapping analysis windows cancels perfectly during reconstruction.

Mathematical Expression

The bit allocation in perceptual audio coding is driven by the signal-to-mask ratio (SMR). For each sub-band i, the masking threshold T_i is computed from the psychoacoustic model. The number of bits assigned to sub-band i is chosen so that the quantization noise power N_i stays below T_i, ensuring the noise is inaudible. The signal-to-noise ratio for b-bit uniform quantization is: SNR = 6.02 * b + 1.76 dB. Adding one bit improves SNR by approximately 6 dB. This formula directly determines how many bits each sub-band needs.

Practical Understanding

MP3 encoding uses a feedback loop where the psychoacoustic model continuously computes masking thresholds for each frame and feeds this information to the bit allocator. The bit allocator assigns more bits to sub-bands with low SMR (loud signal, low masking) and fewer bits to sub-bands with high SMR. Sub-bands where the signal is entirely below the absolute threshold receive zero bits.

After quantization, the non-zero coefficients are coded using Huffman coding, a lossless variable-length entropy coding scheme where frequently occurring values get short codes. This stage further reduces the bit rate by 20-30 percent beyond quantization. The combination of perceptual quantization and Huffman coding produces the final MP3 bitstream.

Example
Given:
PCM audio: fs = 44100 Hz, 16-bit stereo
Target: MP3 at 128 kbps

Why this formula applies:
Uncompressed bit rate = fs * bits_per_sample * channels
SNR per sub-band: SNR = 6.02 * b + 1.76 dB (for b-bit uniform quantization)
Compression ratio = uncompressed_rate / compressed_rate

Formula:
Uncompressed bit rate = fs x bits x channels
Compression ratio = uncompressed / target_rate

Substitution:
Uncompressed = 44100 x 16 x 2 = 1,411,200 bps = 1411 kbps
Compression ratio = 1411 / 128

Calculation:
Compression ratio = 11.02 : 1

For a sub-band assigned b = 6 bits:
SNR = 6.02 x 6 + 1.76 = 36.12 + 1.76 = 37.88 dB

Final Answer:
Compression ratio = approximately 11:1
Sub-band SNR with 6-bit quantization = 37.88 dB
Exam Tip: SNR for b-bit uniform quantizer = 6.02b + 1.76 dB. Each additional bit adds approximately 6 dB SNR. MDCT length 2N produces N real coefficients. MP3 uses 32 sub-bands with MDCT; AAC uses 1024-point MDCT directly.

Mechanism: Quantization Noise vs. Masking Threshold

Perceptual Masking Threshold vs. Quantization Noise in Sub-bandsFrequency (Sub-bands 1 to 32)Level (dB)0+20+40Signal Spectrum S(f)Masking Threshold T(f)Bars: Quantization noise per sub-band. Noise kept BELOW masking threshold = inaudible.More bits allocated to sub-bands where noise must stay well below threshold (low SMR bands).
Figure 2: Quantization noise in each sub-band is shaped to remain below the masking threshold, making it perceptually inaudible.

Mechanism Summary

  • Audio compression exploits frequency masking (loud tone hides nearby quiet tones) and temporal masking (pre-masking ~2 ms, post-masking ~200 ms) to discard or coarsely quantize inaudible components.
  • MP3 uses a 32-band polyphase filter bank combined with MDCT. MDCT of a 2N-sample block produces N real coefficients with 50 percent overlap and perfect reconstruction via TDAC.
  • Psychoacoustic model computes the masking threshold T_i per sub-band. Bit allocation assigns b_i bits so that quantization noise power stays below T_i (noise shaping principle).
  • SNR of a uniform b-bit quantizer: SNR = 6.02b + 1.76 dB. Each additional bit adds roughly 6 dB SNR, directly determining minimum bits required per sub-band.
  • Huffman coding applied to quantized MDCT coefficients provides additional lossless compression. Frequently occurring zero-run patterns in MDCT coefficients are coded efficiently.

Quick Revision

  • Uncompressed PCM stereo 44.1 kHz 16-bit = 1411 kbps. MP3 at 128 kbps gives approximately 11:1 compression ratio.
  • Two masking types: simultaneous (frequency domain, same time instant) and temporal (pre ~2 ms, post ~200 ms around loud sound).
  • SNR formula for b-bit uniform quantizer: SNR = 6.02b + 1.76 dB. One extra bit adds ~6 dB.
  • MDCT: 2N input samples, N real output coefficients, 50 percent overlap between frames, perfect reconstruction via TDAC.
  • MP3 pipeline: PCM input, 32-band filter bank, MDCT, psychoacoustic model, bit allocator, quantizer, Huffman coder, bitstream output.
  • Exam trap: MP3 uses MDCT (not DFT/DCT directly). AAC uses a 1024-point MDCT without the polyphase filter bank, giving better quality at same bit rate.
  • Absolute threshold of hearing: ear most sensitive 1-5 kHz. Components below this threshold are discarded with zero bits allocated.

Audio Compression Quiz

Test your understanding of perceptual audio coding and MP3 compression principles.

Question 1 of 3

Q1.In MP3 (MPEG-1 Layer III) encoding, the Modified Discrete Cosine Transform (MDCT) is applied to overlapping blocks. The overlap is necessary to: