Fixed vs Floating Point
Quantization noise, dynamic range.
In digital signal processing systems, every number must be represented in a finite number of bits. The choice of fixed-point versus floating-point representation directly affects system cost, processing speed, power consumption, and the quality of output signals. Understanding this tradeoff is essential for hardware design decisions and is a recurring topic in GATE and competitive examinations.
Fixed-Point Representation
In fixed-point representation, the binary point (radix point) is fixed at a predetermined position within the word. The programmer or hardware designer decides how many bits represent the integer part and how many represent the fractional part. For a Q-format number written as Q(m.n), there are m integer bits and n fractional bits, giving a total wordlength of m+n+1 bits including the sign bit. The smallest representable nonzero value is 2^(-n), which is the quantization step size or resolution.
The representable range for an unsigned Q(m.n) number is 0 to (2^m - 2^(-n)). The fixed nature of the radix point means that if the true signal value exceeds the range, overflow occurs. If the signal is much smaller than the range, the relative quantization error (noise) is large. This mismatch between signal level and the fixed range is the fundamental limitation of fixed-point arithmetic.
Floating-Point Representation
Floating-point numbers follow the IEEE 754 standard and represent a value as: (-1)^S x 1.M x 2^(E-bias), where S is the sign bit, E is the biased exponent field, and M is the mantissa (fractional part of the significand). The exponent effectively shifts the radix point, giving the format a very wide dynamic range. For single precision (32-bit), the dynamic range spans approximately 10^-38 to 10^38. The mantissa provides about 7 decimal digits of precision regardless of magnitude.
The floating-point unit requires dedicated hardware (FPU) with normalization, rounding, and exponent alignment logic. This makes it significantly more expensive in silicon area and power consumption than fixed-point hardware. However, it eliminates the overflow problem and greatly reduces the programmer's burden of managing scaling.
Quantization Noise
Quantization noise arises because the continuous (or high-precision) signal is rounded to the nearest representable value. For uniform quantization with step size q = 2^(-n) (for n-bit fixed-point fraction), the quantization error e lies in the range [-q/2, q/2]. Assuming the error is uniformly distributed and uncorrelated (white noise model), the quantization noise power is: P_q = q^2 / 12 = 2^(-2n) / 12.
The Signal-to-Quantization-Noise Ratio (SQNR) for a full-scale sinusoidal signal with n-bit uniform quantization is approximately: SQNR = 6.02n + 1.76 dB. This means each additional bit of resolution improves SQNR by approximately 6 dB. Floating-point avoids large SQNR variations because the step size scales with the signal magnitude, keeping relative error roughly constant.
Dynamic Range
Dynamic range is the ratio of the largest to the smallest representable nonzero value. For an n-bit fixed-point format, the dynamic range is 2^n, equivalent to 6n dB. A 16-bit fixed-point system has approximately 96 dB of dynamic range. For IEEE 754 single precision, the dynamic range is about 2^254 approximately 1.5 x 10^76, which is effectively unlimited for most engineering applications. Floating-point achieves this because the exponent can shift the representable range across many orders of magnitude.
Given:
Fixed-point format: Q8.8 (8 integer bits, 8 fraction bits, 17 bits total with sign)
Signal: a sinusoid with amplitude 50.0
Why this formula applies:
SQNR formula for uniform quantization with n fraction bits.
Formula:
SQNR = 6.02 x n + 1.76 dB (n = number of fraction bits)
Quantization step q = 2^(-n)
Substitution:
n = 8 fraction bits
SQNR = 6.02 x 8 + 1.76
Calculation:
SQNR = 48.16 + 1.76 = 49.92 dB
q = 2^(-8) = 0.00390625
Final Answer: SQNR is approximately 49.92 dB with a quantization step size of approximately 0.0039. Increasing to Q8.16 would give SQNR of approximately 98.08 dB.Exam Tip: The SQNR formula SQNR = 6.02n + 1.76 dB applies to full-scale sinusoidal input with n-bit uniform quantization. In GATE, n refers to fraction bits in fixed-point or total significant bits. Each additional bit adds approximately 6 dB. Floating-point does not eliminate quantization noise; it only makes the relative noise roughly uniform across the dynamic range.
- Fixed-point stores numbers with a radix point fixed in hardware; quantization step size is 2^(-n) for n fraction bits.
- Quantization noise power for uniform quantization is q^2/12 where q is the step size.
- Floating-point uses sign, exponent, and mantissa fields; the exponent shifts the radix point dynamically.
- Dynamic range of fixed-point is 6n dB for an n-bit word; floating-point has essentially unlimited dynamic range for engineering use.
- Fixed-point is preferred for low-cost, low-power embedded DSP; floating-point is used when wide dynamic range or ease of programming is required.
Quick Revision
- Fixed-point: radix point is fixed; range and resolution are set at design time by the Q format chosen.
- SQNR = 6.02n + 1.76 dB for n-bit uniform quantization with full-scale sinusoidal input.
- Each additional bit adds approximately 6 dB to SQNR.
- Dynamic range of n-bit fixed-point word: 20 log10(2^n) = 6.02n dB.
- Floating-point IEEE 754: value = (-1)^S x 1.M x 2^(E-127) for single precision.
- Trap: More bits always reduce quantization noise in fixed-point but do not eliminate overflow if the signal exceeds range.
- Floating-point does not eliminate quantization noise; it only keeps relative noise approximately constant across the dynamic range.
Fixed vs Floating Quiz
Test your understanding of quantization noise and dynamic range in fixed and floating-point DSP systems.
Q1.For a B-bit fixed-point uniform quantizer, the signal-to-quantization noise ratio (SQNR) for a full-scale sinusoidal input is approximately:
Related Articles
Wavelet Transform
Short-time Fourier transform limitation, wavelets.
5 min read
Decimation
Downsampling, anti-aliasing filter.
12 min read
Interpolation
Upsampling, anti-imaging filter.
5 min read
Noise Cancellation
Application of adaptive filters in removing noise.
5 min read
Image Processing Basics
2D convolution, edge detection.
9 min read