Differential Entropy
Entropy of continuous random variables.
In information theory, entropy measures the average uncertainty in a random variable. For discrete sources, this is straightforward to compute using probability mass functions. However, when a signal can take any value in a continuous range, such as an analog voltage or a Gaussian noise sample, the discrete entropy formula breaks down and a new formulation called differential entropy becomes necessary. Understanding differential entropy is essential for analyzing continuous-alphabet channels, rate-distortion theory, and the capacity of Gaussian channels.
Core Concept Explanation
Discrete entropy is defined as H(X) = -sum p(xi) log p(xi) over all symbols. When the variable X is continuous, there is no probability mass at any single point and the concept of p(xi) does not apply. Instead, we use the probability density function (PDF) f(x) to describe how probability is distributed. The differential entropy h(X) is defined by replacing the sum with an integral over the support of f(x).
The formal definition is h(X) = -integral f(x) log f(x) dx, where the integral is taken over the entire support of the distribution. The base of the logarithm determines the unit: base 2 gives bits, natural log gives nats. One critical distinction from discrete entropy is that differential entropy can be negative, which is initially surprising but is mathematically consistent because f(x) can exceed 1 for continuous distributions.
Differential entropy is not an absolute measure of uncertainty in the same way discrete entropy is. It is best understood as a relative quantity and it becomes meaningful primarily when comparing two continuous distributions or when used inside formulas such as mutual information I(X;Y) = h(X) - h(X|Y) for continuous random variables.
Mathematical Expression
For a continuous random variable X with PDF f(x), the differential entropy is defined as:
h(X) = -integral_{-inf}^{inf} f(x) log f(x) dx
For a Gaussian random variable X ~ N(mu, sigma^2), the PDF is f(x) = (1 / sqrt(2 pi sigma^2)) exp(-(x-mu)^2 / 2 sigma^2). Substituting into the integral and evaluating gives the closed form result: h(X) = (1/2) log(2 pi e sigma^2). This result is extremely important. It shows that the differential entropy of a Gaussian depends only on the variance sigma^2 and not on the mean mu. Among all distributions with a given variance, the Gaussian has the maximum differential entropy. This property is used directly in deriving the capacity of the Gaussian channel.
For a uniform distribution X ~ Uniform(a, b), the PDF is f(x) = 1/(b-a) for x in [a,b]. Substituting gives h(X) = log(b-a). If b-a < 1, this is negative, confirming that differential entropy can be negative. This occurs because a very narrow uniform distribution concentrates probability into a small region, making the density values large, and the log of a large density is positive, so -f(x) log f(x) becomes negative.
Practical Understanding
In digital communications, the source signal before sampling is often modeled as a continuous random variable. The differential entropy of this source determines the theoretical minimum bit rate required to represent it digitally with a given distortion level, which is the domain of rate-distortion theory. A source with higher differential entropy is harder to compress losslessly and requires more bits per sample.
The Gaussian assumption is widely used because thermal noise in electronic circuits follows a Gaussian distribution. The differential entropy formula h = (1/2) log(2 pi e sigma^2) directly connects the noise power sigma^2 to its information-theoretic uncertainty. When you increase noise power, you increase h(X), meaning the output of the channel carries more noise uncertainty, reducing the useful information that can be recovered.
Conditional differential entropy h(X|Y) = -integral integral f(x,y) log f(x|y) dx dy measures the remaining uncertainty in X after observing Y. Mutual information for continuous variables is then I(X;Y) = h(X) - h(X|Y) = h(Y) - h(Y|X). This quantity is always non-negative and equals zero only when X and Y are statistically independent. It forms the basis of channel capacity calculations.
Given:
X ~ Gaussian with mean mu = 0, variance sigma^2 = 4 (sigma = 2 V)
Logarithm base: natural log (nats)
Why this formula applies:
For a Gaussian random variable, the differential entropy has a closed-form result
derived by evaluating -integral f(x) ln f(x) dx analytically.
Formula:
h(X) = (1/2) ln(2 * pi * e * sigma^2)
Substitution:
h(X) = (1/2) ln(2 * pi * e * 4)
= (1/2) ln(2 * 3.14159 * 2.71828 * 4)
= (1/2) ln(68.33)
Calculation:
ln(68.33) = 4.224
h(X) = (1/2) * 4.224
= 2.112 nats
Final Answer: h(X) = 2.112 nats
Equivalently in bits: 2.112 / ln(2) = 2.112 / 0.693 = 3.05 bitsExam Tip: In GATE, the maximum differential entropy for a fixed variance sigma^2 is always achieved by the Gaussian distribution and equals (1/2) log(2 pi e sigma^2). Do not confuse this with discrete entropy, which is always non-negative. Differential entropy of a uniform distribution over an interval shorter than 1 unit is negative.
Differential Entropy of Common Distributions
- For Gaussian X with variance sigma^2, h(X) = (1/2) log(2 pi e sigma^2). Increasing variance increases entropy, reflecting greater spread.
- For uniform X over [a,b], h(X) = log(b-a). This is negative when (b-a) < 1, which happens for very narrow intervals.
- Scaling property: h(aX) = h(X) + log|a|. Linear transformations shift differential entropy by a logarithm of the scaling factor.
- The Gaussian distribution is the maximum entropy distribution for a fixed variance. This is why Gaussian noise represents the worst-case noise in channel capacity analysis.
- Mutual information I(X;Y) = h(X) - h(X|Y) is always non-negative, preserving the same interpretation as in the discrete case.
Quick Revision
- Differential entropy h(X) = -integral f(x) log f(x) dx generalizes discrete entropy to continuous distributions.
- Unlike discrete entropy, h(X) can be negative. This is mathematically valid for narrow distributions where f(x) > 1.
- Gaussian formula: h(X) = (1/2) log(2 pi e sigma^2). Depends only on variance, not mean.
- Uniform formula: h(X) = log(b-a). Negative when the interval width is less than 1.
- Gaussian distribution achieves maximum differential entropy among all distributions with the same variance. This is the maximum entropy theorem for continuous variables.
- Mutual information for continuous variables: I(X;Y) = h(X) - h(X|Y), always non-negative.
- Exam trap: Differential entropy is not directly comparable to discrete entropy. They have different interpretations and differential entropy can be negative.
Differential Entropy Quiz
Test your grasp of continuous random variable entropy and its properties.
Q1.The differential entropy of a Gaussian random variable X with variance sigma^2 is given by h(X) = (1/2) log(2*pi*e*sigma^2). If the variance is doubled, what is the change in differential entropy?
Related Articles
Source Entropy
Entropy H = -sum P*log2(P), maximum entropy conditions.
7 min read
Joint and Conditional Entropy
H(X,Y), H(Y|X), chain rule for entropy.
5 min read
Information Theory Basics
Entropy, information content I = -log2(P), Shannon.
6 min read
Differential PSK
DPSK modulation and demodulation, non-coherent advantage.
11 min read
Mutual Information
I(X;Y) = H(X) - H(X|Y), channel relationship.
9 min read