Contents

Digital Communication
Other Subjects
Section Progress100%

9 of 9 articles

Differential Entropy

Entropy of continuous random variables.

Darshan N
Updated: 19 March 2026
5 min read

In information theory, entropy measures the average uncertainty in a random variable. For discrete sources, this is straightforward to compute using probability mass functions. However, when a signal can take any value in a continuous range, such as an analog voltage or a Gaussian noise sample, the discrete entropy formula breaks down and a new formulation called differential entropy becomes necessary. Understanding differential entropy is essential for analyzing continuous-alphabet channels, rate-distortion theory, and the capacity of Gaussian channels.

Discrete Entropy vs Differential EntropyDiscrete Entropy H(X)Source: {x1, x2, x3, x4}x1x2x3x4H(X) = -sum p(xi) log p(xi)Finite sum over symbolsAlways non-negativeUnit: bits or natsDifferential Entropy h(X)Source: continuous range (a, b)f(x) continuous densityh(X) = -integral f(x) log f(x) dxIntegral over supportCAN be negativeUnit: bits or nats
Figure 1: Discrete entropy uses a summation over symbols; differential entropy uses integration over a continuous PDF.

Core Concept Explanation

Discrete entropy is defined as H(X) = -sum p(xi) log p(xi) over all symbols. When the variable X is continuous, there is no probability mass at any single point and the concept of p(xi) does not apply. Instead, we use the probability density function (PDF) f(x) to describe how probability is distributed. The differential entropy h(X) is defined by replacing the sum with an integral over the support of f(x).

The formal definition is h(X) = -integral f(x) log f(x) dx, where the integral is taken over the entire support of the distribution. The base of the logarithm determines the unit: base 2 gives bits, natural log gives nats. One critical distinction from discrete entropy is that differential entropy can be negative, which is initially surprising but is mathematically consistent because f(x) can exceed 1 for continuous distributions.

Differential entropy is not an absolute measure of uncertainty in the same way discrete entropy is. It is best understood as a relative quantity and it becomes meaningful primarily when comparing two continuous distributions or when used inside formulas such as mutual information I(X;Y) = h(X) - h(X|Y) for continuous random variables.

Mathematical Expression

For a continuous random variable X with PDF f(x), the differential entropy is defined as:

h(X) = -integral_{-inf}^{inf} f(x) log f(x) dx

For a Gaussian random variable X ~ N(mu, sigma^2), the PDF is f(x) = (1 / sqrt(2 pi sigma^2)) exp(-(x-mu)^2 / 2 sigma^2). Substituting into the integral and evaluating gives the closed form result: h(X) = (1/2) log(2 pi e sigma^2). This result is extremely important. It shows that the differential entropy of a Gaussian depends only on the variance sigma^2 and not on the mean mu. Among all distributions with a given variance, the Gaussian has the maximum differential entropy. This property is used directly in deriving the capacity of the Gaussian channel.

For a uniform distribution X ~ Uniform(a, b), the PDF is f(x) = 1/(b-a) for x in [a,b]. Substituting gives h(X) = log(b-a). If b-a < 1, this is negative, confirming that differential entropy can be negative. This occurs because a very narrow uniform distribution concentrates probability into a small region, making the density values large, and the log of a large density is positive, so -f(x) log f(x) becomes negative.

Practical Understanding

In digital communications, the source signal before sampling is often modeled as a continuous random variable. The differential entropy of this source determines the theoretical minimum bit rate required to represent it digitally with a given distortion level, which is the domain of rate-distortion theory. A source with higher differential entropy is harder to compress losslessly and requires more bits per sample.

The Gaussian assumption is widely used because thermal noise in electronic circuits follows a Gaussian distribution. The differential entropy formula h = (1/2) log(2 pi e sigma^2) directly connects the noise power sigma^2 to its information-theoretic uncertainty. When you increase noise power, you increase h(X), meaning the output of the channel carries more noise uncertainty, reducing the useful information that can be recovered.

Conditional differential entropy h(X|Y) = -integral integral f(x,y) log f(x|y) dx dy measures the remaining uncertainty in X after observing Y. Mutual information for continuous variables is then I(X;Y) = h(X) - h(X|Y) = h(Y) - h(Y|X). This quantity is always non-negative and equals zero only when X and Y are statistically independent. It forms the basis of channel capacity calculations.

Example
Given:
X ~ Gaussian with mean mu = 0, variance sigma^2 = 4 (sigma = 2 V)
Logarithm base: natural log (nats)

Why this formula applies:
For a Gaussian random variable, the differential entropy has a closed-form result
derived by evaluating -integral f(x) ln f(x) dx analytically.

Formula:
h(X) = (1/2) ln(2 * pi * e * sigma^2)

Substitution:
h(X) = (1/2) ln(2 * pi * e * 4)
     = (1/2) ln(2 * 3.14159 * 2.71828 * 4)
     = (1/2) ln(68.33)

Calculation:
ln(68.33) = 4.224
h(X) = (1/2) * 4.224
     = 2.112 nats

Final Answer: h(X) = 2.112 nats
Equivalently in bits: 2.112 / ln(2) = 2.112 / 0.693 = 3.05 bits
Exam Tip: In GATE, the maximum differential entropy for a fixed variance sigma^2 is always achieved by the Gaussian distribution and equals (1/2) log(2 pi e sigma^2). Do not confuse this with discrete entropy, which is always non-negative. Differential entropy of a uniform distribution over an interval shorter than 1 unit is negative.

Differential Entropy of Common Distributions

Differential Entropy: Key Distributions and PropertiesGaussian N(mu, sigma^2)h = (1/2) log(2 pi e sigma^2)Maximum entropy for fixed varianceUniform [a, b]f(x) = 1/(b-a)h = log(b - a)Negative if (b-a) < 1Exponential (lambda)h = 1 - log(lambda)Can be negative if lambda > eKey Properties of Differential Entropy1. h(X) can be negative unlike discrete H(X)2. h(aX) = h(X) + log|a| (scaling changes entropy)3. Gaussian maximizes h(X) among all distributions with the same variance4. Mutual information I(X;Y) = h(X) - h(X|Y) is always non-negative5. Units: bits (log base 2) or nats (natural log)
Figure 2: Differential entropy formulas for the three most common continuous distributions and their key behavioral properties.
  • For Gaussian X with variance sigma^2, h(X) = (1/2) log(2 pi e sigma^2). Increasing variance increases entropy, reflecting greater spread.
  • For uniform X over [a,b], h(X) = log(b-a). This is negative when (b-a) < 1, which happens for very narrow intervals.
  • Scaling property: h(aX) = h(X) + log|a|. Linear transformations shift differential entropy by a logarithm of the scaling factor.
  • The Gaussian distribution is the maximum entropy distribution for a fixed variance. This is why Gaussian noise represents the worst-case noise in channel capacity analysis.
  • Mutual information I(X;Y) = h(X) - h(X|Y) is always non-negative, preserving the same interpretation as in the discrete case.

Quick Revision

  • Differential entropy h(X) = -integral f(x) log f(x) dx generalizes discrete entropy to continuous distributions.
  • Unlike discrete entropy, h(X) can be negative. This is mathematically valid for narrow distributions where f(x) > 1.
  • Gaussian formula: h(X) = (1/2) log(2 pi e sigma^2). Depends only on variance, not mean.
  • Uniform formula: h(X) = log(b-a). Negative when the interval width is less than 1.
  • Gaussian distribution achieves maximum differential entropy among all distributions with the same variance. This is the maximum entropy theorem for continuous variables.
  • Mutual information for continuous variables: I(X;Y) = h(X) - h(X|Y), always non-negative.
  • Exam trap: Differential entropy is not directly comparable to discrete entropy. They have different interpretations and differential entropy can be negative.

Differential Entropy Quiz

Test your grasp of continuous random variable entropy and its properties.

Question 1 of 3

Q1.The differential entropy of a Gaussian random variable X with variance sigma^2 is given by h(X) = (1/2) log(2*pi*e*sigma^2). If the variance is doubled, what is the change in differential entropy?