Mutual Information
I(X;Y) = H(X) - H(X|Y), channel relationship.
Mutual information quantifies exactly how much knowing the output of a channel reduces uncertainty about its input. It is the central measure connecting source statistics to channel behavior, and it appears directly in the definition of channel capacity. For GATE and university exams, mutual information bridges entropy, conditional entropy, and capacity into one unified framework.
Core Concept Explanation
Consider a communication channel with discrete input random variable X and output random variable Y. Before the receiver observes Y, its uncertainty about X is measured by the entropy H(X). After observing Y, the remaining uncertainty about X is the conditional entropy H(X|Y). The reduction in uncertainty that Y provides about X is precisely the mutual information I(X;Y).
Physically, mutual information measures how much the channel output reveals about the channel input. If the channel is noiseless, observing Y tells you X perfectly, so I(X;Y) = H(X). If the channel is pure noise and Y is statistically independent of X, then H(X|Y) = H(X), giving I(X;Y) = 0. Real channels fall between these extremes.
Mutual information is always non-negative and symmetric: I(X;Y) = I(Y;X). This symmetry is non-obvious intuitively but follows from the joint entropy relation. It says that knowing Y reduces uncertainty about X by exactly as much as knowing X reduces uncertainty about Y.
Mathematical Expression
The primary definition using conditional entropy is:
I(X;Y) = H(X) - H(X|Y)
Equivalent forms that appear in exams include:
- I(X;Y) = H(Y) - H(Y|X) — symmetric form using output entropy
- I(X;Y) = H(X) + H(Y) - H(X,Y) — using joint entropy H(X,Y)
- I(X;Y) = sum over x,y of P(x,y) log [ P(x,y) / (P(x)P(y)) ] — KL divergence form
The KL divergence form shows that mutual information measures how far the joint distribution P(x,y) is from the product of marginals P(x)P(y). Independence means P(x,y)=P(x)P(y), making the log term zero and I(X;Y)=0. Strong statistical dependence drives the divergence up, increasing I(X;Y).
The conditional entropy H(X|Y) is computed as H(X|Y) = sum over y of P(y) H(X|Y=y), where each H(X|Y=y) = minus sum over x of P(x|y) log2 P(x|y). This weighted average represents the residual uncertainty about X after averaging over all possible observations of Y.
Practical Understanding
In a practical binary channel, noise causes some input bits to flip at the output. The mutual information in this case equals 1 minus the entropy of the error probability, representing how many reliable bits per channel use the system actually conveys. This directly connects to the concept of channel capacity, which is the maximum of I(X;Y) over all possible input distributions P(x).
For a Gaussian channel under power constraint, the capacity-achieving distribution is Gaussian, and I(X;Y) achieves the Shannon capacity formula. Mutual information is therefore not just a theoretical measure but the operational quantity that determines the maximum data rate a channel can support reliably.
Engineers use mutual information in practice to evaluate modulation schemes, assess the effectiveness of error-correcting codes, and design iterative decoders. In LDPC and turbo code design, the EXIT chart technique directly plots mutual information transfer between decoder components to predict convergence.
Solved Numerical Example
Consider a binary channel where input X is equally likely: P(X=0)=P(X=1)=0.5. The channel transition probabilities are P(Y=0|X=0)=0.9, P(Y=1|X=0)=0.1, P(Y=0|X=1)=0.1, P(Y=1|X=1)=0.9. We compute I(X;Y) using I(X;Y) = H(Y) - H(Y|X).
Given:
P(X=0) = P(X=1) = 0.5
P(Y=1|X=0) = P(Y=0|X=1) = p = 0.1 (crossover probability)
Why this formula applies:
I(X;Y) = H(Y) - H(Y|X) is easiest when output marginals are computed first.
Formula:
I(X;Y) = H(Y) - H(Y|X)
H(Y|X) = H(p) = -p*log2(p) - (1-p)*log2(1-p) [binary entropy of crossover probability]
Substitution:
P(Y=1) = P(Y=1|X=0)*P(X=0) + P(Y=1|X=1)*P(X=1)
= 0.1*0.5 + 0.9*0.5 = 0.5
H(Y) = -0.5*log2(0.5) - 0.5*log2(0.5) = 1 bit
H(Y|X) = H(0.1) = -0.1*log2(0.1) - 0.9*log2(0.9)
Calculation:
H(0.1) = 0.1*3.3219 + 0.9*0.1520 = 0.3322 + 0.1368 = 0.469 bits
I(X;Y) = 1 - 0.469
Final Answer: I(X;Y) = 0.531 bits per channel useExam Tip: For a BSC with crossover probability p, I(X;Y) = 1 - H(p) when input is uniform. H(Y|X) always equals H(p) regardless of input distribution, but H(Y) is maximized at 1 bit only for equal priors. GATE often tests whether students incorrectly compute H(Y|X) using output marginals instead of transition probabilities.
- I(X;Y) equals the entropy of the input minus the conditional entropy remaining after the output is observed, quantifying exactly what the channel reveals.
- The symmetry I(X;Y) = I(Y;X) holds mathematically via the joint entropy formula, even though the physical interpretation differs.
- Channel capacity is defined as C = max_{P(x)} I(X;Y), meaning we optimize the input distribution to squeeze maximum mutual information from the channel.
- When X and Y are statistically independent, I(X;Y) = 0 and the channel conveys no useful information about the source.
- The KL divergence interpretation shows that mutual information is zero only when the joint distribution completely factors into the product of marginals.
Quick Revision
- I(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) = H(X)+H(Y)-H(X,Y). All forms are equivalent and GATE tests all of them.
- I(X;Y) >= 0 always. It equals zero only when X and Y are independent (useless channel).
- I(X;Y) = I(Y;X): mutual information is symmetric even though H(X|Y) is not equal to H(Y|X) in general.
- For a BSC with crossover p and uniform input: I(X;Y) = 1 - H(p) where H(p) = -p log2 p - (1-p) log2(1-p).
- Channel capacity C = max over all input distributions of I(X;Y).
- Exam trap: H(Y|X) is computed using transition probabilities P(y|x), NOT the output marginals P(y). Confusing these is a common error.
- H(X,Y) = H(X) + H(Y|X) = H(Y) + H(X|Y): chain rule connects joint, marginal, and conditional entropies.
Mutual Information Quiz
Test your command of mutual information definitions, symmetry properties, and channel interpretations.
Q1.Mutual information I(X;Y) is correctly expressed as:
Related Articles
Information Theory Basics
Entropy, information content I = -log2(P), Shannon.
6 min read
BEC Channel
Binary Erasure Channel properties.
5 min read
BSC Channel
Binary Symmetric Channel, capacity calculation.
10 min read
Differential Entropy
Entropy of continuous random variables.
5 min read
Source Entropy
Entropy H = -sum P*log2(P), maximum entropy conditions.
7 min read