Contents

Digital Communication
Other Subjects
Section Progress44%

4 of 9 articles

Mutual Information

I(X;Y) = H(X) - H(X|Y), channel relationship.

Darshan N
Updated: 19 March 2026
9 min read

Mutual information quantifies exactly how much knowing the output of a channel reduces uncertainty about its input. It is the central measure connecting source statistics to channel behavior, and it appears directly in the definition of channel capacity. For GATE and university exams, mutual information bridges entropy, conditional entropy, and capacity into one unified framework.

H(X)H(Y)I(X;Y)MutualInformationH(X|Y)(only in X)H(Y|X)(only in Y)Source entropyChannel output entropyVenn Diagram of EntropiesI(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) = H(X)+H(Y)-H(X,Y)
Figure 1: Mutual information I(X;Y) as the shared entropy between source X and channel output Y

Core Concept Explanation

Consider a communication channel with discrete input random variable X and output random variable Y. Before the receiver observes Y, its uncertainty about X is measured by the entropy H(X). After observing Y, the remaining uncertainty about X is the conditional entropy H(X|Y). The reduction in uncertainty that Y provides about X is precisely the mutual information I(X;Y).

Physically, mutual information measures how much the channel output reveals about the channel input. If the channel is noiseless, observing Y tells you X perfectly, so I(X;Y) = H(X). If the channel is pure noise and Y is statistically independent of X, then H(X|Y) = H(X), giving I(X;Y) = 0. Real channels fall between these extremes.

Mutual information is always non-negative and symmetric: I(X;Y) = I(Y;X). This symmetry is non-obvious intuitively but follows from the joint entropy relation. It says that knowing Y reduces uncertainty about X by exactly as much as knowing X reduces uncertainty about Y.

Mathematical Expression

The primary definition using conditional entropy is:

I(X;Y) = H(X) - H(X|Y)

Equivalent forms that appear in exams include:

  • I(X;Y) = H(Y) - H(Y|X) — symmetric form using output entropy
  • I(X;Y) = H(X) + H(Y) - H(X,Y) — using joint entropy H(X,Y)
  • I(X;Y) = sum over x,y of P(x,y) log [ P(x,y) / (P(x)P(y)) ] — KL divergence form

The KL divergence form shows that mutual information measures how far the joint distribution P(x,y) is from the product of marginals P(x)P(y). Independence means P(x,y)=P(x)P(y), making the log term zero and I(X;Y)=0. Strong statistical dependence drives the divergence up, increasing I(X;Y).

The conditional entropy H(X|Y) is computed as H(X|Y) = sum over y of P(y) H(X|Y=y), where each H(X|Y=y) = minus sum over x of P(x|y) log2 P(x|y). This weighted average represents the residual uncertainty about X after averaging over all possible observations of Y.

Practical Understanding

In a practical binary channel, noise causes some input bits to flip at the output. The mutual information in this case equals 1 minus the entropy of the error probability, representing how many reliable bits per channel use the system actually conveys. This directly connects to the concept of channel capacity, which is the maximum of I(X;Y) over all possible input distributions P(x).

For a Gaussian channel under power constraint, the capacity-achieving distribution is Gaussian, and I(X;Y) achieves the Shannon capacity formula. Mutual information is therefore not just a theoretical measure but the operational quantity that determines the maximum data rate a channel can support reliably.

Engineers use mutual information in practice to evaluate modulation schemes, assess the effectiveness of error-correcting codes, and design iterative decoders. In LDPC and turbo code design, the EXIT chart technique directly plots mutual information transfer between decoder components to predict convergence.

Solved Numerical Example

Consider a binary channel where input X is equally likely: P(X=0)=P(X=1)=0.5. The channel transition probabilities are P(Y=0|X=0)=0.9, P(Y=1|X=0)=0.1, P(Y=0|X=1)=0.1, P(Y=1|X=1)=0.9. We compute I(X;Y) using I(X;Y) = H(Y) - H(Y|X).

Example
Given:
P(X=0) = P(X=1) = 0.5
P(Y=1|X=0) = P(Y=0|X=1) = p = 0.1  (crossover probability)

Why this formula applies:
I(X;Y) = H(Y) - H(Y|X) is easiest when output marginals are computed first.

Formula:
I(X;Y) = H(Y) - H(Y|X)
H(Y|X) = H(p) = -p*log2(p) - (1-p)*log2(1-p)  [binary entropy of crossover probability]

Substitution:
P(Y=1) = P(Y=1|X=0)*P(X=0) + P(Y=1|X=1)*P(X=1)
       = 0.1*0.5 + 0.9*0.5 = 0.5
H(Y) = -0.5*log2(0.5) - 0.5*log2(0.5) = 1 bit
H(Y|X) = H(0.1) = -0.1*log2(0.1) - 0.9*log2(0.9)

Calculation:
H(0.1) = 0.1*3.3219 + 0.9*0.1520 = 0.3322 + 0.1368 = 0.469 bits
I(X;Y) = 1 - 0.469

Final Answer: I(X;Y) = 0.531 bits per channel use
Exam Tip: For a BSC with crossover probability p, I(X;Y) = 1 - H(p) when input is uniform. H(Y|X) always equals H(p) regardless of input distribution, but H(Y) is maximized at 1 bit only for equal priors. GATE often tests whether students incorrectly compute H(Y|X) using output marginals instead of transition probabilities.
Source XP(x) knownChannelP(y|x) definedOutput YP(y) derivedH(X): Prior uncertaintybefore observing YI(X;Y): Reduction= H(X) - H(X|Y)H(X|Y): Posteriorremaining uncertaintyEquivalent formsI(X;Y) = H(Y) - H(Y|X) = H(X)+H(Y)-H(X,Y)Always: I(X;Y) >= 0 and I(X;Y) = I(Y;X)Channel Capacity C = max over P(x) of I(X;Y)
Figure 2: Mutual information as the information flow mechanism between source and channel output
  • I(X;Y) equals the entropy of the input minus the conditional entropy remaining after the output is observed, quantifying exactly what the channel reveals.
  • The symmetry I(X;Y) = I(Y;X) holds mathematically via the joint entropy formula, even though the physical interpretation differs.
  • Channel capacity is defined as C = max_{P(x)} I(X;Y), meaning we optimize the input distribution to squeeze maximum mutual information from the channel.
  • When X and Y are statistically independent, I(X;Y) = 0 and the channel conveys no useful information about the source.
  • The KL divergence interpretation shows that mutual information is zero only when the joint distribution completely factors into the product of marginals.

Quick Revision

  • I(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) = H(X)+H(Y)-H(X,Y). All forms are equivalent and GATE tests all of them.
  • I(X;Y) >= 0 always. It equals zero only when X and Y are independent (useless channel).
  • I(X;Y) = I(Y;X): mutual information is symmetric even though H(X|Y) is not equal to H(Y|X) in general.
  • For a BSC with crossover p and uniform input: I(X;Y) = 1 - H(p) where H(p) = -p log2 p - (1-p) log2(1-p).
  • Channel capacity C = max over all input distributions of I(X;Y).
  • Exam trap: H(Y|X) is computed using transition probabilities P(y|x), NOT the output marginals P(y). Confusing these is a common error.
  • H(X,Y) = H(X) + H(Y|X) = H(Y) + H(X|Y): chain rule connects joint, marginal, and conditional entropies.

Mutual Information Quiz

Test your command of mutual information definitions, symmetry properties, and channel interpretations.

Question 1 of 3

Q1.Mutual information I(X;Y) is correctly expressed as: