Joint and Conditional Entropy
H(X,Y), H(Y|X), chain rule for entropy.
When a communication system involves two related random variables, such as the input and output of a noisy channel, the entropy of each variable alone does not capture the full picture of their statistical relationship. Joint entropy and conditional entropy extend the concept of entropy to pairs of random variables, providing the mathematical tools to quantify shared uncertainty, residual uncertainty, and the information gained about one variable by observing the other. These concepts are foundational to deriving channel capacity and understanding the limits of communication over noisy links.
Core Concept: Joint and Conditional Entropy
The joint entropy H(X,Y) of two discrete random variables X and Y measures the total uncertainty contained in the pair (X,Y) together. It is defined as H(X,Y) = -sum over all x,y of P(x,y) * log₂(P(x,y)), where P(x,y) is the joint probability mass function. Joint entropy generalizes single-variable entropy to two dimensions. If X and Y are independent, the joint entropy equals the sum of individual entropies: H(X,Y) = H(X) + H(Y). For dependent variables, knowing one reduces uncertainty about the other, so H(X,Y) < H(X) + H(Y).
The conditional entropy H(Y|X) measures the remaining uncertainty about Y after X is fully known. It is computed as the expected value of the conditional self-information, averaged over all values of X: H(Y|X) = sum over all x of P(x) * H(Y|X=x), which expands to H(Y|X) = -sum over all x,y of P(x,y) * log₂(P(y|x)). Intuitively, H(Y|X) answers: on average, how many more bits are needed to describe Y if we already know X? In a noiseless channel where Y = X exactly, H(Y|X) = 0. In a completely noisy channel where Y is independent of X, H(Y|X) = H(Y).
The relationship between these quantities is captured by the chain rule for entropy: H(X,Y) = H(X) + H(Y|X). This states that the joint uncertainty about both X and Y equals the uncertainty about X alone, plus the remaining uncertainty about Y given X. Equivalently, H(X,Y) = H(Y) + H(X|Y). These two decompositions are equal because H(X,Y) is symmetric and the choice of which variable to condition on first does not matter for the total.
Mathematical Expression
The four key entropy quantities for a pair (X,Y) and their relationships are: H(X,Y) = -sum_{x,y} P(x,y) log₂ P(x,y) for joint entropy; H(Y|X) = -sum_{x,y} P(x,y) log₂ P(y|x) for conditional entropy of Y given X; and H(X|Y) = -sum_{x,y} P(x,y) log₂ P(x|y) for conditional entropy of X given Y. The mutual information I(X;Y), which measures the information shared between X and Y, is defined as I(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) = H(X) + H(Y) - H(X,Y). Mutual information is always non-negative and equals zero only when X and Y are statistically independent.
The chain rule extends to longer sequences. For three variables: H(X,Y,Z) = H(X) + H(Y|X) + H(Z|X,Y). This is extremely useful for analyzing multi-stage communication systems where symbols have memory. For a Markov chain X → Y → Z, the data processing inequality states that I(X;Z) <= I(X;Y), meaning no processing of Y can increase the information about X beyond what Y already contains.
Several inequalities are worth memorizing for GATE. Conditioning never increases entropy: H(X|Y) <= H(X), with equality only when X and Y are independent. Joint entropy is sub-additive: H(X,Y) <= H(X) + H(Y), with equality for independent X and Y. Mutual information satisfies 0 <= I(X;Y) <= min(H(X), H(Y)).
Practical Understanding
In channel analysis, the conditional entropy H(Y|X) represents the equivocation or noise entropy. When a symbol X is sent over a noisy channel and Y is received, H(Y|X) measures how uncertain the received symbol is even after knowing exactly what was sent. This uncertainty arises from channel noise. A high equivocation means the channel is very noisy. The channel capacity is C = max over input distribution of I(X;Y), which is maximized when the input distribution is chosen to be the one that maximizes mutual information.
The reverse conditional entropy H(X|Y) is called the conditional entropy of the input given the output. This measures the uncertainty about what was sent even after observing what was received. It is the irreducible ambiguity from the receiver's perspective. A reliable communication system aims to make H(X|Y) as small as possible. In a binary symmetric channel (BSC) with crossover probability p, H(Y|X) = H(X|Y) = H(p), where H(p) = -p log₂(p) - (1-p) log₂(1-p) is the binary entropy function.
Solved Numerical Example
Consider a simple binary channel where X and Y each take values {0,1}. The joint probability table is given: P(X=0, Y=0) = 0.4, P(X=0, Y=1) = 0.1, P(X=1, Y=0) = 0.1, P(X=1, Y=1) = 0.4. The marginal probabilities are P(X=0) = 0.5, P(X=1) = 0.5, P(Y=0) = 0.5, P(Y=1) = 0.5. Compute H(X), H(Y), H(X,Y), H(Y|X) and I(X;Y).
Given:
P(X=0,Y=0) = 0.4, P(X=0,Y=1) = 0.1
P(X=1,Y=0) = 0.1, P(X=1,Y=1) = 0.4
Marginals: P(X=0)=P(X=1)=0.5, P(Y=0)=P(Y=1)=0.5
Why this formula applies:
Joint entropy uses joint PMF. Conditional entropy uses conditional PMF P(y|x)=P(x,y)/P(x).
Formula:
H(X) = -sum P(x) log₂P(x)
H(X,Y) = -sum_{x,y} P(x,y) log₂P(x,y)
H(Y|X) = H(X,Y) - H(X) [chain rule]
I(X;Y) = H(Y) - H(Y|X) = H(X) + H(Y) - H(X,Y)
Substitution:
H(X) = -(0.5*log₂0.5 + 0.5*log₂0.5) = -(2 * 0.5*(-1)) = 1 bit
H(Y) = 1 bit (same uniform marginal)
H(X,Y) = -[0.4*log₂0.4 + 0.1*log₂0.1 + 0.1*log₂0.1 + 0.4*log₂0.4]
log₂(0.4) = -1.322 → 0.4 * (-1.322) = -0.529
log₂(0.1) = -3.322 → 0.1 * (-3.322) = -0.332
H(X,Y) = -[(-0.529) + (-0.332) + (-0.332) + (-0.529)]
H(X,Y) = -(-1.722) = 1.722 bits
Calculation:
H(Y|X) = H(X,Y) - H(X) = 1.722 - 1 = 0.722 bits
H(X|Y) = H(X,Y) - H(Y) = 1.722 - 1 = 0.722 bits
I(X;Y) = H(X) + H(Y) - H(X,Y) = 1 + 1 - 1.722 = 0.278 bits
Verification: I(X;Y) = H(Y) - H(Y|X) = 1 - 0.722 = 0.278 bits (consistent)
Final Answer with units:
H(X) = 1 bit, H(Y) = 1 bit
H(X,Y) = 1.722 bits
H(Y|X) = H(X|Y) = 0.722 bits (equivocation)
I(X;Y) = 0.278 bits (mutual information)Exam Tip: Always use the chain rule H(X,Y) = H(X) + H(Y|X) to compute conditional entropy from joint and marginal entropies. Avoid computing H(Y|X) directly from conditional probabilities unless the table is very simple. Also remember: conditioning reduces entropy H(Y|X) <= H(Y), and I(X;Y) >= 0 always.
Mechanism Explained
- Joint entropy H(X,Y) = -sum_{x,y} P(x,y) log₂ P(x,y) measures total uncertainty in the pair. It equals H(X) + H(Y) for independent X and Y, and is strictly less for dependent variables.
- Conditional entropy H(Y|X) = H(X,Y) - H(X) measures residual uncertainty about Y after X is observed. In a channel context, this is the equivocation or noise entropy.
- Chain rule H(X,Y) = H(X) + H(Y|X) decomposes joint uncertainty into the uncertainty of X alone plus the remaining uncertainty about Y given X. The order is interchangeable.
- Mutual information I(X;Y) = H(X) + H(Y) - H(X,Y) = H(Y) - H(Y|X) measures information shared between X and Y. For a channel, it is the rate of information transfer and is maximized over input distributions to obtain channel capacity.
- For a binary symmetric channel with crossover probability p, conditional entropy H(Y|X) = H(p) = -p log₂(p) - (1-p) log₂(1-p), and capacity C = 1 - H(p) bits per channel use with uniform input.
Quick Revision
- Joint entropy: H(X,Y) = -sum_{x,y} P(x,y) log₂P(x,y). Total uncertainty of the pair.
- Conditional entropy: H(Y|X) = H(X,Y) - H(X). Residual uncertainty about Y after knowing X.
- Chain rule: H(X,Y) = H(X) + H(Y|X) = H(Y) + H(X|Y). Key tool for computation.
- Mutual information: I(X;Y) = H(X) + H(Y) - H(X,Y) = H(Y) - H(Y|X). Always non-negative.
- Inequalities: H(Y|X) <= H(Y), H(X,Y) <= H(X)+H(Y), 0 <= I(X;Y) <= min(H(X), H(Y)).
- For BSC with crossover p: Capacity C = 1 - H(p), equivocation = H(p) for uniform input.
- GATE trap: H(Y|X) is NOT the same as H(X|Y) in general. Only symmetric when X and Y have identical marginals, as in a BSC with uniform input.
Joint Conditional Entropy Quiz
Test your ability to apply joint entropy, conditional entropy, and the chain rule for entropy.
Q1.The chain rule for entropy states that H(X,Y) equals:
Related Articles
Differential Entropy
Entropy of continuous random variables.
5 min read
Information Theory Basics
Entropy, information content I = -log2(P), Shannon.
6 min read
BEC Channel
Binary Erasure Channel properties.
5 min read
BSC Channel
Binary Symmetric Channel, capacity calculation.
10 min read
Shannon Limit
Implications of Shannon-Hartley theorem.
6 min read