Contents

Digital Communication
Other Subjects
Section Progress33%

3 of 9 articles

Joint and Conditional Entropy

H(X,Y), H(Y|X), chain rule for entropy.

Darshan N
Updated: 19 March 2026
5 min read

When a communication system involves two related random variables, such as the input and output of a noisy channel, the entropy of each variable alone does not capture the full picture of their statistical relationship. Joint entropy and conditional entropy extend the concept of entropy to pairs of random variables, providing the mathematical tools to quantify shared uncertainty, residual uncertainty, and the information gained about one variable by observing the other. These concepts are foundational to deriving channel capacity and understanding the limits of communication over noisy links.

Entropy Venn Diagram: H(X), H(Y), Joint and Conditional EntropyH(X|Y)X aloneI(X;Y)Mutual InfoH(Y|X)Y aloneH(X,Y) = H(X|Y) + I(X;Y) + H(Y|X)H(X,Y) = H(X) + H(Y|X) = H(Y) + H(X|Y)H(X)H(Y)← H(X,Y) Joint →H(X,Y)-ΣΣ P(x,y)log₂P(x,y)H(Y|X)-ΣΣ P(x,y)log₂P(y|x)
Figure 1: Entropy Venn diagram illustrating the relationship between joint entropy, conditional entropy and mutual information for two random variables X and Y

Core Concept: Joint and Conditional Entropy

The joint entropy H(X,Y) of two discrete random variables X and Y measures the total uncertainty contained in the pair (X,Y) together. It is defined as H(X,Y) = -sum over all x,y of P(x,y) * log₂(P(x,y)), where P(x,y) is the joint probability mass function. Joint entropy generalizes single-variable entropy to two dimensions. If X and Y are independent, the joint entropy equals the sum of individual entropies: H(X,Y) = H(X) + H(Y). For dependent variables, knowing one reduces uncertainty about the other, so H(X,Y) < H(X) + H(Y).

The conditional entropy H(Y|X) measures the remaining uncertainty about Y after X is fully known. It is computed as the expected value of the conditional self-information, averaged over all values of X: H(Y|X) = sum over all x of P(x) * H(Y|X=x), which expands to H(Y|X) = -sum over all x,y of P(x,y) * log₂(P(y|x)). Intuitively, H(Y|X) answers: on average, how many more bits are needed to describe Y if we already know X? In a noiseless channel where Y = X exactly, H(Y|X) = 0. In a completely noisy channel where Y is independent of X, H(Y|X) = H(Y).

The relationship between these quantities is captured by the chain rule for entropy: H(X,Y) = H(X) + H(Y|X). This states that the joint uncertainty about both X and Y equals the uncertainty about X alone, plus the remaining uncertainty about Y given X. Equivalently, H(X,Y) = H(Y) + H(X|Y). These two decompositions are equal because H(X,Y) is symmetric and the choice of which variable to condition on first does not matter for the total.

Mathematical Expression

The four key entropy quantities for a pair (X,Y) and their relationships are: H(X,Y) = -sum_{x,y} P(x,y) log₂ P(x,y) for joint entropy; H(Y|X) = -sum_{x,y} P(x,y) log₂ P(y|x) for conditional entropy of Y given X; and H(X|Y) = -sum_{x,y} P(x,y) log₂ P(x|y) for conditional entropy of X given Y. The mutual information I(X;Y), which measures the information shared between X and Y, is defined as I(X;Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) = H(X) + H(Y) - H(X,Y). Mutual information is always non-negative and equals zero only when X and Y are statistically independent.

The chain rule extends to longer sequences. For three variables: H(X,Y,Z) = H(X) + H(Y|X) + H(Z|X,Y). This is extremely useful for analyzing multi-stage communication systems where symbols have memory. For a Markov chain X → Y → Z, the data processing inequality states that I(X;Z) <= I(X;Y), meaning no processing of Y can increase the information about X beyond what Y already contains.

Several inequalities are worth memorizing for GATE. Conditioning never increases entropy: H(X|Y) <= H(X), with equality only when X and Y are independent. Joint entropy is sub-additive: H(X,Y) <= H(X) + H(Y), with equality for independent X and Y. Mutual information satisfies 0 <= I(X;Y) <= min(H(X), H(Y)).

Practical Understanding

In channel analysis, the conditional entropy H(Y|X) represents the equivocation or noise entropy. When a symbol X is sent over a noisy channel and Y is received, H(Y|X) measures how uncertain the received symbol is even after knowing exactly what was sent. This uncertainty arises from channel noise. A high equivocation means the channel is very noisy. The channel capacity is C = max over input distribution of I(X;Y), which is maximized when the input distribution is chosen to be the one that maximizes mutual information.

The reverse conditional entropy H(X|Y) is called the conditional entropy of the input given the output. This measures the uncertainty about what was sent even after observing what was received. It is the irreducible ambiguity from the receiver's perspective. A reliable communication system aims to make H(X|Y) as small as possible. In a binary symmetric channel (BSC) with crossover probability p, H(Y|X) = H(X|Y) = H(p), where H(p) = -p log₂(p) - (1-p) log₂(1-p) is the binary entropy function.

Solved Numerical Example

Consider a simple binary channel where X and Y each take values {0,1}. The joint probability table is given: P(X=0, Y=0) = 0.4, P(X=0, Y=1) = 0.1, P(X=1, Y=0) = 0.1, P(X=1, Y=1) = 0.4. The marginal probabilities are P(X=0) = 0.5, P(X=1) = 0.5, P(Y=0) = 0.5, P(Y=1) = 0.5. Compute H(X), H(Y), H(X,Y), H(Y|X) and I(X;Y).

Example
Given:
P(X=0,Y=0) = 0.4, P(X=0,Y=1) = 0.1
P(X=1,Y=0) = 0.1, P(X=1,Y=1) = 0.4
Marginals: P(X=0)=P(X=1)=0.5,  P(Y=0)=P(Y=1)=0.5

Why this formula applies:
Joint entropy uses joint PMF. Conditional entropy uses conditional PMF P(y|x)=P(x,y)/P(x).

Formula:
H(X)   = -sum P(x) log₂P(x)
H(X,Y) = -sum_{x,y} P(x,y) log₂P(x,y)
H(Y|X) = H(X,Y) - H(X)   [chain rule]
I(X;Y) = H(Y) - H(Y|X)  = H(X) + H(Y) - H(X,Y)

Substitution:
H(X) = -(0.5*log₂0.5 + 0.5*log₂0.5) = -(2 * 0.5*(-1)) = 1 bit
H(Y) = 1 bit  (same uniform marginal)

H(X,Y) = -[0.4*log₂0.4 + 0.1*log₂0.1 + 0.1*log₂0.1 + 0.4*log₂0.4]
log₂(0.4) = -1.322  →  0.4 * (-1.322) = -0.529
log₂(0.1) = -3.322  →  0.1 * (-3.322) = -0.332
H(X,Y) = -[(-0.529) + (-0.332) + (-0.332) + (-0.529)]
H(X,Y) = -(-1.722) = 1.722 bits

Calculation:
H(Y|X) = H(X,Y) - H(X) = 1.722 - 1 = 0.722 bits
H(X|Y) = H(X,Y) - H(Y) = 1.722 - 1 = 0.722 bits
I(X;Y) = H(X) + H(Y) - H(X,Y) = 1 + 1 - 1.722 = 0.278 bits

Verification: I(X;Y) = H(Y) - H(Y|X) = 1 - 0.722 = 0.278 bits (consistent)

Final Answer with units:
H(X) = 1 bit,  H(Y) = 1 bit
H(X,Y) = 1.722 bits
H(Y|X) = H(X|Y) = 0.722 bits  (equivocation)
I(X;Y) = 0.278 bits  (mutual information)
Exam Tip: Always use the chain rule H(X,Y) = H(X) + H(Y|X) to compute conditional entropy from joint and marginal entropies. Avoid computing H(Y|X) directly from conditional probabilities unless the table is very simple. Also remember: conditioning reduces entropy H(Y|X) <= H(Y), and I(X;Y) >= 0 always.
Chain Rule, Inequalities and Channel InterpretationChain Rule for EntropyH(X,Y) = H(X) + H(Y|X)H(X,Y) = H(Y) + H(X|Y)Both forms equal H(X,Y)Chain rule extends to n variablesKey InequalitiesH(Y|X) ≤ H(Y) [conditioning reduces entropy]H(X,Y) ≤ H(X) + H(Y) [sub-additivity]0 ≤ I(X;Y) ≤ min(H(X),H(Y))I(X;Y) = 0 iff X,Y independentChannel Interpretation (Binary Symmetric Channel)Input XOutput YP(correct) = 1-pP(error) = p [crossover]H(Y|X) = H(p) equivocation (noise)H(X|Y) = H(p) ambiguity at receiverI(X;Y) = H(Y) - H(p) = 1 - H(p) [for uniform input]Channel Capacity C = max I(X;Y) = 1 - H(p) bits/use
Figure 2: Entropy chain rule equations, key inequalities, and channel interpretation showing how H(Y|X) represents noise and I(X;Y) gives channel capacity for BSC

Mechanism Explained

  • Joint entropy H(X,Y) = -sum_{x,y} P(x,y) log₂ P(x,y) measures total uncertainty in the pair. It equals H(X) + H(Y) for independent X and Y, and is strictly less for dependent variables.
  • Conditional entropy H(Y|X) = H(X,Y) - H(X) measures residual uncertainty about Y after X is observed. In a channel context, this is the equivocation or noise entropy.
  • Chain rule H(X,Y) = H(X) + H(Y|X) decomposes joint uncertainty into the uncertainty of X alone plus the remaining uncertainty about Y given X. The order is interchangeable.
  • Mutual information I(X;Y) = H(X) + H(Y) - H(X,Y) = H(Y) - H(Y|X) measures information shared between X and Y. For a channel, it is the rate of information transfer and is maximized over input distributions to obtain channel capacity.
  • For a binary symmetric channel with crossover probability p, conditional entropy H(Y|X) = H(p) = -p log₂(p) - (1-p) log₂(1-p), and capacity C = 1 - H(p) bits per channel use with uniform input.

Quick Revision

  • Joint entropy: H(X,Y) = -sum_{x,y} P(x,y) log₂P(x,y). Total uncertainty of the pair.
  • Conditional entropy: H(Y|X) = H(X,Y) - H(X). Residual uncertainty about Y after knowing X.
  • Chain rule: H(X,Y) = H(X) + H(Y|X) = H(Y) + H(X|Y). Key tool for computation.
  • Mutual information: I(X;Y) = H(X) + H(Y) - H(X,Y) = H(Y) - H(Y|X). Always non-negative.
  • Inequalities: H(Y|X) <= H(Y), H(X,Y) <= H(X)+H(Y), 0 <= I(X;Y) <= min(H(X), H(Y)).
  • For BSC with crossover p: Capacity C = 1 - H(p), equivocation = H(p) for uniform input.
  • GATE trap: H(Y|X) is NOT the same as H(X|Y) in general. Only symmetric when X and Y have identical marginals, as in a BSC with uniform input.

Joint Conditional Entropy Quiz

Test your ability to apply joint entropy, conditional entropy, and the chain rule for entropy.

Question 1 of 3

Q1.The chain rule for entropy states that H(X,Y) equals: