Information Theory Basics
Entropy, information content I = -log2(P), Shannon.
Information theory provides the mathematical foundation for quantifying, compressing, and reliably transmitting information through noisy channels. Developed by Claude Shannon in 1948, it answers a fundamental question that had no rigorous answer before: how much information is contained in a message, and what is the maximum rate at which information can be reliably sent over a given channel? For electronics and communication engineering students, information theory is the theoretical backbone behind every practical system from mobile networks to data storage.
Core Concept of Information
The first question information theory answers is: what does it mean to quantify information? Shannon defined the self-information or information content of an event x with probability P(x) as I(x) = -log₂(P(x)) bits. This definition captures a physically meaningful intuition: a rare event carries more information than a common event. If someone tells you that the sun rose today, that message carries essentially no information because it was virtually certain. If someone tells you an earthquake just struck in an unexpected location, that message carries a large amount of information because it was unlikely.
The logarithm base determines the unit of information. Base 2 gives bits, which is the standard in digital communication. Base e gives nats, and base 10 gives Hartleys. For a binary source with equally likely symbols (P = 0.5 each), the information content of each symbol is I = -log₂(0.5) = 1 bit, which aligns with our intuitive understanding of binary digit information. The logarithm also ensures the additivity property: if two independent events A and B occur together, their combined information equals I(A) + I(B) = -log₂(P(A)P(B)) = -log₂(P(A)) - log₂(P(B)).
The three axioms that uniquely determine this measure of information are: information must be a continuous function of probability, a certain event (P=1) must carry zero information, and information from independent events must add. Only the negative logarithmic function satisfies all three conditions simultaneously. This is not a convention but a mathematical necessity.
Mathematical Expression
While self-information measures the information in a single specific outcome, the entropy H(X) of a source X measures the average information per symbol produced by that source. Entropy is the expected value of self-information over all possible symbols: H(X) = -sum over all i of P(xi) log₂P(xi). This average is computed using the probabilities of each symbol as weights. A source that always produces the same symbol has zero entropy. A source with completely unpredictable output has maximum entropy.
Entropy reaches its maximum value of log₂(M) bits per symbol when all M symbols are equally likely, each with probability 1/M. This is the most uncertain source, and it carries the most information per symbol on average. Any non-uniform probability distribution reduces entropy below this maximum. For a binary source with probabilities p and (1-p), the entropy function H(p) = -p log₂(p) - (1-p) log₂(1-p) forms a concave curve peaking at H = 1 bit when p = 0.5, and falling to zero when p = 0 or p = 1.
The Shannon channel capacity theorem states that reliable communication is possible at any rate below C = B log₂(1 + SNR) bits per second, where B is the bandwidth and SNR is the signal-to-noise ratio. At rates above C, reliable communication is theoretically impossible regardless of the coding scheme used. This is the most important result in all of information theory and establishes the theoretical upper limit for any digital communication system.
Practical Understanding
Information theory directly determines how efficiently data can be compressed. A source with entropy H(X) bits per symbol requires at least H(X) bits per symbol on average to encode it without loss. This is the source coding theorem, also called Shannon's first theorem. If a compression algorithm uses fewer than H(X) bits per symbol, it cannot losslessly recover all source messages. If it uses more than H(X), it is suboptimal and wasting bandwidth. This is why entropy is the fundamental limit for lossless data compression.
In practical engineering, the Shannon capacity formula guides system design. Increasing bandwidth B by a factor of 2 doubles capacity. Doubling SNR adds only log₂(2) = 1 bit/s/Hz of spectral efficiency. This means that beyond a certain SNR, adding more transmit power gives diminishing returns, and it becomes more efficient to increase bandwidth instead. Modern communication systems such as 4G LTE and 5G NR are engineered to operate close to the Shannon limit using advanced coding techniques like turbo codes and LDPC codes.
Solved Numerical Example
Consider a discrete memoryless source (DMS) producing four symbols A, B, C, D with probabilities 1/2, 1/4, 1/8, 1/8 respectively. The self-information of each symbol and the source entropy are to be determined. The entropy also represents the minimum average code length achievable for lossless compression of this source.
Given:
Symbols: A, B, C, D
Probabilities: P(A) = 1/2, P(B) = 1/4, P(C) = 1/8, P(D) = 1/8
Why this formula applies:
Self-information I(x) = -log₂P(x) measures surprise per symbol.
Entropy H = -sum P(x) log₂P(x) gives the average information per symbol.
Formula:
I(x) = -log₂(P(x)) [bits]
H(X) = -sum P(xi) * log₂(P(xi)) [bits/symbol]
Substitution (Self-information of each symbol):
I(A) = -log₂(1/2) = log₂(2) = 1 bit
I(B) = -log₂(1/4) = log₂(4) = 2 bits
I(C) = -log₂(1/8) = log₂(8) = 3 bits
I(D) = -log₂(1/8) = log₂(8) = 3 bits
Calculation (Entropy):
H(X) = P(A)*I(A) + P(B)*I(B) + P(C)*I(C) + P(D)*I(D)
H(X) = (1/2)*1 + (1/4)*2 + (1/8)*3 + (1/8)*3
H(X) = 0.5 + 0.5 + 0.375 + 0.375
H(X) = 1.75 bits/symbol
Maximum possible entropy (if all 4 symbols equally likely):
H_max = log₂(4) = 2 bits/symbol
Final Answer with units:
I(A) = 1 bit, I(B) = 2 bits, I(C) = 3 bits, I(D) = 3 bits
Source Entropy H(X) = 1.75 bits/symbol
Minimum average code length achievable = 1.75 bits/symbol (Shannon limit)Exam Tip: For GATE, when all M symbols are equally probable, entropy H = log₂(M). For a binary source, maximum entropy = 1 bit/symbol at p=0.5. Entropy is always non-negative and never exceeds log₂(M). A common trap is forgetting that 0 * log₂(0) = 0 by convention (limit).
Mechanism Explained
- Self-information I(x) = -log₂(P(x)) quantifies the surprise or information content of a single event. Rare events have high information; certain events have zero information.
- Entropy H(X) = -sum P(xi) log₂(P(xi)) is the statistical average of self-information and represents the average uncertainty or information content per symbol of the source.
- Entropy is maximum at H = log₂(M) bits/symbol when all M symbols are equally probable, and falls to zero when the source is deterministic (one symbol has probability 1).
- The source coding theorem establishes that H(X) is the fundamental lower bound on the average code length: no lossless compression scheme can achieve an average of fewer than H(X) bits per source symbol.
- The channel capacity C = B log₂(1 + SNR) is the maximum rate at which information can be reliably transmitted. Error-free communication is theoretically achievable at any rate below C with appropriate coding.
Quick Revision
- Self-information: I(x) = -log₂(P(x)) bits. Rare event: high I. Certain event: I = 0.
- Entropy: H(X) = -sum P(xi) log₂(P(xi)) bits/symbol. Average information of source.
- Maximum entropy: H_max = log₂(M) bits/symbol for M equally likely symbols.
- Source coding theorem: Average code length L_avg >= H(X). Huffman coding achieves H(X) <= L_avg < H(X) + 1.
- Channel capacity: C = B log₂(1 + SNR) bps. Shannon-Hartley law. Rate R < C enables reliable comm.
- GATE trap: 0 * log₂(0) = 0 by convention. Entropy is always non-negative. Entropy is measured in bits when log base 2 is used.
- Units: Self-information and entropy in bits (log base 2), nats (log base e), or Hartleys (log base 10).
Information Theory Basics Quiz
Test your command of self-information, entropy definitions, and Shannon's foundational results.
Q1.A source emits symbol x with probability P(x) = 1/32. What is the self-information of this symbol in bits?
Related Articles
Mutual Information
I(X;Y) = H(X) - H(X|Y), channel relationship.
9 min read
Differential Entropy
Entropy of continuous random variables.
5 min read
Shannon Limit
Implications of Shannon-Hartley theorem.
6 min read
BEC Channel
Binary Erasure Channel properties.
5 min read
BSC Channel
Binary Symmetric Channel, capacity calculation.
10 min read