Contents

Digital Communication
Other Subjects
Section Progress11%

1 of 9 articles

Information Theory Basics

Entropy, information content I = -log2(P), Shannon.

Mohith N
Updated: 19 March 2026
6 min read

Information theory provides the mathematical foundation for quantifying, compressing, and reliably transmitting information through noisy channels. Developed by Claude Shannon in 1948, it answers a fundamental question that had no rigorous answer before: how much information is contained in a message, and what is the maximum rate at which information can be reliably sent over a given channel? For electronics and communication engineering students, information theory is the theoretical backbone behind every practical system from mobile networks to data storage.

Information Theory: Core Concepts at a GlanceSelf-InformationI(x) = -log₂(P(x))Unit: bitsLow P → High IEntropy H(X)H = -ΣP log₂PAverage informationMax when P = 1/MChannel CapacityC = B log₂(1+SNR)Shannon limitMax reliable rateMutual InfoI(X;Y)Info shared byinput and outputInformation Content vs Probability (Intuition)High ILow IP → 1 (Certain)P → 0Rare event: high infoCertain event: zero infoAxioms of Information Measure (Shannon's Requirements)I increases as P decreasesI = 0 for certain eventI is additive for independent eventsSurprisal propertyP=1 → I(x)=0I(AB) = I(A) + I(B)
Figure 1: Overview of core information theory concepts including self-information, entropy, channel capacity and their mathematical definitions

Core Concept of Information

The first question information theory answers is: what does it mean to quantify information? Shannon defined the self-information or information content of an event x with probability P(x) as I(x) = -log₂(P(x)) bits. This definition captures a physically meaningful intuition: a rare event carries more information than a common event. If someone tells you that the sun rose today, that message carries essentially no information because it was virtually certain. If someone tells you an earthquake just struck in an unexpected location, that message carries a large amount of information because it was unlikely.

The logarithm base determines the unit of information. Base 2 gives bits, which is the standard in digital communication. Base e gives nats, and base 10 gives Hartleys. For a binary source with equally likely symbols (P = 0.5 each), the information content of each symbol is I = -log₂(0.5) = 1 bit, which aligns with our intuitive understanding of binary digit information. The logarithm also ensures the additivity property: if two independent events A and B occur together, their combined information equals I(A) + I(B) = -log₂(P(A)P(B)) = -log₂(P(A)) - log₂(P(B)).

The three axioms that uniquely determine this measure of information are: information must be a continuous function of probability, a certain event (P=1) must carry zero information, and information from independent events must add. Only the negative logarithmic function satisfies all three conditions simultaneously. This is not a convention but a mathematical necessity.

Mathematical Expression

While self-information measures the information in a single specific outcome, the entropy H(X) of a source X measures the average information per symbol produced by that source. Entropy is the expected value of self-information over all possible symbols: H(X) = -sum over all i of P(xi) log₂P(xi). This average is computed using the probabilities of each symbol as weights. A source that always produces the same symbol has zero entropy. A source with completely unpredictable output has maximum entropy.

Entropy reaches its maximum value of log₂(M) bits per symbol when all M symbols are equally likely, each with probability 1/M. This is the most uncertain source, and it carries the most information per symbol on average. Any non-uniform probability distribution reduces entropy below this maximum. For a binary source with probabilities p and (1-p), the entropy function H(p) = -p log₂(p) - (1-p) log₂(1-p) forms a concave curve peaking at H = 1 bit when p = 0.5, and falling to zero when p = 0 or p = 1.

The Shannon channel capacity theorem states that reliable communication is possible at any rate below C = B log₂(1 + SNR) bits per second, where B is the bandwidth and SNR is the signal-to-noise ratio. At rates above C, reliable communication is theoretically impossible regardless of the coding scheme used. This is the most important result in all of information theory and establishes the theoretical upper limit for any digital communication system.

Practical Understanding

Information theory directly determines how efficiently data can be compressed. A source with entropy H(X) bits per symbol requires at least H(X) bits per symbol on average to encode it without loss. This is the source coding theorem, also called Shannon's first theorem. If a compression algorithm uses fewer than H(X) bits per symbol, it cannot losslessly recover all source messages. If it uses more than H(X), it is suboptimal and wasting bandwidth. This is why entropy is the fundamental limit for lossless data compression.

In practical engineering, the Shannon capacity formula guides system design. Increasing bandwidth B by a factor of 2 doubles capacity. Doubling SNR adds only log₂(2) = 1 bit/s/Hz of spectral efficiency. This means that beyond a certain SNR, adding more transmit power gives diminishing returns, and it becomes more efficient to increase bandwidth instead. Modern communication systems such as 4G LTE and 5G NR are engineered to operate close to the Shannon limit using advanced coding techniques like turbo codes and LDPC codes.

Solved Numerical Example

Consider a discrete memoryless source (DMS) producing four symbols A, B, C, D with probabilities 1/2, 1/4, 1/8, 1/8 respectively. The self-information of each symbol and the source entropy are to be determined. The entropy also represents the minimum average code length achievable for lossless compression of this source.

Example
Given:
Symbols: A, B, C, D
Probabilities: P(A) = 1/2, P(B) = 1/4, P(C) = 1/8, P(D) = 1/8

Why this formula applies:
Self-information I(x) = -log₂P(x) measures surprise per symbol.
Entropy H = -sum P(x) log₂P(x) gives the average information per symbol.

Formula:
I(x) = -log₂(P(x))   [bits]
H(X) = -sum P(xi) * log₂(P(xi))   [bits/symbol]

Substitution (Self-information of each symbol):
I(A) = -log₂(1/2) = log₂(2) = 1 bit
I(B) = -log₂(1/4) = log₂(4) = 2 bits
I(C) = -log₂(1/8) = log₂(8) = 3 bits
I(D) = -log₂(1/8) = log₂(8) = 3 bits

Calculation (Entropy):
H(X) = P(A)*I(A) + P(B)*I(B) + P(C)*I(C) + P(D)*I(D)
H(X) = (1/2)*1 + (1/4)*2 + (1/8)*3 + (1/8)*3
H(X) = 0.5 + 0.5 + 0.375 + 0.375
H(X) = 1.75 bits/symbol

Maximum possible entropy (if all 4 symbols equally likely):
H_max = log₂(4) = 2 bits/symbol

Final Answer with units:
I(A) = 1 bit,  I(B) = 2 bits,  I(C) = 3 bits,  I(D) = 3 bits
Source Entropy H(X) = 1.75 bits/symbol
Minimum average code length achievable = 1.75 bits/symbol (Shannon limit)
Exam Tip: For GATE, when all M symbols are equally probable, entropy H = log₂(M). For a binary source, maximum entropy = 1 bit/symbol at p=0.5. Entropy is always non-negative and never exceeds log₂(M). A common trap is forgetting that 0 * log₂(0) = 0 by convention (limit).
Shannon's Information Theory: System HierarchySelf-Information I(x) = -log₂P(x)Entropy H(X) = E[I(x)]Joint Entropy H(X,Y)Source Coding TheoremMutual Info I(X;Y)Channel Capacity CMin code lengthL_avg ≥ H(X)Info shared betweenX and YC = B log₂(1+SNR)Max reliable data rateAll lead to Shannon's fundamental limit: Reliable communication at rate R is possible iff R < C
Figure 2: Hierarchical relationship between information theory concepts from self-information to entropy, mutual information and Shannon channel capacity

Mechanism Explained

  • Self-information I(x) = -log₂(P(x)) quantifies the surprise or information content of a single event. Rare events have high information; certain events have zero information.
  • Entropy H(X) = -sum P(xi) log₂(P(xi)) is the statistical average of self-information and represents the average uncertainty or information content per symbol of the source.
  • Entropy is maximum at H = log₂(M) bits/symbol when all M symbols are equally probable, and falls to zero when the source is deterministic (one symbol has probability 1).
  • The source coding theorem establishes that H(X) is the fundamental lower bound on the average code length: no lossless compression scheme can achieve an average of fewer than H(X) bits per source symbol.
  • The channel capacity C = B log₂(1 + SNR) is the maximum rate at which information can be reliably transmitted. Error-free communication is theoretically achievable at any rate below C with appropriate coding.

Quick Revision

  • Self-information: I(x) = -log₂(P(x)) bits. Rare event: high I. Certain event: I = 0.
  • Entropy: H(X) = -sum P(xi) log₂(P(xi)) bits/symbol. Average information of source.
  • Maximum entropy: H_max = log₂(M) bits/symbol for M equally likely symbols.
  • Source coding theorem: Average code length L_avg >= H(X). Huffman coding achieves H(X) <= L_avg < H(X) + 1.
  • Channel capacity: C = B log₂(1 + SNR) bps. Shannon-Hartley law. Rate R < C enables reliable comm.
  • GATE trap: 0 * log₂(0) = 0 by convention. Entropy is always non-negative. Entropy is measured in bits when log base 2 is used.
  • Units: Self-information and entropy in bits (log base 2), nats (log base e), or Hartleys (log base 10).

Information Theory Basics Quiz

Test your command of self-information, entropy definitions, and Shannon's foundational results.

Question 1 of 3

Q1.A source emits symbol x with probability P(x) = 1/32. What is the self-information of this symbol in bits?