CMOS Adders

Ripple carry, Carry lookahead, Manchester carry chain in CMOS.

Darshan N
Updated: 19 March 2026
10 min read

Addition is the most fundamental arithmetic operation in digital hardware. In CMOS VLSI, adder circuits are built from full adder cells, and the way carry is propagated between bit positions determines the overall speed and area of the circuit. Three major architectures: ripple carry, carry lookahead, and Manchester carry chain each represent a distinct engineering trade-off in CMOS.

Understanding these adder topologies is essential for GATE and for designing datapath blocks such as ALUs, multipliers, and accumulators. The delay analysis of each topology directly applies RC delay models and logical effort concepts from foundational VLSI.

Three CMOS Adder Architectures: Speed vs AreaRipple CarryFA bit 0FA bit 1FA bit 2Delay: O(N)Area: smallCarry LookaheadGenerate/PropCLA LogicSum computationDelay: O(log N)Area: largeManchester ChainPass transistorsCarry chain RCBuffered outputDelay: O(N) but fastArea: compact
Figure 1: Overview of three major CMOS adder topologies showing structural differences in carry propagation

Core Concept Explanation

A full adder computes one bit of a binary sum: given inputs A, B, and carry-in Cin, it produces a Sum output and a carry-out Cout. For an N-bit addition, N full adder cells are needed. The challenge is how to generate and propagate the carry bit from position 0 to position N-1 as quickly as possible.

In a ripple carry adder, the carry-out of each stage directly feeds the carry-in of the next. The critical path runs through all N carry stages in series. Total delay scales as O(N), making ripple carry slow for wide word lengths like 32 or 64 bits. However, it uses minimal area and is regular and simple, making it useful for narrow additions or non-critical paths.

The carry lookahead adder (CLA) removes the serial dependency by pre-computing carry signals using generate and propagate expressions. Define G_i = A_i AND B_i (this bit generates a carry regardless of Cin) and P_i = A_i XOR B_i (this bit propagates an incoming carry). Then Cout_i = G_i OR (P_i AND Cin_i). By expanding this recursively, carry at any position can be computed directly from the original inputs in O(log N) time, at the cost of wider logic gates and more area.

The Manchester carry chain takes a different approach. Instead of full static logic for each carry, it uses pass transistors controlled by generate and propagate signals to either pull the carry node directly to a defined voltage or pass the previous carry value. This forms an RC chain where the carry ripples through switched nodes. It is faster than static ripple carry for the same N because the transistors are small and the structure is compact, though it still scales as O(N) and requires periodic buffering to restore signal levels.

Mathematical Expression

For a ripple carry adder, total delay is approximately:

T_ripple = T_FA_carry * (N - 1) + T_FA_sum

where T_FA_carry is the carry propagation delay of one full adder and T_FA_sum is the final sum generation delay.

For a 4-bit CLA group, the carry expressions are:

C1 = G0 + P0*C0

C2 = G1 + P1*G0 + P1*P0*C0

C3 = G2 + P2*G1 + P2*P1*G0 + P2*P1*P0*C0

Each carry is computed in two gate delays regardless of bit position within the group. For hierarchical CLA, the O(log N) scaling is achieved by applying the same look-ahead structure at multiple levels.

For the Manchester carry chain, the delay to bit k is modeled using Elmore delay:

T_manchester(k) = sum from i=0 to k of (Ron_i * sum from j=i to k of C_j)

This is the standard Elmore formula for a series RC chain with pass transistors as resistors and node capacitances as the shunt elements. Due to the quadratic growth with k, buffers are inserted every 4 to 8 bits.

Practical Understanding

In modern standard-cell libraries, the full adder cell is usually implemented using a combination of static CMOS and transmission gates to minimize transistor count and delay simultaneously. The XOR gates needed for sum and propagate computation are efficiently realized with transmission gates, as seen in the 28-transistor full adder common in textbooks, which reduces to 10-12 transistors with TG-based approaches.

In high-performance processors, neither ripple carry nor a single-level CLA is used. Instead, prefix adder structures such as Kogge-Stone or Brent-Kung adders extend the CLA principle using a tree of group generate and propagate operators. These achieve O(log N) delay with better area and fanout characteristics than a naive multi-level CLA.

The Manchester carry chain is particularly favored in custom datapath design where layout is hand-crafted. Its compactness and predictable RC delay make it suitable for bit-slice designs in ALU and multiplier arrays. Modern SoC designs use it in timing-critical 4-bit slices before handing off to a lookahead structure.

Example
Given:
8-bit ripple carry adder
Full adder carry delay: T_FA_carry = 200 ps
Full adder sum delay: T_FA_sum = 300 ps
Also compare: 8-bit CLA with 2-level hierarchy
CLA group delay (4-bit group generate): T_group = 150 ps
CLA inter-level: T_inter = 100 ps, T_sum = 200 ps

Why this formula applies:
Ripple: carry must ripple through 7 FA carry stages, then final sum
CLA: 2-level hierarchy: level-1 covers 4 bits, level-2 selects the group

Formula:
T_ripple = T_FA_carry * (N-1) + T_FA_sum
T_CLA    = T_group + T_inter + T_sum

Substitution:
T_ripple = 200e-12 * 7 + 300e-12
T_CLA    = 150e-12 + 100e-12 + 200e-12

Calculation:
T_ripple = 1400 + 300 = 1700 ps
T_CLA    = 150 + 100 + 200 = 450 ps

Final Answer:
Ripple carry delay = 1700 ps = 1.7 ns
CLA delay         = 450 ps  = 0.45 ns
Speedup = 1700 / 450 = 3.78x improvement
Exam Tip: In GATE questions on adder delay, identify whether it is asking for carry propagation delay or total sum output delay. Ripple carry delay is (N-1)*T_carry + T_sum, not N*T_carry. The last stage only needs to generate sum, not a carry to any further stage.
Manchester Carry Chain: Pass-Transistor StructureCarry node pre-charged to Vdd; pass transistors pull down conditionallyC_inPass_P0ctrl: P0C1Kill_G0GNDPass_P1ctrl: P1C2Kill_G1GNDPass_P2ctrl: P2C3Kill_G2GNDBufOperation: Pi=1 passes Cin through; Gi=1 pulls carry node to 0 (kill)RC Elmore delay grows as O(k^2) for k bits, so buffers inserted every 4-8 bitsCompact layout: pass transistors share diffusion with FA cells in custom datapath
Figure 2: Manchester carry chain schematic showing pass-transistor propagate path and kill transistors for generate logic
  • Each carry node is pre-charged; a Kill transistor (controlled by Gi) pulls it to ground when a carry is generated at that bit.
  • A Pass transistor (controlled by Pi) forms the propagate path, allowing carry from the previous stage to ripple through.
  • Elmore delay grows quadratically with chain length, requiring buffer insertion every 4-8 bits.
  • Manchester chain is compact in layout and widely used in custom 4-bit ALU slices.
  • CLA achieves O(log N) delay by pre-computing all carry signals in parallel using generate-propagate logic.

Quick Revision

  • Ripple carry delay: T = (N-1)*T_carry + T_sum; simple, small area, slow for wide adders.
  • CLA uses G_i = A_i AND B_i and P_i = A_i XOR B_i to compute carries in O(log N).
  • Manchester carry chain uses pass transistors for P and kill transistors for G; compact and fast for short chains.
  • Elmore delay in Manchester chain: quadratic with chain length, so buffer every 4-8 bits.
  • Kogge-Stone and Brent-Kung are prefix adders extending CLA with better fanout control.
  • TG-based full adder cell uses 10-12 transistors versus 28 in full static CMOS.
  • Exam trap: Ripple carry delay uses (N-1) carry stages, not N; the last stage only needs sum, not carry-out.

CMOS Adders Basics

Test your knowledge on carry logic and ripple delays.

Question 1 of 3

Q1.What logical condition defines the carry generate signal in a full adder?