FPGA Architecture Elements

LUT, CLB, Switch Matrix.

Mohith N
Updated: 19 March 2026
7 min read

A Field Programmable Gate Array (FPGA) is an integrated circuit whose logic function is not fixed at manufacturing time but can be programmed by the user after fabrication. Unlike ASICs, FPGAs allow designers to implement and modify digital circuits without any custom silicon steps. Understanding the internal architecture of an FPGA is essential for VLSI design courses, as it explains how a single programmable device can emulate arbitrary logic, memories, and interconnect structures.

FPGA Architecture: Array of CLBs and RoutingCLBLUT+FFCLBLUT+FFCLBLUT+FFCLBLUT+FFCLBLUT+FFCLBSwitchMatrixCLBSwitchMatrixCLBCLBCLBSwitchMatrixCLBCLBI/O BlocksProgrammable Horizontal and Vertical Routing ChannelsCLB = Configurable Logic Block, Switch Matrix connects routing channels
Figure 1: FPGA internal structure - an array of CLBs connected through Switch Matrices via programmable routing channels.

Core Concept: The Three Fundamental Elements

An FPGA chip is organized as a regular two-dimensional array of three main elements: Configurable Logic Blocks (CLBs), a programmable routing network (including Switch Matrices and Connection Blocks), and I/O Blocks (IOBs) at the periphery. Each of these is programmed by loading configuration bits into static RAM cells on the chip, which is why standard FPGAs are called SRAM-based. When power is removed, the configuration is lost and must be reloaded from an external flash memory.

Look-Up Table (LUT): Logic Implementation

The logic inside each CLB is implemented using a Look-Up Table (LUT). An n-input LUT is essentially a small 2^n x 1 SRAM array. For every combination of the n input signals, the SRAM stores a preloaded output bit. Because the SRAM can hold any truth table, an n-input LUT can implement any Boolean function of n variables without any structural change to the hardware. Modern FPGAs typically use 6-input LUTs (64 SRAM cells each), which balance logic density against routing complexity.

A 4-input LUT has 2^4 = 16 SRAM cells. If the desired function is a 4-input AND gate, then only the cell corresponding to input combination 1111 is set to 1 and all other 15 cells are set to 0. If the function is later changed to XOR, the SRAM contents are simply reprogrammed with the XOR truth table. This runtime reprogrammability is the defining advantage of FPGA over fixed-function gates.

Configurable Logic Block (CLB): Complete Slice

A Configurable Logic Block contains one or more LUTs along with flip-flops, carry chains for fast arithmetic, and local routing MUXes. In Xilinx terminology, a CLB contains multiple smaller units called slices, each of which has LUTs, D flip-flops, and carry logic. The flip-flop in the CLB is used to register the LUT output, enabling synchronous sequential logic. The carry chain passes the carry bit directly between adjacent CLBs without using the general routing, making arithmetic operations much faster than they would be if routed through the switch matrix.

Switch Matrix and Routing Architecture

The Switch Matrix is a programmable crossbar-like network placed at the intersection of horizontal and vertical routing channels. It contains a set of pass transistors (or MUXes) controlled by SRAM configuration bits. When a configuration bit is set to 1, the corresponding pass transistor closes, connecting two wire segments. By programming the appropriate bits in the switch matrices along a path, a signal can be routed from one CLB output to any other CLB input anywhere on the chip.

The Connection Block sits between the CLB and the routing channel and determines which wires in the routing channel a given CLB input or output can connect to. The flexibility of the connection block and the density of the switch matrix determine the routing efficiency of the FPGA. Increasing flexibility increases routability but also increases area and capacitance on the interconnect, reducing speed.

Mathematical Expression: LUT Size and Function Coverage

An n-input LUT requires 2^n configuration bits to store the truth table. The number of distinct Boolean functions of n variables is 2^(2^n). For n=4, this is 2^16 = 65,536 distinct functions, all implementable by a single 4-LUT. For n=6, a single 6-LUT can implement any of 2^64 possible 6-input Boolean functions.

Numerical Example

Example
Given:
LUT size: n = 4 inputs
Target function: 4-input majority function (output = 1 if 3 or more inputs are 1)

Why this formula applies:
An n-input LUT stores 2^n bits covering all input combinations.

Formula:
LUT size = 2^n SRAM cells

Substitution:
LUT size = 2^4 = 16 SRAM cells (addresses 0000 to 1111)

Calculation:
Majority function is 1 when sum of inputs >= 3:
Combinations with 3 ones: 0111,1011,1101,1110 (4 combinations)
Combinations with 4 ones: 1111 (1 combination)
Total SRAM cells set to 1: 5 out of 16

Final Answer:
The 4-input LUT is programmed with a 16-bit word where bits at addresses
7 (0111), 11 (1011), 13 (1101), 14 (1110), and 15 (1111) are set to 1,
and all others are 0. Any reprogramming to a new function requires only
changing the 16 SRAM contents, no hardware modification needed.
Exam Tip: For GATE, remember that an n-input LUT uses 2^n SRAM bits and can implement any Boolean function of n variables. Switch matrices connect routing channels but not CLBs directly to each other. The carry chain in a CLB bypasses the switch matrix for fast arithmetic. FPGA is slower than ASIC for the same function due to routing through programmable switches.
CLB Internal Structure: LUT, Flip-Flop, and Carry ChainIn0In1In2In34-input LUT16-cell SRAMtruth table storeOutput = storedbit at address{In3,In2,In1,In0}MUXcomb / regD Flip-FlopCLK, CE, ROutput QCarry Chain (Fast Arithmetic Path)Carry propagates between adjacent CLBs directlyBypasses switch matrix, 4x-8x faster than routed carry
Figure 2: CLB internals - 4-input LUT drives a MUX that selects combinational or registered output; carry chain bypasses routing.
  • The LUT is a small SRAM that stores the truth table; programming it changes the implemented logic function without any hardware modification.
  • CLB contains LUT(s), flip-flop(s), and carry chain; the carry chain is a dedicated fast path for adder and comparator logic.
  • Switch Matrix is a programmable crossbar at the intersection of routing channels; SRAM bits control which wire segments are connected.
  • Connection Blocks link CLB inputs and outputs to the routing channels, controlling which specific channel wires are accessible.
  • FPGA is slower and less power-efficient than ASIC for the same function because signals travel through multiple programmable switch transistors.

Quick Revision

  • LUT: n-input LUT has 2^n SRAM cells; implements any Boolean function of n variables.
  • CLB = LUT + Flip-Flop + Carry Chain + local MUXes; slice is the sub-unit in Xilinx FPGAs.
  • Switch Matrix: programmable crossbar at routing channel intersections; SRAM-controlled pass transistors route signals.
  • Connection Block: interfaces CLB I/O pins to routing channel wires.
  • Carry chain bypasses switch matrix for fast arithmetic, critical for adder performance on FPGA.
  • SRAM-based FPGA: configuration lost on power off; must reload from external flash at startup.
  • Exam trap: the LUT stores truth table bits, not logic gates. More LUT inputs = more area but fewer levels of logic for complex functions.

FPGA Internal Architecture

Test your knowledge on this topic.

Question 1 of 3

Q1.How many SRAM cells are required to implement a k-input Look-Up Table (LUT)?