SIMD Instructions
MMX, SSE, AVX overview.
SIMD (Single Instruction, Multiple Data) is a class of processor instructions that apply one operation simultaneously to multiple data elements packed into a wide register. Rather than executing one addition at a time across multiple loop iterations, a SIMD instruction can perform four, eight, or sixteen additions in a single clock cycle by operating on all elements of a wide vector register in parallel. This makes SIMD critical for signal processing, image processing, machine learning inference, and scientific computing.
Flynn's Taxonomy and SIMD Classification
Flynn's Taxonomy classifies computer architectures by the number of instruction streams and data streams. SIMD (Single Instruction, Multiple Data) is one of the four classes: one instruction is broadcast to multiple execution units, each of which operates on a different data element simultaneously. This is distinct from MIMD (Multiple Instruction, Multiple Data), which describes multi-core and multi-processor systems where each core executes its own instruction stream independently.
MMX Instructions
Intel introduced MMX (MultiMedia eXtensions) in 1996 with the Pentium MMX processor. MMX added eight 64-bit registers (MM0 to MM7) which were aliased to the x87 FPU register stack (a design decision that caused complications since both could not be used simultaneously). MMX operated exclusively on integer data using packed integer arithmetic: a 64-bit MMX register could hold eight 8-bit integers, four 16-bit integers, or two 32-bit integers. One MMX instruction performed the operation on all packed elements in parallel. MMX introduced concepts like saturation arithmetic (clamping results to max/min instead of wrapping on overflow), which is useful for image pixel manipulation.
SSE: Streaming SIMD Extensions
SSE, introduced with the Intel Pentium III in 1999, added 128-bit XMM registers (XMM0 to XMM7, later extended to XMM15 in 64-bit mode) and crucially added support for single-precision floating-point operations. An XMM register can hold four 32-bit floats, and one SSE instruction (such as ADDPS) adds four pairs of floats simultaneously. SSE registers are separate from x87, resolving the aliasing problem of MMX.
Subsequent extensions SSE2, SSE3, SSSE3, and SSE4 added support for double-precision float operations (64-bit), integer operations on 128-bit registers, horizontal operations (adding elements within the same register rather than across two registers), and specialized operations for string processing and dot products. SSE2 became the baseline required for the x86-64 (AMD64) architecture, meaning all 64-bit x86 processors support at least SSE2.
AVX: Advanced Vector Extensions
Intel's AVX (Advanced Vector Extensions), introduced in 2011 with Sandy Bridge processors, extended registers to 256 bits, called YMM registers (YMM0 to YMM15). A YMM register holds eight 32-bit floats or four 64-bit doubles. AVX also changed the instruction encoding to a three-operand form (destination, source1, source2) which avoids the destructive two-operand format of SSE where one source register was also the destination.
AVX2 (introduced with Haswell in 2013) extended 256-bit operations to integers as well. AVX-512, introduced in 2017 with Intel Skylake-X server processors, extended registers to 512 bits (ZMM0 to ZMM31). One AVX-512 instruction processes sixteen 32-bit floats or eight 64-bit doubles simultaneously. AVX-512 also introduced write masking, where each element of the output vector can be conditionally written based on a mask register, enabling efficient conditional vector operations without scalar fallback loops.
FMA: Fused Multiply-Add
A particularly important SIMD operation is Fused Multiply-Add (FMA). The FMA operation computes ab + c in a single instruction with a single rounding step, which is both faster (one instruction vs two) and more numerically accurate (only one rounding error instead of two). FMA is critical for dense linear algebra (matrix multiplication), convolution, and neural network inference. FLOPS (floating-point operations per second) for FMA count as 2 operations per cycle since one instruction does both multiply and add.
Theoretical Peak Throughput Calculation
The peak floating-point throughput of a processor with SIMD can be calculated as: Peak GFLOPS = Clock_GHz x Cores x SIMD_width_per_core x FMA_factor. The SIMD width per core is the number of FP operations per instruction (for AVX-512 float, this is 16). The FMA factor is 2 if FMA is used (since one FMA = 2 FLOPS).
Given:
Processor: 4-core chip
Clock frequency = 3.5 GHz
SIMD width = AVX-256 (8 single-precision floats per instruction)
FMA support: Yes (factor = 2)
FMA units per core = 2 (can execute 2 FMA instructions per cycle)
Why this formula applies:
Each core can execute 2 FMA instructions per cycle.
Each FMA instruction computes 8 multiply-add pairs (AVX-256 float).
Each FMA pair counts as 2 FLOPS.
Formula:
Peak GFLOPS = Cores x Clock_GHz x FMA_units_per_core x SIMD_width x FMA_factor
Substitution:
Peak GFLOPS = 4 x 3.5 x 2 x 8 x 2
Calculation:
Peak GFLOPS = 4 x 3.5 = 14
14 x 2 = 28
28 x 8 = 224
224 x 2 = 448
Final Answer: Peak throughput = 448 GFLOPS (single-precision)Exam Tip: When calculating FLOPS with SIMD, remember that FMA counts as 2 FLOPS per data element per cycle (one multiply + one add). For AVX-512 with FMA, one instruction on one core gives 16 x 2 = 32 FLOPS. Multiply by cores and clock frequency for total peak. GATE may also ask how many floats fit in a 256-bit register: answer is 256/32 = 8 single-precision floats.
SIMD Packing and Data Layout
- MMX: 64-bit registers (MM0-MM7), integer only, holds 8 x Int8 or 4 x Int16 or 2 x Int32. Aliased to x87 FPU registers.
- SSE: 128-bit XMM registers, single-precision float (4 x Float32 per register), separate from x87, introduced MOVAPS, ADDPS, MULPS.
- SSE2: adds Double precision (2 x Float64) and integer operations on 128-bit registers. Mandatory baseline for x86-64.
- AVX: 256-bit YMM registers, 8 x Float32 or 4 x Float64, three-operand non-destructive encoding. Introduced VADDPS, VMULPS.
- AVX-512: 512-bit ZMM registers (32 registers), 16 x Float32 or 8 x Float64, adds write masking per element, scatter/gather support.
- FMA (Fused Multiply-Add): computes ab+c in one instruction with one rounding. Counts as 2 FLOPS. Critical for GEMM and convolution.
- Data layout: Structure of Arrays (SoA) allows contiguous memory loads with a single VMOVAPS, enabling full SIMD width utilization. Array of Structures (AoS) requires slower gather instructions.
Quick Revision
- SIMD: Single Instruction, Multiple Data. One instruction operates on all elements of a vector register in parallel. Flynn's class: SIMD.
- MMX: 64-bit integer only. SSE: 128-bit XMM, float32. SSE2: adds float64 and int. AVX: 256-bit YMM. AVX-512: 512-bit ZMM.
- Number of float32 elements per register: MMX=2, SSE=4, AVX=8, AVX-512=16.
- Peak GFLOPS = Cores x GHz x FMA_units_per_core x SIMD_elements x 2 (FMA factor).
- FMA instruction: a*b+c in one clock, one rounding error. Counts as 2 FLOPS per element.
- SoA data layout preferred for SIMD (contiguous loads). AoS layout requires gather/scatter (slower).
- GATE trap: when asked how many floats fit in AVX register, answer is 256/32 = 8. For AVX-512: 512/32 = 16. Do not confuse bits and bytes.
SIMD Instruction Sets
Assess your understanding of vector processing and instruction extensions.
Q1.Advanced Vector Extensions expands the width of SIMD registers to how many bits?
Related Articles
Pipelining
Instruction pipeline, hazards, superscalar arch.
7 min read
Stack Instructions
PUSH, POP, XTHL, SPHL.
9 min read
Instruction Set Groups
Data transfer, Arithmetic, Logical, Branching.
8 min read
Rotate Instructions
RLC, RRC, RAL, RAR usage.
10 min read
String Instructions
MOVS, CMPS, SCAS, LODS, STOS with REP.
7 min read