Multi-core Processors

SMP, CMP, shared memory consistency.

Darshan N
Updated: 19 March 2026
10 min read

Multi-core processors integrate two or more independent processing cores onto a single chip. Rather than continuing to push single-core clock frequencies higher (which increases power consumption and heat generation disproportionately), semiconductor engineers began placing multiple complete cores on one die. This shift from frequency scaling to parallelism fundamentally changed how software is written and how hardware is designed.

Multi-Core Processor DieCore 0ALU / FPUL1 I-Cache (32KB)L1 D-Cache (32KB)Core 1ALU / FPUL1 I-Cache (32KB)L1 D-Cache (32KB)Core 2ALU / FPUL1 I-Cache (32KB)L1 D-Cache (32KB)Core 3ALU / FPUL1 I$ (32KB)L1 D$ (32KB)Shared L2 Cache (Core 0 and 1) — 256KB eachShared L2 Cache (Core 2 and 3) — 256KB eachShared Last Level Cache (L3) — 8 MB (All Cores)Memory Bus and Cache Coherence InterconnectMain Memory (DRAM)
Figure 1: Typical quad-core processor die with private L1 caches, shared L2 per pair, and a shared LLC (L3)

Symmetric Multiprocessing (SMP)

In Symmetric Multiprocessing (SMP), all processors or cores share a single physical address space and have equal access to memory and I/O devices. Each processor runs any process assigned by the OS, and no processor is designated a master. The OS treats all cores identically, enabling load balancing across cores. This is the dominant architecture in desktop, laptop, and server multi-core chips today.

The key challenge in SMP is maintaining a consistent view of memory across all cores. When Core 0 writes a value to a memory location that is also cached in Core 1's private L1 cache, Core 1 may read a stale value. This is the cache coherence problem and requires hardware protocols to solve it.

Chip Multiprocessor (CMP)

A Chip Multiprocessor (CMP), also called a multi-core processor in common usage, places multiple complete processor cores on a single semiconductor die. This reduces inter-core communication latency compared to discrete multi-chip SMP systems since cores share on-chip buses and caches. Modern Intel Core i7, AMD Ryzen, and ARM Cortex-A series processors are all CMPs. All CMPs can implement SMP, but SMP can also exist across multiple chips.

Cache Coherence Protocols

The MESI protocol is the most widely used cache coherence protocol in multi-core systems. MESI stands for four states a cache line can be in: Modified (dirty copy, only in this cache), Exclusive (clean copy, only in this cache), Shared (clean copy, possibly in other caches), and Invalid (not valid). When one core writes to a cache line, the protocol invalidates copies in all other caches (write-invalidate protocol), or it updates them (write-update protocol). Write-invalidate is more common as it avoids sending the actual data on every write.

Hardware coherence is typically implemented using a snooping bus (where all caches monitor bus transactions) or a directory-based scheme (where a centralized or distributed directory tracks which caches have which lines). Snooping scales well up to 8 to 16 cores. Beyond that, directory-based protocols are preferred.

Shared Memory Consistency

Even with cache coherence, the order in which memory operations appear to other processors is governed by the memory consistency model. Sequential consistency (SC) requires that all processors see all memory operations in the same order, as if they were executed on a single processor. This is the strongest and most intuitive model but limits hardware optimization. Relaxed consistency models (like Total Store Order used in x86 and Partial Store Order used in SPARC) allow certain reorderings in exchange for performance, requiring explicit memory fence (barrier) instructions to enforce ordering where needed.

Amdahl's Law and Parallelism Limit

The maximum speedup achievable by adding more cores is limited by the fraction of the program that must remain serial. Amdahl's Law states: Speedup = 1 / (S + (1-S)/N), where S is the serial fraction, N is the number of cores, and (1-S) is the parallelizable fraction. Even with infinite cores, speedup is bounded by 1/S. For a program with 20% serial code, the maximum speedup is 1/0.2 = 5x regardless of how many cores are added.

Example
Given:
Serial fraction of program (S) = 0.25 (25% cannot be parallelized)
Number of cores (N) = 8

Why this formula applies:
Amdahl's Law bounds speedup by the irreducible serial portion.
As N increases, the parallel portion's time approaches zero, leaving only serial time.

Formula:
Speedup = 1 / (S + (1-S)/N)

Substitution:
Speedup = 1 / (0.25 + (0.75)/8)

Calculation:
Speedup = 1 / (0.25 + 0.09375)
Speedup = 1 / 0.34375

Final Answer: Speedup = 2.91x

Note: Maximum possible speedup (N -> infinity) = 1/S = 1/0.25 = 4x
With 8 cores, we reach only 72.7% of theoretical maximum.
Exam Tip: GATE tests Amdahl's Law numerically and conceptually. A common trap is assuming speedup equals N for N cores. Always identify the serial fraction first. Also, MESI state transitions are asked directly in GATE — know that a write to a Shared line causes an Invalidation broadcast.

MESI Protocol State Transitions

Modified(M)Dirty, only hereExclusive(E)Clean, only hereShared(S)Clean, multi-copyInvalid(I)Not validLocal WriteOther cache readWriteback on evictRemote write -> InvalidateInvalidationLocal WriteLocal Read (from mem)Local Read
Figure 2: MESI protocol state transitions — write to Shared state causes invalidation of all other cached copies
  • Modified (M): Cache line is dirty and valid only in this cache. Must be written back to memory before another cache can read it.
  • Exclusive (E): Cache line is clean and only present in this cache. A local write transitions it to Modified without broadcasting.
  • Shared (S): Cache line is clean and may be present in multiple caches. A write causes an Invalidation message to all other sharers.
  • Invalid (I): Cache line is not valid. A read causes a fetch from memory or from another cache holding the Modified copy (cache-to-cache transfer).
  • False sharing occurs when two cores access different variables that happen to reside in the same cache line. Each write invalidates the other core's line even though they are not sharing data logically.

Quick Revision

  • SMP: all cores share one address space with equal memory access. CMP: multiple cores on one die implementing SMP.
  • Amdahl's Law: Speedup = 1 / (S + (1-S)/N). Maximum speedup = 1/S even with infinite cores.
  • MESI states: Modified (dirty, exclusive), Exclusive (clean, exclusive), Shared (clean, replicated), Invalid (no valid copy).
  • Write-invalidate: on a write to Shared line, all other copies are invalidated. Write-update: all copies are updated. Invalidate is more common.
  • Snooping bus coherence scales to ~16 cores. Directory-based coherence needed for larger core counts.
  • Sequential consistency: all cores see memory operations in same global order. Relaxed models (TSO, PSO) allow reordering for performance; fences restore ordering.
  • GATE trap: False sharing causes performance degradation even when cores access different variables. Cache line granularity is the key cause.

Multicore Processor Architecture

Evaluate your understanding of SMP and memory consistency models.

Question 1 of 3

Q1.In a Symmetric Multiprocessor system, what characterizes the memory access time for all processors?