Hyper-Threading
SMT concepts.
Hyper-Threading Technology (HTT), Intel's commercial implementation of Simultaneous Multithreading (SMT), allows a single physical processor core to appear as two (or more) logical processors to the operating system. The motivation is to improve CPU resource utilization: in a typical workload, one thread often stalls waiting for memory, cache misses, or long-latency instructions, leaving execution units idle. SMT fills these idle slots with instructions from a second thread running concurrently on the same core.
Core Concept: What SMT Does
In a traditional out-of-order processor, a single thread can issue multiple instructions per cycle but cannot always keep all execution units busy. Cache misses, branch mispredictions, and data dependencies create bubbles in the instruction stream. Simultaneous Multithreading (SMT) addresses this by maintaining the complete architectural state (register file, program counter, stack pointer, and status flags) of two or more threads simultaneously within one physical core. Instructions from both threads are interleaved in the issue queues, and the out-of-order engine dispatches whichever instructions are ready, regardless of which thread they belong to.
The key insight is that duplicating the architectural state (register file, PC, etc.) is relatively cheap in terms of die area, while the large, complex execution units (ALU, FPU, load/store units, reorder buffer) are expensive and are shared. The OS sees two logical processors, can schedule two threads, and each thread has its own context, but underneath they time-share the same physical execution resources simultaneously.
What is Duplicated vs What is Shared
Understanding the hardware boundary is essential for GATE. Elements that are duplicated per logical thread include the general-purpose register file, program counter, stack pointer, status/flag registers, and (in some implementations) parts of the reorder buffer and load/store queue. Elements that are shared between threads include all execution units (ALU, FPU, SIMD units), L1 instruction and data caches, L2 and L3 caches, the branch predictor (partially), and the TLB. This sharing means threads can interfere with each other: one thread's cache pressure can evict cache lines that the other thread needs.
Performance Model
The performance benefit of SMT depends strongly on the workload. If both threads are compute-intensive and keep all execution units busy already, SMT adds little benefit and may even hurt due to resource contention. If one thread frequently stalls (memory-bound workloads, or mixed compute and I/O), the second thread can productively use the otherwise idle execution units. The theoretical maximum speedup from 2-way SMT is not 2x because the cores are not doubled; they are better utilized. Typical real-world speedup is 15 to 30% for suitable workloads.
The throughput (IPC or instructions per clock) of an SMT core running two threads is compared against two separate physical cores. Let IPC_single be the IPC of one thread on a core without SMT. For SMT, define IPC_smt_combined as total instructions retired per cycle across both threads. The SMT efficiency is: Efficiency = IPC_smt_combined / (2 x IPC_single). Values above 0.5 indicate SMT is beneficial.
Given:
Core executes Thread A and Thread B simultaneously via 2-way SMT
Thread A alone achieves IPC = 2.4 on this core
Thread B alone achieves IPC = 2.0 on this core
With SMT (both threads together), combined IPC measured = 3.6
Why this formula applies:
If each thread ran on its own core, we'd get 2.4 + 2.0 = 4.4 total IPC from two cores.
With SMT, two threads share one core and get combined IPC = 3.6.
We compare resource utilization against the ideal 2-core scenario.
Formula:
SMT Efficiency = IPC_combined_SMT / (IPC_A + IPC_B)
Substitution:
SMT Efficiency = 3.6 / (2.4 + 2.0)
Calculation:
SMT Efficiency = 3.6 / 4.4
Final Answer: SMT Efficiency = 0.818 (approximately 82%)
SMT recovers 82% of the performance of two separate cores using only one physical core.Exam Tip: GATE and interview questions frequently ask what is duplicated versus shared in SMT. Remember: register file and PC are duplicated (cheap), execution units and caches are shared (expensive). SMT does NOT double the number of physical cores. It increases utilization of one core.
SMT vs Multi-core vs Time-sharing
- Time-sharing (OS-level): threads take turns using one core. One thread runs at a time. Context switches happen at millisecond timescales. No real parallelism at the hardware level.
- SMT: both threads issue instructions to the shared execution units in the same cycle. Thread 1's instructions fill slots left idle by Thread 0's stalls. True cycle-level hardware parallelism within one core.
- Multi-core: each thread runs on a completely separate physical core with its own execution units. No sharing of critical resources (except LLC). Provides stronger isolation and predictable performance.
- SMT can hurt performance when both threads are cache-intensive, as they compete for L1 and L2 cache space, causing higher miss rates than either thread would have alone.
- Security implication: SMT creates microarchitectural side channels (such as Spectre, MDS vulnerabilities) because threads share execution unit state (cache, TLB, buffers). Some security-sensitive deployments disable SMT.
Quick Revision
- SMT (Hyper-Threading) presents N logical cores to the OS from one physical core by duplicating architectural state (registers, PC) while sharing execution units.
- Duplicated: register file, PC, status flags, portions of ROB. Shared: ALU, FPU, L1 cache, L2 cache, TLB, branch predictor.
- Typical SMT speedup: 15 to 30% for suitable (mixed latency) workloads. No benefit or slight regression for purely compute-bound workloads.
- SMT efficiency = Combined IPC of two threads on SMT / Sum of individual IPCs on dedicated cores.
- SMT differs from multi-core: multi-core has physically separate execution units. SMT shares them.
- SMT differs from time-sharing: time-sharing is OS-level context switching (ms granularity). SMT is hardware-level cycle-by-cycle interleaving.
- GATE trap: do not say SMT doubles performance. It improves utilization of existing resources. The OS sees 2 logical CPUs, but there is 1 physical core.
Simultaneous Multithreading Concepts
Test your knowledge of hardware threading and shared execution resources.
Q1.Which processor resources are duplicated in a Simultaneous Multithreading architecture?
Related Articles
Pipelining
Instruction pipeline, hazards, superscalar arch.
7 min read
Virtual Memory
Paging, segmentation with paging, TLB.
5 min read
Protected Mode
Descriptor tables, selectors, protection rings.
10 min read
Cache Memory
L1, L2, L3 cache, hit/miss, mapping techniques.
9 min read
Evolution to Pentium
286, 386, 486, Pentium features.
5 min read