Instruction Set

Thumb-2 instruction set highlights.

Darshan N
Updated: 19 March 2026
8 min read

The instruction set of a processor defines the vocabulary of operations it can execute. For ARM Cortex-M processors, the Thumb-2 instruction set is the foundation that enables both high code density and strong computational performance, making it the preferred choice for embedded systems in resource-constrained microcontrollers.

ARM Thumb-2 Instruction Set ArchitectureMixed 16-bit and 32-bit instructions in a single unified ISA16-bit ThumbCompact encodingHigh code densityReduces Flash usage32-bit ARMFull operand rangeComplex operationsWider immediatesThumb-2 MergedBoth in one modeNo mode switchingBest of both worldsKey Thumb-2 Instruction CategoriesData ProcessingMOV, ADD, SUBAND, ORR, EORLoad / StoreLDR, STRPUSH, POPBranchB, BL, BXConditional BccMultiply / DivideMUL, UDIVSDIV, MLA
Figure 1: Thumb-2 unifies 16-bit and 32-bit instructions without mode switching, balancing code density and performance.

Core Concept Explanation

The original ARM architecture used 32-bit fixed-length instructions. To reduce code size for embedded devices, ARM introduced the Thumb instruction set, which used 16-bit encodings. However, Thumb-only mode had limitations such as restricted register access and no conditional execution for most instructions. Thumb-2 was introduced to solve this by allowing both 16-bit and 32-bit instructions to coexist in a single execution state, without the overhead of switching modes.

In Thumb-2, approximately 60 percent of instructions compile to 16-bit encodings, reducing the memory footprint significantly. The remaining operations that require wider immediates or more complex addressing use 32-bit encodings. The processor does not need an explicit mode-switch instruction between them. This is a fundamental advantage over the older Thumb and ARM interworking approach.

Cortex-M processors such as M0, M3, M4, and M7 all use Thumb-2 as their base ISA. The M0 uses a subset, while M3 and above support the full Thumb-2 instruction set including hardware divide, saturation arithmetic, and DSP extensions on M4 and M7.

Mathematical Expression

Code density is often measured as bytes per operation. For a given program, if N instructions are compiled:

Code Size (Thumb-2) = N16 x 2 bytes + N32 x 4 bytes, where N16 is the count of 16-bit instructions and N32 is the count of 32-bit instructions. For comparison, pure ARM encoding would give Code Size (ARM) = N x 4 bytes. Thumb-2 typically achieves 25 to 35 percent code size reduction over pure 32-bit ARM encoding.

Instruction throughput depends on pipeline stages. The Cortex-M3 uses a 3-stage pipeline: Fetch, Decode, Execute. For most Thumb-2 instructions, the CPI (Cycles Per Instruction) is 1 for register operations and 2 for load/store operations.

Practical Understanding

In practice, the compiler (GCC or LLVM/Clang) targets the Thumb-2 ISA automatically when you specify the Cortex-M target. You do not manually choose between 16-bit and 32-bit opcodes. The assembler selects the most compact valid encoding. When writing inline assembly or studying disassembly, you will notice some instructions prefixed with a dot-wide or dot-narrow suffix to force a specific width.

The IT (If-Then) block is a unique Thumb-2 feature. It allows up to four conditional instructions to execute without a branch, reducing pipeline flushes. For example, ITEEE EQ allows one instruction to execute if equal and the next three to execute if not equal. This is exam-relevant because it is specific to Thumb-2 and not found in ARM state.

Register file in Cortex-M has 16 general-purpose registers (R0 to R15). R13 is the Stack Pointer, R14 is the Link Register (LR) that stores the return address on a BL call, and R15 is the Program Counter. Most 16-bit Thumb-2 instructions can only access R0 to R7, whereas 32-bit encodings can access the full R0 to R12 range.

Example
Given:
A Cortex-M3 program has 200 instructions.
60% compile to 16-bit = 120 instructions
40% compile to 32-bit = 80 instructions

Why this formula applies:
Thumb-2 uses variable width encoding: 2 bytes for 16-bit, 4 bytes for 32-bit.

Formula:
Code Size = (N16 x 2) + (N32 x 4)

Substitution:
Code Size = (120 x 2) + (80 x 4)

Calculation:
Code Size = 240 + 320 = 560 bytes

Comparison (pure ARM):
Code Size = 200 x 4 = 800 bytes

Final Answer:
Thumb-2 saves 240 bytes = 30% reduction over pure ARM encoding.
Exam Tip: GATE and university exams often ask which Cortex-M variant supports hardware divide. M0 and M0+ do NOT support UDIV/SDIV. M3 and above do. Also remember that IT blocks are a Thumb-2 exclusive feature, not available in pure ARM state.
Thumb-2 Instruction Execution Pipeline and Register FileFETCHRead instr from FlashDECODE16 or 32-bit decodeEXECUTEALU / LSU / BranchRegister FileR0 - R7 (Low regs)R0 to R7 - 16/32-bit accessR8 to R12 - 32-bit onlyR13 - Stack Pointer (SP)R14 - Link Register (LR)R15 - Program Counter (PC)xPSR - Program Status RegThumb-2 Instruction Width DecisionInstructionADD R0, R1LDR R0, [R1, #256]MOV R8, R0UDIV R0, R1, R2Width16-bit32-bit32-bit32-bitReasonLow regs, small immLarge offset neededHigh reg accessM3+ only instructionIT Block Example (Thumb-2 Exclusive)ITEEE EQ -> If EQ: exec instr1 | If NE: exec instr2, instr3, instr4Avoids branch and pipeline flush for small conditional blocks
Figure 2: Cortex-M3 3-stage pipeline processes Thumb-2 instructions; register access rules differ by instruction width.

Instruction Categories Explained

  • Data processing instructions (ADD, SUB, MOV, AND, ORR, EOR, LSL, LSR) operate on registers and update the condition flags in xPSR when the S suffix is used, for example ADDS sets the carry and zero flags.
  • Load and store instructions (LDR, STR, LDRB, STRH) transfer data between registers and memory. Cortex-M uses a unified memory map so peripheral registers are accessed using the same LDR/STR instructions.
  • Branch instructions include B for unconditional branch, BL for branch with link (function call, stores return address in LR), BX for branch to register (used for function return via BX LR), and conditional branches like BEQ, BNE, BGT.
  • Multiply and divide instructions such as MUL, UMULL (unsigned 64-bit result), UDIV, and SDIV are available from M3 onwards. M0 does not support hardware divide.
  • DSP and SIMD extensions (QADD, SMULL, SMLAL, SIMD byte operations) are available on Cortex-M4 and M7, enabling digital signal processing without a separate DSP core.

Quick Revision

  • Thumb-2 supports mixed 16-bit and 32-bit instructions in one execution state without mode switching.
  • 16-bit instructions access R0 to R7; 32-bit instructions access R0 to R12 and support larger immediates.
  • Code size formula: Size = (N16 x 2) + (N32 x 4) bytes. Typically 25-35% smaller than pure ARM.
  • IT block allows up to 4 conditional instructions without branching, reducing pipeline flushes.
  • UDIV and SDIV (hardware divide) are NOT available on Cortex-M0. Available from M3 onwards.
  • R13 = SP, R14 = LR (holds return address after BL), R15 = PC.
  • Trap: Do not confuse BL (branch with link, for calls) with BX LR (branch to LR, for return).

Thumb Instruction Set

Test your knowledge on this topic!

Question 1 of 3

Q1.What defines the instruction word length in the Thumb-2 instruction set?