Exception Handling
Entry/Exit sequences, tail-chaining.
Exception handling in ARM Cortex-M defines how the processor responds to events that interrupt normal program flow, such as hardware interrupts, faults, and system calls. Understanding the entry and exit sequences is essential for writing correct interrupt service routines and for GATE-level questions on processor architecture.
Core Concept Explanation
An exception in Cortex-M is any event that causes the processor to deviate from normal sequential execution. This includes hardware interrupts from peripherals (IRQs), system exceptions like HardFault, MemManage, BusFault, UsageFault, and software-triggered events via SVC (Supervisor Call) and PendSV. All exceptions are numbered and managed through the NVIC (Nested Vectored Interrupt Controller) in coordination with the processor core.
When an exception occurs, the hardware automatically saves processor state to the current stack. This is called auto-stacking. Eight registers are pushed: R0, R1, R2, R3, R12, LR (Link Register), PC (return address), and xPSR (program status). These are saved in a specific order and the stack grows downward. After stacking, the processor fetches the exception vector from the Vector Table and begins executing the ISR in Handler Mode.
The vector table is located at address 0x00000000 by default but can be relocated using the VTOR (Vector Table Offset Register). Each entry in the vector table holds the address of the corresponding handler function. The processor reads the appropriate entry based on the exception number and jumps to that address.
Mathematical Expression
Exception latency is the number of cycles from exception trigger to first instruction of the ISR. For Cortex-M3 and M4, the latency without tail-chaining is 12 cycles: 8 cycles for stacking plus a few cycles for vector fetch and pipeline refill. The priority of exceptions is encoded as a numerical value where a lower number means higher priority.
Priority encoding: if an exception has priority P1 and is active, it can be preempted by a pending exception with priority P2 only if P2 < P1 (numerically lower value = higher priority). The number of implemented priority bits varies: M0 implements 2 bits (4 levels), M3 and M4 implement up to 8 bits (256 levels, though most devices implement 4 bits giving 16 levels).
Practical Understanding
The EXC_RETURN value is a special 32-bit value loaded into LR by hardware on exception entry. Its upper 28 bits are all 1s (0xFFFFFFFx), and the lower 4 bits encode the return context: whether to return to Thread or Handler mode, whether to use the Main Stack Pointer (MSP) or Process Stack Pointer (PSP), and whether FPU state was saved. When the ISR executes BX LR with this special value, the hardware recognizes it as an exception return rather than a normal function return.
Tail-chaining is an optimization where, if another exception is pending when the current ISR completes, the processor does not unstack and restack. Instead, it directly fetches the next vector and continues in Handler Mode. This saves the 12 cycle overhead of the full exception entry sequence and makes ISR throughput much more efficient in systems with many back-to-back interrupts.
Late arrival is another optimization. If a higher-priority exception occurs while the processor is already performing the stacking operation for a lower-priority exception, the processor switches to serving the higher-priority handler first. The stacking overhead is shared, so the total latency is not doubled.
Given:
Cortex-M4 system. Normal exception latency = 12 cycles.
Three IRQs fire in rapid succession (back-to-back).
Each ISR body = 20 cycles of actual work.
Why this formula applies:
With tail-chaining, subsequent ISRs skip unstacking+restacking (saves 12 cycles each).
Formula:
Total cycles = First entry (12) + ISR1 (20) + ISR2 entry via tail-chain (6) + ISR2 (20) + ISR3 entry via tail-chain (6) + ISR3 (20)
Tail-chain saves = 12 - 6 = 6 cycles per chained transition.
Substitution:
Without tail-chain: 3 x (12 + 20) = 3 x 32 = 96 cycles
With tail-chain: 12 + 20 + (6 + 20) + (6 + 20) = 84 cycles
Calculation:
Cycles saved = 96 - 84 = 12 cycles
Final Answer:
Tail-chaining saves 12 cycles for 3 back-to-back ISRs (6 cycles saved per chained transition).Exam Tip: In GATE and university exams, remember that auto-stacking saves exactly 8 registers (R0-R3, R12, LR, PC, xPSR). Callee-saved registers (R4-R11) are NOT automatically stacked by hardware. If your ISR uses R4-R11, the compiler adds explicit PUSH/POP instructions. Also, EXC_RETURN in LR starts with 0xFFFFFFx, not a normal address.
Exception Handling Key Mechanisms
- Auto-stacking saves R0, R1, R2, R3, R12, LR, PC, xPSR in that order on the stack. This preserves the calling context without any ISR code. The stack grows downward so xPSR ends up at the highest address and R0 at the lowest.
- LR is loaded with the EXC_RETURN value (0xFFFFFFF9, 0xFFFFFFFD, or 0xFFFFFFE9 depending on context) which tells the processor the return mode when BX LR is executed inside the ISR.
- Tail-chaining allows back-to-back ISR execution without the overhead of full exception exit and entry. The processor remains in Handler Mode and fetches the next vector directly.
- Late-arrival optimization handles the case where a higher-priority exception arrives during the stacking phase. The processor completes stacking once and serves the higher-priority handler first.
- The Vector Table must be aligned to at least 128 bytes (or a power of two that is large enough to hold all vectors). VTOR register allows its relocation to SRAM for dynamic patching of vectors at runtime.
Quick Revision
- Auto-stacking saves exactly 8 registers: R0, R1, R2, R3, R12, LR, PC, xPSR. R4-R11 are NOT automatically saved.
- Exception latency for Cortex-M3/M4 is 12 cycles in the normal case (no tail-chain, no late arrival).
- EXC_RETURN is loaded into LR by hardware on exception entry. Value 0xFFFFFFF9 = return to Thread Mode, use MSP.
- Tail-chaining: when next IRQ is pending on ISR exit, skip unstack+restack, fetch next vector directly.
- Late arrival: higher-priority IRQ arriving during stacking phase gets served first without duplicate stacking overhead.
- Priority: numerically lower value = higher priority. M0 has 2 priority bits (4 levels), M3/M4 up to 8 bits.
- Trap: SVC can only be used in Thread Mode. Using SVC inside an ISR causes a HardFault.
Exception Handling Practice
Test your knowledge on this topic!
Q1.Which optimization technique does the NVIC use when a new exception occurs while the processor is already executing exception exit routines?
Related Articles
ARM Cortex-M Architecture
M3/M4 pipeline, bus matrix.
8 min read
Programmers Model
Registers R0-R15, xPSR.
10 min read
Stack Memory
MSP vs PSP stacks, operation.
8 min read
Memory Map
Code, SRAM, Peripheral bit-band regions.
12 min read
Embedded Systems Overview
Definition, constraints, design metrics.
9 min read