Instruction Queue

Pipelining fetch and execute benefits.

Darshan N
Updated: 19 March 2026
7 min read

The instruction queue is a 6-byte FIFO buffer inside the 8086 that enables the Bus Interface Unit (BIU) to prefetch instruction bytes from memory while the Execution Unit (EU) is processing a previous instruction. This overlapping of fetch and execute cycles is a form of pipelining that significantly improves processor throughput compared to a simple fetch-then-execute design.

8086 Internal Architecture: BIU and EU with Instruction QueueBus Interface Unit (BIU)Segment Registers: CS, DS, SS, ESInstruction Pointer (IP)Instruction Queue (6 bytes)B1B2B3B4B5B6FIFO: filled by BIU from memoryAddress Generation Logic (20-bit)Execution Unit (EU)General Registers: AX, BX, CX, DXPointer/Index: SP, BP, SI, DIALU (Arithmetic Logic Unit)executes decoded instructionsFlag RegisterInstruction DecoderfetchQueue feeds EU continuously — pipelining fetch and execute stages
Figure 1: 8086 BIU and EU with the 6-byte instruction queue enabling fetch-execute overlap

Why the Instruction Queue Exists

Without prefetching, the processor would need to fetch each instruction from memory before it could begin executing, causing the bus to be idle during execution and the ALU to be idle during fetch. The instruction queue breaks this dependency. The Bus Interface Unit (BIU) continuously fills the 6-byte queue from memory whenever the queue has 2 or more empty bytes, using bus cycles that are not required by the EU for data access.

The Execution Unit (EU) fetches instruction bytes from the local queue rather than going to external memory directly. This means the EU rarely has to wait for instruction bytes. Simultaneously, the BIU is fetching the next instruction from memory into the queue. The two units operate concurrently, implementing a two-stage pipeline.

FIFO Operation and Queue Behavior

The queue operates as a First-In First-Out (FIFO) buffer. Bytes enter the queue from the high-address end (filled by BIU) and are consumed from the low-address end (read by EU). Since the 8086 fetches two bytes at a time from memory (using its 16-bit data bus), it fills the queue in 2-byte chunks. The queue can hold up to 6 instruction bytes at any time.

The queue is flushed and restarted when a branch instruction is executed. If the EU executes a JMP, CALL, or any transfer-of-control instruction, the prefetched bytes in the queue are no longer valid because the next instruction will come from a different address. The BIU discards the queue contents and begins fetching from the new address. This queue flush is the main overhead associated with branch instructions in 8086 programs.

Comparison with 8085

The 8085 microprocessor does not have an instruction queue. It uses a simple fetch-decode-execute cycle sequentially. Every instruction requires a memory fetch before execution can begin. The 8086 instruction queue is a key architectural improvement that gives the 8086 better throughput even at the same clock frequency. This is a commonly tested comparison in university exams.

Numerical Example

The efficiency gain from pipelining can be estimated by comparing sequential and pipelined execution times. Assume each bus cycle (fetch or execute) takes one unit of time. For a sequence of three single-byte instructions, sequential execution takes 6 units (3 fetch + 3 execute). With the instruction queue, execution overlaps with fetch.

Example
Given:
3 instructions, each requiring 1 unit fetch time and 1 unit execute time.

Why this formula applies:
With an instruction queue, BIU fetches next instruction while EU executes current one.

Formula:
Sequential time = N x (Fetch + Execute)
Pipelined time = Fetch1 + N x Execute (steady-state overlap)

Substitution:
Sequential: 3 x (1 + 1) = 6 units
Pipelined: 1 + 3 x 1 = 4 units

Calculation:
Speedup = Sequential / Pipelined = 6 / 4 = 1.5x

Final Answer: Instruction queue provides 1.5x throughput improvement for this sequence.
Exam Tip: The 8086 instruction queue is 6 bytes. The 8088 (8-bit external bus version) has a 4-byte queue. Queue is flushed on every branch/jump instruction. This flush penalty is why loops and branch-heavy code slightly reduces pipeline efficiency.

Mechanism: Queue Fill and Flush Cycle

Instruction Queue: Fill, Consume, and Flush BehaviorTimeline: BIU Fetch vs EU ExecuteBIU and EU operate concurrently on separate tasksBIU: Fetch Instr 1 from memoryBIU: Fetch Instr 2 from memoryEU: Idle (waiting first cycle)EU: Execute Instr 1 (from queue)Steady state: BIU pre-fetches next while EU executes current — overlap achievedBranch instruction detected: Queue FLUSHED, BIU refetches from branch target addressBIU refills queue from new address | EU waits (pipeline bubble) | Normal overlap resumes after refillThis is why JUMP instructions cause a performance penalty in 8086
Figure 2: Instruction queue operation — steady-state fetch-execute overlap and branch-caused flush behavior
  • BIU fills the 6-byte FIFO queue whenever 2 or more bytes are empty and the EU does not need the bus for data.
  • EU reads from the queue directly without accessing external memory, eliminating most fetch wait times.
  • BIU fetches 2 bytes at a time using the full 16-bit external data bus.
  • On a branch, the queue is flushed and BIU immediately begins refetching from the branch target. This creates a brief pipeline bubble.
  • The 8088 uses a 4-byte queue because its external bus is only 8 bits wide, fetching 1 byte per cycle.

Quick Revision

  • 8086 has a 6-byte instruction queue (FIFO); 8088 has a 4-byte queue.
  • BIU fills the queue; EU consumes from it — both operate concurrently (2-stage pipeline).
  • Queue is flushed on every branch, call, or jump instruction — this is the main pipeline penalty.
  • BIU refills queue when 2 or more bytes are free and EU does not need the bus.
  • 8085 has no instruction queue — sequential fetch-execute only.
  • Speedup formula: Pipelined time = Fetch1 + N x Execute (for N-instruction sequences).
  • Queue status signals QS0 and QS1 in maximum mode tell external coprocessors the queue state.

Instruction Queue Mechanics

Examine pipelining and prefetch behavior.

Question 1 of 3

Q1.What is the maximum length of the instruction prefetch queue in the 8086 microprocessor?