1.

Given the program:

1:  addi  x5,  x0, 1       # i = 1  
2:  slli  x21, x5, 3       # n = i · 2³ = 8  
3:  Loop: bge   x5, x21, Exit  # if i ≥ n, exit  
4:         slli  x6,  x5, 2   # x6 = i · 4  
5:         add   x7,  x22, x6 # x7 = &B[i]  
6:         lw    x9,  0(x7)   # x9 = B[i]  
7:         slli  x10, x9, 2   # x10 = 4·B[i]  
8:         sw    x10, 0(x7)   # B[i] = x10  
9:         addi  x5,  x5, 1   # i++  
10:        beq   x0,  x0, Loop  # unconditional back  
11: Exit:
  1. Iterations
    • a) The loop body (lines 4–9) runs for through , so 7 iterations.
  2. Conditional‑branch executions
    • b) Line 3 executes once per iteration plus once more at exit, so 8 times.
  3. Body instructions
    • c) Lines 4–9 are 6 instructions/iteration ⇒ .
  4. Unconditional branch
    • d) Line 10 runs once per iteration ⇒ 7 times.
  5. Total dynamic instructions
    • e) instructions.

2.

  • CPI = 1, cycle‑time =
  • Total instructions = 59
  • Time =

3.

a) Data & Control Hazards

BetweenHazard TypeRegister
1 & 2RAWx5
1 & 3RAWx5
2 & 3RAWx21
3 & 4none—
4 & 5RAWx6
5 & 6RAWx7
6 & 7Load–Use RAWx9
7 & 8RAWx10
9 & 3 (next)RAWx5
10 & 3 (next)Control hazard—

b) Pipeline Diagram

Instruction1234567891011121314151617181920
I1 (L1)FDEMW
I2 (L2)FDEMW
I3 (L3: bge)FDEMW
B1 (flush)FDEMW
B2 (flush)FDEMW
I4 (L4)FDEMW
I5 (L5)FDEMW
I6 (L6)FDEMW
B3 (stall)FDEMW
I7 (L7)FDEMW
I8 (L8)FDEMW
I9 (L9)FDEMW
I10 (L10)FDEMW
B4 (flush)FDEMW
B5 (flush)FDEMW
I11 (exit)FDEMW
  • F = IF
  • D = ID
  • E = EX
  • M = MEM
  • W = WB
  • N = pipeline bubble

c) Calculations

  1. CPI

    • Instructions = 8 (lines 3–10)
    • Bubbles = 2 (cond‑branch) + 1 (load‑use) + 2 (uncond‑branch) = 5
  2. Total cycles

    • Pipeline fill = 5 stages − 1 = 4 cycles
    • Iterations = 7
  3. Execution time

  4. Speedup


d)

i. Reordered loop:

Loop:
    bge   x5, x21, Exit     # line 3
    slli  x6, x5, 2         # line 4
    add   x7, x22, x6       # line 5
    lw    x9, 0(x7)         # line 6
    addi  x5, x5, 1         # line 9 (moved up)
    beq   x0, x0, Loop      # line 10 (moved up)
    slli  x10, x9, 2        # line 7
    sw    x10, 0(x7)        # line 8
Exit:
  • By moving the addi and beq into the load‑use slot, we eliminate that 1‑cycle stall.
    ii. performance
  • Instrs∕iter = 8
  • Bubbles∕iter = 2 (cond) + 2 (uncond) = 4
  • CPI∕iter = 8 + 4 = 12
  • Total = 4 + 7×12 = 88 cycles ⇒ 88 ns
  • Speedup =