Lesson 6 / 25

Vectorisation Instead of Loops

Push the loop into C.

Why loops are slow

A Python loop over a million numbers pays interpreter overhead for every element; a vectorised call such as np.dot or (x * x).sum() loops once in compiled code. Rewrite loops as array expressions, masks and reductions; use np.where for conditional values. The results may differ in the last few bits because the order of additions changes, so compare with isclose.

A sum of squares, looped and vectorised, run

I ran this with Python 3.12.3 and NumPy 2.5.3. The two results are not bit-for-bit equal but are close; across three repeats the vectorised version was more than 20 times faster on my machine (exact timings vary, so only the comparison is shown).

import numpy as np, timeit

x = np.arange(1_000_000, dtype=np.float64)

def loop():
    total = 0.0
    for v in x:
        total += v * v
    return total

def vec():
    return float(np.dot(x, x))

print(loop() == vec(), np.isclose(loop(), vec()))   # summation order differs
print(f"{vec():.6e}")
t_loop = min(timeit.repeat(loop, number=1, repeat=3))
t_vec = min(timeit.repeat(vec, number=1, repeat=3))
print("vectorised faster by more than 20x:", t_loop / t_vec > 20)

Output:

False True
3.333328e+17
vectorised faster by more than 20x: True

One bulk order, not a million trips

A loop is a million trips to the shop for one item each; a vectorised call is a single bulk order.

Quick check: Why can a looped sum and np.dot differ slightly?

  • Loops round to integers
  • np.dot is wrong
  • Python ints are inexact
  • They add the numbers in a different order, and floating-point addition is not exactly associative
Answer

They add the numbers in a different order, and floating-point addition is not exactly associative — Compare floats with a tolerance.