Lesson 6 / 25
Vectorisation Instead of Loops
Push the loop into C.
Why loops are slow
A Python loop over a million numbers pays interpreter overhead for every element; a vectorised call such as np.dot or (x * x).sum() loops once in compiled code. Rewrite loops as array expressions, masks and reductions; use np.where for conditional values. The results may differ in the last few bits because the order of additions changes, so compare with isclose.
A sum of squares, looped and vectorised, run
I ran this with Python 3.12.3 and NumPy 2.5.3. The two results are not bit-for-bit equal but are close; across three repeats the vectorised version was more than 20 times faster on my machine (exact timings vary, so only the comparison is shown).
import numpy as np, timeit
x = np.arange(1_000_000, dtype=np.float64)
def loop():
total = 0.0
for v in x:
total += v * v
return total
def vec():
return float(np.dot(x, x))
print(loop() == vec(), np.isclose(loop(), vec())) # summation order differs
print(f"{vec():.6e}")
t_loop = min(timeit.repeat(loop, number=1, repeat=3))
t_vec = min(timeit.repeat(vec, number=1, repeat=3))
print("vectorised faster by more than 20x:", t_loop / t_vec > 20)
Output:
False True 3.333328e+17 vectorised faster by more than 20x: True
One bulk order, not a million trips
A loop is a million trips to the shop for one item each; a vectorised call is a single bulk order.
Quick check: Why can a looped sum and np.dot differ slightly?
- Loops round to integers
- np.dot is wrong
- Python ints are inexact
- They add the numbers in a different order, and floating-point addition is not exactly associative
Answer
They add the numbers in a different order, and floating-point addition is not exactly associative — Compare floats with a tolerance.