Lesson 21 / 25

Memory-Efficient Structures: slots, array, generators and NumPy

Store large amounts of data compactly.

When millions of objects add up

Python objects carry overhead: every int is a full object, every instance normally has a __dict__, and a list stores pointers to separately allocated objects. At scale, this matters. __slots__ (or @dataclass(slots=True)) removes the per-instance __dict__, cutting memory per object substantially and speeding attribute access. The array module stores homogeneous numbers as raw C values (array('i', ...), array('d', ...)), typically several times smaller than a list of Python ints or floats. NumPy arrays go further: compact typed memory plus vectorised operations implemented in C, often 10 to 100 times faster than Python loops for numeric work; pandas and Polars build columnar tables on similar ideas. Generators and iterators (yield, generator expressions, itertools) process data lazily, one item at a time, so memory stays constant regardless of input size. bytes/bytearray/memoryview handle binary data without copies. Also consider interning repeated strings with sys.intern, storing data in SQLite or files when it does not need to be in memory, and choosing integer IDs over long strings as keys.

Comparing memory footprints

Slots, array and a generator pipeline.

import sys
import tracemalloc
from array import array
from dataclasses import dataclass

@dataclass
class ReadingDict:
    sensor: int
    value: float

@dataclass(slots=True)
class ReadingSlots:
    sensor: int
    value: float

def peak_mb(factory, n=200_000):
    tracemalloc.start()
    items = [factory(i % 100, i * 0.5) for i in range(n)]
    _, peak = tracemalloc.get_traced_memory()
    tracemalloc.stop()
    return round(peak / 1e6, 1)

print("with __dict__:", peak_mb(ReadingDict), "MB")
print("with slots:   ", peak_mb(ReadingSlots), "MB")      # noticeably smaller

floats_list = [i * 0.5 for i in range(1_000_000)]
floats_array = array("d", floats_list)                     # 8 bytes per value, no float objects
print(sys.getsizeof(floats_array) / 1e6, "MB for the array")

def read_values(path):                                     # generator: one line at a time
    with open(path) as f:
        for line in f:
            yield float(line.split(",")[1])

# total = sum(read_values("readings.csv"))  # constant memory, whatever the file size

# NumPy (third-party): compact and vectorised
# import numpy as np
# values = np.arange(1_000_000) * 0.5
# print(values.nbytes / 1e6, "MB", values.mean())

Optimise memory where volume is high

Slots and arrays matter when you hold hundreds of thousands or millions of items. For a few hundred objects, clarity matters more than bytes.

Quick check: What does __slots__ (or dataclass(slots=True)) save?

  • Disk space
  • CPU cores
  • The per-instance __dict__, reducing memory and speeding attribute access
  • Import time
Answer

The per-instance __dict__, reducing memory and speeding attribute access — Slotted classes store attributes in fixed slots instead of a dictionary.