पाठ 21 / 25
Memory-Efficient Structures: slots, array, generators and NumPy
Store large amounts of data compactly.
When millions of objects add up
Python objects carry overhead: every int is a full object, every instance normally has a __dict__, and a list stores pointers to separately allocated objects. At scale, this matters. __slots__ (or @dataclass(slots=True)) removes the per-instance __dict__, cutting memory per object substantially and speeding attribute access. The array module stores homogeneous numbers as raw C values (array('i', ...), array('d', ...)), typically several times smaller than a list of Python ints or floats. NumPy arrays go further: compact typed memory plus vectorised operations implemented in C, often 10 to 100 times faster than Python loops for numeric work; pandas and Polars build columnar tables on similar ideas. Generators and iterators (yield, generator expressions, itertools) process data lazily, one item at a time, so memory stays constant regardless of input size. bytes/bytearray/memoryview handle binary data without copies. Also consider interning repeated strings with sys.intern, storing data in SQLite or files when it does not need to be in memory, and choosing integer IDs over long strings as keys.
Comparing memory footprints
Slots, array and a generator pipeline.
import sys
import tracemalloc
from array import array
from dataclasses import dataclass
@dataclass
class ReadingDict:
sensor: int
value: float
@dataclass(slots=True)
class ReadingSlots:
sensor: int
value: float
def peak_mb(factory, n=200_000):
tracemalloc.start()
items = [factory(i % 100, i * 0.5) for i in range(n)]
_, peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
return round(peak / 1e6, 1)
print("with __dict__:", peak_mb(ReadingDict), "MB")
print("with slots: ", peak_mb(ReadingSlots), "MB") # noticeably smaller
floats_list = [i * 0.5 for i in range(1_000_000)]
floats_array = array("d", floats_list) # 8 bytes per value, no float objects
print(sys.getsizeof(floats_array) / 1e6, "MB for the array")
def read_values(path): # generator: one line at a time
with open(path) as f:
for line in f:
yield float(line.split(",")[1])
# total = sum(read_values("readings.csv")) # constant memory, whatever the file size
# NumPy (third-party): compact and vectorised
# import numpy as np
# values = np.arange(1_000_000) * 0.5
# print(values.nbytes / 1e6, "MB", values.mean())Optimise memory where volume is high
Slots and arrays matter when you hold hundreds of thousands or millions of items. For a few hundred objects, clarity matters more than bytes.
त्वरित जाँच: What does __slots__ (or dataclass(slots=True)) save?
- Disk space
- CPU cores
- The per-instance __dict__, reducing memory and speeding attribute access
- Import time
Answer
The per-instance __dict__, reducing memory and speeding attribute access — Slotted classes store attributes in fixed slots instead of a dictionary.