# Memory-Efficient Structures: slots, array, generators and NumPy — Data Structures in Python

Source: https://www.skillbyai.com/en/data-structures-python/a-memory

> Store large amounts of data compactly.

## When millions of objects add up

Python objects carry overhead: every `int` is a full object, every instance normally has a `__dict__`, and a list stores pointers to separately allocated objects. At scale, this matters. **`__slots__`** (or `@dataclass(slots=True)`) removes the per-instance `__dict__`, cutting memory per object substantially and speeding attribute access. The **`array`** module stores homogeneous numbers as raw C values (`array('i', ...)`, `array('d', ...)`), typically several times smaller than a list of Python ints or floats. **NumPy** arrays go further: compact typed memory plus **vectorised** operations implemented in C, often 10 to 100 times faster than Python loops for numeric work; **pandas** and **Polars** build columnar tables on similar ideas. **Generators** and iterators (`yield`, generator expressions, `itertools`) process data lazily, one item at a time, so memory stays constant regardless of input size. **`bytes`/`bytearray`/`memoryview`** handle binary data without copies. Also consider **interning** repeated strings with `sys.intern`, storing data in **SQLite** or files when it does not need to be in memory, and choosing integer IDs over long strings as keys.

## Comparing memory footprints

Slots, array and a generator pipeline.

```python
import sys
import tracemalloc
from array import array
from dataclasses import dataclass

@dataclass
class ReadingDict:
    sensor: int
    value: float

@dataclass(slots=True)
class ReadingSlots:
    sensor: int
    value: float

def peak_mb(factory, n=200_000):
    tracemalloc.start()
    items = [factory(i % 100, i * 0.5) for i in range(n)]
    _, peak = tracemalloc.get_traced_memory()
    tracemalloc.stop()
    return round(peak / 1e6, 1)

print("with __dict__:", peak_mb(ReadingDict), "MB")
print("with slots:   ", peak_mb(ReadingSlots), "MB")      # noticeably smaller

floats_list = [i * 0.5 for i in range(1_000_000)]
floats_array = array("d", floats_list)                     # 8 bytes per value, no float objects
print(sys.getsizeof(floats_array) / 1e6, "MB for the array")

def read_values(path):                                     # generator: one line at a time
    with open(path) as f:
        for line in f:
            yield float(line.split(",")[1])

# total = sum(read_values("readings.csv"))  # constant memory, whatever the file size

# NumPy (third-party): compact and vectorised
# import numpy as np
# values = np.arange(1_000_000) * 0.5
# print(values.nbytes / 1e6, "MB", values.mean())
```

## Optimise memory where volume is high

Slots and arrays matter when you hold hundreds of thousands or millions of items. For a few hundred objects, clarity matters more than bytes.

**Quiz:** What does __slots__ (or dataclass(slots=True)) save?

- [ ] Disk space
- [ ] CPU cores
- [x] The per-instance __dict__, reducing memory and speeding attribute access
- [ ] Import time

*Answer:* The per-instance __dict__, reducing memory and speeding attribute access. Slotted classes store attributes in fixed slots instead of a dictionary.
