# Mini-Batches, Epochs and Devices — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/w-batches

> How data flows and where computation happens.

## Batches trade noise for speed

Training on **mini-batches** (for example 32 to 512 examples) gives a noisy but cheap estimate of the gradient and lets hardware process many examples in parallel. An **epoch** is one pass over the training data. Larger batches are faster per epoch on GPUs but may need a higher learning rate and can generalise slightly differently. Move the model and each batch to the same **device** (`model.to(device)`, `x.to(device)`); mixing CPU and GPU tensors is a common error. Shuffle training data every epoch but not validation data.

## Device-agnostic training code (sketch)

Works on CPU or GPU unchanged. Not run here.

```python
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
for xb, yb in train_loader:
    xb, yb = xb.to(device), yb.to(device)
    optimizer.zero_grad()
    loss = loss_fn(model(xb), yb)
    loss.backward()
    optimizer.step()
```

## Use pin_memory and workers for speed

For GPU training, DataLoader(num_workers=..., pin_memory=True) keeps the GPU fed; measure before tuning.

**Quiz:** What is an epoch?

- [x] One full pass over the training data
- [ ] One weight update
- [ ] One layer
- [ ] One GPU

*Answer:* One full pass over the training data. Many mini-batch steps make up an epoch.
