पाठ 10 / 25
Mini-Batches, Epochs and Devices
How data flows and where computation happens.
Batches trade noise for speed
Training on mini-batches (for example 32 to 512 examples) gives a noisy but cheap estimate of the gradient and lets hardware process many examples in parallel. An epoch is one pass over the training data. Larger batches are faster per epoch on GPUs but may need a higher learning rate and can generalise slightly differently. Move the model and each batch to the same device (model.to(device), x.to(device)); mixing CPU and GPU tensors is a common error. Shuffle training data every epoch but not validation data.
Device-agnostic training code (sketch)
Works on CPU or GPU unchanged. Not run here.
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
for xb, yb in train_loader:
xb, yb = xb.to(device), yb.to(device)
optimizer.zero_grad()
loss = loss_fn(model(xb), yb)
loss.backward()
optimizer.step()Use pin_memory and workers for speed
For GPU training, DataLoader(num_workers=..., pin_memory=True) keeps the GPU fed; measure before tuning.
त्वरित जाँच: What is an epoch?
- One full pass over the training data
- One weight update
- One layer
- One GPU
Answer
One full pass over the training data — Many mini-batch steps make up an epoch.