Lesson 2 / 25

Neurons and Activation Functions

Weighted sum plus non-linearity.

A tiny function with learnable weights

An artificial neuron multiplies each input by a weight, adds a bias, and passes the sum through an activation function. Common activations: ReLU (max(0, z)), the default in most hidden layers because it is simple and trains well; sigmoid (squashes to 0 to 1, used for binary probabilities); tanh (-1 to 1); GELU and similar smooth variants common in transformers. The activation is what makes networks non-linear: without it, stacked layers collapse into one linear function. Training means adjusting all weights and biases.

One neuron by hand, run

I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. Three inputs, three weights and a bias give a weighted sum of -0.720. ReLU turns it into 0.000; sigmoid gives 0.327.

import numpy as np
x = np.array([0.5, -1.2, 3.0])          # inputs
w = np.array([0.8, 0.1, -0.4])          # weights
b = 0.2                                 # bias
z = w @ x + b
relu = max(0.0, z)
sigmoid = 1 / (1 + np.exp(-z))
print(f"weighted sum z = {z:.3f}")
print(f"ReLU(z) = {relu:.3f} | sigmoid(z) = {sigmoid:.3f}")

Output:

weighted sum z = -0.720
ReLU(z) = 0.000 | sigmoid(z) = 0.327

Use ReLU (or GELU) in hidden layers

Start with ReLU for hidden layers and choose the output activation from the task: none for regression, sigmoid for binary, softmax for multi-class.

Quick check: Why are activation functions needed?

  • They make the network non-linear; without them stacked layers equal one linear layer
  • They store the training data
  • They speed up disk access
  • They replace the weights
Answer

They make the network non-linear; without them stacked layers equal one linear layer — Non-linearity gives networks their power.