Lesson 2 / 25
Neurons and Activation Functions
Weighted sum plus non-linearity.
A tiny function with learnable weights
An artificial neuron multiplies each input by a weight, adds a bias, and passes the sum through an activation function. Common activations: ReLU (max(0, z)), the default in most hidden layers because it is simple and trains well; sigmoid (squashes to 0 to 1, used for binary probabilities); tanh (-1 to 1); GELU and similar smooth variants common in transformers. The activation is what makes networks non-linear: without it, stacked layers collapse into one linear function. Training means adjusting all weights and biases.
One neuron by hand, run
I ran this on CPU with Python 3, PyTorch 2.14.1, numpy 2.5.3 and scikit-learn 1.9.1, with fixed seeds and one thread. Three inputs, three weights and a bias give a weighted sum of -0.720. ReLU turns it into 0.000; sigmoid gives 0.327.
import numpy as np
x = np.array([0.5, -1.2, 3.0]) # inputs
w = np.array([0.8, 0.1, -0.4]) # weights
b = 0.2 # bias
z = w @ x + b
relu = max(0.0, z)
sigmoid = 1 / (1 + np.exp(-z))
print(f"weighted sum z = {z:.3f}")
print(f"ReLU(z) = {relu:.3f} | sigmoid(z) = {sigmoid:.3f}")
Output:
weighted sum z = -0.720 ReLU(z) = 0.000 | sigmoid(z) = 0.327
Use ReLU (or GELU) in hidden layers
Start with ReLU for hidden layers and choose the output activation from the task: none for regression, sigmoid for binary, softmax for multi-class.
Quick check: Why are activation functions needed?
- They make the network non-linear; without them stacked layers equal one linear layer
- They store the training data
- They speed up disk access
- They replace the weights
Answer
They make the network non-linear; without them stacked layers equal one linear layer — Non-linearity gives networks their power.