# Learning From Examples: The Perceptron — Artificial Intelligence

Source: https://www.skillbyai.com/en/artificial-intelligence/l-perceptron

> The ancestor of neural networks.

## Adjust weights when wrong

Machine learning lets a system improve from **examples** instead of rules. The **perceptron** (1958) is the simplest learning classifier: it predicts 1 if a weighted sum of inputs plus a bias is positive, and whenever it makes a mistake it nudges the weights towards the correct answer. If the classes can be separated by a straight line, it is guaranteed to converge. Stacking many such units with non-linear activations and training them with gradient descent gives the neural networks behind modern deep learning (covered in depth in the Machine Learning and Deep Learning courses).

## From examples and from rewards

Learning lets agents improve from data and experience instead of hand-written knowledge.

![Three ideas: the perceptron, planning with MDPs, Q-learning.](assets/figures/artificial-intelligence/section-6-map.svg) — Figure 6.1 — Perceptron, MDPs and Q-learning.

## Training a perceptron by hand, run

I ran this with plain Python 3 (standard library only), with fixed random seeds where randomness is used. Six labelled points are separable, so the perceptron converges after 9 epochs with weights [-1, 4] and bias -7. It then classifies the new point (4, 4) as 1 and (1, 0) as 0.

```python
data = [((2, 3), 1), ((3, 3), 1), ((4, 5), 1), ((1, 1), 0), ((2, 1), 0), ((1, 2), 0)]
w = [0.0, 0.0]; b = 0.0
for epoch in range(1, 11):
    errors = 0
    for (x1, x2), y in data:
        pred = 1 if w[0] * x1 + w[1] * x2 + b > 0 else 0
        if pred != y:
            errors += 1
            w[0] += (y - pred) * x1; w[1] += (y - pred) * x2; b += (y - pred)
    if errors == 0:
        print(f"converged after {epoch} epochs: w = {w}, b = {b}"); break
print("prediction for (4, 4):", 1 if w[0] * 4 + w[1] * 4 + b > 0 else 0)
print("prediction for (1, 0):", 1 if w[0] * 1 + w[1] * 0 + b > 0 else 0)
```

Output:

```
converged after 9 epochs: w = [-1.0, 4.0], b = -7.0
prediction for (4, 4): 1
prediction for (1, 0): 0
```

## Know the limits of linear models

A single perceptron cannot learn XOR-like patterns; hidden layers are needed for that.

**Quiz:** When is the perceptron guaranteed to converge?

- [ ] Only on images
- [ ] Always, on any data
- [x] When the classes are linearly separable
- [ ] Only with one example

*Answer:* When the classes are linearly separable. Separable data guarantees convergence.
