Lesson 18 / 25
Recurrent Networks and Their Limits
Reading one step at a time.
State carried through time
A recurrent neural network (RNN) processes a sequence one element at a time, updating a hidden state that summarises what it has seen. LSTM and GRU cells add gates that decide what to keep or forget, easing (but not removing) the problem of remembering long-range information. RNNs are compact and still useful for some streaming and small time-series problems, but they process tokens sequentially (hard to parallelise) and struggle with very long dependencies, which is why transformers replaced them for most language tasks.
Reading through a keyhole
An RNN is like reading a book one word at a time and keeping only a short note of what came before; attention lets you flip back to any page whenever you need it.
Consider simple baselines for time series
For forecasting, compare RNNs with seasonal baselines and gradient boosting on lag features before committing to deep models.
Quick check: Why did transformers replace RNNs for most language tasks?
- They process tokens in parallel and handle long-range context better
- RNNs cannot learn at all
- Transformers need no data
- RNNs only work on images
Answer
They process tokens in parallel and handle long-range context better — Parallel attention scales better.