# Recurrent Networks and Their Limits — Deep Learning & Neural Networks

Source: https://www.skillbyai.com/en/deep-learning/q-rnn

> Reading one step at a time.

## State carried through time

A **recurrent neural network (RNN)** processes a sequence one element at a time, updating a hidden state that summarises what it has seen. **LSTM** and **GRU** cells add gates that decide what to keep or forget, easing (but not removing) the problem of remembering long-range information. RNNs are compact and still useful for some streaming and small time-series problems, but they process tokens sequentially (hard to parallelise) and struggle with very long dependencies, which is why **transformers** replaced them for most language tasks.

## Reading through a keyhole

An RNN is like reading a book one word at a time and keeping only a short note of what came before; attention lets you flip back to any page whenever you need it.

## Consider simple baselines for time series

For forecasting, compare RNNs with seasonal baselines and gradient boosting on lag features before committing to deep models.

**Quiz:** Why did transformers replace RNNs for most language tasks?

- [x] They process tokens in parallel and handle long-range context better
- [ ] RNNs cannot learn at all
- [ ] Transformers need no data
- [ ] RNNs only work on images

*Answer:* They process tokens in parallel and handle long-range context better. Parallel attention scales better.
