Transformers vs. RNNs
Why self-attention replaced recurrence: the parallelism transformers unlock, the quadratic cost they pay for it, and what RNNs got right.
Transformers vs. RNNs
Transformers vs. RNNs
Before transformers, sequence models like RNNs and LSTMs were the default way to process text. Understanding why they were replaced — and what they actually did well — clarifies why the transformer architecture looks the way it does.
RNNs: Sequential by Construction
A recurrent neural network processes a sequence one token at a time, carrying a hidden state forward:
Each hidden state is a compressed summary of everything seen so far. This is elegant and memory-efficient, but it creates two hard problems:
| Problem | Why it happens |
|---|---|
| Vanishing/exploding gradients | Backpropagating through many time steps repeatedly multiplies by the same weight matrix |
| No parallelism across time | can't be computed until exists, so training can't parallelize over sequence length |
LSTMs and GRUs mitigated the gradient problem with gating mechanisms, but the sequential bottleneck remained — training on long documents was slow simply because the computation couldn't be parallelized across the time dimension.
Self-Attention: Parallel by Construction
Transformers replace recurrence with self-attention: every token looks directly at every other token in a single step, rather than through a chain of hidden states.
Because this computation has no dependency on a previous time step, every token's representation can be computed simultaneously on a GPU. That's the single biggest practical reason transformers overtook RNNs: given the same hardware, they use it far more efficiently.
The Tradeoff: Compute vs. Memory
| Aspect | RNN | Transformer |
|---|---|---|
| Per-step compute | per token | per token (attends to all tokens) |
| Total compute for length | ||
| Parallelizable across sequence | No | Yes |
| Long-range dependencies | Degrade with distance | Direct, constant-distance access |
Self-attention's cost is the price paid for direct access to every token — an RNN has to route information for a long-range dependency through every intermediate hidden state, while a transformer connects any two tokens with a single attention weight. That quadratic cost is also why long-context transformers need techniques like sparse attention or sliding windows: the direct-access advantage becomes expensive at very large .
Why This Mattered in Practice
# RNN: must process sequentially
for t in range(seq_len):
h[t] = rnn_cell(h[t-1], x[t])
# Transformer: all positions computed in one batched operation
attention_output = self_attention(Q, K, V) # shape: (seq_len, d_model)The RNN loop is inherently serial in Python and on hardware alike. The transformer's attention call is a handful of matrix multiplications over the whole sequence at once — exactly the kind of operation GPUs are built to accelerate. Given that training LLMs is bottlenecked by how much data can be processed on available hardware, that parallelism, more than any single accuracy improvement, is why transformers became the default architecture for language modeling.
Related notes
Word Embeddings and Why They Work
The distributional hypothesis, the skip-gram objective, and why vector arithmetic on words like king minus man plus woman approximates queen.
How LLMs Work
A plain-language walkthrough of what actually happens between typing a prompt and a large language model typing back an answer: tokens, embeddings, attention, and why they sometimes make things up.
Scaled Dot-Product Attention
A compact note on the attention equation, why scaling matters, and how to read the tensor shapes.
Newsletter
New articles in your inbox
An email when I publish something new. No spam, unsubscribe anytime.
Double opt-in. See the privacy policy for how your email is handled.