Transformers vs. RNNs
Transformers vs. RNNs
Why self-attention replaced recurrence: the parallelism transformers unlock, the quadratic cost they pay for it, and what RNNs got right.
~/articles
Technical writing on artificial intelligence, software architecture, and the practical details behind building modern products.
Search and filter
Transformers vs. RNNs
Why self-attention replaced recurrence: the parallelism transformers unlock, the quadratic cost they pay for it, and what RNNs got right.
The distributional hypothesis, the skip-gram objective, and why vector arithmetic on words like king minus man plus woman approximates queen.
Word Embeddings and Why They Work
How convolution, pooling, and stacked layers let a network build up from edges to full objects with far fewer parameters than a fully-connected net.
Convolutional Neural Networks, Explained
How the chain rule turns a network's forward pass into gradients for every weight, and why the "backward" direction is what makes training tractable.
Backpropagation, Step by Step
Why almost every model learns by walking downhill: the loss landscape, the role of the learning rate, and how mini-batches change the walk.
Gradient Descent, Intuitively
A plain-language walkthrough of what actually happens between typing a prompt and a large language model typing back an answer: tokens, embeddings, attention, and why they sometimes make things up.
How LLMs Work
A beginner-friendly walkthrough of what artificial intelligence actually is, how it differs from ordinary software, and where AI, machine learning, deep learning, and LLMs fit together.
What Is AI?
A compact note on the attention equation, why scaling matters, and how to read the tensor shapes.
Scaled Dot-Product Attention
Newsletter
An email when I publish something new. No spam, unsubscribe anytime.
Double opt-in. See the privacy policy for how your email is handled.