Scaled Dot-Product Attention
A compact note on the attention equation, why scaling matters, and how to read the tensor shapes.
Scaled Dot-Product Attention
Scaled Dot-Product Attention
The core attention operation maps queries, keys, and values into a weighted mixture of value vectors.
Here, asks what each token is looking for, describes what each token offers, and carries the information that gets mixed after the attention weights are computed.
Shape Check
For a sequence length and key dimension :
| Symbol | Shape | Meaning |
|---|---|---|
| Query vectors | ||
| Key vectors | ||
| Value vectors | ||
| Pairwise token scores |
The division by keeps the logits from growing too large as the dimension increases.
Implementation Note
In code, the softmax is usually applied across the last dimension of the score matrix:
const scores = q @ transpose(k);
const weights = softmax(scores / Math.sqrt(dK));
const output = weights @ v;Related notes
How LLMs Work
A plain-language walkthrough of what actually happens between typing a prompt and a large language model typing back an answer: tokens, embeddings, attention, and why they sometimes make things up.
Backpropagation, Step by Step
How the chain rule turns a network's forward pass into gradients for every weight, and why the "backward" direction is what makes training tractable.
Transformers vs. RNNs
Why self-attention replaced recurrence: the parallelism transformers unlock, the quadratic cost they pay for it, and what RNNs got right.
Newsletter
New articles in your inbox
An email when I publish something new. No spam, unsubscribe anytime.
Double opt-in. See the privacy policy for how your email is handled.