Attention explained with real numbers
How the attention mechanism in a GPT works, step by step, with a worked example you can check by hand and the real numbers from a trained model.
Attention is how a GPT decides which earlier words (or letters) matter for the one it is reading now. It is the idea that made GPT models possible, and it comes down to three steps: score, softmax, blend. This page works through all three with small numbers you can check by hand, then shows what they look like inside a real trained model.
The examples come from microgpt, Andrej Karpathy's 200-line GPT, which reads names one letter at a time.
The idea in one picture: a library search
Think of every earlier letter as a book on a shelf. Each book has:
- a key: the label on its spine (what it is about);
- a value: what is inside (the information it can share).
The letter you are reading now has a query: the question it wants answered. Attention compares the query with every key, then takes home a mix of every book's contents, with more from the books whose labels match best. Nothing is picked outright. It is a soft, weighted lookup.
In a GPT, the query, key and value are all lists of numbers (vectors). Each letter makes its own three vectors by multiplying its embedding by three learned matrices.
Step 1: score each key
The match between the query and a key is their dot product: multiply them position by position and add up. Pointing the same way gives a big number; pointing apart gives a small or negative one.
Then divide by the square root of the vector's length. With vectors of 2 numbers, that is √2. Without this, longer vectors would give bigger scores, and the next step would become too sure of itself.
A tiny example: query q = [1, 0] and three keys:
| key | dot product with q | ÷ √2 |
|---|---|---|
| [2, 0] | 2 | 1.414 |
| [0, 2] | 0 | 0 |
| [1, 1] | 1 | 0.707 |
Step 2: softmax the scores into weights
Softmax turns any list of numbers into positive weights that add up to 1 (100%). Raise e to each score, then divide by the total:
- e^1.414 ≈ 4.11, e^0 = 1, e^0.707 ≈ 2.03, total ≈ 7.14
- weights ≈ 0.576, 0.140 and 0.284
The first key matched best, so it gets the most attention: about 58%.
Step 3: blend the values
Each key comes with a value. Say the three values are [10, 0], [0, 10] and [5, 5]. The output is the weighted mix, worked out one number at a time:
- first number: 0.576 × 10 + 0.140 × 0 + 0.284 × 5 = 5.76 + 0 + 1.42 = 7.18
- second number: 0.576 × 0 + 0.140 × 10 + 0.284 × 5 = 0 + 1.40 + 1.42 = 2.82
So the output is [7.18, 2.82]: mostly the first value, because its key matched best. Like mixing paint in those proportions.
The same thing in Python
import math
def attend(q, keys, values):
scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(len(q)) for k in keys]
top = max(scores)
exps = [math.exp(s - top) for s in scores]
weights = [e / sum(exps) for e in exps]
return [sum(w * v[j] for w, v in zip(weights, values)) for j in range(len(values[0]))]
print(attend([1, 0], [[2, 0], [0, 2], [1, 1]], [[10, 0], [0, 10], [5, 5]]))
# about [7.18, 2.82]
Subtracting the largest score before exp does not change the weights. It only stops very large scores from overflowing.
Only looking backwards
When a GPT reads text, a letter may attend only to itself and the letters before it, never the future. microgpt gets this for free: it reads one letter at a time and keeps the keys and values seen so far in a list (the KV cache), so the future simply is not there yet. Bigger models read a whole passage at once and hide the future with a mask instead.
What a trained model actually does
In microgpt, each letter's vector has 16 numbers, split into 4 heads of 4 numbers each. Each head runs its own score, softmax and blend, so a letter can look for several kinds of thing at once. With 4-number heads, the scores are divided by √4 = 2.
Real numbers from the trained model, reading the name "emma": at the final "a", head 1 scores the earlier positions as follows (before softmax).
| position | start marker | e | m | m | a |
|---|---|---|---|---|---|
| score | 0.08 | 0.24 | −0.11 | −0.23 | −0.20 |
A model this small, with one layer and 1,000 training steps, has fairly fuzzy heads. They spread attention over recent letters, the start marker and the letter itself, rather than specialising neatly. In big GPTs, with dozens of layers and far more training, heads specialise much more clearly.
Where attention sits in a GPT
For each letter, microgpt runs: embedding → normalise → attention → add the result back (a residual connection) → a small neural network (the MLP) → 27 scores, one per possible next letter. Attention's job is to pack what the earlier letters say into this letter's vector before the guess is made. Without it, the guess after "emm" could only see the last "m".
Build it yourself
Zero to microGPT builds attention from scratch in Module 8, in your browser, with no installs:
- Position embeddings: where am I in the word?
- Queries, keys, values: the library search.
- Scaled dot-product attention: score, softmax, blend.
- Causal masking and the KV cache: only look backwards.
- RMSNorm and residual connections: keeping it trainable.
Each lesson has a short film, a heatmap of the real trained model's attention you can explore, and a Python lab that checks your code.
Common questions
Is "self-attention" different from attention? Self-attention means the queries, keys and values all come from the same text, which is the case in a GPT. The three steps are the same.
Why divide by the square root? Dot products of longer vectors are bigger on average. Dividing by √(length) keeps the scores in a range where softmax still spreads its attention instead of putting nearly all of it on one item.
Do I need advanced maths for this? No. You need multiplication, addition, exp and the idea of a weighted average. The course teaches each of those from scratch, in Module 4.