Skip to main content

Attention explained with real numbers

How the attention mechanism in a GPT works, step by step, with a worked example you can check by hand and the real numbers from a trained model.

Updated 10 Oct 2026Numbers checked by running the code

Attention is how a GPT decides which earlier words (or letters) matter for the one it is reading now. It is the idea that made GPT models possible, and it comes down to three steps: score, softmax, blend. This page works through all three with small numbers you can check by hand, then shows what they look like inside a real trained model.

The examples come from microgpt, Andrej Karpathy's 200-line GPT, which reads names one letter at a time.

Think of every earlier letter as a book on a shelf. Each book has:

  • a key: the label on its spine (what it is about);
  • a value: what is inside (the information it can share).

The letter you are reading now has a query: the question it wants answered. Attention compares the query with every key, then takes home a mix of every book's contents, with more from the books whose labels match best. Nothing is picked outright. It is a soft, weighted lookup.

In a GPT, the query, key and value are all lists of numbers (vectors). Each letter makes its own three vectors by multiplying its embedding by three learned matrices.

Step 1: score each key

The match between the query and a key is their dot product: multiply them position by position and add up. Pointing the same way gives a big number; pointing apart gives a small or negative one.

Then divide by the square root of the vector's length. With vectors of 2 numbers, that is √2. Without this, longer vectors would give bigger scores, and the next step would become too sure of itself.

A tiny example: query q = [1, 0] and three keys:

keydot product with q÷ √2
[2, 0]21.414
[0, 2]00
[1, 1]10.707

Step 2: softmax the scores into weights

Softmax turns any list of numbers into positive weights that add up to 1 (100%). Raise e to each score, then divide by the total:

  • e^1.414 ≈ 4.11, e^0 = 1, e^0.707 ≈ 2.03, total ≈ 7.14
  • weights ≈ 0.576, 0.140 and 0.284

The first key matched best, so it gets the most attention: about 58%.

Step 3: blend the values

Each key comes with a value. Say the three values are [10, 0], [0, 10] and [5, 5]. The output is the weighted mix, worked out one number at a time:

  • first number: 0.576 × 10 + 0.140 × 0 + 0.284 × 5 = 5.76 + 0 + 1.42 = 7.18
  • second number: 0.576 × 0 + 0.140 × 10 + 0.284 × 5 = 0 + 1.40 + 1.42 = 2.82

So the output is [7.18, 2.82]: mostly the first value, because its key matched best. Like mixing paint in those proportions.

The same thing in Python

import math

def attend(q, keys, values):
    scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(len(q)) for k in keys]
    top = max(scores)
    exps = [math.exp(s - top) for s in scores]
    weights = [e / sum(exps) for e in exps]
    return [sum(w * v[j] for w, v in zip(weights, values)) for j in range(len(values[0]))]

print(attend([1, 0], [[2, 0], [0, 2], [1, 1]], [[10, 0], [0, 10], [5, 5]]))
# about [7.18, 2.82]

Subtracting the largest score before exp does not change the weights. It only stops very large scores from overflowing.

Only looking backwards

When a GPT reads text, a letter may attend only to itself and the letters before it, never the future. microgpt gets this for free: it reads one letter at a time and keeps the keys and values seen so far in a list (the KV cache), so the future simply is not there yet. Bigger models read a whole passage at once and hide the future with a mask instead.

What a trained model actually does

In microgpt, each letter's vector has 16 numbers, split into 4 heads of 4 numbers each. Each head runs its own score, softmax and blend, so a letter can look for several kinds of thing at once. With 4-number heads, the scores are divided by √4 = 2.

Real numbers from the trained model, reading the name "emma": at the final "a", head 1 scores the earlier positions as follows (before softmax).

positionstart markeremma
score0.080.24−0.11−0.23−0.20

A model this small, with one layer and 1,000 training steps, has fairly fuzzy heads. They spread attention over recent letters, the start marker and the letter itself, rather than specialising neatly. In big GPTs, with dozens of layers and far more training, heads specialise much more clearly.

Where attention sits in a GPT

For each letter, microgpt runs: embedding → normalise → attention → add the result back (a residual connection) → a small neural network (the MLP) → 27 scores, one per possible next letter. Attention's job is to pack what the earlier letters say into this letter's vector before the guess is made. Without it, the guess after "emm" could only see the last "m".

Build it yourself

Zero to microGPT builds attention from scratch in Module 8, in your browser, with no installs:

  1. Position embeddings: where am I in the word?
  2. Queries, keys, values: the library search.
  3. Scaled dot-product attention: score, softmax, blend.
  4. Causal masking and the KV cache: only look backwards.
  5. RMSNorm and residual connections: keeping it trainable.

Each lesson has a short film, a heatmap of the real trained model's attention you can explore, and a Python lab that checks your code.

Common questions

Is "self-attention" different from attention? Self-attention means the queries, keys and values all come from the same text, which is the case in a GPT. The three steps are the same.

Why divide by the square root? Dot products of longer vectors are bigger on average. Dividing by √(length) keeps the scores in a range where softmax still spreads its attention instead of putting nearly all of it on one item.

Do I need advanced maths for this? No. You need multiplication, addition, exp and the idea of a weighted average. The course teaches each of those from scratch, in Module 4.