Skip to main content

How transformers work, step by step

How the transformer inside GPT models works, from embeddings through attention, the MLP and residual connections to the next-token scores, traced through a real GPT whose 4,192 parameters you can count by hand.

Updated 11 Oct 2026Numbers checked by running the code

The transformer is the neural network design inside GPT, ChatGPT and most of today's language models. It takes the tokens read so far and turns them into one score for every possible next token. Inside, it does two kinds of work, in turns:

  • attention, where each token talks to the earlier tokens and gathers what matters;
  • an MLP, where each token thinks on its own about what it gathered.

This page traces one full pass through a real transformer: the gpt() function in microgpt, Andrej Karpathy's complete GPT in about 200 lines of Python. It reads names one letter at a time, so its vocabulary is 27 tokens (26 letters plus a start/end marker). Every vector in it has 16 numbers. It is small enough to count every parameter, and it has all the parts a big transformer has. (For the bigger picture of how a language model is trained and used, start with how LLMs work.)

The whole pass at a glance

For each token, microgpt runs:

  1. Embed: token embedding + position embedding → 16 numbers.
  2. Normalise (RMSNorm).
  3. Attention block: queries, keys and values; 4 heads; blend; project; add back (residual).
  4. MLP block: 16 → 64 → 16 numbers with a ReLU in between; add back (residual).
  5. Output: 16 numbers → 27 scores, one per possible next token.

Steps 3 and 4 together are one layer. microgpt has one layer. Bigger models repeat the layer many times; GPT-2 small stacks 12.

1. Embeddings: a token becomes a vector

Each token id picks a row from a learned table called wte: 16 numbers that stand for that letter. A second table, wpe, has one 16-number row per position (0 to 15), and the two rows are added together.

The position row is what lets the model tell "ana" from "naa". Attention by itself ignores order, so position has to be built into the vector. Both tables start random and are learned in training, like every other weight. (The first step of all, cutting text into tokens, is in the tokens guide.)

2. RMSNorm: keep the numbers in range

Before each block, the vector is rescaled so that its typical size is about 1. RMSNorm divides every number by the root mean square of the vector. This keeps values from growing or shrinking as they pass through the network, which keeps training stable.

3. Attention: each token looks back

Each token makes three vectors by multiplying its 16 numbers by three learned 16×16 matrices:

  • a query: what it is looking for;
  • a key: what it offers;
  • a value: what it will share.

Its query is compared with the key of every token so far (dot product, divided by the square root of the head size). Softmax turns those scores into weights that add up to 1, and the output is the weighted mix of the values.

A token may only look at itself and earlier tokens. microgpt gets this for free: it reads one token at a time and keeps the keys and values seen so far in a list (the KV cache), so future tokens don't exist yet.

The 16 numbers are split into 4 heads of 4 numbers each, and each head runs its own attention, so a token can look for several things at once. The four head outputs are joined back into 16 numbers and passed through a fourth 16×16 matrix (the output projection).

Attention explained works all of this through by hand, with real scores from the trained model.

4. Residual connections: add, don't replace

The attention output is not used instead of the vector. It is added to it:

x = [a + b for a, b in zip(x, x_residual)]

Each block only has to learn a correction to what is already there. Gradients can also flow straight back along these additions during training, which is a big part of why deep stacks of layers train at all.

5. The MLP: thinking per token

Next, each token's vector goes through a small two-layer neural network on its own, with no looking at other tokens:

  1. normalise (RMSNorm again);
  2. multiply by a 64×16 matrix to get 64 numbers;
  3. apply ReLU: keep positive numbers, set negative ones to 0;
  4. multiply by a 16×64 matrix to get back to 16 numbers;
  5. add the result back (another residual connection).

The ReLU matters: without it, the two matrix multiplications would collapse into one. In the trained microgpt, only about 7 or 8 of the 64 hidden units fire at each position, on average.

6. Output: 27 scores

Finally, a table called lm_head turns the 16 numbers into 27 scores (logits), one per token in the vocabulary. Softmax turns them into chances, and the model samples the next letter from them. After the start marker, the trained model's most likely first letter is "a", at about 14%.

Counting every parameter

partshapeparameters
token embeddings wte27 × 16432
position embeddings wpe16 × 16256
attention: query, key, value, output4 × (16 × 16)1,024
MLP64 × 16 + 16 × 642,048
output lm_head27 × 16432
total4,192

The MLP holds almost half the parameters, which is also true of big GPTs. microgpt has no bias terms, and RMSNorm has no learned weights, so these nine tables are the whole model.

How microgpt differs from GPT-2

The comment in the file says it follows GPT-2 "with minor differences":

  • RMSNorm instead of LayerNorm;
  • no biases;
  • ReLU instead of GeLU.

The layout is the same: embeddings, then layers of attention + MLP with residual connections, then the output scores. GPT-style models are decoder-only transformers: they only look backwards and predict the next token. The original 2017 transformer, built for translation, also had an encoder that read the whole input sentence at once.

Build it yourself

Zero to microGPT builds this transformer piece by piece in your browser, with a short film and a Python lab for each part:

  1. Embeddings and position embeddings
  2. Queries, keys, values, scaled dot-product attention and the KV cache
  3. RMSNorm and residual connections
  4. Multi-head attention, the MLP block and stacking layers. In the last lab you write the body of gpt() yourself, and the test checks it against the trained model's real outputs.

Common questions

Why is it called a transformer? Each layer transforms every token's vector, mixing in information from other tokens (attention) and processing it (MLP), while keeping the vector the same size. The name comes from the 2017 paper "Attention Is All You Need".

What is the difference between attention and a transformer? Attention is one part. A transformer is the whole stack: embeddings, attention, MLPs, normalisation, residual connections and the output layer.

Do bigger transformers work differently? Not in layout. They have more layers, wider vectors, more heads, a longer context and a far bigger vocabulary, and they process many tokens at once on GPUs. The pass through one layer is the one on this page.