How do LLMs work? Large language models explained from scratch
How a large language model like ChatGPT works, step by step: tokens, next-token prediction, training by gradient descent and the transformer, shown with the real numbers from a GPT small enough to read in full.
A large language model (LLM), such as the model behind ChatGPT, does one thing over and over: given some text, it predicts what comes next. It gives every possible next piece of text a chance, picks one, adds it to the text, and repeats. Everything else (the training, the maths, the giant data centres) is about making those predictions good.
This page walks through how that works, using microgpt, Andrej Karpathy's complete GPT in about 200 lines of Python. It is tiny: it learns from 32,033 first names and invents new ones, one letter at a time. But its author's note at the top of the file says it is "the complete algorithm", and ChatGPT runs on the same ideas at a vastly bigger scale. Every number below comes from running it.
Step 1: text becomes numbers
A model can only do arithmetic, so text is first cut into tokens and each token gets an id. microgpt uses the simplest scheme: each of the 26 letters is a token, plus one special marker that means "start or end of a name". That is a vocabulary of 27 tokens.
Real models use pieces of words instead: GPT-2's vocabulary has 50,257 tokens, built by byte-pair encoding. The tokens guide shows how.
Step 2: predict the next token as a list of chances
Given the tokens so far, the model outputs one score for every token in its vocabulary. A function called softmax turns those scores into chances that add up to 100%.
For example, ask the trained microgpt how a name starts (it has only seen the start marker so far). It gives each of the 27 tokens a chance, and the most likely first letter is "a", at about 14%.
Then it rolls a weighted die: "a" comes up about 14% of the time, other letters less often. That is why the same model gives different answers each time. A setting called temperature sharpens or flattens the die before the roll: low temperature means safer choices, high means wilder ones. microgpt uses 0.5.
Repeat: feed the chosen letter back in, get new chances, roll again, until the model picks the end marker. Out come names it never saw, such as "alilan" and "akalen".
Step 3: the simplest language model is a counting table
You can build a working (if weak) language model without any neural network. Count every pair of neighbouring letters in all 32,033 names (228,146 pairs), and use each row of counts as the weighted die for "what follows this letter".
This bigram model works, but it only remembers one letter. After "emm" it sees just the last "m". An LLM's job is to use all the earlier tokens.
Step 4: score the model with "surprise" (the loss)
To compare models, measure how surprised each one is by real text. If the model gave the true next letter a chance of p, its surprise is −ln(p): zero when it was certain and right, large when it thought the letter unlikely. The average surprise over the data is the loss. Lower is better.
| model | loss |
|---|---|
| know-nothing (every token 1 in 27) | ln 27 ≈ 3.30 |
| microgpt before training (random knobs) | 3.37 at step 1 |
| counting table (bigram) | 2.45 |
| microgpt after 1,000 training steps | 2.28 (average of the last 100 steps) |
The trained GPT beats the counting table because it can look at every earlier letter, not just the last one.
Step 5: a neural network full of knobs
Instead of a table of counts, an LLM computes its scores with a neural network: lots of multiplying and adding, controlled by numbers called parameters. Think of them as knobs. microgpt has 4,192 of them, arranged in nine tables. Before training they are small random numbers, so the model's guesses are no better than know-nothing.
The arrangement of those knobs is the transformer:
- Each token's id picks out a learned list of 16 numbers (its embedding), plus a second list for its position in the text.
- Attention lets each token look back at the earlier ones and pull in what matters.
- A small MLP (a two-layer neural network) works on each token on its own.
- A final table turns the result into one score per possible next token.
The transformers guide walks through each part, and attention explained works attention through by hand.
Step 6: training is walking downhill
Training sets the knobs. microgpt repeats the following steps 1,000 times:
- Take one name, run it through the model, and measure the loss.
- Work out, for every one of the 4,192 knobs, which way to turn it to lower the loss. That number is the knob's gradient, and computing all of them at once is backpropagation, which uses the chain rule from calculus. microgpt does it with a small autograd engine, the same idea as Karpathy's micrograd.
- Nudge every knob a little in the downhill direction. microgpt uses the Adam optimizer, which adapts the step size for each knob.
The loss starts at 3.37, first averages below 2.5 around step 254, and ends near 2.28. No one writes any rules about names into the model. Whatever patterns it picks up, it picks up because they lower the loss.
From microgpt to ChatGPT
Karpathy says the differences between microgpt and production models don't "alter the core algorithm". What changes is:
- Tokens: word pieces instead of letters, so the same text takes far fewer steps.
- Speed: the same maths done on huge grids of numbers (tensors) on GPUs, instead of one number at a time.
- Scale: more layers, wider vectors, a longer context, and vastly more text.
| parameters | context (tokens it can see) | |
|---|---|---|
| microgpt | 4,192 | 16 |
| GPT-2 (2019) | 1.5 billion | 1,024 |
| GPT-3 (2020) | 175 billion | 2,048 |
- More training stages: next-token training on a huge pile of text is called pretraining. Chat assistants are then trained further on example conversations (supervised fine-tuning) and with feedback on their answers (reinforcement learning), so they follow instructions instead of just continuing the text.
This also explains hallucination. The model rolls a die over plausible next tokens. It doesn't look facts up. microgpt invents believable names that don't exist, and a big model can likewise write a fluent sentence that isn't true.
Build it yourself
Zero to microGPT is a free course that takes you from your first line of Python to writing this whole model from a blank file, in your browser:
- Counting letter pairs, temperature and loss as surprise: language as probability.
- Gradient descent and embeddings: learning by walking downhill.
- Backpropagation: the autograd engine.
- Attention and the full transformer.
- Adam and reading loss curves.
- microgpt from a blank file, then from microgpt to ChatGPT.
New to Python? Start at lesson 0.1, or take the placement check to skip what you know.
Common questions
Does an LLM understand what it says? It is trained only to predict the next token well. Doing that across huge amounts of text forces it to pick up a lot of structure: grammar, facts, styles of reasoning. Whether that counts as understanding is debated, but the mechanism itself is next-token prediction.
Is an LLM the same as a GPT? GPT ("generative pre-trained transformer") is one family of LLMs. Most of today's chat models use the same basic design: a transformer trained to predict the next token.
Do I need advanced maths to learn this? No. You need multiplication, addition, exp and log, the idea of a slope, and weighted averages. The course teaches each of them from scratch before it uses them.