11.1 microgpt from a blank file
This is it: an empty file and nine section headings. You have built every piece that goes under them, lesson by lesson; a few fixed choices that tie them together are listed below. Write it once more, start to finish, and your microgpt will match Karpathy’s training run number for number.
1Watch
Allow 2 to 4 hours, over several sittings. Every earlier lab had one or two blanks; this one is about 150 lines. Your code saves automatically, so stop whenever you like.
The starter file keeps only Karpathy’s header and his section comments. Fill each section in order. If you get stuck, the lesson that taught it is listed below. Re-open it rather than copying: the full program is in the lesson 0.1 and 11.2 labs and in the sidebar, and you could copy it. Nothing stops you. But the point of this lesson is to find that you can write it yourself.
- Tables, in this orderwte (27×16), wpe (16×16), lm_head (27×16), then for each layer attn_wq, attn_wk, attn_wv, attn_wo (16×16 each), mlp_fc1 (64×16), mlp_fc2 (16×64). Every knob is random.gauss(0, 0.08), drawn row by row; params lists them in the same order.
- Which name each step usesdoc = docs[step % len(docs)], not random.choice: the names were already shuffled once.
- TokensBOS on both ends: [BOS] + letters + [BOS]. Then n = min(block_size, len(tokens) - 1) positions.
- softmax on ValuesLesson 6.4’s softmax used plain numbers. Here the scores are Values: subtract the max using .data (max(v.data for v in logits)), then call .exp() on each Value.
- Residuals on listsx = [a + b for a, b in zip(x, x_residual)]; x + x_residual would glue the lists end to end.
- Adamlearning_rate 0.01 (the 10.1 lab used 0.1), beta1 0.85, beta2 0.99, eps 1e-8, bias correction with step + 1.
- Learning-rate decaylr_t = learning_rate * (1 - step / num_steps). It depends on num_steps, so a 20-step run is not the first 20 steps of a 1,000-step run.
- After the updatep.grad = 0 for every knob. Gradients add up (+=), so without the reset each step would also carry the old ones.
- Python you may wanta if cond else b picks a or b in one expression (Value wraps plain numbers with it). [p for mat in tables for row in mat for p in row] reads its fors left to right, like nested loops: it flattens the tables into one list. break leaves a loop at once (stop sampling at BOS). .index and enumerate are from 8.4.
- Two lines for the checksparams_at_start = [p.data for p in params] right after building params, and step_losses.append(loss.data) at the end of every step (the starter already has step_losses = []).
Other harmless differences still match: sum(losses) / n or (1 / n) * sum(losses), head_dim ** 0.5 or math.sqrt(head_dim).
The check is strict on purpose. microgpt is deterministic: random.seed(42) fixes the shuffle, the starting knobs and the samples. So a correct microgpt with num_steps = 20 lands on Karpathy’s exact loss after step 20 (2.7749), to 9 decimal places. One wrong sign anywhere and the numbers drift.
The checks go in four stages, in the order of the file: the shuffled names, the starting knobs, the step-1 loss, then all 20 losses. Step 1 (3.3660) runs on the starting knobs, before any update. So if step 1 is wrong, look at the setup and the model; if step 1 matches and a later step differs, look at the training loop and Adam. Press Check after each section to see how far you have got.
When it passes, set num_steps = 1000 and press Run for the real thing. It takes a few minutes in the browser. Then read the names it invents: every line that made them is yours.
2Explore
- Datasetlessons 1.5, 3.4
- Tokenizerlessons 2.2, 5.1
- Autograd: Value and backward()lessons 3.3, 7.4, 7.5
- Parameterslessons 2.3, then 9.1 to 9.3 for the layer tables
- Model: linear, softmax, rmsnormlessons 4.5, 6.4, 8.5
- gpt(): attention, MLP, layerslessons 8.2 to 9.3
- Adamlessons 10.1
- Training looplessons 7.5 (backward), 10.1 to 10.3
- Inferencelessons 5.2, 6.4 (logits ÷ temperature)
3Build
Write microgpt under the section comments, following the conventions card. Keep num_steps = 20 while checking. Reading names.txt replaces Karpathy’s download (lines 15 to 18), since the file is already in your browser.
Milestones as you go: num docs: 32033, vocab size: 27, num params: 4192, then step 1 / 20 | loss 3.3660 and step 20 / 20 | loss 2.7749.
"""
YOUR microgpt, from a blank file. Every section below is one you have built
in an earlier lesson. Write them in order, following the conventions card on
the page. With num_steps = 20 the checks compare your run with Karpathy's own
20-step run in four stages: the shuffled names, the starting knobs, the step-1
loss, then every step's loss. Check after each section to see how far you got.
The most atomic way to train and run inference for a GPT in pure, dependency-free Python.
This file is the complete algorithm.
Everything else is just efficiency.
@karpathy
"""
import os # (Karpathy uses os.path.exists to download names; not needed here)
import math # math.log, math.exp
import random # random.seed, random.choices, random.gauss, random.shuffle
random.seed(42) # Let there be order among chaos
# Let there be a Dataset `docs`: list[str] of documents (e.g. a list of names)
# TODO
# Let there be a Tokenizer to translate strings to sequences of integers ("tokens") and back
# TODO
# Let there be Autograd to recursively apply the chain rule through a computation graph
# TODO
# Initialize the parameters, to store the knowledge of the model
# TODO
# For the checks, end this section with: params_at_start = [p.data for p in params]
# Define the model architecture: a function mapping tokens and parameters to logits over what comes next
# TODO
# Follow GPT-2, blessed among the GPTs, with minor differences: layernorm -> rmsnorm, no biases, GeLU -> ReLU
# TODO
# Let there be Adam, the blessed optimizer and its buffers
# TODO
# Repeat in sequence
num_steps = 20 # 20 while checking; set 1000 for the real run
step_losses = [] # for the checks: end every step with step_losses.append(loss.data)
# TODO
# Inference: may the model babble back to us
# TODO
4Check yourself
5Unlocked in microgpt
The header (lines 1 to 8) and the rest of the inference section (lines 186, 188 to 194 and 197 to 200): start at BOS, run gpt(), divide by the temperature, roll, stop at BOS. With these, every line on the map is lit.
Progress is saved in this browser. Sign in to keep it across devices.
The code is hidden so you can write it yourself. Try first; peek only when you’re stuck.