Course map
Part A builds the thinking tools. Part B climbs Karpathy’s own ladder,train0.py → train5.py, until you’ve written every line of microgpt.
Already know Python? Take the 12-question placement check and skip the foundations.
Part A · Thinking like a computer
M0
The destination
M1
Computational thinking & Python
Module check:See what you already know (3 min)
- 1.1StartThinking like a computerDecomposition, patterns, abstraction, algorithms as recipes.
- 1.2StartValues and variablesNumbers, strings, names for things.
- 1.3StartDecisions and loops`if`, `for`, `while`.
- 1.4StartFunctionsPackaging an idea so you can reuse it.
- 1.5StartReading a datasetLoad 32,033 names from a file.
M2
Data structures
Module check:See what you already know (3 min)
- 2.1StartListsOrdered collections, indexing, slicing.
- 2.2StartDicts and setsLook things up by name; keep only unique items.
- 2.3StartLists of lists = matricesGrids of numbers, rows and columns.
- 2.4StartClasses and objectsBundle data with the operations on it.
- 2.5StartGraphsNodes and edges: the shape of every computation.
M3
Algorithms
Module check:See what you already know (3 min)
- 3.1StartRecursionA function that calls itself.
- 3.2StartDepth-first searchExplore a graph all the way down first.
- 3.3StartTopological sortSocks before shoes: an order that respects every dependency.
- 3.4StartRandomness and weighted diceSeeds, shuffles and `random.choices`.
- 3.5StartHow much work?Counting steps, and why GPUs exist.
M4
Math toolkit (visual first)
Module check:See what you already know (3 min)
- 4.1StartFunctions and graphsInputs, outputs, and their pictures.
- 4.2StartSlope by nudgingThe derivative is “nudge the input, watch the output.”
- 4.3Startexp and logTurning multiplying into adding.
- 4.4StartVectors and the dot productSimilarity as a number.
- 4.5StartMatrix × vectorMany dot products at once.
- 4.6StartProbability distributionsNumbers that add up to 1.
Part B · Building the GPT
M5
Language as probability
train0.pyModule check:See what you already know (3 min)
M6
Learning = walking downhill
train1.pyModule check:See what you already know (3 min)
M7
Autograd = automatic chain rule
train2.pyModule check:See what you already know (3 min)
- 7.1StartComputation graphsEvery calculation is a graph.
- 7.2StartThe chain ruleMultiplying exchange rates.
- 7.3StartWhen paths branchGradients from different paths add up.
- 7.4StartBuilding Valueadd, mul, pow, log, exp, relu, one at a time.
- 7.5Startbackward(): let the graph do calculusReverse topological order plus the chain rule, in 14 lines.
M8
Attention
train3.pyModule check:See what you already know (3 min)
- 8.1StartPosition embeddingsWhere am I in the word?
- 8.2StartQueries, keys, valuesA library search between tokens.
- 8.3StartScaled dot-product attentionWho should I listen to?
- 8.4StartCausal masking & the KV cacheOnly look backwards.
- 8.5StartRMSNormKeep the numbers in a healthy range.
- 8.6StartResidual connectionsA gradient highway.
M9
The full Transformer
train4.pyModule check:See what you already know (3 min)
M10
Training like a pro
train5.pyModule check:See what you already know (3 min)