nanoGPT vs nanochat vs microgpt: Karpathy's GPT projects compared
What nanoGPT, nanochat and microgpt each are, how big they are, what they need to run, and which one to start with if you want to understand how GPTs work.
Andrej Karpathy has published several small GPT projects. They share one algorithm but are built for different jobs. In short:
- microgpt is for understanding: the whole algorithm in one file of plain Python.
- nanoGPT was for training real GPTs with PyTorch. Karpathy now marks it as deprecated.
- nanochat is its successor: a complete pipeline, from raw text to a model you can chat with.
The quick comparison
| microgpt | nanoGPT | nanochat | |
|---|---|---|---|
| What it is | a GPT in one file, from scratch | a training repo for medium-sized GPTs | a full harness for training a chat model |
| Code | about 200 lines, pure Python, no libraries | train.py and model.py, about 300 lines each, built on PyTorch | a larger repo built on PyTorch |
| Trains on | 32,033 first names, one letter at a time | e.g. OpenWebText, or Shakespeare for a quick start | web text, then conversations |
| Model size | 4,192 parameters | GPT-2 small (124M) and up | GPT-2 level and up, set by one --depth setting |
| Needs | any computer; runs in minutes | a GPU; the GPT-2 (124M) reproduction took about 4 days on 8 A100 GPUs | a GPU node; GPT-2 level in about 2 hours on 8 H100 GPUs (about $48) |
| Status | published 2026 | deprecated since November 2025 | the recommended one for training |
The nanoGPT and nanochat figures come from their READMEs, as of October 2026.
microgpt: the whole algorithm, nothing else
microgpt trains a tiny GPT to invent new names. Its single file contains:
- a dataset loader and a tokenizer that treats each letter as a token;
- an autograd engine (a
Valueclass, like micrograd); - the transformer itself: embeddings, attention, an MLP and RMSNorm;
- the Adam optimizer, a training loop and a sampler.
Karpathy's own summary at the top of the file: "This file is the complete algorithm. Everything else is just efficiency."
That is why it is the best starting point if you want to understand a GPT. Every number is a plain Python float you can print, and there is no library hiding the maths. The cost is speed: working one number at a time is far too slow for anything bigger than a toy.
nanoGPT: the same model, at real scale
nanoGPT is "the simplest, fastest repository for training/finetuning medium-sized GPTs". It is a rewrite of Karpathy's earlier minGPT. Its train.py reproduces GPT-2 (124M parameters) on the OpenWebText dataset, and its model.py can load OpenAI's GPT-2 weights. A quick-start config trains a small model on Shakespeare, one character at a time.
Compared with microgpt, what changes is efficiency, not the idea:
- Tensors instead of single numbers. PyTorch multiplies whole grids of numbers at once on a GPU, and does the autograd for you.
- Word pieces instead of letters. On web text it uses GPT-2's byte-pair encoding (BPE) tokenizer, with a vocabulary of 50,257 tokens.
- Bigger everything: a longer context, more layers, more heads, many more parameters.
Since November 2025 the nanoGPT README carries a note: it is "very old and deprecated", and you probably want nanochat instead.
nanochat: from raw text to a chat model
nanochat describes itself as "the simplest experimental harness for training LLMs". Where nanoGPT stops at pretraining, nanochat covers every major stage: tokenization, pretraining, finetuning, evaluation and inference, ending with a simple chat interface. Its README says it can train a model with GPT-2-level ability for about $48 (around two hours on a node of eight H100 GPUs). Most settings are derived from a single --depth option, the number of transformer layers.
Which one should you start with?
If you want to understand how a GPT works, start with microgpt. It is the only one of the three that runs anywhere and hides nothing. When you know what each of its 200 lines does, nanoGPT's model.py reads as the same model written with tensors, and nanochat's extra stages make sense.
If you want to train a model that writes real text, use nanochat. You'll need access to GPUs. Knowing microgpt first makes its code far easier to follow.
Learning microgpt from zero
Zero to microGPT is a free course that builds microgpt line by line. It is written for high-school students and needs no prior coding. Every lesson has a short film, an interactive picture and a Python lab that runs in your browser. The course:
- starts with microgpt running in your browser;
- teaches the Python and maths it needs (lists, dictionaries, recursion, slopes, logs, dot products);
- builds the model in the order Karpathy's own ladder does: counting, gradient descent, autograd, attention, the full transformer and Adam;
- ends with you writing microgpt from a blank file and matching Karpathy's numbers exactly;
- finishes with what changes on the way to ChatGPT, the same scale-up nanoGPT and nanochat make.
Common questions
Is nanoGPT still worth reading? Yes, as a short, clean PyTorch GPT. For training new models, its README points you to nanochat.
Can microgpt write sentences? No. It is trained on names, one letter at a time, with 4,192 parameters. It invents plausible new names, which is enough to show every part of the algorithm working.
Do I need a GPU for microgpt? No. It is pure Python with no libraries, so it runs on any computer, and even in a web browser.
What is minGPT? Karpathy's earlier, education-focused PyTorch GPT. nanoGPT is a rewrite of it that prioritises speed.