Skip to main content

nanochat explained: Karpathy's pipeline from raw text to a chat model

What Andrej Karpathy's nanochat is, what it trains, what it costs and needs, how it relates to nanoGPT and microgpt, and how to understand it from scratch first.

Updated 10 Oct 2026Facts from the project’s own README

nanochat is Andrej Karpathy's small, readable project for training a chat model end to end. In its own words, it is "the simplest experimental harness for training LLMs". One script takes you from raw text to a model you can talk to:

  • tokenization
  • pretraining
  • finetuning
  • evaluation
  • inference, with a simple chat interface

It replaced Karpathy's older nanoGPT, whose README now points people to nanochat. Everything below comes from the nanochat README, as of October 2026.

What it can train, and what it costs

nanochat's headline result is a model with GPT-2-level ability. Training GPT-2 cost about $43,000 in 2019. With nanochat, the README says you can do it for about $48: roughly two hours on one machine with eight H100 GPUs, at about $24 an hour. On a cheaper spot instance it can be closer to $15.

The project keeps a public leaderboard for this "GPT-2 speedrun": how fast a model reaches GPT-2's score on a benchmark called CORE. The reference run is the script runs/speedrun.sh, and the best times have kept falling. Check the README for the current record.

Once training finishes, you chat with your model from the command line:

bash runs/speedrun.sh            # train (on an 8-GPU machine)
python -m scripts.chat_cli       # then talk to it

Karpathy describes talking to the result as "a bit like talking to a kindergartener". It writes fluent sentences that are often confidently wrong, which is a good way to see what pretraining alone does and doesn't give you.

One dial: depth

Most training setups have dozens of settings. nanochat sets almost all of them from a single number, --depth: the number of layers in the transformer. Width, number of attention heads, learning rates, training length and weight decay are all calculated from it automatically. GPT-2-level ability arrives at around depth 24 to 26. Smaller depths give a whole series of smaller models.

What you need to run it

  • A GPU machine for the real thing. The speedrun is designed for an 8×H100 node. It also runs on 8×A100 (a bit slower), or on a single GPU, which gives about the same result but takes about 8 times longer. GPUs with less than 80 GB of memory need a smaller batch size.
  • A laptop for a toy version. runs/runcpu.sh shrinks everything to train on a CPU or Apple Silicon in a few tens of minutes. The README is clear that you won't get strong results this way.
  • Python and PyTorch, installed with the uv tool.

nanochat vs nanoGPT vs microgpt

microgptnanoGPTnanochat
Goalunderstand the algorithmtrain a GPTtrain a chat model end to end
Codeone 200-line file, no libraries~600 lines of PyTorcha larger PyTorch project
Runs onany computer, even a browsera GPUa GPU node (or a toy CPU run)
Statuspublished 2026deprecatedKarpathy's current project

The three share one algorithm. microgpt has all of it in plain Python, one number at a time. nanoGPT and nanochat add what makes it fast and big: tensors, GPUs, a real word-piece tokenizer, more layers. nanochat also adds the steps after pretraining that turn a text predictor into something you can chat with. For more on the differences, see nanoGPT vs nanochat vs microgpt.

How to understand nanochat from scratch

nanochat's code is short for what it does, but it assumes you already know what attention, an optimizer and a training loop are. The quickest way to get there is to build the same model in miniature first:

  1. Learn what a token is and how a model predicts the next one (What is a token?).
  2. Build autograd, the engine that computes every gradient (micrograd explained).
  3. Build attention (attention explained), then the full transformer and Adam.
  4. Write the whole model from a blank file.

Zero to microGPT does exactly that, in your browser, starting from your first line of Python. It ends with what changes on the way to ChatGPT, the same scale-up that nanochat makes.

Common questions

Can nanochat make something like ChatGPT? It trains a small chat model with GPT-2-level ability, far smaller than today's assistants. It is a learning and research harness, not a product.

Do I need to pay for GPUs to learn from it? To train the real speedrun, yes: you rent a GPU machine for a few hours. To understand the ideas, no: microgpt runs the same algorithm on any computer.

Is nanoGPT still worth reading? As a short, clean example of a GPT in PyTorch, yes. For training new models, its README points you to nanochat.