Skip to main content

micrograd explained: Karpathy's tiny autograd engine

What micrograd is, how its backward pass works, how it differs from the Value class in microgpt, and how to build it yourself in Python.

micrograd is a tiny autograd engine written by Andrej Karpathy. In about 100 lines of Python it does the job at the heart of every neural network library: work out how much each number in a calculation affected the result. A second file of about 50 lines builds a small neural network library on top of it. It is MIT-licensed, it installs with pip install micrograd, and you can read the whole thing in an afternoon.

This page explains what it does, how it works, and how it relates to microgpt, Karpathy's 200-line GPT, which uses a newer version of the same idea.

What "autograd" means

Training a neural network means nudging thousands (or billions) of numbers so that a score called the loss goes down. To know which way to nudge each number, you need its gradient: how much the loss changes when that number changes a tiny bit.

You could measure every gradient by nudging each number and re-running the whole calculation, but with thousands of numbers that is far too slow. Autograd (automatic differentiation) gets every gradient in one backward sweep, using the chain rule.

A small example with numbers. Let a = 2 and b = 3, and compute:

c = a * b      # 6
L = c + a      # 8
  • Changing b by a tiny amount changes L by a times as much, so the gradient of L with respect to b is 2.
  • a reaches L along two paths: through c (worth b = 3) and directly (worth 1). The two paths add up, so the gradient with respect to a is 4.

Autograd gets these answers automatically, for any calculation you write.

How micrograd works

micrograd wraps every number in a Value object. Each Value remembers three things:

  1. its number (data),
  2. the Values it was made from (its children), and
  3. how to pass a gradient back to those children.

Every +, * or ** you write creates a new Value and records it in a graph. When you call .backward() on the final result:

  1. It sorts the graph so that every node comes after the nodes it depends on (a topological sort).
  2. It sets the result's own gradient to 1.
  3. It walks the sorted list backwards, and each node passes gradient to its children using its local rule. For multiplication, the gradient for a in a * b is b, and the other way round.

Gradients are added (+=), never overwritten. That is what makes the two paths from a in the example above add up to 4.

Here is micrograd's multiplication, from its engine.py:

def __mul__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    out = Value(self.data * other.data, (self, other), '*')

    def _backward():
        self.grad += other.data * out.grad
        other.grad += self.data * out.grad
    out._backward = _backward

    return out

micrograd vs the Value class in microgpt

microgpt's Value does the same job, but stores things differently. micrograd's README says microgpt "builds on a more efficient and better version of the autograd engine here (storing local gradients at forward time instead of per-op backward closures)". In code:

# microgpt
def __mul__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    return Value(self.data * other.data, (self, other), (other.data, self.data))

The third argument holds the local gradients: other.data for self and self.data for other. Nothing has to be remembered as a function. One shared backward() then does the whole pass:

for v in reversed(topo):
    for child, local_grad in zip(v._children, v._local_grads):
        child.grad += local_grad * v.grad

That last line is the chain rule: the gradient that reaches a child is its local gradient times the gradient of the node above it.

microgradmicrogpt's Value
Sizeabout 100 lines (engine.py) plus about 50 (nn.py)part of one 200-line file
Backward rule stored asa small function per operationa tuple of local gradients
What it trainsa small neural network that sorts 2-D points into two groups (the "moons" demo)a full GPT that invents new names
Works onsingle numbers (scalars)single numbers (scalars)

Both work on one number at a time, which is why they are slow. Real libraries like PyTorch do the same calculation on large grids of numbers (tensors) on a GPU. The idea is identical; only the speed changes.

Build it yourself

The clearest way to understand micrograd is to build its pieces in order. Module 7 of the free Zero to microGPT course does that in your browser, with no installs:

  1. Computation graphs: a formula as a graph of small steps.
  2. The chain rule: how rates multiply along a path.
  3. When paths branch: why gradients add up (+=).
  4. The Value class: the local gradient for each operation.
  5. backward(): topological sort, then one reverse sweep.

Each lesson has a short film, an interactive picture and a Python lab that checks your code. If you'd like the calculus background first, Module 4 builds slopes from "nudge it and see" with high-school maths.

Karpathy also explains micrograd in a long video, The spelled-out intro to neural networks and backpropagation: building micrograd, the first part of his Neural Networks: Zero to Hero series.

Common questions

Is micrograd used for real models? No. It is a teaching tool: working on one number at a time is far too slow for real networks. Its value is that the whole idea fits on one screen.

Do I need calculus to understand it? You need the idea of a slope and the chain rule, both of which can be learned from examples. The local gradients for +, * and ** are short rules you can check by nudging numbers.

What should I read after micrograd? microgpt. It takes the same autograd idea and puts a complete GPT on top of it: tokens, attention, training and sampling, in about 200 lines. See nanoGPT, nanochat and microgpt compared for how it fits with Karpathy's other projects.