7.4 Building Value
Time to build microgpt’s engine room. Each operation, whether add, multiply, power, log, exp or relu, makes a new Value and writes down one tiny fact: its own local rate. Five lines, five rates, and you’ve written most of an autograd system.
1Watch
A Value stores data (its number), _children (its inputs) and _local_grads: one local rate per child, worked out at the moment it is created, while the numbers are known.
a + b: rates (1, 1). Given as an example.a × b: rates (b, a). Nudge a, and the product moves by b.x ** n: rate n × xn−1.log(x): rate 1/x.exp(x): rate exp(x).relu(x)(keep positives, zero out negatives): rate 1 when x > 0, else 0.
You don’t have to take these on trust. Each one is just the slope you’d measure by nudging, as in lesson 4.2. Pick an operation in the widget below, move x, and check that “nudging says” agrees with the rule.
Subtraction, division and negation reuse these (lines 51 to 57: a − b is a + (−1)×b, a / b is a × b−1), so they need no rates of their own.
2Explore
3Build
This is microgpt’s real Value class with five local rates blanked out (__mul__, __pow__, log, exp, relu). Fill them in. The last check compares every rate against nudging.
The formulas say x and n, but the code doesn’t have those names. Inside these methods the input’s number is self.data, and in __pow__ the exponent n is called other (a plain number). Local grads must be plain numbers, not Values, so always use .data: other.data, never other in __mul__. (A Value there would build a new node inside the node, and in __pow__ that never stops.)
A tuple with one item needs a comma: (1/self.data,), just like the (0,) already there. Without it the brackets are just brackets and you get a single number.
The rest of the class is given. __slots__ only saves memory, and isinstance wraps a plain number into a Value. The r versions (__radd__, __rmul__, …) handle a plain number on the left, like 2 * x. You don’t need to touch them.
import math
class Value:
__slots__ = ('data', 'grad', '_children', '_local_grads') # Python optimization for memory usage
def __init__(self, data, children=(), local_grads=()):
self.data = data # scalar value of this node calculated during forward pass
self.grad = 0 # derivative of the loss w.r.t. this node, calculated in backward pass
self._children = children # children of this node in the computation graph
self._local_grads = local_grads # local derivative of this node w.r.t. its children
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
return Value(self.data + other.data, (self, other), (1, 1))
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
return Value(self.data * other.data, (self, other), (0, 0)) # TODO: d(a*b)/da = b, d(a*b)/db = a (a is self.data, b is other.data)
def __pow__(self, other): return Value(self.data**other, (self,), (0,)) # TODO: d(x**n)/dx = n * x**(n-1) (here n is other, a plain number; x is self.data)
def log(self): return Value(math.log(self.data), (self,), (0,)) # TODO: d(log x)/dx = 1/x (x is self.data; keep the comma: (...,))
def exp(self): return Value(math.exp(self.data), (self,), (0,)) # TODO: d(exp x)/dx = exp(x) (x is self.data)
def relu(self): return Value(max(0, self.data), (self,), (0.0,)) # TODO: 1.0 if x > 0, else 0.0 (x is self.data)
def __neg__(self): return self * -1
def __radd__(self, other): return self + other
def __sub__(self, other): return self + (-other)
def __rsub__(self, other): return other + (-self)
def __rmul__(self, other): return self * other
def __truediv__(self, other): return self * other**-1
def __rtruediv__(self, other): return other * self**-1
# Check a few local gradients against nudging the input a tiny bit.
x = Value(3.0)
y = x * Value(4.0)
print('3 * 4 =', y.data, ' local grads:', y._local_grads)
print('log(2): local grad', Value(2.0).log()._local_grads)
print('relu(-2):', Value(-2.0).relu().data, Value(-2.0).relu()._local_grads)
4Check yourself
5Unlocked in microgpt
Lines 39 to 57 are now yours (log and exp on lines 48 and 49 first appeared in lesson 4.3): every operation microgpt’s model uses, each recording its local rates. Line 29’s comment says it all: “Let there be Autograd.” One method is still missing, the one that walks the graph backwards and multiplies all these rates. That’s backward(), the next lesson.
Progress is saved in this browser. Sign in to keep it across devices.