10.1 Adam
Plain gradient descent uses one stride for every knob. But some knobs sit on steep cliffs and others on gentle plains. One stride can’t suit them all. microgpt uses Adam, which gives every knob its own.
1Watch
Picture a long, narrow valley: 10 times narrower than it is long, so its walls curve 100 times more sharply across than along. A stride big enough to make progress along the valley sends you flying across it. Plain descent on x² + 100y² explodes once the learning rate passes 0.01, and below that it creeps.
Adam keeps two running averages for every knob:
- m, the average gradient (momentum): keep rolling in the direction you’ve been going, smoothing out noisy steps.
- v, the average squared gradient: how big this knob’s slopes usually are.
The step is lr × m / √v. A knob with huge slopes gets divided by a huge √v; a knob with tiny slopes gets a boost. Because m and v start at 0, early averages are too small, so bias correction divides by (1 − βt) to fix that. As a result the very first step is lr in size for any ordinary gradient (only gradients near the tiny eps, 1e-8, behave differently).
What a running average is. Each step, m keeps 85% of its old value and mixes in 15% of the new gradient g: m = 0.85 × m + 0.15 × g. That 0.85 is β₁, “how much to keep”. v does the same with g² and keeps 99% (β₂ = 0.99). Old gradients fade a little each step instead of being forgotten at once.
Worked first step, with g = 5. m starts at 0, so m = 0.85 × 0 + 0.15 × 5 = 0.75: far too small, only because m started at 0. Bias correction divides by 1 − 0.85¹ = 0.15, giving m̂ = 5. Likewise v = 0.01 × 25 = 0.25, divided by 1 − 0.99¹ = 0.01, gives v̂ = 25 and √v̂ = 5. The step is lr × 5 / 5 = lr. By step 50, 0.85⁵⁰ is almost 0, so the correction fades away.
microgpt’s settings (line 146): lr = 0.01, β₁ = 0.85, β₂ = 0.99, eps = 1e-8. The lab and widget give Adam lr 0.1 instead: since Adam’s step is about lr in size, 0.1 means “move about 0.1 per step”, which suits a valley 3 units wide. Line 175 also shrinks lr over training; that’s the next lesson.
Then reset. After the update, line 182 sets every p.grad = 0. Gradients add up (+=, lesson 7.3), so without the reset each step would also carry all the old gradients, and the knobs would be pushed by slopes from names long gone.
2Explore
The valley curves 100× more sharply across than along (it is 10× narrower than it is long). Plain descent must use a tiny stride or it explodes across the steep direction (push it past 0.0100). Adam gives each knob its own stride. microgpt itself runs Adam at lr 0.01.
3Build
Two blanks in adam: update the running averages m[i] and v[i], as described above. Everything else is microgpt’s exact Adam.
# A long, narrow valley: f(x, y) = x*x + 100*y*y. Steep across (y), gentle
# along (x) - like many directions of a real network's loss.
def f(p):
x, y = p
return x * x + 100 * y * y
def grad(p):
x, y = p
return [2 * x, 200 * y]
def sgd(steps, lr):
p = [3.0, 1.0]
for _ in range(steps):
g = grad(p)
p = [pi - lr * gi for pi, gi in zip(p, g)]
return p
# Adam, exactly as microgpt lines 177 to 181 (inside the update loop, 174 to 182), for every knob i:
# m: a running average of the gradient (momentum: keep rolling the same way)
# v: a running average of the squared gradient (how big this knob's slopes are)
# divide by sqrt(v) so every knob gets a sensible step size of its own
def adam(steps, lr, beta1=0.85, beta2=0.99, eps=1e-8):
p = [3.0, 1.0]
m = [0.0, 0.0]
v = [0.0, 0.0]
for step in range(steps):
g = grad(p)
for i in range(len(p)):
m[i] = 0.0 # TODO: keep beta1 of the old m[i], mix in (1 - beta1) of this gradient
v[i] = 1.0 # TODO: the same for v[i], with beta2 and the gradient squared
m_hat = m[i] / (1 - beta1 ** (step + 1)) # bias correction:
v_hat = v[i] / (1 - beta2 ** (step + 1)) # m and v start at 0
p[i] -= lr * m_hat / (v_hat ** 0.5 + eps)
return p
print('plain descent, lr 0.009 :', f(sgd(200, 0.009)))
print('plain descent, lr 0.0101:', f(sgd(200, 0.0101)), ' <- one notch more and it explodes')
print('Adam, lr 0.1 :', f(adam(200, 0.1)))
# (microgpt itself uses lr 0.01; here 0.1 means "move about 0.1 per step")
4Check yourself
5Unlocked in microgpt
Lines 146 to 149 set Adam’s settings and create the m and v buffers; lines 174 to 182 are the update you just wrote, run for all 4,192 knobs every step. Line 175 shrinks lr over training (next lesson), line 181 is the downhill step you first met in 6.2, and line 182 resets p.grad to 0 for the next step.
Progress is saved in this browser. Sign in to keep it across devices.