Skip to main content

Module 9 · after

The full Transformer

The same 6 questions you saw before the module. No running the code: predict, then compare with your first score.

1. Predict: two heads of 2 numbers each, reading a cache of two values. What does this print?
def blend(weights, values):
    return [sum(w * v[j] for w, v in zip(weights, values)) for j in range(len(values[0]))]

values = [[1, 2, 3, 4], [5, 6, 7, 8]]
out = []
for h, weights in [(0, [1, 0]), (1, [0, 1])]:
    v_h = [v[h * 2:h * 2 + 2] for v in values]
    out.extend(blend(weights, v_h))
print(out)
2. Predict: a tiny MLP that widens 2 numbers to 4, applies ReLU, then squeezes back to 2. What does it print?
def linear(x, w):
    return [sum(wi * xi for wi, xi in zip(row, x)) for row in w]

fc1 = [[1, 1], [1, -1], [-1, 1], [-1, -1]]
fc2 = [[1, 1, 1, 1], [1, -1, 0, 0]]
x = [3, 1]
hidden = [max(0, v) for v in linear(x, fc1)]
print(linear(hidden, fc2))
3. You change microgpt to n_head = 8, keeping n_embd = 16. What happens inside attention?
4. microgpt has 4,192 knobs with n_layer = 1. Each layer holds 4 attention tables of 16 × 16 plus mlp_fc1 (64 × 16) and mlp_fc2 (16 × 64). How many knobs with n_layer = 2?
5. Which order does gpt() run for one token?
6. Predict: a stand-in block inside a 2-layer loop with residuals. What does it print?
def block(x):
    return [max(0, v - 1) for v in x]

x = [0.5, 3.0]
for li in range(2):
    x = [a + b for a, b in zip(block(x), x)]
print(x)

Your answers are saved with a random id, not your name, so we can see which lessons work. Privacy