Video lesson 5
Building makemore Part 4: Becoming a Backprop Ninja
Unplug loss.backward(). Backpropagate through the entire MLP-with-BatchNorm by hand - cross-entropy, linear layers, tanh, BatchNorm's full anatomy, the embedding lookup - twenty-six gradients checked against autograd one by one. Then collapse cross-entropy and BatchNorm each to a single derived expression, and train the network on your gradients alone.
The chain rule at tensor scaleMatmul gradientsBroadcasting in reverseGradient accumulationSoftmax/cross-entropy backwardBatchNorm backward