Neural Networks: Zero to Hero · Lesson 4

Building makemore Part 3: Activations & Gradients, BatchNorm

Video lesson 4

Building makemore Part 3: Activations & Gradients, BatchNorm

Open the black box. Your Module 3 MLP trained - but it started at loss 27 when 'no idea' scores 3.3, and most of its tanh layer arrived dead. Diagnose both, fix the init down to Kaiming's formula, build BatchNorm from scratch, then rebuild the whole network as your own mini torch.nn with the telemetry professionals watch: activation histograms, gradient stats, update-to-data ratios.

InitializationTanh saturationDead neuronsVanishing gradientsKaiming initBatchNorm