Neural Networks: Zero to Hero · Lesson 9

Let's reproduce GPT-2 (124M)

Video lesson 9

Let's reproduce GPT-2 (124M)

The bonus module. Rewrite your Module 7 GPT the way OpenAI shipped it - fused attention, tanh GELU, weight tying, their exact init - then load OpenAI's actual 124M checkpoint into your own class and hear the real GPT-2 speak through your code. Build the GPT-3-grade training loop and the HellaSwag eval around it, and finish by fine-tuning the real weights on Shakespeare.

Fused multi-head attentionWeight tyingGELU and frozen approximationsResidual-stream init scalingState-dict surgeryCheckpoint fingerprinting