Writing the post that I wished I'd found when I started learning whatever it was...

GPT-2 mysteries

After my LLM from scratch series, there were a few mysteries to resolve -- in particular, why my own models were worse than OpenAI's on instruction-following. These occasional posts are my attempts to solve them.