the city skyline as seen from a boat on the water in miami beach, florida
Education Elevation

Grokking

Your neural net memorized everything and learned nothing — then suddenly, it got it.

Grounded in the research on grokking (machine learning)

Wait, it already memorized the data. Why is it still trainin

Because memorization and generalization are not the same thing. A model can hit 100% training accuracy — perfect recall, zero loss — and still be useless on new inputs. Grokking is what happens next: you keep training past that point, for thousands of extra steps, and then generalization suddenly clicks. Not gradually. Suddenly. Test accuracy jumps from chance to near-perfect with almost no warning.

Who caught this? When?

Alethea Power and colleagues at OpenAI published the paper in January 2022. They were training small transformers on modular arithmetic — things like (a + b) mod 113. The models memorized the training set fast, then plateaued. But when they kept training well past memorization, generalization kicked in late. They named the phenomenon grokking, borrowing from Heinlein. The dataset was tiny. The effect was enormous.

More mind games