Pre-training and scaling laws
How a model learns by predicting the next word over trillions of tokens, and what scaling laws have revealed.
ObjectiveDescribe pre-training and explain what scaling laws say about the size of models and data.
Pre-training is the longest and most expensive phase. The model reads a very large corpus of text and, at each position, tries to predict the next token. When it gets it wrong, its parameters are slightly corrected. Repeated trillions of times, this correction gives rise to grammar, knowledge and ways of reasoning.
- PretrainingPredict the next word over a very large body of text.
- Fine-tuningLearn to follow instructions from written examples.
- AlignmentPrefer the answers people judge useful and safe.
- UseThe model is frozen. It keeps nothing from your conversations.
What scaling laws have shown
In 2020, researchers measured that a model’s error falls steadily and predictably as model size, the amount of data and compute increase. This regularity shaped the whole race for size that followed.
In 2022, the so-called Chinchilla study corrected course. For the same compute budget, many models were too large and undertrained. A smaller model fed with more data did better. The rule of thumb often cited is about twenty training tokens per parameter.
References
- Kaplan, McCandlish, Henighan et al. (2020). Scaling Laws for Neural Language Models. arXiv preprint. arxiv.org/abs/2001.08361
- Hoffmann, Borgeaud, Mensch et al. (2022). Training Compute-Optimal Large Language Models. arXiv preprint. arxiv.org/abs/2203.15556
- Brown, Mann, Ryder et al. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33, p. 1877–1901. arxiv.org/abs/2005.14165