Lesson 5 of 8 · 6 min

Stacked layers and billions of parameters

What a model’s parameters are, and why their number has mattered so much.

ObjectiveDescribe how a transformer’s layers are stacked and what a parameter represents.

A transformer stacks identical blocks. Each block contains an attention step, which moves information between tokens, followed by a small network that transforms each token separately. A token’s vector is thus reworked layer after layer, carrying more and more context.

Parameters are the numbers tuned during training. They define the embeddings, the attention computations and the small networks. The GPT-3 model, published in 2020, had 175 billion parameters spread across 96 layers. Its authors showed that at this scale, a model could perform a task from a few examples given in the request, without retraining.

What parameters are not

  • They are not knowledge cards. No parameter contains “the capital of Switzerland”.
  • Knowledge is spread across millions of parameters at once, which makes it robust but hard to correct point by point.
  • More parameters do not guarantee a better model. The quantity and quality of the data matter just as much.

References

  1. Vaswani, Shazeer, Parmar et al. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30. arxiv.org/abs/1706.03762
  2. Brown, Mann, Ryder et al. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33, p. 1877–1901. arxiv.org/abs/2005.14165