Lesson 6 of 8 · 3 min

Choosing the next word: probabilities and temperature

The model does not produce a word, it produces a probability for every possible word. The final choice depends on a setting.

ObjectiveExplain how the next token is chosen and what temperature changes.

At the end of all its layers, the model computes, for each token in the vocabulary, a probability of being the next one. After “The capital of Switzerland is”, “Bern” receives a very high probability, “Zurich” a low but non-zero one.

Then a choice has to be made. Always taking the most probable token gives safe but repetitive texts. Drawing at random according to the probabilities gives more varied texts. Temperature sets this balance. When it is low, the model sticks to the most probable choices. When it is high, it takes more risks.

  • To extract figures, classify or answer a factual question, a low temperature limits deviations.
  • To suggest titles, ideas or variants, a higher temperature widens the range.
  • The same question can give two different answers. This is not a fault, it is sampling.

References

  1. Brown, Mann, Ryder et al. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33, p. 1877–1901. arxiv.org/abs/2005.14165