Lesson 1 of 8 · 7 min

Tokens, the model’s basic unit

Why a model reads neither letters nor words but tokens, and what that changes for you.

ObjectiveExplain what a token is, how a text is split, and why the length of a text is counted in tokens.

A language model handles nothing but numbers. Before any computation, the text is therefore split into pieces called tokens, and each token receives a number in a fixed vocabulary. This splitting is called tokenisation.

Tokens are neither letters nor whole words. A common word often fits in a single token. A rare word, a proper name or a technical term is cut into several frequent pieces. The most widespread method, called byte-pair encoding, starts from characters and merges, step by step, the pairs that occur most often in a large corpus.

Sp43aced9811␣rev306ision22109␣work1891s7356␣rel12340iably710.13
An illustrative split. Each piece receives a number, and the model sees only these numbers.

What this explains

  • The size of a document, the limit of a conversation and often the cost are counted in tokens, not in words.
  • A text in French or German generally needs more tokens than the same text in English, because the vocabularies were built on mostly English corpora.
  • Counting the letters of a word or writing it backwards is hard for a model. It does not see the letters, it sees pieces.
Less effective

How many “r”s are there in “strawberry”?

More effective

Spell “strawberry” letter by letter, then count the “r”s.

By asking for the spelling first, you make the model produce the letters one by one, which it can do, before it counts.

References

  1. Sennrich, Haddow, Birch (2016). Neural Machine Translation of Rare Words with Subword Units. Proceedings of the 54th Annual Meeting of the ACL, p. 1715–1725. doi.org/10.18653/v1/P16-1162