Inside a language model
Tokens, vectors, attention and training: what really happens between your question and the answer.
8 lessons · 1 quiz · 45 min · completion badge
- Explain how a text becomes a sequence of tokens, then of vectors.
- Describe the role of attention in a transformer, with an example.
- Distinguish pre-training, instruction tuning and alignment.
- Relate these mechanisms to the errors you observe day to day.
- Read a reference paper in the field without getting lost in the vocabulary.
- Who it is for
- AI champions, trainers, technical profiles and anyone who wants to understand how models work internally.
- Prerequisites
- The course “What AI can do, and its limits”, or some first experience with assistants.
- Duration
- 45 min
- Level
- Advanced
What the course covers.
- 01From text to numbers3 lessons
- Tokens, the model’s basic unit7 min
Why a model reads neither letters nor words but tokens, and what that changes for you.
- Words become vectors7 min
How each token becomes a list of numbers that encodes its meaning, and why geometry replaces the dictionary.
- Meaning as closeness: semantic search3 min
Vectors are also used to find a document by its meaning, without sharing a single word with the question.
- Tokens, the model’s basic unit7 min
- 02The transformer3 lessons
- Attention, or reading the whole text at once3 min
The mechanism that lets each word take all the others into account, at the heart of every large model today.
- Stacked layers and billions of parameters6 min
What a model’s parameters are, and why their number has mattered so much.
- Choosing the next word: probabilities and temperature3 min
The model does not produce a word, it produces a probability for every possible word. The final choice depends on a setting.
- Attention, or reading the whole text at once3 min
- 03Training a model2 lessons
- Pre-training and scaling laws3 min
How a model learns by predicting the next word over trillions of tokens, and what scaling laws have revealed.
- Following instructions: tuning and alignment3 min
Why a base model completes text instead of answering, and how it is taught to be helpful and careful.
- Pre-training and scaling laws3 min
- ✓Final quiz8 questions