Lesson 4 of 8 · 3 min

Attention, or reading the whole text at once

The mechanism that lets each word take all the others into account, at the heart of every large model today.

ObjectiveExplain with an example what attention computes and why it replaced earlier architectures.

In 2017, a Google team published “Attention Is All You Need” and introduced the transformer. Its key idea is attention. For each token, the model computes how useful each of the other tokens in the text is for interpreting it, then blends their information accordingly.

Take “The cat sleeps because it is tired”. To understand “it”, the model has to know that it refers to the cat. Attention gives a high weight to “cat” when it processes “it”. None of this is programmed by hand. The weights are learned.

What does “it” pay attention to?

The4 %cat62 %sleeps9 %because3 %it12 %is4 %tired6 %
Illustrative attention weights for the word “it”. A real model computes such weights in dozens of heads and layers.

Why it is decisive

  • All tokens are processed in parallel, which makes training on huge corpora possible with graphics processors (GPUs).
  • A word can link directly to another one placed very far away in the text.
  • Several attention heads work side by side. One can track grammar, another references, another the topic.

References

  1. Vaswani, Shazeer, Parmar et al. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30. arxiv.org/abs/1706.03762