Lesson 4 of 5 · 3 min

Mixture of experts

How a model can have a great many parameters while using only some of them for each token.

ObjectiveExplain the principle of a mixture of experts and its consequences for cost and speed.

In a classic transformer, each token passes through all the parameters of every layer. A mixture of experts, or MoE, replaces some blocks with several sub-networks called experts. A small network, the router, chooses for each token the few experts that will process it.

Token “work”RouterExpert 1Expert 2Expert 3Expert 4Expert 5Expert 6Expert 7Expert 82 of 8 experts compute this token
For each token, the router activates only some of the experts.

The idea dates from 2017 and was simplified in 2022 with Switch Transformers, which send each token to a single expert. The open model Mixtral 8x7B, published in 2024, uses eight experts per layer and activates two per token. It has about 47 billion parameters in total, but uses only about 13 billion for each token.

What this changes

  • The computation per token is that of a smaller model, so the answer is faster and cheaper.
  • The memory required stays that of the whole model, because all the experts must be loaded.
  • Experts are not readable specialists such as “law” or “medicine”. Their division of labour is learned and often finer-grained.

References

  1. Shazeer, Mirhoseini, Maziarz et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. ICLR 2017. arxiv.org/abs/1701.06538
  2. Fedus, Zoph, Shazeer (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research 23(120), p. 1–39. jmlr.org/papers/v23/21-0998.html
  3. Jiang, Sablayrolles, Roux et al. (2024). Mixtral of Experts. arXiv preprint. arxiv.org/abs/2401.04088