Following instructions: tuning and alignment
Why a base model completes text instead of answering, and how it is taught to be helpful and careful.
ObjectiveDistinguish instruction tuning, learning from human preferences and principle-based alignment.
A model that has only been pre-trained does not answer a question. It continues it. Given “Write a follow-up email”, it may add more instructions instead of writing the email. It therefore has to be taught the role of an assistant.
- Instruction tuning. People write thousands of pairs of a request and a good answer, and the model is trained on them.
- Preference learning. For the same request, people rank several answers. A second model learns to predict these preferences, then guides the first.
- Principle-based alignment. The model critiques and rewrites its own answers against a list of written principles, which reduces the need for human annotation.
The effect is considerable. In the 2022 InstructGPT study, evaluators preferred the answers of an aligned model with 1.3 billion parameters to those of the 175-billion-parameter base model.
What alignment does not do
- It does not make the model infallible. It makes it more cooperative, which can also make it too compliant.
- It reflects the choices of the people and the principles that guided it.
- It does not replace your guardrails. Learnya’s approvals remain the last barrier before an action.
References
- Ouyang, Wu, Jiang et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35, p. 27730–27744. arxiv.org/abs/2203.02155
- Bai, Kadavath, Kundu et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv preprint. arxiv.org/abs/2212.08073