Search before answering: retrieval-augmented generation
The principle that lets a model answer from your documents and cite its sources.
ObjectiveDescribe the stages of retrieval-augmented generation and identify where it can fail.
A model knows neither your contracts nor your procedures. Retrieval-augmented generation, or RAG, brings it this knowledge at the moment of answering. The principle was formalised in 2020: first the relevant passages are retrieved, then they are given to the model along with the question.
- Question“What does our contract say about termination?”
- RetrievalThe closest passages are found in your documents.
- ContextThese passages are placed before the question.
- AnswerThe model answers and cites the passages it used.
This setup reduces fabrication, because the answer relies on a text present in the context, and it makes it possible to cite the source. This is how Learnya’s answers about your documents and the web work.
Where it can fail
- The retrieval brings back the wrong passage. The answer will then be faithful to an off-topic text.
- The right passage is found but badly chunked, and the decisive sentence is missing.
- The model mixes the passage with its own knowledge. A 2023 survey of hallucinations ranks this case among the most common.
References
- Lewis, Perez, Piktus et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems 33, p. 9459–9474. arxiv.org/abs/2005.11401
- Ji, Lee, Frieske et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12), p. 1–38. doi.org/10.1145/3571730