Lesson 3 of 5 · 3 min

Search before answering: retrieval-augmented generation

The principle that lets a model answer from your documents and cite its sources.

ObjectiveDescribe the stages of retrieval-augmented generation and identify where it can fail.

A model knows neither your contracts nor your procedures. Retrieval-augmented generation, or RAG, brings it this knowledge at the moment of answering. The principle was formalised in 2020: first the relevant passages are retrieved, then they are given to the model along with the question.

  1. Question“What does our contract say about termination?”
  2. RetrievalThe closest passages are found in your documents.
  3. ContextThese passages are placed before the question.
  4. AnswerThe model answers and cites the passages it used.
The four stages of retrieval-augmented generation.

This setup reduces fabrication, because the answer relies on a text present in the context, and it makes it possible to cite the source. This is how Learnya’s answers about your documents and the web work.

Where it can fail

  • The retrieval brings back the wrong passage. The answer will then be faithful to an off-topic text.
  • The right passage is found but badly chunked, and the decisive sentence is missing.
  • The model mixes the passage with its own knowledge. A 2023 survey of hallucinations ranks this case among the most common.

References

  1. Lewis, Perez, Piktus et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems 33, p. 9459–9474. arxiv.org/abs/2005.11401
  2. Ji, Lee, Frieske et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12), p. 1–38. doi.org/10.1145/3571730