# RAG (retrieval-augmented generation)

> Looking up relevant documents first and putting them in the prompt, so the model answers from your material instead of from memory.

A model only knows what was in its training data, which stops at a fixed date and never included your contracts, your product documentation or last quarter's numbers. RAG closes that gap without retraining anything: when a question arrives, a search step finds the handful of passages most likely to contain the answer, those passages go into the prompt alongside the question, and the model answers from them.

The retrieval step is where these systems succeed or fail, and it is usually the part that gets the least attention. If the search returns the wrong three paragraphs, a better model will not save the answer — it will produce a fluent, confident response built on the wrong source. Most disappointing RAG deployments are retrieval problems wearing a generation costume.

It is also the honest way to get citations. Because the system knows which documents it put in the prompt, it can show them, and a reader can check. A model answering from training data alone cannot do this, which is why answers that cite nothing should be treated as claims rather than facts.

## Related terms

- https://dfieldsolutions.com/en/glossary/embedding.md
- https://dfieldsolutions.com/en/glossary/vector-database.md
- https://dfieldsolutions.com/en/glossary/hallucination.md
- https://dfieldsolutions.com/en/glossary/context-window.md

---

Source: https://dfieldsolutions.com/en/glossary/rag
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
