A model only knows what was in its training data, which stops at a fixed date and never included your contracts, your product documentation or last quarter's numbers. RAG closes that gap without retraining anything: when a question arrives, a search step finds the handful of passages most likely to contain the answer, those passages go into the prompt alongside the question, and the model answers from them.
The retrieval step is where these systems succeed or fail, and it is usually the part that gets the least attention. If the search returns the wrong three paragraphs, a better model will not save the answer — it will produce a fluent, confident response built on the wrong source. Most disappointing RAG deployments are retrieval problems wearing a generation costume.
It is also the honest way to get citations. Because the system knows which documents it put in the prompt, it can show them, and a reader can check. A model answering from training data alone cannot do this, which is why answers that cite nothing should be treated as claims rather than facts.
Related terms
Embedding
A list of numbers representing a piece of text, arranged so that texts about similar things end up close together.
Vector database
A database built to find the stored items whose embeddings are closest to a query's, quickly, across millions of rows.
Hallucination
A confident, fluent, entirely invented answer — a citation, a figure or an API that does not exist.
Context window
The maximum amount of text a model can take into account at once, counting the instructions, the conversation and the answer together.
The bench this belongs to
AI automationThe repetitive half of your week, handed to software that does not get bored. Inbox triage, follow-ups, reporting, data entry between tools that were never meant to talk.
