What is RAG?

RAG stands for retrieval-augmented generation. It is a way of getting an AI language model to answer from your own documents instead of only from what it learned in training.

The problem it solves

A language model does not know your company handbook, your product manual or last week's report. Asked about them, it either says it does not know or makes something up. Pasting the whole document into every question is slow, costly and often too long.

How it works

  1. Split. The documents are cut into short passages, often called chunks.
  2. Turn into vectors. Each passage is converted to a list of numbers that captures what it is about.
  3. Retrieve. When a question arrives, it is converted the same way, and the passages whose numbers are closest are picked.
  4. Generate. Those passages are placed in the prompt with the question, and the model answers from them.

What decides the quality

Most poor RAG answers come from retrieval, not from the model. If the right passage is not among those retrieved, the model cannot answer well. Passage size, how the document is split and how many passages are retrieved all change what is found.

When to use it

Try RAG on your own document to see each step, or read the chunk size guide.