What Is RAG? Retrieval-Augmented Generation for Beginners
RAG lets an AI answer questions about your own documents by finding the relevant pieces first and handing them to the model. How retrieval-augmented generation works, step by step, when you need it, when you don't, and the mistakes that make it give bad answers.
You want an AI feature that answers questions about your stuff — your help docs, your product catalogue, your company handbook. The model doesn't know any of it. The most common fix has a clunky name: RAG, short for retrieval-augmented generation.
The idea in one paragraph
Before asking the AI a question, look up the relevant pieces of your own content and paste them into the prompt. Then ask the model to answer using those pieces. "Retrieval" is the looking-up; "augmented generation" is the model writing an answer with that extra material.
It's an open-book exam. The model is a smart student who never read your documents; RAG hands it the right pages just before each question.
Why not just give the model everything?
Two reasons:
- Size. Models have a context window — a limit on how much text they can read at once. Modern windows are large, but a big help centre, years of support tickets or a product catalogue can still exceed it.
- Cost and speed. You pay per token read. Sending a whole library with every question is slow and expensive. Sending the five relevant paragraphs is cheap.
If your content is small — a few dozen pages — you may not need RAG at all. Put it in the prompt (with prompt caching, repeated text costs much less). Start there, and reach for RAG when the content outgrows it.
How RAG works, step by step
Once, ahead of time: index your content
- Split documents into chunks — a few paragraphs each, ideally along headings.
- Turn each chunk into an embedding — a list of numbers that represents its meaning. See what are embeddings?
- Store the chunks and their embeddings in a database that can search by similarity — a vector database, or Postgres with pgvector.
Every time someone asks a question
- Embed the question the same way.
- Search for the chunks whose embeddings are closest to the question's — the most similar meanings.
- Build a prompt: instructions, the top few chunks, and the question.
- Ask the model to answer using only those chunks, and to say when they don't contain the answer.
- Show the answer, ideally with links to the source chunks so people can check.
A tiny example
A user asks: "Can I get a refund after 30 days?"
The search finds two chunks from your refund policy. The prompt becomes:
Answer the customer's question using only the policy excerpts below. If they don't answer it, say you're not sure and suggest contacting support.
Excerpt 1: Refunds are available within 14 days of purchase... Excerpt 2: After 14 days, we offer account credit instead...
Question: Can I get a refund after 30 days?
The model answers from your actual policy instead of guessing.
Search doesn't have to be "vector" search
Embedding search finds text with similar meaning ("money back" matches "refund"). Classic keyword search finds exact words (product codes, error messages, names). The best RAG systems often use both — "hybrid search". If your content is full of exact identifiers, start with keyword search; Postgres has good built-in full-text search.
Why RAG gives bad answers
Most RAG problems are retrieval problems, not model problems. If the right chunk wasn't found, no model can answer well.
- Chunks that are too big or too small. Huge chunks bury the answer; tiny ones lose context. Split along headings and keep the heading with the chunk.
- Stale content. Change a document and forget to re-index it, and the AI quotes the old version.
- No "I don't know". Without explicit instructions, models fill gaps confidently. Tell it to admit when the excerpts don't cover the question.
- No testing. Write down 20–30 real questions with known answers and check them after every change. It's the only way to know whether a tweak helped.
- Leaking private data. If different users can see different documents, filter the search by the user's permissions before anything reaches the model.
The summary
- RAG = find the relevant pieces of your content, put them in the prompt, ask the model to answer from them.
- It saves tokens and works when your content is too large to send every time.
- Small content? Skip RAG and put it straight in the prompt.
- Most failures are bad retrieval; test with real questions.
EasySpawn servers come with PostgreSQL ready to go, so your app, its RAG index and its data can live in one database on one server — and Claude Code can build and test the whole pipeline there. See how it works or join the waitlist.
Related: What Is an LLM? · What Are Embeddings? · pgvector Tutorial · How to Add an AI Chatbot to Your App · Fine-Tuning vs RAG
Keep reading
What Is an SDK? SDK vs API, Explained for Beginners
An SDK is a ready-made toolkit for using a service from your code. What an SDK contains, how it differs from an API, real examples from Stripe and Anthropic, how to install one, and why AI tools sometimes invent SDK functions that don't exist.
Why Does AI Hallucinate? And How to Reduce It in Your App
AI models sometimes state false things with complete confidence. Why it happens — they predict plausible text rather than look up facts — the kinds of hallucination you'll meet in coding and apps, and practical ways to reduce it: grounding, tools, structure, checks and room to say 'I don't know'.