Fine-Tuning vs RAG: How to Teach an AI About Your Data
You want an AI model to know about your products, documents or customers. RAG looks things up and adds them to the prompt; fine-tuning retrains the model. What each actually does, what each is good at, costs and pitfalls, and why most apps should start with neither.
A general AI model doesn't know your refund policy, your product catalogue, or last week's support tickets. There are two famous ways to fix that — RAG and fine-tuning — and a third, simpler way most people should try first.
Option 0: Just put it in the prompt
Modern models have large context windows. If your knowledge fits in a few thousand words — a policy document, a product list, an FAQ — paste it into the system prompt. With prompt caching, repeating it on every request is cheap.
This is the right answer surprisingly often. Try it before building anything.
RAG: look it up, then answer
RAG (retrieval-augmented generation) is for when your knowledge is too big for the prompt — thousands of documents, a help centre, a database.
- Split your documents into chunks and store them in a searchable index (often with embeddings in a vector database like pgvector).
- When a question comes in, search for the most relevant chunks.
- Add those chunks to the prompt and ask the model to answer using them.
It's an open-book exam: the model doesn't memorise your data, it reads the relevant pages each time. (What is RAG?)
Good at: facts, documents, anything that changes, answers with citations, data that differs per user.
Fine-tuning: change the model itself
Fine-tuning means training a model further on your own examples — usually hundreds or thousands of input/output pairs — so its behaviour changes permanently.
It's teaching, not looking things up. The model learns patterns: a style, a format, a classification, a specialised way of responding.
Good at: consistent tone or format, narrow repetitive tasks (classify this ticket, extract these fields), getting a smaller, cheaper model to do one job as well as a bigger one.
Bad at: teaching facts. Fine-tuned models still make things up about details, and you can't easily update or remove what they "learned." Retraining every time your docs change isn't practical.
Side by side
| Prompt | RAG | Fine-tuning | |
|---|---|---|---|
| Best for | Small, stable knowledge | Large or changing knowledge | Behaviour, style, format |
| Updating knowledge | Edit the prompt | Re-index documents | Retrain |
| Shows sources | Can | Yes, naturally | No |
| Per-user data | Yes | Yes | No |
| Setup effort | Minutes | Days | Days to weeks, plus data |
| Hallucination on facts | Low if info is present | Low if retrieval works | Still happens |
| Cost profile | Tokens per request | Tokens + search infrastructure | Training cost + often cheaper inference |
Common mistakes
- Fine-tuning to teach facts. The most common misconception. Use RAG.
- RAG with bad retrieval. If search returns the wrong chunks, the model answers confidently from the wrong text. Most RAG quality problems are search problems — test what's retrieved, not just the final answer.
- Building RAG for ten pages of docs. Just put them in the prompt.
- Skipping evaluation. Whatever you choose, keep a set of real questions with known good answers and check them after every change.
Can you combine them?
Yes. A fine-tuned model can answer in your house style while RAG supplies the facts. But it's rare for a small app to need both. Start simple and add complexity when tests show you need it.
The decision, simply
- Does the knowledge fit in the prompt? → Put it there.
- Too big or changes often? → RAG.
- The model knows enough but behaves wrong (format, tone, narrow task) and prompting doesn't fix it? → Consider fine-tuning.
The summary
- Try the prompt first; it's often enough.
- RAG looks things up at question time — best for facts and changing data.
- Fine-tuning changes behaviour — best for style, format and narrow tasks.
- Don't fine-tune to teach facts.
EasySpawn servers come with PostgreSQL ready for pgvector, so your app, its documents and its RAG index can live in one database on one server. See how it works or join the waitlist.
Related: What Is RAG? · What Are Embeddings? · pgvector Tutorial · What Is a Context Window?
Keep reading
Why Does AI Hallucinate? And How to Reduce It in Your App
AI models sometimes state false things with complete confidence. Why it happens — they predict plausible text rather than look up facts — the kinds of hallucination you'll meet in coding and apps, and practical ways to reduce it: grounding, tools, structure, checks and room to say 'I don't know'.
What Is a System Prompt? How to Give an AI Its Instructions
The system prompt is the standing instruction that shapes every answer an AI model gives in your app — its role, rules, tone and format. How it differs from user messages, what to put in one, a template you can adapt, and why it's not a security boundary.