RAG vs fine-tuning: which one does your product need?
Retrieval fixes missing knowledge; fine-tuning fixes missing behaviour. Most products need retrieval first, fine-tuning rarely, and both only when an evaluation proves a specific failure survives prompt and retrieval improvements.
RAG
Retrieval-augmented generation: the model is given the right passages from your documents at answer time, with citations.
Fine-tuning
Training a model further on examples so it learns a behaviour, format or vocabulary. It does not reliably learn facts.
Verdict by criterion
| Criterion | RAG | Fine-tuning |
|---|---|---|
| Knowledge that changes | Refresh the index in minutes | Requires retraining |
| Per-user permissions | Enforced at retrieval time | Impossible — weights cannot forget |
| Citations | Returns the passage | Cannot cite |
| Strict output format | Prompting may drift | Learned reliably from examples |
| Latency and cost at scale | Retrieval adds a step | A small tuned model can beat a large general one |
| Data needed | Your documents as they are | Hundreds to thousands of reviewed examples |
| Time to production | 3–8 weeks | 4–10 weeks, dominated by labelling |
| Reversibility | Change the index or prompt | Retrain |
The question that decides it
Does the model lack knowledge or behaviour? Missing facts about your products, policies, customers and documents is a retrieval problem. A house style, a rigid output format or a domain vocabulary the model keeps misreading is a fine-tuning problem.
Where both are needed
- Extraction at volume: a tuned small model extracts fields fast; retrieval supplies the reference data it validates against.
- Voice and chat in regional languages: behaviour is tuned; the facts the agent speaks come from retrieval and tools.
- Strict-format generation over changing sources: tune for the template, retrieve for the content.
What it costs to choose wrong
Fine-tuning to "teach the model our company" produces a model that is confidently out of date within weeks and cannot forget a departed customer's data. Skipping fine-tuning where a parser depends on an exact format produces a feature that works in demos and breaks in production. An evaluation on a golden set answers it before either mistake is paid for. The long version is in RAG vs fine-tuning.
Frequently asked questions
Can fine-tuning replace RAG?
▾
Not for knowledge that changes or needs permissions and citations. It is for behaviour, not facts.
How much data does fine-tuning need?
▾
Hundreds of clean, reviewed examples for a narrow behaviour; thousands for broader ones. Labelling is usually the largest cost.
Which should we build first?
▾
Retrieval, because it is faster and reversible, and because it tells you whether a behaviour problem exists at all.