azyware
Technology

What is RAG? Retrieval-augmented generation explained for business

EZ
Eazyware
· 7 min read
Quick answer

What is RAG, and what should a business know about retrieval-augmented generation before building with it?

RAG grounds an AI's answers in your documents by retrieving relevant passages first, so responses are specific, current and citeable. It is the standard way to make a language model answer from your policies, contracts and tickets rather than from its training data, and most enterprise AI assistants are built on it.

What is RAG? Retrieval-augmented generation is a way of building AI applications where the model does not answer from memory. When a user asks a question, the system first searches your documents for the passages most likely to contain the answer, then hands those passages to the language model with the question, and the model writes a reply grounded in them, with citations. The model's general knowledge supplies the language; your documents supply the facts. This article explains how it works, why it became the default architecture for enterprise AI, where it fails, what it costs, and how to tell whether it is what your project needs.

Why retrieval-augmented generation exists

A large language model knows what was in its training data up to a cut-off date. It does not know your returns policy, last week's price list, the clause in your supplier contract or the ticket a customer raised yesterday. Asked about them, it will produce fluent, plausible and often wrong text. Two responses are possible: teach the model your content by fine-tuning it, or give it the relevant content at the moment of the question. RAG is the second. It is cheaper, it updates the moment a document changes, it can cite its sources, and it can respect who is allowed to see what. The trade-off against fine-tuning is covered in RAG vs fine-tuning.

How RAG works, in five stages

StageWhat happensBusiness decision hidden inside it
IngestionDocuments are parsed: PDFs, wikis, tickets, spreadsheets, scansWhich sources are in scope and who owns them
Chunking and indexingContent is split into passages, each stored with its meaning (an embedding), keywords and metadataGranularity: clause, section, page, ticket
RetrievalThe question is matched against the index to find the best passagesKeyword, vector or hybrid; how many passages; permission filtering
GenerationThe model answers using only the retrieved passages and cites themTone, refusal behaviour when nothing is found, citation format
EvaluationA golden set of questions checks retrieval and answer quality on every changeWhat accuracy is acceptable, and who signs it off

RAG explained with a concrete question

A customer asks a support assistant: "Can I return a mattress after thirty days if it is unopened?" Without RAG, the model guesses from what retailers generally do. With RAG, the system retrieves the returns policy section on bedding, the exception clause on sealed products and, if the customer is logged in, their order date. The model reads those three passages and answers: unopened mattresses can be returned within forty-five days, the order was placed twenty-eight days ago, here is the link to start the return, and here is the policy clause. If the policy changes tomorrow, the answer changes tomorrow, because retrieval reads the current document. That is what grounded AI answers means in practice.

What RAG is good for

  • Internal knowledge assistants over policies, procedures and wikis
  • Customer support agents that answer from the help centre and account data
  • Contract and document question-answering for legal, procurement and compliance teams
  • Product copilots that explain features from documentation and release notes
  • Research assistants over large report libraries
  • Any application where the answer must be traceable to a source

What RAG is not good for

  • Questions whose answer is a calculation over structured data; that is natural-language data querying, a different architecture
  • Changing how the model writes or reasons; that is prompting or fine-tuning
  • Questions about relationships across many documents ("which suppliers depend on this component?"), where GraphRAG does better
  • Corpora that are mostly images or audio without transcription
  • Situations where nobody can say what a correct answer looks like, because then quality cannot be measured

Where RAG fails, and why the demo misleads

A RAG demo with ten clean PDFs works almost every time. Production corpora are not clean: tables become word soup when parsed badly, one chunk size splits contract clauses in half, vector search misses exact part numbers, indexes go stale, and permissions get forgotten so a user retrieves a document they should never see. Each of these is an engineering problem with a known fix, not a model problem, and the fixes are listed in Why basic RAG fails in production. The practical consequence for a buyer is that the cost of RAG is mostly in ingestion, retrieval quality and evaluation, not in the model.

Retrieval quality is the product

If the right passage is not retrieved, the best model in the world cannot answer correctly. Most production systems use hybrid search, which combines keyword matching for exact terms with vector similarity for meaning, followed by a re-ranker that scores the candidates against the question. The hybrid search guide explains why vectors alone are not enough. Retrieval quality is measured on a golden set of real questions: did the right passage appear, and how much irrelevant material came with it. Those two numbers, recall and precision, tell you more about a RAG system than any demo.

Freshness is the other half of retrieval quality. An index built once answers from the past; policies, prices and tickets change, so the index has to be refreshed on a schedule or on events, with each chunk carrying the version it came from. Citations can then show the document date, and a user can tell a current policy from last year's without asking.

Grounded answers, citations and refusals

The generation stage should do three things beyond writing well: use only the retrieved passages, cite them so the user can check, and say "I could not find that" when retrieval returns nothing relevant. The third is the one demos skip and users notice. A system that improvises when it has nothing loses trust quickly; one that refuses gracefully and offers a hand-off keeps it. Groundedness, whether the answer stayed inside the sources, is measured alongside retrieval; the method is in How to measure RAG quality.

Security and permissions

Because RAG retrieves from your documents, it can retrieve the wrong ones for the wrong person. Every chunk should carry the access rights of its source, and every query should be filtered by the user's entitlements before ranking. This must be enforced in the retrieval layer; asking the model politely not to reveal things is not a control. Permission-aware retrieval covers the design. Data residency is a separate question: RAG can run entirely inside your own cloud with open-weight models when the documents cannot leave, which is the private agentic AI service.

A worked example

A non-banking lender needed to answer questions from branch staff about lending policy, which lived in a mix of PDFs, circulars and emails, and changed often. The first attempt, a vector index over everything with a chat interface, impressed in a demo and was quietly abandoned by staff because it cited outdated circulars and could not find specific clause numbers. The rebuild started with a corpus audit and a golden set of real staff questions, added layout-aware parsing for the circulars, hybrid search so clause numbers matched exactly, versioning so answers named the current circular, and a refusal path when nothing current applied. The same retrieval foundation later fed a document-intelligence pipeline described in the KYC document intelligence case study. Staff adoption followed the citations: once answers showed the clause and the circular date, people trusted them.

Team and timeline

A production RAG system is typically an AI engineer for retrieval and generation, a data engineer for ingestion and refresh, and, where permissions are complex, an architect for the access model. Four to ten weeks depending on document types and permissions, with the golden set built in the first two. The retrieval and knowledge engineering service starts at $14,000 / ₹8.8L; a three-week ProofRun from $6,250 is the right first step when you want to see retrieval quality on your own documents before committing. Ongoing refresh and evaluation run under a Care Plan. All prices are on the pricing page.

Before you start: a checklist

  • List the document sources, their formats and how often each changes
  • Collect 100–300 real questions with verified answers and source passages
  • Decide who may see what and get the permission model signed off
  • Choose whether the answer must cite sources (usually yes)
  • Define what the assistant should say when it finds nothing
  • Decide where it runs: vendor API, your cloud, or fully self-hosted
  • Agree the accuracy threshold that counts as done
  • Name an owner for the knowledge base after launch

Glossary

  • Retrieval-augmented generation (RAG): answering with a language model after retrieving relevant passages from your own documents
  • Embedding: a numeric representation of a passage's meaning, used for similarity search
  • Chunk: a passage of a document stored and retrieved as a unit
  • Hybrid search: keyword and vector retrieval combined, usually with a re-ranker
  • Groundedness: whether an answer stays within the retrieved passages
  • Golden set: real questions with verified answers used to evaluate every change
  • Citation: the source passage shown with the answer so the user can verify it

Why basic RAG fails in production, What does a RAG system cost? and the pricing page. The original retrieval-augmented generation paper by Lewis and colleagues is the primary source for the technique.

RAG is not a product you buy; it is a retrieval system you engineer and measure, and the answers are only as good as the passages it finds.

Frequently asked questions

Is RAG the same as training the AI on our documents?

▾

No. Training changes the model; RAG leaves the model alone and gives it your documents at question time. RAG is cheaper, updates instantly, cites sources and can respect permissions, which is why it is the default for enterprise assistants.

Does RAG stop hallucinations?

▾

It reduces them substantially when retrieval is good and the answer layer refuses to go beyond the sources. It does not eliminate them, which is why groundedness is measured on every change.

How much does a RAG system cost?

▾

Retrieval and knowledge engineering starts at $14,000 / ₹8.8L, and a three-week ProofRun from $6,250 proves retrieval quality on your own documents first. See the pricing page.