azyware
Business

RAG development services in India: costs, delivery models and data rules

EZ
Eazyware
· 7 min read
Quick answer

What does RAG development services cost in India?

RAG development services in India cost $14,000 to $49,000, or ₹8.8 lakh to ₹32 lakh, for a fixed-price build, depending on source count, whether answers must respect access rights, and whether the system runs inside your own infrastructure. Support adds ₹68,000 to ₹3,40,000 a month.

RAG development services in India cost $14,000 to $49,000, or ₹8.8 lakh to ₹32 lakh, for a fixed-price build, depending on how many sources you connect, whether answers must respect access rights, and whether the system runs inside your own infrastructure. Support adds ₹68,000 to ₹3,40,000 a month.

Price is the question most buyers open with, but it is rarely the one that decides the project. Data residency, the delivery model you choose and whether your sector regulator has an opinion all move the budget more than the day rate does. This article covers the Indian numbers, the delivery shapes available in this market, and the rules that constrain where your documents may sit.

What does RAG development services cost in India?

Indian retrieval engineering is priced by scope rather than by headcount, and serious firms publish their bands. Below are ours, quoted in both currencies because Indian clients are invoiced in INR with GST and international clients in USD.

What you are buyingUSDINRDuration
Discovery sprint: corpus audit, question set, thresholds$3,250₹2,00,000Ten days
ProofRun: hardest subset tested against your thresholdsFrom $6,250From ₹4,00,000Three weeks
Single-corpus retrieval build with citationsFrom $14,000From ₹8,80,000Six to eight weeks
Multi-source, permission-aware programmeUp to $49,000Up to ₹32,00,000Twelve to sixteen weeks
Care Plan, Essential$1,000 per month₹68,000 per monthRolling
Care Plan, Enterprise with named engineer$5,250 per month₹3,40,000 per monthRolling
AI system add-on: evals, re-indexing, prompt regression$750 per month₹40,000 per monthRolling

Model API usage sits outside these figures and is billed to your own provider accounts, so you see the real number without a markup. Every starting figure appears on the pricing page, and the scope behind each band is set out on the retrieval and knowledge engineering page.

What the price actually buys at each band

At the entry band you get one corpus, one or two connectors, a retrieval pipeline with hybrid keyword and vector search, reranking, an interface that shows citations, and an evaluation harness with a golden question set your team can re-run. That is a complete system for a support knowledge base, an internal policy assistant or product documentation search.

The upper band is not a bigger version of the same thing. It buys access-filtered retrieval that mirrors your existing groups, six or more connectors with their own authentication, evaluation sets segmented by department, an audit trail of what was retrieved for whom, and usually a deployment inside your own cloud account or data centre. The extra money goes into governance and integration, not into a better embedding model.

Indian pricing sits meaningfully below equivalent US or UK engagements for the same scope, which is the structural reason this market exists at all; software development pricing in India vs the US works through the comparison. The gap is labour cost, not standards, and the evaluation numbers a system must hit are the same either way.

Two adjustments are worth making to any Indian quote before you compare it. The first is GST, which applies to domestic invoices and is often quoted separately, so confirm whether the figure in front of you is inclusive. The second is the internal cost of your own people: the subject-matter experts who decide which documents are authoritative and write the reference answers. That time is the largest unbilled line in most Indian retrieval projects, and it is the one that determines whether the system is accurate.

Delivery models available in the Indian market

  • Fixed price against a locked scope. Predictable and comparable. Works when the corpus and question set are known, which is why a discovery step usually precedes it.
  • Discovery then fixed price. A short paid engagement produces the inputs, and the build is then quoted firmly. This is our default and the reason our quotes hold.
  • Dedicated pod. Two to four engineers for a quarter or more, billed monthly. Appropriate when retrieval is one stream inside a longer platform programme.
  • Time and materials. Common in the Indian market and appropriate only for genuinely open-ended research. For a defined retrieval build it transfers all estimation risk to you.
  • Build and transfer. The vendor builds, then trains and hands over to your team on a fixed date, with the eval harness and runbook as the deliverables that matter.
  • Staff augmentation. Cheapest per hour and the weakest fit for retrieval, because the work needs judgement built from previous corpora rather than extra hands.

The two that fail most often are time and materials on an undefined scope, and staff augmentation on a project with no internal owner. Both end with a system nobody can evaluate. Our stance on commercial shape is set out in fixed price vs time and materials.

Where your data sits, and the rules that decide

The DPDP Act

The Digital Personal Data Protection Act, 2023 is India's general data protection statute and is administered by the Ministry of Electronics and Information Technology. For a retrieval system the practical consequences are specific: you must know which documents in the index contain personal data, you must be able to delete a person's data from the index as well as the source, and you must be able to say which processors, including model providers, saw it. A vector index is a copy of your content, and deletion has to reach it. The obligations are summarised in DPDP Act 2023 and AI.

Sectoral rules that override your preference

Banking, insurance and healthcare clients rarely get to choose freely. RBI outsourcing expectations bring audit rights, exit plans and localisation questions into scope, as discussed in RBI guidelines and AI. Hospital groups have their own constraints on where clinical documents may be processed. In these cases data residency is a design input from week one rather than a deployment decision at the end.

What residency actually costs you

Keeping everything in India is achievable: Indian regions exist for the major clouds, and open-weight models can be self-hosted so no document leaves your boundary. The trade is real. Self-hosting means GPU capacity you pay for whether or not anyone asks a question, and the strongest proprietary models may be unavailable to you, which usually costs a few points of answer quality on hard questions. Decide with numbers from your own evaluation set, not from a policy instinct.

Multilingual content is an Indian design constraint

A great many Indian corpora are not monolingual. Policy documents in English sit beside customer correspondence in Hindi, forms filled in Kannada or Tamil, and support tickets that switch language mid-sentence. This affects the embedding model you choose, the chunking of mixed-script documents, and the golden question set, which must contain questions in the languages people will actually type. Treating multilingual behaviour as a later enhancement is how Indian retrieval projects end up rebuilding the index.

Judging a domestic partner against a global one

Judge both on the same evidence: a production evaluation report with a failed metric in it, a written running-cost model at your query volume, per-role access test fixtures, and a defined exit package. Location changes none of those requirements.

Where an Indian partner has a genuine edge is proximity to Indian content and Indian constraints: mixed-language documents, regional-format scans, GST and regulatory paperwork, and a working knowledge of what an RBI inspection asks for. Where a global firm may have an edge is a long track record in your specific vertical. Our own team works from Bengaluru across IST, UK and US East hours, which is described on the Bangalore page and, for overseas buyers, in working with an Indian AI company from the UK and EU.

When an Indian partner is the wrong choice

If your contract or regulator forbids any processing outside a jurisdiction that India is not in, the question is settled and no amount of price advantage changes it. If your corpus is classified, or export-controlled, the same applies.

There is a subtler case. If your internal stakeholders are unwilling to spend time with a team in another timezone, a retrieval project will suffer more than most, because it depends on subject-matter experts reviewing questions and reference answers week after week. A partner in your own city who gets that attention will beat a cheaper one who does not. And if your requirement is generic search over a small, tidy corpus, buy a product rather than commissioning anyone.

A worked example

An NBFC needed identity and income documents read accurately during onboarding, with nothing leaving their own infrastructure. The constraint set the architecture: self-hosted models, an index inside their boundary, and an evaluation set built from their own rejected applications rather than from clean samples. The cost premium of self-hosting was justified by the residency requirement rather than by accuracy. The programme is described in the KYC document intelligence case study, and the wider pattern for regulated buyers is covered in our fintech and BFSI work.

RAG development services, security and the DPDP Act is the compliance checklist to work through before you sign, outsourcing AI development to India covers how this market has changed, and a ten-day discovery sprint at $3,250 or ₹2,00,000, credited to the build, produces the corpus audit and thresholds that make any of these quotes firm.

Price the residency constraint first; in India it decides the architecture, and the architecture decides the budget.

Frequently asked questions

What does a RAG system cost to build in India?

▾

A fixed-price build runs from $14,000 or ₹8,80,000 for a single corpus with one or two connectors, to $49,000 or ₹32,00,000 for a multi-source, permission-aware programme deployed inside your own infrastructure. Model API usage is billed separately to your own provider accounts.

Does the DPDP Act require RAG data to stay in India?

▾

The Act does not impose blanket localisation, but it does require you to know where personal data is processed, to honour deletion across every copy including the vector index, and to account for processors. Sectoral regulators such as the RBI impose stricter expectations on regulated entities.

Can a RAG system run entirely inside our own infrastructure in India?

▾

Yes. Open-weight models can be self-hosted alongside the index in an Indian cloud region or your own data centre, so no document leaves your boundary. You pay for GPU capacity whether it is used or not, and the strongest proprietary models become unavailable, which usually costs some accuracy on hard questions.