azyware
Business

Self-hosted LLMs for BFSI: a practical checklist

EZ
Eazyware
· Updated · 6 min read
Quick answer

What should a BFSI company check before self-hosting an LLM?

Banks, NBFCs and insurers can run AI agents without sending customer data to a public API. This is the checklist we use before, during and after every private deployment, with GPU sizing, benchmarking and the regulatory obligations that shape the design.

Banks, NBFCs and insurers ask the same first question: can this run without customer data leaving our environment? Yes. Open-weight models, private vector stores and agent runtimes deploy inside a VPC or on-prem, with no data egress. The work is in doing it properly: classifying data, benchmarking models on your actual tasks, sizing hardware from measurement, integrating identity, and logging every decision in a form auditors accept. This checklist is what we use on private deployments, expanded with the sizing table, the benchmark method, the regulatory obligations and a cost comparison at volume.

Why BFSI needs private AI, in one paragraph

Customer identity documents, statements, transaction histories and case notes are exactly the data that public API terms, cross-border transfers and outsourcing rules make hard to send outside your perimeter. Regulators expect accountability for automated decisions, data localisation where mandated, and an audit trail. A private deployment satisfies all of that without giving up the capability: for retrieval, extraction and workflow tasks, open-weight models match public APIs closely enough that the difference is measured, not felt.

Before you deploy

  • Classify the data each workflow touches: public, internal, confidential, regulated. Decide which tasks may fall back to a public model, if any, and which may never.
  • Benchmark two or three open-weight models on your actual tasks with a golden set. The leaderboard is not your workload; a model that tops reasoning benchmarks may lose on Indian bank-statement extraction.
  • Size the hardware from measured throughput, not from parameter counts. Run the benchmark on a rented GPU first and measure tokens per second at your concurrency.
  • Map identity: SSO, roles, and which documents each role may retrieve. Permission-aware retrieval is designed here, not after.
  • Agree the audit requirement with compliance: what is logged, where, for how long, and who can query it.

GPU sizing: a starting table

WorkloadModel classStarting hardwareNotes
Document extraction and classification, moderate volume7–9B, quantised1× L4 or A10 (24 GB)Quantisation often halves memory with little accuracy loss
Retrieval-augmented assistant for a few hundred users7–14B1–2× A10/L4 or 1× A100 40 GBBatching with vLLM matters more than raw GPU size
Multi-agent workflows with planning steps30–70B2–4× A100/H100 80 GBRoute planning to the large model, workers to small ones
Vision-language extraction from photographed documentsVLM 7–12B1× A100 40 GBImage tokens dominate; measure per page

Treat this as a starting point; the benchmark decides. Serving with vLLM or a comparable engine, with continuous batching, usually delivers two to four times the throughput of a naive setup on the same hardware.

During the build

  • Log every model call with inputs, outputs, model version, user and reviewer, to an append-only store the audit team can query.
  • Put spend and rate limits on agents from day one, and approval gates on any action with financial effect.
  • Run evals on every model or prompt change and keep the reports; they are your evidence when a regulator or auditor asks how you know it works.
  • Keep retrieval permission-aware: a relationship manager and a collections agent must not retrieve the same documents.
  • Route private-first: if a public fallback exists for non-sensitive tasks, make it explicit, logged and reversible.

Regulatory obligations that shape the design

RegimeWhat it means for a private AI system
DPDP Act 2023 (India)Lawful purpose, consent or legitimate use, retention limits, breach notification, a grievance officer, and rights to access and erasure that the system must be able to honour
RBI outsourcing and localisation expectationsAccountability for outsourced technology, payment data stored in India, audit access, and controls over third-party model providers
GDPR / UK GDPR (for EU or UK customers)Lawful basis, DPIA for high-risk processing, processor terms, transfer safeguards; private deployment removes most transfer questions
Internal model-risk policyModel inventory, validation before use, monitoring, and documented change control; the eval suite is the validation artefact

Cost comparison at volume

Public APIs are cheaper at low volume; private deployment wins as usage grows. A rough shape for a lender processing 40,000 document pages a month with a vision-language model plus a retrieval assistant for 300 staff: public API inference in the low thousands of dollars per month and rising with volume; a private stack on one or two rented A100-class GPUs at a comparable monthly figure that does not rise with volume, plus the engineering to run it. Above that scale, private is usually cheaper as well as compliant. We cover the running-cost mechanics in How much does an AI agent cost? and the offering in Agentic AI Solutions (self-hosted).

After launch

Private deployments need care: model updates, index refreshes, cost tuning as usage grows, quarterly reviews of thresholds. Treat it like any other production system with an owner, a monitoring dashboard and a change process, and it will pass the audit. A Care Plan with the AI add-on covers evaluation regression, cost review and re-indexing; the KYC document-intelligence case study shows what the operating rhythm looks like a year in.

A one-page checklist

  • Data classified; egress rules written
  • Golden set built; two or three models benchmarked
  • Hardware sized from measured throughput
  • SSO, roles and document permissions mapped
  • Audit log design approved by compliance
  • Spend limits and approval gates on agents
  • Evals run on every change; reports kept
  • Runbooks, monitoring and an owner in place
  • Quarterly review scheduled

Benchmark method: how we choose the model

We take a golden set of two hundred real cases, including the ugliest documents and the rarest bank formats, and run two or three open-weight candidates plus one commercial API against it on rented GPUs. Each is scored on field-level accuracy by document type, latency at the target concurrency and cost per thousand tasks. The memo that results says which documents each model handles reliably and which need a human queue, so the exception path is designed rather than discovered. The commercial API is included only as a reference point; for regulated data it is not a deployment option.

Common mistakes in private deployments

  • Choosing a model from a leaderboard rather than from your tasks
  • Under-sizing GPUs from optimistic assumptions and discovering it under month-end load
  • Treating security as a checkbox after deployment instead of designing it with the information security lead
  • Building a stack nobody in the organisation can operate; runbooks and training are part of the deliverable
  • Skipping evaluation regression when the model or serving engine is updated
  • Forgetting data residency for backups and logs, not just the model

What the operating rhythm looks like

Monthly: patching, dependency updates, cost review, eval regression on any model or prompt change. Quarterly: threshold review with compliance, golden-set growth from corrections, capacity check against volume. Annually: model refresh benchmark, security review, audit evidence pack. This is what a Care Plan with the AI add-on covers, and it is why private AI is an operating commitment rather than a project. The broader private-agent offering, including agents built on the stack, is on the self-hosted agentic AI page, and the FinTech industry page lists the plays we run for lenders and insurers.

Team and timeline

A private deployment is a platform engineer with GPU and Kubernetes experience, an AI engineer for model evaluation and agents, and a security-minded architect who works with your information security lead, over ten to sixteen weeks. Your side needs an infrastructure owner, a compliance contact for the design review and someone who can produce anonymised samples quickly. The schedule risk is procurement of GPU capacity and security sign-off, both of which start in week one.

Before you start: a checklist

  • Classify the data and write the egress rules
  • Assign an infrastructure owner and a compliance reviewer
  • Produce anonymised samples of the ugliest documents
  • Confirm cloud GPU quota or on-prem hardware
  • Map SSO, roles and document permissions
  • Agree the audit log design with compliance
  • Plan the care rhythm before launch

A note on insurers and NBFC field operations

Insurers face the same constraints with claims documents and medical reports; NBFCs with field agents add offline collection apps and voice reminders in regional languages. The private stack serves all of them: the same models and retrieval layer, different agents on top, each with scoped tools and its own evaluation set. Sizing should account for the peak, month-end for lenders, catastrophe events for insurers, rather than the average day, and the capacity plan should be reviewed quarterly against volume. What does not change across any of them is the audit requirement: every automated decision traceable, every model version pinned, every evaluation report kept.

Related reading: How much does AI development cost? for the deployment-model cost differences, What is an AI agent? for the agents that run on the stack, and AI readiness assessment for the constraints question that leads here.

Frequently asked questions

Are open-weight models good enough for banking tasks?

▾

For retrieval, extraction and workflow tasks, yes, and a three-week ProofRun on your data proves it; hard reasoning tasks may need routing to a larger model.

Do we need to buy GPUs?

▾

No. Cloud GPU instances inside your VPC work; on-prem is for organisations whose policy requires it.

How long does a private deployment take?

▾

Typically ten to sixteen weeks from compliance sign-off to a hardened stack with the first agent live.