azyware
Technology

Zero data egress: designing AI that never leaves your VPC

EZ
Eazyware
· 7 min read
Quick answer

What does zero data egress mean for an AI deployment?

Zero egress means models, vector stores and logs inside your network, with private-first routing and no public fallback for sensitive data. It is enforced by network policy, not by a vendor contract, and done properly it covers inference, retrieval, observability and the developer tooling around them.

Zero data egress AI is a deployment where no prompt, document, embedding, log or trace crosses the boundary of your network. The model runs on GPUs you control, the vector store and cache sit beside it, observability stays internal, and the routing layer has no path to a public API for data classified as sensitive. The guarantee comes from egress rules on the VPC and the cluster, not from a clause in a supplier's terms. This article sets out how we design that for private agentic AI clients, what usually leaks in a first attempt, and what it costs relative to an API-first build.

What zero data egress AI actually requires

Most teams think of the model first, but the model is one of six components that handle your data. Each needs to live inside the boundary, and each has a habit of quietly phoning home unless configured otherwise.

ComponentWhere data leaks in a naive buildZero-egress design
InferenceCalls to a public model APISelf-hosted model on private GPUs; routing has no public route for sensitive classes
EmbeddingsHosted embedding endpointOpen-weight embedding model served in the same cluster
Vector store and cacheManaged vector database in the vendor's cloudSelf-hosted store in your VPC, encrypted at rest with your keys
ObservabilityTrace SaaS receiving full promptsSelf-hosted tracing with prompts stored internally
Developer toolingIDE assistants and notebooks with default telemetryTelemetry disabled or proxied; no production data in developer environments
EvaluationLLM-as-judge via a public APIJudge model served internally, or human review on samples

VPC AI deployment: the network design

The design starts with a private subnet for the inference and retrieval services, security groups that allow inbound traffic only from your application tier, and an egress policy that denies all outbound traffic by default. Package registries, model weight downloads and OS updates go through an internal mirror or a tightly scoped allow-list that is reviewed. On Kubernetes, network policies enforce the same rule per namespace, so a new service added later inherits the restriction rather than opting into it. Cloud providers document VPC endpoints and private connectivity for their managed services; for example, AWS's VPC documentation covers interface endpoints that keep traffic off the public internet. The point is that the rule is expressed in infrastructure code, tested in CI, and visible to an auditor.

Private-first routing

Model-agnostic routing is useful even in a private deployment, because you may run two or three open-weight models sized for different tasks. The routing layer classifies each request by data sensitivity and task, then chooses a model. For sensitive classes the only permitted destinations are internal. If your policy allows public models for non-sensitive work such as marketing copy, that route exists as a separate, explicitly listed path with its own logging, and it is the exception rather than the default. We have written about the routing itself in multi-model routing.

Air-gapped LLM: when the network is not enough

Some environments, typically defence, certain banking systems and research facilities, require an air-gapped LLM with no network path to the outside at all. The engineering is the same but the operational burden is higher: model weights, container images and updates arrive through a controlled transfer process, evaluation datasets must be built inside, and the team needs a rehearsed procedure for patching. We advise clients to reserve true air-gapping for the systems that genuinely require it and to use a strict-egress VPC for everything else, because the maintenance cost of air-gapping is substantial and it is easy to let a system fall behind on security updates.

Data residency AI: keeping it in-region

Residency is a related but different requirement: the data must stay within a jurisdiction, not necessarily within your network. A GPU instance in an Indian region inside your VPC satisfies both. What catches teams out is the supporting services: a managed vector store defaulting to a US region, a logging service replicating across continents, or a backup bucket configured for cross-region redundancy. Every storage location in the architecture should be enumerated with its region, and the list should be part of the security review. For Indian regulated entities, the RBI's guidance on outsourcing and data storage sets the expectations that most private deployments are built to meet.

What it costs compared with an API build

Private deployment trades a variable per-token bill for a fixed infrastructure line and an engineering line. For a mid-size workload the infrastructure is typically one or two GPU instances, a small Kubernetes cluster and storage, which is a predictable monthly figure. The engineering cost is in the initial build and in ownership: someone must update models, patch the serving stack and keep the evals current. Our GPU sizing guide shows how to keep the hardware line honest. The break-even against an API depends on volume and on how much the compliance requirement is worth; for many regulated buyers the second factor decides it before the first is calculated.

Observability without leaking

The most common leak we find in an audit is the tracing tool. Teams adopt a hosted LLM observability product during the pilot, it receives full prompts and responses by design, and nobody removes it before production. The zero-egress version is a self-hosted tracing stack, with prompts stored internally and retention set by policy. Metrics that carry no content, such as latency and token counts, can be exported if you choose, but the default is to keep everything inside. The evaluation loop should also run inside: an internally served judge model, or human review on samples, rather than a public API scoring your private data.

Developer environments and the pilot-to-production gap

The pilot was probably built with a public API key on a laptop. Moving it into a zero-egress environment is a rebuild of the integration layer, not a configuration change, and it should be planned as such. The safest pattern is to build the pilot against the private stack from the start, even at small scale, so that nothing has to be retrofitted. Developer tooling gets the same treatment as production: no real data in notebooks, telemetry disabled or proxied, and synthetic or masked datasets for day-to-day work. This is one of the reasons we scope production constraints in the discovery sprint rather than after the pilot.

A worked example

A lender wanted an assistant that reads onboarding documents, extracts fields and flags inconsistencies, with a hard rule that customer documents never leave its environment. The first architecture proposal used a public model for extraction and a hosted vector store, both of which failed the rule. The redesign placed an open-weight extraction model and an embedding model on GPU instances in a private subnet within an Indian region, a self-hosted vector store beside them, and an internal tracing stack. The egress policy denied everything by default with an allow-list for the OS mirror. During the review, the client's security team used the infrastructure code and a network test to confirm that a request carrying a document could reach nothing outside the boundary. That confirmation, rather than a vendor assurance, is what satisfied their auditor. The pattern matches our KYC document intelligence case study.

Team and timeline

A zero-egress build sits within our private agentic AI service, starting at $31,500 or ₹20.8 lakh plus infrastructure, and takes eight to twelve weeks for a first use case including the platform. Your side provides a cloud or platform engineer who owns the account, a security reviewer, and the data owner for the use case. We provide the infrastructure code, the serving stack, retrieval, evaluation and the runbook. If the model or the use case is still uncertain, a ProofRun built against the private stack answers the question in three weeks. Ongoing ownership is covered by a Standard or Enterprise Care Plan, priced on the pricing page, and you own all code, weights and documentation.

Before you start: a checklist

  • A data classification that says which classes must never leave the boundary
  • An agreed cloud region or data centre, with residency requirements written down
  • A list of every component that touches data: model, embeddings, store, cache, traces, backups
  • Default-deny egress in infrastructure code, with an allow-list that is reviewed
  • A decision on whether any non-sensitive route to a public model is permitted, and how it is logged
  • Self-hosted observability chosen before the pilot, not after
  • Masked or synthetic data for developer environments
  • A named owner for patching and model updates

Questions clients ask

  • Is a private endpoint from a cloud AI service the same as zero egress? No. Traffic stays off the public internet, but your data still reaches a third party's model service. Whether that satisfies your policy depends on your classification and your regulator.
  • Can we still use the best frontier models? For non-sensitive work, through an explicit route. For sensitive data, open-weight models served internally, which now cover most business tasks when evaluated properly.
  • How do we prove it to an auditor? Infrastructure code, network policy tests and an inventory of storage locations by region. Show the test, not the diagram.
  • Does zero egress slow the system down? Not materially; a well-sized private model often has lower latency than a public API because there is no internet hop.

Self-hosted LLMs for BFSI: a practical checklist, AI audit trails: what regulators will ask to see, and what a six-week AI MVP contains.

Zero egress is a network policy you can test, applied to every component that touches your data, and it is far cheaper to design in than to retrofit.

Frequently asked questions

What is zero data egress in AI?

▾

A deployment where prompts, documents, embeddings, logs and traces never cross your network boundary. Models, vector stores and observability run inside your VPC or data centre, enforced by default-deny egress rules rather than by contract.

Is an air-gapped LLM necessary for compliance?

▾

Rarely. Most regulated workloads are satisfied by a strict-egress VPC in the right region. True air-gapping suits defence and similar environments and carries a heavy patching and update burden.

Which component most often leaks data in private AI builds?

▾

Observability and developer tooling. Hosted tracing tools receive full prompts by design, and IDE assistants send telemetry; both must be self-hosted or disabled before production.