azyware
Business

HIPAA-aligned AI for healthcare providers

EZ
Eazyware
· 7 min read
Quick answer

What should you know about HIPAA AI development before deploying AI in a healthcare organisation?

HIPAA-aligned AI keeps PHI in controlled environments, signs BAAs, logs access and never sends patient data to unvetted models. The rules were written for records systems and apply unchanged to a model that reads a chart, transcribes a consultation or answers a patient's call; the work is vetting vendors.

HIPAA AI development is ordinary HIPAA work applied to a new class of system. The Health Insurance Portability and Accountability Act, through its Privacy Rule, Security Rule and Breach Notification Rule, governs how covered entities and their business associates handle protected health information. An AI assistant that summarises clinical notes, a voice agent that books appointments, a model that extracts fields from referral faxes: each handles PHI, and each has to meet the same standards a records system does. The AI-specific questions are which vendors may see PHI, how PHI moves through prompts and logs, and whether the system's outputs are safe enough to act on. This guide sets out how we approach each, for providers in the US and for organisations elsewhere that adopt HIPAA as a benchmark.

What HIPAA requires, applied to AI

The Privacy Rule limits use and disclosure of PHI to treatment, payment, operations and other permitted purposes, and applies the minimum necessary standard. The Security Rule requires administrative, physical and technical safeguards for electronic PHI, including access controls, audit controls, integrity controls and transmission security, backed by a documented risk analysis. The Breach Notification Rule requires notice to affected individuals and to the Department of Health and Human Services when unsecured PHI is compromised. Any vendor that creates, receives, maintains or transmits PHI on your behalf is a business associate and must sign a Business Associate Agreement before seeing any of it. The rules and guidance are published by HHS. None of this mentions AI, and none of it needs to.

Where PHI and LLMs meet

TouchpointRiskControl
Prompts to a model APIPHI disclosed to the vendor; a business associate relationship whether you intended it or notBAA in place, or de-identify before the call, or run the model in your own environment
Retrieval over clinical documentsAssistant surfaces a record to a user without a treatment relationshipRetrieval filtered by the user's role and the patient's care team, enforced below the model
Observability tracesPrompts and outputs containing PHI stored in a third-party toolSelf-hosted tracing, or redaction before storage; access logged
Voice transcriptionAudio and transcripts are PHI; speech vendors are business associatesVetted speech vendor under BAA, or on-premise speech models
Evaluation datasetsReal patient conversations kept indefinitely for testingDe-identified sets under the Safe Harbor or Expert Determination method; retention limits
Agent actionsA model updates a chart or sends a message with the wrong contentPolicy-gated actions; clinician confirmation for anything entering the record
Fine-tuningPHI baked into weights that cannot be selectively removedAvoid; use retrieval, or fine-tune only on de-identified data

BAA AI vendor decisions

The first gate for any component is whether the vendor will sign a BAA that covers the specific service you intend to use. The major model providers and cloud platforms offer BAAs on enterprise tiers, often only for named services and regions, and often with conditions such as disabling data retention for abuse monitoring. Read the BAA and the service terms together: a platform BAA that excludes the particular API you plan to call does not help. Speech-to-text vendors, observability tools, vector database services and messaging providers all need the same check. Where a vendor will not sign, or the terms are unacceptable, the component runs inside your environment instead. For the model layer that means an open-weight model on your own infrastructure, delivered through our private agentic AI service; the wider decision is covered in self-hosted LLMs: when running your own model beats an API.

Controlled environments and access logging

The Security Rule's audit control requirement means every access to ePHI must be recordable and reviewable. For an AI system that means logging who asked what, which records were retrieved to answer, what the model returned and what action followed, with the log itself protected as ePHI. It also means the retrieval layer, not the prompt, enforces who can see what: a nurse on one ward should not be able to ask the assistant about a patient on another, and the system should refuse at the query level rather than rely on the model to decline. This is the permission-aware retrieval pattern we use in every regulated deployment, and it is what makes an audit answerable. Network boundaries matter as much: PHI processing stays in a defined environment with encryption in transit and at rest, and the environment is part of the documented risk analysis.

De-identification as a design tool

Much AI work in healthcare does not need identified data. Summarising a discharge note, classifying a referral, drafting a patient-facing explanation of a procedure: the model needs the clinical content, not the name and date of birth. De-identifying under the Safe Harbor method, which removes eighteen listed identifier types, before the model call takes the data outside the definition of PHI and widens the choice of vendors and tools. The de-identification step must be reliable, tested and logged, and re-identification happens only in your own environment after the model returns. Where dates or ages matter clinically, the Expert Determination method allows more nuance at the cost of a documented statistical assessment. Either way, the eval set for the system should be built from de-identified data from the start.

Output safety: the part HIPAA does not cover

HIPAA governs data handling, not clinical accuracy. An assistant that keeps PHI perfectly secure and summarises a medication list wrongly is compliant and dangerous. Our evals-over-demos stance matters most here: a golden set of real, de-identified cases with clinician-checked expected outputs, run on every prompt or model change, with error categories that clinicians defined. Shadow mode before autonomy means the assistant drafts and a clinician reviews for as long as it takes to establish the error rate per task, and some tasks stay reviewed permanently. Nothing enters the medical record without a human confirmation. For patient-facing voice, the same applies to booking, triage routing and information given over the phone, as described in the voice agent for appointment booking article.

Breach readiness for AI components

  • Include prompt injection, retrieval over-exposure and trace leakage in the risk analysis and the incident runbook
  • Know which vendors hold PHI so notification can be coordinated under the BAAs
  • Log enough to reconstruct which records a compromised session could have reached
  • Test the deletion path across index, cache and traces so a containment order can be executed
  • Keep the encryption and key management that makes a breach of stored traces a breach of secured PHI

A worked example

A hospital network wanted a multilingual voice agent for appointment booking and follow-up reminders, and an internal assistant for clinical documentation. Its policy did not allow patient data to leave its own infrastructure, which settled the model question: open-weight models on GPUs in its data centre, with speech models deployed alongside. The voice agent identified callers and looked up appointments through the hospital's own systems with role-based access enforced at the API. Transcripts were stored as ePHI with access logging. The documentation assistant drafted summaries that clinicians reviewed and confirmed before anything entered the record, and its golden set was built from de-identified cases the clinical team checked. Non-PHI tasks such as terminology lookups routed to a public API with no patient data in the prompt. The programme is described in the multilingual voice agent case study.

Team and timeline

Healthcare AI needs a compliance or privacy officer on your side who owns the risk analysis and BAAs, a clinical lead who defines error categories and reviews outputs, and an IT owner for the environment. A Sprint Zero discovery sprint at $3,250 / ₹2,00,000 delivers the PHI flow map, the vendor and BAA assessment, the environment decision and the de-identification design. The build is typically a private agentic AI engagement from $31,500 / ₹20.8L plus infrastructure, or a voice agent from $17,500 / ₹11.2L plus per-minute usage where the phone channel is the focus. Enterprise Care Plan cover at $5,250 / ₹3,40,000 a month, with 24×7 and a one-hour critical response, is the usual fit for clinical systems. Bands are on the pricing page.

Before you start: a checklist

  • Map every PHI flow the AI system will touch, from capture to deletion
  • Confirm a BAA covering the specific service and region for every vendor that will see PHI
  • Decide which tasks can run on de-identified data and how de-identification is tested
  • Enforce access at the retrieval and API layer by role and care relationship
  • Plan access logging for prompts, retrievals, outputs and actions, protected as ePHI
  • Define clinical error categories and build a de-identified golden set
  • Require human confirmation for anything entering the record or reaching a patient
  • Update the risk analysis and incident runbook to include the AI components

Glossary

  • PHI: protected health information, individually identifiable health data held by a covered entity or business associate
  • Covered entity: a provider, health plan or clearinghouse subject to HIPAA
  • Business associate: a vendor that handles PHI on a covered entity's behalf under a BAA
  • Minimum necessary: the standard limiting PHI use and disclosure to what the purpose requires
  • Safe Harbor: the de-identification method that removes eighteen identifier types
  • Risk analysis: the Security Rule's required assessment of threats to ePHI and the safeguards in place

Continue with GDPR and AI systems: a builder's guide for EU patients, permission-aware retrieval, and the healthcare industry page. HHS publications are the primary source; this article is an engineering guide, not legal advice.

Keep PHI inside environments and vendors you have vetted, log every access, de-identify where you can, and let clinicians confirm anything that matters.

Frequently asked questions

Can we use ChatGPT or a public LLM API with patient data?

▾

Only on an enterprise tier where the vendor signs a BAA covering that specific service, with data retention for training disabled. Consumer products and default API tiers do not qualify. Otherwise de-identify first or run the model in your own environment.

Is de-identified data still covered by HIPAA?

▾

Data de-identified under the Safe Harbor or Expert Determination method is not PHI and falls outside HIPAA's restrictions. The de-identification step itself must be reliable and documented, and re-identification keys stay in your environment.

Does HIPAA require a human to review AI outputs?

▾

HIPAA does not regulate clinical accuracy. Patient safety, professional standards and your own risk analysis do, and in practice anything entering the record or reaching a patient should have a clinician's confirmation until error rates are established.