azyware
Technology

Private clinical assistants: AI over discharge notes and protocols

EZ
Eazyware
· 7 min read
Quick answer

What should a hospital know about a clinical AI assistant before deploying one?

Clinical assistants retrieve and summarise notes and protocols inside the hospital's environment, with clinicians deciding. The model runs on hospital-controlled infrastructure, every answer cites its source, access follows the clinician's existing rights, and nothing enters the record unsigned.

A clinical AI assistant, in the form we build, is a private retrieval system over the hospital's own documents: discharge summaries, progress notes, care protocols, formulary rules, departmental guidelines and consent templates. A clinician asks a question, the assistant retrieves the relevant passages from documents that clinician is permitted to see, and it answers with citations. It drafts a discharge summary from the notes; it does not sign it. It finds the protocol for a situation; it does not choose the treatment. The model runs inside the hospital's environment, so no note leaves the perimeter.

This article sets out what a clinical assistant should do, why the deployment must be private, how retrieval over clinical documents differs from retrieval over a wiki, how it is evaluated with clinicians, and what a first deployment looks like. It is written for CMIOs, CIOs, medical directors and quality heads at hospitals and hospital networks.

Why a clinical AI assistant needs to be private

Clinical notes are the most sensitive data a hospital holds. Sending them to a hosted model endpoint creates consent, retention and third-party questions that most hospitals cannot resolve quickly, and in India the DPDP Act and sectoral expectations push in the same direction. A private deployment, with an open-weight model and a vector index running in the hospital's own cloud account or data centre, removes the data egress question and leaves ordinary IT controls: identity, access, logging and retention. The healthcare industry page describes our approach to the sector, and the perimeter design is the same one we use for banks in private AI for banks.

What the assistant does, and what it never does

TaskAssistant doesClinician does
Discharge summaryDrafts from admission notes, progress notes, investigations and medication lists, with each statement linked to its source noteReviews, edits, signs; the draft never enters the record unsigned
Protocol lookupFinds the current version of the relevant protocol and quotes the applicable sectionDecides whether and how it applies to the patient
Handover summarySummarises the last 24 hours of notes for a patient or a ward, with citationsConfirms and adds what the notes do not capture
Formulary and dosing rulesRetrieves the hospital's own formulary text; never computes a dosePrescribes
Patient history questionAnswers only from the record the clinician can access, with citationsInterprets
Diagnosis or treatment recommendationNothing; declines and points to the protocolAll of it

Retrieval over clinical documents is not retrieval over a wiki

Three things make healthcare RAG harder than an internal knowledge assistant. Documents are patient-scoped, so the index must enforce which patient a clinician may query, not just which document types. Protocols have versions and effective dates, so retrieval must return the current version and say so, and superseded versions must be excluded or clearly marked. And clinical notes use abbreviations, templated sections and copy-forward text, so chunking has to respect note structure or the assistant will confidently summarise a section that was pasted from three days earlier. We build the index with patient, encounter, document type, version and author as first-class metadata and filter on them before any semantic search runs.

The retrieval fundamentals, including why naive chunking fails, are in why basic RAG fails in production and chunking strategies for RAG.

Permission-aware access

The assistant inherits the clinician's rights from the HIS or EMR. A consultant sees their patients and their department's protocols; a nurse sees their ward; a quality auditor sees de-identified records for the period under review. The index carries these labels and the query is filtered before retrieval, so the model never receives a passage the user could not open directly. Break-glass access, where the HIS permits it, is mirrored and logged the same way. The design pattern is described in permission-aware retrieval.

Citations, groundedness and refusal

Every factual sentence in an answer links to the note or protocol section it came from, and the clinician can open the source in one click. If the assistant cannot find support in the retrieved documents, it says so rather than filling the gap. Questions that ask for a diagnosis, a treatment choice or a dose are declined with a pointer to the relevant protocol. These behaviours are tested, not hoped for: the evaluation suite includes questions designed to tempt the model into unsupported answers, and a release that fails them does not ship.

Evaluation with clinicians

Before build, a small group of clinicians from the departments that will use the assistant write a golden set: real questions with the answer they would accept and the source they would expect. Draft discharge summaries are evaluated against clinician-written ones for completeness, accuracy and unsupported statements. The assistant runs in shadow mode for several weeks, producing drafts and answers that clinicians review without relying on them, and the review findings drive prompt, retrieval and chunking changes. Go-live is a decision the medical director makes on evidence, not a date on a plan.

Where the draft goes

A discharge summary draft is written into a draft area, attributed to the assistant, and presented to the responsible clinician for editing and signature. The signed version enters the record under the clinician's identity, with a marker that a draft was used, and the draft is retained for audit. The assistant never writes to the signed record. This keeps accountability where the regulator and the hospital's own governance expect it.

Keeping the index current

A clinical assistant is only as good as the freshness of its index. Notes are ingested from the HIS as they are signed, so a handover summary reflects the last entry, not last night's batch. Protocols are ingested from the quality team's document system with version and effective date, and a superseded protocol is removed from retrieval the day the new one takes effect. Deletions follow the record: if a note is corrected or a patient exercises a data right, the index and the logs are updated through the same pipeline. Model and prompt changes are versioned, evaluated against the golden set and recorded, so the quality team can say which version produced any given draft.

A worked example

A multi-site hospital network wanted to reduce the time consultants spent writing discharge summaries and to make departmental protocols easier to find at the point of care. The assistant was deployed privately in the network's cloud account, with an open-weight model, a local embedding model and an index over notes and protocols labelled by patient, encounter, department and version. The first departments were general medicine and orthopaedics, chosen because their discharge summaries were structured and their protocols current. Consultants reviewed drafts in shadow mode for several weeks; the main changes were to how the assistant handled copy-forward text and medication reconciliation. After go-live the qualitative outcomes were that consultants described drafts as "a good first pass to correct rather than a page to write", protocol lookups replaced hunting through shared drives, and the quality team could show, per summary, which notes had been used and who had signed. The voice front-desk work at the same network is described in the multilingual voice agent case study.

Team and timeline

Clinical assistants are built under retrieval and knowledge engineering from $14,000 / ₹8.8L for the retrieval layer, delivered as a private deployment under private and self-hosted agentic AI from $31,500 / ₹20.8L plus infrastructure when the model must run inside the hospital. Your side provides a clinical lead, two or three clinicians per department for the golden set, HIS or EMR access, an information security contact and a cloud account or GPU host. We provide an AI engineer, a retrieval engineer and a lead who owns evaluation and documentation. A first deployment for one or two departments typically takes ten to fourteen weeks including shadow mode. Care Plan tiers for model updates and evaluation reruns are on the pricing page.

Before you start: a checklist

  • Choose one or two departments with structured notes and current protocols
  • Confirm how the HIS or EMR exposes notes and access rights
  • Decide the hosting boundary and reserve GPU capacity
  • Agree the retention and access rules for the index and the logs
  • Recruit clinicians to write the golden set and review drafts
  • Define the draft workflow: where drafts live and how they are signed
  • List the question types the assistant must refuse
  • Name the medical director or CMIO who makes the go-live decision

Glossary

  • RAG: retrieval-augmented generation; the model answers from retrieved documents rather than memory
  • Groundedness: whether each statement is supported by a retrieved source
  • Patient-scoped index: an index that filters on patient and encounter before searching
  • Copy-forward: note text pasted from an earlier note, common in progress notes
  • Break-glass access: emergency access to a record outside normal rights, always logged
  • Shadow mode: the assistant produces output that clinicians review but do not rely on

The Ayushman Bharat Digital Mission is the primary reference for India's health data standards and consent architecture, and HL7 publishes the FHIR standard most modern EMRs expose. Related posts: HIPAA-aligned AI for healthcare providers, patient data and AI in India and the healthcare industry page.

Keep the model inside the hospital, cite every sentence, filter by the clinician's rights, and let clinicians decide; that is a clinical assistant a medical director can sign off.

Frequently asked questions

Can a clinical AI assistant make diagnoses or recommend treatment?

▾

Not in the systems we build. It retrieves and summarises the hospital's own notes and protocols with citations, drafts documents for clinician signature, and declines diagnostic or treatment questions with a pointer to the relevant protocol.

Does patient data leave the hospital?

▾

No. The model, the embedding model, the index and the logs run inside the hospital's own cloud account or data centre. There is no data egress to a third-party model provider.

How do you know the assistant's summaries are accurate?

▾

Clinicians write a golden set before build, drafts are compared with clinician-written summaries in shadow mode for several weeks, and every release is tested for unsupported statements before it ships. Go-live is a clinical decision on that evidence.