AI audit trails: what regulators will ask to see
What should an AI audit trail record so that it satisfies a regulator?
Log inputs, outputs, model version, user and reviewer for every automated decision in an append-only store auditors can query. Add the prompt version, retrieved sources, policy checks and the action taken, keep it for the retention period your regulator sets, and make sure a non-engineer can reconstruct any decision.
An AI audit trail is a record that lets someone who was not in the room reconstruct a single automated decision: what came in, what the system knew, which model and prompt version produced the output, what policy checks ran, what action followed, and which human, if any, reviewed it. It lives in an append-only store, it is queryable by case rather than by log line, and it is kept for as long as your regulator says. This article describes the audit trail we build into private agentic AI deployments, the questions regulators and internal auditors actually ask, and how to keep the logging useful rather than merely voluminous.
Why an AI audit trail is different from application logging
Ordinary application logs are written for engineers debugging a failure. An audit trail is written for a reviewer establishing what happened and whether it was permitted. The difference shows in three ways. First, the unit of record is the decision, not the request; a single decision may involve several model calls, retrievals and tool actions that must be tied together. Second, the record must be immutable, because its value depends on nobody being able to tidy it up later. Third, it must be readable by a compliance officer, a lawyer or an examiner, which means plain fields and a viewer, not a search across JSON.
What to log for every automated decision
| Field | What it contains | Why an auditor wants it |
|---|---|---|
| Decision ID and timestamp | Unique reference, time, channel, session | Ties every other record together |
| Input | The user's request or the document, or a reference to it, with sensitive fields masked as policy requires | Shows what the system was asked |
| Context retrieved | Document IDs, chunk references and versions used | Shows what the system knew and whether it was current |
| Model and prompt version | Model name, provider or self-hosted build, prompt template hash, parameters | Lets the decision be reproduced and changes be traced |
| Output | The response or structured result, with confidence where available | The decision itself |
| Policy checks | Which rules ran, their result, and any block or escalation | Proves controls operated |
| Action taken | Tool called, parameters, external system reference | Shows the effect in the real world |
| Human involvement | Reviewer identity, decision, time, reason | Shows accountability and oversight |
Most of these fields already exist somewhere in a well-built system; the work is in collecting them under one decision ID and storing them somewhere that cannot be edited. We treat the trail as a product feature with its own tests, not as a side effect of tracing.
AI logging compliance: what regulators actually ask
The questions are more practical than the frameworks suggest. Can you show me this specific customer's case? Which version of the model was running that week? Who approved the change to the prompt? What did the system do when its confidence was low? Was a person involved, and did they have time to consider it? Which records did the model rely on, and were they the right ones? How long do you keep this, and who can read it? The NIST AI Risk Management Framework frames these under governance, mapping, measurement and management, and the EU AI Act sets logging obligations for high-risk systems that are worth reading even if you are not in scope, because Indian regulators borrow from both. In India, the RBI's expectations on outsourcing, model governance and grievance handling apply to lenders using automation; its guidance is the reference for BFSI clients.
Retention and access
Retention follows the underlying regulation for the decision type: lending and KYC records carry their own periods, health records theirs, and employment decisions theirs. The trail inherits the longest applicable period. Access is role-based: engineers see technical fields, compliance sees everything for a case, and customers may have a right to an explanation under data protection law, which means the trail must be exportable in a form a person can read. The DPDP Act adds obligations on purpose and retention that shape the design.
Explainable AI audit: what explanation can honestly be given
Large language models do not produce a faithful explanation of their own reasoning, and a trail that stores a model's self-description as the explanation is misleading. What can be given honestly is the evidence: the inputs, the retrieved sources with the passages used, the policy rules that applied, the confidence score if calibrated, and the human decision where one was made. For classification and scoring models, feature attributions can be logged. For generative decisions, the retrieved sources and the structured checks are the explanation, and the trail should present them as such. This is the position we take in permission-aware retrieval work too: cite what was used, not what the model says it thought.
Model decision logs: the store and the viewer
The store must be append-only. An object store with write-once retention, a database table with insert-only permissions and periodic hashing, or a dedicated ledger all work; what matters is that deletion and edit are impossible for the application and logged for administrators. The viewer is where most projects under-invest. A compliance officer needs to enter a customer reference or a date range and see decisions as cases, with each field labelled, sources expandable and the human review visible. Without the viewer, the audit trail exists in theory but every request from an examiner becomes an engineering task, and the trail is quietly considered a failure.
Change control belongs in the same trail
Every prompt change, model swap, policy edit and retrieval index rebuild is itself an event that affects decisions. Record it with author, approver, evaluation results and effective time, so that a decision from a given day can be matched to the exact configuration in force. We describe the release discipline in prompt versioning and evaluation; the audit trail is where that discipline becomes evidence.
Keeping the trail private
The trail contains the most sensitive data in the system, by construction. It belongs inside the same boundary as the model and the vector store, encrypted with your keys, and it must not be forwarded to a hosted tracing product. Masking sensitive fields at write time is prudent where the regulation permits it, with the unmasked record available under a stricter role. The zero data egress design covers the network side of this.
A worked example
A non-banking lender automated part of its document verification for onboarding. Before launch, the compliance head asked one question: if a customer disputes a rejection six months from now, what will we show the ombudsman? The answer at the time was a set of application logs and a tracing dashboard. The build added a decision record for every file processed, with the pages read, the fields extracted, the confidence per field, the rules that flagged inconsistencies, the model and prompt version, and the analyst who reviewed the flagged cases and their decision. A viewer let compliance search by application number. When an internal audit sampled cases the following quarter, the team answered each query in minutes. The engagement resembled our KYC document intelligence work, and the audit trail was the part the client valued most.
Team and timeline
The audit trail is designed in the first two weeks of a private agentic AI build and delivered with the first use case. It requires your compliance owner to define retention, access roles and the fields that must be masked, and our engineers to build the decision record, the store and the viewer. For an existing system without a trail, a retrofit typically takes four to six weeks and can be scoped in a discovery sprint. Ongoing changes to prompts and models are recorded as part of a Care Plan. Costs for the build and the plans are on the pricing page, and the trail, like everything else, is code and data you own.
Before you start: a checklist
- A list of the automated decisions in scope, with the regulation that governs each
- The retention period per decision type, from your compliance owner
- Which input fields must be masked at write time and who can see them unmasked
- An append-only store chosen and its immutability tested
- A decision ID that ties model calls, retrievals, checks and actions together
- Change events for prompts, models and policies recorded in the same trail
- A viewer that a compliance officer has used before launch
- An exercise: reconstruct one decision end to end from the trail alone
Glossary
- Decision record: the complete set of fields for one automated decision, tied by an ID
- Append-only store: storage where records can be added but not edited or deleted by the application
- Prompt hash: a fingerprint of the exact prompt template used, so versions can be matched
- Policy check: a deterministic rule applied before or after the model, with its result logged
- Human-in-the-loop record: reviewer identity, decision and reason for a case a person handled
- Retention period: how long the record must be kept, set by the governing regulation
Related reading
AI governance for mid-size companies, policy-gated actions, and self-hosted LLMs for BFSI.
Build the decision record before the first real decision, keep it where nobody can edit it, and give compliance a viewer, and the regulator's questions become routine.
Frequently asked questions
What must an AI audit trail contain?
▾
For each decision: input, retrieved context, model and prompt version, output, policy checks, action taken, and any human reviewer with their decision, all tied to one decision ID in an append-only store.
Can an LLM explain its decision for an audit?
▾
Not reliably. Log the evidence instead: the sources used, the rules that ran, the confidence where calibrated, and the human review. Present those as the explanation rather than the model's own account.
How long should AI decision logs be kept?
▾
As long as the regulation for the underlying decision requires; lending, health and employment each have their own periods. The trail inherits the longest applicable one, set by your compliance owner.