azyware
Technology

A security questionnaire for AI vendors

EZ
Eazyware
· 6 min read
Quick answer

What should you ask an AI vendor in a security review?

A security review of an AI vendor needs questions that generic IT questionnaires miss: where prompts and retrieved documents go, whether retrieval respects user permissions, how prompt injection is tested, and what the audit trail records. Thirty questions, grouped by the risk each one covers.

Most enterprise security questionnaires were written for software that stores and retrieves data. AI systems also send your data to third-party models, assemble answers from documents the asking user may not be entitled to see, take actions in your systems, and behave differently after a provider updates a model. None of that is covered by asking whether data is encrypted at rest. These are the thirty questions that actually separate a vendor who has run AI in production from one who has not.

1. Data handling and model providers

  • Which model providers will process our data, and in which regions?
  • Do you have zero-retention or no-training terms in writing with each provider, and can we see them?
  • Is our data ever used to train or improve models, yours or anyone else's?
  • What exactly is sent to the model on each request: the prompt, retrieved documents, user identifiers, conversation history?
  • How long are prompts and completions retained in your logs, and who can read them?
  • Can we run entirely inside our own environment if residency requires it?

The fourth question is the revealing one. Many teams have never audited their own payload and will answer vaguely. The right answer names each field and explains what is redacted before it leaves your perimeter.

2. Retrieval and permissions

  • Does retrieval filter documents by the asking user's entitlements before ranking, or after?
  • Where do permissions live: in the index, in the application, or in the prompt?
  • When someone's access is revoked in the source system, how long until the index reflects it?
  • What happens to documents that are deleted at source?
  • Can you demonstrate a user seeing different answers to the same question based on their role?

If permissions are enforced anywhere other than the retrieval layer, the system is one clever question away from a leak. Ask for the demonstration; it takes five minutes and it is the single most informative thing in the whole review. The pattern to look for is permission-aware retrieval.

3. Guardrails and abuse

  • What actions can the system take without human approval, and where is that list enforced?
  • Are guardrails implemented in code or written into the prompt?
  • How do you test for prompt injection, and is it part of the release gate?
  • What limits exist on spend, rate and blast radius if the system misbehaves?
  • How does the system handle instructions embedded in documents or tickets it reads?

Rules written into a prompt are requests, not controls. Anything a regulator or a customer could be harmed by needs to be enforced by the application before an action executes. OWASP maintains a public list of the top risks for LLM applications that is a reasonable basis for this section, and a vendor who has not read it will say so in their answers.

4. Evaluation and change

  • Do you maintain an evaluation suite on our real cases, and does it gate releases?
  • What happens when a provider deprecates or updates the model we launched on?
  • How would we learn that accuracy had degraded, and how quickly?
  • Is there a rollback path for a prompt change, and how fast is it?
  • Who owns the evaluation set at the end of the engagement?

5. Audit and accountability

  • What is recorded for each automated decision: inputs, outputs, model version, user, timestamp, reviewer?
  • How long are those records kept, and can we query them?
  • If a customer disputes an automated decision from four months ago, can you reconstruct it?
  • Who is accountable when the system is wrong, and what is the escalation path?
  • Can we export the full audit trail if we leave?

The dispute question is the practical test. Systems built without an audit trail cannot answer it at all, and the gap only becomes visible when a complaint arrives. Our own answers to this section are published on the trust centre rather than kept for a call.

6. People, access and exit

  • Who on your team can access our systems and data, and how is that reviewed?
  • Is multi-factor authentication enforced on every account that touches our environment?
  • How is access revoked when someone rotates off our project?
  • On termination, what is deleted, from where, within what window, and is deletion confirmed in writing?

How to read the answers

SignalWhat it usually means
Names a specific region, retention window and provider termThey have been through a real security review before
Answers permissions with 'the prompt tells it not to'Treat as a critical finding, not a nuance
Has never run an injection testThe system has not been attacked yet; that is not the same as safe
Cannot say what happens at model deprecationTheir obligation ends at launch, whatever the contract says
Offers a demonstration rather than a documentThe strongest signal in the whole review

Two practical notes. First, do not send this as a 200-row spreadsheet and accept yes or no answers; four of these questions asked live, with a screen shared, will tell you more than the whole document. Second, weight the answers by what you are actually deploying. A read-only knowledge assistant over public policy documents carries a fraction of the risk of an agent that issues refunds, and a review that treats them identically will either block the first or wave through the second.

The four questions that do the most work

If a full review is not proportionate, or you have one call rather than a procurement cycle, these four cover most of the real risk.

QuestionWhy it mattersA weak answer sounds like
Show me the same question answered by two users with different rolesProves permissions are enforced in retrieval, not suggested in a promptWe restrict it in the system instructions
What leaves our network on a single request?Reveals whether anyone has audited the actual payloadJust the query, essentially
Show me your last evaluation reportProves accuracy is measured rather than assertedWe test thoroughly before each release
Reconstruct one automated decision from last quarterProves an audit trail exists and is queryableWe can check the application logs

Each of these asks for an artefact rather than a statement, which is why they work. A team running AI properly produces all four inside an hour; a team that has built a demo cannot produce any of them at any notice, because the artefacts were never created.

Scoring it without turning it into theatre

Score each section red, amber or green against the risk of the specific deployment, and write down what would move an amber to green. A vendor with no injection testing but a read-only assistant over public content is an amber with a clear remedy. The same gap on an agent that issues payments is a red that blocks the project until it is fixed.

Then re-run the questionnaire at renewal rather than only at onboarding. AI systems change underneath you: providers update models, scopes expand from drafting to acting, and a system approved as read-only quietly gains write access in its third quarter. The DPDP Act expects you to know what your processors are doing on a continuing basis, not once at signature.

What a good vendor does unprompted

The better responses to this questionnaire tend to arrive with things you did not ask for: a data-flow diagram showing exactly what leaves your perimeter, a redaction list, an example of the audit record for one decision, and the evaluation report from their last release. Those artefacts exist or they do not; they cannot be produced during a sales cycle.

If you are early in the process, a paid discovery sprint is a cheap way to test all of this for real, because you see the working practices rather than the answers about them. The security posture that shows up in a ten-day engagement is the one you will live with.

One last habit worth adopting: keep the completed questionnaire, the artefacts and the date in the same place as the contract. When a new CISO arrives, or an auditor asks how the vendor was assessed, the answer is a folder rather than a memory. It also makes the renewal review an hour's work instead of a fortnight's, because you are comparing against a baseline rather than starting again.

Frequently asked questions

Do we need ISO 27001 from an AI vendor?

▾

It helps, and it is not sufficient. Certification covers the organisation's controls, not whether retrieval respects user permissions or whether guardrails are enforced in code. Ask the AI-specific questions regardless of what certificate is on the wall.

What is the single most important question?

▾

Whether retrieval filters by the asking user's entitlements before ranking. Getting that wrong is the failure most likely to cause real harm, and it is invisible until the day it surfaces something private.

How do we assess a small vendor with no certifications?

▾

Ask for demonstrations rather than documents: a permissions test, an injection test, an audit record, an evaluation report. A small team that has shipped production AI can produce all four in an hour.