azyware
Business

Questions to ask a multi-agent system development vendor before you sign

EZ
Eazyware
· 7 min read
Quick answer

What should you ask a multi-agent system development vendor?

Ask a multi-agent system development vendor five things: how they evaluate agents before release, who owns the code and prompts, what the system costs to run monthly, how it escalates to a human, and what happens if you leave. Vendors who have shipped answer in numbers. Vendors who have not answer in adjectives.

Ask a multi-agent system development vendor five things: how they evaluate agents before release, who owns the code and the prompts, what the system will cost to run each month, how an agent escalates to a human, and what happens to you if the relationship ends. Vendors who have shipped answer these in numbers. Vendors who have not answer in adjectives.

What follows is the question bank we would use if we were on your side of the table, grouped into five areas, with the shape of a strong answer and the shape of an evasive one next to each. It is written for the second or third conversation, after the demo, when you are deciding whether to sign.

Why multi-agent system development needs harder due diligence than a normal build

A multi-agent system is several specialised agents coordinated by an orchestrator, each with its own tools, prompts and permissions, working together on a task that no single agent completes cleanly. That structure changes what can go wrong. A conventional web application fails visibly: a page errors, a queue backs up, an alert fires. A multi-agent system fails quietly and plausibly, because a planner can hand a worker the wrong sub-task and the worker will complete it competently.

Vendor selection therefore has to test for engineering discipline you cannot see in a demo. A demo shows the happy path with a prompt the vendor chose. What you need to know is how the system behaves on the two hundred cases nobody chose, what it does when a tool returns nothing, and who finds out when it degrades three weeks after launch.

The OWASP Top 10 for Large Language Model Applications names excessive agency and prompt injection among the principal risks of systems that let a model call tools, which is precisely what a multi-agent system does by design. Any vendor whose answer to that is a content filter has not built one in production.

The five areas, and what a strong answer sounds like

Question areaEvasive answerAnswer from a vendor who has shipped
EvaluationWe test thoroughly before releaseHere is our scenario suite: 240 cases, pass thresholds per agent, run on every prompt change in CI
OwnershipYou get the deliverablesCode, prompts, eval sets, infrastructure definitions and docs in your repository from week one
Running costToken costs are minimalCost per completed task modelled at your volume, with routing and caching assumptions written down
EscalationThere is a human fallbackNamed confidence and policy triggers, an escalation queue with context attached, and a weekly review owner
ExitWe would hand over the projectA handover pack, a named engineer for the transition, and your team running a release before we leave
Failure historyOur projects go live successfullyA description of what went wrong on the last build and what changed in the method afterwards

Questions about evidence, not demos

Start here, because everything else follows from it. Multi-agent system development due diligence is mostly a test of whether a vendor can show you evidence that predates the sales conversation.

  • Show me the eval suite from your last agent build. Redacted is fine. You are looking for real scenarios, expected outcomes and a pass threshold, not a spreadsheet of prompts that someone tried by hand.
  • What is your task completion rate in production, and how is it defined? A vendor who cannot define completion for their own system will not define it for yours.
  • How do you decide when an agent should refuse? Good multi-agent systems have an explicit abstain path. Systems without one fabricate.
  • What proportion of runs escalate to a human today? Any number is credible. No number means nobody is measuring.
  • Which model do you use for the planner, and why that one? The answer should reference a benchmark on their task, not a brand preference.
  • What happens when one worker agent fails mid-task? Listen for retries, compensation, idempotency and a dead-letter path rather than for the word robust.

Our own stance on this is written up in Evals over demos, and the reason is not purity. It is that an eval suite is the only artefact that lets you hold a vendor to a standard after the invoice is paid.

What should multi-agent system development cost, and how should a vendor quote it?

A production multi-agent build sits in the mid five figures in dollars, and a vendor should quote it as a fixed scope with named agents rather than as a monthly rate. Our multi-agent systems programme runs from $24,500 to $84,000, or ₹16,00,000 to ₹56,00,000, depending on how many agents, tools and systems of record are in scope. Every starting figure is published on the pricing page rather than quoted on request.

Ask how the vendor handles the case where the first estimate turns out wrong. A ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, exists exactly so that scoping error surfaces before the contract rather than after it. A three-week ProofRun at $6,250 or ₹4,00,000 proves the hardest agent path on your real data. If a vendor will quote a full multi-agent programme with no discovery and no proof, they are pricing optimism.

Then ask about running cost separately, because it is a different budget line. API usage should sit in your own accounts, not resold through the vendor with a margin. Ask them to model cost per completed task at your expected volume; you can sanity-check the arithmetic with the LLM inference cost calculator before the meeting.

Questions about ownership, security and exit

Ownership

Ask who owns the prompts. Code ownership is now standard in contracts; prompt and eval ownership often is not, and prompts are where most of the behaviour lives. The correct answer is that you own the code, the prompts, the eval sets, the infrastructure definitions and the documentation, and that they live in your repository from the first week rather than arriving as a zip file at the end.

Security and data handling

Ask which systems each agent can write to, and how that permission is enforced. The answer should describe scoped tools with narrow contracts rather than a service account with broad rights. Ask where data goes: whether prompts leave your boundary, whether a model provider retains them, and whether a self-hosted option exists for the sensitive parts. A security questionnaire for AI vendors covers the full list worth sending in writing.

Exit

Ask the uncomfortable one: if we end this in month four, what do we have? A credible vendor describes a handover pack, a transition engineer and the point at which your team ships a release themselves. An incredible one talks about how unlikely that is.

Where this question list is the wrong tool

There are three situations where grilling a multi-agent vendor wastes everyone's time.

The first is when you do not yet have a workflow that needs multiple agents. If the job is one agent calling three tools, a multi-agent architecture adds coordination cost for nothing, and the honest vendor will tell you so. Read multi-agent systems explained and check that your workflow genuinely decomposes before you run a procurement.

The second is when the real blocker is data. If the records the agents would read are inconsistent, undocumented or locked inside a system with no API, no vendor answer to a due diligence question fixes that. Spend the discovery budget on the data layer instead.

The third is when you are buying a product rather than a build. If a packaged tool covers eighty per cent of the workflow, the relevant questions are about integration, limits and pricing tiers, not about eval methodology. Our own Eazy Chat AI is a product for that reason, and we will say so when it fits better than a custom programme.

What a reference call should actually cover

Ask for a reference whose project is at least six months live, because the interesting failures appear after launch, not during it. On the call, ask the client three things: what the agent got wrong in the first month, how quickly the vendor found out, and who fixed it. You are testing the operating relationship, not the build.

Ask the vendor to walk you through a system in production rather than a prototype. Our in-app copilot case study describes a build where scoped tools, approval thresholds and a shadow-mode period preceded autonomy, and it is the shape of answer you want: specific intents, specific gates, specific review cadence.

The checklist to take into the meeting

  • Ask for a redacted eval suite from a previous multi-agent build
  • Ask for the production task completion rate and its definition
  • Confirm in writing that you own code, prompts, eval sets and infrastructure
  • Confirm API usage runs through your own provider accounts
  • Get cost per completed task modelled at your volume, not a token price
  • Get the escalation design: triggers, queue, context, review owner
  • Agree the shadow-mode period and the evidence needed to shorten it
  • Get the exit and handover terms before you sign, not at the end

Questions to ask before hiring an AI agency covers the general version of this conversation, what a fixed-price AI quote should contain shows what a good proposal looks like line by line, and red flags when hiring an AI development partner lists the signals worth walking away from. The OWASP Top 10 for LLM applications is the primary source we point clients to for the risk categories a multi-agent vendor should be able to discuss without prompting. If you want a second opinion on a proposal already in front of you, send it to us.

The vendor worth signing is the one whose answers get more specific the harder you push, not less.

Frequently asked questions

What should you ask a multi-agent system development vendor first?

▾

Ask to see the evaluation suite from their last agent build, redacted if necessary. It reveals more than any other question: whether they define correct behaviour in advance, whether they test failure paths, and whether they can hold themselves to a measurable standard once your project is live.

Who should own the prompts in a multi-agent system contract?

▾

You should. Code ownership is standard in most contracts, but prompts, eval sets and infrastructure definitions often are not, and that is where most agent behaviour lives. Insist they sit in your repository from the first week rather than being handed over at project close.

How can you tell if a multi-agent vendor has shipped in production?

▾

Production experience shows up as numbers and scars: a stated task completion rate with a definition, an escalation percentage, a description of what broke on a previous build, and a named review cadence after launch. Vendors without production history describe process instead of outcomes.