Multi-agent systems explained: planner, worker and reviewer patterns
What should you know about multi-agent systems?
A multi-agent system splits a workflow across specialised agents, a planner, workers and a reviewer, coordinated by an orchestrator with shared state. It earns its complexity only when a single agent's tool list or context has become unmanageable; before that, one well-built agent is cheaper and easier to evaluate.
Multi-agent systems are the architecture you reach for when one agent has too many tools, too long a context, or too many ways to fail. Instead of a single model juggling everything, the work is split: a planner decomposes the goal, workers each handle one kind of task with a short tool list, and a reviewer checks the result before anything leaves the system. An orchestrator holds the shared state and decides who runs next.
That is the whole idea. The rest of this article is about when it is worth the extra parts, which patterns we actually ship, and where teams go wrong.
What a multi-agent architecture is and why it matters
A single agent is a loop: read the goal, pick a tool, observe the result, repeat. It works well up to a point. Past that point three things degrade at once. The prompt grows until the model starts ignoring instructions. The tool list grows until the model picks the wrong one. And the failure modes multiply until you cannot write an evaluation suite that covers them.
Splitting the work fixes all three. Each worker has a short prompt, a handful of tools and a narrow job, so it is easier to test and cheaper to run on a smaller model. The planner sees the whole goal but touches no tools. The reviewer sees the output and the original goal and nothing else. Specialisation is not a stylistic choice; it is how you keep the system evaluable.
The cost is coordination. Shared state, hand-offs, retries and tracing all have to be engineered. That is why the first question is always whether you need it yet. The plain-English version of a single agent is in What is an AI agent?.
The three roles
Planner
The planner receives the goal and produces a structured plan: an ordered list of sub-tasks, each with the worker type that should handle it, the inputs it needs and the condition for success. It does not call business tools. It may call a retrieval tool to understand context. A good planner outputs a plan the orchestrator can execute mechanically, which means the plan is data, not prose.
Workers
Each worker is a small agent with one job: extract fields from a document, query the order system, draft a message, reconcile two ledgers. It has the two to five tools that job needs and nothing else. Workers are the place to use cheaper or open-weight models, because their tasks are narrow enough to evaluate precisely.
Reviewer
The reviewer checks the assembled result against the original goal and against policy: are the required fields present, do the numbers add up, is the proposed action inside limits. It can approve, send a sub-task back to a worker with a reason, or escalate to a human. A reviewer is what turns a plausible output into a trustworthy one, and it should run on a different prompt, and often a different model, from the workers it is checking.
Agent orchestration patterns compared
| Pattern | How it works | Best for | Watch out for |
|---|---|---|---|
| Single agent with tools | One loop, one prompt, all tools | Fewer than eight tools, short tasks | Prompt bloat, wrong-tool selection |
| Planner, workers, reviewer | Planner decomposes; orchestrator dispatches; reviewer gates | Multi-step work across systems | Plan quality; state design |
| Router plus specialists | A classifier sends each request to one specialist agent | Support and intake with distinct intents | Misrouting; overlapping specialists |
| Pipeline | Fixed sequence of agents, each transforming the output of the last | Document processing, ETL-like work | No recovery if a stage fails silently |
| Debate or ensemble | Several agents answer; a judge picks or merges | High-stakes reasoning, research | Cost multiplies; judge bias |
| Hierarchical | Planners managing planners | Very large, long-running programmes | Usually a sign the scope is too big |
Most business workflows fit the second or third row. Hierarchical designs usually mean the problem should be split into separate products.
Shared state: the part that decides whether it works
Agents do not talk to each other in free text if you can help it. They read and write a shared state object: the goal, the plan, each sub-task's status and result, the tool calls made, and the reviewer's verdicts. The orchestrator owns the state; agents are functions over it.
This matters for three reasons. Tracing: every step is in the state, so a failed run can be replayed. Recovery: if a worker crashes, the orchestrator restarts that sub-task, not the whole job. Evaluation: you can assert on the state at any point, which is how scenario tests check that the planner produced the right plan even when a worker later failed.
Frameworks like LangGraph model this explicitly as a graph with typed state, and their multi-agent documentation is the clearest description of the pattern. Whether to use a framework or a workflow engine is a separate decision, covered in AI agent orchestration: LangGraph vs custom vs workflow engines.
Where multi-agent systems go wrong
- Splitting too early. A single agent with six tools and a good prompt beats a five-agent system that nobody can debug. Split when evals show the single agent failing on tool selection or instruction following, not before.
- Free-text hand-offs. If workers pass prose to each other, errors compound and nothing is testable. Use structured outputs with schemas.
- No reviewer. Without a checking step, the system is only as reliable as its weakest worker. The reviewer is the cheapest reliability you can buy.
- Same model everywhere. Workers on a frontier model when a smaller one scores identically on their eval wastes money and latency. Route per role.
- Unbounded loops. A planner that replans on every failure can run forever. Every loop needs a step budget and a spend limit.
- Evaluating the whole, never the parts. Each worker needs its own eval set. End-to-end tests alone cannot tell you which agent regressed.
Evaluating a multi-agent system
We evaluate at three levels. Each worker has a unit eval: fifty to two hundred inputs with expected outputs, run on every prompt or model change. The planner has a plan eval: given a goal, does the plan match the reference decomposition. The whole system has a scenario eval: a set of end-to-end cases with expected final state, escalations and tool calls. Passing thresholds are set per level, and a change is not released until all three pass. This is the same evals-over-demos stance we apply to everything, and it is the only way to expand autonomy with confidence; the metrics are set out in How to measure an AI agent.
A worked example
A non-banking financial company needed to process KYC document packs: several document types per applicant, scanned at varying quality, with cross-checks between them. A single agent with extraction, validation and case-management tools was tried first and was unreliable on the cross-checks, because the prompt could not hold every rule and every document at once.
The rebuilt system used a planner that inspected the pack and produced a per-document plan, extraction workers specialised by document type, a validation worker that ran cross-checks against structured outputs, and a reviewer that compared the assembled case against policy and routed anything ambiguous to an analyst with a summary. Each worker had its own eval set drawn from real packs. Autonomy was granted per document type as evals stabilised. The qualitative outcome and the shape of the build are on the KYC document intelligence case study.
Team and timeline
A production multi-agent system is typically an AI engineer who owns the planner and reviewer, a second engineer for workers and tool contracts, and a backend engineer for state, orchestration and tracing, over eight to sixteen weeks. Weeks one and two produce the workflow map, the state schema and the eval sets. Weeks three to eight build workers individually against their evals, then the planner and reviewer. The remaining weeks run shadow mode and expand autonomy. This is the multi-agent systems service, from $24,500 or ₹16 lakh, with the range depending on the number of workers and systems. If the workflow is not yet mapped, a ten-day Sprint Zero does the decomposition first and is credited to the build. Full details are on the pricing page, and a Care Plan keeps evals running as models change.
Before you start: a checklist
- Confirm a single agent has actually failed on tool selection or context, with eval evidence
- Map the workflow into sub-tasks with clear inputs, outputs and success conditions
- Design the shared state schema before writing any prompts
- Give each worker a tool list of five or fewer
- Decide which model runs each role and why
- Write unit evals per worker and scenario evals end to end
- Set step budgets and spend limits on every loop
- Plan the reviewer's escalation path to a named human
Glossary
- Orchestrator: the code that owns shared state and decides which agent runs next
- Planner: the agent that decomposes a goal into structured sub-tasks
- Worker: a narrow agent with a short tool list that executes one sub-task type
- Reviewer: the agent that checks assembled output against goal and policy
- Shared state: the typed object all agents read and write instead of passing prose
- Hand-off: the transfer of a sub-task between agents, ideally as structured data
- Step budget: the maximum number of loop iterations before forced escalation
Related reading
How to build an AI agent that is safe to run unattended covers the permission and limit design every role needs, MCP explained covers how workers reach your systems, and the private agentic AI page covers running the whole thing inside your own infrastructure.
Split the work when the evals tell you to, keep the state typed and the reviewer independent, and a multi-agent system becomes something you can trust rather than something you hope about.
Frequently asked questions
When do you need a multi-agent system instead of a single agent?
▾
When evaluation shows a single agent failing on tool selection, instruction following or context length. Typically that happens past eight tools or when a task spans several systems with different rules. Before that, one agent is cheaper and easier to test.
What are the main multi-agent architecture patterns?
▾
Planner-worker-reviewer for multi-step work, router plus specialists for intake with distinct intents, pipelines for document processing, and ensemble or debate patterns for high-stakes reasoning. Most business workflows use the first two.
How long does a multi-agent system take to build?
▾
Eight to sixteen weeks including shadow mode, starting at $24,500 for the multi-agent systems service. The first two weeks map the workflow, design the shared state and build evaluation sets before any prompts are written.