azyware
Business

How to scope an AI agent so it ships in six weeks

EZ
Eazyware
· 7 min read
Quick answer

How do you scope an AI agent so that it ships in six weeks?

Scope an AI agent to ship in six weeks by fixing one intent, no more than three tools, a frozen evaluation set and a written cut list before week one. The date holds because the scope is decided in advance, not because the team works faster than physics allows.

You scope an AI agent to ship in six weeks by deciding, before week one, the single intent it will complete, the three tools it may call, the evaluation set that defines done, and the list of things explicitly cut. Six weeks is a scope decision, not a speed decision. Everything that misses the date was added after the start.

What follows is the plan we work backwards from: a week-by-week schedule with a testable condition at the end of each week, the cut list that makes it survivable, and the honest cases where six weeks is not the right target.

Working backwards from the ship date

An AI agent is software that takes a goal, plans steps, calls tools in your systems and returns a completed task. That definition contains the reason agent projects overrun: tools. Every system the agent touches is an integration, an authentication story, a permission decision and a failure mode, and integrations are where estimates die. A six-week plan is therefore mostly a plan about how few systems you will touch.

Start from the end. On the last Friday you need a system running in shadow mode on real traffic, with evaluation scores on a frozen set, an approval gate on anything consequential, and traces you can inspect. Work backwards from that and the weeks fill themselves in. If a task does not serve that Friday, it is not in the six weeks.

This is the same discipline we apply to fast MVPs, described in Scope lock. The difference with agents is that the cut list has to include actions as well as features, because each action carries governance work that features do not.

The six-week plan, week by week

Each week ends with a condition that is either true or false on Friday afternoon. No partial credit; a week that ends amber has cost you the date and you cut scope on Monday rather than hoping.

WeekWhat must be true by FridayWho owns it on your side
Week 0One intent chosen, cut list signed, systems and credentials confirmedProduct owner and a systems administrator
Week 1Evaluation set of 100 to 200 real cases labelled with correct outcomesThe person who does this work today
Week 2Tool contracts built and callable in a test environment, with limitsBackend engineer with API access
Week 3Agent completes the happy path end to end and scores on the eval setEazyware, reviewed by the product owner
Week 4Approval gates live, edge cases handled, escalation path workingProduct owner signs the thresholds
Week 5Shadow mode on real traffic, proposals reviewed dailyOperations lead and the team doing the work
Week 6Acceptance rate reviewed, first intent released with a gate, handover doneEveryone, in one room

Week 1 surprises people. Labelling a hundred real cases feels like administration until the first disagreement about what the correct outcome actually was, which is the most valuable argument in the project and the one you want in week one rather than week five.

What to cut, and cut in writing

The cut list is the artefact that makes the date real. It is written before work starts, signed by the person who can say no, and reviewed only at the end. These are the cuts that buy the most time for the least value lost.

  • Every intent but one. Pick the intent with the highest volume and the clearest correct answer. The second intent is cheaper after the first ships, because the tools and evals already exist.
  • Autonomy. Ship with a human approving the consequential actions. Moving the threshold later is a configuration change; retrofitting a gate is a rebuild.
  • Systems beyond three. Each additional system adds authentication, rate limits, sandbox access and an owner who is on leave. Read-only access to a fourth system is acceptable; write access is not.
  • Channels beyond one. Web, WhatsApp, voice and email each have their own handling. Ship one, add the rest afterwards.
  • Multilingual support, unless the intent is genuinely bilingual from day one, in which case it belongs in the eval set rather than as a later phase.
  • A new interface. Put the agent inside the tool people already use. A new console is a separate adoption problem.
  • Historical backfill. Running the agent over two years of past cases is interesting and is not shipping.

Write next to each cut the week it is scheduled for instead. A cut list without a follow-up plan reads as a refusal; with one, it reads as sequencing, and stakeholders accept it.

How do you define done for an agent?

Done is a completion rate on a frozen set of scenarios, not a good demo. Before week two ends, agree three numbers: the share of cases the agent completes correctly without help, the share it escalates cleanly, and the share it gets wrong. The third number is the one that matters, because an agent that escalates too often is annoying while an agent that acts wrongly is expensive.

Set the release threshold against the process you have today, not against perfection. Human teams make mistakes, reverse them and move on; the fair comparison is like for like, with the advantage that every agent decision is logged. Define the escalation path before you define the accuracy target, because a clean escalation converts a failure into a slower success.

Shadow mode is how the numbers get earned. The agent proposes, a person accepts or corrects, and the acceptance rate becomes your evidence. The mechanics are in Shadow mode: the right way to launch AI agents, and it is why week five exists in the plan rather than being an optional extra at the end.

What a six-week agent costs

A single-intent agent with three tools, approval gates and an eval suite sits at the lower end of our published ranges. Multi-agent systems and workflow orchestration start at $24,500 or ₹16,00,000, and an AI customer service agent starts at $12,500 or ₹8,00,000. Six-week delivery is our Launch 6 programme shape; all starting prices sit on the pricing page.

If the intent is not yet chosen or the systems are unmapped, do not start the six weeks. A ten-day AI Discovery Sprint at $3,250 or ₹2,00,000, credited against the build, produces the intent, the cut list, the tool inventory and the eval plan, which is exactly the week 0 column above. A three-week ProofRun at $6,250 or ₹4,00,000 goes further and proves the hardest step works before you commit to the programme. Both are cheaper than a six-week build that discovers in week three that the ordering system has no write API.

Running cost is separate and usage-based: you pay your own model provider through your own accounts, and we set budgets, routing and dashboards so the figure stays predictable. Estimate the return before you start with the AI agent ROI calculator.

When six weeks is the wrong target

We turn down six-week agent builds regularly, and these are the reasons.

The systems have no API. If the agent's work requires a screen that only a person can drive, you are looking at an automation project with a different shape, or an integration project first. Six weeks does not fit around building the interface the agent needs to use.

The decision rules are genuinely undefined. If three experienced people give three different answers to the same case and cannot reconcile them, no model will resolve it. That is a policy problem, and it wants a workshop, not an agent.

The action is irreversible and high value. Moving money, changing medication records or committing legal positions can still be agent work, but the approval design, audit requirements and sign-off cycle add time that has nothing to do with engineering. Plan twelve weeks and be pleased if it is ten.

The work is pure question answering. If nobody opens another system or changes anything, you want retrieval over your documents, which is a smaller and cheaper project. The distinction is set out in AI agent vs chatbot.

A worked example

For a field-service SaaS company we built an in-app copilot that could act rather than only answer, described in the in-app copilot case study. The scoping decision that made it tractable was choosing job reassignment as the first intent: high volume, a clear correct outcome, and three systems involved rather than seven. Reassignments above a certain scale routed to a human for approval, and the system ran alongside dispatchers while they accepted or corrected its proposals. Later intents reused the same tool contracts and eval harness, which is why the second one took a fraction of the time.

Before week one

  • Name the single intent and the person who does it today
  • Confirm every system has an API and that someone can issue credentials this week
  • Agree who signs off approval thresholds, and book that person's time
  • Export three to six months of real cases for the eval set
  • Write the cut list and get it signed
  • Agree the three numbers that define done: completed, escalated, wrong
  • Block the operations team's time for shadow mode in week five

How much does an AI agent cost to build in 2026 breaks the budget down line by line, and From POC to production: the checklist covers what happens after week six. For tool contracts specifically, the Model Context Protocol documents an open standard for exposing systems to models with explicit, scoped capabilities rather than raw database access.

Six weeks is not a promise about velocity; it is a promise about restraint, and the cut list is where that promise is kept.

Frequently asked questions

Can an AI agent really be built in six weeks?

▾

Yes, for one intent with up to three tool integrations, human approval on consequential actions and an evaluation set agreed in week one. Six weeks does not fit multiple intents, systems without APIs, or workflows where the correct decision has never been written down. Those need discovery first.

What is the most common reason an agent build misses its date?

▾

Integration surprises. A system turns out to be read-only, rate limited, or owned by a team that cannot grant access for a month. Confirming API access and credentials before week one, rather than during week two, removes the single largest source of slippage in agent projects.

How many people do we need on our side?

▾

Three, part time: a product owner who can sign off approval thresholds, a backend engineer with API access for tool contracts, and the person who does the work today to label the evaluation set and review shadow-mode proposals. About half our agent work is paired with an internal team this way.