azyware
Business

What to expect in the first 30 days with an AI development partner

EZ
Eazyware
· 7 min read
Quick answer

What should you expect in the first 30 days of working with an AI agency?

Expect access setup, data sampling, a scope lock, weekly demos and an evaluation baseline within the first month. If any of those five has not happened by day 30, something is wrong with the engagement, and the earlier you name it the cheaper it is to fix.

Working with an AI agency for the first time is easier if you know what a healthy first month looks like. Expect five things: access set up in the first week, real data sampled and assessed, the scope locked in writing, a weekly demo you attend, and an evaluation baseline that says how good the system is before anyone tunes it. If any of those has not happened by day 30, the engagement has a problem, and naming it early is far cheaper than discovering it at the deadline.

This article walks through the first thirty days week by week, what the agency should be doing, what you will be asked to do, and the signs that things are going well or badly. It is based on how our own programs run, but the shape is common to any competent AI development partner.

The first month of an AI engagement at a glance

WeekAgency doesYou doOutput
1Kick-off, access requests, environment setup, stakeholder interviewsGrant access, name a decision-maker, provide data samplesWorking environments; access log; interview notes
2Data assessment, use-case refinement, first evaluation set draftedReview data findings, answer domain questions, validate example casesData assessment; draft golden set; scope proposal
3Scope lock, architecture, first end-to-end slice, baseline eval runSign off scope, attend demo, comment on the sliceSigned scope; architecture; baseline results
4Second slice, evaluation tuning, deployment path, running-cost forecastAttend demo, review costs, plan internal users for shadow modeWorking slice; eval trend; cost forecast; shadow-mode plan

Week one: access and the AI project kickoff

The kick-off is a working session, not a presentation. A good agency arrives with a list of what they need: repository access, cloud account or a sandbox in yours, read access to the systems that hold the data, a channel for daily questions, and the names of the people who know the domain. They should also ask who can say yes and who can say no, because a project with three approvers and no decision-maker will stall in week three.

Access is the most common cause of a lost first week. Security reviews, procurement holds and cloud permissions all take longer than anyone plans. The agency should send the access list before the kick-off so your IT team can start early, and should be able to work on public samples or synthetic data while the real access clears. If you are two weeks in and the team still cannot see real data, that is the first warning sign.

What you will be asked in week one

  • Who uses the system, and what they do today without it
  • Which systems hold the data, who owns them, and how stale the data is
  • What a bad outcome looks like: the mistake that would get someone fired
  • Which decisions the AI may make alone and which need a person
  • Who attends the weekly demo and who signs off scope

Week two: data sampling and the first evaluation set

The agency should be looking at real data by now, and the findings will shape everything after. Documents are messier than the spec said; the CRM has three fields for the same concept; half the historical tickets have no resolution recorded. None of this is unusual. What matters is that the agency reports it plainly and adjusts the plan, rather than discovering it in week five.

The other week-two deliverable is the draft evaluation set: a collection of real examples with the correct answer or action recorded. You will be asked to validate it, because only your team knows what correct means. This is the single most valuable hour you will spend in the first month. A partner who does not ask for it is planning to judge the system by how the demo feels. Golden question sets explains how these sets are built.

Week three: scope lock and the baseline

By the end of week three the scope should be written down as a list of deliverables and signed by both sides. This is the scope lock: it defines what the fixed price and fixed date cover, and it is the reference for every change conversation afterwards. Expect the agency to have removed things from your original wish list. That is what a scope lock is for, and the scope lock discipline is what makes a six-week build possible.

Week three should also produce a baseline: the evaluation set run against a first, untuned version of the system. The number will be unimpressive, and that is the point. It is the reference against which every later improvement is measured, and it is how you will know whether a change made things better or only made the demo look better. Ask for the baseline in writing.

The first weekly demo

Demos should show a thin end-to-end slice on real data, not a slide deck. The right response to a demo is a list of things that are wrong, which the agency will fold into the next week. If you find yourself being asked to admire rather than critique, push back. Evals over demos is a stance you should expect the partner to hold too.

Week four: a working slice and the running-cost forecast

By day 30 you should be able to use something, even if it is narrow: one intent of a support agent, one document type in an extraction pipeline, one question type in a data assistant. The evaluation trend should be visible from the baseline. And the agency should have a forecast of running costs, covering inference, hosting, channel fees and support, so nobody is surprised after launch. LLM inference cost forecasting covers what that forecast should contain.

Week four is also when the deployment path should be settled: where the system will run, how it will be released, and how shadow mode will work. Agents that take actions should launch in shadow mode, producing outputs that people review but that do not execute, and the partner should be identifying which of your team will do the reviewing.

Signs the first month is going well or badly

  • Good: you have seen real data on screen by day 10
  • Good: the agency has told you something you did not want to hear about your data
  • Good: the evaluation set exists and your team validated it
  • Good: the scope has shrunk and both sides signed it
  • Bad: demos are decks, or run on invented examples
  • Bad: the same engineers are not on every call
  • Bad: the team is waiting on access and nobody has escalated
  • Bad: nobody can tell you the baseline number

A worked example

A last-mile logistics operator engaged us to build a dispatch platform with an offline-first driver app. Week one was almost entirely access: their order data sat in a system managed by a third party, and it took eight working days to get read access. The team built against a sample export in the meantime. Week two's data assessment found that pickup addresses were free text with no geocoding, which changed the architecture and became the first deliverable. Scope was locked in week three with a narrower first release than the operator had asked for: dispatch for one city, with exception handling deferred. Week four's demo showed assignments running against a day of real orders, with the baseline for assignment quality recorded. The dispatch platform case study describes where it went from there; the first month looked exactly like the table above, including the access delay.

Team and timeline

The first thirty days typically map to Sprint Zero (ten working days, $3,250 or ₹2,00,000, credited to the build) followed by the opening weeks of Launch 6 (six fixed weeks from $26,500 or ₹17,60,000) or ProofRun (three weeks from $6,250) if there is a hard technical question to answer first. From the agency side you get a lead engineer on every call, one or two applied-AI engineers and a designer where needed. From your side, budget three to four hours a week: a decision-maker at the weekly demo, a domain expert for evaluation-set review, and someone from IT for access in week one. The pricing page lists all programs and the AI agents service describes a typical build; to talk through a kick-off, contact us.

Before you start: a checklist

  • Name one decision-maker and one domain expert, and clear their calendars for the weekly demo
  • Start access requests before the kick-off: repositories, cloud, data systems
  • Prepare a sample of real data, anonymised if necessary, for day one
  • Write down the mistake the system must never make
  • Decide which actions the system may take alone and which need a person
  • Agree the demo day and time for the whole program
  • Ask for the evaluation baseline and the scope lock by name
  • Identify the internal reviewers for shadow mode

Questions clients ask

  • What if we cannot give data access in week one? Say so early. The partner can work on samples or synthetic data, but the plan should be adjusted, not quietly slipped.
  • How much of our time does it take? Three to four hours a week for the named people, more in weeks two and three when the evaluation set is being validated.
  • Should the demo be recorded? Yes, and the notes written. Decisions made in demos are the audit trail for scope changes.
  • Can we change scope after the lock? Yes, by trading: something of equal size comes out.

See how to brief an AI development company, what a six-week AI MVP actually contains and shadow mode: the right way to launch AI agents. For evaluation practice at the model level, OpenAI's evaluation guidance is a useful primary reference for what a baseline run involves.

A good first month is access, data, scope, demos and a baseline; if you have those, the rest of the engagement is engineering.

Frequently asked questions

What should happen in the first week with an AI development partner?

▾

Kick-off as a working session, an access list sent in advance, stakeholder interviews, environment setup, and a named decision-maker on your side. Access delays are the most common cause of a lost first week.

What is an evaluation baseline and why does it matter?

▾

It is the evaluation set run against the first untuned version of the system. It gives a reference number so every later change can be measured rather than judged by how a demo feels.

How much client time does the first month need?

▾

Three to four hours a week from a decision-maker and a domain expert, plus IT time in week one for access. Weeks two and three are heaviest because the evaluation set needs validation.