How long does AI proof of concept development take? A realistic timeline
How long does AI proof of concept development take?
Three weeks, for a proof of concept that ends in a scored decision on one workflow. Eazyware runs ProofRun as fifteen working days from $6,250 or ₹4 lakh. Extra integrations add roughly a week each, and missing evaluation data adds a fortnight before you start.
Three weeks. That is how long AI proof of concept development takes when it is scoped to one workflow and ends in a scored decision rather than a demo. Eazyware runs it as ProofRun, fifteen working days at a fixed price from $6,250 or ₹4,00,000. Each additional system integration adds roughly a week, and absent evaluation data adds a fortnight before day one.
Below is the actual calendar: what happens in each week, which activities genuinely run in parallel, the four things that push the schedule out, and why compressing the work below two weeks produces something that looks finished and decides nothing.
What the three weeks actually contain
The shape is the same across retrieval, extraction and agent work. What varies is the depth of week two, not the structure.
| Week | Focus | What we do | What you provide | Gate at the end |
|---|---|---|---|---|
| Week 0 | Mobilisation | Access requests, NDA, data-handling policy, environment setup | Credentials, sandbox, named workflow owner | Environment live, scope locked |
| Week 1 | Evaluation set and baseline | Freeze 60 to 200 labelled cases, run the simplest possible system, record the score | A domain expert for half a day to two days of labelling | A baseline number everyone accepts |
| Week 2 | Iteration | One mechanism at a time: retrieval, structure, verification, routing, each scored | Answers to domain questions within a day | Best configuration and its cost per case |
| Week 3 | Production shaping and the memo | Failure analysis, cost at real volume, integration assessment, written recommendation | Two hours from the decision-maker | Go, narrow, or stop, signed |
Week 0 is not padding. It is where projects are usually lost: access to a sandbox that needs a security review, a data export that needs a second approval, a workflow owner on leave. We run it in parallel with contracting where the client allows it, which is how fifteen working days stays fifteen.
What runs in parallel, and what does not
Three strands genuinely overlap. Integration scaffolding can be built while the evaluation set is being labelled. Cost instrumentation is wired in on day two and runs silently throughout. Security review of the deployment pattern proceeds alongside the modelling, because the architecture is agreed in week 0 rather than discovered later.
Two strands cannot overlap, however much anyone wants them to. You cannot iterate before the evaluation set is frozen, because you have nothing to iterate against and every improvement becomes a matter of opinion. And you cannot write the memo before the cost-per-case figure is measured on production-shaped inputs, because that number changes conclusions more often than accuracy does. OpenAI documents the general shape of this loop, scoring model outputs against a stored dataset, in its evals guide.
What adds weeks
Four things account for nearly every proof of concept that runs long, and all four are visible before the start if anyone asks.
- No labelled data. If nobody has ever written down what a correct output looks like, add one to two weeks for a domain expert to build the set. This is the single most common delay.
- More than two integrations. Each additional system with authentication, rate limits and a sandbox of its own costs about a week, most of it waiting rather than coding.
- Regulated data. A data protection impact assessment, redaction tooling and a deployment inside your own perimeter add one to two weeks, though much of it runs in week 0 if started early.
- Undecided scope. Two candidate workflows instead of one does not take one and a half times as long; it takes twice as long and produces two weak answers.
- Absent decision-makers. A memo nobody is scheduled to read turns a fifteen-day engagement into a two-month one at no extra effort to anyone.
Does the shape of the work change the timeline?
Less than people expect. The three shapes we prove most often all fit fifteen working days, but the pressure lands in different places, and knowing where helps you staff your side of the engagement.
Retrieval over documents is the most predictable. Week 1 is dominated by getting a clean document set and writing eighty good questions; week 2 is chunking, hybrid search and reranking, each scored in turn. The risk is document quality, not modelling, and it shows up on day three or not at all.
Extraction and classification needs the largest labelling effort, because correctness is field by field rather than answer by answer. Budget two days of a domain expert rather than half a day, and expect week 2 to be spent on the exception cases that carry most of the cost.
Tool-using agents need working sandboxes for every system they touch, which is why they are the shape most likely to need a fourth week. The modelling is rarely the constraint; waiting for a test account with the right permissions almost always is.
Who is on the engagement
One AI engineer full time for three weeks, a backend engineer for roughly a third of that to open integrations, and a delivery lead who holds scope and writes the memo. On your side: a workflow owner for about two hours a week, a domain expert for the labelling, and an administrator who can grant access without escalating. Six people, none of them full time except ours.
How much does the timeline cost?
Eazyware prices the AI POC sprint at $6,250 to $10,500, or ₹4,00,000 to ₹6,80,000, fixed, for three weeks. The range reflects integration count and data sensitivity, not the number of hours we expect to spend, which is the point of a fixed-price engagement. Model API usage runs through your own accounts with a budget set in week one. All starting prices sit on the pricing page.
If the workflow is not yet chosen, the ten-day AI discovery sprint at $3,250 or ₹2,00,000 comes first and is credited against what follows, which makes the combined path five weeks from first conversation to a decision with evidence behind it. If the answer is proceed, a six-week Launch 6 MVP starts at $26,500 or ₹17,60,000 and reuses everything the proof of concept produced.
The week before week one
Access is the long pole
Credentials, sandbox environments and VPN access routinely take longer than any modelling task in the engagement. Start the requests the day the contract is signed, not the day the team arrives. We send the access list with the proposal for exactly this reason.
Data extraction is somebody's real job
Producing a few hundred real records, redacted where needed, is half a day to two days of a data engineer's time, and that person already has a sprint commitment. Book them formally rather than asking a favour, or the evaluation set slips and week one slips with it.
Scope has to be locked, not agreed
A written scope naming one workflow, one metric and one quality bar is what keeps three weeks to three weeks. The discipline is the same one that makes fixed-date MVPs possible, described in scope lock.
When three weeks is the wrong answer
Compressing below two weeks does not produce a faster proof of concept; it produces a demo. There is no time to build an evaluation set, so the system is judged by whoever is in the room, and the conclusion reflects the demo rather than the data. If you only have a week, spend it on discovery instead and come back with a scoped question.
Stretching beyond four weeks is a different failure. Past that point the team is building a product without anyone having decided to, and the sunk cost quietly makes the stop recommendation unsayable. If a workflow genuinely needs six weeks of engineering to evaluate, it is not a proof of concept candidate; it is an MVP, and it belongs on the AI-accelerated MVP track with the governance that comes with it.
There is also the case where a proof of concept is simply unnecessary. When the pattern is well understood and the only real risk is integration with your systems, three weeks of scoring tells you what you already know. Go straight to a scoped build with a fixed price and put the evaluation work inside it.
A realistic calendar with a regulated client
A hospital network wanted multilingual appointment handling by voice, where latency, language mix and clinical safety all needed evidence before anyone committed. Week 0 covered consent language and call-recording policy. Week 1 froze a set of real call transcripts across languages. Week 2 tuned the speech stack and measured turn latency, which is the metric that decides whether callers stay on the line. Week 3 produced the cost per minute at projected volume and the escalation design. The system that followed is described in the multilingual voice agent case study.
The engagement stayed within its fifteen days because the consent review started in week 0 rather than surfacing in week 3, which is the general lesson: compliance work does not extend a timeline, discovering compliance work does.
Checklist to protect the three weeks
- Contract and NDA signed at least one week before the planned start
- Access list issued and every credential requested on signature day
- One workflow, one metric, one quality bar, written down and locked
- Domain expert booked in the calendar for labelling, with dates
- Sandbox available for every system the workflow touches
- Model budget and dashboard agreed before the first API call
- Decision meeting scheduled for the last day of week three
- Named signatory who can accept a stop recommendation
Related reading
Fixed price, fixed date: how we make it work explains the delivery discipline behind a schedule you can plan around, and what a six-week AI MVP actually contains sets out the stage that follows a proceed decision.
Three weeks is not a constraint imposed on the work; it is the shortest honest path from an open question to a decision you can defend.
Frequently asked questions
Can an AI proof of concept be done in one week?
▾
You can build something demonstrable in a week, but not something decisive. A week leaves no time to freeze an evaluation set, so the system gets judged on whoever's examples are to hand. If one week is all you have, run discovery instead and return with a scoped question worth three weeks.
What makes an AI proof of concept run over schedule?
▾
Four causes dominate: no labelled data to score against, more than two system integrations, regulated data whose handling was not agreed up front, and an undecided scope carrying two candidate workflows. All four are visible before the start, which is why we lock scope and issue the access list on signature day.
How soon after a proof of concept can the build start?
▾
Immediately, if the decision meeting is scheduled for the final day. Eazyware's six-week MVP track starts at $26,500 or ₹17,60,000 and reuses the evaluation set, prompts, cost model and architecture unchanged, so there is no rediscovery phase between the proof of concept ending and the build beginning.