The AI discovery sprint: ten days to a straight answer
What should you know about an AI discovery sprint before you commission one?
A ten-day sprint benchmarks models on your data, sketches the architecture, estimates cost and ends with a go or no-go and a fixed price. It replaces months of vendor conversations with evidence: which model clears your accuracy bar, what the monthly bill looks like, and whether the build is worth funding at all.
An AI discovery sprint is a fixed, short piece of work whose only job is to answer one question honestly: should you build this, and if so, what will it cost and how long will it take? Ten working days is enough. In that time a small team benchmarks candidate models on a sample of your real data, sketches the architecture, prices the running cost and hands back a written go or no-go with a fixed quote for the next phase. If the answer is no, you have spent very little to learn it.
This article explains what happens inside those ten days, what you get at the end, how it differs from a longer AI feasibility study, and how to prepare so the sprint is not spent waiting for data access.
Why the discovery phase of an AI project matters
Most AI projects that fail do so before a line of production code is written. The failure is a decision made on a demo: a model looked impressive on a vendor's curated examples, a budget was approved, and six months later the team discovers the real documents are scanned, the real questions are ambiguous and the real cost per query is ten times the estimate. A discovery phase exists to move that discovery to week one, where it is cheap.
AI scoping is harder than ordinary software scoping for one reason: the behaviour of the system is not fully known until you test it on your data. You cannot read a specification and know whether a model will extract the right field from your invoices. You have to run it. A discovery sprint is the shortest structured way to run it.
What a ten-day AI discovery sprint contains
| Days | Activity | Output |
|---|---|---|
| 1–2 | Stakeholder interviews, decision mapping, data access | One-page problem statement; list of the decisions the AI will change |
| 3–5 | Data audit and sample assembly | A labelled sample of real inputs with expected outputs; data-quality notes |
| 5–8 | Model benchmarking on the sample | Accuracy, latency and cost per candidate model, in a table you can read |
| 7–9 | Architecture sketch and integration survey | Diagram of systems touched, permissions needed, retrieval and routing choices |
| 9–10 | Cost model and recommendation | Monthly running-cost estimate, build estimate, risks, go or no-go, fixed quote |
The days overlap deliberately. Benchmarking starts as soon as a usable sample exists, and the architecture sketch is revised as benchmark results come in. Nothing in the sprint depends on a perfect dataset; it depends on a representative one.
The benchmark is the heart of it
The single most valuable artefact from the sprint is a benchmark table: three to five candidate models, run against a few hundred real examples, scored on the metric that matters for your task. For document extraction that is field-level accuracy. For a support agent it is resolution on a set of real conversations. For natural-language querying it is whether the generated query returns the right rows.
We run the same sample through models from OpenAI, Anthropic, Google and at least one open-weight option, because the answer is often not the most expensive one. Cost and latency go in the same table as accuracy, so the recommendation can say "model B is two points less accurate and a fifth of the price, and the two points fall in a category humans review anyway." That sentence is what a buyer needs and what a demo can never give.
What the benchmark cannot do
It cannot predict production behaviour on inputs unlike the sample, and it cannot substitute for evals during the build. Its job is to remove the largest uncertainty, not every uncertainty. We say so in the report.
The architecture sketch and the cost model
The sketch is a one-page diagram: where data comes from, where the model sits, what it can read and write, where a human reviews, and which systems need integration. It is enough to price the build and to expose the awkward questions early: does the CRM have an API, who approves write access, is the data allowed to leave the country.
The cost model separates build cost from running cost. Running cost is inference (tokens per task multiplied by expected volume), retrieval and storage, any per-minute or per-message channel fees, and support. We show it as a monthly figure at three volumes so finance can see how it scales. A longer treatment of the components is in How much does AI development cost in 2026? and the companion post on total cost of ownership for AI systems.
Go, no-go, or not yet
The recommendation has three possible shapes. Go: the benchmark clears the bar, the integration is feasible, the running cost is acceptable, and here is a fixed price and date for the next phase. No-go: the models do not clear the bar on your data, or the cost per task exceeds the value per task; here is what would have to change. Not yet: the AI is feasible but the data or a dependent system is not ready; here is the preparation work, which is usually ordinary engineering, not AI.
A sprint that says no-go has done its job. It is far cheaper than a proof of concept that drifts for a quarter. The fee is credited to the next build if you proceed, so the sprint costs nothing extra when the answer is go.
Discovery sprint vs AI feasibility study
A traditional feasibility study is longer, more document-heavy and often run by people who will not build the system. It tends to produce a slide deck with options. A discovery sprint is run by the engineers who would build it, on your data, and produces a decision with a price. The difference is not rigour; it is that the sprint is forced to be concrete because it has to end in a quote someone will be held to.
If you are still deciding whether AI belongs on the roadmap at all, start with the questions in AI readiness assessment: the ten questions before you build. The sprint assumes you have a candidate use case.
A worked example
A lending business wanted to automate the first pass of KYC document checks. The sprint started with two days of interviews that narrowed the goal from "automate KYC" to "extract and cross-check six fields from four document types and flag mismatches." The data audit found that a large share of documents were phone photographs, not scans, which changed the sample and the model shortlist. Benchmarking showed two models clearing the accuracy bar on the clean documents and only one holding up on the photographs. The architecture sketch placed a human review step on every flagged mismatch. The cost model priced the running cost per application at a figure well below the operations cost it replaced. The recommendation was go, with a fixed quote for a proof of concept on the full document set. The eventual build is described in the KYC document intelligence case study.
Team and timeline
Ten working days. From our side: a lead engineer, an AI engineer and a part-time architect. From yours: a sponsor who can make the go or no-go decision, a domain expert available for two to three hours in the first week, and someone who can grant read access to a data sample and the systems in scope. Access delays are the only thing that stretches a sprint, so we ask for them before day one.
The sprint is sold as Sprint Zero, our AI discovery sprint, at a fixed price of $3,250 or ₹2,00,000, credited to the next build. It sits inside our AI product strategy service and usually leads to a ProofRun proof of concept (three weeks, $6,250–10,500) or straight to a Launch 6 MVP. Current figures are on the pricing page.
Before you start: a checklist
- Write the decision the AI should change in one sentence, and who makes it today
- Identify a sample of real inputs (documents, tickets, queries) with known correct outputs
- Confirm who can grant read access to the sample and to the systems involved
- Agree the accuracy metric and the threshold that would count as good enough
- Know your expected monthly volume so the cost model has something to multiply
- Name the sponsor who will receive the recommendation and decide
- Note any constraints on where data may be processed or stored
- Set a date for the read-out and put it in the sponsor's diary now
Glossary
- Discovery sprint: a fixed, short engagement that ends in a go or no-go and a priced next step
- Benchmark: candidate models run against the same labelled sample and scored on the same metric
- Labelled sample: real inputs paired with the output a competent human would produce
- Running cost: the monthly cost of operating the system, dominated by inference and channel fees
- Architecture sketch: a one-page diagram of data flow, permissions and integrations, enough to price a build
- Go / no-go / not yet: the three possible recommendations, each with a reason and a next step
Related reading
What a six-week AI MVP actually contains describes the build that usually follows a sprint, and How to choose an AI development company: a 12-point checklist covers what to ask any vendor before commissioning one. For a primary source on how to evaluate models systematically, OpenAI's evaluation guidance is a sound starting point.
Ten days, your data, a table of results and a price: that is the whole point of a discovery sprint, and it is the cheapest way to find out whether an AI project deserves a budget.
Frequently asked questions
How much does an AI discovery sprint cost?
▾
Our Sprint Zero is a fixed $3,250 or ₹2,00,000 for ten working days, and the fee is credited against the next build if you proceed. Details are on the pricing page.
What if we do not have labelled data?
▾
You almost always have real inputs; labelling a few hundred of them is part of the first week. A domain expert's two or three hours is usually enough to build a sample that makes the benchmark meaningful.
Can the sprint be done remotely?
▾
Yes. Interviews are video calls, data access is granted to a controlled environment, and the read-out is a written report plus a call. Studios in Bengaluru, New York and London cover most time zones.