Prove the model works on your data before you bet the roadmap on it.
A three-week, engineering-led proof of concept. Real data, real model, measured accuracy. Not a demo.
What is an AI proof of concept?
An AI proof of concept is a working build of the riskiest slice of your idea, run on your real data and measured against agreed accuracy, latency and cost thresholds. Eazyware's ProofRun delivers it in three weeks as production-style code with an evaluation report, so you know whether to proceed before committing a full budget.
| Service line | AI Strategy & Discovery |
|---|---|
| Engagement | Fixed-price program (ProofRun) |
| Duration | 3 weeks |
| Starting price | $6,250 |
| Typical range | $6,250 – $10,500 |
| Deliverables | 4 listed below |
| Delivered from | Bengaluru, India (IST, UK and US East hours) |
| Code ownership | Client owns code, infrastructure, prompts and documentation |
What problem does it solve?
Vendor demos work on vendor data. The question is whether it works on yours: your document formats, your edge cases, your languages, your latency budget.
How do we approach it?
ProofRun exists because demos lie. We build the riskiest slice of the system end to end, on your real data, and measure it against thresholds agreed before we start. Week one gets the data pipeline working and a baseline number on the table, however bad. Week two is iteration: prompts, retrieval strategy, model choice, pre-processing, whatever moves the number. Week three is evaluation and honesty: we run the final system against the golden set, write up where it succeeds and fails by category, and assess what it would take to make it production-ready. The code is written the way production code is written, in a repository you keep, so the proof of concept is the first sprint of the build rather than something to throw away.
What do clients use it for?
- Extraction from your invoices, contracts or forms
- Retrieval over your knowledge base with measured relevance
- A single agent step against your live tools
- Classification or scoring on your historical data
Is it the right fit?
Good fit when
- Teams with a Sprint Zero or clear spec
- Buyers who need evidence before a full build
- Products where accuracy thresholds are non-negotiable
Probably not when
- Ideas without any usable sample data
- Full-feature MVPs (use Launch 6)
What do we build?
- Build the riskiest slice end to end: extraction, classification, retrieval or an agent step
- Run it on a meaningful sample of your data
- Measure accuracy, latency and cost against agreed thresholds
- Document what worked, what failed, and what it takes to productionise
What you get
- Working POC: repository plus hosted demo
- Evaluation report with metrics against your thresholds
- Production readiness assessment
- Updated build proposal
How does the engagement work?
- 01
Week 1: data pipeline and baseline
- 02
Week 2: iterate model, prompt and retrieval
- 03
Week 3: evals, report, readout
What does good look like?
A working system you can demonstrate on your own data, a report with the accuracy, latency and cost figures against your thresholds, and a clear statement of what is production-ready and what is not. The best ProofRuns end with a decision that would have been expensive to get wrong: build, build with a human in the loop for these categories, or do not build because the ceiling is too low.
How does it compare?
| Eazyware | Typical agency | In-house hire | |
|---|---|---|---|
| Time to first result | 3 weeks | 6–12 weeks of discovery before a proposal | 3–6 months to hire, then ramp |
| Pricing model | Fixed scope, milestone billing, INR or USD | Time and materials, open-ended | Salaries, tooling, management overhead |
| AI depth | Multi-model, evals, cost routing, observability as standard | Often a single vendor API and a prompt | Depends entirely on who you can hire |
| Ownership | Client owns code, infra, prompts and docs | Sometimes retained or licensed back | Owned, but concentrated in one or two people |
| After launch | Care Plans with SLA and AI add-on | Change requests at hourly rates | Ongoing headcount whether or not there is work |
Which pitfalls do we design around?
Proof-of-concept work goes wrong when it optimises for the demo: a curated sample, a cherry-picked prompt, a metric chosen after the fact. We fix the sample and the thresholds up front, include the ugly documents, and report by category so a strong average cannot hide a weak segment. We also resist scope creep into a mini-MVP; a ProofRun proves one thing, and proving it properly is the whole point.
What do we measure?
Every engagement is instrumented. These are the numbers you see in the dashboard and the monthly report, not claims on a website.
- Accuracy, precision and recall against thresholds
- p50 and p95 latency
- Cost per task at projected volume
Which technologies do we use?
- Node.js / Python
- Eval harness
- Candidate models benchmarked side by side
Who does the work?
Two AI engineers for three weeks, with an architect reviewing the production-readiness assessment and the principal signing off the report.
What do you need to bring?
A data sample that includes the ugly cases, agreed accuracy, latency and cost thresholds, and a person who can judge whether an output is correct. Access to the target system's API or a staging copy if the slice involves an action. A weekly hour for the review.
Frequently asked questions
POC versus Sprint Zero?
Sprint Zero estimates. ProofRun builds and measures.
Can the POC code be reused?
Yes. It's production-style code, not a notebook.
What if it fails the threshold?
You get the report and the reasons. No build is pushed.
Where does this fit?
AI POC Sprint is part of our AI Strategy & Discovery line. See all pricing or talk to an engineer.