ProofRun
Also: AI POC Sprint
What is ProofRun?
ProofRun is our three-week AI POC Sprint, priced from $6,250 to $10,500, that tests one use case against your real data with a written eval suite, so you know whether it works before funding a build.
What ProofRun means
ProofRun takes the top use case from a Sprint Zero or from your own analysis and answers one question in three weeks: does this work on your data, to a standard you would accept in production? We build a thin working version, connect it to a representative slice of your real documents, tickets, calls or records, and measure it against a golden dataset that you help define in week one.
The deliverable is a proof of concept you can run, plus an eval report with accuracy, cost per unit and failure categories, and a fixed quote for a Launch 6 build if the numbers justify it. Price ranges from $6,250 to $10,500 depending on integration depth and data preparation.
ProofRun is not a demo. A demo shows the happy path on chosen examples; a ProofRun reports the numbers across a sample you did not cherry-pick, including the cases where the system fails. That difference is our engineering stance in practice: evals over demos. It is also not an MVP, because it is not built to be operated, only to be measured.
Who it really matters to
- CTO / Head of Engineering: you get measured accuracy and cost figures on your own data before any architecture is committed.
- Founder / CEO: a three-week, fixed-price answer beats a six-month pilot that never reports a number.
- CFO: the cost-per-unit figure from the eval lets you model the economics of the full build.
- Data lead: building the golden dataset exposes data quality problems early, when they are cheap to fix.
- Compliance officer: the ProofRun runs on a controlled data slice with agreed handling terms, so it can be approved without a full production review.
Why it exists
ProofRun exists because AI capability varies wildly by data: a document extractor that scores well on clean PDFs may collapse on scanned forms from a branch office. Committing to a full build before measuring is how pilots stall. A short, bounded sprint with a written eval suite gives a defensible answer at a small fraction of build cost. The trade-off is that a ProofRun is thrown away rather than extended; it is built for measurement, not operation. We accept that because a measured number is worth more than a fragile prototype that someone tries to ship.
Where it is applied
- A bank testing bank-statement extraction accuracy across formats from a dozen lenders before automating credit decisions.
- A SaaS company measuring whether a text-to-SQL analyst answers its customers' golden questions correctly.
- An insurer running claims-triage classification against a labelled sample of historical claims.
- A hospital measuring speech recognition accuracy for a Kannada and Hindi appointment-booking agent on recorded calls.
- A retailer checking whether semantic search improves catalogue hit rates against a query log.
Is ProofRun a skill?
Eazyware programAn Eazyware program: three weeks, fixed price from $6,250 to $10,500, delivered under the AI POC Sprint service. It usually follows a Sprint Zero and, where the eval numbers justify it, leads to a fixed-price Launch 6 build.
Eazyware service that covers it: AI POC Sprint. Starting prices are on the pricing page.
Frequently asked questions
Can the ProofRun code be extended into the production build?
Parts of it, such as the eval suite and data connectors, usually carry over. The core is built for measurement, not operation, so the production build in Launch 6 starts from a proper architecture rather than stretching the prototype.
What if the ProofRun shows the use case does not work?
Then you have saved a build budget. The eval report explains the failure categories, and often points to a fix such as better data, a narrower scope or a different model that can be tested in a short follow-up.