Computer vision for documents and quality inspection
What should you know about computer vision for business use cases like documents and inspection?
Computer vision reads documents and inspects products or images; modern vision-language models make many cases viable without custom training. The question is no longer whether you have enough labelled images but whether a general model, prompted well and checked with evals, is accurate enough for the decision.
Computer vision business cases fall into two families. The first is reading: invoices, KYC documents, delivery proofs, forms, meter photos, anything where the image contains text or structure a system needs. The second is judging: is this product defective, is this shelf stocked, is this site safe, does this photo show what the claim says it shows. Five years ago both needed a custom model trained on thousands of labelled images. Today a vision-language model handles most reading cases and a good share of judging cases out of the box, and a custom model is the fallback for the narrow, high-volume, high-precision cases rather than the starting point. This article explains how to tell which you have, what each costs, and how we build them.
What computer vision business teams can do now
A vision-language model takes an image and a prompt and returns text or structured data. Point it at a scanned invoice with a schema and it returns supplier, line items and totals. Point it at a photo of a packed order and ask whether the seal is intact and it answers, with a reason. This is a different economics from the classifier era: no labelling project, no training run, a working prototype in days. The trade-off is that a general model is slower and more expensive per image than a small specialised one, and its accuracy on your specific edge cases is unknown until you measure it. Measuring is the work.
Choosing the approach
| Case | First choice | When to go custom | Typical accuracy check |
|---|---|---|---|
| Document OCR AI: invoices, forms, IDs | Vision-language model with a JSON schema | Very high volume where per-page cost matters, or a fixed layout that a small model reads faster | Field-level exact match against a hand-checked set |
| Handwriting and poor scans | Vision-language model, with a confidence flag for review | Rarely; models keep improving here | Field-level match with a review-rate target |
| Visual inspection AI: defects on a line | Vision-language model for prototype and rare defects | High-speed lines, sub-second decisions, or defects too subtle for a general model | Defect recall per defect type on a held-out set |
| Image classification business cases: shelf, site, vehicle condition | Vision-language model with a fixed label set | When the label set is stable and volume justifies a cheaper model | Confusion matrix on a labelled sample |
| Counting and measurement | Detection model (open-weight) with a small fine-tune | Almost always; general models count poorly | Count error against manual counts |
| Video and continuous monitoring | Frame sampling plus a small detector | Always custom for the detector; VLM for the escalation summary | Event precision and recall over a recorded shift |
Document reading: the case most companies start with
Document OCR AI used to mean a text extraction step followed by fragile rules to find the fields. Now the model reads the page and fills the schema directly, which handles layout variation far better. What has not changed is the need for a ground-truth set. Before we tune anything, someone on your team checks a few hundred documents by hand and records the correct values. That set becomes the eval: every prompt change, model change or preprocessing change is scored against it, field by field. Numeric fields are checked for exact match; names and addresses for normalised match. Confidence is not something the model reports reliably, so we derive it from agreement between two passes or between the model and a checksum, and anything below the threshold goes to a human queue. Our KYC document intelligence work for an NBFC follows exactly this shape, and we cover the wider pattern in document intelligence.
Where documents still go wrong
- Multi-page documents where the field you need is on page four and the model was shown page one
- Tables that wrap or merge cells; line items are the hardest field to get right
- Stamps, signatures and overlays that hide text
- Mixed languages and scripts on one page, common in Indian documents
- Photos of documents taken at an angle, with glare, rather than scans
Visual inspection: when a general model is enough and when it is not
Visual inspection AI on a production line has a hard constraint that documents do not: time. If the line moves a unit every second, a general model that takes several seconds per image cannot keep up, and a small detector running on an edge device is the only option. Where inspection happens offline, at the end of a batch or on a returned item, the constraint disappears and a vision-language model prompted with examples of good and bad units is often accurate enough to start. The pattern we recommend is to prototype with the general model to learn which defect types matter and how often they occur, then decide whether the volume and speed justify a custom model. Many clients discover the general model plus a human review queue is the whole answer.
Custom models: smaller, faster, and still needed
When a custom model is right, it is usually an open-weight detector or classifier from a hub such as Hugging Face, fine-tuned on your images. A few hundred labelled examples per class is often enough to fine-tune well, far fewer than training from scratch. The model is small enough to run on a modest GPU or an edge box, decisions take milliseconds, and the per-image cost is negligible at volume. The costs move to labelling, which your domain experts have to do, and to the MLOps loop of monitoring and retraining when the camera, lighting or product changes. Both approaches sit within our AI/ML development service, and a project often uses both: a fast detector for the line and a vision-language model to describe what it caught.
Evals before autonomy
Whatever the approach, the system runs in shadow mode first: it reads or judges every image, its output is compared to what a person decided, and the disagreement rate per field or per defect type is reviewed weekly. Only when the numbers hold for a few weeks does the system start acting without a check on the categories where it is reliable. High-stakes categories, such as a rejected KYC document or a scrapped unit, keep a person in the loop for longer or permanently. This is the same evals-over-demos stance we apply to every AI product; a vision demo on ten cherry-picked images tells you nothing about the thousandth photo taken on a phone in bad light.
A worked example
A retail distributor wanted to check proof-of-delivery photos that drivers uploaded from the field: was the delivery at the right location, was the parcel visible, was the signature or stamp present. A vision-language model with a short structured prompt handled the first pass and returned a verdict with reasons for each check. Six hundred photos checked by the operations team formed the eval set. Poor-light photos and partial parcels were the main failure cases; adding examples to the prompt and a rule to request a second photo when the parcel check failed brought the review rate down to something the team could handle. No custom model was trained. The work was closer to prompt engineering and evaluation than to machine learning, and it shipped inside a Launch 6 MVP alongside the driver app features described in our dispatch platform case study.
Team and timeline
A document or inspection prototype with an eval set is a good fit for a ProofRun POC sprint (3 weeks, $6,250–10,500 / from ₹4,00,000): you supply a few hundred checked examples, we deliver a working pipeline with measured field-level or defect-level accuracy. A production build, with the review queue, integration to your system of record and monitoring, is an AI/ML development engagement from $17,500 / ₹11.2L, typically six to ten weeks. A custom detector for a high-speed line adds labelling time on your side and edge deployment on ours. Ongoing retraining when cameras or products change sits under a Care Plan. The pricing page has the bands.
Before you start: a checklist
- Decide which decision the system makes and what happens when it is unsure
- Collect a few hundred real images, including the bad ones, from the actual capture path (phone, scanner, line camera)
- Have domain experts record the correct answer for each, field by field or defect by defect
- Agree the accuracy floor per field or defect type and the acceptable review rate
- Check the time budget: seconds per image is fine offline, not on a moving line
- Confirm where images can be processed, given data protection obligations on identity documents
- Identify the system of record the results flow into
Glossary
- Vision-language model (VLM): a model that takes images and text prompts and returns text or structured output
- OCR: optical character recognition, the older text-extraction step that VLMs now largely absorb
- Detector: a model that finds and locates objects in an image, used for counting and line inspection
- Fine-tuning: adapting an open-weight model to your images with a modest labelled set
- Field-level accuracy: the share of extracted fields that exactly match the checked value
- Review rate: the share of images sent to a person because confidence was below threshold
Related reading
See how much data do you need to train a model for the labelling question, LLM or ML: choosing the right tool for the general-versus-custom decision, and the retail industry page for related work.
Start with a general model and a checked set of your own images; train a custom one only when the numbers tell you to.
Frequently asked questions
Do we need to label thousands of images for computer vision?
▾
Not usually. A vision-language model needs a few hundred checked examples for evaluation, not training. A fine-tuned detector for a narrow case needs a few hundred labelled images per class, which domain experts can produce in days.
Can computer vision read handwritten forms and poor scans?
▾
Modern vision-language models read most handwriting and degraded scans reasonably, but accuracy varies by field. Measure it on your documents and route low-confidence fields to a review queue rather than assuming.
Is computer vision suitable for real-time production line inspection?
▾
Yes, with a small custom detector on an edge device, because general models are too slow for a moving line. Use the general model to prototype and to explain what the detector caught.