Document intelligence: extracting data from PDFs, scans and forms with AI
What should you know about document intelligence before automating data extraction from PDFs, scans and forms?
Document intelligence parses layout, tables and handwriting from PDFs and photos, extracts fields with confidence scores and routes exceptions. It replaces manual keying for invoices, KYC documents, claims and contracts, and the design decision that matters most is what happens when the model is unsure.
Document intelligence is the use of AI to read documents the way a trained clerk does: recognise what kind of document it is, find the fields that matter, read them accurately whether they are typed, printed, photographed or handwritten, extract them into structured data with a confidence score per field, and send anything doubtful to a person. It is the modern form of what used to be called OCR plus templates, and it is different in kind because vision-language models read layout and context rather than matching boxes. This article explains the pipeline, where accuracy comes from, how exceptions are handled, what it costs, and the checks to do before committing.
Why document intelligence matters now
Most organisations still key data from documents by hand: invoices into the ERP, KYC documents into onboarding, claim forms into case systems, delivery notes into logistics platforms. Template-based OCR handled the clean, fixed-layout minority and failed on the rest. Vision-language models changed that: they read a photographed utility bill or a multi-page contract without a template, and they can be asked questions about the page. The work that remains is engineering: validation, confidence, exception routing, and measurement, which is what separates a demo that reads one invoice from a pipeline that processes thousands a day.
The intelligent document processing pipeline
| Stage | What it does | Failure it prevents |
|---|---|---|
| Capture and pre-processing | Deskew, denoise, split multi-document PDFs, detect page orientation | Rotated photos, merged scans, blank pages |
| Classification | Identify document type: invoice, PAN card, bank statement, contract, unknown | Applying the wrong extraction schema |
| Layout and text recognition | OCR or a vision-language model reads text, tables, checkboxes and handwriting with positions | Tables collapsing into word soup; handwriting skipped |
| Field extraction | Map the document to a schema: vendor, invoice number, line items, totals, dates | Fields missed or mislabelled |
| Validation | Cross-checks: totals add up, dates are plausible, IDs match a checksum or a registry lookup | Confident but wrong values |
| Confidence and routing | Per-field confidence; low-confidence or failed-validation documents go to a review queue | Bad data entering systems silently |
| Output and audit | Structured record into the target system, with the source image, coordinates and reviewer decisions retained | No way to explain a value later |
OCR AI versus vision-language models
Classic OCR turns pixels into characters and leaves structure to you. Vision-language models read the page as a whole: they know a table is a table, that the number next to "Total" is the total, that a signature block is not text to extract. For clean printed documents both work; for photographs, multi-column layouts, stamps over text and handwriting, the model is markedly better. The trade-off is cost and speed per page, so production pipelines often use a fast OCR layer for text and layout and a vision-language model for fields the OCR layer is unsure about, or for document types where structure matters. Model choice is per document type, decided by evaluation, and model-agnostic routing lets it change as better models appear. Anthropic's documentation on PDF and image inputs and Google's Document AI are the primary references for the current capabilities.
PDF data extraction AI: where accuracy comes from
A schema per document type
Extraction is only as good as the definition of what to extract. An invoice schema names every field, its type, whether it is required, and how to handle line items. Ambiguity in the schema becomes inconsistency in the output. The schema is written with the people who key the data today, because they know which fields are always present, which are sometimes handwritten, and which vendors format things oddly.
Validation rules
Models make confident mistakes. Rules catch them: line items must sum to the subtotal, tax must be a valid rate, a GSTIN must pass its checksum, an invoice date cannot be in the future, a PAN must match the format. Where a registry exists, look it up. Validation converts a model's guess into a checked value or an exception, and it is the part of the pipeline that most improves trust.
Confidence per field, not per document
A document with twenty fields at high confidence and one doubtful field should not go to a reviewer whole; it should go with that one field highlighted. Per-field confidence, calibrated against a labelled set so that "90%" means roughly ninety in a hundred correct, lets the review queue focus on what needs a person. Calibration is checked on the golden set and re-checked when the model changes.
Exception handling is the product
Every document intelligence system has a review queue, and its design decides whether the system saves time or moves it. The reviewer should see the source image with the extracted field highlighted at its position, the model's value, the validation failure if any, and a one-click accept or correct. Corrections are logged and become evaluation cases. Documents the classifier cannot identify go to a separate queue, because they are often a new document type that needs a schema rather than a one-off fix. Over time the review rate per document type is the health metric: falling is good, rising means a vendor changed their format or the model regressed.
Handwriting, stamps and Indian documents
Handwritten fields on printed forms, stamps and signatures over text, and the wide variation in Indian identity and address documents are where template OCR failed and where vision-language models are strong but not perfect. Expect lower confidence on handwriting and design the review queue for it. Regional-language documents need a model evaluated on that script; a Kannada-language property document is a different evaluation from an English invoice. Regulated extraction such as KYC has its own rules on what may be stored and for how long, covered in the DPDP guide.
Document intelligence and RAG
The same parsing layer that extracts fields also feeds retrieval. A contract parsed with its clauses, tables and headings intact is a far better source for a question-answering assistant than the same contract flattened to text, which is why document intelligence and retrieval and knowledge engineering are one service. Extracted fields become metadata on the chunks (vendor, date, contract value), so retrieval can filter on them. The RAG failure guide lists parsing as the first thing that goes wrong; document intelligence is the fix.
A worked example
A non-banking financial company onboarded borrowers with a KYC pack of identity documents, address proofs, bank statements and photographs, all keyed by a back-office team from images uploaded through field agents' phones. Images were rotated, cropped and photographed under poor light. The pipeline classified each image, ran a vision-language model for extraction with per-field confidence, validated identifiers against format rules and registry checks, and routed low-confidence fields to a review screen that showed the image with the field highlighted. Bank statements, the hardest type, went through table-aware parsing so transactions stayed in rows. The back-office team moved from keying every document to reviewing flagged fields, and the review rate per document type became the weekly quality number. The build is described in the KYC document intelligence case study.
Team and timeline
A document intelligence pipeline needs an AI engineer for extraction and evaluation, a data engineer for the pipeline, integrations and review tooling, and a subject-matter reviewer on your side who labels the golden set and reviews exceptions during the pilot. Four to eight weeks for two or three document types, longer as types are added. It is delivered under retrieval and knowledge engineering from $14,000 / ₹8.8L, or under AI/ML development from $17,500 / ₹11.2L when custom vision models are involved. A three-week ProofRun from $6,250 measures accuracy on a labelled sample of your real documents before the wider build, and a Care Plan handles model changes and new vendor formats afterwards. Prices are on the pricing page.
Before you start: a checklist
- Sample 200–500 real documents per type, including the ugly ones
- Label the golden set with the people who key the data today
- Write the schema per document type, with required fields and line-item rules
- List validation rules and available registry lookups
- Decide the target system and how corrected records flow into it
- Design the review screen before the extraction model is chosen
- Set retention rules for source images under DPDP or your regulator
- Agree the accuracy and review-rate thresholds that count as done
Questions clients ask
- Do we still need OCR? Often, as a fast first layer. Vision-language models read the hard cases; OCR keeps cost down on the easy ones.
- What accuracy should we expect? It depends on document type and image quality, and it is measured on your labelled sample rather than promised in advance. Clean printed documents score far higher than handwritten or photographed ones.
- Can the model read scanned Hindi or Tamil documents? Current models handle major Indian scripts, with accuracy that must be evaluated per script and per document type.
- Where do the documents go? Wherever you decide. The pipeline can run in your cloud with self-hosted models when documents cannot leave.
- Does this replace the back-office team? It changes their work from keying to reviewing exceptions and new document types.
Related reading
Computer vision for documents and quality inspection for custom vision models, What is RAG? for how parsed documents feed retrieval, and the pricing page.
Extract with a schema, validate with rules, score every field, and design the review queue first; the model is the easy part.
Frequently asked questions
What is document intelligence?
▾
AI that classifies documents, reads text, tables and handwriting from PDFs and images, extracts fields into structured data with per-field confidence, validates them and routes doubtful cases to a person for review.
How is it different from OCR?
▾
OCR converts pixels to characters and leaves structure to you. Document intelligence reads layout and meaning, handles photographs and handwriting, validates the values and manages exceptions.
How long before it is processing real documents?
▾
A three-week ProofRun measures accuracy on your labelled sample; a production pipeline for two or three document types takes four to eight weeks after that.