Skip to content

Document intelligence

Document intelligence puts frontier models on the contracts, invoices, claims, forms and scans that arrive every day: pulling out the fields, sorting and routing, checking one document against another, and holding back anything it is unsure about for a person to look at. Every value it produces points back to the page it came from. Before any of it touches a live process we measure how often it is right, field by field, on your own documents with the bad scans left in, and you decide against that number whether it goes live.

How the system is put together.

What goes in, what the system is allowed to touch, where a person decides, and the number it gets measured on. We draw every system before we build it, which is the cheapest place to have the argument about what it should do.

Working drawing of document intelligence pipelineFIG. 02DOCUMENT INTELLIGENCE PIPELINEINPUTSScans and PDFsEmails and attachmentsReference recordsFRONTIERMODEL LAYERselected by evaluationTOOLS AND INTEGRATIONSExtraction schemaMatching engineReview queueHUMAN CHECKPOINTLow-confidence fieldsOUTCOME, MEASUREDPrecision and recall
Working drawing of the document intelligence pipeline. Inputs: scans and PDFs; emails and attachments; reference records. These feed a frontier model layer, selected by evaluation. The model layer works through three tools and integrations: extraction schema; matching engine; review queue. Below the model layer there is a human checkpoint on low-confidence fields. The outcome measured is precision and recall.

What actually changes once it is live.

These are the changes we measure. One of them becomes the number in the contract, and the monthly report is written against it for as long as we run the system.

  • Hours of reading taken out of every case
  • Accuracy measured field by field on your own documents, before go-live and after it
  • Every value traceable back to the page it came from
  • Reviewers see the documents that genuinely need a reviewer

What document intelligence gets used for.

Contracts and obligations

Clauses extracted, deviations from your standard terms flagged, and renewal dates and obligations tracked across an estate nobody has read end to end.

Invoices and remittances

Line-level extraction, three-way matching, and exceptions routed to the person who can clear them.

Claims, applications and forms

Mixed-quality scans, photographs and the occasional handwritten page, classified, checked for completeness and turned into fields somebody can work with.

Regulatory and technical documents

Requirements extracted, two versions of a standard compared to show exactly what changed, and evidence mapped for an audit.

Onboarding and due diligence packs

Identity documents, certificates, registers and proofs of address checked against each other and against what the file already says.

Correspondence

What the message is about, what it refers to, who should have it, and a draft reply for them to work from.

From your process to a running system.

  1. 01Take a real sample, including the scans nobody wants to open, and set the baseline on that.
  2. 02Design the fields and the review rules with the people who do the work today.
  3. 03Build the pipeline: confidence thresholds, a review queue for whatever falls below them, and a link back to the page behind every value.
  4. 04Measure precision and recall field by field. You take the go-live decision against those figures, and they keep being measured once it is live.

Good fit

If thousands of documents a month pass through people who are mostly transcribing them, this is the shortest route to a number that moves.

Discuss this use case

FAQ

Document intelligence: questions

Bring the frontier into production.

Tell us about the process you want to change. We reply within one business day.