Service · documents

Intelligent document processing

Structured, validated data out of messy documents, with a known error rate and a person in the loop exactly where the system is unsure.

The problem

Somewhere in most organisations, people retype data from PDFs, scans and email attachments into another system. Generic OCR gets the characters right and the meaning wrong: the total ends up in the tax field, a table row splits across pages, and a contract’s renewal clause is missed because it was phrased differently this time.

What gets built

A pipeline in which every stage can be inspected and measured:

  1. Intake: files arrive by upload, email, API or a watched folder. Duplicates and unreadable files are caught early.
  2. Classification: what kind of document is this, and which processing path applies? Rules where rules are reliable, a model where they are not.
  3. Parsing: text, layout and tables, using OCR only for pages that need it.
  4. Extraction: the fields you care about, into a fixed schema rather than free text. Language models are used where layouts vary too much for templates, and they are constrained to the schema.
  5. Validation: totals that must add up, dates that must be in range, IDs that must exist in your systems. Most errors are caught by plain code, not by more AI.
  6. Confidence and review: each field carries a confidence score. Low-confidence or invalid records go to a review screen, and corrections feed back into evaluation.
  7. Output: clean records in your database, ERP, spreadsheet or API, with a link back to the source page for every value.

What you get to know

Before anything goes into production you get a measured error rate per field on a sample of your own documents, and a clear statement of which fields are fully automated and which still need a person. That number is the basis for deciding whether automation pays off.

When this is the wrong tool

If the documents come from a system that can export structured data, use the export. If volume is a few dozen documents a month, a well-designed manual form may cost less than any automation. If the source is a single fixed template, a deterministic parser beats a language model on cost, speed and predictability.

How an engagement starts

With a fixed-scope feasibility assessment on a representative sample of your documents. It covers achievable accuracy, which fields need review, and what running it would cost per document.

Have a project like this?

  1. I read your message and reply personally, normally within two working days.
  2. Any questions are clarified by email. A short call only if required.
  3. A written proposal for a fixed-scope first step, with a fixed price.

Evidence and further reading

  • LAB/002

    Document Intelligence Engine

    Receipt extraction with an open-weight vision-language model, checked by plain-code arithmetic and routed to a person when the checks fail. Measured on 200 public receipts.

  • The cheapest way to trust AI document extraction

    A vision-language model read 200 real receipts. Simple arithmetic checks decided which results could skip human review, and every receipt they let through had the right total.

Updated