LAB/002 · Building · Self-directed, not client work

Document Intelligence Engine

Receipt extraction with an open-weight vision-language model, checked by plain-code arithmetic and routed to a person when the checks fail. Measured on 200 public receipts.

Average extraction accuracy is the wrong number for document automation. What matters is how much of the work can be handed to the system without a person, and how often the system is wrong on exactly that part. This Lab measures both.

The pipeline

  1. Extract: an open-weight vision-language model (Qwen3-VL-4B, running locally on one consumer GPU) reads the receipt image and fills a fixed JSON schema: line items, subtotal, tax, service, discount, total, cash and change.
  2. Validate: plain code checks the arithmetic that every receipt must satisfy. Line prices add up to the subtotal; subtotal plus tax and service minus discount equals the total; cash minus change equals the total.
  3. Route: a receipt is accepted automatically only if every applicable check passes. Everything else goes to a person.

No second model, no confidence scores from the model itself. Just arithmetic.

The data is CORD-v2: real shop and restaurant receipts from Indonesia, photographed and labelled, published under CC BY 4.0. 200 receipts from the validation and test splits were used. The prompt was written before the run and not tuned on these receipts.

Results

200 receipts. A receipt is auto-accepted only if every applicable arithmetic check passes (72.5% were).
All receiptsAuto-acceptedSent to review
Receipts20014555
Fully correct65.5%80.7%25.5%
Total correct97.4%100.0%90.6%
Line items F10.880.940.76
LAB/002 · measured 2026-10-06 · NVIDIA GeForce RTX 4070 · median 3.4 s per receipt

72.5% of receipts were accepted automatically, and every one of them had the correct total. Of the 5 receipts where the model got the total wrong, all 5 were sent to review. The review queue is where the problems concentrate: only a quarter of the receipts sent there were fully correct, against 81% of the auto-accepted ones.

The checks are not a guarantee. 28 auto-accepted receipts still had some error, and 23 of those were line-item problems: a misread item name, or a line split or merged while the prices still added up. Arithmetic cannot see a wrong name. If item names matter downstream, they need their own check, for example against a product catalogue.

A receipt the checks caught

Scanned receipt validation/2 from the CORD-v2 dataset

RECEIPT VALIDATION/2 · ROUTED: HUMAN REVIEW

Extracted values against ground truth
FieldModel outputGround truth
subtotal18,18118,181
tax10% Tax Included1,818
discount50%—
total20,00020,000
cash100,000100,000
change80,00080,000
itemsS-Ovaltine: 20,000
  • Failed: sum of line prices = subtotal (or total if no subtotal)
  • Failed: subtotal + tax + service − |discount| = total
  • Passed: cash − change = total
Receipt images: CORD-v2 (Park et al., 2019), CC BY 4.0, resized.

The model copied the printed text “10% Tax Included” into the tax field instead of the tax amount, and read “50%” from the item name (“S-Ovaltine 50%” is a sugar level) as a discount. Both arithmetic checks failed, so the receipt went to a person. Nothing about the output looked wrong at a glance. It was well-formed JSON with plausible values.

A receipt that went straight through

Scanned receipt validation/0 from the CORD-v2 dataset

RECEIPT VALIDATION/0 · ROUTED: AUTO-ACCEPT

Extracted values against ground truth
FieldModel outputGround truth
subtotal45,500—
total45,50045,500
cash50,00050,000
change4,5004,500
itemsREAL GANACHE: 16,500EGG TART: 13,000PIZZA TOAST: 16,000
  • Passed: sum of line prices = subtotal (or total if no subtotal)
  • Passed: subtotal + tax + service − |discount| = total
  • Passed: cash − change = total
Receipt images: CORD-v2 (Park et al., 2019), CC BY 4.0, resized.

Field by field

Extraction accuracy by field, all receiptsScale 0 – 100% · n = receipts with that field
  1. total n=19397%
  2. subtotal n=13299%
  3. cash n=13196%
  4. change n=12095%
  5. tax n=8883%
  6. service n=2564%
  7. discount n=1292%
  8. line items (F1) n=47288%

Totals, subtotals, cash and change are read reliably. Tax and especially service charges are weaker. They are printed in many formats (a rate, an amount, “included”) and appear on fewer receipts.

The checks have limits too

Receipts are messy, and so is ground truth. Measured on the correct labels themselves, the checks hold 91–94% of the time:

How often each check holds on the correct (ground-truth) data. A check that fails on correct data sends good extractions to review.
CheckApplies toHolds on ground truth
cash − change = total106 receipts94%
sum of line prices = subtotal (or total if no subtotal)192 receipts92%
subtotal + tax + service − |discount| = total129 receipts91%

So even a perfect extractor would only be auto-accepted on about 82% of these receipts with these rules. The rest would go to review needlessly. Loosening a rule raises the automation rate and lets more errors through. That trade-off should be chosen per field and per business, not hidden inside a model.

Cost and speed

Median 3.4 seconds per receipt (95th percentile 7.2 s) on one RTX 4070, processing one receipt at a time, with no data leaving the machine. Batching or a smaller image size would raise throughput. A larger model would likely improve item names at the cost of speed.

Method and environment
model
Qwen/Qwen3-VL-4B-Instruct
decoding
greedy, max_new_tokens 1024, bf16, image longest side ≤ 1280 px
prompt
Extract the data from this receipt as JSON with exactly these keys: { "items": [{"name": string, "quantity": number or null, "price": string}], "subtotal": string or null, "tax": string or null, "service": string or null, "discount": string or null, "total": string or null, "cash": string or null, "change": string or null } Rules: one entry in "items" per purchased line; "price" is the line amount as printed. Copy amounts exactly as printed. Use null for anything not printed on the receipt. Return only the JSON.
normalisation
amounts reduced to digits (sign kept); item match = equal price and (name similarity ≥ 0.8 or one normalised name contains the other)
rules
{"items_sum":"sum of line prices = subtotal (or total if no subtotal)","subtotal_to_total":"subtotal + tax + service − |discount| = total","cash_change":"cash − change = total"}
routing
auto-accept only if JSON parsed, a total exists and every applicable rule passes
environment
gpu: NVIDIA GeForce RTX 4070 · cpu: 13th Gen Intel(R) Core(TM) i5-13400F · python: 3.12.14 · torch: 2.14.1+cu130 · transformers: 5.19.0
measured
2026-10-06

Limits

One model, one prompt, one document type from one country, and 200 receipts. CORD labels contain their own mistakes, and “fully correct” ignores extra fields the model filled in where the receipt printed none (115 across 200 receipts, mostly a subtotal equal to the total, or a tax of zero). The code is a standalone package (labs/document-intelligence in this site’s repository) and runs end to end with one command.