Document Intelligence Engine
Receipt extraction with an open-weight vision-language model, checked by plain-code arithmetic and routed to a person when the checks fail. Measured on 200 public receipts.
Average extraction accuracy is the wrong number for document automation. What matters is how much of the work can be handed to the system without a person, and how often the system is wrong on exactly that part. This Lab measures both.
The pipeline
- Extract: an open-weight vision-language model (Qwen3-VL-4B, running locally on one consumer GPU) reads the receipt image and fills a fixed JSON schema: line items, subtotal, tax, service, discount, total, cash and change.
- Validate: plain code checks the arithmetic that every receipt must satisfy. Line prices add up to the subtotal; subtotal plus tax and service minus discount equals the total; cash minus change equals the total.
- Route: a receipt is accepted automatically only if every applicable check passes. Everything else goes to a person.
No second model, no confidence scores from the model itself. Just arithmetic.
The data is CORD-v2: real shop and restaurant receipts from Indonesia, photographed and labelled, published under CC BY 4.0. 200 receipts from the validation and test splits were used. The prompt was written before the run and not tuned on these receipts.
Results
| All receipts | Auto-accepted | Sent to review | |
|---|---|---|---|
| Receipts | 200 | 145 | 55 |
| Fully correct | 65.5% | 80.7% | 25.5% |
| Total correct | 97.4% | 100.0% | 90.6% |
| Line items F1 | 0.88 | 0.94 | 0.76 |
72.5% of receipts were accepted automatically, and every one of them had the correct total. Of the 5 receipts where the model got the total wrong, all 5 were sent to review. The review queue is where the problems concentrate: only a quarter of the receipts sent there were fully correct, against 81% of the auto-accepted ones.
The checks are not a guarantee. 28 auto-accepted receipts still had some error, and 23 of those were line-item problems: a misread item name, or a line split or merged while the prices still added up. Arithmetic cannot see a wrong name. If item names matter downstream, they need their own check, for example against a product catalogue.
A receipt the checks caught

| Field | Model output | Ground truth |
|---|---|---|
| subtotal | 18,181 | 18,181 |
| tax | 10% Tax Included | 1,818 |
| discount | 50% | — |
| total | 20,000 | 20,000 |
| cash | 100,000 | 100,000 |
| change | 80,000 | 80,000 |
| items | S-Ovaltine: 20,000 | |
- Failed: sum of line prices = subtotal (or total if no subtotal)
- Failed: subtotal + tax + service − |discount| = total
- Passed: cash − change = total
The model copied the printed text “10% Tax Included” into the tax field instead of the tax amount, and read “50%” from the item name (“S-Ovaltine 50%” is a sugar level) as a discount. Both arithmetic checks failed, so the receipt went to a person. Nothing about the output looked wrong at a glance. It was well-formed JSON with plausible values.
A receipt that went straight through

| Field | Model output | Ground truth |
|---|---|---|
| subtotal | 45,500 | — |
| total | 45,500 | 45,500 |
| cash | 50,000 | 50,000 |
| change | 4,500 | 4,500 |
| items | REAL GANACHE: 16,500EGG TART: 13,000PIZZA TOAST: 16,000 | |
- Passed: sum of line prices = subtotal (or total if no subtotal)
- Passed: subtotal + tax + service − |discount| = total
- Passed: cash − change = total
Field by field
- total n=19397%
- subtotal n=13299%
- cash n=13196%
- change n=12095%
- tax n=8883%
- service n=2564%
- discount n=1292%
- line items (F1) n=47288%
Totals, subtotals, cash and change are read reliably. Tax and especially service charges are weaker. They are printed in many formats (a rate, an amount, “included”) and appear on fewer receipts.
The checks have limits too
Receipts are messy, and so is ground truth. Measured on the correct labels themselves, the checks hold 91–94% of the time:
| Check | Applies to | Holds on ground truth |
|---|---|---|
| cash − change = total | 106 receipts | 94% |
| sum of line prices = subtotal (or total if no subtotal) | 192 receipts | 92% |
| subtotal + tax + service − |discount| = total | 129 receipts | 91% |
So even a perfect extractor would only be auto-accepted on about 82% of these receipts with these rules. The rest would go to review needlessly. Loosening a rule raises the automation rate and lets more errors through. That trade-off should be chosen per field and per business, not hidden inside a model.
Cost and speed
Median 3.4 seconds per receipt (95th percentile 7.2 s) on one RTX 4070, processing one receipt at a time, with no data leaving the machine. Batching or a smaller image size would raise throughput. A larger model would likely improve item names at the cost of speed.
Method and environment
- model
Qwen/Qwen3-VL-4B-Instruct- decoding
greedy, max_new_tokens 1024, bf16, image longest side ≤ 1280 px- prompt
Extract the data from this receipt as JSON with exactly these keys: { "items": [{"name": string, "quantity": number or null, "price": string}], "subtotal": string or null, "tax": string or null, "service": string or null, "discount": string or null, "total": string or null, "cash": string or null, "change": string or null } Rules: one entry in "items" per purchased line; "price" is the line amount as printed. Copy amounts exactly as printed. Use null for anything not printed on the receipt. Return only the JSON.- normalisation
amounts reduced to digits (sign kept); item match = equal price and (name similarity ≥ 0.8 or one normalised name contains the other)- rules
{"items_sum":"sum of line prices = subtotal (or total if no subtotal)","subtotal_to_total":"subtotal + tax + service − |discount| = total","cash_change":"cash − change = total"}- routing
auto-accept only if JSON parsed, a total exists and every applicable rule passes- environment
gpu: NVIDIA GeForce RTX 4070 · cpu: 13th Gen Intel(R) Core(TM) i5-13400F · python: 3.12.14 · torch: 2.14.1+cu130 · transformers: 5.19.0- measured
- 2026-10-06
Limits
One model, one prompt, one document type from one country, and 200 receipts. CORD labels contain their own mistakes, and “fully correct” ignores extra fields the model filled in where the receipt printed none (115 across 200 receipts, mostly a subtotal equal to the total, or a tax of zero). The code is a standalone package (labs/document-intelligence in this site’s repository) and runs end to end with one command.