Invoice Extraction in Four Languages, with Plain-Code Checks
A local vision model read 400 invoices in Lithuanian, German, Polish and English, and five plain-code checks decided which ones a person must see. 95.5% went straight through correctly. The 1.8% that slipped past the checks show where to add the next one.
- Can a local model enter our supplier invoices in four languages, and can simple checks make sure the wrong ones reach a person instead of the ledger?
- 95.5% of invoices were extracted fully correctly and passed every check. 1.8% were wrong but passed every check; most of these were German invoices where the model read the thousands dot (1.601,08) as a decimal point, shrinking every amount about 1,000 times.
- Automated entry works, but checks that only test whether the numbers add up cannot catch a mistake that is made consistently. Test on invoices in each language you receive, and add a check for each failure you find.
Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.
Supplier invoices are one of the most common automation requests, and in Lithuania they arrive in several languages. The useful question is not only “how accurate is the model” but “how often does a wrong invoice get into the books without anyone noticing?” This Lab measures exactly that.
The test
- Invoices: 400 synthetic B2B invoices, 100 each in Lithuanian, German, Polish and English (Irish VAT), with each country’s VAT rates, number format, VAT ID format and valid IBANs. Three layouts; half rendered clean, half like a scan. All companies are fictional, so the exact correct answer for every field is known. The generator is published with the Lab.
- Extraction: Qwen3-VL-8B, a vision language model, running locally on one RTX 4070. Its output is forced into a fixed JSON format. The prompt was written before the run and not changed afterwards.
- Five plain-code checks: the lines add up to the net total; VAT per line adds up to the VAT total; net plus VAT equals the total; the IBAN check digits are valid; the VAT IDs have the right national format.
Every invoice ends in one of four outcomes: straight-through (all checks pass and everything is right), caught (a check fails and something is wrong), silent error (all checks pass but something is wrong) and false reject (a check fails although everything is right). “Everything” means the nine header fields and the quantity, price, net amount and VAT rate of every line; item descriptions were not scored.
Results
| Straight-through | Caught | Silent error | False reject | |
|---|---|---|---|---|
| All | 95.5% | 2.8% | 1.8% | 0.0% |
| Lithuanian | 99.0% | 1.0% | 0.0% | 0.0% |
| German | 87.0% | 7.0% | 6.0% | 0.0% |
| Polish | 99.0% | 0.0% | 1.0% | 0.0% |
| English (Irish VAT) | 97.0% | 3.0% | 0.0% | 0.0% |
| Clean images | 95.0% | 2.5% | 2.5% | 0.0% |
| Scan-like images | 96.0% | 3.0% | 1.0% | 0.0% |
Lithuanian
- straight-through99.0%
- caught1.0%
- silent error0.0%
- false reject0.0%
German
- straight-through87.0%
- caught7.0%
- silent error6.0%
- false reject0.0%
Polish
- straight-through99.0%
- caught0.0%
- silent error1.0%
- false reject0.0%
English (Irish VAT)
- straight-through97.0%
- caught3.0%
- silent error0.0%
- false reject0.0%
Lithuanian, Polish and English invoices were close to perfect. German invoices were the exception, and the reason is one specific failure.
The failure: a dot that means “thousand”
German writes 1.601,08 where English writes 1,601.08. On 12 of the 92 German invoices with four-digit prices, the model read the dot as a decimal point and every amount on the invoice came out about 1,000 times too small. Because the mistake was made consistently, the totals still added up. Six of these were caught only because the shortened amounts no longer added up to the cent; the other six passed every check:
- line 1: 25 × 1.601,08 = 40.027,00 €; total 69.395,44 €
- line 1: quantity 25, unit_price 1.60, net 40.02; total 69.39
- all five passed
No other language had this problem: Lithuanian and Polish use a space for thousands, and Irish invoices use a comma for thousands and a dot for decimals, the decimal format the prompt asked for.
The one remaining silent error is a different kind: a digit dropped from a unit price while the line total was read correctly. A sixth check, that quantity times unit price equals the line amount, would have caught it:
- line 1: 10 × 2 226,93 = 22 269,30 zł; total 37 140,35 zł
- line 1: quantity 10, unit_price 226.93, net 22269.30; total 37140.35
- all five passed
Field by field, and what the checks caught
| Correct | |
|---|---|
| invoice number | 100.0% |
| issue date | 100.0% |
| due date | 100.0% |
| seller vat id | 99.8% |
| buyer vat id | 100.0% |
| iban | 99.0% |
| net total | 97.0% |
| vat total | 97.0% |
| total | 97.0% |
| line items (recall) | 96.1% |
| Wrong invoices caught | |
|---|---|
| lines sum | 6 |
| vat by rate | 3 |
| net plus vat | 0 |
| iban checksum | 4 |
| vat id format | 1 |
The IBAN checksum caught every misread IBAN, usually a dropped digit. There were no false rejects: the checks never sent a correct invoice to a person.
What this means in practice
- Silent errors are the number to manage, not accuracy. Here, 95.5% straight-through with no human touch, 2.8% correctly sent to a person, and 1.8% wrong without warning.
- Consistency checks miss consistent mistakes. A number-format misread keeps the arithmetic intact. Add checks that compare with something outside the invoice: the purchase order, the expected price range for that supplier, or the payment amount.
- Test every language and format you actually receive. The model was near-perfect in three languages and had a systematic weakness in the fourth. An average over all invoices would have hidden it.
- Each failure found is a check added. Both silent-error patterns here have a simple, plain-code fix.
On one GPU, extraction took about 6.5 seconds per invoice, roughly 550 invoices an hour, with no data leaving the machine.
Method and environment
- data
400 synthetic B2B invoices, 100 each in Lithuanian, German, Polish and English (Irish VAT), half rendered clean and half scan-like; fictional companies; exact ground truth; generator in the repository- model
Qwen3-VL-8B Q4_K_M (llama.cpp), output constrained to a JSON schema, greedy- prompt
Extract this invoice as JSON. Copy text values exactly as printed. Write amounts as plain decimals with a dot and no thousands separators (for example 1234.56), dates as YYYY-MM-DD, the VAT rate of each line as a number (for example 21), and the IBAN and VAT IDs without spaces. Include every line item.- validators
line nets sum to net total; VAT per line and rate sums to VAT total (±0.02); net + VAT = total; IBAN mod-97; VAT ID format per country- outcomes
straight-through = all checks pass and every field right; caught = a check fails and something is wrong; silent error = all checks pass but something is wrong; false reject = a check fails but everything is right- environment
gpu: RTX 4070 12 GB- measured
- 2026-10-09
Limits
Synthetic invoices: clean, consistent layouts with no stamps, handwriting, multi-page documents or supplier-specific quirks, so real invoices will produce more errors, and different ones. One model, one prompt and one run. The JSON format requires amounts with at most two decimals; whether this contributed to the German misreads (by cutting “1.601” to “1.60”) was not tested separately. The checks were designed before the run; the sixth check suggested above was not tested. The code and the invoice generator are in labs/invoice-extraction in this site’s repository.
New Lab results by email
A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.