LAB/014 · Running · Self-directed, not client work

Invoice Extraction in Four Languages, with Plain-Code Checks

A local vision model read 400 invoices in Lithuanian, German, Polish and English, and five plain-code checks decided which ones a person must see. 95.5% went straight through correctly. The 1.8% that slipped past the checks show where to add the next one.

The question
Can a local model enter our supplier invoices in four languages, and can simple checks make sure the wrong ones reach a person instead of the ledger?
What it showed
95.5% of invoices were extracted fully correctly and passed every check. 1.8% were wrong but passed every check; most of these were German invoices where the model read the thousands dot (1.601,08) as a decimal point, shrinking every amount about 1,000 times.
What it means for you
Automated entry works, but checks that only test whether the numbers add up cannot catch a mistake that is made consistently. Test on invoices in each language you receive, and add a check for each failure you find.

Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.

Supplier invoices are one of the most common automation requests, and in Lithuania they arrive in several languages. The useful question is not only “how accurate is the model” but “how often does a wrong invoice get into the books without anyone noticing?” This Lab measures exactly that.

The test

  • Invoices: 400 synthetic B2B invoices, 100 each in Lithuanian, German, Polish and English (Irish VAT), with each country’s VAT rates, number format, VAT ID format and valid IBANs. Three layouts; half rendered clean, half like a scan. All companies are fictional, so the exact correct answer for every field is known. The generator is published with the Lab.
  • Extraction: Qwen3-VL-8B, a vision language model, running locally on one RTX 4070. Its output is forced into a fixed JSON format. The prompt was written before the run and not changed afterwards.
  • Five plain-code checks: the lines add up to the net total; VAT per line adds up to the VAT total; net plus VAT equals the total; the IBAN check digits are valid; the VAT IDs have the right national format.

Every invoice ends in one of four outcomes: straight-through (all checks pass and everything is right), caught (a check fails and something is wrong), silent error (all checks pass but something is wrong) and false reject (a check fails although everything is right). “Everything” means the nine header fields and the quantity, price, net amount and VAT rate of every line; item descriptions were not scored.

Results

What happens to each invoice: processed with no human touch, sent to a person, or wrong without anyone knowing.
Straight-throughCaughtSilent errorFalse reject
All95.5%2.8%1.8%0.0%
Lithuanian99.0%1.0%0.0%0.0%
German87.0%7.0%6.0%0.0%
Polish99.0%0.0%1.0%0.0%
English (Irish VAT)97.0%3.0%0.0%0.0%
Clean images95.0%2.5%2.5%0.0%
Scan-like images96.0%3.0%1.0%0.0%
LAB/014 · measured 2026-10-09 · RTX 4070 12 GB
Outcome per languageShare of invoices · silent errors are the ones to drive to zero

Lithuanian

  1. straight-through99.0%
  2. caught1.0%
  3. silent error0.0%
  4. false reject0.0%

German

  1. straight-through87.0%
  2. caught7.0%
  3. silent error6.0%
  4. false reject0.0%

Polish

  1. straight-through99.0%
  2. caught0.0%
  3. silent error1.0%
  4. false reject0.0%

English (Irish VAT)

  1. straight-through97.0%
  2. caught3.0%
  3. silent error0.0%
  4. false reject0.0%

Lithuanian, Polish and English invoices were close to perfect. German invoices were the exception, and the reason is one specific failure.

The failure: a dot that means “thousand”

German writes 1.601,08 where English writes 1,601.08. On 12 of the 92 German invoices with four-digit prices, the model read the dot as a decimal point and every amount on the invoice came out about 1,000 times too small. Because the mistake was made consistently, the totals still added up. Six of these were caught only because the shortened amounts no longer added up to the cent; the other six passed every check:

German thousands dot read as a decimal point: every check passedWrong
Printed on the invoice
line 1: 25 × 1.601,08 = 40.027,00 €; total 69.395,44 €
Qwen3-VL-8B extracted
line 1: quantity 25, unit_price 1.60, net 40.02; total 69.39
Checks
all five passed

Synthetic invoice de-0066 (fictional, generated for this Lab).

No other language had this problem: Lithuanian and Polish use a space for thousands, and Irish invoices use a comma for thousands and a dot for decimals, the decimal format the prompt asked for.

The one remaining silent error is a different kind: a digit dropped from a unit price while the line total was read correctly. A sixth check, that quantity times unit price equals the line amount, would have caught it:

One digit dropped from a unit price: every check passedWrong
Printed on the invoice
line 1: 10 × 2 226,93 = 22 269,30 zł; total 37 140,35 zł
Qwen3-VL-8B extracted
line 1: quantity 10, unit_price 226.93, net 22269.30; total 37140.35
Checks
all five passed

Synthetic invoice pl-0070 (fictional, generated for this Lab).

Field by field, and what the checks caught

Field accuracy, all 400 invoices (exact match after removing spaces).
Correct
invoice number100.0%
issue date100.0%
due date100.0%
seller vat id99.8%
buyer vat id100.0%
iban99.0%
net total97.0%
vat total97.0%
total97.0%
line items (recall)96.1%
LAB/014 · measured 2026-10-09 · RTX 4070 12 GB
Which plain-code check caught the wrong invoices (one invoice can fail several).
Wrong invoices caught
lines sum6
vat by rate3
net plus vat0
iban checksum4
vat id format1
LAB/014 · measured 2026-10-09 · RTX 4070 12 GB

The IBAN checksum caught every misread IBAN, usually a dropped digit. There were no false rejects: the checks never sent a correct invoice to a person.

What this means in practice

  • Silent errors are the number to manage, not accuracy. Here, 95.5% straight-through with no human touch, 2.8% correctly sent to a person, and 1.8% wrong without warning.
  • Consistency checks miss consistent mistakes. A number-format misread keeps the arithmetic intact. Add checks that compare with something outside the invoice: the purchase order, the expected price range for that supplier, or the payment amount.
  • Test every language and format you actually receive. The model was near-perfect in three languages and had a systematic weakness in the fourth. An average over all invoices would have hidden it.
  • Each failure found is a check added. Both silent-error patterns here have a simple, plain-code fix.

On one GPU, extraction took about 6.5 seconds per invoice, roughly 550 invoices an hour, with no data leaving the machine.

Method and environment
data
400 synthetic B2B invoices, 100 each in Lithuanian, German, Polish and English (Irish VAT), half rendered clean and half scan-like; fictional companies; exact ground truth; generator in the repository
model
Qwen3-VL-8B Q4_K_M (llama.cpp), output constrained to a JSON schema, greedy
prompt
Extract this invoice as JSON. Copy text values exactly as printed. Write amounts as plain decimals with a dot and no thousands separators (for example 1234.56), dates as YYYY-MM-DD, the VAT rate of each line as a number (for example 21), and the IBAN and VAT IDs without spaces. Include every line item.
validators
line nets sum to net total; VAT per line and rate sums to VAT total (±0.02); net + VAT = total; IBAN mod-97; VAT ID format per country
outcomes
straight-through = all checks pass and every field right; caught = a check fails and something is wrong; silent error = all checks pass but something is wrong; false reject = a check fails but everything is right
environment
gpu: RTX 4070 12 GB
measured
2026-10-09

Limits

Synthetic invoices: clean, consistent layouts with no stamps, handwriting, multi-page documents or supplier-specific quirks, so real invoices will produce more errors, and different ones. One model, one prompt and one run. The JSON format requires amounts with at most two decimals; whether this contributed to the German misreads (by cutting “1.601” to “1.60”) was not tested separately. The checks were designed before the run; the sixth check suggested above was not tested. The code and the invoice generator are in labs/invoice-extraction in this site’s repository.

New Lab results by email

A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.