LAB/011 · Running · Self-directed, not client work

Local OCR on Scans and Photos, Lithuanian vs English

Classic OCR and two vision language models, all running locally, read the same pages in Lithuanian and English, plus 248 real English documents. No tool won everywhere, and the losers fail in very different ways.

The question
Which local tool turns our scans and photos into usable text, how well does it handle Lithuanian letters, and what does it cost per 1,000 pages?
What it showed
On clean scans, Tesseract with the Lithuanian pack was best (0.08% character errors, no diacritic errors) and 6× faster. On phone photos it fell apart (82% errors) while Qwen3-VL-8B stayed under 1%. Vision models got 1.6–2.6% of Lithuanian diacritics wrong even on clean pages.
What it means for you
Route documents by quality: classic OCR for clean scans, a vision model for photos and bad scans, and always the right language pack. Measure on your own documents, because the right tool depends on what you receive.

Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.

Every document project starts with the same step: turning images of pages into text. It is also where quality is decided, because nothing downstream can recover a word the OCR got wrong. This Lab measures three local tools on the documents a Lithuanian organisation actually receives: clean PDFs, office scans, poor scans and phone photos, in Lithuanian and in English, on identical content.

The test

  • Pages: 60 page pairs built from professionally translated passages (Belebele, CC BY-SA 4.0), so each Lithuanian page has an English twin with the same content. Each was rendered in four conditions, from a clean 300 dpi page down to a ~120 dpi phone photo with a shadow across it: 480 pages, with the exact text known.
  • Tools, all local: Tesseract 5, the classic open-source OCR engine, with the right language pack and, as a common mistake, with English only; Qwen3-VL-8B, a general vision language model; and olmOCR-2-7B, a vision model tuned specifically for reading documents. The vision models ran on one RTX 4070.
  • Scores: character error rate, and, for Lithuanian, the share of letters with diacritics (ą č ę ė į š ų ū ž) read wrongly.

The “phone photo” condition was added after a first Tesseract test showed the poor scans were still too easy to separate the tools. It was fixed before any vision model was run.

Lithuanian pages

Character error rate (%), Lithuanian pages. Lower is better. Same passages and same image degradation in both languages.
clean, 300 dpioffice scan, 200 dpipoor scan, 150 dpiphone photo with shadow
Tesseract 5, right language pack0.080.192.8181.78
Tesseract 5, English pack only6.736.749.4684.62
Qwen3-VL-8B (vision LLM)0.420.400.450.62
olmOCR-2-7B (OCR-tuned vision LLM)0.480.490.451.94
LAB/011 · measured 2026-10-08 · RTX 4070 12 GB
Lithuanian letters with diacritics (ą č ę ė į š ų ū ž) read wrongly (%).
clean, 300 dpioffice scan, 200 dpipoor scan, 150 dpiphone photo with shadow
Tesseract 5, right language pack0.00.22.782.3
Tesseract 5, English pack only100.0100.0100.0100.0
Qwen3-VL-8B (vision LLM)1.71.61.81.9
olmOCR-2-7B (OCR-tuned vision LLM)2.62.72.26.1
LAB/011 · measured 2026-10-08 · RTX 4070 12 GB

On clean and office scans, classic OCR wins. Tesseract with the Lithuanian pack made almost no errors and got every diacritic right on clean pages. Both vision models misread 1.6–2.6% of diacritics even on perfect pages, and the errors look plausible: ą read as a, į as i or j, ę as ė, ė as é. A person skimming the text would not notice; a search for a name or a legal term would miss it.

On a phone photo with a shadow, classic OCR falls apart. Tesseract read only the unshadowed parts (82% character errors), while Qwen3-VL-8B stayed below 1%:

Lithuanian phone photos: character error rateLower is better · % of characters wrong
  1. Tesseract 5, right language pack81.8%
  2. Tesseract 5, English pack only84.6%
  3. Qwen3-VL-8B (vision LLM)0.6%
  4. olmOCR-2-7B (OCR-tuned vision LLM)1.9%

The wrong language pack is a silent disaster. With only the English pack, Tesseract lost every single diacritic, and nothing in its output signals a problem.

The same pages in English

Character error rate (%), English pages. Lower is better. Same passages and same image degradation in both languages.
clean, 300 dpioffice scan, 200 dpipoor scan, 150 dpiphone photo with shadow
Tesseract 5, right language pack0.090.103.5475.78
Tesseract 5, English pack only0.090.103.5475.78
Qwen3-VL-8B (vision LLM)0.030.020.020.12
olmOCR-2-7B (OCR-tuned vision LLM)0.040.060.090.17
LAB/011 · measured 2026-10-08 · RTX 4070 12 GB

The vision models made more than ten times fewer errors on the English twins of the same pages. Lithuanian also took them almost twice as long per page, because the same text needs more output tokens in Lithuanian.

Real English documents: olmOCR-bench

The rendered pages above have exact ground truth but simple layouts. To test real documents, the same tools read 248 PDFs from olmOCR-bench (Allen Institute for AI, ODC-BY): old typewritten scans, tables, multi-column pages and pages with headers and footers, scored by the benchmark’s own unit tests.

olmOCR-bench subset (248 English PDFs: old scans, tables, multi-column, headers/footers), official scorer: unit tests passed (%).
Allheaders footersmulti columnold scanstables
Tesseract 5, right language pack41.836.460.020.20.0
Qwen3-VL-8B (vision LLM)54.838.378.940.117.6
olmOCR-2-7B (OCR-tuned vision LLM)81.794.885.946.282.1
Subset and our image rendering, so not directly comparable with the published leaderboard · LAB/011 · measured 2026-10-08 · RTX 4070 12 GB

Here the specialist wins clearly. olmOCR-2 passed 81.7% of the tests, close to its published result on the full benchmark, and it was far ahead on tables (82%) and on leaving out page headers and footers (95%). Tesseract produces no table structure at all, so it fails every table test; Qwen3-VL was asked for plain text, which also loses table structure. So the model that is best on English document layout is also the one that is weakest on Lithuanian letters: which tool is “best” depends on your documents and your language.

Speed and cost

Speed and energy on one RTX 4070 (vision LLMs) or 4 CPU threads (Tesseract), one page at a time.
Pages per minuteSeconds per LT pageSeconds per EN pageGPU kWh per 1,000 pages€ per 1,000 pages at €0.20/kWh
Tesseract 5, right language pack57.11.20.9——
Tesseract 5, English pack only66.80.90.9——
Qwen3-VL-8B (vision LLM)9.18.54.70.350.07
olmOCR-2-7B (OCR-tuned vision LLM)9.48.14.70.340.07
Energy is GPU power only (nvidia-smi); the electricity price is an assumption · LAB/011 · measured 2026-10-08 · RTX 4070 12 GB

At about 9 pages a minute, one consumer GPU reads roughly 13,000 pages a day; the electricity for 1,000 pages costs a few cents. Tesseract does 57 pages a minute on four CPU threads, with no GPU at all.

What this means in practice

  • Route by document quality. Clean scans to classic OCR with the right language pack; photos and bad scans to a vision model. A quick quality check decides.
  • Check diacritics explicitly. For Lithuanian, measure diacritic errors on your own documents; overall accuracy hides them.
  • Never run OCR with the wrong language pack. It is the cheapest mistake to make and the hardest to see.
  • Private is affordable. None of these pages left the machine, and the cost is in cents per thousand pages.
Method and environment
pages
60 page pairs from Belebele passages (CC BY-SA 4.0), identical content in Lithuanian and English, rendered in 4 fonts and 4 conditions; 480 pages
conditions
clean (300 dpi); scan (200 dpi, skew ≤0.8°, blur, noise, JPEG 70); poor (150 dpi, skew ≤1.8°, more blur, JPEG 45); phone (~120 dpi, skew ≤3°, blur, a hard-edged shadow, JPEG 35; added after the Tesseract pilot showed 'poor' was too easy, before any vision model ran)
tools
Tesseract 5 (tessdata_best lit/eng, --psm 3); Qwen3-VL-8B and olmOCR-2-7B as Q4_K_M GGUF in llama.cpp, image longest side 1288 px, greedy
scoring
NFC, markdown and front matter stripped, quotes/dashes unified, whitespace collapsed; CER by Levenshtein; diacritic errors by edit alignment
bench
olmOCR-bench (ODC-BY): first page of 248 PDFs, 200 dpi, official scorer (allenai/olmocr)
environment
gpu: RTX 4070 12 GB · cpu: i5-13400F
measured
2026-10-08

Limits

Rendered text pages, not real scans of real documents: no tables, stamps, handwriting or multi-column layouts in the Lithuanian test (the benchmark section covers some of those in English). The benchmark subset uses page one of each PDF and our own rendering, so its scores are not directly comparable with the published leaderboard; the vision models also received different prompts (olmOCR-2 its own, which asks for tables as HTML). The phone condition uses one simulated hard-edged shadow, which may be harsher than many real photos. Two vision models and one OCR engine, with default settings. The code is in labs/ocr-shootout in this site’s repository.

New Lab results by email

A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.