Local OCR on Scans and Photos, Lithuanian vs English
Classic OCR and two vision language models, all running locally, read the same pages in Lithuanian and English, plus 248 real English documents. No tool won everywhere, and the losers fail in very different ways.
- Which local tool turns our scans and photos into usable text, how well does it handle Lithuanian letters, and what does it cost per 1,000 pages?
- On clean scans, Tesseract with the Lithuanian pack was best (0.08% character errors, no diacritic errors) and 6× faster. On phone photos it fell apart (82% errors) while Qwen3-VL-8B stayed under 1%. Vision models got 1.6–2.6% of Lithuanian diacritics wrong even on clean pages.
- Route documents by quality: classic OCR for clean scans, a vision model for photos and bad scans, and always the right language pack. Measure on your own documents, because the right tool depends on what you receive.
Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.
Every document project starts with the same step: turning images of pages into text. It is also where quality is decided, because nothing downstream can recover a word the OCR got wrong. This Lab measures three local tools on the documents a Lithuanian organisation actually receives: clean PDFs, office scans, poor scans and phone photos, in Lithuanian and in English, on identical content.
The test
- Pages: 60 page pairs built from professionally translated passages (Belebele, CC BY-SA 4.0), so each Lithuanian page has an English twin with the same content. Each was rendered in four conditions, from a clean 300 dpi page down to a ~120 dpi phone photo with a shadow across it: 480 pages, with the exact text known.
- Tools, all local: Tesseract 5, the classic open-source OCR engine, with the right language pack and, as a common mistake, with English only; Qwen3-VL-8B, a general vision language model; and olmOCR-2-7B, a vision model tuned specifically for reading documents. The vision models ran on one RTX 4070.
- Scores: character error rate, and, for Lithuanian, the share of letters with diacritics (ą č ę ė į š ų ū ž) read wrongly.
The “phone photo” condition was added after a first Tesseract test showed the poor scans were still too easy to separate the tools. It was fixed before any vision model was run.
Lithuanian pages
| clean, 300 dpi | office scan, 200 dpi | poor scan, 150 dpi | phone photo with shadow | |
|---|---|---|---|---|
| Tesseract 5, right language pack | 0.08 | 0.19 | 2.81 | 81.78 |
| Tesseract 5, English pack only | 6.73 | 6.74 | 9.46 | 84.62 |
| Qwen3-VL-8B (vision LLM) | 0.42 | 0.40 | 0.45 | 0.62 |
| olmOCR-2-7B (OCR-tuned vision LLM) | 0.48 | 0.49 | 0.45 | 1.94 |
| clean, 300 dpi | office scan, 200 dpi | poor scan, 150 dpi | phone photo with shadow | |
|---|---|---|---|---|
| Tesseract 5, right language pack | 0.0 | 0.2 | 2.7 | 82.3 |
| Tesseract 5, English pack only | 100.0 | 100.0 | 100.0 | 100.0 |
| Qwen3-VL-8B (vision LLM) | 1.7 | 1.6 | 1.8 | 1.9 |
| olmOCR-2-7B (OCR-tuned vision LLM) | 2.6 | 2.7 | 2.2 | 6.1 |
On clean and office scans, classic OCR wins. Tesseract with the Lithuanian pack made almost no errors and got every diacritic right on clean pages. Both vision models misread 1.6–2.6% of diacritics even on perfect pages, and the errors look plausible: ą read as a, į as i or j, ę as ė, ė as é. A person skimming the text would not notice; a search for a name or a legal term would miss it.
On a phone photo with a shadow, classic OCR falls apart. Tesseract read only the unshadowed parts (82% character errors), while Qwen3-VL-8B stayed below 1%:
- Tesseract 5, right language pack81.8%
- Tesseract 5, English pack only84.6%
- Qwen3-VL-8B (vision LLM)0.6%
- olmOCR-2-7B (OCR-tuned vision LLM)1.9%
The wrong language pack is a silent disaster. With only the English pack, Tesseract lost every single diacritic, and nothing in its output signals a problem.
The same pages in English
| clean, 300 dpi | office scan, 200 dpi | poor scan, 150 dpi | phone photo with shadow | |
|---|---|---|---|---|
| Tesseract 5, right language pack | 0.09 | 0.10 | 3.54 | 75.78 |
| Tesseract 5, English pack only | 0.09 | 0.10 | 3.54 | 75.78 |
| Qwen3-VL-8B (vision LLM) | 0.03 | 0.02 | 0.02 | 0.12 |
| olmOCR-2-7B (OCR-tuned vision LLM) | 0.04 | 0.06 | 0.09 | 0.17 |
The vision models made more than ten times fewer errors on the English twins of the same pages. Lithuanian also took them almost twice as long per page, because the same text needs more output tokens in Lithuanian.
Real English documents: olmOCR-bench
The rendered pages above have exact ground truth but simple layouts. To test real documents, the same tools read 248 PDFs from olmOCR-bench (Allen Institute for AI, ODC-BY): old typewritten scans, tables, multi-column pages and pages with headers and footers, scored by the benchmark’s own unit tests.
| All | headers footers | multi column | old scans | tables | |
|---|---|---|---|---|---|
| Tesseract 5, right language pack | 41.8 | 36.4 | 60.0 | 20.2 | 0.0 |
| Qwen3-VL-8B (vision LLM) | 54.8 | 38.3 | 78.9 | 40.1 | 17.6 |
| olmOCR-2-7B (OCR-tuned vision LLM) | 81.7 | 94.8 | 85.9 | 46.2 | 82.1 |
Here the specialist wins clearly. olmOCR-2 passed 81.7% of the tests, close to its published result on the full benchmark, and it was far ahead on tables (82%) and on leaving out page headers and footers (95%). Tesseract produces no table structure at all, so it fails every table test; Qwen3-VL was asked for plain text, which also loses table structure. So the model that is best on English document layout is also the one that is weakest on Lithuanian letters: which tool is “best” depends on your documents and your language.
Speed and cost
| Pages per minute | Seconds per LT page | Seconds per EN page | GPU kWh per 1,000 pages | € per 1,000 pages at €0.20/kWh | |
|---|---|---|---|---|---|
| Tesseract 5, right language pack | 57.1 | 1.2 | 0.9 | — | — |
| Tesseract 5, English pack only | 66.8 | 0.9 | 0.9 | — | — |
| Qwen3-VL-8B (vision LLM) | 9.1 | 8.5 | 4.7 | 0.35 | 0.07 |
| olmOCR-2-7B (OCR-tuned vision LLM) | 9.4 | 8.1 | 4.7 | 0.34 | 0.07 |
At about 9 pages a minute, one consumer GPU reads roughly 13,000 pages a day; the electricity for 1,000 pages costs a few cents. Tesseract does 57 pages a minute on four CPU threads, with no GPU at all.
What this means in practice
- Route by document quality. Clean scans to classic OCR with the right language pack; photos and bad scans to a vision model. A quick quality check decides.
- Check diacritics explicitly. For Lithuanian, measure diacritic errors on your own documents; overall accuracy hides them.
- Never run OCR with the wrong language pack. It is the cheapest mistake to make and the hardest to see.
- Private is affordable. None of these pages left the machine, and the cost is in cents per thousand pages.
Method and environment
- pages
60 page pairs from Belebele passages (CC BY-SA 4.0), identical content in Lithuanian and English, rendered in 4 fonts and 4 conditions; 480 pages- conditions
clean (300 dpi); scan (200 dpi, skew ≤0.8°, blur, noise, JPEG 70); poor (150 dpi, skew ≤1.8°, more blur, JPEG 45); phone (~120 dpi, skew ≤3°, blur, a hard-edged shadow, JPEG 35; added after the Tesseract pilot showed 'poor' was too easy, before any vision model ran)- tools
Tesseract 5 (tessdata_best lit/eng, --psm 3); Qwen3-VL-8B and olmOCR-2-7B as Q4_K_M GGUF in llama.cpp, image longest side 1288 px, greedy- scoring
NFC, markdown and front matter stripped, quotes/dashes unified, whitespace collapsed; CER by Levenshtein; diacritic errors by edit alignment- bench
olmOCR-bench (ODC-BY): first page of 248 PDFs, 200 dpi, official scorer (allenai/olmocr)- environment
gpu: RTX 4070 12 GB · cpu: i5-13400F- measured
- 2026-10-08
Limits
Rendered text pages, not real scans of real documents: no tables, stamps, handwriting or multi-column layouts in the Lithuanian test (the benchmark section covers some of those in English). The benchmark subset uses page one of each PDF and our own rendering, so its scores are not directly comparable with the published leaderboard; the vision models also received different prompts (olmOCR-2 its own, which asks for tables as HTML). The phone condition uses one simulated hard-edged shadow, which may be harsher than many real photos. Two vision models and one OCR engine, with default settings. The code is in labs/ocr-shootout in this site’s repository.
New Lab results by email
A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.