LAB/006 · Running · Self-directed, not client work

Lithuanian AI Readiness

The same 900 questions in English and Lithuanian, put to seven local language models and three search models. How much is lost in Lithuanian, and which choices make the loss small.

Most AI models are trained mostly on English. Lithuanian is a small language with complex grammar: seven cases and rich word endings. Before building a document system for a Lithuanian organisation, the obvious question is how much worse does it work in Lithuanian, and what can be done about it? This Lab measures that on equal terms: the same questions, on the same texts, in both languages, with models that run privately on one consumer GPU (the setup from LAB/005).

The test material

Belebele is a public benchmark from Meta AI (CC BY-SA 4.0): 900 multiple-choice reading-comprehension questions on 488 short passages, professionally translated into over a hundred languages. Because the English and Lithuanian versions are translations of each other, any difference in results comes from the language, not from the content.

Reading comprehension

Each model read a passage, a question and four possible answers, and replied with a letter. The instruction was in English in both cases; only the language of the material changed.

Accuracy in English and Lithuanian, same questionsScale 0–100% · chance 25%

Qwen3-4B Instruct-2507 Q4_K_M

  1. English91.1%
  2. Lithuanian78.6%

Qwen3-4B Instruct-2507 Q8_0

  1. English91.3%
  2. Lithuanian79.1%

Qwen3-8B Q4_K_M

  1. English84.4%
  2. Lithuanian79.2%

Qwen3-8B Q8_0

  1. English86.0%
  2. Lithuanian78.9%

Qwen3-14B Q4_K_M

  1. English94.7%
  2. Lithuanian87.7%

Phi-4-mini Q4_K_M

  1. English88.8%
  2. Lithuanian62.2%

Phi-4-mini Q8_0

  1. English89.2%
  2. Lithuanian64.6%
Reading comprehension, Belebele: 900 four-option questions on the same passages in English and Lithuanian. Chance is 25%.
EnglishLithuanianGap (points)
Qwen3-4B Instruct-2507 Q4_K_M91.1%78.6%12.5
Qwen3-4B Instruct-2507 Q8_091.3%79.1%12.2
Qwen3-8B Q4_K_M84.4%79.2%5.2
Qwen3-8B Q8_086.0%78.9%7.1
Qwen3-14B Q4_K_M94.7%87.7%7.0
Phi-4-mini Q4_K_M88.8%62.2%26.6
Phi-4-mini Q8_089.2%64.6%24.6
LAB/006 · measured 2026-10-07 · RTX 4070 12 GB

Every model is worse in Lithuanian, but by very different amounts. The Qwen3 models lose 5–13 points; Microsoft’s Phi-4-mini loses about 25 and gets roughly one question in three wrong. The best Lithuanian result, Qwen3-14B at 88%, is about where the mid-sized models are in English.

Two details worth knowing:

  • Qwen3-4B Instruct-2507 does better in English than the larger Qwen3-8B. It is a newer release (July 2025) trained further for instruction following. In Lithuanian, the two are level.
  • 4-bit and 8-bit versions of the same model differ by at most 2.4 points: within the noise for 900 questions. Compression is not where the Lithuanian loss comes from.

Search: finding the right passage

A document assistant first has to find the right text. Here each question had to find the passage it was written for, among all 488, using an embedding model (a model that turns text into vectors for search).

Lithuanian questions, Lithuanian passages: right passage ranked firstRecall@1 · scale 0–100%
  1. bge-m3 (multilingual)87.0%
  2. multilingual-e5-base (multilingual)83.4%
  3. bge-base-en (English-only)45.0%
Search: find the passage a question was written for among 488 passages. Recall@1 = the right passage ranked first.
EN→ENLT→LTLT→ENEN→LT
bge-m3 (multilingual)91.3%87.0%83.1%84.8%
multilingual-e5-base (multilingual)92.1%83.4%65.8%77.9%
bge-base-en (English-only)89.7%45.0%17.6%28.9%
Recall@1; LT→EN = Lithuanian question, English passages · LAB/006 · measured 2026-10-07 · RTX 4070 12 GB

This is the clearest result in the Lab. An English-only search model fails on Lithuanian: it ranks the right passage first less than half the time, against 90% in English, and nothing in the system reports that anything is wrong. A multilingual model (bge-m3, MIT licence, small enough to run beside the language model) loses only 4 points.

The cross-language columns matter for real organisations, where Lithuanian staff often search English documents, or the reverse. bge-m3 handles this well (83–85%). multilingual-e5-base, also multilingual, is clearly weaker at it (66% from Lithuanian questions to English passages).

With the five best matches instead of one, bge-m3 finds the right Lithuanian passage 97% of the time:

Search: find the passage a question was written for among 488 passages. Recall@5 = the right passage in the top five.
EN→ENLT→LTLT→ENEN→LT
bge-m3 (multilingual)98.3%96.8%94.9%95.6%
multilingual-e5-base (multilingual)98.0%94.3%81.7%91.0%
bge-base-en (English-only)96.7%60.3%28.6%44.6%
LAB/006 · measured 2026-10-07 · RTX 4070 12 GB

What this means in practice

  • Lithuanian works with local models, with a measurable cost. Expect several points lower accuracy than the same system in English, and test with Lithuanian material from the start, not translated afterwards.
  • The search model matters more than the language model. An English-only embedding model is the most common and most damaging mistake: it fails quietly. Use a multilingual one such as bge-m3.
  • Pick the language model on Lithuanian results, not English ones. Phi-4-mini is competitive in English and the weakest in Lithuanian by a wide margin.
  • Larger helps. Qwen3-14B, the largest that fits a 12 GB card, was clearly best in Lithuanian.
Method and environment
dataset
Belebele (Bandarkar et al., 2023), lit_Latn and eng_Latn, 900 parallel questions on 488 FLORES-200 passages; CC BY-SA 4.0
reading
zero-shot, one user message, greedy, answer letter parsed; instruction in English for both languages: Read the passage and answer the multiple-choice question. Reply with a single letter (A, B, C or D) and nothing else.
models
llama.cpp b11461 (CUDA 13.4), same GGUF files as LAB/005; Qwen3-8B/14B with thinking disabled
retrieval
exact cosine search, 900 questions against the 488 unique passages, per direction; models: bge-m3 (BAAI/bge-m3), multilingual-e5-base (intfloat/multilingual-e5-base), bge-base-en (BAAI/bge-base-en-v1.5)
environment
gpu: RTX 4070 12 GB · llama.cpp: b11461 (CUDA 13.4)
measured
2026-10-07

Limits

Belebele passages are short, edited, general-interest texts translated from English. Real Lithuanian documents (legal texts, municipal decisions, invoices, emails in everyday language) will be harder, and terminology is not tested here. Multiple choice measures understanding, not how well a model writes Lithuanian, which deserves its own test. Search was over 488 passages; larger collections make every model’s job harder. The code is a standalone package (labs/lithuanian-ai in this site’s repository) and reruns with three commands.