Offline Multilingual Triage (LT / EN / RU)
Sorting and tagging Lithuanian, English and Russian text on one consumer GPU, with nothing leaving the machine. A small classifier trained on English examples sorted all three languages as well as or better than the language models; local LLMs extracted names and places at similar quality in each language.
Analysts in the Baltic region read in at least three languages. Triage, deciding what a text is about and who and what it mentions, is the first step before anyone reads it closely. When the material is sensitive, that step has to run on hardware the organisation controls. This Lab measures how well it works offline, on one consumer GPU (RTX 4070), in Lithuanian, English and Russian, using public data only.
The test
- Topic triage: SIB-200 (CC BY-SA 4.0), 204 test sentences in 7 topics (science and technology, travel, politics, sports, health, entertainment, geography). The English, Lithuanian and Russian versions are translations of the same sentences, so the languages are compared on identical content.
- Entity extraction: WikiANN, 300 Wikipedia sentences per language with people, organisations and locations marked. These labels were generated automatically and contain errors, so absolute scores understate quality; the comparison across languages and models is what counts.
Two approaches:
- Local language models (the four from LAB/005), asked directly, with output forced into a fixed JSON format (the lesson from LAB/008).
- A small classifier: a multilingual embedding model plus logistic regression, trained on the 701 labelled English training sentences only, then applied to all three languages.
Topic triage
| English | Lithuanian | Russian | |
|---|---|---|---|
| Qwen3-4B Instruct-2507 Q4_K_M · zero-shot | 82.4% | 81.4% | 84.3% |
| Qwen3-8B Q4_K_M · zero-shot | 82.4% | 76.0% | 84.8% |
| Qwen3-14B Q4_K_M · zero-shot | 86.8% | 86.3% | 87.7% |
| Phi-4-mini Q4_K_M · zero-shot | 85.3% | 78.4% | 84.3% |
| bge-m3 + logistic regression · trained on English only | 89.2% | 86.8% | 89.2% |
| multilingual-e5-base + logistic regression · trained on English only | 89.7% | 83.8% | 87.7% |
| bge-base-en + logistic regression · trained on English only | 89.2% | 43.1% | 28.9% |
The small classifier won, in every language. Trained on English examples alone, bge-m3 with a simple classifier sorted Lithuanian and Russian sentences almost as well as English (87–89%), as well as or better than any of the language models asked directly (76–88%). It is also a much smaller and simpler component to run than a language model.
That needs a multilingual embedding model. The English-only model scores the same on English and collapses on the others (43% Lithuanian, 29% Russian, against 14% by chance), the same failure LAB/006 found for search. Nothing in the output warns you; it just sorts wrongly.
Among the language models, Qwen3-14B was the most even across languages (86–88%). The smaller models lost the most on Lithuanian.
Entity extraction
| English | Lithuanian | Russian | Median time (s) | Replies cut off | |
|---|---|---|---|---|---|
| Qwen3-4B Instruct-2507 Q4_K_M | 76.4 / 54.9 | 69.8 / 49.5 | 71.2 / 28.0 | 0.24 | 0 |
| Qwen3-8B Q4_K_M | 79.4 / 62.2 | 73.9 / 63.6 | 73.7 / 37.6 | 0.51 | 0 |
| Qwen3-14B Q4_K_M | 79.3 / 62.9 | 72.1 / 64.9 | 71.2 / 36.4 | 0.50 | 0 |
| Phi-4-mini Q4_K_M | 57.8 / 43.2 | 44.4 / 34.7 | 45.1 / 23.7 | 0.17 | 1 |
Qwen3-4B Instruct-2507 Q4_K_M
- English76.4%
- Lithuanian69.8%
- Russian71.2%
Qwen3-8B Q4_K_M
- English79.4%
- Lithuanian73.9%
- Russian73.7%
Qwen3-14B Q4_K_M
- English79.3%
- Lithuanian72.1%
- Russian71.2%
Phi-4-mini Q4_K_M
- English57.8%
- Lithuanian44.4%
- Russian45.1%
With relaxed matching (same type, and one name contains the other), the Qwen models reach about 70–80% in all three languages, so Lithuanian and Russian are 5–8 points behind English. Exact matching shows a bigger gap for Russian, mostly for reasons that are not errors: Russian names change their endings with grammatical case, and the reference labels write some names as “Surname , Name”. Any downstream system needs normalisation that understands this, for Lithuanian too.
Phi-4-mini was clearly weaker on entities, and once got stuck repeating itself until it hit the length limit: a failure the output format cannot prevent, but plain code can detect.
What this means in practice
- Offline triage in Lithuanian, English and Russian is practical on one consumer GPU. No data has to leave the building.
- Use the simplest tool that works. For sorting text into known categories, a small classifier on multilingual embeddings matched or beat language models, and needed examples in only one language.
- Never use an English-only model on Lithuanian or Russian text. It fails silently.
- Use language models where categories are open-ended, such as extracting names and places, and normalise their output for each language’s grammar before matching or counting.
Method and environment
- topic data
SIB-200 (Adelani et al., 2024; CC BY-SA 4.0), test split, eng_Latn / lit_Latn / rus_Cyrl, parallel- ner data
WikiANN (Pan et al., 2017), test split, 300 sentences per language (seed 3); silver labels- llm
llama.cpp b11461 (CUDA 13.4); output constrained to a JSON schema (topic: enum of the 7 labels; entities: list of text + PER/ORG/LOC); greedy; Qwen3-8B/14B thinking off- topic prompt
Classify the topic of this text. Choose exactly one of: science/technology, travel, politics, sports, health, entertainment, geography. Reply as JSON: {"topic": "..."}. Text: {text}- ner prompt
List the named entities in this text: people (PER), organisations (ORG) and locations (LOC). Copy each entity exactly as written in the text. Reply as JSON: {"entities": [{"text": "...", "type": "PER" | "ORG" | "LOC"}]}. Use an empty list if there are none. Text: {text}- embeddings
bge-m3, multilingual-e5-base, bge-base-en-v1.5 (English-only reference); logistic regression (C = 4) on the 701 English training sentences- scoring
topic accuracy; entity-level precision/recall/F1 after whitespace/punctuation/case normalisation; relaxed = same type and containment- environment
gpu: RTX 4070 12 GB · llama.cpp: b11461 (CUDA 13.4)- measured
- 2026-10-08
Limits
Short, clean sentences from encyclopedia, travel and news sources, not live news, social media or transcripts. WikiANN’s labels are automatic and imperfect, and its sentences differ between languages, so the entity results compare languages only approximately. One prompt per task, written before the runs. The code is in labs/multilingual-triage in this site’s repository.
Public data only; this Lab covers text analysis, not collection or surveillance.