Private LLM on One 12 GB GPU
Seven open-weight model setups on one consumer graphics card, measured for speed, memory and answering from documents, including whether they admit when the answer isn't there.
Can a company answer questions from its own documents with a language model that runs on one ordinary graphics card, with nothing sent to an outside service? This Lab measures what that hardware actually delivers: how fast, how much memory, how accurate, and, most importantly for business use, how often the model admits that the answer isn’t in the documents.
Hardware: one NVIDIA RTX 4070 with 12 GB of memory, the kind of card found in a gaming PC or a small workstation. Runtime: llama.cpp, with every model fully on the GPU.
The models
Three sizes of Qwen3 (4, 8 and 14 billion parameters, Apache-2.0) and Microsoft’s Phi-4-mini (3.8 billion, MIT), each compressed (“quantised”) to 4 bits (Q4_K_M) and, where it fits, 8 bits (Q8_0). All are free for commercial use. Qwen3-8B and 14B can “think” before answering; thinking was switched off so speeds are comparable and answers stay short.
Speed and memory
| File (GB) | GPU memory at 8k context (GB) | Reading (prompt tok/s, 4k) | Writing (tok/s) | Longest context that fits | Load (s) | |
|---|---|---|---|---|---|---|
| Qwen3-4B Instruct-2507 Q4_K_M | 2.5 | 3.7 | 7,077 | 151 | 49,152 | 0.7 |
| Qwen3-4B Instruct-2507 Q8_0 | 4.3 | 5.4 | 7,371 | 99 | 49,152 | 0.8 |
| Qwen3-8B Q4_K_M | 5.0 | 5.8 | 4,362 | 91 | 40,960 | 0.9 |
| Qwen3-8B Q8_0 | 8.7 | 8.9 | 4,581 | 56 | 24,576 | 1.2 |
| Qwen3-14B Q4_K_M | 9.0 | 9.6 | 2,375 | 51 | 16,384 | 1.3 |
| Phi-4-mini Q4_K_M | 2.5 | 3.6 | 8,475 | 162 | 65,536 | 0.8 |
| Phi-4-mini Q8_0 | 4.1 | 5.1 | 8,899 | 106 | 49,152 | 0.9 |
- Qwen3-4B Instruct-2507 Q4_K_M151
- Qwen3-4B Instruct-2507 Q8_099
- Qwen3-8B Q4_K_M91
- Qwen3-8B Q8_056
- Qwen3-14B Q4_K_M51
- Phi-4-mini Q4_K_M162
- Phi-4-mini Q8_0106
Every setup writes faster than a person reads, and reads a page of context in a fraction of a second. Speed is not the constraint. Memory is. The 14B model takes 9.6 GB of the 12 GB at an 8,000-token context and stops fitting at about 16,000 tokens. Smaller models leave room for long documents and for the rest of a pipeline. In the end-to-end test below, Qwen3-14B and the search model did not fit together at 8,000 tokens; the context had to come down to 4,000.
The writing speeds are 75–96% of the limit set by the card’s memory bandwidth (504 GB/s divided by the model’s size), which is the expected ceiling for this kind of workload. That check is how a speed benchmark gets caught running on the CPU by mistake.
Answering from documents
Each of 500 questions from SQuAD 2.0 (a public reading-comprehension dataset built on Wikipedia) came with five numbered paragraphs: the one the question was written against and four others from the same article. The model had to answer with a short phrase and the passage number, or reply NOT FOUND. In 167 of the questions the answer is not in the text: these questions were written to look answerable, which is exactly how real “not in our documents” questions tend to look.
| Overall correct | Answer F1 (answerable) | Cites the right passage | Says NOT FOUND when it should | NOT FOUND was right | Median answer time (s) | |
|---|---|---|---|---|---|---|
| Qwen3-4B Instruct-2507 Q4_K_M | 74.6% | 75.0% | 95.8% | 70.1% | 84.2% | 0.22 |
| Qwen3-4B Instruct-2507 Q8_0 | 73.2% | 74.0% | 95.3% | 69.5% | 87.2% | 0.23 |
| Qwen3-8B Q4_K_M | 72.0% | 78.1% | 96.5% | 55.7% | 84.5% | 0.33 |
| Qwen3-8B Q8_0 | 74.8% | 81.5% | 97.5% | 55.1% | 88.5% | 0.37 |
| Qwen3-14B Q4_K_M | 75.6% | 85.9% | 97.3% | 50.3% | 98.8% | 0.59 |
| Phi-4-mini Q4_K_M | 63.6% | 80.5% | 74.5% | 23.4% | 100.0% | 0.19 |
| Phi-4-mini Q8_0 | 63.4% | 80.9% | 88.6% | 24.0% | 100.0% | 0.21 |
- Qwen3-4B Instruct-2507 Q4_K_M74.6%
- Qwen3-4B Instruct-2507 Q8_073.2%
- Qwen3-8B Q4_K_M72.0%
- Qwen3-8B Q8_074.8%
- Qwen3-14B Q4_K_M75.6%
- Phi-4-mini Q4_K_M63.6%
- Phi-4-mini Q8_063.4%
Bigger models give better answers: answer F1 rises from 75% (Qwen3-4B) to 86% (Qwen3-14B). But overall, the best setups are within about a point of each other, because bigger models are also more willing to answer when they shouldn’t:
- Qwen3-4B Instruct-2507 Q4_K_M70.1%
- Qwen3-4B Instruct-2507 Q8_069.5%
- Qwen3-8B Q4_K_M55.7%
- Qwen3-8B Q8_055.1%
- Qwen3-14B Q4_K_M50.3%
- Phi-4-mini Q4_K_M23.4%
- Phi-4-mini Q8_024.0%
Qwen3-4B Instruct says NOT FOUND on 70% of the unanswerable questions; Qwen3-14B on 50%; Phi-4-mini on under a quarter. Qwen3-14B’s NOT FOUND is almost always right (99%), so when it refuses it can be trusted. It simply refuses too rarely. For a system that answers customers or staff from internal documents, a confident wrong answer usually costs more than “I couldn’t find that”. That trade-off, not the benchmark rank, should decide the model.
Phi-4-mini’s weak citation score is mostly a format problem: in 4-bit form it left out the passage number in 60 of 333 answers. Plain code can catch that and ask again. The Qwen models always included it, and pointed to the right passage 95–98% of the time.
With 500 questions, differences of less than about 4 percentage points are within sampling noise. On that basis, 4-bit and 8-bit versions of the same model performed the same: the 4-bit file is about 40% smaller and 1.5–1.6 times faster at writing.
Two examples
Chosen by a fixed rule: the first unanswerable question in the sample that the best setup answered, and the first one it correctly declined.
- What rule didn't some native live under?
- In between the French and the British, large areas were dominated by native tribes. To the north, the Mi'kmaq and the Abenaki were engaged in Father Le Loutre's War and still held sway in parts of Nova Scotia, Acadia, and the eastern portions of the province of Canada, as well as much of present-day Maine. The Iroquois Confederation dominated much of present-day Upstate New York and the Ohio Country, although the latter also included Algonquian-speaking populations of Delaware and Shawnee, as well as Iroquoian-speaking Mingo. These tribes were formally under Iroquois rule, and were limited by them in authority to make agreements.
- Iroquois rule [1]
- 0 of 7 configurations replied NOT FOUND.
The paragraph says the tribes were under Iroquois rule; the question asks what rule they didn’t live under. The answer looks perfectly grounded, with a correct-looking citation, and is wrong.
- During what years was William Ellery Channing President of Harvard?
- Charles W. Eliot, president 1869–1909, eliminated the favored position of Christianity from the curriculum while opening it to student self-direction. While Eliot was the most crucial figure in the secularization of American higher education, he was motivated not by a desire to secularize education, but by Transcendentalist Unitarian convictions. Derived from William Ellery Channing and Ralph Waldo Emerson, these convictions were focused on the dignity and worth of human nature, the right and ability of each person to perceive truth, and the indwelling God in each person.
- NOT FOUND.
- 7 of 7 configurations replied NOT FOUND.
End to end: search, then answer
The more realistic setup: no one hands the model the right paragraph. The question searches all 1,204 paragraphs with an open embedding model (bge-base, on the same GPU), the top five go to the language model, and the same prompt applies.
| Right paragraph retrieved | Overall correct | Answer F1 | NOT FOUND when it should | Search (ms) | Answer (ms) | Total p95 (s) | |
|---|---|---|---|---|---|---|---|
| Qwen3-14B Q4_K_M | 90.0% | 69.0% | 77.4% | 50.9% | 4 | 597 | 1.08 |
| Qwen3-4B Instruct-2507 Q4_K_M | 90.0% | 71.4% | 67.8% | 76.0% | 4 | 216 | 0.42 |
Search found the right paragraph 90% of the time and took about 4 milliseconds. The language model is 99% of the waiting time. With imperfect retrieval in the loop, the 4B model’s caution paid off: it ended slightly ahead of the 14B model overall, at a third of the response time, while the 14B gave better answers when it did answer. Either one answers in about a second or less on this card.
What this means in practice
- A single 12 GB card is enough for private question answering over company documents at interactive speed. (This Lab measured one request at a time; concurrent users are a separate test.)
- Choose the model for its failure mode, not its benchmark rank. Measure how often it answers when it shouldn’t, on your own documents.
- Use 4-bit quantisation unless your own tests show a loss; here there was none.
- Leave memory headroom for long documents and for the search model; a model that just fits will not fit in the pipeline.
Method and environment
- runtime
llama.cpp b11461 (CUDA 13.4), llama-server, all layers on the GPU, flash attention, f16 KV cache- speed
llama-bench: prompt 512 and 4096 tokens, generation 128 tokens, 3 repetitions; GPU memory read from nvidia-smi minus the idle desktop- dataset
SQuAD 2.0 dev (CC BY-SA 4.0): 500 questions, seed 7, one third unanswerable- context
gold paragraph + 4 random paragraphs from the same article, shuffled and numbered- prompt
You answer questions using only the numbered passages provided. If the passages contain the answer, reply with the shortest exact phrase from the passage that answers it, followed by the passage number in square brackets, for example: Paris [2]. If the passages do not contain the answer, reply exactly: NOT FOUND. Do not use outside knowledge. Reply with one line and nothing else.- decoding
greedy (temperature 0), max 64 tokens, Qwen3-8B/14B with thinking disabled- scoring
official SQuAD normalisation; EM/F1 against all gold answers; NOT FOUND scores 0 on answerable questions- end_to_end
BAAI/bge-base-en-v1.5 cosine search over 1204 paragraphs, top 5, same prompt; 4096-token context and fp16 embedder so both fit on the GPU with Qwen3-14B- environment
gpu: RTX 4070 12 GB · cpu: i5-13400F · ram: 32 GB · os: Arch Linux · llama.cpp: b11461 (CUDA 13.4)- measured
- 2026-10-07
Limits
One GPU, one runtime and one prompt; English only (Lithuanian is the subject of LAB/006). Short factual answers only: summaries, long answers and multi-step reasoning are not measured. SQuAD’s unanswerable questions are deliberately tricky, so abstention rates on real documents may be higher. A random extra paragraph can occasionally contain an answer to an “unanswerable” question; this was not corrected for. The prompt was written before the runs and not tuned on these questions. The code is a standalone package (labs/private-llm in this site’s repository) and reruns with four commands.