LAB/005 · Running · Self-directed, not client work

Private LLM on One 12 GB GPU

Seven open-weight model setups on one consumer graphics card, measured for speed, memory and answering from documents, including whether they admit when the answer isn't there.

Can a company answer questions from its own documents with a language model that runs on one ordinary graphics card, with nothing sent to an outside service? This Lab measures what that hardware actually delivers: how fast, how much memory, how accurate, and, most importantly for business use, how often the model admits that the answer isn’t in the documents.

Hardware: one NVIDIA RTX 4070 with 12 GB of memory, the kind of card found in a gaming PC or a small workstation. Runtime: llama.cpp, with every model fully on the GPU.

The models

Three sizes of Qwen3 (4, 8 and 14 billion parameters, Apache-2.0) and Microsoft’s Phi-4-mini (3.8 billion, MIT), each compressed (“quantised”) to 4 bits (Q4_K_M) and, where it fits, 8 bits (Q8_0). All are free for commercial use. Qwen3-8B and 14B can “think” before answering; thinking was switched off so speeds are comparable and answers stay short.

Speed and memory

Speed and memory on one RTX 4070 (12 GB), llama.cpp, all layers on the GPU, flash attention. Generation speed is what a person waiting for an answer feels.
File (GB)GPU memory at 8k context (GB)Reading (prompt tok/s, 4k)Writing (tok/s)Longest context that fitsLoad (s)
Qwen3-4B Instruct-2507 Q4_K_M2.53.77,07715149,1520.7
Qwen3-4B Instruct-2507 Q8_04.35.47,3719949,1520.8
Qwen3-8B Q4_K_M5.05.84,3629140,9600.9
Qwen3-8B Q8_08.78.94,5815624,5761.2
Qwen3-14B Q4_K_M9.09.62,3755116,3841.3
Phi-4-mini Q4_K_M2.53.68,47516265,5360.8
Phi-4-mini Q8_04.15.18,89910649,1520.9
Context steps tested: 4k, 8k, 16k, 24k, 32k, 40k, 48k, 64k, 96k, 128k, 256k, capped at the model's native length. Load time with the file already in the OS cache · LAB/005 · measured 2026-10-07 · RTX 4070 12 GB
Writing speed, tokens per secondHigher is faster · reading speed is not shown: it is 2,000–9,000 tok/s for every model
  1. Qwen3-4B Instruct-2507 Q4_K_M151
  2. Qwen3-4B Instruct-2507 Q8_099
  3. Qwen3-8B Q4_K_M91
  4. Qwen3-8B Q8_056
  5. Qwen3-14B Q4_K_M51
  6. Phi-4-mini Q4_K_M162
  7. Phi-4-mini Q8_0106

Every setup writes faster than a person reads, and reads a page of context in a fraction of a second. Speed is not the constraint. Memory is. The 14B model takes 9.6 GB of the 12 GB at an 8,000-token context and stops fitting at about 16,000 tokens. Smaller models leave room for long documents and for the rest of a pipeline. In the end-to-end test below, Qwen3-14B and the search model did not fit together at 8,000 tokens; the context had to come down to 4,000.

The writing speeds are 75–96% of the limit set by the card’s memory bandwidth (504 GB/s divided by the model’s size), which is the expected ceiling for this kind of workload. That check is how a speed benchmark gets caught running on the CPU by mistake.

Answering from documents

Each of 500 questions from SQuAD 2.0 (a public reading-comprehension dataset built on Wikipedia) came with five numbered paragraphs: the one the question was written against and four others from the same article. The model had to answer with a short phrase and the passage number, or reply NOT FOUND. In 167 of the questions the answer is not in the text: these questions were written to look answerable, which is exactly how real “not in our documents” questions tend to look.

Grounded answering on 500 SQuAD 2.0 questions: the right paragraph plus 4 others from the same article. 167 questions have no answer in the text.
Overall correctAnswer F1 (answerable)Cites the right passageSays NOT FOUND when it shouldNOT FOUND was rightMedian answer time (s)
Qwen3-4B Instruct-2507 Q4_K_M74.6%75.0%95.8%70.1%84.2%0.22
Qwen3-4B Instruct-2507 Q8_073.2%74.0%95.3%69.5%87.2%0.23
Qwen3-8B Q4_K_M72.0%78.1%96.5%55.7%84.5%0.33
Qwen3-8B Q8_074.8%81.5%97.5%55.1%88.5%0.37
Qwen3-14B Q4_K_M75.6%85.9%97.3%50.3%98.8%0.59
Phi-4-mini Q4_K_M63.6%80.5%74.5%23.4%100.0%0.19
Phi-4-mini Q8_063.4%80.9%88.6%24.0%100.0%0.21
Overall correct: answerable questions need an answer with F1 ≥ 0.5, unanswerable ones need NOT FOUND · LAB/005 · measured 2026-10-07 · RTX 4070 12 GB
Overall correct on grounded questionsScale 0–100% · includes saying NOT FOUND when the answer is missing
  1. Qwen3-4B Instruct-2507 Q4_K_M74.6%
  2. Qwen3-4B Instruct-2507 Q8_073.2%
  3. Qwen3-8B Q4_K_M72.0%
  4. Qwen3-8B Q8_074.8%
  5. Qwen3-14B Q4_K_M75.6%
  6. Phi-4-mini Q4_K_M63.6%
  7. Phi-4-mini Q8_063.4%

Bigger models give better answers: answer F1 rises from 75% (Qwen3-4B) to 86% (Qwen3-14B). But overall, the best setups are within about a point of each other, because bigger models are also more willing to answer when they shouldn’t:

Unanswerable questions: how often the model says NOT FOUNDScale 0–100% · the rest got a confident, wrong answer
  1. Qwen3-4B Instruct-2507 Q4_K_M70.1%
  2. Qwen3-4B Instruct-2507 Q8_069.5%
  3. Qwen3-8B Q4_K_M55.7%
  4. Qwen3-8B Q8_055.1%
  5. Qwen3-14B Q4_K_M50.3%
  6. Phi-4-mini Q4_K_M23.4%
  7. Phi-4-mini Q8_024.0%

Qwen3-4B Instruct says NOT FOUND on 70% of the unanswerable questions; Qwen3-14B on 50%; Phi-4-mini on under a quarter. Qwen3-14B’s NOT FOUND is almost always right (99%), so when it refuses it can be trusted. It simply refuses too rarely. For a system that answers customers or staff from internal documents, a confident wrong answer usually costs more than “I couldn’t find that”. That trade-off, not the benchmark rank, should decide the model.

Phi-4-mini’s weak citation score is mostly a format problem: in 4-bit form it left out the passage number in 60 of 333 answers. Plain code can catch that and ask again. The Qwen models always included it, and pointed to the right passage 95–98% of the time.

With 500 questions, differences of less than about 4 percentage points are within sampling noise. On that basis, 4-bit and 8-bit versions of the same model performed the same: the 4-bit file is about 40% smaller and 1.5–1.6 times faster at writing.

Two examples

Chosen by a fixed rule: the first unanswerable question in the sample that the best setup answered, and the first one it correctly declined.

No answer in the text, but the model gave oneWrong
Question
What rule didn't some native live under?
Passage [1] (the one the question was written against; 4 other passages not shown)
In between the French and the British, large areas were dominated by native tribes. To the north, the Mi'kmaq and the Abenaki were engaged in Father Le Loutre's War and still held sway in parts of Nova Scotia, Acadia, and the eastern portions of the province of Canada, as well as much of present-day Maine. The Iroquois Confederation dominated much of present-day Upstate New York and the Ohio Country, although the latter also included Algonquian-speaking populations of Delaware and Shawnee, as well as Iroquoian-speaking Mingo. These tribes were formally under Iroquois rule, and were limited by them in authority to make agreements.
Qwen3-14B Q4_K_M replied
Iroquois rule [1]
Other models
0 of 7 configurations replied NOT FOUND.

SQuAD 2.0 dev, question 5ad3a3fc604f3c001a3fea7b, article “French and Indian War”. CC BY-SA 4.0.

The paragraph says the tribes were under Iroquois rule; the question asks what rule they didn’t live under. The answer looks perfectly grounded, with a correct-looking citation, and is wrong.

No answer in the text, and the model said soCorrect
Question
During what years was William Ellery Channing President of Harvard?
Passage [1] (the one the question was written against; 4 other passages not shown)
Charles W. Eliot, president 1869–1909, eliminated the favored position of Christianity from the curriculum while opening it to student self-direction. While Eliot was the most crucial figure in the secularization of American higher education, he was motivated not by a desire to secularize education, but by Transcendentalist Unitarian convictions. Derived from William Ellery Channing and Ralph Waldo Emerson, these convictions were focused on the dignity and worth of human nature, the right and ability of each person to perceive truth, and the indwelling God in each person.
Qwen3-14B Q4_K_M replied
NOT FOUND.
Other models
7 of 7 configurations replied NOT FOUND.

SQuAD 2.0 dev, question 5a8205d131013a001a3350f7, article “Harvard University”. CC BY-SA 4.0.

End to end: search, then answer

The more realistic setup: no one hands the model the right paragraph. The question searches all 1,204 paragraphs with an open embedding model (bge-base, on the same GPU), the top five go to the language model, and the same prompt applies.

End to end: the question searches all 1,204 paragraphs (BAAI/bge-base-en-v1.5, top 5), then the model answers from what was found.
Right paragraph retrievedOverall correctAnswer F1NOT FOUND when it shouldSearch (ms)Answer (ms)Total p95 (s)
Qwen3-14B Q4_K_M90.0%69.0%77.4%50.9%45971.08
Qwen3-4B Instruct-2507 Q4_K_M90.0%71.4%67.8%76.0%42160.42
Stage times are medians · LAB/005 · measured 2026-10-07 · RTX 4070 12 GB

Search found the right paragraph 90% of the time and took about 4 milliseconds. The language model is 99% of the waiting time. With imperfect retrieval in the loop, the 4B model’s caution paid off: it ended slightly ahead of the 14B model overall, at a third of the response time, while the 14B gave better answers when it did answer. Either one answers in about a second or less on this card.

What this means in practice

  • A single 12 GB card is enough for private question answering over company documents at interactive speed. (This Lab measured one request at a time; concurrent users are a separate test.)
  • Choose the model for its failure mode, not its benchmark rank. Measure how often it answers when it shouldn’t, on your own documents.
  • Use 4-bit quantisation unless your own tests show a loss; here there was none.
  • Leave memory headroom for long documents and for the search model; a model that just fits will not fit in the pipeline.
Method and environment
runtime
llama.cpp b11461 (CUDA 13.4), llama-server, all layers on the GPU, flash attention, f16 KV cache
speed
llama-bench: prompt 512 and 4096 tokens, generation 128 tokens, 3 repetitions; GPU memory read from nvidia-smi minus the idle desktop
dataset
SQuAD 2.0 dev (CC BY-SA 4.0): 500 questions, seed 7, one third unanswerable
context
gold paragraph + 4 random paragraphs from the same article, shuffled and numbered
prompt
You answer questions using only the numbered passages provided. If the passages contain the answer, reply with the shortest exact phrase from the passage that answers it, followed by the passage number in square brackets, for example: Paris [2]. If the passages do not contain the answer, reply exactly: NOT FOUND. Do not use outside knowledge. Reply with one line and nothing else.
decoding
greedy (temperature 0), max 64 tokens, Qwen3-8B/14B with thinking disabled
scoring
official SQuAD normalisation; EM/F1 against all gold answers; NOT FOUND scores 0 on answerable questions
end_to_end
BAAI/bge-base-en-v1.5 cosine search over 1204 paragraphs, top 5, same prompt; 4096-token context and fp16 embedder so both fit on the GPU with Qwen3-14B
environment
gpu: RTX 4070 12 GB · cpu: i5-13400F · ram: 32 GB · os: Arch Linux · llama.cpp: b11461 (CUDA 13.4)
measured
2026-10-07

Limits

One GPU, one runtime and one prompt; English only (Lithuanian is the subject of LAB/006). Short factual answers only: summaries, long answers and multi-step reasoning are not measured. SQuAD’s unanswerable questions are deliberately tricky, so abstention rates on real documents may be higher. A random extra paragraph can occasionally contain an answer to an “unanswerable” question; this was not corrected for. The prompt was written before the runs and not tuned on these questions. The code is a standalone package (labs/private-llm in this site’s repository) and reruns with four commands.