What a 12 GB GPU can run for private document AI
Seven open-weight model setups on one consumer graphics card. All were fast enough. The real differences were memory, and how often each model answered a question it should have refused.
A common first question about private AI is “what hardware do we need?”. Another is “will a model small enough to run in-house be good enough?”. I measured both on one NVIDIA RTX 4070 (12 GB), a consumer graphics card, with seven setups of freely licensed models: Qwen3 at 4, 8 and 14 billion parameters and Microsoft’s Phi-4-mini, in 4-bit and 8-bit versions. The full method and every number are on the LAB/005 page. Here is what matters for a decision.
Speed is not the problem
- Qwen3-4B Instruct-2507 Q4_K_M151
- Qwen3-4B Instruct-2507 Q8_099
- Qwen3-8B Q4_K_M91
- Qwen3-8B Q8_056
- Qwen3-14B Q4_K_M51
- Phi-4-mini Q4_K_M162
- Phi-4-mini Q8_0106
Even the largest setup writes about 50 tokens per second, several times faster than anyone reads, and every setup reads a page of context in well under a second. For one user at a time, any of these models feels responsive on this card.
Memory is the problem
The 14B model uses 9.6 GB of the 12 GB at an 8,000-token context, and stops fitting at about 16,000 tokens. That sounds like headroom until the rest of the system needs the same card. In the end-to-end test, the 14B model and a small search model could not share the GPU at 8,000 tokens. Plan memory for the whole pipeline, not for the model alone.
4-bit quantisation helps. In this test, the 4-bit files were about 40% smaller, 1.5–1.6 times faster at writing, and no less accurate than the 8-bit ones (within the test’s margin of error).
The question that decides the model: does it admit “not found”?
Each model got 500 questions with five numbered paragraphs, and had to answer with a short phrase and a passage number, or say NOT FOUND. A third of the questions had no answer in the text, though they were written to look like they did.
Larger models gave better answers when there was one: answer F1 rose from 75% to 86% between the 4B and 14B models. But they were also more willing to answer when they should have refused:
- Qwen3-4B Instruct-2507 Q4_K_M70.1%
- Qwen3-4B Instruct-2507 Q8_069.5%
- Qwen3-8B Q4_K_M55.7%
- Qwen3-8B Q8_055.1%
- Qwen3-14B Q4_K_M50.3%
- Phi-4-mini Q4_K_M23.4%
- Phi-4-mini Q8_024.0%
The smaller Qwen3-4B said NOT FOUND on 70% of the unanswerable questions; the 14B on half of them; Phi-4-mini on under a quarter. Overall, the 4B and 14B ended within about a point of each other, and the 4B answers in a third of the time.
Here is what a confident wrong answer looks like:
- What rule didn't some native live under?
- In between the French and the British, large areas were dominated by native tribes. To the north, the Mi'kmaq and the Abenaki were engaged in Father Le Loutre's War and still held sway in parts of Nova Scotia, Acadia, and the eastern portions of the province of Canada, as well as much of present-day Maine. The Iroquois Confederation dominated much of present-day Upstate New York and the Ohio Country, although the latter also included Algonquian-speaking populations of Delaware and Shawnee, as well as Iroquoian-speaking Mingo. These tribes were formally under Iroquois rule, and were limited by them in authority to make agreements.
- Iroquois rule [1]
- 0 of 7 configurations replied NOT FOUND.
The answer is copied from the right paragraph, cites the right passage, and is wrong. Nothing about the format gives it away. A system that answers staff or customers from internal documents will produce answers like this, and a reviewer skimming them will not catch it.
What I’d do with this
- Start with a small model on one card. For grounded answers from your own documents, a 4B model on a 12 GB GPU is a serious option, not a toy.
- Choose the model for how it fails. Build a test set from your own documents that includes questions whose answers aren’t there, and measure how often each candidate answers anyway. That rate matters more than any leaderboard.
- Check the output format in code. Missing citations, for example, are easy to detect and retry; one model here omitted them in 18% of its answers.
- Leave memory for the pipeline: search, longer documents and later growth.
The limits: one card, English only, short factual answers, and one prompt that was not tuned. Lithuanian is measured separately, and on your own documents the numbers will differ. That is the point of measuring them.