When not to use an LLM
A practical test for deciding whether a step in your system should be a language model, ordinary code, a search index, or a person.
Language models are now the default answer to almost any software question. That is a problem, because a model is a component with unusual properties. It is slow, it costs money on every call, its behaviour changes when the vendor updates it, and it is occasionally wrong in fluent, convincing ways. Sometimes that trade is worth it. Often it isn’t.
This is the test I use, one step of a system at a time.
1. Is there a correct answer that code can compute?
If the output is determined by rules (tax calculations, date arithmetic, routing by a field value, validation against a schema), write the code. It will be faster, free to run, testable and identical every time. A model asked to do arithmetic or apply a rulebook will be right most of the time, which is a worse property than being right all of the time.
A common anti-pattern is asking a model to “check whether this invoice total matches the line items”. Parse the numbers and add them.
2. Does the user actually need to find something rather than generate something?
Many “AI assistant” requests are search problems in disguise. If the answer exists verbatim in a document, policy or database, the user is usually better served by a good search result that links to the source than by a paraphrase of it. Generation adds latency and a chance of error to an answer that already existed.
Retrieval often belongs in the system even when generation does not. A strong lexical or hybrid search with good filters solves more internal-knowledge problems than its reputation suggests.
3. Is the input genuinely unstructured and variable?
This is where models earn their place: free-text emails, documents with no stable layout, user questions phrased a hundred different ways, classification by meaning rather than keyword. If a regular expression or a template would cover 98% of inputs, use them and send only the remaining 2% to a model or a person.
4. Can you check the output cheaply?
The most useful question of all. A model’s output is safe to use automatically when something cheaper than the model can verify it:
- extracted fields that must satisfy a schema and add up;
- generated code that must pass tests;
- a classification that a downstream rule can sanity-check;
- a citation that must point to a real passage containing the claim.
If nothing can check the output and a mistake is expensive, the model should draft and a person should decide.
5. What does a wrong answer cost, and who notices?
A wrong product description that a marketer reviews is cheap. A wrong figure in a compliance report that nobody reviews is not. Set the level of automation by the cost of an unnoticed error, not by how good the demo looked.
6. Do the economics hold at your real volume?
Multiply the per-call cost and latency by the actual number of calls. A model in the hot path of a high-traffic feature can cost more than the feature earns, and it adds seconds that users feel. Caching, smaller models, batching or moving the step offline often fix this. So does not using a model at all.
A worked example
“Read incoming supplier invoices and enter them into the accounting system.”
- Detecting duplicates, matching supplier IDs, checking totals and VAT: code.
- Finding the invoice among email attachments: rules first (file type, sender), with a small classifier for the ambiguous rest.
- Extracting fields from layouts that vary by supplier: a model, constrained to a schema.
- Deciding whether to pay an invoice that failed validation: a person, with the source document and the failed check in front of them.
One model call in a pipeline of mostly ordinary software. That is what most good AI systems look like up close.
The short version
Use a language model for the steps that need to interpret unstructured input, and make sure something else checks its work. Use code for anything with a correct answer, search for anything that already exists, and people for decisions where an unnoticed error is expensive.