LLM and AI integration
Language-model features inside your existing product or internal tools, connected to your data with permissions, tests and cost control.
The problem
Adding a chat box that calls a model API takes an afternoon. Making that feature reliable is the actual work: answering from your data rather than the model’s memory, respecting who is allowed to see what, behaving predictably when the model is slow or wrong, and costing a known amount per request.
What gets built
- Data access done properly: the model reaches your database, documents or internal APIs through narrow, typed tools with the same permission checks as the rest of your software. It never gets raw credentials or open-ended query access.
- Structured outputs: responses validated against a schema before your code acts on them, with retries and a defined failure path.
- Evaluation: a test set of real inputs with expected behaviour, run on every prompt or model change, so “it seems better” becomes a number.
- Operational limits: timeouts, rate limits, caching, per-user and per-month cost ceilings, and logging that lets you see what the model was asked and what it answered.
- Model choice: hosted models where that is acceptable, open-weight models on your own infrastructure where it is not (see private AI), behind one interface so the choice can change later.
- MCP servers where an assistant needs to reach your systems through the Model Context Protocol, built with the same constraints as any other API.
When this is the wrong tool
If the task has a correct answer that code can compute, write the code. If users mostly need to find existing information, better search usually beats generated text (see retrieval systems). If a mistake would be expensive and you cannot check the output, a language model should not be making that decision on its own.
How an engagement starts
Usually with a short assessment of one concrete feature: the data it needs, the failure modes, an evaluation plan and an estimate of running cost. Then a prototype inside your actual software, not a separate demo.
Have a project like this?
- I read your message and reply personally, normally within two working days.
- Any questions are clarified by email. A short call only if required.
- A written proposal for a fixed-scope first step, with a fixed price.
Evidence and further reading
When not to use an LLM
A practical test for deciding whether a step in your system should be a language model, ordinary code, a search index, or a person.
How an engagement works
Every project starts small and fixed in scope, so you see evidence before you commit to more.
Feasibility assessment
A short, fixed-scope look at your problem, your data and the options.
A written assessment and a clear recommendation, including "don't build this" when that is the honest answer.
Prototype
A working system on your real data, not a demo on someone else's.
The prototype plus an evaluation report with measured results.
Build
The production system, integrated with what you already run.
Tested, documented code that you own, with a proper handover.
Support
Monitoring, evaluation reruns and improvements as your data changes.
Agreed scope and response times, no lock-in.