Service · cross-cutting

Private and local AI

AI systems that keep your data on infrastructure you control, using open-weight models sized realistically for the job and the hardware.

Why private

Some data cannot be sent to a third-party model API: client files, health records, unreleased financials, or anything a contract or regulation keeps in-house. For many European organisations, GDPR and confidentiality obligations make “where does the data go?” the first question, before “which model is best?”.

What gets built

The same systems as elsewhere on this site, including document search and question answering, document processing and AI features in your software, deployed so that documents, prompts and outputs stay on infrastructure you control:

  • On-premises servers or workstations with GPUs, or a private cloud tenancy in a region you choose.
  • Open-weight models for generation, embeddings and reranking, chosen for the task, the languages involved and the hardware available.
  • Access control and audit logs integrated with your existing identity system.
  • Measured quality: the same evaluation set run against private and hosted models, so you know the price of privacy in accuracy, if there is one.

Realistic tradeoffs

Smaller open-weight models running on one or two GPUs handle retrieval, extraction and classification well, and often match hosted models on narrow tasks. They are usually weaker at long, open-ended reasoning. Hardware cost is front-loaded but predictable. The honest answer for your case comes from measuring on your data, not from benchmarks.

When this is the wrong tool

If the data is not sensitive, a hosted API is cheaper to start and simpler to operate. A hybrid is common too: private models for sensitive steps, hosted models for the rest.

How an engagement starts

With a feasibility assessment that measures candidate models on a sample of your actual task and recommends hardware, or a hosted alternative if privacy does not require one.

Have a project like this?

  1. I read your message and reply personally, normally within two working days.
  2. Any questions are clarified by email. A short call only if required.
  3. A written proposal for a fixed-scope first step, with a fixed price.

Evidence and further reading

  • LAB/002

    Document Intelligence Engine

    Receipt extraction with an open-weight vision-language model, checked by plain-code arithmetic and routed to a person when the checks fail. Measured on 200 public receipts.

Updated