Service · retrieval

RAG development and retrieval systems

Search and question-answering over your own documents that returns the right sources and shows where every answer came from.

The problem

Most organisations have the answer somewhere in their documents, but nobody can find it. A quick retrieval-augmented generation (RAG) prototype often makes this worse: it answers confidently, cites the wrong passage, and nobody can tell how often it is wrong.

What gets built

  • Ingestion: parsing PDFs, office documents and web pages into clean text with structure and metadata preserved.
  • Retrieval: lexical, vector and hybrid search, tuned on your real queries, with reranking where it pays for its latency.
  • Answering: context selection, grounded answers with citations to the exact source passage, and an honest “not found”.
  • Evaluation: a test set built from real questions, with retrieval and answer quality measured before and after every change.
  • Deployment: in your cloud, on-premises, or with local models when data must not leave your infrastructure.

When RAG is the wrong tool

If the data is already structured, a database query or a well-designed search interface is often cheaper, faster and more reliable than an LLM. If the documents are few and stable, better information architecture may be enough. A feasibility assessment answers this before anything is built.

How an engagement starts

With a feasibility assessment on a sample of your documents and real questions: how well retrieval works today, what would improve it, and what a production system would cost to run.

Have a project like this?

  1. I read your message and reply personally, normally within two working days.
  2. Any questions are clarified by email. A short call only if required.
  3. A written proposal for a fixed-scope first step, with a fixed price.

Evidence and further reading

  • LAB/001

    Retrieval Observatory

    A measured look at the retrieval stages RAG systems are built from (lexical, dense, hybrid and reranked search) on public benchmark data, with every number reproducible.

  • How to evaluate a RAG system without fooling yourself

    A small, honest evaluation set beats a large, vague one. How to build one, which numbers to track, and the mistakes that make RAG systems look better than they are.

  • Hybrid search and reranking, measured

    The standard advice for RAG retrieval is BM25 plus embeddings plus a reranker. On three public benchmarks, a single good embedding model was hard to beat, and the reranker's value depended on the domain.

  • Why RAG retrieval fails

    Most bad RAG answers are retrieval failures in disguise. A field guide to the common failure modes, how to recognise each one, and what usually fixes it.

Updated