Writing

Technical notes, experiments and postmortems. Measurements state their method, dataset and date.

Writing

  • How to evaluate a RAG system without fooling yourself

    A small, honest evaluation set beats a large, vague one. How to build one, which numbers to track, and the mistakes that make RAG systems look better than they are.

  • Hybrid search and reranking, measured

    The standard advice for RAG retrieval is BM25 plus embeddings plus a reranker. On three public benchmarks, a single good embedding model was hard to beat, and the reranker's value depended on the domain.

  • The cheapest way to trust AI document extraction

    A vision-language model read 200 real receipts. Simple arithmetic checks decided which results could skip human review, and every receipt they let through had the right total.

  • When not to use an LLM

    A practical test for deciding whether a step in your system should be a language model, ordinary code, a search index, or a person.

  • Why RAG retrieval fails

    Most bad RAG answers are retrieval failures in disguise. A field guide to the common failure modes, how to recognise each one, and what usually fixes it.