Writing
Technical notes, experiments and postmortems. Measurements state their method, dataset and date.
Writing
How to evaluate a RAG system without fooling yourself
A small, honest evaluation set beats a large, vague one. How to build one, which numbers to track, and the mistakes that make RAG systems look better than they are.
Hybrid search and reranking, measured
The standard advice for RAG retrieval is BM25 plus embeddings plus a reranker. On three public benchmarks, a single good embedding model was hard to beat, and the reranker's value depended on the domain.
The cheapest way to trust AI document extraction
A vision-language model read 200 real receipts. Simple arithmetic checks decided which results could skip human review, and every receipt they let through had the right total.
When not to use an LLM
A practical test for deciding whether a step in your system should be a language model, ordinary code, a search index, or a person.
Why RAG retrieval fails
Most bad RAG answers are retrieval failures in disguise. A field guide to the common failure modes, how to recognise each one, and what usually fixes it.