LAB/001 · Building · Self-directed, not client work

Retrieval Observatory

A measured look at the retrieval stages RAG systems are built from (lexical, dense, hybrid and reranked search) on public benchmark data, with every number reproducible.

Most advice about retrieval-augmented generation is folklore: always use hybrid search, always add a reranker. Those are reasonable defaults, but they are claims about data, and they can be measured. This Lab project measures them.

What is measured

Each query goes through the stages a typical RAG retriever is built from, and quality is measured at the output of every stage:

SYSTEM LAB/001DATASET SCIFACTQUERIES 300RECORDED 2026-10-06

  1. LEXBM250.1 msnDCG@10 0.686

    query terms → top 100 of 5,183 docsExact term matching with stemming. Strong on identifiers, names and rare terms.

  2. VECDense (bge-base)3.6 msnDCG@10 0.744

    query embedding → top 100 by cosineEmbeds query and documents in one vector space. Strong on paraphrase and synonyms.

  3. FUSEReciprocal rank fusion< 0.1 msnDCG@10 0.735

    2 × top 100 → fused top 100Combines ranks, not scores, so the two systems need no calibration (k = 60).

  4. RERANKCross-encoder (bge-reranker-v2-m3)661.4 msnDCG@10 0.736

    100 (query, document) pairs → reorderedReads query and document together. Most accurate in principle, and the slowest stage.

Latency: p50 per query, one query at a time, NVIDIA GeForce RTX 4070. nDCG@10 is measured at the output of each stage.
  • BM25: classic lexical search with stemming.
  • Dense: two bi-encoder embedding models (bge-small and bge-base), exact cosine search.
  • Hybrid: reciprocal rank fusion of BM25 and the bge-base dense results.
  • Reranked: cross-encoders of three sizes reordering the hybrid top 100. The largest also reranks the dense top 100, to separate the effect of fusion from the effect of reranking.

The data is three public test sets from the BEIR benchmark, chosen because they differ: scientific claims (SciFact), medical questions in plain language (NFCorpus) and financial opinion questions (FiQA).

Results

nDCG@10 by retrieval systemScale 0 – 0.8 · higher is better · measured 2026-10-06

scifact · 5,183 docs

  1. BM250.686
  2. Dense (bge-small-en-v1.5)0.713
  3. Dense (bge-base-en-v1.5)0.744
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.735
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.686
  6. Hybrid → rerank top-100 (bge-reranker-base)0.716
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.736
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.733

nfcorpus · 3,633 docs

  1. BM250.322
  2. Dense (bge-small-en-v1.5)0.343
  3. Dense (bge-base-en-v1.5)0.374
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.364
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.356
  6. Hybrid → rerank top-100 (bge-reranker-base)0.318
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.344
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.341

fiqa · 57,638 docs

  1. BM250.251
  2. Dense (bge-small-en-v1.5)0.403
  3. Dense (bge-base-en-v1.5)0.406
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.366
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.371
  6. Hybrid → rerank top-100 (bge-reranker-base)0.332
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.431
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.436

What the numbers say

1. A good embedding model alone was the strongest single choice. Dense retrieval with bge-base beat BM25 on all three datasets, by a wide margin on FiQA (0.406 against 0.251 nDCG@10), and had the best nDCG@10 of any system on SciFact and NFCorpus.

2. Hybrid fusion did not beat dense retrieval on ranking quality anywhere. Per query it mostly reshuffled results: on SciFact it improved 53 queries and worsened 53; on FiQA it improved 156 and worsened 216. With equal-weight fusion, a much weaker retriever pulls good dense results down. Hybrid did raise Recall@100 on SciFact (0.970 against 0.963), which matters if a reranker follows, but not on the other two.

3. Rerankers depend on the domain. The two smaller cross-encoders lowered nDCG@10 in five of six cases. The largest, bge-reranker-v2-m3, lifted FiQA from 0.406 to 0.436 (about 7% relative) but lowered SciFact (0.744 → 0.733) and NFCorpus (0.375 → 0.341). It also costs roughly 540–670 ms per query to rerank 100 candidates on this GPU, against about 4 ms for the dense search itself.

The practical conclusion: the textbook pipeline of hybrid search plus a reranker is a hypothesis to test, not a default to ship. Start from a strong embedding model, measure on your own questions, and add each stage only where it pays for its latency.

scifact: 5,183 documents, 300 test queries. Higher is better; latency is per query.
SystemnDCG@10Recall@10Recall@100MRR@10Latency p50 / p95
BM250.6860.8190.9130.6490.1 / 0.18 ms
Dense (bge-small-en-v1.5)0.7130.8360.9420.6823.57 / 4.12 ms
Dense (bge-base-en-v1.5)0.7440.8780.9630.7083.64 / 3.93 ms
Hybrid RRF (BM25 + bge-base-en-v1.5)0.7350.8640.9700.701—
Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.6860.8060.9700.65757.99 / 63.66 ms
Hybrid → rerank top-100 (bge-reranker-base)0.7160.8440.9700.683217.42 / 237.5 ms
Hybrid → rerank top-100 (bge-reranker-v2-m3)0.7360.8460.9700.707661.43 / 718.9 ms
Dense → rerank top-100 (bge-reranker-v2-m3)0.7330.8470.9630.702—
LAB/001 · measured 2026-10-06 · NVIDIA GeForce RTX 4070
nfcorpus: 3,633 documents, 323 test queries. Higher is better; latency is per query.
SystemnDCG@10Recall@10Recall@100MRR@10Latency p50 / p95
BM250.3220.1470.2470.5280.1 / 0.14 ms
Dense (bge-small-en-v1.5)0.3430.1620.3110.5293.46 / 3.57 ms
Dense (bge-base-en-v1.5)0.3740.1780.3370.5663.42 / 3.55 ms
Hybrid RRF (BM25 + bge-base-en-v1.5)0.3640.1730.3290.568—
Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.3560.1670.3290.58658.73 / 62.48 ms
Hybrid → rerank top-100 (bge-reranker-base)0.3180.1480.3290.545218.53 / 230.5 ms
Hybrid → rerank top-100 (bge-reranker-v2-m3)0.3440.1710.3290.544667.61 / 698.7 ms
Dense → rerank top-100 (bge-reranker-v2-m3)0.3410.1680.3370.539—
LAB/001 · measured 2026-10-06 · NVIDIA GeForce RTX 4070
fiqa: 57,638 documents, 648 test queries. Higher is better; latency is per query.
SystemnDCG@10Recall@10Recall@100MRR@10Latency p50 / p95
BM250.2510.3180.5590.3070.21 / 1.21 ms
Dense (bge-small-en-v1.5)0.4030.4630.6950.4883.6 / 3.69 ms
Dense (bge-base-en-v1.5)0.4060.4810.7410.4863.7 / 3.83 ms
Hybrid RRF (BM25 + bge-base-en-v1.5)0.3660.4490.7190.437—
Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.3710.4520.7190.44150.68 / 54.71 ms
Hybrid → rerank top-100 (bge-reranker-base)0.3320.4120.7190.403178.5 / 188.49 ms
Hybrid → rerank top-100 (bge-reranker-v2-m3)0.4310.4960.7190.520541.77 / 571.26 ms
Dense → rerank top-100 (bge-reranker-v2-m3)0.4360.5070.7410.519—
LAB/001 · measured 2026-10-06 · NVIDIA GeForce RTX 4070

Averages hide a lot. The table below counts, query by query, how often a change actually helped:

Per-query nDCG@10: how many queries the second system improved / left unchanged / made worse.
Comparisonscifactnfcorpusfiqa
Dense (bge-base) vs BM2586 / 171 / 43131 / 110 / 82332 / 232 / 84
Hybrid vs BM2577 / 208 / 15130 / 143 / 50312 / 303 / 33
Hybrid vs dense (bge-base)53 / 194 / 53103 / 118 / 102156 / 276 / 216
Hybrid + bge-reranker-v2-m3 vs hybrid48 / 201 / 5187 / 123 / 113252 / 270 / 126
Dense + bge-reranker-v2-m3 vs dense46 / 192 / 6290 / 106 / 127212 / 257 / 179
Method and environment
splits
BEIR test split; queries with at least one positive judgement
metrics
nDCG@10 (graded), Recall@10, Recall@100, MRR@10; own implementation, trec_eval definitions
depth
100
bm25
{"k1":1.5,"b":0.75,"method":"lucene","stemmer":"english","stopwords":"en"}
dense
{"models":["BAAI/bge-small-en-v1.5","BAAI/bge-base-en-v1.5"],"search":"exact cosine (brute force) on GPU, fp16","query_prompt":"model-recommended instruction for bge"}
hybrid
{"method":"reciprocal rank fusion","k":60,"inputs":["bm25","BAAI/bge-base-en-v1.5"]}
rerankers
["cross-encoder/ms-marco-MiniLM-L6-v2","BAAI/bge-reranker-base","BAAI/bge-reranker-v2-m3"]
rerank_input
hybrid top-100 (and dense top-100 for BAAI/bge-reranker-v2-m3), max_length 512, fp16 (checked: identical nDCG@10 to fp32 for bge-reranker-base on SciFact)
latency
200 randomly sampled queries, one at a time, wall clock, after warm-up
seed
1234
environment
gpu: NVIDIA GeForce RTX 4070 · cpu: 13th Gen Intel(R) Core(TM) i5-13400F · python: 3.12.14 · torch: 2.14.1+cu130 · sentence_transformers: 6.1.0 · bm25s: 0.3.12
measured
2026-10-06

Reproducing it

The code is a standalone Python package (labs/retrieval-observatory in this site’s repository). After downloading the public datasets and models, one command runs every stage and writes the JSON file these figures are rendered from. The metric implementation was checked against published figures: BM25 and both embedding models reproduce reported SciFact and NFCorpus scores to two or three decimal places.

Limits

Three English datasets, one machine, exact search rather than an approximate index, and models from 22 to 568 million parameters, and the BEIR relevance judgements, which are known to be incomplete (an unjudged relevant document counts as a miss). Larger rerankers, domain-tuned models or weighted fusion could change the picture. On your own documents, the only way to know is to measure, which is exactly what this harness is for.