Writing · 3 min read

Hybrid search and reranking, measured

The standard advice for RAG retrieval is BM25 plus embeddings plus a reranker. On three public benchmarks, a single good embedding model was hard to beat, and the reranker's value depended on the domain.

Ask how to build retrieval for a RAG system and you will usually hear the same recipe: combine keyword search (BM25) with vector search, fuse the results, then rerank the top candidates with a cross-encoder. Every stage sounds like an improvement. Each one also adds latency, cost and moving parts.

I measured the recipe stage by stage on three public test sets from the BEIR benchmark: scientific claims (SciFact), medical questions in plain language (NFCorpus) and financial questions from forums (FiQA). Full method, code and limits are on the LAB/001 page. This article is about what the numbers mean for building real systems.

The short version

nDCG@10 by retrieval systemScale 0 – 0.8 · higher is better · measured 2026-10-06

scifact · 5,183 docs

  1. BM250.686
  2. Dense (bge-small-en-v1.5)0.713
  3. Dense (bge-base-en-v1.5)0.744
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.735
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.686
  6. Hybrid → rerank top-100 (bge-reranker-base)0.716
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.736
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.733

nfcorpus · 3,633 docs

  1. BM250.322
  2. Dense (bge-small-en-v1.5)0.343
  3. Dense (bge-base-en-v1.5)0.374
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.364
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.356
  6. Hybrid → rerank top-100 (bge-reranker-base)0.318
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.344
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.341

fiqa · 57,638 docs

  1. BM250.251
  2. Dense (bge-small-en-v1.5)0.403
  3. Dense (bge-base-en-v1.5)0.406
  4. Hybrid RRF (BM25 + bge-base-en-v1.5)0.366
  5. Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.371
  6. Hybrid → rerank top-100 (bge-reranker-base)0.332
  7. Hybrid → rerank top-100 (bge-reranker-v2-m3)0.431
  8. Dense → rerank top-100 (bge-reranker-v2-m3)0.436
  1. A strong embedding model on its own was the best system on two of the three datasets, and the best single-stage system on all three.
  2. Reciprocal rank fusion with BM25 never improved ranking quality over that model.
  3. Reranking helped on one dataset and hurt on two, including with the largest reranker tested, at more than 100 times the latency of the retrieval itself.

None of this means hybrid search or reranking is useless. It means they are hypotheses to test on your data, not defaults to ship.

Why hybrid fusion didn’t help here

Hybrid search exists because lexical and dense retrieval fail differently. BM25 finds exact identifiers, names and rare terms; embeddings find paraphrases. Fusing them should cover both blind spots.

That only works when both retrievers are reasonably good. Reciprocal rank fusion gives each list equal weight. When one retriever is much weaker (here BM25 trails the embedding model by 0.15 nDCG@10 on FiQA), fusion pulls weak results into the top ranks about as often as it rescues good ones. The per-query count shows it directly:

Per-query nDCG@10: how many queries the second system improved / left unchanged / made worse.
Comparisonscifactnfcorpusfiqa
Dense (bge-base) vs BM2586 / 171 / 43131 / 110 / 82332 / 232 / 84
Hybrid vs BM2577 / 208 / 15130 / 143 / 50312 / 303 / 33
Hybrid vs dense (bge-base)53 / 194 / 53103 / 118 / 102156 / 276 / 216
Hybrid + bge-reranker-v2-m3 vs hybrid48 / 201 / 5187 / 123 / 113252 / 270 / 126
Dense + bge-reranker-v2-m3 vs dense46 / 192 / 6290 / 106 / 127212 / 257 / 179

On SciFact, fusion improved 53 queries and worsened 53. On FiQA it worsened more queries (216) than it improved (156).

Where hybrid still makes sense:

  • Queries full of identifiers: part numbers, error codes, legal references, product SKUs. These are the cases embeddings handle worst, and they are underrepresented in these three datasets.
  • As a recall stage before a good reranker: hybrid raised Recall@100 on SciFact, so a reranker had more relevant candidates to work with.
  • With weighting: a weighted fusion, or a lexical stage that only contributes when the query contains rare tokens, avoids diluting a strong dense ranking. I did not test this here; it is the obvious next experiment.

Why reranking was a coin flip

Cross-encoders read the query and each candidate together, which in principle makes them more accurate than retrieval models that encode them separately. In practice they are trained on particular kinds of data, mostly web search questions and answers, and their advantage does not always transfer.

  • The two smaller rerankers (22 and 278 million parameters) lowered nDCG@10 in five of six cases.
  • The largest, bge-reranker-v2-m3 (568 million parameters), lifted FiQA from 0.406 to 0.436 nDCG@10, about 7% relative. Financial forum questions are close to the question-answer style rerankers are trained on.
  • The same model lowered SciFact from 0.744 to 0.733 and NFCorpus from 0.375 to 0.341. Scientific claims and medical topic queries are a different task from “find the passage that answers this question”.

And the cost: reranking 100 candidates took a median of 540–670 ms per query on an RTX 4070, against roughly 4 ms for the dense search. Fewer candidates or a smaller model cut that, at some cost in quality. Either way it is a real trade-off that belongs in the decision.

fiqa: 57,638 documents, 648 test queries. Higher is better; latency is per query.
SystemnDCG@10Recall@10Recall@100MRR@10Latency p50 / p95
BM250.2510.3180.5590.3070.21 / 1.21 ms
Dense (bge-small-en-v1.5)0.4030.4630.6950.4883.6 / 3.69 ms
Dense (bge-base-en-v1.5)0.4060.4810.7410.4863.7 / 3.83 ms
Hybrid RRF (BM25 + bge-base-en-v1.5)0.3660.4490.7190.437—
Hybrid → rerank top-100 (ms-marco-MiniLM-L6-v2)0.3710.4520.7190.44150.68 / 54.71 ms
Hybrid → rerank top-100 (bge-reranker-base)0.3320.4120.7190.403178.5 / 188.49 ms
Hybrid → rerank top-100 (bge-reranker-v2-m3)0.4310.4960.7190.520541.77 / 571.26 ms
Dense → rerank top-100 (bge-reranker-v2-m3)0.4360.5070.7410.519—
LAB/001 · measured 2026-10-06 · NVIDIA GeForce RTX 4070

What I would do on a real project

  1. Build an evaluation set first. Fifty to a hundred real questions with the passages that answer them (see how to evaluate a RAG system). Without it, none of these choices can be made on evidence.
  2. Start with a strong embedding model and measure it.
  3. Look at the failures. If they involve exact identifiers or rare terms, try adding lexical search, ideally weighted, and check the per-query counts, not just the average.
  4. Try a reranker only if ranking is the problem: high Recall@100 but low nDCG@10. Test it on your queries, and decide whether the gain is worth the latency.
  5. Re-measure when anything changes: documents, models or the kind of questions users ask.

The common advice is a sensible list of things to try. Measured on your own data, some of them will pay for themselves and some won’t, and you can only find out which by testing.

Method and environment
splits
BEIR test split; queries with at least one positive judgement
metrics
nDCG@10 (graded), Recall@10, Recall@100, MRR@10; own implementation, trec_eval definitions
depth
100
bm25
{"k1":1.5,"b":0.75,"method":"lucene","stemmer":"english","stopwords":"en"}
dense
{"models":["BAAI/bge-small-en-v1.5","BAAI/bge-base-en-v1.5"],"search":"exact cosine (brute force) on GPU, fp16","query_prompt":"model-recommended instruction for bge"}
hybrid
{"method":"reciprocal rank fusion","k":60,"inputs":["bm25","BAAI/bge-base-en-v1.5"]}
rerankers
["cross-encoder/ms-marco-MiniLM-L6-v2","BAAI/bge-reranker-base","BAAI/bge-reranker-v2-m3"]
rerank_input
hybrid top-100 (and dense top-100 for BAAI/bge-reranker-v2-m3), max_length 512, fp16 (checked: identical nDCG@10 to fp32 for bge-reranker-base on SciFact)
latency
200 randomly sampled queries, one at a time, wall clock, after warm-up
seed
1234
environment
gpu: NVIDIA GeForce RTX 4070 · cpu: 13th Gen Intel(R) Core(TM) i5-13400F · python: 3.12.14 · torch: 2.14.1+cu130 · sentence_transformers: 6.1.0 · bm25s: 0.3.12
measured
2026-10-06