Main comparison
Results
All six models, evaluated identically against the same real NFCorpus test split (n=323 queries) and the same qrels. No number below is estimated — every cell is generated from results/metrics/*.json.
| Model | P@10 | Recall@100 | MAP | MRR@10 | nDCG@10 | Latency (ms/query) |
|---|---|---|---|---|---|---|
| TF-IDF | 0.2167 | 0.2372 | 0.1372 | 0.5062 | 0.3050 | 3.631 |
| BM25 | 0.2071 | 0.2295 | 0.1333 | 0.4939 | 0.2954 | 2.155 |
| BGE | 0.2796 | 0.3368 | 0.1831 | 0.5556 | 0.3712 | 2.678 |
| MedCPT | 0.2697 | 0.3488 | 0.1824 | 0.5487 | 0.3654 | 1.871 |
| BM25+MedCPT (RRF) | 0.2598 | 0.3389 | 0.1809 | 0.5678 | 0.3620 | 4.069 |
| Hybrid+Reranker (pool=50) | 0.2765 | 0.2782 | 0.1760 | 0.5670 | 0.3731 | 1945.6 |
Bold/accent = best value in that column. Reranked row's Recall@100/MAP are capped by its 50-document candidate pool — see the note below and Evaluation for the full explanation and statistical significance tests.
Reading the reranker row correctly
A cross-encoder can only reorder the candidates it is given. With a candidate pool of 50, the reranked run never contains more than 50 documents per query, so Recall@100 == Recall@50 exactly by construction — confirmed mechanistically: at pool=100, Recall@100 exactly matches hybrid RRF's own Recall@100 (reordering an unchanged 100-document set cannot change how many relevant documents it contains). This is a genuine methodological property, not a bug — see the pool-size ablation on Efficiency.
Progression across the study
Ordered by real nDCG@10, lowest to highest.