IR

Main comparison

Results

All six models, evaluated identically against the same real NFCorpus test split (n=323 queries) and the same qrels. No number below is estimated — every cell is generated from results/metrics/*.json.

ModelP@10Recall@100MAPMRR@10nDCG@10Latency (ms/query)
TF-IDF0.21670.23720.13720.50620.30503.631
BM250.20710.22950.13330.49390.29542.155
BGE0.27960.33680.18310.55560.37122.678
MedCPT0.26970.34880.18240.54870.36541.871
BM25+MedCPT (RRF)0.25980.33890.18090.56780.36204.069
Hybrid+Reranker (pool=50)0.27650.27820.17600.56700.37311945.6

Bold/accent = best value in that column. Reranked row's Recall@100/MAP are capped by its 50-document candidate pool — see the note below and Evaluation for the full explanation and statistical significance tests.

Reading the reranker row correctly

A cross-encoder can only reorder the candidates it is given. With a candidate pool of 50, the reranked run never contains more than 50 documents per query, so Recall@100 == Recall@50 exactly by construction — confirmed mechanistically: at pool=100, Recall@100 exactly matches hybrid RRF's own Recall@100 (reordering an unchanged 100-document set cannot change how many relevant documents it contains). This is a genuine methodological property, not a bug — see the pool-size ablation on Efficiency.

Progression across the study

BM25<TF-IDF<Hybrid RRF<MedCPT<BGE<Hybrid+Reranker

Ordered by real nDCG@10, lowest to highest.