IR

M7

Evaluation & Statistical Analysis

Every model evaluated with the same qrels via pytrec_eval, cross-checked against from-scratch metric implementations. Significance: paired bootstrap at the query level.

Metrics computed

  • Precision@{1,5,10}
  • Recall@{10,20,50,100}
  • MRR, MRR@10
  • MAP, MAP@100
  • nDCG@{5,10,20}

Source: src/biomedical_ir/evaluation.py, cross-checked in tests/test_metrics.py

Graded relevance handling

NFCorpus qrels are graded (0/1/2). Empirically confirmed (not assumed): binary-style measures treat any grade >0 as relevant; nDCG uses linear gain (gain(rel)=rel), not exponential gain; a query with zero relevant documents scores 0.0 for every measure.

Statistical significance (paired_bootstrap, n_resamples=10000, seed=42)

ComparisonMetricDiff (A−B)95% CIpSig.
TF-IDF vs BM25P@10+0.0096[0.0019, 0.0183]0.0170✓
TF-IDF vs BM25Recall@100+0.0077[0.0014, 0.0160]0.0100✓
TF-IDF vs BM25MAP+0.0038[-0.0007, 0.0085]0.0954–
TF-IDF vs BM25MRR@10+0.0123[-0.0107, 0.0354]0.2946–
TF-IDF vs BM25nDCG@10+0.0096[0.0010, 0.0183]0.0296✓
BM25 vs BGEP@10-0.0724[-0.0920, -0.0536]0.0000✓
BM25 vs BGERecall@100-0.1073[-0.1285, -0.0861]0.0000✓
BM25 vs BGEMAP-0.0498[-0.0645, -0.0364]0.0000✓
BM25 vs BGEMRR@10-0.0618[-0.1015, -0.0236]0.0030✓
BM25 vs BGEnDCG@10-0.0757[-0.0988, -0.0536]0.0000✓
BM25 vs MedCPTP@10-0.0625[-0.0817, -0.0443]0.0000✓
BM25 vs MedCPTRecall@100-0.1193[-0.1421, -0.0969]0.0000✓
BM25 vs MedCPTMAP-0.0491[-0.0641, -0.0354]0.0000✓
BM25 vs MedCPTMRR@10-0.0549[-0.0940, -0.0162]0.0046✓
BM25 vs MedCPTnDCG@10-0.0700[-0.0931, -0.0480]0.0000✓
BGE vs MedCPTP@10+0.0099[-0.0034, 0.0235]0.1512–
BGE vs MedCPTRecall@100-0.0120[-0.0279, 0.0028]0.1238–
BGE vs MedCPTMAP+0.0007[-0.0134, 0.0139]0.9130–
BGE vs MedCPTMRR@10+0.0069[-0.0260, 0.0387]0.6824–
BGE vs MedCPTnDCG@10+0.0057[-0.0133, 0.0241]0.5408–
MedCPT vs Hybrid RRFP@10+0.0099[-0.0028, 0.0229]0.1262–
MedCPT vs Hybrid RRFRecall@100+0.0099[0.0004, 0.0199]0.0396✓
MedCPT vs Hybrid RRFMAP+0.0015[-0.0054, 0.0085]0.6622–
MedCPT vs Hybrid RRFMRR@10-0.0191[-0.0498, 0.0116]0.2288–
MedCPT vs Hybrid RRFnDCG@10+0.0034[-0.0106, 0.0178]0.6270–
Hybrid RRF vs Hybrid+RerankerP@10-0.0167[-0.0300, -0.0034]0.0150✓
Hybrid RRF vs Hybrid+RerankerRecall@100+0.0607[0.0496, 0.0735]0.0000✓
Hybrid RRF vs Hybrid+RerankerMAP+0.0049[-0.0036, 0.0148]0.2912–
Hybrid RRF vs Hybrid+RerankerMRR@10+0.0009[-0.0298, 0.0325]0.9624–
Hybrid RRF vs Hybrid+RerankernDCG@10-0.0111[-0.0275, 0.0061]0.1944–

Source: results/tables/statistical_tests.json — src/biomedical_ir/statistics.py (paired_bootstrap_test)

Per-hypothesis verdicts

H1

BM25 will outperform TF-IDF

REJECTED

TF-IDF significantly beats BM25 on P@10 (p=0.017), Recall@100 (p=0.010), nDCG@10 (p=0.030).

H2

MedCPT will outperform BGE (biomedical vs. general dense)

NOT SUPPORTED

No significant difference on any of 5 metrics (all p≥0.12) — statistically indistinguishable.

H3

Hybrid RRF will outperform BM25 and MedCPT individually

NOT SUPPORTED

MedCPT significantly beats Hybrid RRF on Recall@100 (p=0.040) — the opposite direction.

H4

Cross-encoder reranking improves nDCG@10, increases latency

PARTIALLY SUPPORTED

Latency increase unambiguous. nDCG@10 gain is the best point estimate in the study but not significant (p=0.194). P@10 improves significantly (p=0.015).

Note: "not supported" means no significant difference was found — this is not the same claim as "proven equal" (absence of evidence is not evidence of absence).