IR

CAP 6776 — Information Retrieval

Hybrid Biomedical Information Retrieval with Lexical, Dense, and Cross-Encoder Reranking

A reproducible study on NFCorpus comparing six retrieval systems — TF-IDF, BM25, BGE, MedCPT, a BM25+MedCPT hybrid, and biomedical cross-encoder reranking — under one evaluation protocol and real paired statistical significance testing.

Corpus

3,633

documents

Test queries

323

with qrels

Models compared

6

lexical → dense → hybrid → reranked

Best nDCG@10

0.3731

Hybrid + Cross-Encoder

Pipeline

NFCorpus→Validation→TF-IDF / BM25→BGE / MedCPT→RRF Fusion→Cross-Encoder→Evaluation

See Pipeline for the full architecture diagram and design rationale.

Headline finding

Both dense retrievers (BGE, MedCPT) significantly outperform BM25 on every primary metric (e.g. nDCG@10: p=0.0000) — the most robust finding across all statistical tests run. Biomedical-domain training (MedCPT) did not significantly outperform a general-purpose embedding model (BGE), and cross-encoder reranking's best raw nDCG@10 number in the study is not statistically significant relative to hybrid RRF.

See the full statistical analysis →

Explore

Experiment status

MilestoneDescriptionStatus
M0Repository foundation, packaging, configs, CI skeletonComplete
M1Dataset ingestion + validationComplete
M2TF-IDF + BM25 baselinesComplete
M3BGE general dense retrievalComplete
M4MedCPT biomedical dense retrievalComplete
M5Hybrid RRFComplete
M6MedCPT cross-encoder rerankingComplete
M7Full evaluation, statistical tests, efficiency analysisComplete
M8Error analysisComplete
M9Colab notebooksComplete
M10Paper artifacts (figures, BibTeX)Complete
M11Next.js research portalComplete
M12Vercel deploymentIn progress