CAP 6776 — Information Retrieval
Hybrid Biomedical Information Retrieval with Lexical, Dense, and Cross-Encoder Reranking
A reproducible study on NFCorpus comparing six retrieval systems — TF-IDF, BM25, BGE, MedCPT, a BM25+MedCPT hybrid, and biomedical cross-encoder reranking — under one evaluation protocol and real paired statistical significance testing.
Corpus
3,633
documents
Test queries
323
with qrels
Models compared
6
lexical → dense → hybrid → reranked
Best nDCG@10
0.3731
Hybrid + Cross-Encoder
Pipeline
See Pipeline for the full architecture diagram and design rationale.
Headline finding
Both dense retrievers (BGE, MedCPT) significantly outperform BM25 on every primary metric (e.g. nDCG@10: p=0.0000) — the most robust finding across all statistical tests run. Biomedical-domain training (MedCPT) did not significantly outperform a general-purpose embedding model (BGE), and cross-encoder reranking's best raw nDCG@10 number in the study is not statistically significant relative to hybrid RRF.
See the full statistical analysis →Explore
Dataset
NFCorpus schema, stats, provenance
Pipeline
The full retrieval architecture
Models
TF-IDF, BM25, BGE, MedCPT, RRF, Cross-Encoder
Experiments
Real manifests for every run
Results
Main comparison table
Evaluation
MAP, MRR, nDCG, P, R, significance
Error Analysis
Query-level win/loss examples
Efficiency
Latency, index size trade-offs
Search Demo
Precomputed retrieval examples
Paper
Abstract, methodology, full results
Experiment status
| Milestone | Description | Status |
|---|---|---|
| M0 | Repository foundation, packaging, configs, CI skeleton | Complete |
| M1 | Dataset ingestion + validation | Complete |
| M2 | TF-IDF + BM25 baselines | Complete |
| M3 | BGE general dense retrieval | Complete |
| M4 | MedCPT biomedical dense retrieval | Complete |
| M5 | Hybrid RRF | Complete |
| M6 | MedCPT cross-encoder reranking | Complete |
| M7 | Full evaluation, statistical tests, efficiency analysis | Complete |
| M8 | Error analysis | Complete |
| M9 | Colab notebooks | Complete |
| M10 | Paper artifacts (figures, BibTeX) | Complete |
| M11 | Next.js research portal | Complete |
| M12 | Vercel deployment | In progress |