IR

CAP 6776 — Information Retrieval

Hybrid Biomedical Information Retrieval with Lexical, Dense, and Cross-Encoder Reranking: A Reproducible Study on NFCorpus

Abstract

Biomedical information retrieval sits between two demands that are often in tension: exact lexical precision (drug names, gene symbols, precise terminology) and semantic understanding of paraphrased, layperson queries against expert-authored literature. We present a reproducible study on NFCorpus (3,633 documents, 323 test queries, verified against the live Hugging Face Hub) comparing six retrieval systems along the progression classical lexical retrieval (TF-IDF, BM25) → general-purpose dense retrieval (BGE) → biomedical-domain dense retrieval (MedCPT) → hybrid lexical-semantic fusion (Reciprocal Rank Fusion) → biomedical cross-encoder reranking, under one evaluation protocol, one qrels set, and paired query-level significance testing.

Contrary to a common textbook expectation, TF-IDF significantly outperforms BM25 on this corpus. Both dense retrievers significantly and substantially outperform BM25 on every primary metric (p<0.005 each) — the study's most robust finding — but biomedical-domain training (MedCPT) does not significantly outperform a strong general-purpose embedding model (BGE). Cross-encoder reranking achieves the highest nDCG@10 of all six systems at a real, substantial latency cost, but that specific improvement is not statistically significant. Query-level error analysis surfaces concrete mechanisms behind these aggregate numbers.

Read the full abstract on GitHub ↗

Full paper sections

Every section is complete and written only from real, completed results — no placeholder language remains. Full text is in the repository (rendered here would duplicate content that's already maintained as the single source of truth):