IR

Architecture

Retrieval Pipeline

NFCorpus → validation → lexical + dense retrieval → RRF fusion → cross-encoder reranking → evaluation → statistical analysis → research portal.

Full pipeline

NFCorpus (BeIR/nfcorpus + BeIR/nfcorpus-qrels)
Data Validation
Lexical IR — TF-IDF / BM25
+
Dense IR — BGE / MedCPT
RRF Fusion (BM25 + MedCPT, rank-based)
MedCPT Cross-Encoder Reranking
Ranked Documents
Qrels (test split)
→
Evaluation (pytrec_eval)
MAP / MRR / nDCG / Precision / Recall
Statistical Analysis (paired bootstrap)
Research Results (this portal)

Stage → module → script

#StageModuleScript
1Dataset ingestionbiomedical_ir/data.pyscripts/download_data.py
2Dataset validationbiomedical_ir/validation.pyscripts/audit_dataset.py
3TF-IDF (M1)biomedical_ir/tfidf.pyscripts/run_tfidf.py
4BM25 (M2)biomedical_ir/bm25.pyscripts/run_bm25.py
5General dense retrieval (M3, BGE)biomedical_ir/dense.pyscripts/run_bge.py
6Biomedical dense retrieval (M4, MedCPT)biomedical_ir/medcpt.pyscripts/run_medcpt.py
7Hybrid RRF (M5)biomedical_ir/fusion.pyscripts/run_hybrid.py
8Cross-encoder reranking (M6)biomedical_ir/reranker.pyscripts/run_reranker.py
9Evaluation + statisticsbiomedical_ir/{evaluation,statistics,efficiency}.pyscripts/evaluate_all.py
10Error analysisbiomedical_ir/error_analysis.pyscripts/error_analysis.py
11Figures / tables—scripts/generate_figures.py
12Web export—scripts/export_web_results.py

Why RRF operates on rank, not raw score

BM25 scores are unbounded and MedCPT similarities are raw (unnormalized) dot products — not on a comparable scale. Summing them directly would implicitly and arbitrarily weight whichever retriever happens to produce larger-magnitude scores, not whichever is more accurate. Reciprocal Rank Fusion instead combines rank positions, which are already on a common scale by construction.

Why Vercel never runs the models

BGE, MedCPT, and the MedCPT cross-encoder are never invoked inside a Vercel serverless function. All embeddings/rankings are computed offline (locally or in Colab), evaluated, and exported as compact JSON artifacts. This portal only reads those artifacts — the /search page is explicitly labeled as showing precomputed experiment output, never a live model call.