Architecture
Retrieval Pipeline
NFCorpus → validation → lexical + dense retrieval → RRF fusion → cross-encoder reranking → evaluation → statistical analysis → research portal.
Full pipeline
Stage → module → script
| # | Stage | Module | Script |
|---|---|---|---|
| 1 | Dataset ingestion | biomedical_ir/data.py | scripts/download_data.py |
| 2 | Dataset validation | biomedical_ir/validation.py | scripts/audit_dataset.py |
| 3 | TF-IDF (M1) | biomedical_ir/tfidf.py | scripts/run_tfidf.py |
| 4 | BM25 (M2) | biomedical_ir/bm25.py | scripts/run_bm25.py |
| 5 | General dense retrieval (M3, BGE) | biomedical_ir/dense.py | scripts/run_bge.py |
| 6 | Biomedical dense retrieval (M4, MedCPT) | biomedical_ir/medcpt.py | scripts/run_medcpt.py |
| 7 | Hybrid RRF (M5) | biomedical_ir/fusion.py | scripts/run_hybrid.py |
| 8 | Cross-encoder reranking (M6) | biomedical_ir/reranker.py | scripts/run_reranker.py |
| 9 | Evaluation + statistics | biomedical_ir/{evaluation,statistics,efficiency}.py | scripts/evaluate_all.py |
| 10 | Error analysis | biomedical_ir/error_analysis.py | scripts/error_analysis.py |
| 11 | Figures / tables | — | scripts/generate_figures.py |
| 12 | Web export | — | scripts/export_web_results.py |
Why RRF operates on rank, not raw score
BM25 scores are unbounded and MedCPT similarities are raw (unnormalized) dot products — not on a comparable scale. Summing them directly would implicitly and arbitrarily weight whichever retriever happens to produce larger-magnitude scores, not whichever is more accurate. Reciprocal Rank Fusion instead combines rank positions, which are already on a common scale by construction.
Why Vercel never runs the models
BGE, MedCPT, and the MedCPT cross-encoder are never invoked inside a Vercel serverless function. All embeddings/rankings are computed offline (locally or in Colab), evaluated, and exported as compact JSON artifacts. This portal only reads those artifacts — the /search page is explicitly labeled as showing precomputed experiment output, never a live model call.