IR

Section 24 of the project spec

Reproducibility

Exact instructions to regenerate every number on this site from scratch.

One-command reproduction

git clone https://github.com/Arungharami/biomedical-hybrid-ir
cd biomedical-hybrid-ir
python3.11 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt && pip install -e .

python scripts/reproduce.py --config configs/default.yaml

Runs, in order: dataset audit → TF-IDF → BM25 → BGE → MedCPT → hybrid RRF → cross-encoder reranking → statistical/efficiency analysis → error analysis → web export. Each step is skipped if its expected output artifact already exists — pass --force to recompute everything, or --from <step> to resume after a failure.

Environment

  • Python 3.11 (PyTorch/FAISS wheels lag behind newer CPython releases)
  • seed = 42 throughout
  • Library versions recorded per-run in every results/manifests/*.json — never hand-maintained, so they can't drift out of sync

Device handling

Prefers CUDA (Colab / cloud GPUs), then Apple MPS (used for local development on this project's M1 Pro), then CPU. Configurable via configs/default.yaml → device.preference.

Colab notebooks

13 notebooks total. 8 were executed end-to-end in a real Jupyter kernel during development — not just validated as well-formed JSON.

NotebookStatus
00_environment_setup.ipynbVerified (executed)
01_nfcorpus_dataset_audit.ipynbVerified (M1)
02_tfidf_baseline.ipynbVerified (executed)
03_bm25_baseline.ipynbVerified (executed)
04_bge_dense_retrieval.ipynbComplete (~2 min runtime)
05_medcpt_dense_retrieval.ipynbComplete (~3 min runtime)
06_hybrid_rrf.ipynbVerified (executed)
07_cross_encoder_reranking.ipynbComplete (~40 min for all pools)
08_evaluation.ipynbVerified (executed)
09_statistical_analysis.ipynbVerified (executed)
10_error_analysis.ipynbVerified (executed)
11_export_research_results.ipynbVerified (executed)
Biomedical_Hybrid_IR_Full_Pipeline.ipynbComplete (master notebook)

Testing

pytest -q — 209 tests, all passing (unit tests for TF-IDF/BM25/RRF/metrics/statistics/error-analysis, plus notebook and figure validity checks). CI (.github/workflows/python-ci.yml) runs install → lint → test on every push, using CPU-only PyTorch and no network calls — it intentionally does not download NFCorpus or any transformer model.