IR

M1 – M6

Models

Every model's pooling, normalization, and input-format choices are verified against the official Hugging Face model card before use — never assumed.

Classical lexical

TF-IDF

scikit-learn TfidfVectorizer (sublinear TF, L2 norm, smoothed IDF) + cosine similarity.

configs/tfidf.yaml

P@10

0.2167

Recall@100

0.2372

MAP

0.1372

MRR@10

0.5062

nDCG@10

0.3050

Classical lexical

BM25

From-scratch Okapi BM25 (k1=1.2, b=0.75, frozen literature defaults). Robertson/Sparck-Jones IDF.

configs/bm25.yaml

P@10

0.2071

Recall@100

0.2295

MAP

0.1333

MRR@10

0.4939

nDCG@10

0.2954

General dense retrieval

BGE

BAAI/bge-base-en-v1.5

CLS pooling, L2-normalized embeddings, query-only instruction prefix. FAISS IndexFlatIP.

configs/bge.yaml

P@10

0.2796

Recall@100

0.3368

MAP

0.1831

MRR@10

0.5556

nDCG@10

0.3712

Biomedical dense retrieval

MedCPT

ncbi/MedCPT-Query-Encoder + ncbi/MedCPT-Article-Encoder

Two-tower biomedical retriever, CLS pooling, no normalization (raw dot product). Articles as [title, text] pairs.

configs/medcpt.yaml

P@10

0.2697

Recall@100

0.3488

MAP

0.1824

MRR@10

0.5487

nDCG@10

0.3654

Hybrid retrieval

BM25 + MedCPT (RRF)

Reciprocal Rank Fusion of BM25 and MedCPT rankings, k=60. Operates on rank positions, not raw scores.

configs/hybrid.yaml

P@10

0.2598

Recall@100

0.3389

MAP

0.1809

MRR@10

0.5678

nDCG@10

0.3620

Biomedical reranking

Hybrid + MedCPT Cross-Encoder

ncbi/MedCPT-Cross-Encoder

Reranks the hybrid RRF run's top candidates (default pool=50) with a biomedical cross-encoder.

configs/reranker.yaml

P@10

0.2765

Recall@100

0.2782

MAP

0.1760

MRR@10

0.5670

nDCG@10

0.3731

RQ1–RQ6 and models

The primary progression this study tells its research story through: TF-IDF → BM25 → General Dense (BGE) → Biomedical Dense (MedCPT) → BM25 + Biomedical Dense (Hybrid RRF) → Cross-Encoder Reranking. See the paper for the full research questions, hypotheses, and per-hypothesis verdicts.