IR

M7 · Section 19

Efficiency

Latency, index size, and the effectiveness/efficiency trade-off, measured on Apple M1 Pro (MPS) — not representative of GPU-optimized production latency, but real measurements on the hardware actually used.

ModelDeviceEmbedding dimIndex sizeLatencynDCG@10
TF-IDFcpu——3.63 ms0.3050
BM25cpu——2.16 ms0.2954
BGEmps76810899.0 KB2.68 ms0.3712
MedCPTmps76810899.0 KB1.87 ms0.3654
Hybrid RRFcpu——4.07 ms0.3620
Hybrid + Rerankermps——1.95 s0.3731

Source: results/tables/efficiency.json — src/biomedical_ir/efficiency.py

nDCG@10 vs. latency

Scatter plot of nDCG@10 versus query latency (log scale) for all six models

Cross-encoder reranking sits roughly three orders of magnitude further right than any other system for a nDCG@10 gain that is numerically real but not statistically significant (see Evaluation).

Reranker candidate-pool ablation (A6)

PoolP@10Recall@100MAPnDCG@10Latency
20 0.26220.21740.15860.3640765.41 ms
50 (default)0.27650.27820.17600.37311.94 s
100 0.26900.33890.18460.36643.92 s

Recall@k for k exceeding the pool size equals Recall@pool exactly — a reranker cannot recover documents outside its candidate pool. Effectiveness vs. pool size is non-monotonic for nDCG@10 (peaks at 50, not 100).