M7 · Section 19
Efficiency
Latency, index size, and the effectiveness/efficiency trade-off, measured on Apple M1 Pro (MPS) — not representative of GPU-optimized production latency, but real measurements on the hardware actually used.
| Model | Device | Embedding dim | Index size | Latency | nDCG@10 |
|---|---|---|---|---|---|
| TF-IDF | cpu | — | — | 3.63 ms | 0.3050 |
| BM25 | cpu | — | — | 2.16 ms | 0.2954 |
| BGE | mps | 768 | 10899.0 KB | 2.68 ms | 0.3712 |
| MedCPT | mps | 768 | 10899.0 KB | 1.87 ms | 0.3654 |
| Hybrid RRF | cpu | — | — | 4.07 ms | 0.3620 |
| Hybrid + Reranker | mps | — | — | 1.95 s | 0.3731 |
Source: results/tables/efficiency.json — src/biomedical_ir/efficiency.py
nDCG@10 vs. latency

Cross-encoder reranking sits roughly three orders of magnitude further right than any other system for a nDCG@10 gain that is numerically real but not statistically significant (see Evaluation).
Reranker candidate-pool ablation (A6)
| Pool | P@10 | Recall@100 | MAP | nDCG@10 | Latency |
|---|---|---|---|---|---|
| 20 | 0.2622 | 0.2174 | 0.1586 | 0.3640 | 765.41 ms |
| 50 (default) | 0.2765 | 0.2782 | 0.1760 | 0.3731 | 1.94 s |
| 100 | 0.2690 | 0.3389 | 0.1846 | 0.3664 | 3.92 s |
Recall@k for k exceeding the pool size equals Recall@pool exactly — a reranker cannot recover documents outside its candidate pool. Effectiveness vs. pool size is non-monotonic for nDCG@10 (peaks at 50, not 100).