NFCorpus
Dataset
Biomedical/nutrition IR collection: layperson health queries paired with graded relevance judgments against PubMed-indexed scientific articles.
Corpus
3,633
documents
Queries (total)
3,237
Train qrels
2590
queries
Test qrels
323
queries
Splits and their role
- traindevelopment (not used for final evaluation)
- devparameter tuning / model selection only
- testfinal evaluation only - never used for tuning
Source: The qrels HF split is literally named "validation", mapped internally to "dev". See docs/dataset.md.
Validation (automatic)
- corpus size
- 3633
- query count
- 3237
- empty documents
- 0
- empty queries
- 0
- duplicate document ids
- 0
- duplicate query ids
- 0
- missing doc references in qrels
- 0
- missing query references in qrels
- 0
Source: data/processed/dataset_stats.json (scripts/audit_dataset.py)
Corpus statistics
- Documents
- 3,633
- Approx. vocabulary size
- 66,399
- Title word count (mean / median)
- 12.79 / 13
- Body word count (mean / median)
- 220.98 / 224
- Body word count (min–max)
- 13–1460
Source: data/processed/corpus_stats.json
Provenance
- hf dataset
- BeIR/nfcorpus
- hf qrels dataset
- BeIR/nfcorpus-qrels
- corpus config
- corpus
- queries config
- queries
- splits requested
- ["train","dev","test"]
- corpus size
- 3633
- query count
- 3237
- qrel counts
- {"train":2590,"dev":324,"test":323}
- raw corpus row count
- 3633
- raw query row count
- 3237
Source: Verified directly against the live Hugging Face Hub — BeIR/nfcorpus (corpus, queries) + BeIR/nfcorpus-qrels.
Document composition
When both fields exist, a document is composed as {title} [SEP] {text} for TF-IDF/BM25/BGE. MedCPT's article encoder instead consumes documents as a structured [title, text] pair — the format its tokenizer call expects, verified against the official model card.