CAP 6776 — Information Retrieval
About this project
Course project
Built for CAP 6776 (Information Retrieval). The primary deliverable is a scientifically reproducible information retrieval experiment on NFCorpus — this website presents the experiment and its verified results; it is not the deliverable itself.
Scientific-integrity statement
This project reports only genuine experimental output. No metric, dataset count, citation, or statistical result is fabricated. Pending experiments are labeled "pending", never filled with plausible-looking numbers; failed runs are labeled "failed" with their error preserved. Every retrieval model choice (pooling, normalization, instruction prefixes, input formats) was verified against the official Hugging Face model card before use, not assumed.
Architecture
Hugging Face → Google Colab / local Python experiments → verified JSON/CSV artifacts → GitHub → Next.js → Vercel. Large models and embeddings stay in Colab/local environments — this website reads only compact exported results and never runs a transformer model itself.
Technology
- Research pipeline: Python 3.11, PyTorch, Transformers, sentence-transformers, FAISS, pytrec_eval
- Research portal: Next.js (App Router), TypeScript, Tailwind CSS
- Deployment: Vercel
Citation
@misc{gharami2026hybridir,
title = {Hybrid Biomedical Information Retrieval with Lexical, Dense,
and Cross-Encoder Reranking: A Reproducible Study on NFCorpus},
author = {Gharami, Arun},
year = {2026},
note = {CAP 6776 -- Information Retrieval},
url = {https://github.com/Arungharami/biomedical-hybrid-ir}
}